Open Access. Powered by Scholars. Published by Universities.®
- Institution
- Keyword
-
- Directivity (5)
- Automatic speech recognition (3)
- Radiation (3)
- Speech processing systems (3)
- Deaf people (2)
-
- Electro-larynx (2)
- Hearing impaired (2)
- Humans (2)
- Intelligibility (2)
- Laryngectomy (2)
- Speech (2)
- Speech synthesis (2)
- Acoustic surface wave devices (1)
- Articulation (1)
- Articulator motion (1)
- Articulatory synthesis (1)
- Auditory feedback (1)
- Binary control systems (1)
- Biomechanical Phenomena (1)
- Brain (1)
- Brain-computer interface (BCI) (1)
- Child (1)
- Child, Preschool (1)
- Children (1)
- Closing phase (1)
- Computational linguistics (1)
- Drug delivery (1)
- Dynamics (1)
- Dysarthria (1)
- Electrocorticographic gamma activity (1)
- Publication Year
- Publication
- Publication Type
Articles 1 - 24 of 24
Full-Text Articles in Speech and Hearing Science
Guitar Amplifier Directivity, Rachel C. Edelman, Brian E. Anderson, Samuel D. Bellows, Timothy W. Leishman
Guitar Amplifier Directivity, Rachel C. Edelman, Brian E. Anderson, Samuel D. Bellows, Timothy W. Leishman
Directivity
No abstract provided.
Revealing Spatiotemporal Neural Activation Patterns In Electrocorticography Recordings Of Human Speech Production By Mutual Information, Julio Kovacs, Dean Krusienski, Minu Maninder, Willy Wriggers
Revealing Spatiotemporal Neural Activation Patterns In Electrocorticography Recordings Of Human Speech Production By Mutual Information, Julio Kovacs, Dean Krusienski, Minu Maninder, Willy Wriggers
Mechanical & Aerospace Engineering Faculty Publications
Background
Spatiotemporal mapping of neural activity during continuous speech production has been traditionally approached using correlation coefficient (CC) analysis between cortical signals and speech recordings. A prior study employed this approach using electrocorticography (ECoG) data from participants who underwent invasive intracranial monitoring for epilepsy. However, CC cannot detect nonlinear relationships and is dominated by the correspondence between periods of silence and of non-silence.
New Method
We introduce the mutual information (MI) measure, which can capture both linear and nonlinear dependencies. We validated CC and MI on the sub-second spatiotemporal brain activity recorded during continuous speech tasks. To refine the results, …
Trumpet Directivity From A Rotating Semicircular Array, Samuel D. Bellows, Joseph E. Avila, Timothy W. Leishman
Trumpet Directivity From A Rotating Semicircular Array, Samuel D. Bellows, Joseph E. Avila, Timothy W. Leishman
Directivity
The directivity function of a played musical instrument describes the angular dependence of its acoustic radiation and diffraction about the instrument, musician, and musician’s chair. Directivity influences sound in rehearsal, performance, and recording environments and signals in audio systems. Because high-resolution, spherically comprehensive measurements of played musical instruments have been unavailable in the past, the authors have undertaken research to produce and share such data for studies of musical instruments, simulations of acoustical environments, optimizations of microphone placements, and other applications. The authors acquired the data from repeated chromatic scales produced by a trumpet played at mezzo-forte in an anechoic …
Gamelan Gong Directivity Dataset, Samuel D. Bellows, Dallin T. Harwood, Kent L. Gee, Micah R. Shepherd
Gamelan Gong Directivity Dataset, Samuel D. Bellows, Dallin T. Harwood, Kent L. Gee, Micah R. Shepherd
Directivity
No abstract provided.
Kemar Hats Head Orientation Directivity, Samuel D. Bellows, Timothy W. Leishman
Kemar Hats Head Orientation Directivity, Samuel D. Bellows, Timothy W. Leishman
Directivity
This directivity data set for a KEMAR head head-and-torso simulator (HATS) includes head orientations in 14 directions in 5° steps starting from 0° to 40° and then in 10° steps from 40° to 90°. The full spherical measurements followed at an a = 0.97 m radius with the mouth aperture at the spherical center. The sampling density and distribution followed the AES 5° dual-equiangular sampling standard, omitting the south pole (θ = 180°). Thus, each spherical directivity assessment included 36 polar-angle θ samples and 72 azimuthal-angle ϕ samples. The presented data include 22 1/3-octave bands, ranging from 80 Hz …
Speaker Encoding For Zero-Shot Speech Synthesis, Tristin W. Cory
Speaker Encoding For Zero-Shot Speech Synthesis, Tristin W. Cory
Graduate Theses/Dissertations
Spoken communication, for many, is an essential part of everyday life. Some individuals can lose or not be born with the ability to speak. To function on a day-to-day basis, these individuals find other ways of communication. Adaptive speech synthesis is one of those ways. It recreates a user’s previous voice or creates a voice that blends with their regional dialect. Current adaptive speech synthesis techniques that achieve human-like speech require thirty minutes, to a few hours of high-quality audio recordings of a target speaker. This amount of recorded audio is not commonly possessed by people in need of a …
Average Speech Directivity, Samuel D. Bellows, Claire M. Pincock, Jennifer K. Whiting, Timothy W. Leishman
Average Speech Directivity, Samuel D. Bellows, Claire M. Pincock, Jennifer K. Whiting, Timothy W. Leishman
Directivity
Speech directivity describes the angular dependence of acoustic radiation from a talker’s mouth and nostrils and diffraction about his or her body and chair (if seated). It is an essential physical aspect of communication affecting sounds and signals in acoustical environments, audio, and telecommunication systems. Because high-resolution, spherically comprehensive measurements of live, phonetically balanced speech have been unavailable in the past, the authors have undertaken research to produce and share such data for simulations of acoustical environments, optimizations of microphone placements, speech studies, and other applications. The measurements included three male and three female talkers who repeated phonetically balanced passages …
Effects Of Vocal Fold Nodules On Glottal Cycle Measurements Derived From High-Speed Videoendoscopy In Children, Rita R. Patel, Harikrishnan Unnikrishnan, Kevin D. Donohue
Effects Of Vocal Fold Nodules On Glottal Cycle Measurements Derived From High-Speed Videoendoscopy In Children, Rita R. Patel, Harikrishnan Unnikrishnan, Kevin D. Donohue
Electrical and Computer Engineering Faculty Publications
The goal of this study is to quantify the effects of vocal fold nodules on vibratory motion in children using high-speed videoendoscopy. Differences in vibratory motion were evaluated in 20 children with vocal fold nodules (5–11 years) and 20 age and gender matched typically developing children (5–11 years) during sustained phonation at typical pitch and loudness. Normalized kinematic features of vocal fold displacements from the mid-membranous vocal fold point were extracted from the steady-state high-speed video. A total of 12 kinematic features representing spatial and temporal characteristics of vibratory motion were calculated. Average values and standard deviations (cycle-to-cycle variability) of …
Spatio-Temporal Progression Of Cortical Activity Related To Continuous Overt And Covert Speech Production In A Reading Task, Jonathan S. Brumberg, Dean J. Krusienski, Shreya Chakrabarti, Aysegul Gunduz, Peter Brunner, Anthony L. Ritaccio, Gerwin Schalk
Spatio-Temporal Progression Of Cortical Activity Related To Continuous Overt And Covert Speech Production In A Reading Task, Jonathan S. Brumberg, Dean J. Krusienski, Shreya Chakrabarti, Aysegul Gunduz, Peter Brunner, Anthony L. Ritaccio, Gerwin Schalk
Electrical & Computer Engineering Faculty Publications
How the human brain plans, executes, and monitors continuous and fluent speech has remained largely elusive. For example, previous research has defined the cortical locations most important for different aspects of speech function, but has not yet yielded a definition of the temporal progression of involvement of those locations as speech progresses either overtly or covertly. In this paper, we uncovered the spatio-temporal evolution of neuronal population-level activity related to continuous overt speech, and identified those locations that shared activity characteristics across overt and covert speech. Specifically, we asked subjects to repeat continuous sentences aloud or silently while we recorded …
The Electromagnetic Articulography Mandarin Accented English (Ema-Mae) Corpus Of Acoustic And 3d Articulatory Kinematic Data, Jeffrey J. Berry, An Ji, Michael T. Johnson
The Electromagnetic Articulography Mandarin Accented English (Ema-Mae) Corpus Of Acoustic And 3d Articulatory Kinematic Data, Jeffrey J. Berry, An Ji, Michael T. Johnson
Speech Pathology and Audiology Faculty Research and Publications
There is a significant need for more comprehensive electromagnetic articulography (EMA) datasets that can provide matched acoustics and articulatory kinematic data with good spatial and temporal resolution. The Marquette University Electromagnetic Articulography Mandarin Accented English (EMA-MAE) corpus provides kinematic and acoustic data from 40 gender and dialect balanced speakers representing 20 Midwestern standard American English L1 speakers and 20 Mandarin Accented English (MAE) L2 speakers, half Beijing region dialect and half are Shanghai region dialect. Three dimensional EMA data were collected at a 400 Hz sampling rate using the NDI Wave system, with articulatory sensors on the midsagittal lips, lower …
Sensorimotor Adaptation Of Speech Using Real-Time Articulatory Resynthesis, Jeffrey J. Berry, Cassandra North, Michael T. Johnson
Sensorimotor Adaptation Of Speech Using Real-Time Articulatory Resynthesis, Jeffrey J. Berry, Cassandra North, Michael T. Johnson
Speech Pathology and Audiology Faculty Research and Publications
Sensorimotor adaptation is an important focus in the study of motor learning for non-disordered speech, but has yet to be studied substantially for speech rehabilitation. Speech adaptation is typically elicited experimentally using LPC resynthesis to modify the sounds that a speaker hears himself producing. This method requires that the participant be able to produce a robust speech-acoustic signal and is therefore not well-suited for talkers with dysarthria. We have developed a novel technique using electromagnetic articulography (EMA) to drive an articulatory synthesizer. The acoustic output of the articulatory synthesizer can be perturbed experimentally to study auditory feedback effects on sensorimotor …
Developing A Drug Delivery System For Treatment Of Vocal Fold Scarring, Aaron Michael Kosinski
Developing A Drug Delivery System For Treatment Of Vocal Fold Scarring, Aaron Michael Kosinski
Open Access Dissertations
Vocal fold scarring is an affliction that results in the formation of a disorganized and stiff extracellular matrix (ECM) with abnormal ECM component densities & structures including a significant increase in collagen deposition. It is caused by improper healing post injury and results in profound changes in the biomechanical properties of the vocal folds impairing their ability to generate a normal mucosal wave during phonation.
Finding an effective treatment for vocal fold scarring has been elusive. Currently, treatments seek temporary solutions that correct glottal incompetence and reduce stiffness caused by the scar through the augmentation of the vocal folds using …
Individual Articulator's Contribution To Phoneme Production, Jun Wang, Jordan R. Green, Ashok Samal
Individual Articulator's Contribution To Phoneme Production, Jun Wang, Jordan R. Green, Ashok Samal
School of Computing: Conference and Workshop Papers
Speech sounds are the result of coordinated movements of individual articulators. Understanding each articulator’s role in speech is fundamental not only for understanding how speech is produced, but also for optimizing speech assessments and treatments. In this paper, we studied the individual contributions of six articulators, tongue tip, tongue blade, tongue body front, tongue body back, upper lip, and lower lip to phoneme classification. A total of 3,838 vowel and consonant production samples were collected from eleven native English speakers. The results of speech movement classification using a support vector machine indicated that the tongue encoded significantly more information than …
Augmented Control Of A Hands-Free Electrolarynx, Brian Madden, James Condron, Eugene Coyle
Augmented Control Of A Hands-Free Electrolarynx, Brian Madden, James Condron, Eugene Coyle
Conference Papers
During voiced speech, the larynx acts as the sound source, providing a quasi-periodic excitation of the vocal tract. Following a total laryngectomy, some people speak using an electrolarynx which employs an electromechanical actuator to perform the excitatory function of the absent larynx. Drawbacks of conventional electrolarynx designs include the monotonic sound emitted, the need for a free-hand to operate the device, and the difficulty experienced by many laryngectomees in adapting to its use. One improvement to the electrolarynx, which clinicians and users frequently suggest, is the provision of a convenient hands-free control facility. This would allow more natural use of …
Intelligibility Of Electrolarynx Speech Using A Novel Actuator, Brian Madden, Mark Nolan, Ted Burke, James Condron, Eugene Coyle
Intelligibility Of Electrolarynx Speech Using A Novel Actuator, Brian Madden, Mark Nolan, Ted Burke, James Condron, Eugene Coyle
Conference Papers
During voiced speech, the larynx provides quasi-periodic acoustic excitation of the vocal tract. Following a laryngectomy, some people speak using an electrolarynx which replaces the excitatory function of the absent larynx. Drawbacks of conventional electrolarynx designs include the buzzing monotonic sound emitted, the need for a free hand to operate the device, and difficulty experienced by many laryngectomees in adapting to its use. Despite these shortcomings, it remains the preferred method of speech rehabilitation for a substantial minority of laryngectomees. In most electrolarynxes, mechanical vibrations are produced by a linear electromechanical actuator, the armature of which percusses against a metal …
Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally
Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally
Electrical & Computer Engineering Theses & Dissertations
Automatic Speaker Recognition is the process of automatically recognizing who is speaking on the basis of individual information contained in speech signals. This technique of Automatic Speaker Recognition makes it possible to use the speaker's voice to verify their identity and control access to services such as voice dialing, banking by telephone, telephone shopping, database access services, information services, voice mail, security control for confidential information areas, and remote access to computers.
In this thesis, the techniques of Gaussian Mixture Models and Neural Networks for Automatic Speaker Identification are presented. Algorithms for Speaker Identification using Gaussian Mixture Models were developed, …
A Computer-Based Articulation Training Aid For Short Words (Cata), Mukund Devarajan
A Computer-Based Articulation Training Aid For Short Words (Cata), Mukund Devarajan
Electrical & Computer Engineering Theses & Dissertations
Several improvements in the vowel articulation training aid (VATA) are described, as well as the efforts to extend the visual feedback system to operate with short words in the form of consonant, vowel and consonant (CVC). The extended version of the visual feedback system is referred to as CATA (Computer-based Articulation Training Aid); the vowel version of the aid (VATA) only operates with ten American English monopthong vowels. Improvements in VATA include the use of a neural network (NN) recognizer method to prune a large database of vowel recordings to eliminate noisy and/or mispronounced tokens. The spectral jitter problem, previously …
Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi
Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi
Electrical & Computer Engineering Theses & Dissertations
This thesis presents a pitch detection algorithm that is extremely robust for both high quality and telephone speech. The kernel method for this algorithm is the Normalized Cross Correlation (NCCF) reported by David Talkin [16]. Major innovations include: processing of the original acoustic signal and a nonlinearly processed version of the signal to partially restore very weak F0 components; intelligent peak picking to select multiple F0 candidates and assign merit factors; and, incorporation of highly robust pitch contours obtained from smoothed versions of low frequency portions of spectrograms. Dynamic programming is used to find the ''best" pitch track among all …
Real-Time Visual Speech Articulation Training Aid, Neiyer S. Correal
Real-Time Visual Speech Articulation Training Aid, Neiyer S. Correal
Electrical & Computer Engineering Theses & Dissertations
A real-time visual articulation training aid has been implemented. It provides instantaneous visual feedback of vowel and stop-consonant production on a computer screen. The vowel training system corresponds to an improved floating-point implementation of a previous fixed-point system developed by Beck (1992). The new implementation provides better accuracy and an approximate five-fold increase in speed. Acoustic features computed from global short-time spectral shape are used for classification of vowels. Temporal spectral trajectories timed to begin with burst onset are used for stop consonants. A neural network is used to transform measurements of auditory stimuli from the feature space to a …
Text Independent Speaker Verification Using Binary-Pair Partitioned Neural Networks, Claude A. Norton Iii
Text Independent Speaker Verification Using Binary-Pair Partitioned Neural Networks, Claude A. Norton Iii
Electrical & Computer Engineering Theses & Dissertations
A method is presented for the application of binary-pair partitioned neural networks to the task of speaker verification. This technique is based on a previously developed neural network classifier for speaker identification.
The main focus of this research was the development and testing of the algorithms necessary to extend the binary-pair partitioning approach from speaker identification to speaker verification. The method is based on the development of a user profile which is obtained from discriminative data provided by the binary-pair partitioned neural networks.
Experimental results are provided which demonstrate the viability of this approach, using the TIMIT speech corpus for …
Formant Estimation From Dctc's Using A Feedforward Neural Network, Shubhangi U. Kelkar
Formant Estimation From Dctc's Using A Feedforward Neural Network, Shubhangi U. Kelkar
Electrical & Computer Engineering Theses & Dissertations
Formants are the natural frequencies of the human vocal tract. Existing methods for estimating formants from speech signals are computationally complex and subject to errors for certain type of speech sounds. This thesis describes a method for estimating vowel formant frequencies from Discrete Cosine Transform Coefficients (DCTC's), a form of cepstral coefficients, using a feedforward neural network with back-propagation training. Experimental results are based on a large multispeaker data base. The results are obtained for both a linear transformation and a feedforward neural network with a nonlinear hidden layer. In general, the neural network transformation is superior to the linear …
Visual Speech Training Aid For The Deaf, Subhashri Venkat
Visual Speech Training Aid For The Deaf, Subhashri Venkat
Electrical & Computer Engineering Theses & Dissertations
A computer-based vowel articulation training aid has been developed. A "continuous" acoustic-phonetic transformation is performed to map speech parameters to a lower dimensionality display space. There are two possible approaches to this transformation problem. The transformation could be either linear or a combination nonlinear/linear. The nonlinear transformation is performed using a multi-layered feedforward neural network with linear output layers. Speech parameters are extracted either from an analog filter bank arrangement (band energies) or by a digital signal processing procedure (Discrete Cosine Transform Coefficients). The speech parameters obtained from both methods correspond to the spectral envelope of the speech signals. The …
An Investigation To Improve Linear Predictive Vocoder Pulse/Noise Excitation Models, Elizabeth Annella Martina Effer
An Investigation To Improve Linear Predictive Vocoder Pulse/Noise Excitation Models, Elizabeth Annella Martina Effer
Electrical & Computer Engineering Theses & Dissertations
The quality of synthetic speech from Linear Predictive (LP) vocoders is known to be degraded due to the lack of detail in the commonly used pulse/noise excitation model. In this investigation, it was hypothesized that this degradation is due to the lack of precise timing information in the pulses and to the constraint that each short-time segment of excitation be either an impulse train or white noise. Accordingly, more complex excitation models were implemented using precise timing from peaks in the residual and a mixture of pulses and noise. Since the LP residual is known to be the perfect excitation …
Color Display Of Vowel Spectra As A Training Aid For The Deaf, Amir Jalali Jagharghi
Color Display Of Vowel Spectra As A Training Aid For The Deaf, Amir Jalali Jagharghi
Electrical & Computer Engineering Theses & Dissertations
The objective of this research was to develop a transformation for mapping speech parameters to color parameter. This transformation is done in real-time, and the resulting color parameter are continuously displayed on a color monitor. This visual speech display is to be used as a speech articulation training aid for the deaf. The conversion of speech acoustic signals into speech parameter was accomplished using special -purpose electronics. The real-time conversion of speech parameter to display parameter was controlled by an 8086/8088 microprocessor operating in an S-100 bus structure. The coefficients of the Karhunen-Loeve series expansion of speech power spectra were …