Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences

Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 23 of 23

Full-Text Articles in Speech and Hearing Science

Revealing Spatiotemporal Neural Activation Patterns In Electrocorticography Recordings Of Human Speech Production By Mutual Information, Julio Kovacs, Dean Krusienski, Minu Maninder, Willy Wriggers Jan 2025

Revealing Spatiotemporal Neural Activation Patterns In Electrocorticography Recordings Of Human Speech Production By Mutual Information, Julio Kovacs, Dean Krusienski, Minu Maninder, Willy Wriggers

Mechanical & Aerospace Engineering Faculty Publications

Background

Spatiotemporal mapping of neural activity during continuous speech production has been traditionally approached using correlation coefficient (CC) analysis between cortical signals and speech recordings. A prior study employed this approach using electrocorticography (ECoG) data from participants who underwent invasive intracranial monitoring for epilepsy. However, CC cannot detect nonlinear relationships and is dominated by the correspondence between periods of silence and of non-silence.

New Method

We introduce the mutual information (MI) measure, which can capture both linear and nonlinear dependencies. We validated CC and MI on the sub-second spatiotemporal brain activity recorded during continuous speech tasks. To refine the results, …


Amplification Vs The Natural Ear: A Test On The Effectiveness Of The Natural Ear On Adults Ability To Match Pitch In Song, Celeste Orozco Jan 2019

Amplification Vs The Natural Ear: A Test On The Effectiveness Of The Natural Ear On Adults Ability To Match Pitch In Song, Celeste Orozco

Open Access Theses & Dissertations

Background: Singing is a natural enjoyment of life; however, individuals tend to isolate themselves from this enjoyment due to their inability to match pitch accurately. A new technology, the Natural Ear provides altered auditory feedback to the user while singing. It is hypothesized that this feedback may aid in the userâ??s ability to match pitch.

Purpose: The purpose of this study is to compare the effects of the Natural Ear to amplification and no amplification conditions on pitch matching accuracy in song.

Study Design: This study used a complex counterbalance within-subjects design.

Methods: 50 adults from the El Paso Metropolitan …


Across-Speaker Articulatory Normalization For Speaker-Independent Silent Speech Recognition, Jun Wang, Ashok Samal, Jordan Green Sep 2014

Across-Speaker Articulatory Normalization For Speaker-Independent Silent Speech Recognition, Jun Wang, Ashok Samal, Jordan Green

School of Computing: Conference and Workshop Papers

Silent speech interfaces (SSIs), which recognize speech from articulatory information (i.e., without using audio information), have the potential to enable persons with laryngectomy or a neurological disease to produce synthesized speech with a natural sounding voice using their tongue and lips. Current approaches to SSIs have largely relied on speaker-dependent recognition models to minimize the negative effects of talker variation on recognition accuracy. Speaker-independent approaches are needed to reduce the large amount of training data required from each user; only limited articulatory samples are often available for persons with moderate to severe speech impairments, due to the logistic difficulty of …


Articulatory Distinctiveness Of Vowels And Consonants: A Data-Driven Approach, Jun Wang, Jordan R. Green, Ashok Samal, Yana Yunusova Oct 2013

Articulatory Distinctiveness Of Vowels And Consonants: A Data-Driven Approach, Jun Wang, Jordan R. Green, Ashok Samal, Yana Yunusova

School of Computing: Faculty Publications

Purpose: To quantify the articulatory distinctiveness of 8 major English vowels and 11 English consonants based on tongue and lip movement time series data using a data-driven approach.

Method: Tongue and lip movements of 8 vowels and 11 consonants from 10 healthy talkers were collected. First, classification accuracies were obtained using 2 complementary approaches: (a) Procrustes analysis and (b) a support vector machine. Procrustes distance was then used to measure the articulatory distinctiveness among vowels and consonants. Finally, the distance (distinctiveness) matrices of different vowel pairs and consonant pairs were used to derive articulatory vowel and consonant spaces …


Word Recognition From Continuous Articulatory Movement Time-Series Data Using Symbolic Representations, Jun Wang, Arvind Balasubramanian, Luis Mojica De La Vega, Jordan R. Green, Ashok Samal, Balakrishnan Prabhakaran Aug 2013

Word Recognition From Continuous Articulatory Movement Time-Series Data Using Symbolic Representations, Jun Wang, Arvind Balasubramanian, Luis Mojica De La Vega, Jordan R. Green, Ashok Samal, Balakrishnan Prabhakaran

School of Computing: Conference and Workshop Papers

Although still in experimental stage, articulation-based silent speech interfaces may have significant potential for facilitating oral communication in persons with voice and speech problems. An articulation-based silent speech interface converts articulatory movement information to audible words. The complexity of speech production mechanism (e.g., co-articulation) makes the conversion a formidable problem. In this paper, we reported a novel, real-time algorithm for recognizing words from continuous articulatory movements. This approach differed from prior work in that (1) it focused on word-level, rather than phoneme-level; (2) online segmentation and recognition were conducted at the same time; and (3) a symbolic representation (SAX) was …


Whole-Word Recognition From Articulatory Movements For Silent Speech Interfaces, Jun Wang, Ashok Samal, Jordan R. Green, Frank Rudzicz Sep 2012

Whole-Word Recognition From Articulatory Movements For Silent Speech Interfaces, Jun Wang, Ashok Samal, Jordan R. Green, Frank Rudzicz

Department of Special Education and Communication Disorders: Faculty Publications

Articulation-based silent speech interfaces convert silently produced speech movements into audible words. These systems are still in their experimental stages, but have significant potential for facilitating oral communication in persons with laryngectomy or speech impairments. In this paper, we report the result of a novel, real-time algorithm that recognizes whole-words based on articulatory movements. This approach differs from prior work that has focused primarily on phoneme-level recognition based on articulatory features. On average, our algorithm missed 1.93 words in a sequence of twenty-five words with an average latency of 0.79 seconds for each word prediction using a data set of …


Sentence Recognition From Articulatory Movements For Silent Speech Interfaces, Jun Wang, Ashok Samal, Jordan R. Green, Frank Rudzicz Mar 2012

Sentence Recognition From Articulatory Movements For Silent Speech Interfaces, Jun Wang, Ashok Samal, Jordan R. Green, Frank Rudzicz

Department of Special Education and Communication Disorders: Faculty Publications

Recent research has demonstrated the potential of using an articulation-based silent speech interface for command-and-control systems. Such an interface converts articulation to words that can then drive a text-to-speech synthesizer. In this paper, we have proposed a novel near-time algorithm to recognize whole-sentences from continuous tongue and lip movements. Our goal is to assist persons who are aphonic or have a severe motor speech impairment to produce functional speech using their tongue and lips. Our algorithm was tested using a functional sentence data set collected from ten speakers (3012 utterances). The average accuracy was 94.89% with an average latency of …


Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally Jul 2006

Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally

Electrical & Computer Engineering Theses & Dissertations

Automatic Speaker Recognition is the process of automatically recognizing who is speaking on the basis of individual information contained in speech signals. This technique of Automatic Speaker Recognition makes it possible to use the speaker's voice to verify their identity and control access to services such as voice dialing, banking by telephone, telephone shopping, database access services, information services, voice mail, security control for confidential information areas, and remote access to computers.

In this thesis, the techniques of Gaussian Mixture Models and Neural Networks for Automatic Speaker Identification are presented. Algorithms for Speaker Identification using Gaussian Mixture Models were developed, …


A Computer-Based Articulation Training Aid For Short Words (Cata), Mukund Devarajan Oct 2003

A Computer-Based Articulation Training Aid For Short Words (Cata), Mukund Devarajan

Electrical & Computer Engineering Theses & Dissertations

Several improvements in the vowel articulation training aid (VATA) are described, as well as the efforts to extend the visual feedback system to operate with short words in the form of consonant, vowel and consonant (CVC). The extended version of the visual feedback system is referred to as CATA (Computer-based Articulation Training Aid); the vowel version of the aid (VATA) only operates with ten American English monopthong vowels. Improvements in VATA include the use of a neural network (NN) recognizer method to prune a large database of vowel recordings to eliminate noisy and/or mispronounced tokens. The spectral jitter problem, previously …


Automatic Speaker Identification Using Reusable And Retrainable Binary-Pair Partitioned Neural Networks, Ashutosh Mishra Apr 2003

Automatic Speaker Identification Using Reusable And Retrainable Binary-Pair Partitioned Neural Networks, Ashutosh Mishra

Electrical & Computer Engineering Theses & Dissertations

This thesis presents an extension of the work previously done on speaker identification using Binary Pair Partitioned (BPP) neural networks. In the previous work, a separate network was used for each pair of speakers in the speaker population. Although the basic BPP approach did perform well and had a simple underlying algorithm, it had the obvious disadvantage of requiring an extremely large number of networks for speaker identification with large speaker populations. It also requires training of networks proportional to the square of the number of speakers under consideration, leading to a very large number of networks to be trained …


Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi Oct 2002

Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi

Electrical & Computer Engineering Theses & Dissertations

This thesis presents a pitch detection algorithm that is extremely robust for both high quality and telephone speech. The kernel method for this algorithm is the Normalized Cross Correlation (NCCF) reported by David Talkin [16]. Major innovations include: processing of the original acoustic signal and a nonlinearly processed version of the signal to partially restore very weak F0 components; intelligent peak picking to select multiple F0 candidates and assign merit factors; and, incorporation of highly robust pitch contours obtained from smoothed versions of low frequency portions of spectrograms. Dynamic programming is used to find the ''best" pitch track among all …


Minimum Mean Square Error Spectral Peak Envelope Estimation For Automatic Vowel Classification, Jaishree Venugopal Jul 2001

Minimum Mean Square Error Spectral Peak Envelope Estimation For Automatic Vowel Classification, Jaishree Venugopal

Electrical & Computer Engineering Theses & Dissertations

Spectral feature computations continue to be a very difficult problem for accurate machine recognition of speech. In this work, which focuses on vowels, a new spectral peak envelope method for vowel classification is developed, based on a missing frequency components model of speech recognition. According to the missing frequency components model, vowel recognition depends only on the spectral (harmonic) peaks. Smoothing and interpolation of the spectra, performed in the standard cepstral analysis method commonly used in automatic speech recognition, actually loses valuable information and results in reduced recognition accuracy. The new method for feature extraction presented in this thesis is …


Variability Analysis Of Discrete Cosine Transform Coefficient (Dctc) Features For Speech Processing, Bingjun Dai Oct 1998

Variability Analysis Of Discrete Cosine Transform Coefficient (Dctc) Features For Speech Processing, Bingjun Dai

Electrical & Computer Engineering Theses & Dissertations

In this research, the variability of Discrete Cosine Transform Coefficient (DCTC) features was investigated. Additionally, a new pitch-synchronous processing method was explored to increase the stability of features and to reduce window effects when compared to the regular method. The noise sources that lead to feature variability were analyzed, and different smoothing methods were tested. It was found that longer frames, frequency warping, time smoothing of the log spectrum, and DCS level time smoothing, all help reduce DCTC variability and increase classification performance. The pitch­ synchronous method was implemented with Matlab. Important processing methods, including pitch period estimation, time­ domain …


Text Independent Speaker Verification Using Binary-Pair Partitioned Neural Networks, Claude A. Norton Iii Oct 1995

Text Independent Speaker Verification Using Binary-Pair Partitioned Neural Networks, Claude A. Norton Iii

Electrical & Computer Engineering Theses & Dissertations

A method is presented for the application of binary-pair partitioned neural networks to the task of speaker verification. This technique is based on a previously developed neural network classifier for speaker identification.

The main focus of this research was the development and testing of the algorithms necessary to extend the binary-pair partitioning approach from speaker identification to speaker verification. The method is based on the development of a user profile which is obtained from discriminative data provided by the binary-pair partitioned neural networks.

Experimental results are provided which demonstrate the viability of this approach, using the TIMIT speech corpus for …


Graduate Bulletin, 1995-1996 (1995), Moorhead State University Jan 1995

Graduate Bulletin, 1995-1996 (1995), Moorhead State University

Graduate Bulletins (Catalogs)

No abstract provided.


Graduate Bulletin, 1993-1995, Moorhead State University Jan 1993

Graduate Bulletin, 1993-1995, Moorhead State University

Graduate Bulletins (Catalogs)

No abstract provided.


Formant Estimation From Dctc's Using A Feedforward Neural Network, Shubhangi U. Kelkar Apr 1992

Formant Estimation From Dctc's Using A Feedforward Neural Network, Shubhangi U. Kelkar

Electrical & Computer Engineering Theses & Dissertations

Formants are the natural frequencies of the human vocal tract. Existing methods for estimating formants from speech signals are computationally complex and subject to errors for certain type of speech sounds. This thesis describes a method for estimating vowel formant frequencies from Discrete Cosine Transform Coefficients (DCTC's), a form of cepstral coefficients, using a feedforward neural network with back-propagation training. Experimental results are based on a large multispeaker data base. The results are obtained for both a linear transformation and a feedforward neural network with a nonlinear hidden layer. In general, the neural network transformation is superior to the linear …


Graduate Bulletin, 1991-1993, Moorhead State University Jan 1991

Graduate Bulletin, 1991-1993, Moorhead State University

Graduate Bulletins (Catalogs)

No abstract provided.


Visual Speech Training Aid For The Deaf, Subhashri Venkat Jul 1990

Visual Speech Training Aid For The Deaf, Subhashri Venkat

Electrical & Computer Engineering Theses & Dissertations

A computer-based vowel articulation training aid has been developed. A "continuous" acoustic-phonetic transformation is performed to map speech parameters to a lower dimensionality display space. There are two possible approaches to this transformation problem. The transformation could be either linear or a combination nonlinear/linear. The nonlinear transformation is performed using a multi-layered feedforward neural network with linear output layers. Speech parameters are extracted either from an analog filter bank arrangement (band energies) or by a digital signal processing procedure (Discrete Cosine Transform Coefficients). The speech parameters obtained from both methods correspond to the spectral envelope of the speech signals. The …


Color Display Of Vowel Spectra As A Training Aid For The Deaf, Amir Jalali Jagharghi Jul 1985

Color Display Of Vowel Spectra As A Training Aid For The Deaf, Amir Jalali Jagharghi

Electrical & Computer Engineering Theses & Dissertations

The objective of this research was to develop a transformation for mapping speech parameters to color parameter. This transformation is done in real-time, and the resulting color parameter are continuously displayed on a color monitor. This visual speech display is to be used as a speech articulation training aid for the deaf. The conversion of speech acoustic signals into speech parameter was accomplished using special -purpose electronics. The real-time conversion of speech parameter to display parameter was controlled by an 8086/8088 microprocessor operating in an S-100 bus structure. The coefficients of the Karhunen-Loeve series expansion of speech power spectra were …


Graduate Bulletin, 1985-1987 (1985), Moorhead State University Jan 1985

Graduate Bulletin, 1985-1987 (1985), Moorhead State University

Graduate Bulletins (Catalogs)

No abstract provided.


Block Encoding Of Speech Spectral Principal Components, James R. Holland Jr. Jul 1984

Block Encoding Of Speech Spectral Principal Components, James R. Holland Jr.

Electrical & Computer Engineering Theses & Dissertations

A Karhunen-Loeve series expansion was used to block encode speech spectral principal components as a function of time. Each of ten principal components was first obtained as a linear combination of 2© speech spectral band energies. Using a fixed block length of 10 frames (0.128 s), the K-L basis vectors were computed separately for various speakers for each principal component. In all cases the resulting basis vectors were essentially a set of discrete cosine basis vectors. Synthesis of speech from the block encoded parameters showed that very little information is lost with up to 70% data reduction. The block encoding …


Graduate Bulletin, 1982-1984 (1982), Moorhead State University Jan 1982

Graduate Bulletin, 1982-1984 (1982), Moorhead State University

Graduate Bulletins (Catalogs)

No abstract provided.