Open Access. Powered by Scholars. Published by Universities.®
- Discipline
- Institution
- Publication
- Publication Type
Articles 1 - 13 of 13
Full-Text Articles in Signal Processing
Spoken Language Processing And Modeling For Aviation Communications, Aaron Van De Brook
Spoken Language Processing And Modeling For Aviation Communications, Aaron Van De Brook
Doctoral Dissertations and Master's Theses
With recent advances in machine learning and deep learning technologies and the creation of larger aviation-specific corpora, applying natural language processing technologies, especially those based on transformer neural networks, to aviation communications is becoming increasingly feasible. Previous work has focused on machine learning applications to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained checkpoints and those trained from their base weight initializations (trained from scratch). The results suggest that transformer language …
Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally
Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally
Electrical & Computer Engineering Theses & Dissertations
Automatic Speaker Recognition is the process of automatically recognizing who is speaking on the basis of individual information contained in speech signals. This technique of Automatic Speaker Recognition makes it possible to use the speaker's voice to verify their identity and control access to services such as voice dialing, banking by telephone, telephone shopping, database access services, information services, voice mail, security control for confidential information areas, and remote access to computers.
In this thesis, the techniques of Gaussian Mixture Models and Neural Networks for Automatic Speaker Identification are presented. Algorithms for Speaker Identification using Gaussian Mixture Models were developed, …
Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi
Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi
Electrical & Computer Engineering Theses & Dissertations
This thesis presents a pitch detection algorithm that is extremely robust for both high quality and telephone speech. The kernel method for this algorithm is the Normalized Cross Correlation (NCCF) reported by David Talkin [16]. Major innovations include: processing of the original acoustic signal and a nonlinearly processed version of the signal to partially restore very weak F0 components; intelligent peak picking to select multiple F0 candidates and assign merit factors; and, incorporation of highly robust pitch contours obtained from smoothed versions of low frequency portions of spectrograms. Dynamic programming is used to find the ''best" pitch track among all …
Aurora Working Group: Dsr Front End Lvcsr Evaluation — Baseline Recognition System Description, Naveen Parihar, Joseph Picone
Aurora Working Group: Dsr Front End Lvcsr Evaluation — Baseline Recognition System Description, Naveen Parihar, Joseph Picone
Publications
In this document we describe the features of the baseline system to be used in the Distributed Speech Recognition (DSR) front end large vocabulary continuous speech recognition (LVCSR) evaluations being conducted by the Aurora Working Group of the European Telecommunications Standards Institute (ETSI). The objective of these evaluations is to determine the robustness of different front ends for use in client/server type telecommunications applications. As such, our experiments are designed to test the following focus conditions on the DARPA Wall Street Journal (WSJ0) corpus using a 5000-word closed-loop vocabulary and a bigram language model:
- Additive Noise: six noise conditions …
Voice Command Controller, Hoang Nghia Nguyen
Voice Command Controller, Hoang Nghia Nguyen
Theses : Honours
Signal processing technology has been strongly developed and it has attracted interest from scientists and engineers around the world from the last decade. Speech synthesis and speech recognition are particular topic in the field that have been widely used and developed in many different area such as business, controlling, education and entertainment. The project's main objective is to study and develop an application program with the Speech SDK through design and implementation of Tele-Control system based on the commercial product of National Semiconductor: Carrier-Current Transceiver (LM 1893) and Speech development kit (Speech SDK4.0) from Microsoft Corporation. The project is suitable …
Visual Speech Recognition Using Multiple Deformable Lip Models, Devi Chandramohan
Visual Speech Recognition Using Multiple Deformable Lip Models, Devi Chandramohan
Electrical & Computer Engineering Theses & Dissertations
Motivated by the fact that human speech perception is a bimodal process (auditory and visual), several researchers have designed and implemented automatic speech recognition (ASR) systems consisting of both audio and visual subsystems, and shown improved performance relative to traditional purely auditory systems. Several visual speech reading approaches have used deformable templates to model the shape of a speaker's lips. Deformable templates are models of image objects, which can be deformed by adjusting a set of parameters to match the object in some optimal way, as defined by a cost function. Using a single deformable lip model has disadvantages such …
Adaptive Integration Of Audio And Visual Information Using Discrete And Semi-Continuous Hidden Markov Models In Audiovisual Automatic Speech Recognition, Qin Su
Electrical & Computer Engineering Theses & Dissertations
An audiovisual semi-continuous hidden Markov model (HMM)-based Automatic Speech Recognition (ASR) system and an improved method of integrating audio and visual information in an audiovisual discrete HMM-based ASR system are investigated.
In the audiovisual discrete HMM, an adaptive integration formulation is employed, which incorporates the integration into the HMM at a pre-categorical stage. A visual weighting parameter is determined automatically, which allows the relative contribution of audio and visual information to be adjusted adaptively. Using an adaptive weight, the accuracy increased by 13% compared to the same model with no adaptive weight.
The semi-continuous HMM is a class of models …
Text-Independent, Open-Set Speaker Recognition, Stephen V. Pellissier
Text-Independent, Open-Set Speaker Recognition, Stephen V. Pellissier
Theses and Dissertations
Speaker recognition, like other biometric personal identification techniques, depends upon a person's intrinsic characteristics. A realistically viable system must be capable of dealing with the open-set task. This effort attacks the open-set task, identifying the best features to use, and proposes the use of a fuzzy classifier followed by hypothesis testing as a model for text-independent, open-set speaker recognition. Using the TIMIT corpus and Rome Laboratory's GREENFLAG tactical communications corpus, this thesis demonstrates that the proposed system succeeded in open-set speaker recognition. Considering the fact that extremely short utterances were used to train the system (compared to other closed-set speaker …
Generalized Hidden Filter Markov Models Applied To Speaker Recognition, John M. Colombi
Generalized Hidden Filter Markov Models Applied To Speaker Recognition, John M. Colombi
Theses and Dissertations
Classification of time series has wide Air Force, DoD and commercial interest, from automatic target recognition systems on munitions to recognition of speakers in diverse environments. The ability to effectively model the temporal information contained in a sequence is of paramount importance. Toward this goal, this research develops theoretical extensions to a class of stochastic models and demonstrates their effectiveness on the problem of text-independent (language constrained) speaker recognition. Specifically within the hidden Markov model architecture, additional constraints are implemented which better incorporate observation correlations and context, where standard approaches fail. Two methods of modeling correlations are developed, and their …
Clustering Techniques In Speaker Recognition, Douglas N. Prescott
Clustering Techniques In Speaker Recognition, Douglas N. Prescott
Theses and Dissertations
This thesis presents a comparison based on identification rate, of three clustering techniques applied to cepstral features for speaker identification. LBG vector quantization as developed by Linde, Buzo and Gray; is used to provide benchmark performance for comparison with Fuzzy clustering (based on the unsupervised fuzzy partition-optimal number of classes, UFP-ONC algorithm by Gath and Geva) and an Artificial Neural Network, the Multilayer Perceptron. Cepstral features from the TIMIT, King and AFIT93 corpus speaker databases are used to produce speaker-identification classifiers using each of the clustering algorithms. The experiment reported evaluates the speaker identification performance using the 20-dimensional cepstral features …
Cepstral And Auditory Model Features For Speaker Recognition, John M. Colombi
Cepstral And Auditory Model Features For Speaker Recognition, John M. Colombi
Theses and Dissertations
The TIMIT and KING databases, as well as a ten day AFIT speaker corpus, are used to compare proven spectral processing techniques to an auditory neural representation for speaker identification. The feature sets compared were Linear Predictive Coding (LPC) cepstral coefficients and auditory nerve firing rates using the Payton model. This auditory model provides for the mechanisms found in the human middle and inner auditory periphery as well as neural transduction. Clustering algorithms were used to generate speaker specific codebooks - one statistically based and the other a neural approach. These algorithms are the Linde-Buzo-Gray (LBG) algorithm and a Kohonen …
Speech Recognition Using Visible And Infrared Detectors, Patrick T. Marshall
Speech Recognition Using Visible And Infrared Detectors, Patrick T. Marshall
Theses and Dissertations
A system has been developed that tracks lip motion using infrared (IR) or visible detectors. The purpose of this study was to determine if the additional information obtained from the IR or visible detectors can be used to increase the recognition rate of audio Automatic Speech Recognition (ASR) systems. To accomplish this goal, several hardware analog prototypes had to be designed, built and tested. Different detectors (IR and visible) and modes of operation (active and passive) were tried before a reliable and useful signal was found. An analog-to-digital (A/D) board was then designed and built that digitized both the microphone …
Encoding Phonetic Knowledge For Use In Hidden Markov Models Of Speech Recognition, Danming Qian
Encoding Phonetic Knowledge For Use In Hidden Markov Models Of Speech Recognition, Danming Qian
Electrical & Computer Engineering Theses & Dissertations
Hidden Markov models (HMM's) have achieved considerable success for isolated-word speaker-independent automatic speech recognition. However, the performance of an HMM algorithm is limited by its inability to discriminate between similar sounding words. The problem arises because all differences between speech patterns are treated as equally important. Thus the algorithm is particularly susceptible to confusions caused by phonetically-irrelevant differences. This thesis presents two types of preprocessing schemes as candidates for improving HMM performance. The aim is to maximize the differences between phonologically-distinct speech sounds while minimizing the effect of variations in phonologically-equivalent speech sounds. The preprocessors presented are a discrete cosine …