Open Access. Powered by Scholars. Published by Universities.®

Signal Processing Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 13 of 13

Full-Text Articles in Signal Processing

Spoken Language Processing And Modeling For Aviation Communications, Aaron Van De Brook Oct 2023

Spoken Language Processing And Modeling For Aviation Communications, Aaron Van De Brook

Doctoral Dissertations and Master's Theses

With recent advances in machine learning and deep learning technologies and the creation of larger aviation-specific corpora, applying natural language processing technologies, especially those based on transformer neural networks, to aviation communications is becoming increasingly feasible. Previous work has focused on machine learning applications to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained checkpoints and those trained from their base weight initializations (trained from scratch). The results suggest that transformer language …


Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally Jul 2006

Gaussian Mixture Models And Neural Networks For Automatic Speaker Identification, Usha Gayatri Chalkapally

Electrical & Computer Engineering Theses & Dissertations

Automatic Speaker Recognition is the process of automatically recognizing who is speaking on the basis of individual information contained in speech signals. This technique of Automatic Speaker Recognition makes it possible to use the speaker's voice to verify their identity and control access to services such as voice dialing, banking by telephone, telephone shopping, database access services, information services, voice mail, security control for confidential information areas, and remote access to computers.

In this thesis, the techniques of Gaussian Mixture Models and Neural Networks for Automatic Speaker Identification are presented. Algorithms for Speaker Identification using Gaussian Mixture Models were developed, …


Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi Oct 2002

Yet Another Algorithm For Pitch Tracking (Yaapt), Kavita Kasi

Electrical & Computer Engineering Theses & Dissertations

This thesis presents a pitch detection algorithm that is extremely robust for both high quality and telephone speech. The kernel method for this algorithm is the Normalized Cross Correlation (NCCF) reported by David Talkin [16]. Major innovations include: processing of the original acoustic signal and a nonlinearly processed version of the signal to partially restore very weak F0 components; intelligent peak picking to select multiple F0 candidates and assign merit factors; and, incorporation of highly robust pitch contours obtained from smoothed versions of low frequency portions of spectrograms. Dynamic programming is used to find the ''best" pitch track among all …


Aurora Working Group: Dsr Front End Lvcsr Evaluation — Baseline Recognition System Description, Naveen Parihar, Joseph Picone Jul 2001

Aurora Working Group: Dsr Front End Lvcsr Evaluation — Baseline Recognition System Description, Naveen Parihar, Joseph Picone

Publications

In this document we describe the features of the baseline system to be used in the Distributed Speech Recognition (DSR) front end large vocabulary continuous speech recognition (LVCSR) evaluations being conducted by the Aurora Working Group of the European Telecommunications Standards Institute (ETSI). The objective of these evaluations is to determine the robustness of different front ends for use in client/server type telecommunications applications. As such, our experiments are designed to test the following focus conditions on the DARPA Wall Street Journal (WSJ0) corpus using a 5000-word closed-loop vocabulary and a bigram language model:

  1. Additive Noise: six noise conditions …


Voice Command Controller, Hoang Nghia Nguyen Jan 2000

Voice Command Controller, Hoang Nghia Nguyen

Theses : Honours

Signal processing technology has been strongly developed and it has attracted interest from scientists and engineers around the world from the last decade. Speech synthesis and speech recognition are particular topic in the field that have been widely used and developed in many different area such as business, controlling, education and entertainment. The project's main objective is to study and develop an application program with the Speech SDK through design and implementation of Tele-Control system based on the commercial product of National Semiconductor: Carrier-Current Transceiver (LM 1893) and Speech development kit (Speech SDK4.0) from Microsoft Corporation. The project is suitable …


Visual Speech Recognition Using Multiple Deformable Lip Models, Devi Chandramohan Jul 1996

Visual Speech Recognition Using Multiple Deformable Lip Models, Devi Chandramohan

Electrical & Computer Engineering Theses & Dissertations

Motivated by the fact that human speech perception is a bimodal process (auditory and visual), several researchers have designed and implemented automatic speech recognition (ASR) systems consisting of both audio and visual subsystems, and shown improved performance relative to traditional purely auditory systems. Several visual speech reading approaches have used deformable templates to model the shape of a speaker's lips. Deformable templates are models of image objects, which can be deformed by adjusting a set of parameters to match the object in some optimal way, as defined by a cost function. Using a single deformable lip model has disadvantages such …


Adaptive Integration Of Audio And Visual Information Using Discrete And Semi-Continuous Hidden Markov Models In Audiovisual Automatic Speech Recognition, Qin Su Apr 1996

Adaptive Integration Of Audio And Visual Information Using Discrete And Semi-Continuous Hidden Markov Models In Audiovisual Automatic Speech Recognition, Qin Su

Electrical & Computer Engineering Theses & Dissertations

An audiovisual semi-continuous hidden Markov model (HMM)-based Automatic Speech Recognition (ASR) system and an improved method of integrating audio and visual information in an audiovisual discrete HMM-based ASR system are investigated.

In the audiovisual discrete HMM, an adaptive integration formulation is employed, which incorporates the integration into the HMM at a pre-categorical stage. A visual weighting parameter is determined automatically, which allows the relative contribution of audio and visual information to be adjusted adaptively. Using an adaptive weight, the accuracy increased by 13% compared to the same model with no adaptive weight.

The semi-continuous HMM is a class of models …


Text-Independent, Open-Set Speaker Recognition, Stephen V. Pellissier Mar 1996

Text-Independent, Open-Set Speaker Recognition, Stephen V. Pellissier

Theses and Dissertations

Speaker recognition, like other biometric personal identification techniques, depends upon a person's intrinsic characteristics. A realistically viable system must be capable of dealing with the open-set task. This effort attacks the open-set task, identifying the best features to use, and proposes the use of a fuzzy classifier followed by hypothesis testing as a model for text-independent, open-set speaker recognition. Using the TIMIT corpus and Rome Laboratory's GREENFLAG tactical communications corpus, this thesis demonstrates that the proposed system succeeded in open-set speaker recognition. Considering the fact that extremely short utterances were used to train the system (compared to other closed-set speaker …


Generalized Hidden Filter Markov Models Applied To Speaker Recognition, John M. Colombi Mar 1996

Generalized Hidden Filter Markov Models Applied To Speaker Recognition, John M. Colombi

Theses and Dissertations

Classification of time series has wide Air Force, DoD and commercial interest, from automatic target recognition systems on munitions to recognition of speakers in diverse environments. The ability to effectively model the temporal information contained in a sequence is of paramount importance. Toward this goal, this research develops theoretical extensions to a class of stochastic models and demonstrates their effectiveness on the problem of text-independent (language constrained) speaker recognition. Specifically within the hidden Markov model architecture, additional constraints are implemented which better incorporate observation correlations and context, where standard approaches fail. Two methods of modeling correlations are developed, and their …


Clustering Techniques In Speaker Recognition, Douglas N. Prescott Mar 1994

Clustering Techniques In Speaker Recognition, Douglas N. Prescott

Theses and Dissertations

This thesis presents a comparison based on identification rate, of three clustering techniques applied to cepstral features for speaker identification. LBG vector quantization as developed by Linde, Buzo and Gray; is used to provide benchmark performance for comparison with Fuzzy clustering (based on the unsupervised fuzzy partition-optimal number of classes, UFP-ONC algorithm by Gath and Geva) and an Artificial Neural Network, the Multilayer Perceptron. Cepstral features from the TIMIT, King and AFIT93 corpus speaker databases are used to produce speaker-identification classifiers using each of the clustering algorithms. The experiment reported evaluates the speaker identification performance using the 20-dimensional cepstral features …


Cepstral And Auditory Model Features For Speaker Recognition, John M. Colombi Dec 1992

Cepstral And Auditory Model Features For Speaker Recognition, John M. Colombi

Theses and Dissertations

The TIMIT and KING databases, as well as a ten day AFIT speaker corpus, are used to compare proven spectral processing techniques to an auditory neural representation for speaker identification. The feature sets compared were Linear Predictive Coding (LPC) cepstral coefficients and auditory nerve firing rates using the Payton model. This auditory model provides for the mechanisms found in the human middle and inner auditory periphery as well as neural transduction. Clustering algorithms were used to generate speaker specific codebooks - one statistically based and the other a neural approach. These algorithms are the Linde-Buzo-Gray (LBG) algorithm and a Kohonen …


Speech Recognition Using Visible And Infrared Detectors, Patrick T. Marshall Sep 1992

Speech Recognition Using Visible And Infrared Detectors, Patrick T. Marshall

Theses and Dissertations

A system has been developed that tracks lip motion using infrared (IR) or visible detectors. The purpose of this study was to determine if the additional information obtained from the IR or visible detectors can be used to increase the recognition rate of audio Automatic Speech Recognition (ASR) systems. To accomplish this goal, several hardware analog prototypes had to be designed, built and tested. Different detectors (IR and visible) and modes of operation (active and passive) were tried before a reliable and useful signal was found. An analog-to-digital (A/D) board was then designed and built that digitized both the microphone …


Encoding Phonetic Knowledge For Use In Hidden Markov Models Of Speech Recognition, Danming Qian Jul 1990

Encoding Phonetic Knowledge For Use In Hidden Markov Models Of Speech Recognition, Danming Qian

Electrical & Computer Engineering Theses & Dissertations

Hidden Markov models (HMM's) have achieved considerable success for isolated-word speaker-independent automatic speech recognition. However, the performance of an HMM algorithm is limited by its inability to discriminate between similar sounding words. The problem arises because all differences between speech patterns are treated as equally important. Thus the algorithm is particularly susceptible to confusions caused by phonetically-irrelevant differences. This thesis presents two types of preprocessing schemes as candidates for improving HMM performance. The aim is to maximize the differences between phonologically-distinct speech sounds while minimizing the effect of variations in phonologically-equivalent speech sounds. The preprocessors presented are a discrete cosine …