Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (5)
- Physical Sciences and Mathematics (5)
- Arts and Humanities (3)
- Morphology (2)
- Phonetics and Phonology (2)
-
- Anthropological Linguistics and Sociolinguistics (1)
- Artificial Intelligence and Robotics (1)
- Communication Sciences and Disorders (1)
- Comparative and Historical Linguistics (1)
- Computer Engineering (1)
- Databases and Information Systems (1)
- Digital Communications and Networking (1)
- Discourse and Text Linguistics (1)
- Electrical and Computer Engineering (1)
- Engineering (1)
- Language Interpretation and Translation (1)
- Medicine and Health Sciences (1)
- Other Computer Sciences (1)
- Other Linguistics (1)
- Public Affairs, Public Policy and Public Administration (1)
- Public Policy (1)
- Russian Linguistics (1)
- Signal Processing (1)
- Slavic Languages and Societies (1)
- Spanish Linguistics (1)
- Spanish and Portuguese Language and Literature (1)
- Speech and Hearing Science (1)
- Institution
- Publication Year
- Publication
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Computational Linguistics
Topics For He But Not For She: Quantifying And Classifying Gender Bias In The Media, Tyler J. Lanni
Topics For He But Not For She: Quantifying And Classifying Gender Bias In The Media, Tyler J. Lanni
Dissertations, Theses, and Capstone Projects
In this study, we used computational techniques to analyze the language used in news articles to describe female and male politicians. Our corpus included 370 subtexts for male candidates and 374 subtexts for female candidates, gathered through the New York Times API. We conducted two experiments: an LDA topic analysis to explore the data, and a logistic regression to classify the subtexts as either male or female. Our analysis revealed some noteworthy findings that suggest the possibility of developing a gender bias classifier in the future. However, to create a more robust understanding of bias, additional research and data are …
From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur
From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur
Dissertations, Theses, and Capstone Projects
Nearly thirty years ago, the United States Supreme Court revaluated the criteria for accepting forensic science and expert testimony, challenging Forensic Linguistics to assert itself as a reputable science. Much work has been produced in the interim to that end, but much still needs to be accomplished to satisfy the judicial standards. Computational linguistics has the potential to provide that necessary analytical framework. This paper’s intent is two-fold. First, there are two competing theories on the proper features necessary to identify an unknown author. Four features were drawn from the syntactic computational linguistics tradition and four from computational stylometry to …
A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco
A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco
Dissertations, Theses, and Capstone Projects
Code-switching is the linguistic phenomenon where a multilingual person alternates between two or more languages in a conversation, whether that be spoken or written. This thesis studies the automatic detection of code-switching occurring specifically between English and Spanish in two corpora.
Twitter and other social media sites have provided an abundance of linguistic data that is available to researchers to perform countless experiments. Collecting the data is fairly easy if a study is on monolingual text, but if a study requires code-switched data, this becomes a complication as APIs only accept one language as a parameter. This thesis focuses on …
Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn
Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn
Library Faculty Publications
This year’s iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems’ predictions. Transformer-based …
Empirical Analysis Of Cbow And Skip Gram Nlp Models, Tejas Menon
Empirical Analysis Of Cbow And Skip Gram Nlp Models, Tejas Menon
University Honors Theses
CBOW and Skip Gram are two NLP techniques to produce word embedding models that are accurate and performant. They were invented in the seminal paper by T. Mikolov et al. and have since observed optimizations such as negative sampling and subsampling. This paper implements a fully-optimized version of these models using Py-Torch and runs them through a toy sentiment/subject analysis. It is weakly observed that different corpus types affect the skew of word embeddings such that fictional corpus are better suited for sentiment analysis and non-fictional for subject analysis.
Automatic Keyphrase Extraction From Russian-Language Scholarly Papers In Computational Linguistics, Yves Wienecke
Automatic Keyphrase Extraction From Russian-Language Scholarly Papers In Computational Linguistics, Yves Wienecke
University Honors Theses
The automatic extraction of keyphrases from scholarly papers is a necessary step for many Natural Language Processing (NLP) tasks, including text retrieval, machine translation, and text summarization. However, due to the different grammatical and semantic intricacies of languages, this is a highly language-dependent task. Many free and open source implementations of state-of-the-art keyphrase extraction techniques exist, but they are not adapted for processing Russian text. Furthermore, the multi-linguistic character of scholarly papers in the field of Russian computational linguistics and NLP introduces additional complexity to keyphrase extraction. This paper describes a free and open source program as a proof of …
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Dissertations, Theses, and Capstone Projects
The UMETRICS database (Universities: Measuring the Effects of Research on Innovation, Competitiveness, and Science) contains rich information on grants from sponsored federal and non-federal research for 32 universities over a 15-year period. It is hosted at IRIS (Institute for Research on Innovation and Science, University of Michigan) and serves as a rich source of university administrative data; however, it does not contain information on research fields. Categorizing grants data by research field can help to measure results of investment in research and science and provide evidence for the data-driven policy-making; yet administrative data often lacks this type of categorization. In …
Obfuscating Authorship: Results Of A User Study On Nondescript, A Digital Privacy Tool, Robin Camille Davis
Obfuscating Authorship: Results Of A User Study On Nondescript, A Digital Privacy Tool, Robin Camille Davis
Publications and Research
For those who write anonymously, particularly for safety reasons, authorship attribution poses a threat. Nondescript, my web app, guides writers in achieving stylometric obfuscation in order to preserve anonymity. The app runs simulations of authorship attribution scenarios by analyzing the user’s linguistic features. In this paper, I will describe the conception of the Nondescript app; discuss related work; and present the results of a user study. Most users in the study were able to anonymize their writing in at least 5 out of 10 authorship attribution scenarios. Users rated the anonymization process an average of 3.6 out of 5 in …
Intergroup Variability In Personality Recognition, Arundhati Sengupta
Intergroup Variability In Personality Recognition, Arundhati Sengupta
Dissertations, Theses, and Capstone Projects
Automatic Identification of personality in conversational speech has many applications in natural language processing such as leader identification in a meeting, adaptive dialogue systems, and dating websites. However, the widespread acceptance of automatic personality recognition through lexical and vocal characteristics is limited by the variability of error rate in a general purpose model among speakers from different demographic groups. While other work reports accuracy, we explored error rates of automatic personality recognition task using classification models for different genders and native language groups (L1). We also present a statistical experiment showing the influence of gender and L1 on the relation …
A Markedly Different Approach: Investigating Pie Stops Using Modern Empirical Methods, Phillip Barnett
A Markedly Different Approach: Investigating Pie Stops Using Modern Empirical Methods, Phillip Barnett
Theses and Dissertations--Linguistics
In this thesis, I investigate a decades-old problem found in the stop system of Proto-Indo-European (PIE). More specifically, I will be investigating the paucity of */b/ in the forms reconstructed for the ancient, hypothetical language. As cross-linguistic evidence and phonological theory alone have fallen short of providing a satisfactory answer, herein will I employ modern empirical methods of linguistic investigation, namely laboratory phonology experiments and computational database analysis. Following Byrd 2015, I advocate for an examination of synchronic phenomena and behavior as a method for investigating diachronic change.
In Chapter 1, I present an overview of the various proposed phonological …
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Dissertations, Theses, and Capstone Projects
The process of producing written text is complex and constrained by pressures that range from physical to psychological. In a series of three sets of experiments, this thesis demonstrates the effects of linguistic context on the timing patterns of the production of keystrokes. We elucidate the effect of linguistic context at three different levels of granularity: The first set of experiments illustrate how the nontraditional syntax of a single linguistic construct, the multi-word expression, can create significant changes in keystroke production patterns. This set of experiments is followed by a set of experiments that test the hypothesis on the entire …
Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale
Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale
Theses and Dissertations--Linguistics
In Paradigm Function Morphology, it is usual to model affix position classes with an ordered sequence of inflectional rule blocks. Each rule block determines how (or whether) a particular affix position is filled. In this model, competition among inflectional rules is assumed to be limited to members of the same rule block; thus, the appearance of an affix in one position cannot be precluded by the appearance of an affix in another position. I present evidence that apparently disconfirms this restriction and suggests that a more general conception of rule competition is necessary. The data appear to imply that an …
Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier
Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier
Publications and Research
We investigate how novel English-derived words (anglicisms) are used in a German-language Internet hip hop forum, and what factors contribute to their uptake.
Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner
Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner
Department of Linguistics Faculty Scholarship and Creative Works
The Penn Forced Aligner automates the alignment process using the Hidden Markov Model Toolkit (HTK). The core of Prosodylab-Aligner is align.py, a script which performs acoustic model training and alignment. This script automates calls to HTK and SoX, an open-source command-line tool which is capable of resampling audio. The included README file provides instructions for installing HTK and SoX on Linux and Mac OS X, and can also be run on Windows. During training, the model is initialized with flat-start monophones, which are then submitted to a single round of model estimation. Then, a tied-state 'small pause' model is inserted …
Statistical Machine Translation Of Japanese, Erik A. Chapla
Statistical Machine Translation Of Japanese, Erik A. Chapla
Theses and Dissertations
The purpose of this research was to find ways to improve the performance of a statistical machine translation system that translates text from Japanese to English. Methods included altering the training and test data by adding a prior linguistic knowledge, altering sentence structures, and looking for better ways to statistically alter the way words align between the two languages. In addition, methods for properly segmenting words in Japanese text through statistical methods were examined. Finally, experiments were conducted on Japanese speech to produce the best text transcription of the speech. The best statistical machine translation methods implemented resulted in improvements …
A Classifier To Evaluate Language Specificity In Medical Documents, Trudi Miller '08, Gondy A. Leroy, Samir Chatterjee, Jie Fan, Brian Thoms '09
A Classifier To Evaluate Language Specificity In Medical Documents, Trudi Miller '08, Gondy A. Leroy, Samir Chatterjee, Jie Fan, Brian Thoms '09
CGU Faculty Publications and Research
Consumer health information written by health care professionals is often inaccessible to the consumers it is written for. Traditional readability formulas examine syntactic features like sentence length and number of syllables, ignoring the target audience's grasp of the words themselves. The use of specialized vocabulary disrupts the understanding of patients with low reading skills, causing a decrease in comprehension. A naive Bayes classifier for three levels of increasing medical terminology specificity (consumer/patient, novice health learner, medical professional) was created with a lexicon generated from a representative medical corpus. Ninety-six percent accuracy in classification was attained. The classifier was then applied …
Resolving Automatic Prepositional Phrase Attachments By Non-Statistical Means, Deryle W. Lonsdale, Michael B. Manookin
Resolving Automatic Prepositional Phrase Attachments By Non-Statistical Means, Deryle W. Lonsdale, Michael B. Manookin
Faculty Publications
Prepositional-phrase attachment is a topic of active research in the field of computational linguistics. Properly attaching prepositional phrases to their pertinent constituent proves straightforward for humans, but inferring these attachments in a cognitive modeling system becomes difficult. For example, in the sentence, ‘Ralph threw the frisbee to John,’ the prepositional phrase ‘to John’ will attach to the verb phrase ‘threw’. In another example, ‘Joe saw the dog with fur,’ the prepositional phrase ‘with fur’ will attach directly to the noun phrase ‘the dog.’ Humans would have little difficulty resolving these examples, but for computers this would be difficult.
Visual Speech Training Aid For The Deaf, Subhashri Venkat
Visual Speech Training Aid For The Deaf, Subhashri Venkat
Electrical & Computer Engineering Theses & Dissertations
A computer-based vowel articulation training aid has been developed. A "continuous" acoustic-phonetic transformation is performed to map speech parameters to a lower dimensionality display space. There are two possible approaches to this transformation problem. The transformation could be either linear or a combination nonlinear/linear. The nonlinear transformation is performed using a multi-layered feedforward neural network with linear output layers. Speech parameters are extracted either from an analog filter bank arrangement (band energies) or by a digital signal processing procedure (Discrete Cosine Transform Coefficients). The speech parameters obtained from both methods correspond to the spectral envelope of the speech signals. The …