Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Arts and Humanities (6)
- Computer Sciences (5)
- First and Second Language Acquisition (5)
- Physical Sciences and Mathematics (5)
- Anthropological Linguistics and Sociolinguistics (4)
-
- Artificial Intelligence and Robotics (4)
- Psychology (4)
- Syntax (4)
- Discourse and Text Linguistics (3)
- Phonetics and Phonology (3)
- Psycholinguistics and Neurolinguistics (3)
- Semantics and Pragmatics (3)
- Cognitive Psychology (2)
- Morphology (2)
- Public Affairs, Public Policy and Public Administration (2)
- Appalachian Studies (1)
- Applied Linguistics (1)
- Business (1)
- Clinical Psychology (1)
- Cognition and Perception (1)
- Comparative and Historical Linguistics (1)
- Developmental Psychology (1)
- Digital Humanities (1)
- Experimental Analysis of Behavior (1)
- Feminist, Gender, and Sexuality Studies (1)
- Graphics and Human Computer Interfaces (1)
- Jewish Studies (1)
- Keyword
-
- Computational linguistics (6)
- Natural language processing (5)
- Machine learning (4)
- Natural Language Processing (4)
- Machine Learning (3)
-
- Twitter (3)
- Automatic Speech Recognition (2)
- Computational Linguistics (2)
- Deep learning (2)
- Entrainment (2)
- Grapheme-to-phoneme conversion (2)
- Language acquisition (2)
- Linguistics (2)
- Neural Networks (2)
- Neural network (2)
- Prosody (2)
- Russian (2)
- Sentiment analysis (2)
- Speech perception in noise (2)
- Stress (2)
- Text classification (2)
- Topic modeling (2)
- "bubble" noise (1)
- AI (1)
- ASR (1)
- Accessibility (1)
- Accommodation (1)
- Adverbs (1)
- Appalachian English (1)
- Arabic (1)
Articles 31 - 57 of 57
Full-Text Articles in Computational Linguistics
When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller
When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller
Dissertations, Theses, and Capstone Projects
As a subdomain of author profiling, gender prediction (sometimes called gender inference) has received a substantial amount of attention—both as a task in itself, and for other downstream analyses. Throughout the existing literature various statistical and machine learning methods have been applied to extract features in order to either characterize and differentiate female and male writing styles, or simply to achieve maximum accuracy on gender prediction as a binary classification task. However, researchers often do not disclose how they conceptualize gender nor do they consider the implications that gender prediction has for non-binary and trans individuals. Along with an overview …
Mitigating Gender Bias In Neural Machine Translation Using Counterfactual Data, Alan Wong
Mitigating Gender Bias In Neural Machine Translation Using Counterfactual Data, Alan Wong
Dissertations, Theses, and Capstone Projects
Recent advances in deep learning have greatly improved the ability of researchers to develop effective machine translation systems. In particular, the application of modern neural architectures, such as the Transformer, has achieved state-of-the-art BLEU scores in many translation tasks. However, it has been found that even state-of-the-art neural machine translation models can suffer from certain implicit biases, such as gender bias (Lu et al., 2019). In response to this issue, researchers have proposed various potential solutions: some have proposed approaches that inject missing gender information into models, while others have attempted modifying the training data itself. We focus on mitigating …
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Dissertations, Theses, and Capstone Projects
This thesis presents experiments with using representation learning to explore how neural networks learn. Neural networks which take text as input create internal representations of the text during their training. Recent work has found that these representations can be used to perform other downstream linguistic tasks, such as part-of-speech (POS) tagging. This demonstrates that the neural networks are learning linguistic information and storing this information in the representations. We focus on the representations created by neural machine translation (NMT) models and whether they can be used in POS tagging. We train 5 NMT models including an auto-encoder. We extract the …
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Dissertations, Theses, and Capstone Projects
The UMETRICS database (Universities: Measuring the Effects of Research on Innovation, Competitiveness, and Science) contains rich information on grants from sponsored federal and non-federal research for 32 universities over a 15-year period. It is hosted at IRIS (Institute for Research on Innovation and Science, University of Michigan) and serves as a rich source of university administrative data; however, it does not contain information on research fields. Categorizing grants data by research field can help to measure results of investment in research and science and provide evidence for the data-driven policy-making; yet administrative data often lacks this type of categorization. In …
Genderlects In Social Media, Alina Korovatskaya
Genderlects In Social Media, Alina Korovatskaya
Dissertations, Theses, and Capstone Projects
Many studies have found significant differences in ways men and women use language; some argue that these differences occur as a result of culture differences, and others suggest that they are influenced by differences in social status and power between the genders. However, some of the major studies were concluded decades ago and do not reflect changes in gender relations in recent years. In this study, we analyze modern conversations using two social media platforms, Twitter and Reddit, to determine whether substantial differences between men and women’s use of language were preserved between the genders.
Doing Away With Defaults: Motivation For A Gradient Parameter Space, Katherine Howitt
Doing Away With Defaults: Motivation For A Gradient Parameter Space, Katherine Howitt
Dissertations, Theses, and Capstone Projects
In this thesis, I propose a reconceptualization of the traditional syntactic parameter space of the principles and parameters framework (Chomsky, 1981). In lieu of binary parameter settings, parameter values exist on a gradient plane where a learner’s knowledge of their language is encoded in their confidence that a particular parametric target value, and thus grammatical construction of an encountered sentence, is likely to be licensed by their target grammar. First, I discuss other learnability models in the classic parameter space which lack either psychological plausibility, theoretical consistency, or some combination of the two. Then, I argue for the Gradient Parameter …
Computational Approaches To The Syntax–Prosody Interface: Using Prosody To Improve Parsing, Hussein M. Ghaly
Computational Approaches To The Syntax–Prosody Interface: Using Prosody To Improve Parsing, Hussein M. Ghaly
Dissertations, Theses, and Capstone Projects
Prosody has strong ties with syntax, since prosody can be used to resolve some syntactic ambiguities. Syntactic ambiguities have been shown to negatively impact automatic syntactic parsing, hence there is reason to believe that prosodic information can help improve parsing. This dissertation considers a number of approaches that aim to computationally examine the relationship between prosody and syntax of natural languages, while also addressing the role of syntactic phrase length, with the ultimate goal of using prosody to improve parsing.
Chapter 2 examines the effect of syntactic phrase length on prosody in double center embedded sentences in French. Data collected …
Phonologically-Informed Speech Coding For Automatic Speech Recognition-Based Foreign Language Pronunciation Training, Anthony J. Vicario
Phonologically-Informed Speech Coding For Automatic Speech Recognition-Based Foreign Language Pronunciation Training, Anthony J. Vicario
Dissertations, Theses, and Capstone Projects
Automatic speech recognition (ASR) and computer-assisted pronunciation training (CAPT) systems used in foreign-language educational contexts are often not developed with the specific task of second-language acquisition in mind. Systems that are built for this task are often excessively targeted to one native language (L1) or a single phonemic contrast and are therefore burdensome to train. Current algorithms have been shown to provide erroneous feedback to learners and show inconsistencies between human and computer perception. These discrepancies have thus far hindered more extensive application of ASR in educational systems.
This thesis reviews the computational models of the human perception of American …
Ghost Peppers: Using Ensemble Models To Detect Professor Attractiveness Commentary On Ratemyprofessors.Com, Angie Waller
Ghost Peppers: Using Ensemble Models To Detect Professor Attractiveness Commentary On Ratemyprofessors.Com, Angie Waller
Dissertations, Theses, and Capstone Projects
In June 2018, RateMyProfessors.com (RMP), a popular website for students to leave professor reviews, removed a controversial feature known as the “chili pepper” which allowed students to rate their professors as “hot” or “not hot.” Though past research has rigorously analyzed the correlation of the chili pepper with higher ratings in other categories (Felton, Mitchell, and Stinson, 2004; Felton et al., 2008), none has measured the effect of the removal of the chili pepper on the text content submitted by students. While it is a positive step that the chili pepper has been removed, text commentary on teacher attractiveness persists …
Demographic Factors As Domains For Adaptation In Linguistic Preprocessing, Sara Morini
Demographic Factors As Domains For Adaptation In Linguistic Preprocessing, Sara Morini
Dissertations, Theses, and Capstone Projects
Classic natural language processing resources such as the Penn Treebank (Marcus et al. 1993) have long been used both as evaluation data for many linguistic tasks and as training data for a variety of off-the-shelf language processing tools. Recent work has highlighted a gender imbalance in the authors of this text data (Garimella et al. 2019) and hypothesized that tools created with such resources will privilege users from particular demographic groups (Hovy and Søgaard 2015). Domain adaptation is typically employed as a strategy in machine learning to adjust models trained and evaluated with data from different genres. However, the present …
Do It Like A Syntactician: Using Binary Gramaticality Judgements To Train Sentence Encoders And Assess Their Sensitivity To Syntactic Structure, Pablo Gonzalez Martinez
Do It Like A Syntactician: Using Binary Gramaticality Judgements To Train Sentence Encoders And Assess Their Sensitivity To Syntactic Structure, Pablo Gonzalez Martinez
Dissertations, Theses, and Capstone Projects
The binary nature of grammaticality judgments and their use to access the structure of syntax are a staple of modern linguistics. However, computational models of natural language rarely make use of grammaticality in their training or application. Furthermore, developments in modern neural NLP have produced a myriad of methods that push the baselines in many complex tasks, but those methods are typically not evaluated from a linguistic perspective. In this dissertation I use grammaticality judgements with artificially generated ungrammatical sentences to assess the performance of several neural encoders and propose them as a suitable training target to make models learn …
Analyzing Prosody With Legendre Polynomial Coefficients, Rachel Rakov
Analyzing Prosody With Legendre Polynomial Coefficients, Rachel Rakov
Dissertations, Theses, and Capstone Projects
This investigation demonstrates the effectiveness of Legendre polynomial coefficients representing prosodic contours within the context of two different tasks: nativeness classification and sarcasm detection. By making use of accurate representations of prosodic contours to answer fundamental linguistic questions, we contribute significantly to the body of research focused on analyzing prosody in linguistics as well as modeling prosody for machine learning tasks. Using Legendre polynomial coefficient representations of prosodic contours, we answer prosodic questions about differences in prosody between native English speakers and non-native English speakers whose first language is Mandarin. We also learn more about prosodic qualities of sarcastic speech. …
The Perception Of Mandarin Tones In "Bubble" Noise By Native And L2 Listeners, Mengxuan Zhao
The Perception Of Mandarin Tones In "Bubble" Noise By Native And L2 Listeners, Mengxuan Zhao
Dissertations, Theses, and Capstone Projects
Previous studies have revealed the complexity of Mandarin Tones. For example, similarities in the pitch contours of tones 2 and 3 and tones 3 and 4 cause confusion for listeners. The realization of a tone's contour is highly dependent on its context, especially the preceding pitch. This is known as the coarticulation effect. Researchers have demonstrated the robustness of tone perception by both native and non-native listeners, even with incomplete acoustic information or in noisy environment. However, non-native listeners were observed to behave differently from native listeners in their use of contextual information. For example, the disagreement between the end …
Generative Adversarial Networks And Word Embeddings For Natural Language Generation, Robert D. Schultz Jr
Generative Adversarial Networks And Word Embeddings For Natural Language Generation, Robert D. Schultz Jr
Dissertations, Theses, and Capstone Projects
We explore using image generation techniques to generate natural language. Generative Adversarial Networks (GANs), normally used for image generation, were used for this task. To avoid using discrete data such as one-hot encoded vectors, with dimensions corresponding to vocabulary size, we instead use word embeddings as training data. The main motivation for this is the fact that a sentence translated into a sequence of word embeddings (a “word matrix”) is an analogue to a matrix of pixel values in an image. These word matrices can then be used to train a generative adversarial model. The output of the model’s generator …
Recursive Neural Networks For Semantic Sentence Representation, Liam S. Geron
Recursive Neural Networks For Semantic Sentence Representation, Liam S. Geron
Dissertations, Theses, and Capstone Projects
Semantic representation has a rich history rife with both complex linguistic theory and computational models. Though this history stretches back almost 50 years (Salton, 1971), recently the field has undergone an unexpected shift in paradigm thanks to the work of Mikolov et al., 2013(a & b) which has proven that vector-space semantic models can capture large amounts of semantic information. As of yet, these semantic representations are computed at the word level, and finding a semantic representation of a phrase is a much more difficult challenge. Mikolov et al., 2013(a&b) proved that their word vectors can be composed arithmetically to …
Multimodal Depression Detection: An Investigation Of Features And Fusion Techniques For Automated Systems, Michelle Renee Morales
Multimodal Depression Detection: An Investigation Of Features And Fusion Techniques For Automated Systems, Michelle Renee Morales
Dissertations, Theses, and Capstone Projects
Depression is a serious illness that affects a large portion of the world’s population. Given the large effect it has on society, it is evident that depression is a serious health issue. This thesis evaluates, at length, how technology may aid in assessing depression. We present an in-depth investigation of features and fusion techniques for depression detection systems. We also present OpenMM: a novel tool for multimodal feature extraction. Lastly, we present novel techniques for multimodal fusion. The contributions of this work add considerably to our knowledge of depression detection systems and have the potential to improve future systems by …
Speech Perception In “Bubble” Noise: Korean Fricatives And Affricates By Native And Non-Native Korean Listeners, Jiyoung Choi
Speech Perception In “Bubble” Noise: Korean Fricatives And Affricates By Native And Non-Native Korean Listeners, Jiyoung Choi
Dissertations, Theses, and Capstone Projects
The current study examines acoustic cues used by second language learners of Korean to discriminate between Korean fricatives and affricates in noise and how these cues relate to those used by native Korean listeners. Stimuli consist of naturally-spoken consonant-vowel-consonant-vowel (CVCV) syllables: /sɑdɑ/, /s*ɑdɑ/, /tʃɑdɑ/, /tʃhɑdɑ/, and /tʃ*ɑdɑ/. In this experiment, the “bubble noise” methodology of Mandel at al. (2016) was used to identify the time-frequency locations of important cues in each utterance, i.e., where audibility of the location is significantly correlated with correct identification of the utterance in noise. Results show that non-native Korean listeners can discriminate between …
Describing Doggo-Speak: Features Of Doggo Meme Language, Jennifer Bivens
Describing Doggo-Speak: Features Of Doggo Meme Language, Jennifer Bivens
Dissertations, Theses, and Capstone Projects
Doggo-speak is a specialized way of writing most commonly associated with captions on Doggo memes, humorous images of dogs shared in online communities. This paper will explore linguistic features of Doggo-speak through analysis of social media posts by Doggo fan pages. It will use the discussed features as inputs to five machine learning classifiers and will show, through this classification task, that the discussed features are sufficient for distinguishing between Doggo-speak and more general English text.
Intergroup Variability In Personality Recognition, Arundhati Sengupta
Intergroup Variability In Personality Recognition, Arundhati Sengupta
Dissertations, Theses, and Capstone Projects
Automatic Identification of personality in conversational speech has many applications in natural language processing such as leader identification in a meeting, adaptive dialogue systems, and dating websites. However, the widespread acceptance of automatic personality recognition through lexical and vocal characteristics is limited by the variability of error rate in a general purpose model among speakers from different demographic groups. While other work reports accuracy, we explored error rates of automatic personality recognition task using classification models for different genders and native language groups (L1). We also present a statistical experiment showing the influence of gender and L1 on the relation …
A Sentiment Analysis Of Language & Gender Using Word Embedding Models, Ellyn Rolleston Keith
A Sentiment Analysis Of Language & Gender Using Word Embedding Models, Ellyn Rolleston Keith
Dissertations, Theses, and Capstone Projects
Since Robin Lakoff started the conversation around language and gender with her 1975 essay “Language and Woman’s Place,” extensive work has been done on analyzing sociolinguistics associated with gender. While much work has been done on the differences between how men and women use language, there is less research to be found on language about women as opposed to language about men. In this work, I build a word embedding model from a corpus of Wikipedia film summaries and use this model to create lists of words associated with men and words associated with women. I then use sentiment analysis …
Es-Esa: An Information Retrieval Prototype Using Explicit Semantic Analysis And Elasticsearch, Brian D. Sloan
Es-Esa: An Information Retrieval Prototype Using Explicit Semantic Analysis And Elasticsearch, Brian D. Sloan
Dissertations, Theses, and Capstone Projects
Many modern information retrieval systems work by using keyword search to locate documents in an inverted index by matching those documents based on terms in a user’s query. While highly effective for many use-cases, one notable drawback to simple keyword-based searching is that the contextual knowledge surrounding the user’s underlying information need may be lost, particularly if the user’s query terms are ambiguous or have multiple meanings. Research in the field of semantic search aims to make progress towards resolving this. One methodology in particular, explicit semantic analysis, works by modeling a document not only as a set of …
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
Dissertations, Theses, and Capstone Projects
In recent years, Twitter has become a popular testing ground for techniques in authorship attribution. This is due to both the ease of building large corpora as well as the challenges associated with the character limit imposed by the service and the writing styles that have developed as a result. As both false and genuine claims of hacked Twitter accounts have made international news, there is an increasing need for this type of work. For newer Twitter accounts, however, there is little training data. Thus, this study looks to lay the groundwork for cross-domain authorship attribution: training on one source …
An Evaluation Of Pos Taggers For The Childes Corpus, Rui Huang
An Evaluation Of Pos Taggers For The Childes Corpus, Rui Huang
Dissertations, Theses, and Capstone Projects
This project evaluates four mainstream taggers on a representative collection of child-adult’s dialogues from Child Language Data Exchange System. The nine children’s files from Valian corpora and part of Eve corpora have been manually labeled, and rewrote with LARC tagset. They served as gold standard corpora in the training and testing process. Four taggers: CLAN MOR tagger, ACOPOST trigram tagger, Stanford parser, and Ver. 1.14 of Brill tagger have been tested by 10-fold cross validation. By analyzing what kinds of assumptions the tagger made about category assignment lead to failing, we identify several problematic cases of tagging. By comparing the …
Event Parsing In Narrative: Trials And Tribulations Of Archaic English Fairy Tales, Rebecca Lovering
Event Parsing In Narrative: Trials And Tribulations Of Archaic English Fairy Tales, Rebecca Lovering
Dissertations, Theses, and Capstone Projects
While event extraction and automatic summarization have taken great strides in the realm of news stories, fictional narratives like fairy tales have not been so fortunate. A number of challenges arise from the literary elements present in fairy tales that are not found in more straightforward corpora of natural language, such as archaic expressions and sentence structures. To aid in summarization of fictional texts, I created an class - a template for a digital object, in this case a semantic and story event - that captures elements predicted to help classify events as important for inclusion. I wrote a processor …
Nondescript: A Web Tool To Aid Subversion Of Authorship Attribution, Robin Davis
Nondescript: A Web Tool To Aid Subversion Of Authorship Attribution, Robin Davis
Dissertations, Theses, and Capstone Projects
A person’s writing style is uniquely quantifiable and can serve reliably as a biometric. A writer who wishes to remain anonymous can use a number of privacy technologies but can still be identified simply by the words they choose to use — how frequently they use common words like “of,” for instance. Nondescript is a web tool designed first to identify the user’s writing style in terms of word frequency from a given writing sample and document, then to suggest how the author can change their document to lessen its probability of being attributed to them. While Nondescript does not …
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Dissertations, Theses, and Capstone Projects
The process of producing written text is complex and constrained by pressures that range from physical to psychological. In a series of three sets of experiments, this thesis demonstrates the effects of linguistic context on the timing patterns of the production of keystrokes. We elucidate the effect of linguistic context at three different levels of granularity: The first set of experiments illustrate how the nontraditional syntax of a single linguistic construct, the multi-word expression, can create significant changes in keystroke production patterns. This set of experiments is followed by a set of experiments that test the hypothesis on the entire …
Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri
Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri
Dissertations, Theses, and Capstone Projects
Technology to automatically synthesize linguistically accurate and natural-looking animations of American Sign Language (ASL) would make it easier to add ASL content to websites and media, thereby increasing information accessibility for many people who are deaf and have low English literacy skills. State-of-art sign language animation tools focus mostly on accuracy of manual signs rather than on the facial expressions. We are investigating the synthesis of syntactic ASL facial expressions, which are grammatically required and essential to the meaning of sentences. In this thesis, we propose to: (1) explore the methodological aspects of evaluating sign language animations with facial expressions, …