Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

2021

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 18 of 18

Full-Text Articles in Computational Linguistics

Computational Methods For Comparative Analyses Of Discourses, Zachary Stine Dec 2021

Computational Methods For Comparative Analyses Of Discourses, Zachary Stine

Theses and Dissertations

An underexplored aspect of the abundant linguistic data available for computational analysis is the potential for conducting large-scale comparative analyses between different discourses as a more rigorous complement to existing methods. A significant obstacle in comparative studies is the requirement that a researcher distinguish between discursive distinctions that are superficial and those that reflect deeper, structural relationships between the discourses being compared. In this dissertation, I describe two methodologies for comparing discourses in terms of their underlying structures, drawing on distributional semantic models and information theory. These methodologies are explored within three case studies in which the discourses of several …


Exploring The Personality Of Virtual Tutors In Conversational Foreign Language Practice, Johanna Dobbriner, Cathy Ennis, Robert J. Ross Sep 2021

Exploring The Personality Of Virtual Tutors In Conversational Foreign Language Practice, Johanna Dobbriner, Cathy Ennis, Robert J. Ross

Conference papers

Fluid interaction between virtual agents and humans requires the understanding of many issues of conversational pragmatics. One such issue is the interaction between communication strategy and personality. As a step towards developing models of personality driven pragmatics policies, in this paper, we present our initial experiment to explore differences in user interaction with two contrasting avatar personalities. Each user saw a single personality in a video-call setting and gave feedback on the interaction. Our expectations, that a more extroverted outgoing positive personality would be a more successful tutor, were only partially confirmed. While this personality did induce longer conversations in …


From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur Sep 2021

From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur

Dissertations, Theses, and Capstone Projects

Nearly thirty years ago, the United States Supreme Court revaluated the criteria for accepting forensic science and expert testimony, challenging Forensic Linguistics to assert itself as a reputable science. Much work has been produced in the interim to that end, but much still needs to be accomplished to satisfy the judicial standards. Computational linguistics has the potential to provide that necessary analytical framework. This paper’s intent is two-fold. First, there are two competing theories on the proper features necessary to identify an unknown author. Four features were drawn from the syntactic computational linguistics tradition and four from computational stylometry to …


Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale Sep 2021

Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale

Dissertations, Theses, and Capstone Projects

This dissertation presents the first implementation of label imputation for the task of homograph disambiguation using 1) transcribed audio, and 2) parallel, or translated, corpora. For label imputation from parallel corpora, a hypothesis of interlingual alignment between homograph pronunciations and text word forms is developed and formalized. Both audio and parallel corpora label imputation techniques are tested empirically in experiments that compare homograph disambiguation model performance using: 1) hand-labeled training data, and 2) hand-labeled training data augmented with label-imputed data. Regularized, multinomial logistic regression and pre-trained ALBERT, BERT, and XLNet language models fine-tuned as token classifiers are developed for homograph …


Detection And Morphological Analysis Of Novel Russian Loanwords, Yulia Spektor Sep 2021

Detection And Morphological Analysis Of Novel Russian Loanwords, Yulia Spektor

Dissertations, Theses, and Capstone Projects

This paper investigates recent English loanwords in Russian and explores ways in which computational methods can help further theoretical research. The goal of the study is two-fold: to find new, previously unattested loanwords borrowed over the last decade and to examine the rate of adaptation of the new borrowings, attested by the degree to which they conform to the constraints of the Russian language. First, we train a finite-state pipeline that combines character n-gram language models, which encode phonotactic and lexical properties of loanwords, with a binary classifier to detect loanwords. The model achieves state-of-the-art performance results during evaluation, surpassing …


Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby Jun 2021

Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby

Dissertations, Theses, and Capstone Projects

In a quick search online, one can find many tools which use information from news headlines to make predictions concerning the trajectory of a given stock. But what if we went further, looking instead into the text of the article, to extract this and other information? Here, the goal is to extract the sentence in which a stock ticker symbol is mentioned from a news article, then determine sentiment and subjectivity values from that sentence, and finally make a prediction on whether or not the value of that stock will go up or not in a 24-hour timespan. Bloomberg News …


The Public Innovations Explorer: A Geo-Spatial & Linked-Data Visualization Platform For Publicly Funded Innovation Research In The United States, Seth Schimmel Jun 2021

The Public Innovations Explorer: A Geo-Spatial & Linked-Data Visualization Platform For Publicly Funded Innovation Research In The United States, Seth Schimmel

Dissertations, Theses, and Capstone Projects

The Public Innovations Explorer (https://sethsch.github.io/innovations-explorer/app/index.html) is a web-based tool created using Node.js, D3.js and Leaflet.js that can be used for investigating awards made by Federal agencies and departments participating in the Small Business Innovation Research (SBIR) and Small Business Technology Transfer (STTR) grant-making programs between 2008 and 2018. By geocoding the publicly available grants data from SBIR.gov, the Public Innovations Explorer allows users to identify companies performing publicly-funded innovative research in each congressional district and obtain dynamic district-level summaries of funding activity by agency and year. Applying spatial clustering techniques on districts' employment levels across major economic sectors provides users …


Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos May 2021

Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos

Electronic Theses and Dissertations

This study investigates one of the Polish language’s most arbitrary cases: the genitive masculine inanimate singular. It collects and ranks several guidelines to help language learners discern its proper usage and also introduces a framework to provide detailed feedback regarding arbitrary cases. The study tests this framework by implementing and evaluating a hybrid grammar checker called PLPrepare. PLPrepare performs similarly to other grammar checkers and is able to detect genitive case usages and provide feedback based on a number of error classifications.


Introducing Señal, A Computational Tool For The Linguistic Analysis Of Spanish L2 Compositions, Falcon Restrepo-Ramos Apr 2021

Introducing Señal, A Computational Tool For The Linguistic Analysis Of Spanish L2 Compositions, Falcon Restrepo-Ramos

World Languages & Cultures Department Publications

SEÑAL is a modular program that automatizes and facilitates the assessment of Spanish L2 compositions. The tool can extract syntactic and lexical information, while also assessing grammar, and Spanish L2 proficiency. A computational tool that can process written learners’ corpora and extract measures of language development has enormous practical value in Spanish and modern language departments alike.


An Interactive Visual Database For American Sign Language Reveals How Signs Are Organized In The Mind, Zed Sevcikova Sehyr, Ariel Goldberg, Karen Emmory, Naomi Caselli Apr 2021

An Interactive Visual Database For American Sign Language Reveals How Signs Are Organized In The Mind, Zed Sevcikova Sehyr, Ariel Goldberg, Karen Emmory, Naomi Caselli

Communication Sciences and Disorders Faculty Articles and Research

"We are four researchers who study psycholinguistics, linguistics, neuroscience and deaf education. Our team of deaf and hearing scientists worked with a group of software engineers to create the ASL-LEX database that anyone can use for free. We cataloged information on nearly 3,000 signs and built a visual, searchable and interactive database that allows scientists and linguists to work with ASL in entirely new ways."


Otrouha: A Corpus Of Arabic Etds And A Framework For Automatic Subject Classification, Eman Abdelrahman, Fatimah Alotaibi, Edward A. Fox, Osman Balci Mar 2021

Otrouha: A Corpus Of Arabic Etds And A Framework For Automatic Subject Classification, Eman Abdelrahman, Fatimah Alotaibi, Edward A. Fox, Osman Balci

The Journal of Electronic Theses and Dissertations

Although the Arabic language is spoken by more than 300 million people and is one of the six official languages of the United Nations (UN), there has been less research done on Arabic text data (compared to English) in the realm of machine learning, especially in text classification. In the past decade, Arabic data such as news, tweets, etc. have begun to receive some attention. Although automatic text classification plays an important role in improving the browsability and accessibility of data, Electronic Theses and Dissertations (ETDs) have not received their fair share of attention, in spite of the huge number …


An Automated Tool For Comparing Phonetic Transcriptions, Dallin J. Bailey, Marisha Speights Atkins, Ishaan Mishra, Sicheng Li, Yaoxuan Luan, Cheryl Seals Mar 2021

An Automated Tool For Comparing Phonetic Transcriptions, Dallin J. Bailey, Marisha Speights Atkins, Ishaan Mishra, Sicheng Li, Yaoxuan Luan, Cheryl Seals

Faculty Publications

Many computerized tools for comparing phonetic transcriptions have been proposed and shared in the past; however, previous tools are relatively difficult to access and incorporate into clinical and research practice, or require users to learn additional phonetic symbol systems. The purpose of this project was to develop and test a readily available web-based application for quantitatively comparing phonetic transcriptions that are input using International Phonetic Alphabet (IPA) symbols. A web-based computer application was developed to allow for IPA phonetic transcription comparison. A point-and-click keyboard was developed to provide support for character input of the full IPA, as well as most …


The Asl-Lex 2.0 Project: A Database Of Lexical And Phonological Properties For 2,723 Signs In American Sign Language, Zed Sevcikova Sehyr, Naomi Caselli, Ariel M. Cohen-Goldberg, Karen Emmory Feb 2021

The Asl-Lex 2.0 Project: A Database Of Lexical And Phonological Properties For 2,723 Signs In American Sign Language, Zed Sevcikova Sehyr, Naomi Caselli, Ariel M. Cohen-Goldberg, Karen Emmory

Communication Sciences and Disorders Faculty Articles and Research

ASL-LEX is a publicly available, large-scale lexical database for American Sign Language (ASL). We report on the expanded database (ASL-LEX 2.0) that contains 2,723 ASL signs. For each sign, ASL-LEX now includes a more detailed phonological description, phonological density and complexity measures, frequency ratings (from deaf signers), iconicity ratings (from hearing non-signers and deaf signers), transparency (“guessability”) ratings (from non-signers), sign and videoclip durations, lexical class, and more. We document the steps used to create ASL-LEX 2.0 and describe the distributional characteristics for sign properties across the lexicon and examine the relationships among lexical and phonological properties of signs. Correlation …


A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco Feb 2021

A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco

Dissertations, Theses, and Capstone Projects

Code-switching is the linguistic phenomenon where a multilingual person alternates between two or more languages in a conversation, whether that be spoken or written. This thesis studies the automatic detection of code-switching occurring specifically between English and Spanish in two corpora.

Twitter and other social media sites have provided an abundance of linguistic data that is available to researchers to perform countless experiments. Collecting the data is fairly easy if a study is on monolingual text, but if a study requires code-switched data, this becomes a complication as APIs only accept one language as a parameter. This thesis focuses on …


When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller Feb 2021

When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller

Dissertations, Theses, and Capstone Projects

As a subdomain of author profiling, gender prediction (sometimes called gender inference) has received a substantial amount of attention—both as a task in itself, and for other downstream analyses. Throughout the existing literature various statistical and machine learning methods have been applied to extract features in order to either characterize and differentiate female and male writing styles, or simply to achieve maximum accuracy on gender prediction as a binary classification task. However, researchers often do not disclose how they conceptualize gender nor do they consider the implications that gender prediction has for non-binary and trans individuals. Along with an overview …


Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn Jan 2021

Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn

Library Faculty Publications

This year’s iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems’ predictions. Transformer-based …


What You Do Or What You Say? An Examination Of Analyst Reactions To Prototypical And Non-Prototypical Ceos Linguistic And Competitive Behaviors, Courtney Hart Jan 2021

What You Do Or What You Say? An Examination Of Analyst Reactions To Prototypical And Non-Prototypical Ceos Linguistic And Competitive Behaviors, Courtney Hart

Theses and Dissertations--Management

Non-prototypical CEOs are those that process different demographic characteristics from a target reference group. In the US, a non-prototypical CEO is both white and male. While the negative responses to non-prototypical leaders based on race and gender have been well documented, we know less on what these leaders do that may influence biased evaluations. In this dissertation I took an impression management view to examine analysts’ evaluative bias (AEB) on prototypical and non-prototypical CEOs hiding linguistic behaviors and competitive aggressiveness. Specifically, I examined hiding linguistic behaviors on quarterly conference calls and two attributes of competitive repertoire will be researched. Drawing …


Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman Jan 2021

Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

We present the results and the main findings of the NLP4IF-2021 shared tasks. Task 1 focused on fighting the COVID-19 infodemic in social media, and it was offered in Arabic, Bulgarian, and English. Given a tweet, it asked to predict whether that tweet contains a verifiable claim, and if so, whether it is likely to be false, is of general interest, is likely to be harmful, and is worthy of manual fact-checking; also, whether it is harmful to society, and whether it requires the attention of policy makers. Task 2 focused on censorship detection, and was offered in Chinese. A …