Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Arts and Humanities (7)
- Computer Sciences (3)
- Phonetics and Phonology (3)
- Physical Sciences and Mathematics (3)
- American Sign Language (2)
-
- Discourse and Text Linguistics (2)
- Language Description and Documentation (2)
- Library and Information Science (2)
- Morphology (2)
- Other Linguistics (2)
- Science and Technology Studies (2)
- Sign Languages (2)
- Slavic Languages and Societies (2)
- Artificial Intelligence and Robotics (1)
- Business (1)
- Business Administration, Management, and Operations (1)
- Business and Corporate Communications (1)
- Cataloging and Metadata (1)
- Digital Humanities (1)
- Feminist, Gender, and Sexuality Studies (1)
- First and Second Language Acquisition (1)
- Lesbian, Gay, Bisexual, and Transgender Studies (1)
- Management Sciences and Quantitative Methods (1)
- Public Affairs, Public Policy and Public Administration (1)
- Russian Linguistics (1)
- Scholarly Communication (1)
- Science and Technology Policy (1)
- Institution
- Keyword
-
- Computational linguistics (3)
- Machine learning (2)
- Natural Language Processing (2)
- Natural language processing (2)
- Analysts’ Evaluative Bias (1)
-
- Arabic ETDs (1)
- Arbitrary Cases (1)
- Artificial intelligence (1)
- Authorship identification (1)
- CEO Non-Prototypicality (1)
- Classification (1)
- Competitive Aggressiveness (1)
- Compositions (1)
- Computational Linguistics (1)
- Computational humanities (1)
- Computational social science (1)
- Computer Assisted Language Learning (CALL) (1)
- Computer mediated communication (1)
- Computerized tools (1)
- Cultural analytics (1)
- Deep learning (1)
- Dependency Parsing (1)
- Dictionary-based keyword extraction (1)
- Digital Libraries (1)
- Doc2Vec (1)
- Economic geography (1)
- Ethical AI (1)
- ExtIPA (1)
- Finite-state transducer (1)
- Foreign Language Education (1)
- Publication
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Computational Linguistics
Computational Methods For Comparative Analyses Of Discourses, Zachary Stine
Computational Methods For Comparative Analyses Of Discourses, Zachary Stine
Theses and Dissertations
An underexplored aspect of the abundant linguistic data available for computational analysis is the potential for conducting large-scale comparative analyses between different discourses as a more rigorous complement to existing methods. A significant obstacle in comparative studies is the requirement that a researcher distinguish between discursive distinctions that are superficial and those that reflect deeper, structural relationships between the discourses being compared. In this dissertation, I describe two methodologies for comparing discourses in terms of their underlying structures, drawing on distributional semantic models and information theory. These methodologies are explored within three case studies in which the discourses of several …
Exploring The Personality Of Virtual Tutors In Conversational Foreign Language Practice, Johanna Dobbriner, Cathy Ennis, Robert J. Ross
Exploring The Personality Of Virtual Tutors In Conversational Foreign Language Practice, Johanna Dobbriner, Cathy Ennis, Robert J. Ross
Conference papers
Fluid interaction between virtual agents and humans requires the understanding of many issues of conversational pragmatics. One such issue is the interaction between communication strategy and personality. As a step towards developing models of personality driven pragmatics policies, in this paper, we present our initial experiment to explore differences in user interaction with two contrasting avatar personalities. Each user saw a single personality in a video-call setting and gave feedback on the interaction. Our expectations, that a more extroverted outgoing positive personality would be a more successful tutor, were only partially confirmed. While this personality did induce longer conversations in …
From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur
From An Art To A Science: Features And Methodology In Computational Authorship Identification, Jonathan I. Manczur
Dissertations, Theses, and Capstone Projects
Nearly thirty years ago, the United States Supreme Court revaluated the criteria for accepting forensic science and expert testimony, challenging Forensic Linguistics to assert itself as a reputable science. Much work has been produced in the interim to that end, but much still needs to be accomplished to satisfy the judicial standards. Computational linguistics has the potential to provide that necessary analytical framework. This paper’s intent is two-fold. First, there are two competing theories on the proper features necessary to identify an unknown author. Four features were drawn from the syntactic computational linguistics tradition and four from computational stylometry to …
Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale
Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale
Dissertations, Theses, and Capstone Projects
This dissertation presents the first implementation of label imputation for the task of homograph disambiguation using 1) transcribed audio, and 2) parallel, or translated, corpora. For label imputation from parallel corpora, a hypothesis of interlingual alignment between homograph pronunciations and text word forms is developed and formalized. Both audio and parallel corpora label imputation techniques are tested empirically in experiments that compare homograph disambiguation model performance using: 1) hand-labeled training data, and 2) hand-labeled training data augmented with label-imputed data. Regularized, multinomial logistic regression and pre-trained ALBERT, BERT, and XLNet language models fine-tuned as token classifiers are developed for homograph …
Detection And Morphological Analysis Of Novel Russian Loanwords, Yulia Spektor
Detection And Morphological Analysis Of Novel Russian Loanwords, Yulia Spektor
Dissertations, Theses, and Capstone Projects
This paper investigates recent English loanwords in Russian and explores ways in which computational methods can help further theoretical research. The goal of the study is two-fold: to find new, previously unattested loanwords borrowed over the last decade and to examine the rate of adaptation of the new borrowings, attested by the degree to which they conform to the constraints of the Russian language. First, we train a finite-state pipeline that combines character n-gram language models, which encode phonotactic and lexical properties of loanwords, with a binary classifier to detect loanwords. The model achieves state-of-the-art performance results during evaluation, surpassing …
Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby
Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby
Dissertations, Theses, and Capstone Projects
In a quick search online, one can find many tools which use information from news headlines to make predictions concerning the trajectory of a given stock. But what if we went further, looking instead into the text of the article, to extract this and other information? Here, the goal is to extract the sentence in which a stock ticker symbol is mentioned from a news article, then determine sentiment and subjectivity values from that sentence, and finally make a prediction on whether or not the value of that stock will go up or not in a 24-hour timespan. Bloomberg News …
The Public Innovations Explorer: A Geo-Spatial & Linked-Data Visualization Platform For Publicly Funded Innovation Research In The United States, Seth Schimmel
Dissertations, Theses, and Capstone Projects
The Public Innovations Explorer (https://sethsch.github.io/innovations-explorer/app/index.html) is a web-based tool created using Node.js, D3.js and Leaflet.js that can be used for investigating awards made by Federal agencies and departments participating in the Small Business Innovation Research (SBIR) and Small Business Technology Transfer (STTR) grant-making programs between 2008 and 2018. By geocoding the publicly available grants data from SBIR.gov, the Public Innovations Explorer allows users to identify companies performing publicly-funded innovative research in each congressional district and obtain dynamic district-level summaries of funding activity by agency and year. Applying spatial clustering techniques on districts' employment levels across major economic sectors provides users …
Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos
Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos
Electronic Theses and Dissertations
This study investigates one of the Polish language’s most arbitrary cases: the genitive masculine inanimate singular. It collects and ranks several guidelines to help language learners discern its proper usage and also introduces a framework to provide detailed feedback regarding arbitrary cases. The study tests this framework by implementing and evaluating a hybrid grammar checker called PLPrepare. PLPrepare performs similarly to other grammar checkers and is able to detect genitive case usages and provide feedback based on a number of error classifications.
Introducing Señal, A Computational Tool For The Linguistic Analysis Of Spanish L2 Compositions, Falcon Restrepo-Ramos
Introducing Señal, A Computational Tool For The Linguistic Analysis Of Spanish L2 Compositions, Falcon Restrepo-Ramos
World Languages & Cultures Department Publications
SEÑAL is a modular program that automatizes and facilitates the assessment of Spanish L2 compositions. The tool can extract syntactic and lexical information, while also assessing grammar, and Spanish L2 proficiency. A computational tool that can process written learners’ corpora and extract measures of language development has enormous practical value in Spanish and modern language departments alike.
An Interactive Visual Database For American Sign Language Reveals How Signs Are Organized In The Mind, Zed Sevcikova Sehyr, Ariel Goldberg, Karen Emmory, Naomi Caselli
An Interactive Visual Database For American Sign Language Reveals How Signs Are Organized In The Mind, Zed Sevcikova Sehyr, Ariel Goldberg, Karen Emmory, Naomi Caselli
Communication Sciences and Disorders Faculty Articles and Research
"We are four researchers who study psycholinguistics, linguistics, neuroscience and deaf education. Our team of deaf and hearing scientists worked with a group of software engineers to create the ASL-LEX database that anyone can use for free. We cataloged information on nearly 3,000 signs and built a visual, searchable and interactive database that allows scientists and linguists to work with ASL in entirely new ways."
Otrouha: A Corpus Of Arabic Etds And A Framework For Automatic Subject Classification, Eman Abdelrahman, Fatimah Alotaibi, Edward A. Fox, Osman Balci
Otrouha: A Corpus Of Arabic Etds And A Framework For Automatic Subject Classification, Eman Abdelrahman, Fatimah Alotaibi, Edward A. Fox, Osman Balci
The Journal of Electronic Theses and Dissertations
Although the Arabic language is spoken by more than 300 million people and is one of the six official languages of the United Nations (UN), there has been less research done on Arabic text data (compared to English) in the realm of machine learning, especially in text classification. In the past decade, Arabic data such as news, tweets, etc. have begun to receive some attention. Although automatic text classification plays an important role in improving the browsability and accessibility of data, Electronic Theses and Dissertations (ETDs) have not received their fair share of attention, in spite of the huge number …
An Automated Tool For Comparing Phonetic Transcriptions, Dallin J. Bailey, Marisha Speights Atkins, Ishaan Mishra, Sicheng Li, Yaoxuan Luan, Cheryl Seals
An Automated Tool For Comparing Phonetic Transcriptions, Dallin J. Bailey, Marisha Speights Atkins, Ishaan Mishra, Sicheng Li, Yaoxuan Luan, Cheryl Seals
Faculty Publications
Many computerized tools for comparing phonetic transcriptions have been proposed and shared in the past; however, previous tools are relatively difficult to access and incorporate into clinical and research practice, or require users to learn additional phonetic symbol systems. The purpose of this project was to develop and test a readily available web-based application for quantitatively comparing phonetic transcriptions that are input using International Phonetic Alphabet (IPA) symbols. A web-based computer application was developed to allow for IPA phonetic transcription comparison. A point-and-click keyboard was developed to provide support for character input of the full IPA, as well as most …
The Asl-Lex 2.0 Project: A Database Of Lexical And Phonological Properties For 2,723 Signs In American Sign Language, Zed Sevcikova Sehyr, Naomi Caselli, Ariel M. Cohen-Goldberg, Karen Emmory
The Asl-Lex 2.0 Project: A Database Of Lexical And Phonological Properties For 2,723 Signs In American Sign Language, Zed Sevcikova Sehyr, Naomi Caselli, Ariel M. Cohen-Goldberg, Karen Emmory
Communication Sciences and Disorders Faculty Articles and Research
ASL-LEX is a publicly available, large-scale lexical database for American Sign Language (ASL). We report on the expanded database (ASL-LEX 2.0) that contains 2,723 ASL signs. For each sign, ASL-LEX now includes a more detailed phonological description, phonological density and complexity measures, frequency ratings (from deaf signers), iconicity ratings (from hearing non-signers and deaf signers), transparency (“guessability”) ratings (from non-signers), sign and videoclip durations, lexical class, and more. We document the steps used to create ASL-LEX 2.0 and describe the distributional characteristics for sign properties across the lexicon and examine the relationships among lexical and phonological properties of signs. Correlation …
A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco
A Computational Study In The Detection Of English–Spanish Code-Switches, Yohamy C. Polanco
Dissertations, Theses, and Capstone Projects
Code-switching is the linguistic phenomenon where a multilingual person alternates between two or more languages in a conversation, whether that be spoken or written. This thesis studies the automatic detection of code-switching occurring specifically between English and Spanish in two corpora.
Twitter and other social media sites have provided an abundance of linguistic data that is available to researchers to perform countless experiments. Collecting the data is fairly easy if a study is on monolingual text, but if a study requires code-switched data, this becomes a complication as APIs only accept one language as a parameter. This thesis focuses on …
When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller
When Misclassification Is Misgendering: Gender Prediction In The Context Of Trans Identities, Sean Miller
Dissertations, Theses, and Capstone Projects
As a subdomain of author profiling, gender prediction (sometimes called gender inference) has received a substantial amount of attention—both as a task in itself, and for other downstream analyses. Throughout the existing literature various statistical and machine learning methods have been applied to extract features in order to either characterize and differentiate female and male writing styles, or simply to achieve maximum accuracy on gender prediction as a binary classification task. However, researchers often do not disclose how they conceptualize gender nor do they consider the implications that gender prediction has for non-binary and trans individuals. Along with an overview …
Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn
Sigmorphon 2021 Shared Task On Morphological Reinflection: Generalization Across Languages, Tiago Pimentel, Maria Ryskina, Christopher Straughn
Library Faculty Publications
This year’s iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems’ predictions. Transformer-based …
What You Do Or What You Say? An Examination Of Analyst Reactions To Prototypical And Non-Prototypical Ceos Linguistic And Competitive Behaviors, Courtney Hart
Theses and Dissertations--Management
Non-prototypical CEOs are those that process different demographic characteristics from a target reference group. In the US, a non-prototypical CEO is both white and male. While the negative responses to non-prototypical leaders based on race and gender have been well documented, we know less on what these leaders do that may influence biased evaluations. In this dissertation I took an impression management view to examine analysts’ evaluative bias (AEB) on prototypical and non-prototypical CEOs hiding linguistic behaviors and competitive aggressiveness. Specifically, I examined hiding linguistic behaviors on quarterly conference calls and two attributes of competitive repertoire will be researched. Drawing …
Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman
Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
We present the results and the main findings of the NLP4IF-2021 shared tasks. Task 1 focused on fighting the COVID-19 infodemic in social media, and it was offered in Arabic, Bulgarian, and English. Given a tweet, it asked to predict whether that tweet contains a verifiable claim, and if so, whether it is likely to be false, is of general interest, is likely to be harmful, and is worthy of manual fact-checking; also, whether it is harmful to society, and whether it requires the attention of policy makers. Task 2 focused on censorship detection, and was offered in Chinese. A …