Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

2022

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 20 of 20

Full-Text Articles in Computational Linguistics

Technology In The Classroom: The Features Language Teachers Should Consider, Sophie Cuocci, Padideh Fattahi Marnani Dec 2022

Technology In The Classroom: The Features Language Teachers Should Consider, Sophie Cuocci, Padideh Fattahi Marnani

Journal of English Learner Education

The fast development of technology and the new generation of highly computer literate students led to consider the integration of technology in school as essential. Throughout the last two decades, research has identified multiple factors leading to the successful and unsuccessful integration of technology in the classroom. Educators must consider these factors when deciding on which technology tools to use and how to integrate them to their lessons. Simultaneously, the increasing number of English learners in the United States calls for the identification of teaching strategies that will best support their needs. Many language teachers now rely on teaching techniques …


Creating Data From Unstructured Text With Context Rule Assisted Machine Learning (Craml), Stephen Meisenbacher, Peter Norlander Dec 2022

Creating Data From Unstructured Text With Context Rule Assisted Machine Learning (Craml), Stephen Meisenbacher, Peter Norlander

School of Business: Faculty Publications and Other Works

Popular approaches to building data from unstructured text come with limitations, such as scalability, interpretability, replicability, and real-world applicability. These can be overcome with Context Rule Assisted Machine Learning (CRAML), a method and no-code suite of software tools that builds structured, labeled datasets which are accurate and reproducible. CRAML enables domain experts to access uncommon constructs within a document corpus in a low-resource, transparent, and flexible manner. CRAML produces document-level datasets for quantitative research and makes qualitative classification schemes scalable over large volumes of text. We demonstrate that the method is useful for bibliographic analysis, transparent analysis of proprietary data, …


Is Neuro-Symbolic Ai Meeting Its Promises In Natural Language Processing? A Structured Review, Kyle Hamilton, Kyle Hamilton, Aparna Nayak, Bojan Bozic, Luca Longo Nov 2022

Is Neuro-Symbolic Ai Meeting Its Promises In Natural Language Processing? A Structured Review, Kyle Hamilton, Kyle Hamilton, Aparna Nayak, Bojan Bozic, Luca Longo

Articles

Advocates for Neuro-Symbolic Artificial Intelligence (NeSy) assert that combining deep learning with symbolic reasoning will lead to stronger AI than either paradigm on its own. As successful as deep learning has been, it is generally accepted that even our best deep learning systems are not very good at abstract reasoning. And since reasoning is inextricably linked to language, it makes intuitive sense that Natural Language Processing (NLP), would be a particularly well-suited candidate for NeSy. We conduct a structured review of studies implementing NeSy for NLP, with the aim of answering the question of whether NeSy is indeed meeting its …


Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell Oct 2022

Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell

Data and Test Instruments

This paper outlines the complexity of the psychological construct of individuals' subjective well-being (SWB) and argues for the importance of examining behaviours and linguistic expression of individuals online social interactions in relation to self-reported SWB. This paper calls for a systematic review of the psychology research which examines SWB and its association with various character strengths, personality traits, and behaviours. While the Big Five personality traits (OCEAN) have an underlying neuropsychological basis and are considered as universal dimensions of personality along which humans differ one from another, minimal research has attempted to evaluate the relationship between personality traits, SWB, and …


Predicting Stress In Russian Using Modern Machine-Learning Tools, John Schriner Sep 2022

Predicting Stress In Russian Using Modern Machine-Learning Tools, John Schriner

Dissertations, Theses, and Capstone Projects

In the Russian language, stress on a word is determined via often complex patterns and rules. In this paper, after examining nearly a century of research in stress rules and methods in Russian, we turn to see if modern machine learning tools can aid in predicting stress. Using A.A. Zaliznyak’s dictionary grammar and over 300,000 word forms, we derived stress codes to aid in predicting which syllable primary stress falls on. We trained an LSTM neural network on the data and conducted eight experiments with added features such as lemma, part of speech, and morphology. While the model performed better …


A Study Of Entrainment In Speech As A Possible Predictor Of Perceived Trust, Mariana Graterol Fuenmayor Sep 2022

A Study Of Entrainment In Speech As A Possible Predictor Of Perceived Trust, Mariana Graterol Fuenmayor

Dissertations, Theses, and Capstone Projects

This thesis explores the possibility of using features of speech as possible predictors of perceived trust. It specifically focuses on entrainment, the tendency of participants in a conversation to unconsciously adapt their manner of talking to become more or less similar to each other. We compile a corpus of interviews conducted in English, with different levels of formality and discussing different topics. With it, we distribute a set of surveys to assess raters’ judgment of the participants, focusing on whether they believe the interviewee is trusted by their conversational partner, and whether they find the interlocutors themselves trustworthy. We also …


Towards Explaining Variation In Entrainment, Andreas Weise Sep 2022

Towards Explaining Variation In Entrainment, Andreas Weise

Dissertations, Theses, and Capstone Projects

Entrainment refers to the tendency of human speakers to adapt to their interlocutors to become more similar to them. This affects various dimensions and occurs in many contexts, allowing for rich applications in human-computer interaction. However, it is not exhibited by every speaker in every conversation but varies widely across features, speakers, and contexts, hindering broad application. This variation, whose guiding principles are poorly understood even after decades of entrainment research, is the subject of this thesis. We begin with a comprehensive literature review that serves as the foundation of our own work and provides a reference to guide future …


From Sesame Street To Beyond: Multi-Domain Discourse Relation Classification With Pretrained Bert, Isaac R. Raff Sep 2022

From Sesame Street To Beyond: Multi-Domain Discourse Relation Classification With Pretrained Bert, Isaac R. Raff

Dissertations, Theses, and Capstone Projects

Research efforts in transfer learning have gained massive popularity in recent years. Pretrained language models have demonstrated the most successful results in producing high quality neural networks capable of quality inference after training across domains via transfer learning. This study expands on the domain transfer introduced in \cite{ferracane-etal-2019-news} exploring neural methods for transfer learning of discourse parsing between a news source domain and a medical target domain. \cite{ferracane-etal-2019-news} specifically discuss transfer learning from news articles to PubMed medical journal articles. Experiments in transfer learning in the current work expand to include three domains: Wall Street Journal articles previously annotated with …


Linguistic Abstractions In Children’S Very Early Utterances, Qihui Xu Sep 2022

Linguistic Abstractions In Children’S Very Early Utterances, Qihui Xu

Dissertations, Theses, and Capstone Projects

How early do children produce multiword utterances? Do children's early utterances reflect abstract syntactic knowledge or are they the result of data-driven learning? We examine this issue through corpus analysis, computational modeling, and adult simulation experiments. Chapter 1 investigates when children start producing multiword utterances; we use corpora to establish the development of multiword utterances and a probabilistic computational model to account for the quantitative change of early multiword utterances. We find that multiword utterances of different lengths appear early in acquisition and increase together, and the length growth pattern can be viewed as a probabilistic and dynamic process.

Chapter …


Spectral Analysis Of Multiscale Cultural Traits On Twitter, Chandler Squires, Nikhil Kunapuli, Yaneer Bar-Yam, Alfredo Morales Aug 2022

Spectral Analysis Of Multiscale Cultural Traits On Twitter, Chandler Squires, Nikhil Kunapuli, Yaneer Bar-Yam, Alfredo Morales

Northeast Journal of Complex Systems (NEJCS)

Understanding and mapping the emergence and boundaries of cultural areas is a challenge for social sciences. In this paper, we present a method for analyzing the cultural composition of regions via Twitter hashtags. Cultures can be described as distinct combination of traits which we capture via principal component analysis (PCA). We investigate the top 8 PCA components of an area including France, Spain, and Portugal, in terms of the geographic distribution of their hashtag composition. We also discuss relationships between components and the insights those relationships can provide into the structure of a cultural space. Finally, we compare the spatial …


Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin Jun 2022

Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin

Data and Test Instruments

How do we detect, deter, and prevent the spread of mis- and disinformationwith the human eye and AI? How does theory inform the practice, and how do theevidence-based research and best practices in lie-catching and truth-seekingprofessions—inform AI? The book looks into well-established human practicessuch as the routines and processes used in detective work, journalism, and scientificinquiry, and how they contribute toward innovative AI solutions. The book explainsthe principles, inner workings, and recent evolution of five types of state-of-the-artAI technologies suitable for curtailing the spread of mis- and disinformation:automated deception detectors, clickbait detectors, satirical fake detectors, rumordebunkers, and computational fact-checking tools.


Generic Ab Initio, James A. Heilpern, Earl Kjar Brown, William G. Eggington, Zachary D. Smith Jun 2022

Generic Ab Initio, James A. Heilpern, Earl Kjar Brown, William G. Eggington, Zachary D. Smith

Buffalo Law Review

From comic conventions to disbanded dioceses, courts continue to struggle with a unique but puzzling question of trademark law. Federal law protects certain terms that refer to a product or service from a specific producer instead of to a product generally. Terms that refer to products are considered generic and cannot receive protection. Courts have also held that a term that was generic at the time the party adopted the mark cannot receive protection, even if the public later views it as being specific to a particular producer. But, many marks were adopted decades or centuries ago. As a result, …


A Machine Learning Approach To Text-Based Sarcasm Detection, Lara I. Novic Jun 2022

A Machine Learning Approach To Text-Based Sarcasm Detection, Lara I. Novic

Dissertations, Theses, and Capstone Projects

Sarcasm and indirect language are commonplace for humans to produce and recognize but difficult for machines to detect. While artificial intelligence can accurately analyze sentiment and emotion in speech and text, it may struggle with insincere and sardonic content, although it is possible to train a machine to identify uttered and written sarcasm. This paper aims to detect sarcasm using logistic regression and a support vector machine (SVM) and compare their results to a baseline.

The models are trained on headlines from a Kaggle dataset containing headlines from the satirical news website The Onion and serious news website Huffpost (formerly …


Covert Determiners In Appalachian English Narrative Declarative Sentences, William Oliver Jun 2022

Covert Determiners In Appalachian English Narrative Declarative Sentences, William Oliver

Dissertations, Theses, and Capstone Projects

In this thesis, I explore the syntax and semantics of covert determiners (Ds) in matrix subject determiner phrases (DPs) with definite specific interpretations. To conduct my investigation, I used the Audio-Aligned and Parsed Corpus of Appalachian English (AAPCAppE), a million-word Penn Treebank corpus, and the software CorpusSearch, a Java program that searches Penn Treebank corpora. My research shows that Appalachian English contains a linguistic phenomenon where speakers drop the D, replacing overt Ds with covert Ds, in definite specific DPs. For example, where Standard English speakers say The doctor came by horseback, Appalachian speakers may use a covert D …


Metaphor Detection In Poems In Misurata Arabic Sub-Dialect : An Lstm Model, Azza Abugharsa May 2022

Metaphor Detection In Poems In Misurata Arabic Sub-Dialect : An Lstm Model, Azza Abugharsa

Theses, Dissertations and Culminating Projects

Natural Language Processing (NLP) in Arabic is witnessing an increasing interest in investigating different topics in the field. One of the topics that have drawn attention is the automatic processing of Arabic figurative language. The focus in previous projects is on detecting and interpreting metaphors in comments from social media as well as phrases and/or headlines from news articles. The current project focuses on metaphor detection in poems written in the Misurata Arabic sub-dialect spoken in Misurata, located in the North African region. The dataset is initially annotated by a group of linguists, and their annotation is treated as the …


“I Can See The Forest For The Trees”: Examining Personality Traits With Trasformers, Alexander Moore May 2022

“I Can See The Forest For The Trees”: Examining Personality Traits With Trasformers, Alexander Moore

All Dissertations

Our understanding of Personality and its structure is rooted in linguistic studies operating under the assumptions made by the Lexical Hypothesis: personality characteristics that are important to a group of people will at some point be codified in their language, with the number of encoded representations of a personality characteristic indicating their importance. Qualitative and quantitative efforts in the dimension reduction of our lexicon throughout the mid-20th century have played a vital role in the field’s eventual arrival at the widely accepted Five Factor Model (FFM). However, there are a number of presently unresolved conflicts regarding the breadth and …


Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian Apr 2022

Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian

Northeast Journal of Complex Systems (NEJCS)

In this study, we introduce a new network feature for detecting suicidal ideation from clinical texts and conduct various additional experiments to enrich the state of knowledge. We evaluate statistical features with and without stopwords, use lexical networks for feature extraction and classification, and compare the results with standard machine learning methods using a logistic classifier, a neural network, and a deep learning method. We utilize three text collections. The first two contain transcriptions of interviews conducted by experts with suicidal (n=161 patients that experienced severe ideation) and control subjects (n=153). The third collection consists of interviews conducted by experts …


Searching For Pets: Using Distributional And Sentiment-Based Methods To Find Potentially Euphemistic Terms, Patrick Lee, Martha Gavidia, Anna Feldman, Jing Peng Jan 2022

Searching For Pets: Using Distributional And Sentiment-Based Methods To Find Potentially Euphemistic Terms, Patrick Lee, Martha Gavidia, Anna Feldman, Jing Peng

Department of Computer Science Faculty Scholarship and Creative Works

This paper presents a linguistically driven proof of concept for finding potentially euphemistic terms, or PETs. Acknowledging that PETs tend to be commonly used expressions for a certain range of sensitive topics, we make use of distributional similarities to select and filter phrase candidates from a sentence and rank them using a set of simple sentiment-based metrics. We present the results of our approach tested on a corpus of sentences containing euphemisms, demonstrating its efficacy for detecting single and multi-word PETs from a broad range of topics. We also discuss future potential for sentiment-based methods on this task.


A Report On The Euphemisms Detection Shared Task, Patrick Lee, Anna Feldman, Jing Peng Jan 2022

A Report On The Euphemisms Detection Shared Task, Patrick Lee, Anna Feldman, Jing Peng

Department of Computer Science Faculty Scholarship and Creative Works

This paper presents The Shared Task on Euphemism Detection for the Third Workshop on Figurative Language Processing (FigLang 2022) held in conjunction with EMNLP 2022. Participants were invited to investigate the euphemism detection task: given input text, identify whether it contains a euphemism. The input data is a corpus of sentences containing potentially euphemistic terms (PETs) collected from the GloWbE corpus (Davies and Fuchs, 2015), and are human-annotated as containing either a euphemistic or literal usage of a PET. In this paper, we present the results and analyze the common themes, methods and findings of the participating teams.


Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng Jan 2022

Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng

Department of Computer Science Faculty Scholarship and Creative Works

Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nonetheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, …