Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 61 - 90 of 250

Full-Text Articles in Computational Linguistics

Single-Case Pilot Study For Longitudinal Analysis Of Referential Failures And Sentiment In Schizophrenic Speech From Client-Centered Psychotherapy Recordings, Travis A. Musich Apr 2023

Single-Case Pilot Study For Longitudinal Analysis Of Referential Failures And Sentiment In Schizophrenic Speech From Client-Centered Psychotherapy Recordings, Travis A. Musich

Dissertations

Though computational linguistic analyses have revealed the presence of distinctly characteristic language features in schizophrenic disordered speech, the relative stability of these language features in longitudinal samples is still unknown. This longitudinal pilot study analyzed schizophrenic disordered speech data from the archival therapy audio recordings of one patient spanning 23 years. End-to-end Neural Coreference Resolution software was used to analyze transcribed speech data from three therapy sessions to identify ambiguous pronouns, referred to as referential failures, which were reviewed and confirmed by multiple raters. Speech samples were analyzed using Google Cloud Natural Language API software for sentiment variables (i.e., score, …


Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian) Mar 2023

Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)

Library Philosophy and Practice (e-journal)

Abstract

Purpose: The purpose of this research paper is to explore ChatGPT’s potential as an innovative designer tool for the future development of artificial intelligence. Specifically, this conceptual investigation aims to analyze ChatGPT’s capabilities as a tool for designing and developing near about human intelligent systems for futuristic used and developed in the field of Artificial Intelligence (AI). Also with the helps of this paper, researchers are analyzed the strengths and weaknesses of ChatGPT as a tool, and identify possible areas for improvement in its development and implementation. This investigation focused on the various features and functions of ChatGPT that …


Analysis Of The Inter-Annotators Agreement And Its Effect On The System Of Anti-Asian Hate Crime Detection On Twitter During Covid-19, Amir Toliyat Feb 2023

Analysis Of The Inter-Annotators Agreement And Its Effect On The System Of Anti-Asian Hate Crime Detection On Twitter During Covid-19, Amir Toliyat

Dissertations, Theses, and Capstone Projects

Coronavirus disease 2019 (COVID-19) started in Wuhan, China, in late 2019, and after being utterly contagious in Asian countries, it rapidly spread to other countries. This disease caused governments worldwide to declare a public health crisis with severe measures taken to reduce the speed of the spread of the disease. This pandemic affected the lives of millions of people. Many citizens that lost their loved ones and jobs experienced a wide range of emotions, such as disbelief, shock, concerns about health, fear about food supplies, anxiety, and panic. All of the aforementioned phenomena led to the spread of racism and …


A Sentiment Analysis Of "Filipinx" On Twitter Using A Multinomial Naïve Bayes Classification Model, Clarisse Taboy Feb 2023

A Sentiment Analysis Of "Filipinx" On Twitter Using A Multinomial Naïve Bayes Classification Model, Clarisse Taboy

Dissertations, Theses, and Capstone Projects

On social media, the use of “Filipinx” as a gender neutral, inclusive term for “Filipino” tends to generate high user engagement, at times without regard for the original context in which the word appears. This project applies computational methods to collect a large dataset in English/Filipino from Twitter containing “Filipinx”, and to train a Naïve Bayes model to classify tweets into three sentiments: positive, neutral, and negative. My methodology takes inspiration from that of four related studies that similarly conducted sentiment analysis on English/Filipino tweets involving various topics, and whose resulting accuracy scores were compared side-by-side. Conducting sentiment analysis on …


Simulating The Machine Translation Of Low-Resource Languages By Designing A Translator Between English And An Artificially Constructed Language, Michaela Snyder Jan 2023

Simulating The Machine Translation Of Low-Resource Languages By Designing A Translator Between English And An Artificially Constructed Language, Michaela Snyder

Mahurin Honors College Capstone Experience/Thesis Projects

Natural language processing (NLP), or the use of computers to analyze natural language, is a field that relies heavily on syntax. It would seem intuitive that computers would thrive in this area due to their strict syntax requirements, but the syntax of natural languages leaves them unable to properly parse and generate sentences that seem normal to the average speaker. A subfield of NLP, machine translation, works mainly to computerize translation between different languages. Unfortunately, such translation is not without its weaknesses; language documentation is not created equal, and many low-resource languages—languages with relatively few kinds of documentation, most often …


Context And Coherence, By Una Stojnic, Robert Stainton, Arthur Sullivan Jan 2023

Context And Coherence, By Una Stojnic, Robert Stainton, Arthur Sullivan

Philosophy Publications

No abstract provided.


Brazilian Portuguese-Russian (Braporus) Corpus: Automatic Transcription And Acoustic Quality Of Elderly Speech During Covid-19 Pandemic, Irina A. Sekerina, Anna Smirnova Henriques, Aleksandra Skorobogatova, Natalia Tyulina, Tatiana V. Kachkovskaia, Svetlana Ruseishvili, Sandra Madureira Jan 2023

Brazilian Portuguese-Russian (Braporus) Corpus: Automatic Transcription And Acoustic Quality Of Elderly Speech During Covid-19 Pandemic, Irina A. Sekerina, Anna Smirnova Henriques, Aleksandra Skorobogatova, Natalia Tyulina, Tatiana V. Kachkovskaia, Svetlana Ruseishvili, Sandra Madureira

Publications and Research

This article presents the Brazilian Portuguese-Russian (BraPoRus) corpus, whose goal is to collect, analyze, and preserve for posterity the spoken heritage Russian still used today in Brazil by approximately 1,500 elderly bilingual heritage Russian–Brazilian Portuguese speakers. Their unique 100-year-old variety of moribund Russian is disappearing because it has not been passed to their descendants born in Brazil. During the COVID-19 pandemic, we remotely collected 170 h of speech samples in heritage Russian from 26 participants (Mage = 75.7 years) in naturalistic settings using Zoom or a phone call. To estimate the quality of collected data, we focus on two methodological …


Automatic Transcription Of Northern Prinmi Oral Art: Approaches And Challenges To Automatic Speech Recognition For Language Documentation, Connor Bechler Jan 2023

Automatic Transcription Of Northern Prinmi Oral Art: Approaches And Challenges To Automatic Speech Recognition For Language Documentation, Connor Bechler

Theses and Dissertations--Linguistics

One significant issue facing language documentation efforts is the transcription bottleneck: each documented recording must be transcribed and annotated, and these tasks are extremely labor intensive (Ćavar et al., 2016). Researchers have sought to accelerate these tasks with partial automation via forced alignment, natural language processing, and automatic speech recognition (ASR) (Neubig et al., 2020). Neural network—especially transformer-based—approaches have enabled large advances in ASR over the last decade. Models like XLSR-53 promise improved performance on under-resourced languages by leveraging massive data sets from many different languages (Conneau et al., 2020). This project extends these efforts to a novel context, applying …


N-Gram Text Classification On Standard Croatian, Bosnian And Serbian, Kegan Messmer Jan 2023

N-Gram Text Classification On Standard Croatian, Bosnian And Serbian, Kegan Messmer

Honors Theses and Capstones

This study attempts to use three different kinds of n-gram text classification models to differentiate the standard forms of Croatian, Bosnian, and Serbian. These three languages, along with Montenegrin, were once considered one language, collectively termed “Serbo-Croatian”. These languages share a common South Slavic ancestry, and there is an argument to be made that their novel status as distinct languages is due to non-linguistic factors, such as culture and politics. This study uses 300,000 sentences from each language, sourced from Wikipedia pages written in the standard forms of each language. Three different classifiers were used: one unigram, one bigram, and …


‘A Category Of Their Own’: Quantitative Methods In The Use Of Pile-Sort Data In Perceptual Dialectology, Zachary Ty Gill Jan 2023

‘A Category Of Their Own’: Quantitative Methods In The Use Of Pile-Sort Data In Perceptual Dialectology, Zachary Ty Gill

Theses and Dissertations--Linguistics

The purpose of this study is to investigate how Mississippi Gulf Coast Creoles perceive language differences in their home area. A pile-sort task was carried out in which respondents were given stacks of cards with local communities written on them and instructed to stack together the regions where people “talk the same.” Once the piles were made, the fieldworker discussed their sortings with the respondents. The stacks were analyzed by means of a hierarchal agglomerative cluster analysis and non-parametric multidimensional scaling with k-means cluster analysis overlays to extract the perceived dialect areas. The groupings reveal that respondent strategies are based …


Evaluation Of Different Machine Learning, Deep Learning And Text Processing Techniques For Hate Speech Detection, Nabil Shawkat Jan 2023

Evaluation Of Different Machine Learning, Deep Learning And Text Processing Techniques For Hate Speech Detection, Nabil Shawkat

Graduate Theses/Dissertations

Social media has become a domain that involves a lot of hate speech. Some users feel entitled to engage in abusive conversations by sending abusive messages, tweets, or photos to other users. It is critical to detect hate speech and prevent innocent users from becoming victims. In this study, I explore the effectiveness and performance of various machine learning methods employing text processing techniques to create a robust system for hate speech identification. I assess the performance of Naïve Bayes, Support Vector Machines, Decision Trees, Random Forests, Logistic Regression, and K Nearest Neighbors using three distinct datasets sourced from social …


Technology In The Classroom: The Features Language Teachers Should Consider, Sophie Cuocci, Padideh Fattahi Marnani Dec 2022

Technology In The Classroom: The Features Language Teachers Should Consider, Sophie Cuocci, Padideh Fattahi Marnani

Journal of English Learner Education

The fast development of technology and the new generation of highly computer literate students led to consider the integration of technology in school as essential. Throughout the last two decades, research has identified multiple factors leading to the successful and unsuccessful integration of technology in the classroom. Educators must consider these factors when deciding on which technology tools to use and how to integrate them to their lessons. Simultaneously, the increasing number of English learners in the United States calls for the identification of teaching strategies that will best support their needs. Many language teachers now rely on teaching techniques …


Creating Data From Unstructured Text With Context Rule Assisted Machine Learning (Craml), Stephen Meisenbacher, Peter Norlander Dec 2022

Creating Data From Unstructured Text With Context Rule Assisted Machine Learning (Craml), Stephen Meisenbacher, Peter Norlander

School of Business: Faculty Publications and Other Works

Popular approaches to building data from unstructured text come with limitations, such as scalability, interpretability, replicability, and real-world applicability. These can be overcome with Context Rule Assisted Machine Learning (CRAML), a method and no-code suite of software tools that builds structured, labeled datasets which are accurate and reproducible. CRAML enables domain experts to access uncommon constructs within a document corpus in a low-resource, transparent, and flexible manner. CRAML produces document-level datasets for quantitative research and makes qualitative classification schemes scalable over large volumes of text. We demonstrate that the method is useful for bibliographic analysis, transparent analysis of proprietary data, …


Is Neuro-Symbolic Ai Meeting Its Promises In Natural Language Processing? A Structured Review, Kyle Hamilton, Kyle Hamilton, Aparna Nayak, Bojan Bozic, Luca Longo Nov 2022

Is Neuro-Symbolic Ai Meeting Its Promises In Natural Language Processing? A Structured Review, Kyle Hamilton, Kyle Hamilton, Aparna Nayak, Bojan Bozic, Luca Longo

Articles

Advocates for Neuro-Symbolic Artificial Intelligence (NeSy) assert that combining deep learning with symbolic reasoning will lead to stronger AI than either paradigm on its own. As successful as deep learning has been, it is generally accepted that even our best deep learning systems are not very good at abstract reasoning. And since reasoning is inextricably linked to language, it makes intuitive sense that Natural Language Processing (NLP), would be a particularly well-suited candidate for NeSy. We conduct a structured review of studies implementing NeSy for NLP, with the aim of answering the question of whether NeSy is indeed meeting its …


Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell Oct 2022

Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell

Data and Test Instruments

This paper outlines the complexity of the psychological construct of individuals' subjective well-being (SWB) and argues for the importance of examining behaviours and linguistic expression of individuals online social interactions in relation to self-reported SWB. This paper calls for a systematic review of the psychology research which examines SWB and its association with various character strengths, personality traits, and behaviours. While the Big Five personality traits (OCEAN) have an underlying neuropsychological basis and are considered as universal dimensions of personality along which humans differ one from another, minimal research has attempted to evaluate the relationship between personality traits, SWB, and …


Predicting Stress In Russian Using Modern Machine-Learning Tools, John Schriner Sep 2022

Predicting Stress In Russian Using Modern Machine-Learning Tools, John Schriner

Dissertations, Theses, and Capstone Projects

In the Russian language, stress on a word is determined via often complex patterns and rules. In this paper, after examining nearly a century of research in stress rules and methods in Russian, we turn to see if modern machine learning tools can aid in predicting stress. Using A.A. Zaliznyak’s dictionary grammar and over 300,000 word forms, we derived stress codes to aid in predicting which syllable primary stress falls on. We trained an LSTM neural network on the data and conducted eight experiments with added features such as lemma, part of speech, and morphology. While the model performed better …


A Study Of Entrainment In Speech As A Possible Predictor Of Perceived Trust, Mariana Graterol Fuenmayor Sep 2022

A Study Of Entrainment In Speech As A Possible Predictor Of Perceived Trust, Mariana Graterol Fuenmayor

Dissertations, Theses, and Capstone Projects

This thesis explores the possibility of using features of speech as possible predictors of perceived trust. It specifically focuses on entrainment, the tendency of participants in a conversation to unconsciously adapt their manner of talking to become more or less similar to each other. We compile a corpus of interviews conducted in English, with different levels of formality and discussing different topics. With it, we distribute a set of surveys to assess raters’ judgment of the participants, focusing on whether they believe the interviewee is trusted by their conversational partner, and whether they find the interlocutors themselves trustworthy. We also …


Towards Explaining Variation In Entrainment, Andreas Weise Sep 2022

Towards Explaining Variation In Entrainment, Andreas Weise

Dissertations, Theses, and Capstone Projects

Entrainment refers to the tendency of human speakers to adapt to their interlocutors to become more similar to them. This affects various dimensions and occurs in many contexts, allowing for rich applications in human-computer interaction. However, it is not exhibited by every speaker in every conversation but varies widely across features, speakers, and contexts, hindering broad application. This variation, whose guiding principles are poorly understood even after decades of entrainment research, is the subject of this thesis. We begin with a comprehensive literature review that serves as the foundation of our own work and provides a reference to guide future …


From Sesame Street To Beyond: Multi-Domain Discourse Relation Classification With Pretrained Bert, Isaac R. Raff Sep 2022

From Sesame Street To Beyond: Multi-Domain Discourse Relation Classification With Pretrained Bert, Isaac R. Raff

Dissertations, Theses, and Capstone Projects

Research efforts in transfer learning have gained massive popularity in recent years. Pretrained language models have demonstrated the most successful results in producing high quality neural networks capable of quality inference after training across domains via transfer learning. This study expands on the domain transfer introduced in \cite{ferracane-etal-2019-news} exploring neural methods for transfer learning of discourse parsing between a news source domain and a medical target domain. \cite{ferracane-etal-2019-news} specifically discuss transfer learning from news articles to PubMed medical journal articles. Experiments in transfer learning in the current work expand to include three domains: Wall Street Journal articles previously annotated with …


Linguistic Abstractions In Children’S Very Early Utterances, Qihui Xu Sep 2022

Linguistic Abstractions In Children’S Very Early Utterances, Qihui Xu

Dissertations, Theses, and Capstone Projects

How early do children produce multiword utterances? Do children's early utterances reflect abstract syntactic knowledge or are they the result of data-driven learning? We examine this issue through corpus analysis, computational modeling, and adult simulation experiments. Chapter 1 investigates when children start producing multiword utterances; we use corpora to establish the development of multiword utterances and a probabilistic computational model to account for the quantitative change of early multiword utterances. We find that multiword utterances of different lengths appear early in acquisition and increase together, and the length growth pattern can be viewed as a probabilistic and dynamic process.

Chapter …


Spectral Analysis Of Multiscale Cultural Traits On Twitter, Chandler Squires, Nikhil Kunapuli, Yaneer Bar-Yam, Alfredo Morales Aug 2022

Spectral Analysis Of Multiscale Cultural Traits On Twitter, Chandler Squires, Nikhil Kunapuli, Yaneer Bar-Yam, Alfredo Morales

Northeast Journal of Complex Systems (NEJCS)

Understanding and mapping the emergence and boundaries of cultural areas is a challenge for social sciences. In this paper, we present a method for analyzing the cultural composition of regions via Twitter hashtags. Cultures can be described as distinct combination of traits which we capture via principal component analysis (PCA). We investigate the top 8 PCA components of an area including France, Spain, and Portugal, in terms of the geographic distribution of their hashtag composition. We also discuss relationships between components and the insights those relationships can provide into the structure of a cultural space. Finally, we compare the spatial …


Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin Jun 2022

Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin

Data and Test Instruments

How do we detect, deter, and prevent the spread of mis- and disinformationwith the human eye and AI? How does theory inform the practice, and how do theevidence-based research and best practices in lie-catching and truth-seekingprofessions—inform AI? The book looks into well-established human practicessuch as the routines and processes used in detective work, journalism, and scientificinquiry, and how they contribute toward innovative AI solutions. The book explainsthe principles, inner workings, and recent evolution of five types of state-of-the-artAI technologies suitable for curtailing the spread of mis- and disinformation:automated deception detectors, clickbait detectors, satirical fake detectors, rumordebunkers, and computational fact-checking tools.


Generic Ab Initio, James A. Heilpern, Earl Kjar Brown, William G. Eggington, Zachary D. Smith Jun 2022

Generic Ab Initio, James A. Heilpern, Earl Kjar Brown, William G. Eggington, Zachary D. Smith

Buffalo Law Review

From comic conventions to disbanded dioceses, courts continue to struggle with a unique but puzzling question of trademark law. Federal law protects certain terms that refer to a product or service from a specific producer instead of to a product generally. Terms that refer to products are considered generic and cannot receive protection. Courts have also held that a term that was generic at the time the party adopted the mark cannot receive protection, even if the public later views it as being specific to a particular producer. But, many marks were adopted decades or centuries ago. As a result, …


A Machine Learning Approach To Text-Based Sarcasm Detection, Lara I. Novic Jun 2022

A Machine Learning Approach To Text-Based Sarcasm Detection, Lara I. Novic

Dissertations, Theses, and Capstone Projects

Sarcasm and indirect language are commonplace for humans to produce and recognize but difficult for machines to detect. While artificial intelligence can accurately analyze sentiment and emotion in speech and text, it may struggle with insincere and sardonic content, although it is possible to train a machine to identify uttered and written sarcasm. This paper aims to detect sarcasm using logistic regression and a support vector machine (SVM) and compare their results to a baseline.

The models are trained on headlines from a Kaggle dataset containing headlines from the satirical news website The Onion and serious news website Huffpost (formerly …


Covert Determiners In Appalachian English Narrative Declarative Sentences, William Oliver Jun 2022

Covert Determiners In Appalachian English Narrative Declarative Sentences, William Oliver

Dissertations, Theses, and Capstone Projects

In this thesis, I explore the syntax and semantics of covert determiners (Ds) in matrix subject determiner phrases (DPs) with definite specific interpretations. To conduct my investigation, I used the Audio-Aligned and Parsed Corpus of Appalachian English (AAPCAppE), a million-word Penn Treebank corpus, and the software CorpusSearch, a Java program that searches Penn Treebank corpora. My research shows that Appalachian English contains a linguistic phenomenon where speakers drop the D, replacing overt Ds with covert Ds, in definite specific DPs. For example, where Standard English speakers say The doctor came by horseback, Appalachian speakers may use a covert D …


Metaphor Detection In Poems In Misurata Arabic Sub-Dialect : An Lstm Model, Azza Abugharsa May 2022

Metaphor Detection In Poems In Misurata Arabic Sub-Dialect : An Lstm Model, Azza Abugharsa

Theses, Dissertations and Culminating Projects

Natural Language Processing (NLP) in Arabic is witnessing an increasing interest in investigating different topics in the field. One of the topics that have drawn attention is the automatic processing of Arabic figurative language. The focus in previous projects is on detecting and interpreting metaphors in comments from social media as well as phrases and/or headlines from news articles. The current project focuses on metaphor detection in poems written in the Misurata Arabic sub-dialect spoken in Misurata, located in the North African region. The dataset is initially annotated by a group of linguists, and their annotation is treated as the …


“I Can See The Forest For The Trees”: Examining Personality Traits With Trasformers, Alexander Moore May 2022

“I Can See The Forest For The Trees”: Examining Personality Traits With Trasformers, Alexander Moore

All Dissertations

Our understanding of Personality and its structure is rooted in linguistic studies operating under the assumptions made by the Lexical Hypothesis: personality characteristics that are important to a group of people will at some point be codified in their language, with the number of encoded representations of a personality characteristic indicating their importance. Qualitative and quantitative efforts in the dimension reduction of our lexicon throughout the mid-20th century have played a vital role in the field’s eventual arrival at the widely accepted Five Factor Model (FFM). However, there are a number of presently unresolved conflicts regarding the breadth and …


Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian Apr 2022

Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian

Northeast Journal of Complex Systems (NEJCS)

In this study, we introduce a new network feature for detecting suicidal ideation from clinical texts and conduct various additional experiments to enrich the state of knowledge. We evaluate statistical features with and without stopwords, use lexical networks for feature extraction and classification, and compare the results with standard machine learning methods using a logistic classifier, a neural network, and a deep learning method. We utilize three text collections. The first two contain transcriptions of interviews conducted by experts with suicidal (n=161 patients that experienced severe ideation) and control subjects (n=153). The third collection consists of interviews conducted by experts …


Searching For Pets: Using Distributional And Sentiment-Based Methods To Find Potentially Euphemistic Terms, Patrick Lee, Martha Gavidia, Anna Feldman, Jing Peng Jan 2022

Searching For Pets: Using Distributional And Sentiment-Based Methods To Find Potentially Euphemistic Terms, Patrick Lee, Martha Gavidia, Anna Feldman, Jing Peng

Department of Computer Science Faculty Scholarship and Creative Works

This paper presents a linguistically driven proof of concept for finding potentially euphemistic terms, or PETs. Acknowledging that PETs tend to be commonly used expressions for a certain range of sensitive topics, we make use of distributional similarities to select and filter phrase candidates from a sentence and rank them using a set of simple sentiment-based metrics. We present the results of our approach tested on a corpus of sentences containing euphemisms, demonstrating its efficacy for detecting single and multi-word PETs from a broad range of topics. We also discuss future potential for sentiment-based methods on this task.


A Report On The Euphemisms Detection Shared Task, Patrick Lee, Anna Feldman, Jing Peng Jan 2022

A Report On The Euphemisms Detection Shared Task, Patrick Lee, Anna Feldman, Jing Peng

Department of Computer Science Faculty Scholarship and Creative Works

This paper presents The Shared Task on Euphemism Detection for the Third Workshop on Figurative Language Processing (FigLang 2022) held in conjunction with EMNLP 2022. Participants were invited to investigate the euphemism detection task: given input text, identify whether it contains a euphemism. The input data is a corpus of sentences containing potentially euphemistic terms (PETs) collected from the GloWbE corpus (Davies and Fuchs, 2015), and are human-annotated as containing either a euphemistic or literal usage of a PET. In this paper, we present the results and analyze the common themes, methods and findings of the participating teams.