Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (5)
- Physical Sciences and Mathematics (5)
- Artificial Intelligence and Robotics (3)
- Psychology (3)
- Anthropological Linguistics and Sociolinguistics (2)
-
- Applied Linguistics (2)
- Arts and Humanities (2)
- Communication (2)
- Computer Engineering (2)
- Digital Humanities (2)
- Engineering (2)
- Library and Information Science (2)
- Psycholinguistics and Neurolinguistics (2)
- Social Influence and Political Communication (2)
- Anthropology (1)
- Clinical Psychology (1)
- Cognitive Psychology (1)
- Cognitive Science (1)
- Commercial Law (1)
- Communication Technology and New Media (1)
- Computational Engineering (1)
- Criminology (1)
- Data Science (1)
- Digital Communications and Networking (1)
- Discourse and Text Linguistics (1)
- Education (1)
- Film and Media Studies (1)
- Institution
- Publication
- Publication Type
Articles 1 - 10 of 10
Full-Text Articles in Computational Linguistics
Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma
Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma
Publications and Research
Propagandistic content increasingly circulates through online news and social media, where readers often encounter it with limited scrutiny, highlighting the need for reliable and fine-grained detection. This paper introduces Propasafe-Hybrid, a sentence-level system that integrates a fine-tuned transformer classifier with LLM-based technique classification to identify, label, and explain specific propaganda strategies. The pipeline generates actionable outputs, including highlighted sentences, technique assignments, and concise rationales, so users can immediately understand why a sentence was flagged and how each label was determined. To control inference cost, Propasafe-Hybrid employs a cost-aware pre-filtering stage that forwards only high-likelihood sentences to LLMs, reducing token usage …
Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy
Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy
Publications and Research
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …
Revitalization Of Endangered Languages With Ai, Ivory Yang
Revitalization Of Endangered Languages With Ai, Ivory Yang
Dartmouth College Master’s Theses
The preservation and revitalization of endangered languages, particularly those with minimal digital presence, presents significant challenges for computational linguistics. This thesis addresses these challenges by proposing novel methods for language identification and data generation, focusing on underrepresented Indigenous languages, specifically Nüshu, Native American and Native Alaskan languages.
In the first study, a COLING 2025 paper, we present NüshuRescue, an AI-driven framework designed to facilitate the preservation of Nüshu, an endangered script used exclusively by Yao women in China. Using minimal seed data, we demonstrate how GPT-4-Turbo can generate new translations, expanding a publicly available Nüshu-Chinese corpus, achieving 48.69% accuracy in …
Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell
Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell
Data and Test Instruments
This paper outlines the complexity of the psychological construct of individuals' subjective well-being (SWB) and argues for the importance of examining behaviours and linguistic expression of individuals online social interactions in relation to self-reported SWB. This paper calls for a systematic review of the psychology research which examines SWB and its association with various character strengths, personality traits, and behaviours. While the Big Five personality traits (OCEAN) have an underlying neuropsychological basis and are considered as universal dimensions of personality along which humans differ one from another, minimal research has attempted to evaluate the relationship between personality traits, SWB, and …
Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin
Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin
Data and Test Instruments
How do we detect, deter, and prevent the spread of mis- and disinformationwith the human eye and AI? How does theory inform the practice, and how do theevidence-based research and best practices in lie-catching and truth-seekingprofessions—inform AI? The book looks into well-established human practicessuch as the routines and processes used in detective work, journalism, and scientificinquiry, and how they contribute toward innovative AI solutions. The book explainsthe principles, inner workings, and recent evolution of five types of state-of-the-artAI technologies suitable for curtailing the spread of mis- and disinformation:automated deception detectors, clickbait detectors, satirical fake detectors, rumordebunkers, and computational fact-checking tools.
Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng
Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nonetheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, …
Phonologically-Informed Speech Coding For Automatic Speech Recognition-Based Foreign Language Pronunciation Training, Anthony J. Vicario
Phonologically-Informed Speech Coding For Automatic Speech Recognition-Based Foreign Language Pronunciation Training, Anthony J. Vicario
Dissertations, Theses, and Capstone Projects
Automatic speech recognition (ASR) and computer-assisted pronunciation training (CAPT) systems used in foreign-language educational contexts are often not developed with the specific task of second-language acquisition in mind. Systems that are built for this task are often excessively targeted to one native language (L1) or a single phonemic contrast and are therefore burdensome to train. Current algorithms have been shown to provide erroneous feedback to learners and show inconsistencies between human and computer perception. These discrepancies have thus far hindered more extensive application of ASR in educational systems.
This thesis reviews the computational models of the human perception of American …
Quantifying Coherence In A Transdiagnostic Sample: A Methodological Investigation Of Computationally-Derived Coherence Using Ambulatory Assessment, Taylor L. Fedechko
Quantifying Coherence In A Transdiagnostic Sample: A Methodological Investigation Of Computationally-Derived Coherence Using Ambulatory Assessment, Taylor L. Fedechko
LSU Master's Theses
Schizophrenia is a clinical diagnosis assigned to individuals that experience positive (e.g., hallucinations and delusions), negative (e.g., blunted affect), and disorganized (e.g., incoherent speech) symptoms. One particularly disabling symptom is incoherence, which is defined as the meaning-based relationship between ideas. This symptom can drastically affect an individual’s quality of life by affecting areas such as social and occupational functioning. Currently, the mechanism behind this symptom is unknown and requires further study. One way to examine incoherence is to understand its level of expression in other clinical populations. With the advent of computationally-derived natural language processing (NLP), coherence can be quantified …
Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin
Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin
Data and Test Instruments
Automatic clickbait detection is a relatively novel task in natural language processing (NLP) and machine learning (ML). “Clickbait” is a hyperlink created primarily to attract attention to its target content. This article introduces a binary classifier, the Language and Information Technology Research Lab (LiT.RL, pronounced “literal”) Clickbait Detector, which automatically distinguishes clickbait from nonclickbait. We used NLP and ML for 38 textual features, contrasting clickbait with “headlinese.” When tested on 11,000 hyperlinks, it achieves 94 per cent accuracy using a support vector machine. Integrated with the LiT.RL News Verification Browser, a downloadable stand-alone research tool, the Clickbait Detector user interface …
A Study On The Efficacy Of Sentiment Analysis In Author Attribution, Michael J. Schneider
A Study On The Efficacy Of Sentiment Analysis In Author Attribution, Michael J. Schneider
Electronic Theses and Dissertations
The field of authorship attribution seeks to characterize an author’s writing style well enough to determine whether he or she has written a text of interest. One subfield of authorship attribution, stylometry, seeks to find the necessary literary attributes to quantify an author’s writing style. The research presented here sought to determine the efficacy of sentiment analysis as a new stylometric feature, by comparing its performance in attributing authorship against the performance of traditional stylometric features. Experimentation, with a corpus of sci-fi texts, found sentiment analysis to have a much lower performance in assigning authorship than the traditional stylometric features.