Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

250 Full-Text Articles 420 Authors 150,666 Downloads 65 Institutions

All Articles in Computational Linguistics

Faceted Search

250 full-text articles. Page 1 of 12.

Continuing Christopher Tolkien’S Work In A Digital Age, James K. Tauber 2026 FAU Erlangen-Nürnberg

Continuing Christopher Tolkien’S Work In A Digital Age, James K. Tauber

Journal of Tolkien Research

Christopher Tolkien’s 12-volume History of Middle-earth is a remarkable achievement, yet Christopher repeatedly acknowledged how difficult the material was to present in print form and that “that there was no really satisfactory solution”.  This paper contends that a printed book is not scholarship but just one way of presenting its results. It explores how digital philology can overcome some of the limitations of print, giving examples from the Digital Tolkien Project. It shows how structural markup, version alignment, parallel reading environments, and other approaches can make relationships between drafts and published texts clearer. Importantly, these methods do not replace Christopher’s …


Automatic Glossing In Under-Resourced Languages: Case Studies In Bribri And Cook Islands Māori, Carter D. Anderson 2026 Dartmouth College

Automatic Glossing In Under-Resourced Languages: Case Studies In Bribri And Cook Islands Māori, Carter D. Anderson

Linguistics Undergraduate Senior Theses

Interlinear glossing is a major task in Indigenous language documentation. In this paper, I explore how effectively two Large Language Models, ByT5 and Gemini 2.5 Flash, can produce interlinear glossed text. I also examine how prompting an LLM with different types of information (dictionary entries, other training samples, and translations) can augment model performance. I apply these models to two under-resourced Indigenous languages: Bribri, which is morphologically complex from Costa Rica, and Cook Islands Māori, which has a simpler morphology and is from the Cook Islands in the Pacific Ocean. ByT5 exhibits much better performance when glossing Cook Islands Māori …


The Unspoken And The Unseen: An Analysis Of Victim Gender And Linguistic Framing Of Sexual Assault In Judicial Discourse, Sarnika Ali 2026 Dartmouth College

The Unspoken And The Unseen: An Analysis Of Victim Gender And Linguistic Framing Of Sexual Assault In Judicial Discourse, Sarnika Ali

Quantitative Social Science Undergraduate Senior Theses

Sexual assault is a profound legal and social crisis. However, it is also fundamentally a linguistic one. The words used, or conspicuously not used, to describe victims, perpetrators, and their actions are not neutral arbiters of fact. They are powerful mechanisms that shape perceptions of harm, attributions of blame, and assignments of credibility. The central battleground for survivors is credibility, and while a “credibility discount” is often applied to female victims, the male victim is rendered nearly invisible. This research is therefore guided by one central, overarching question: how does a sexual assault victim’s gender influence the judicial language used, …


The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala 2026 CUNY John Jay College

The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala

Student Theses

The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …


A Machine Learning Approach To Disentangling Developmental Language Disorder From Typical Development In Russian-Speaking Children, Katsiaryna Aharodnik 2026 CUNY Graduate Center

A Machine Learning Approach To Disentangling Developmental Language Disorder From Typical Development In Russian-Speaking Children, Katsiaryna Aharodnik

Dissertations, Theses, and Capstone Projects

This study investigated a machine learning (ML) approach to identifying Developmental Language Disorder (DLD) in Russian-speaking children using narrative data. ML methods can capture subtle linguistic patterns that distinguish typical and atypical development, which is especially important in cross-linguistic contexts where morphosyntactic variation affects the manifestation of DLD. Diagnosis remains challenging in less-studied languages due to limited knowledge of language-specific deficits and a lack of validated assessment tools. This study evaluated whether ML algorithms can provide a more efficient alternative to traditional screening methods.

Two binary classification studies were conducted using corpus data: 1) classification of narratives told by Russian …


G&P2p: A Multi-Source Approach To Grapheme To Phoneme Conversion, Chun-Yi Peng 2026 CUNY Graduate Center

G&P2p: A Multi-Source Approach To Grapheme To Phoneme Conversion, Chun-Yi Peng

Dissertations, Theses, and Capstone Projects

This thesis introduces G&P2P, a multi-source framework for grapheme-to-phoneme (G2P) conversion that integrates side pronunciations from multiple lexical resources. Unlike traditional single-source approaches, G&P2P fuses data from multi-sourced pronunciation dictionaries—including CELEX, PronLex, NETTalk, and WikiPron—through several fusion strategies. The goal is to improve model performance on out-of-vocabulary words through multi-source learning. Experiments were conducted with attentive LSTM, pointer-generator LSTM, and pointer-generator Transformer architectures. Models were trained on combinations of datasets and evaluated using word error rate (WER) across five random seeds.

Results show that fusing expert-curated dictionaries such as CELEX and PronLex consistently improves accuracy, achieving an 11.81-point absolute error …


Understanding Behavioral And Representational Divergences Of Humans And Machines, Thomas Lasman Botch 2026 Dartmouth College

Understanding Behavioral And Representational Divergences Of Humans And Machines, Thomas Lasman Botch

Dartmouth College Ph.D Dissertations

Human behavior and cognition are strikingly variable: people differ from one another in their preferences and abilities, and even from themselves across situations. Yet this diversity arises from common neural machinery shaped by the complex environments humans inhabit. A central objective of cognitive neuroscience is to understand how this varied experience emerges from interactions between brains, agents, and environments. In this dissertation, I argue that comparing behavior and neural computation across humans, artificial systems, and contexts is essential for understanding the flexibility of human cognition. Across three chapters, I use this comparative approach to examine how the rich, multimodal contexts …


Interpreting American Sign Language: A Literature Review Of Assistive Technologies, Natalie Louise Paradiso, Emma Grace Kochenderfer 2026 Chapman University

Interpreting American Sign Language: A Literature Review Of Assistive Technologies, Natalie Louise Paradiso, Emma Grace Kochenderfer

Student Scholar Symposium Abstracts and Posters

American Sign Language (ASL) is a visually elaborate, spatially oriented linguistic methodology that relies on combinations of hand movements, body positioning, facial expressions, and motion/spatial perception, aspects of which make interpretation difficult for automated machine recognition. Current assistive technology approaches to ASL interpretation are generally within the categories of computer vision models (including deep learning, multi-focus image fusion, and keypoint tracking) and wearable, multimodal/sensor-based approaches (such as smart glasses and inertial-sensor gloves). Within controlled environments, computer vision models perform well. However, when applied to conditions such as non-manual signs/features, signer variability, and rapid assimilation, they falter in processing all aspects …


Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma 2026 CUNY John Jay College

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

Propagandistic content increasingly circulates through online news and social media, where readers often encounter it with limited scrutiny, highlighting the need for reliable and fine-grained detection. This paper introduces Propasafe-Hybrid, a sentence-level system that integrates a fine-tuned transformer classifier with LLM-based technique classification to identify, label, and explain specific propaganda strategies. The pipeline generates actionable outputs, including highlighted sentences, technique assignments, and concise rationales, so users can immediately understand why a sentence was flagged and how each label was determined. To control inference cost, Propasafe-Hybrid employs a cost-aware pre-filtering stage that forwards only high-likelihood sentences to LLMs, reducing token usage …


Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma 2026 CUNY John Jay College

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

This presentation introduces Propasafe-Hybrid, a hybrid system for sentence-level propaganda detection that combines offline transformer-based classification with selective large language model (LLM) explainability. The system employs a two-stage pipeline in which a local BERT-based classifier evaluates all input text and filters non-propagandistic content, while only high-confidence candidates are forwarded to an LLM for rhetorical technique labeling and explanation. This design enables cost-aware, privacy-conscious, and scalable analysis by reducing unnecessary reliance on external models.

Propasafe-Hybrid identifies propagandistic techniques such as loaded language, obfuscation, and appeal to fear, and generates concise natural language rationales that make these techniques interpretable to users. By …


Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy 2026 CUNY John Jay College

Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy

Publications and Research

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …


Understanding Perceptions Of A Peruvian Local Market Program: A Reflexive Thematic Analysis, Rosmery Ramos-Sandoval Dr., Jano Ramos-Diaz 2026 Universidad Tecnológica del Perú

Understanding Perceptions Of A Peruvian Local Market Program: A Reflexive Thematic Analysis, Rosmery Ramos-Sandoval Dr., Jano Ramos-Diaz

The Qualitative Report

Despite growing global interest in short food supply chains (SFSCs), little is known about how consumers in developing countries perceive these models, especially through digital platforms like social media. This study investigates how Twitter users represent and perceive SFSCs in the context of the Peruvian government´s “De la Chacra a la Olla” program. The study analyzed 1,167 tweets from Peruvian Twitter users referencing the hashtag #DeLaChacraALaOlla between 2014 and 2020, using reflexive thematic analysis within an exploratory case study framework to examine consumer perceptions of SFSCs. The analysis revealed three key themes in consumer perceptions of SFSCs on Twitter: direct …


Mind The Gap: Morphological Defectivity And Suffixal Competition In Polish, Szymon E. Zuberek 2026 CUNY Graduate Center

Mind The Gap: Morphological Defectivity And Suffixal Competition In Polish, Szymon E. Zuberek

Dissertations, Theses, and Capstone Projects

This study investigates morphological defectivity and suffixal competition in the genitive singular of Polish masculine inanimate nouns. Drawing on survey-based grammaticality judgments from native speakers, it examines how respondents select between the suffixes -a and -u or reject both as unacceptable, thereby signaling defectivity. Mixed-effects logistic regressions revealed that defectivity was rare overall but patterned systematically by age, education, and region, with additional baseline variability across lexical items. Suffix choice showed a strong preference for -a, modulated by age, education, and region, with additional baseline variability across lexical items. These findings inform our understanding of paradigm structure, sociolinguistic variation, and …


A Treasure Hunt: The Interdisciplinary Challenge Of A Digital Humanities Hackathon, Marie Barras, Adélaïde Quenson, Adrien Jeanrenaud, Angela Allemand, Marina Berazategui, Levyn Bürki, Michel Capot, Simon Gabay, Vestin Hategekimana, Pauline Jacsont, Bokar Lamine N'Diaye, Elina Leblanc, Clara May, Anne-Laure Oberson, Margherita Parigini, Lara Pitteloud, Fassaleh Taal, Cédric Viaccoz 2026 Université de Genève

A Treasure Hunt: The Interdisciplinary Challenge Of A Digital Humanities Hackathon, Marie Barras, Adélaïde Quenson, Adrien Jeanrenaud, Angela Allemand, Marina Berazategui, Levyn Bürki, Michel Capot, Simon Gabay, Vestin Hategekimana, Pauline Jacsont, Bokar Lamine N'Diaye, Elina Leblanc, Clara May, Anne-Laure Oberson, Margherita Parigini, Lara Pitteloud, Fassaleh Taal, Cédric Viaccoz

Artl@s Bulletin

Abstract: This article tells the story of an interdisciplinary hackathon emphasizing exploration, collaboration, and creative engagement with globalization-related cultural datasets. Participants worked in teams to produce research posters, which were then presented at the opening of the international conference Image Deluge & Globalization in Geneva. The article discusses three projects: one on 19th-century Spanish chapbooks using topic modeling and network analysis; another on leopard motifs in print imagery; and a third on olfactory heritage. The hackathon highlighted both the potential and the limitations of working with digital cultural data, offering insights into interdisciplinary methods.

Résumé: Cet article raconte un hackathon …


Ctrl + Alt + Inner Speech: A Verbal–Cognitive Scaffold (Vcs) Model Of Pathways To Computational Thinking, Daisuke Akiba 2025 CUNY Queens College; CUNY Graduate Center

Ctrl + Alt + Inner Speech: A Verbal–Cognitive Scaffold (Vcs) Model Of Pathways To Computational Thinking, Daisuke Akiba

Publications and Research

This theoretical paper introduces the Verbal–Cognitive Scaffold (VCS) Model, a cognitively inclusive framework which proposes the cognitive architectures underlying computational thinking (CT). Moving beyond monolithic theories of cognition (e.g., executive-function and metacognitive control models), the VCS Model posits inner speech (InSp) as the predominant cognitive pathway supporting CT operations in neurotypical populations. Synthesizing interdisciplinary scholarship across cognitive science, computational theory, neurodiversity research, and others, this framework articulates distinct mechanisms through which InSp supports CT. The model specifies four primary pathways linking InSp to CT components: verbal working memory supporting decomposition, symbolic representation facilitating pattern recognition and abstraction, sequential processing enabling …


Syntax-Enhanced Boundary-Aware Named Entity Recognition Model, Chuanming YU, Bin DENG, Zhengang ZHANG 2025 School of Information Engineering, Zhongnan University of Economics and Law, Wuhan 430073

Syntax-Enhanced Boundary-Aware Named Entity Recognition Model, Chuanming Yu, Bin Deng, Zhengang Zhang

Journal of Scientific Information Research

[Purpose/significance] This study addresses the issue of inadequate perception of entity boundaries in traditional character-level modeling-based named entity recognition models by integrating syntax information containing entity boundary features into the task using a multi-head graph attention network with dense connections. This integration enhances the effectiveness of named entity recognition.

[Method/process] This study proposes a Syntax-enhanced Boundary-aware Named Entity Recognition Model (SynBNER), which utilizes BERT for text semantic representation and integrates syntax information using a dense-connected graph attention network. This integration incorporates implicit entity boundary information from syntax information into word representations, thereby enhancing the model's entity boundary perception capability.

[Result/conclusion] …


Revitalization Of Endangered Languages With Ai, Ivory Yang 2025 Dartmouth College

Revitalization Of Endangered Languages With Ai, Ivory Yang

Dartmouth College Master’s Theses

The preservation and revitalization of endangered languages, particularly those with minimal digital presence, presents significant challenges for computational linguistics. This thesis addresses these challenges by proposing novel methods for language identification and data generation, focusing on underrepresented Indigenous languages, specifically Nüshu, Native American and Native Alaskan languages.

In the first study, a COLING 2025 paper, we present NüshuRescue, an AI-driven framework designed to facilitate the preservation of Nüshu, an endangered script used exclusively by Yao women in China. Using minimal seed data, we demonstrate how GPT-4-Turbo can generate new translations, expanding a publicly available Nüshu-Chinese corpus, achieving 48.69% accuracy in …


Alle Or Elle: Automatic Speech Recognition On Louisiana French, Emily Chiu 2025 CUNY Graduate Center

Alle Or Elle: Automatic Speech Recognition On Louisiana French, Emily Chiu

Dissertations, Theses, and Capstone Projects

Applications of automatic speech recognition largely serve the most commonly spoken languages, but can cause harm through bias when used for speakers of underrepresented language varieties who are not adequately supported. This experiment’s goal is to reveal how state-of-the-art end-to-end ASR systems perform with Louisiana French, a nonstandard variety of French that has suffered a history of state-sanctioned language oppression in Louisiana.


Fact-Checking As A Multi-Step Process: From Ambiguity Resolution To Claim Validation, Wenbo Wang 2025 New Jersey Institute of Technology

Fact-Checking As A Multi-Step Process: From Ambiguity Resolution To Claim Validation, Wenbo Wang

Dissertations

The spread of misinformation and disinformation has become a major concern, particularly with the rise of social media as a primary source of information for many people. Fact-checking—the process of verifying claims against credible evidence—has emerged as a critical safeguard against misinformation. Yet, the task is fraught with challenges: claims are often ambiguous, context-dependent, or composed of multiple intertwined assertions, while automated systems struggle to replicate the nuanced reasoning of human experts. This dissertation addresses these challenges by reimagining fact-checking as a multi-step, knowledge-guided process that systematically resolves ambiguity, decomposes complexity, and validates claims through structured reasoning. Additionally, the proposed …


From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie 2025 Bellarmine University

From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie

Undergraduate Theses

Adversarial attacks pose a significant threat to the reliability of machine learning-based spam detection systems in social media. This undergraduate thesis, "From Adversarial Attacks to Robust Classifiers: A Study in Social Media Spam Detection – Black Box & White Box," systematically examines the impact of both black-box and white-box adversarial attacks on a range of spam classifiers, including Logistic Regression, Decision Trees, Random Forests, K-Nearest Neighbors, Bagging, Gradient Boosting, and Support Vector Machines. Leveraging a novel dataset derived from Twitter spam messages and enhanced with adversarial perturbations such as synonym replacement and character-level modifications, this study evaluates classifier performance under …


Digital Commons powered by bepress