Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

2026

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 1 - 14 of 14

Full-Text Articles in Computational Linguistics

Continuing Christopher Tolkien’S Work In A Digital Age, James K. Tauber Jul 2026

Continuing Christopher Tolkien’S Work In A Digital Age, James K. Tauber

Journal of Tolkien Research

Christopher Tolkien’s 12-volume History of Middle-earth is a remarkable achievement, yet Christopher repeatedly acknowledged how difficult the material was to present in print form and that “that there was no really satisfactory solution”.  This paper contends that a printed book is not scholarship but just one way of presenting its results. It explores how digital philology can overcome some of the limitations of print, giving examples from the Digital Tolkien Project. It shows how structural markup, version alignment, parallel reading environments, and other approaches can make relationships between drafts and published texts clearer. Importantly, these methods do not replace Christopher’s …


Automatic Glossing In Under-Resourced Languages: Case Studies In Bribri And Cook Islands Māori, Carter D. Anderson Jun 2026

Automatic Glossing In Under-Resourced Languages: Case Studies In Bribri And Cook Islands Māori, Carter D. Anderson

Linguistics Undergraduate Senior Theses

Interlinear glossing is a major task in Indigenous language documentation. In this paper, I explore how effectively two Large Language Models, ByT5 and Gemini 2.5 Flash, can produce interlinear glossed text. I also examine how prompting an LLM with different types of information (dictionary entries, other training samples, and translations) can augment model performance. I apply these models to two under-resourced Indigenous languages: Bribri, which is morphologically complex from Costa Rica, and Cook Islands Māori, which has a simpler morphology and is from the Cook Islands in the Pacific Ocean. ByT5 exhibits much better performance when glossing Cook Islands Māori …


The Unspoken And The Unseen: An Analysis Of Victim Gender And Linguistic Framing Of Sexual Assault In Judicial Discourse, Sarnika Ali Jun 2026

The Unspoken And The Unseen: An Analysis Of Victim Gender And Linguistic Framing Of Sexual Assault In Judicial Discourse, Sarnika Ali

Quantitative Social Science Undergraduate Senior Theses

Sexual assault is a profound legal and social crisis. However, it is also fundamentally a linguistic one. The words used, or conspicuously not used, to describe victims, perpetrators, and their actions are not neutral arbiters of fact. They are powerful mechanisms that shape perceptions of harm, attributions of blame, and assignments of credibility. The central battleground for survivors is credibility, and while a “credibility discount” is often applied to female victims, the male victim is rendered nearly invisible. This research is therefore guided by one central, overarching question: how does a sexual assault victim’s gender influence the judicial language used, …


The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala Jun 2026

The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala

Student Theses

The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …


G&P2p: A Multi-Source Approach To Grapheme To Phoneme Conversion, Chun-Yi Peng Jun 2026

G&P2p: A Multi-Source Approach To Grapheme To Phoneme Conversion, Chun-Yi Peng

Dissertations, Theses, and Capstone Projects

This thesis introduces G&P2P, a multi-source framework for grapheme-to-phoneme (G2P) conversion that integrates side pronunciations from multiple lexical resources. Unlike traditional single-source approaches, G&P2P fuses data from multi-sourced pronunciation dictionaries—including CELEX, PronLex, NETTalk, and WikiPron—through several fusion strategies. The goal is to improve model performance on out-of-vocabulary words through multi-source learning. Experiments were conducted with attentive LSTM, pointer-generator LSTM, and pointer-generator Transformer architectures. Models were trained on combinations of datasets and evaluated using word error rate (WER) across five random seeds.

Results show that fusing expert-curated dictionaries such as CELEX and PronLex consistently improves accuracy, achieving an 11.81-point absolute error …


A Machine Learning Approach To Disentangling Developmental Language Disorder From Typical Development In Russian-Speaking Children, Katsiaryna Aharodnik Jun 2026

A Machine Learning Approach To Disentangling Developmental Language Disorder From Typical Development In Russian-Speaking Children, Katsiaryna Aharodnik

Dissertations, Theses, and Capstone Projects

This study investigated a machine learning (ML) approach to identifying Developmental Language Disorder (DLD) in Russian-speaking children using narrative data. ML methods can capture subtle linguistic patterns that distinguish typical and atypical development, which is especially important in cross-linguistic contexts where morphosyntactic variation affects the manifestation of DLD. Diagnosis remains challenging in less-studied languages due to limited knowledge of language-specific deficits and a lack of validated assessment tools. This study evaluated whether ML algorithms can provide a more efficient alternative to traditional screening methods.

Two binary classification studies were conducted using corpus data: 1) classification of narratives told by Russian …


Understanding Behavioral And Representational Divergences Of Humans And Machines, Thomas Lasman Botch May 2026

Understanding Behavioral And Representational Divergences Of Humans And Machines, Thomas Lasman Botch

Dartmouth College Ph.D Dissertations

Human behavior and cognition are strikingly variable: people differ from one another in their preferences and abilities, and even from themselves across situations. Yet this diversity arises from common neural machinery shaped by the complex environments humans inhabit. A central objective of cognitive neuroscience is to understand how this varied experience emerges from interactions between brains, agents, and environments. In this dissertation, I argue that comparing behavior and neural computation across humans, artificial systems, and contexts is essential for understanding the flexibility of human cognition. Across three chapters, I use this comparative approach to examine how the rich, multimodal contexts …


Interpreting American Sign Language: A Literature Review Of Assistive Technologies, Natalie Louise Paradiso, Emma Grace Kochenderfer May 2026

Interpreting American Sign Language: A Literature Review Of Assistive Technologies, Natalie Louise Paradiso, Emma Grace Kochenderfer

Student Scholar Symposium Abstracts and Posters

American Sign Language (ASL) is a visually elaborate, spatially oriented linguistic methodology that relies on combinations of hand movements, body positioning, facial expressions, and motion/spatial perception, aspects of which make interpretation difficult for automated machine recognition. Current assistive technology approaches to ASL interpretation are generally within the categories of computer vision models (including deep learning, multi-focus image fusion, and keypoint tracking) and wearable, multimodal/sensor-based approaches (such as smart glasses and inertial-sensor gloves). Within controlled environments, computer vision models perform well. However, when applied to conditions such as non-manual signs/features, signer variability, and rapid assimilation, they falter in processing all aspects …


Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma May 2026

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

Propagandistic content increasingly circulates through online news and social media, where readers often encounter it with limited scrutiny, highlighting the need for reliable and fine-grained detection. This paper introduces Propasafe-Hybrid, a sentence-level system that integrates a fine-tuned transformer classifier with LLM-based technique classification to identify, label, and explain specific propaganda strategies. The pipeline generates actionable outputs, including highlighted sentences, technique assignments, and concise rationales, so users can immediately understand why a sentence was flagged and how each label was determined. To control inference cost, Propasafe-Hybrid employs a cost-aware pre-filtering stage that forwards only high-likelihood sentences to LLMs, reducing token usage …


Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma May 2026

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

This presentation introduces Propasafe-Hybrid, a hybrid system for sentence-level propaganda detection that combines offline transformer-based classification with selective large language model (LLM) explainability. The system employs a two-stage pipeline in which a local BERT-based classifier evaluates all input text and filters non-propagandistic content, while only high-confidence candidates are forwarded to an LLM for rhetorical technique labeling and explanation. This design enables cost-aware, privacy-conscious, and scalable analysis by reducing unnecessary reliance on external models.

Propasafe-Hybrid identifies propagandistic techniques such as loaded language, obfuscation, and appeal to fear, and generates concise natural language rationales that make these techniques interpretable to users. By …


Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy Apr 2026

Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy

Publications and Research

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …


Understanding Perceptions Of A Peruvian Local Market Program: A Reflexive Thematic Analysis, Rosmery Ramos-Sandoval Dr., Jano Ramos-Diaz Feb 2026

Understanding Perceptions Of A Peruvian Local Market Program: A Reflexive Thematic Analysis, Rosmery Ramos-Sandoval Dr., Jano Ramos-Diaz

The Qualitative Report

Despite growing global interest in short food supply chains (SFSCs), little is known about how consumers in developing countries perceive these models, especially through digital platforms like social media. This study investigates how Twitter users represent and perceive SFSCs in the context of the Peruvian government´s “De la Chacra a la Olla” program. The study analyzed 1,167 tweets from Peruvian Twitter users referencing the hashtag #DeLaChacraALaOlla between 2014 and 2020, using reflexive thematic analysis within an exploratory case study framework to examine consumer perceptions of SFSCs. The analysis revealed three key themes in consumer perceptions of SFSCs on Twitter: direct …


Mind The Gap: Morphological Defectivity And Suffixal Competition In Polish, Szymon E. Zuberek Feb 2026

Mind The Gap: Morphological Defectivity And Suffixal Competition In Polish, Szymon E. Zuberek

Dissertations, Theses, and Capstone Projects

This study investigates morphological defectivity and suffixal competition in the genitive singular of Polish masculine inanimate nouns. Drawing on survey-based grammaticality judgments from native speakers, it examines how respondents select between the suffixes -a and -u or reject both as unacceptable, thereby signaling defectivity. Mixed-effects logistic regressions revealed that defectivity was rare overall but patterned systematically by age, education, and region, with additional baseline variability across lexical items. Suffix choice showed a strong preference for -a, modulated by age, education, and region, with additional baseline variability across lexical items. These findings inform our understanding of paradigm structure, sociolinguistic variation, and …


A Treasure Hunt: The Interdisciplinary Challenge Of A Digital Humanities Hackathon, Marie Barras, Adélaïde Quenson, Adrien Jeanrenaud, Angela Allemand, Marina Berazategui, Levyn Bürki, Michel Capot, Simon Gabay, Vestin Hategekimana, Pauline Jacsont, Bokar Lamine N'Diaye, Elina Leblanc, Clara May, Anne-Laure Oberson, Margherita Parigini, Lara Pitteloud, Fassaleh Taal, Cédric Viaccoz Jan 2026

A Treasure Hunt: The Interdisciplinary Challenge Of A Digital Humanities Hackathon, Marie Barras, Adélaïde Quenson, Adrien Jeanrenaud, Angela Allemand, Marina Berazategui, Levyn Bürki, Michel Capot, Simon Gabay, Vestin Hategekimana, Pauline Jacsont, Bokar Lamine N'Diaye, Elina Leblanc, Clara May, Anne-Laure Oberson, Margherita Parigini, Lara Pitteloud, Fassaleh Taal, Cédric Viaccoz

Artl@s Bulletin

Abstract: This article tells the story of an interdisciplinary hackathon emphasizing exploration, collaboration, and creative engagement with globalization-related cultural datasets. Participants worked in teams to produce research posters, which were then presented at the opening of the international conference Image Deluge & Globalization in Geneva. The article discusses three projects: one on 19th-century Spanish chapbooks using topic modeling and network analysis; another on leopard motifs in print imagery; and a third on olfactory heritage. The hackathon highlighted both the potential and the limitations of working with digital cultural data, offering insights into interdisciplinary methods.

Résumé: Cet article raconte un hackathon …