Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Arts and Humanities (3)
- Discourse and Text Linguistics (3)
- Computer Sciences (2)
- Other Linguistics (2)
- Physical Sciences and Mathematics (2)
-
- African Languages and Societies (1)
- Anthropological Linguistics and Sociolinguistics (1)
- Artificial Intelligence and Robotics (1)
- Communication (1)
- Computer Engineering (1)
- Computer and Systems Architecture (1)
- Digital Humanities (1)
- Engineering (1)
- Language Description and Documentation (1)
- Latin American Languages and Societies (1)
- Morphology (1)
- Pacific Islands Languages and Societies (1)
- Polynesian Studies (1)
- Rhetoric (1)
- Rhetoric and Composition (1)
- Robotics (1)
- Institution
- Publication
- Publication Type
- File Type
Articles 1 - 9 of 9
Full-Text Articles in Computational Linguistics
Automatic Glossing In Under-Resourced Languages: Case Studies In Bribri And Cook Islands Māori, Carter D. Anderson
Automatic Glossing In Under-Resourced Languages: Case Studies In Bribri And Cook Islands Māori, Carter D. Anderson
Linguistics Undergraduate Senior Theses
Interlinear glossing is a major task in Indigenous language documentation. In this paper, I explore how effectively two Large Language Models, ByT5 and Gemini 2.5 Flash, can produce interlinear glossed text. I also examine how prompting an LLM with different types of information (dictionary entries, other training samples, and translations) can augment model performance. I apply these models to two under-resourced Indigenous languages: Bribri, which is morphologically complex from Costa Rica, and Cook Islands Māori, which has a simpler morphology and is from the Cook Islands in the Pacific Ocean. ByT5 exhibits much better performance when glossing Cook Islands Māori …
Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma
Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma
Publications and Research
This presentation introduces Propasafe-Hybrid, a hybrid system for sentence-level propaganda detection that combines offline transformer-based classification with selective large language model (LLM) explainability. The system employs a two-stage pipeline in which a local BERT-based classifier evaluates all input text and filters non-propagandistic content, while only high-confidence candidates are forwarded to an LLM for rhetorical technique labeling and explanation. This design enables cost-aware, privacy-conscious, and scalable analysis by reducing unnecessary reliance on external models.
Propasafe-Hybrid identifies propagandistic techniques such as loaded language, obfuscation, and appeal to fear, and generates concise natural language rationales that make these techniques interpretable to users. By …
A Computational Investigation Of English Spelling, John Winstead
A Computational Investigation Of English Spelling, John Winstead
Theses and Dissertations--Linguistics
This thesis examines the predictability and regularity of English orthography through computational methods. The primary objective is to use n-gram models to predict missing letters in English words by exploiting contextual information from adjacent letters. The study evaluates the impact of dataset size, word length, letter position, and vowel presence on the predictive accuracy of these models, uncovering patterns and structures inherent to English spelling.
The research utilizes a range of datasets, including the Carnegie Mellon University Pronouncing Dictionary, the Brown Corpus, the Corpus of Late Modern English Texts, the Lampeter Corpus of Early Modern English Tracts, and the Open …
‘A Category Of Their Own’: Quantitative Methods In The Use Of Pile-Sort Data In Perceptual Dialectology, Zachary Ty Gill
‘A Category Of Their Own’: Quantitative Methods In The Use Of Pile-Sort Data In Perceptual Dialectology, Zachary Ty Gill
Theses and Dissertations--Linguistics
The purpose of this study is to investigate how Mississippi Gulf Coast Creoles perceive language differences in their home area. A pile-sort task was carried out in which respondents were given stacks of cards with local communities written on them and instructed to stack together the regions where people “talk the same.” Once the piles were made, the fieldworker discussed their sortings with the respondents. The stacks were analyzed by means of a hierarchal agglomerative cluster analysis and non-parametric multidimensional scaling with k-means cluster analysis overlays to extract the perceived dialect areas. The groupings reveal that respondent strategies are based …
Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby
Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby
Dissertations, Theses, and Capstone Projects
In a quick search online, one can find many tools which use information from news headlines to make predictions concerning the trajectory of a given stock. But what if we went further, looking instead into the text of the article, to extract this and other information? Here, the goal is to extract the sentence in which a stock ticker symbol is mentioned from a news article, then determine sentiment and subjectivity values from that sentence, and finally make a prediction on whether or not the value of that stock will go up or not in a 24-hour timespan. Bloomberg News …
Generating Amharic Present Tense Verbs: A Network Morphology & Datr Account, T. Michael W. Halcomb
Generating Amharic Present Tense Verbs: A Network Morphology & Datr Account, T. Michael W. Halcomb
Theses and Dissertations--Linguistics
In this thesis I attempt to model, that is, computationally reproduce, the natural transmission (i.e. inflectional regularities) of twenty present tense Amharic verbs (i.e. triradicals beginning with consonants) as used by the language’s speakers. I root my approach in the linguistic theory of network morphology (NM) and model it using the DATR evaluator. In Chapter 1, I provide an overview of Amharic and discuss the fidel as an abugida, the verb system’s root-and-pattern morphology, and how radicals of each lexeme interacts with prefixes and suffixes. I offer an overview of NM in Chapter 2 and DATR in Chapter 3. In …
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
Dissertations, Theses, and Capstone Projects
In recent years, Twitter has become a popular testing ground for techniques in authorship attribution. This is due to both the ease of building large corpora as well as the challenges associated with the character limit imposed by the service and the writing styles that have developed as a result. As both false and genuine claims of hacked Twitter accounts have made international news, there is an increasing need for this type of work. For newer Twitter accounts, however, there is little training data. Thus, this study looks to lay the groundwork for cross-domain authorship attribution: training on one source …
The Effect Of Sensor Errors In Situated Human-Computer Dialogue, Niels Schütte, John D. Kelleher, Brian Mac Namee
The Effect Of Sensor Errors In Situated Human-Computer Dialogue, Niels Schütte, John D. Kelleher, Brian Mac Namee
Conference papers
Errors in perception are a problem for computer systems that use sensors to perceive the environment. If a computer system is engaged in dialogue with a human user, these problems in perception lead to problems in the dialogue. We present two experiments, one in which participants interact through dialogue with a robot with perfect perception to fulfil a simple task, and a second one in which the robot is affected by sensor errors and compare the resulting dialogues to determine whether the sensor problems have an impact on dialogue success.
Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee
Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee
Articles
In a situated dialogue, misunderstandings may arise if the participants perceive or interpret the environment in different ways. In human-computer dialogue this may be due the sensor errors. We present an experiment system and a series of experiments in which we investigate this problem.