Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

250 Full-Text Articles 420 Authors 150,666 Downloads 65 Institutions

All Articles in Computational Linguistics

Faceted Search

250 full-text articles. Page 2 of 12.

On The Provenance Of Software Systems: Automating Software Traceability With Knowledge Graph And Large Language Model Synergy, Tyler Procko 2025 Embry-Riddle Aeronautical University

On The Provenance Of Software Systems: Automating Software Traceability With Knowledge Graph And Large Language Model Synergy, Tyler Procko

Doctoral Dissertations and Master's Theses

The present dissertation delineates a system that enables those engaged in software development to automatically generate and maintain project life cycle provenance. All projects are implemented and made manifest with the development of artifacts, e.g., papers, code files, etc. Tools exist to accelerate artifact creation, but little focus is paid to the processes that produce them. In terms of Ontology, or, from Ancient Greek, the study of being, the two most basic entities in reality are Continuant and Occurrent, or, roughly, “Artifact” and “Process”. This dissertation posits that for any created artifact, its process of creation, i.e., its life …


Improving Low-Resource Translation With Finite State Grammars, Nicholas J. Uva 2025 CUNY Graduate Center

Improving Low-Resource Translation With Finite State Grammars, Nicholas J. Uva

Dissertations, Theses, and Capstone Projects

Scarcity of training data continues to pose a problem for the development of neural machine translation systems for low-resource languages. This study develops a method for the incorporation of linguistic information into the training of neural machine translation models for low-resource languages, using morphological grammars created using finite state transducers. This study explores the benefits, historical background, and effectiveness of this approach. This study incorporates morphological tags into a pre-trained multilingual neural machine translation model using a dual encoder structure. This study finds an improvement in performance in the Irish-English translation scenario. This method offers promising results with low computational …


A Proposed Ehrenfeucht-Fraïssé Game Model For Natural Language Processing Generative Adversarial Networks, Don Li 2025 Portland State University

A Proposed Ehrenfeucht-Fraïssé Game Model For Natural Language Processing Generative Adversarial Networks, Don Li

Anthós

Large Language Models (LLM’s) (e.g., ChatGPT) constitute both a significant research area and commercial application of AI. Current major LLM’s are built on Generative Pre-Trained Transformer (GPT) neural network architecture to perform natural language processing (NLP) tasks. Generative Adversarial Network (GAN) is another popular neural network architecture, which leverages a zero-sum game between constituent neural networks within the architecture to train the GAN, and is widely used for visual data applications. This article proposes a new GAN architecture for NLP: an EF-GAN whose underlying algorithm uses Ehrenfeucht–Fraïssé (EF) games, a game-theoretic approach from model theory to determine elementary equivalence of …


Characterizing Language Use In Online Accessibility Discussion Forums, Nithiya Venkatraman, Anand Ravi Aiyer, Yash Prakash, Sampath Jayarathna, Hae-Na Lee, Vikas Ashok 2025 Old Dominion University

Characterizing Language Use In Online Accessibility Discussion Forums, Nithiya Venkatraman, Anand Ravi Aiyer, Yash Prakash, Sampath Jayarathna, Hae-Na Lee, Vikas Ashok

Computer Science Faculty Publications

Discussion forums are one of the favored platforms for knowledge sharing. Given their popularity, copious research exists on understanding the linguistic and behavioral characteristics of forum conversations, so as to inform the design of many downstream applications including discourse visualization, sentiment analysis, and question answering. However, prior investigations have mainly focused on general forums designed primarily for sighted users, and as such the applicability of their findings to dedicated accessibility discussion forums frequented by blind screen reader users remains unanswered. To bridge this knowledge gap and facilitate the development of better-informed assistive technologies for blind people, we investigated language use …


Investigating Post-Adoption Abandonment Of Mental Health Mobile Applications Among Young Adults, Donald Harris 2025 University at Albany, State University of New York

Investigating Post-Adoption Abandonment Of Mental Health Mobile Applications Among Young Adults, Donald Harris

Electronic Theses & Dissertations (2024 - present)

The rising prevalence of mental health issues among young adults has driven increased interest in Mental Health Mobile Applications (MHMAs), which offer accessible and cost-effective solutions to traditional barriers such as financial limitations, stigma, and restricted healthcare access. Despite their promise, MHMAs frequently experience high rates of attrition and abandonment, significantly limiting their long-term effectiveness. Employing a mixed-methods, multi-stage research design, this dissertation explores the determinants of MHMA abandonment among young adults, emphasizing the interplay between technological inhibitors and enablers, individual user characteristics, and the mediating roles of user satisfaction and perceived usefulness.

Study 1 utilized quantitative text analysis, including …


Modeling Context And The Characteristica Universalis, John Kausch 2024 Western University of Ontario

Modeling Context And The Characteristica Universalis, John Kausch

Proceedings from the Document Academy

This paper proposes prototypes for the exploration of the context of terms in a knowledge organization system by visualizing machine learning produced word embeddings. It puts this work in the context of the search for a universal language, typified by Leibniz’s characteristica universalis. This tradition of the search for universal languages is put in the context of universalizing tendencies in taxonomic classification in library and information science. Following this there is a discussion of the use of machine learning models to represent context. These two concerns inform the construction of prototypes for exploring the contextual spaces produced by word embeddings …


Controlling Emotional Text To Speech Using Complex Adverbial Phrases, Zainab T. Akande 2024 CUNY Graduate Center

Controlling Emotional Text To Speech Using Complex Adverbial Phrases, Zainab T. Akande

Dissertations, Theses, and Capstone Projects

This study investigates the usage of adverbial modifiers from audiobook data as a resource for training speech synthesizers with a greater range of speech descriptions. The Tacotron2 text-to-speech (TTS) model was used for the purposes of this study. Utilizing the LibriTTS dataset, the Tacotron2 model is trained under two experimental conditions: one incorporating adverbial modifiers into the input text and the other without. The dataset preprocessing involves embedding descriptions using word embeddings and encoding speaker IDs with machine learning techniques. Additionally, the model architecture includes a prosody encoder inspired by prior research. Evaluation of the trained models involves subjective assessments …


Expanding The Corpus Of Vocalized Hebrew Text: Compiling An Unvocalized Text Corpus And Building An Online Interface For Vocalization Annotation, Rachel Shanblatt Bloch 2024 CUNY Graduate Center

Expanding The Corpus Of Vocalized Hebrew Text: Compiling An Unvocalized Text Corpus And Building An Online Interface For Vocalization Annotation, Rachel Shanblatt Bloch

Dissertations, Theses, and Capstone Projects

Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of …


Predicting Language Proficiency Using A Multiple Regression Model, Madisen Barrieau 2024 Liberty University

Predicting Language Proficiency Using A Multiple Regression Model, Madisen Barrieau

Senior Honors Theses

Businesses that design language learning products have a common goal to impart language proficiency to users. Many variables play a role in reaching language proficiency, including time spent in study, method of learning, use of technology, motivation, and more. This study involved creating several multiple regression models in R on a dataset featuring many of these variables. This research sought to identify the relative weight and measurability of these variables unto the goal of predicting language proficiency. Several regression models were created, and the best model showed that language similarity and length of residence in target culture, among other factors, …


Uncovering The Mimicry Of Online Review Breadth And Depth And Its Subsequent Effect On Consumer Responses, Andrea Pelaez Martinez 2024 CUNY Graduate Center

Uncovering The Mimicry Of Online Review Breadth And Depth And Its Subsequent Effect On Consumer Responses, Andrea Pelaez Martinez

Dissertations, Theses, and Capstone Projects

Word-of-mouth (WOM) in marketing occurs when consumers discuss a company's product or service or any consumption experience with their friends, family, and others with whom they have any relationship. With the advent of social media, this phenomenon has expanded rapidly into virtual environments where consumer conversation is enabled through chats, forums, social media posts, and online reviews. In response to this rapid growth of online WOM, academics and practitioners have focused their interest on this phenomenon and its implications on consumers, firms, and society. So far, the evidence of the critical role that online WOM plays in helping consumers make …


Computational Approaches To Linguistic Challenges In Arabic Speech Recognition, Enas Albasiri 2024 CUNY Graduate Center

Computational Approaches To Linguistic Challenges In Arabic Speech Recognition, Enas Albasiri

Dissertations, Theses, and Capstone Projects

This dissertation aims to document the linguistic features of Arabic that pose challenges to speech and language technologies and advance these technologies by developing state-of-the-art computational tools focusing on automatic speech recognition (ASR), text normalization (TN), and corpus development. TN converts expressions such as numbers, dates, and times—named semiotic classes—from their written to their spoken domain, such as converting ‘$84.00’ to ‘eighty-four dollars’, while inverse text normalization (ITN) converts verbalized text to its written form. This conversion is an essential preprocessing step for text-to-speech (TTS), and post-processing step for ASR. Arabic presents a challenge for TN and ITN because one …


Gpt Assisted Annotation Of Rhetorical And Linguistic Features For Interpretable Propaganda Technique Detection In News Text., Kyle Hamilton, Bojan Bozic, Luca Longo 2024 Technological University Dublin

Gpt Assisted Annotation Of Rhetorical And Linguistic Features For Interpretable Propaganda Technique Detection In News Text., Kyle Hamilton, Bojan Bozic, Luca Longo

Articles

While the use of machine learning for the detection of propaganda techniques in text has garnered considerable attention, most approaches focus on "black-box'' solutions with opaque inner workings. Interpretable approaches provide a solution, however, they depend on careful feature engineering and costly expert annotated data. Additionally, language features specific to propagandistic text are generally the focus of rhetoricians or linguists, and there is no data set labeled with such features suitable for machine learning. This study codifies 22 rhetorical and linguistic features identified in literature related to the language of persuasion for the purpose of annotating an existing data set …


Exploring Asynchronous Pronunciation Training Through Context-Aware Pronunciation Applications, Claire L. Schweikert 2024 University of Nebraska at Omaha

Exploring Asynchronous Pronunciation Training Through Context-Aware Pronunciation Applications, Claire L. Schweikert

Theses/Capstones/Creative Projects

This paper provides a survey of various research articles on context-aware asynchronous pronunciation training applications. First, a set of seven articles is reviewed and summarized. Next, they are synthesized over the three main topics of 1) automated speech recognition, 2) non-native speaker considerations in language learning, and 3) future directions for research and development within computer-assisted pronunciation training (CAPT). Research in the areas of acoustic and pronunciation modeling (both implicit and explicit), pedagogical considerations for CAPT application design, Goodness of Pronunciation algorithm scoring, accent recognition and neutralization, and more are discussed.


Bugsy The Bee's Big Adventure, Jonah L., Zachery Irvin, Jake Wildstrom, Zachariah Robinson, Hilaria Cruz 2024 University of Louisville

Bugsy The Bee's Big Adventure, Jonah L., Zachery Irvin, Jake Wildstrom, Zachariah Robinson, Hilaria Cruz

LING 590/Internet Language

Bugsy the bee goes on a quest for pollen.


Skyler's Lunch, Noah Sherman, Autumn Boone, Hilaria Cruz 2024 University of Louisville

Skyler's Lunch, Noah Sherman, Autumn Boone, Hilaria Cruz

LING 590/Internet Language

Our class was studying the use of emojis across different platforms and wanted to explore how stories using emojis could impact young readers. Here, we try to translate the story of Skyler into emoji, providing translations along the way. We replace words completely with emoji, represent phrases with a few emoji, and use additional emoji to make sense of the content, including punctuation. In this book, we explore the character of Skyler, who is a picky eater. But they learn to eat the nutritious food that is good for them. In the end, they even get a reward!


Research On The Application Of Part-Of-Speech Tagging Of Ancient Books Under The Domain Large Language Model, Danhao ZHU, ZHAO Zhixiao, Die HU, Wenhua ZHAO 2024 1.Department of Criminal Science and Technology, Jiangsu Police Institute, Nanjing 210031

Research On The Application Of Part-Of-Speech Tagging Of Ancient Books Under The Domain Large Language Model, Danhao Zhu, Zhao Zhixiao, Die Hu, Wenhua Zhao

Journal of Scientific Information Research

[Purpose/significance]The development of the large language model has brought new ideas for ancient text mining, and combining the large language model with the digitisation and intelligence of ancient books is a necessary path for the work of ancient books in the new era. [Methods/process]This paper uses the lexically annotated corpus of Zuozhuan to construct a batch of high-quality lexically annotated instruction data through data cleaning and preprocessing, on the basis of which 500, 1 000, 2 000, and 5 000 pieces of data are used to fine-tune the instructions of the large language model, and the performance test is carried …


Retórica Intercultural En El Discurso Académico Universitario: Las Funciones Retóricas De La Citación En Los Trabajos De Fin De Máster Escritos En Español Y En Inglés Por Hablantes Nativos Y No Nativos, David Sanchez-Jimenez 2024 City University of New York (CUNY)

Retórica Intercultural En El Discurso Académico Universitario: Las Funciones Retóricas De La Citación En Los Trabajos De Fin De Máster Escritos En Español Y En Inglés Por Hablantes Nativos Y No Nativos, David Sanchez-Jimenez

Publications and Research

This research derives from the interest in learning the cultural differences in citation practices in the academic genre of Master's thesis of native Spanish (Ee), non-native Filipino writers of Spanish (Fe), native Filipino writers of English (Fi), and American writers of English. A total of thirty-two (32) master´s theses – eight (8) for each group – were analyzed. A quantitative and qualitative methodology was used to study this phenomenon based on the computerized textual analysis of the rhetorical function of citations arranged in typological classification that modified the outline proposed by Petrić in his 2007 article. The results obtained from …


How Do We Learn What We Cannot Say?, Daniel Yakubov 2024 CUNY Graduate Center

How Do We Learn What We Cannot Say?, Daniel Yakubov

Dissertations, Theses, and Capstone Projects

The contributions of this thesis are two-fold. First, this thesis presents UDTube, an easily usable software developed to perform morphological analysis in a multi-task fashion. This work shows the strong performance of UDTube versus the current state-of-the-art, UDPipe, across eight languages, primarily in the annotation of morphological features. The second contribution of this thesis is a exploration into the study of defectivity. UDTube is used to annotate a large amount of data in Greek and Russian which is ultimately used to investigate the plausibility of Indirect Negative Evidence (INE), a popular approach to the acquisition of morphological defectivity. The reported …


Consonant (De)Gradation In Ingrian?, Andrea M. Harrison 2024 CUNY Graduate Center

Consonant (De)Gradation In Ingrian?, Andrea M. Harrison

Dissertations, Theses, and Capstone Projects

This paper will present a dual method toward data enrichment for low-resource languages. Using Yoyodyne -- a Fairseq-inspired neural library for small-vocabulary sequence-to-sequence generation -- a morphological generation task was tested across labeled data encompassing multiple stages of enrichment for the low-resource language Ingrian. Due to limitations in the available data for Ingrian, weighted finite-state transducers (WFSTs) were used to generate an expanded vocabulary via HFST's toolkit for Uralic languages, and GiellaLT, a source for FST-driven lexica for low-resource languages. Further stages of experimentation used labeled data from related, higher-resource languages (Finnish, Estonian) to encourage cross-lingual transfer in the interest …


The Ring Cycle: Journeying Through The Language Of Tolkien’S Third Age With Corpus Linguistics, Michael Livesey 2024 University of Sheffield

The Ring Cycle: Journeying Through The Language Of Tolkien’S Third Age With Corpus Linguistics, Michael Livesey

Journal of Tolkien Research

This article explores the journey taken by the One Ring across J.R.R. Tolkien’s Third Age writings. It employs a digital humanities approach to analyse linguistic patterns in Tolkien’s use of the word ring, across The Hobbit and The Lord of the Rings. Specifically, the article employs corpus linguistic methods to track shifts in the quantities and qualities of the Ring’s appearance across these texts. It uses techniques of keyness and collocation analysis to trace transformations in these quantities/qualities, including: a) the Ring’s transition from a central to a peripheral place in the Third Age’s narrative arc; and b) …


Digital Commons powered by bepress