Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

2024

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 20 of 20

Full-Text Articles in Computational Linguistics

Modeling Context And The Characteristica Universalis, John Kausch Dec 2024

Modeling Context And The Characteristica Universalis, John Kausch

Proceedings from the Document Academy

This paper proposes prototypes for the exploration of the context of terms in a knowledge organization system by visualizing machine learning produced word embeddings. It puts this work in the context of the search for a universal language, typified by Leibniz’s characteristica universalis. This tradition of the search for universal languages is put in the context of universalizing tendencies in taxonomic classification in library and information science. Following this there is a discussion of the use of machine learning models to represent context. These two concerns inform the construction of prototypes for exploring the contextual spaces produced by word embeddings …


Controlling Emotional Text To Speech Using Complex Adverbial Phrases, Zainab T. Akande Sep 2024

Controlling Emotional Text To Speech Using Complex Adverbial Phrases, Zainab T. Akande

Dissertations, Theses, and Capstone Projects

This study investigates the usage of adverbial modifiers from audiobook data as a resource for training speech synthesizers with a greater range of speech descriptions. The Tacotron2 text-to-speech (TTS) model was used for the purposes of this study. Utilizing the LibriTTS dataset, the Tacotron2 model is trained under two experimental conditions: one incorporating adverbial modifiers into the input text and the other without. The dataset preprocessing involves embedding descriptions using word embeddings and encoding speaker IDs with machine learning techniques. Additionally, the model architecture includes a prosody encoder inspired by prior research. Evaluation of the trained models involves subjective assessments …


Predicting Language Proficiency Using A Multiple Regression Model, Madisen Barrieau Jun 2024

Predicting Language Proficiency Using A Multiple Regression Model, Madisen Barrieau

Senior Honors Theses

Businesses that design language learning products have a common goal to impart language proficiency to users. Many variables play a role in reaching language proficiency, including time spent in study, method of learning, use of technology, motivation, and more. This study involved creating several multiple regression models in R on a dataset featuring many of these variables. This research sought to identify the relative weight and measurability of these variables unto the goal of predicting language proficiency. Several regression models were created, and the best model showed that language similarity and length of residence in target culture, among other factors, …


Uncovering The Mimicry Of Online Review Breadth And Depth And Its Subsequent Effect On Consumer Responses, Andrea Pelaez Martinez Jun 2024

Uncovering The Mimicry Of Online Review Breadth And Depth And Its Subsequent Effect On Consumer Responses, Andrea Pelaez Martinez

Dissertations, Theses, and Capstone Projects

Word-of-mouth (WOM) in marketing occurs when consumers discuss a company's product or service or any consumption experience with their friends, family, and others with whom they have any relationship. With the advent of social media, this phenomenon has expanded rapidly into virtual environments where consumer conversation is enabled through chats, forums, social media posts, and online reviews. In response to this rapid growth of online WOM, academics and practitioners have focused their interest on this phenomenon and its implications on consumers, firms, and society. So far, the evidence of the critical role that online WOM plays in helping consumers make …


Computational Approaches To Linguistic Challenges In Arabic Speech Recognition, Enas Albasiri Jun 2024

Computational Approaches To Linguistic Challenges In Arabic Speech Recognition, Enas Albasiri

Dissertations, Theses, and Capstone Projects

This dissertation aims to document the linguistic features of Arabic that pose challenges to speech and language technologies and advance these technologies by developing state-of-the-art computational tools focusing on automatic speech recognition (ASR), text normalization (TN), and corpus development. TN converts expressions such as numbers, dates, and times—named semiotic classes—from their written to their spoken domain, such as converting ‘$84.00’ to ‘eighty-four dollars’, while inverse text normalization (ITN) converts verbalized text to its written form. This conversion is an essential preprocessing step for text-to-speech (TTS), and post-processing step for ASR. Arabic presents a challenge for TN and ITN because one …


Expanding The Corpus Of Vocalized Hebrew Text: Compiling An Unvocalized Text Corpus And Building An Online Interface For Vocalization Annotation, Rachel Shanblatt Bloch Jun 2024

Expanding The Corpus Of Vocalized Hebrew Text: Compiling An Unvocalized Text Corpus And Building An Online Interface For Vocalization Annotation, Rachel Shanblatt Bloch

Dissertations, Theses, and Capstone Projects

Written modern Hebrew presents a unique challenge for training computational models for language processing because modern Hebrew text often lacks vocalization. The lack of available vocalized Hebrew data can lead to ambiguity in training these models and generally hinders work on natural language processing problems. The goal of this project is to contribute to the collection of vocalized Hebrew text by collecting and preprocessing a large corpus of unvocalized Hebrew text and building an online annotation tool. The annotation tool allows people to upload unvocalized Hebrew text, to annotate by adding Hebrew vocalization, and to download comma-separated values files of …


Gpt Assisted Annotation Of Rhetorical And Linguistic Features For Interpretable Propaganda Technique Detection In News Text., Kyle Hamilton, Bojan Bozic, Luca Longo May 2024

Gpt Assisted Annotation Of Rhetorical And Linguistic Features For Interpretable Propaganda Technique Detection In News Text., Kyle Hamilton, Bojan Bozic, Luca Longo

Articles

While the use of machine learning for the detection of propaganda techniques in text has garnered considerable attention, most approaches focus on "black-box'' solutions with opaque inner workings. Interpretable approaches provide a solution, however, they depend on careful feature engineering and costly expert annotated data. Additionally, language features specific to propagandistic text are generally the focus of rhetoricians or linguists, and there is no data set labeled with such features suitable for machine learning. This study codifies 22 rhetorical and linguistic features identified in literature related to the language of persuasion for the purpose of annotating an existing data set …


Exploring Asynchronous Pronunciation Training Through Context-Aware Pronunciation Applications, Claire L. Schweikert May 2024

Exploring Asynchronous Pronunciation Training Through Context-Aware Pronunciation Applications, Claire L. Schweikert

Theses/Capstones/Creative Projects

This paper provides a survey of various research articles on context-aware asynchronous pronunciation training applications. First, a set of seven articles is reviewed and summarized. Next, they are synthesized over the three main topics of 1) automated speech recognition, 2) non-native speaker considerations in language learning, and 3) future directions for research and development within computer-assisted pronunciation training (CAPT). Research in the areas of acoustic and pronunciation modeling (both implicit and explicit), pedagogical considerations for CAPT application design, Goodness of Pronunciation algorithm scoring, accent recognition and neutralization, and more are discussed.


Bugsy The Bee's Big Adventure, Jonah L., Zachery Irvin, Jake Wildstrom, Zachariah Robinson, Hilaria Cruz Apr 2024

Bugsy The Bee's Big Adventure, Jonah L., Zachery Irvin, Jake Wildstrom, Zachariah Robinson, Hilaria Cruz

LING 590/Internet Language

Bugsy the bee goes on a quest for pollen.


Skyler's Lunch, Noah Sherman, Autumn Boone, Hilaria Cruz Apr 2024

Skyler's Lunch, Noah Sherman, Autumn Boone, Hilaria Cruz

LING 590/Internet Language

Our class was studying the use of emojis across different platforms and wanted to explore how stories using emojis could impact young readers. Here, we try to translate the story of Skyler into emoji, providing translations along the way. We replace words completely with emoji, represent phrases with a few emoji, and use additional emoji to make sense of the content, including punctuation. In this book, we explore the character of Skyler, who is a picky eater. But they learn to eat the nutritious food that is good for them. In the end, they even get a reward!


Research On The Application Of Part-Of-Speech Tagging Of Ancient Books Under The Domain Large Language Model, Danhao Zhu, Zhao Zhixiao, Die Hu, Wenhua Zhao Apr 2024

Research On The Application Of Part-Of-Speech Tagging Of Ancient Books Under The Domain Large Language Model, Danhao Zhu, Zhao Zhixiao, Die Hu, Wenhua Zhao

Journal of Scientific Information Research

[Purpose/significance]The development of the large language model has brought new ideas for ancient text mining, and combining the large language model with the digitisation and intelligence of ancient books is a necessary path for the work of ancient books in the new era. [Methods/process]This paper uses the lexically annotated corpus of Zuozhuan to construct a batch of high-quality lexically annotated instruction data through data cleaning and preprocessing, on the basis of which 500, 1 000, 2 000, and 5 000 pieces of data are used to fine-tune the instructions of the large language model, and the performance test is carried …


Retórica Intercultural En El Discurso Académico Universitario: Las Funciones Retóricas De La Citación En Los Trabajos De Fin De Máster Escritos En Español Y En Inglés Por Hablantes Nativos Y No Nativos, David Sanchez-Jimenez Feb 2024

Retórica Intercultural En El Discurso Académico Universitario: Las Funciones Retóricas De La Citación En Los Trabajos De Fin De Máster Escritos En Español Y En Inglés Por Hablantes Nativos Y No Nativos, David Sanchez-Jimenez

Publications and Research

This research derives from the interest in learning the cultural differences in citation practices in the academic genre of Master's thesis of native Spanish (Ee), non-native Filipino writers of Spanish (Fe), native Filipino writers of English (Fi), and American writers of English. A total of thirty-two (32) master´s theses – eight (8) for each group – were analyzed. A quantitative and qualitative methodology was used to study this phenomenon based on the computerized textual analysis of the rhetorical function of citations arranged in typological classification that modified the outline proposed by Petrić in his 2007 article. The results obtained from …


Consonant (De)Gradation In Ingrian?, Andrea M. Harrison Feb 2024

Consonant (De)Gradation In Ingrian?, Andrea M. Harrison

Dissertations, Theses, and Capstone Projects

This paper will present a dual method toward data enrichment for low-resource languages. Using Yoyodyne -- a Fairseq-inspired neural library for small-vocabulary sequence-to-sequence generation -- a morphological generation task was tested across labeled data encompassing multiple stages of enrichment for the low-resource language Ingrian. Due to limitations in the available data for Ingrian, weighted finite-state transducers (WFSTs) were used to generate an expanded vocabulary via HFST's toolkit for Uralic languages, and GiellaLT, a source for FST-driven lexica for low-resource languages. Further stages of experimentation used labeled data from related, higher-resource languages (Finnish, Estonian) to encourage cross-lingual transfer in the interest …


How Do We Learn What We Cannot Say?, Daniel Yakubov Feb 2024

How Do We Learn What We Cannot Say?, Daniel Yakubov

Dissertations, Theses, and Capstone Projects

The contributions of this thesis are two-fold. First, this thesis presents UDTube, an easily usable software developed to perform morphological analysis in a multi-task fashion. This work shows the strong performance of UDTube versus the current state-of-the-art, UDPipe, across eight languages, primarily in the annotation of morphological features. The second contribution of this thesis is a exploration into the study of defectivity. UDTube is used to annotate a large amount of data in Greek and Russian which is ultimately used to investigate the plausibility of Indirect Negative Evidence (INE), a popular approach to the acquisition of morphological defectivity. The reported …


The Ring Cycle: Journeying Through The Language Of Tolkien’S Third Age With Corpus Linguistics, Michael Livesey Jan 2024

The Ring Cycle: Journeying Through The Language Of Tolkien’S Third Age With Corpus Linguistics, Michael Livesey

Journal of Tolkien Research

This article explores the journey taken by the One Ring across J.R.R. Tolkien’s Third Age writings. It employs a digital humanities approach to analyse linguistic patterns in Tolkien’s use of the word ring, across The Hobbit and The Lord of the Rings. Specifically, the article employs corpus linguistic methods to track shifts in the quantities and qualities of the Ring’s appearance across these texts. It uses techniques of keyness and collocation analysis to trace transformations in these quantities/qualities, including: a) the Ring’s transition from a central to a peripheral place in the Third Age’s narrative arc; and b) …


Using Ai For Qualitative Labeling: Consistency And Comparisons, James Temple Jan 2024

Using Ai For Qualitative Labeling: Consistency And Comparisons, James Temple

Honors Program Theses

This paper continues research that evaluates the capacity of artificial intelligence (AI) to perform qualitative coding tasks. The previous study found that AI models lacked consistency with themselves and did not agree with human coded data. Since that study, AI’s general level of intelligence has increased. Hence, this study re-evaluates how well the newest set of AI models (Claude 3 and Gemini) can perform qualitative coding tasks. When tested, the new AI models perform about the same or better than previous models depending on the metric tested. While Gemini and Claude 3 do not agree with human output any more …


Linguistic Inquiry And Word Count (Liwc) For Text Analysis, Isabella Lenzini, Destiny Fore Msw, Anna W. Wright Phd Jan 2024

Linguistic Inquiry And Word Count (Liwc) For Text Analysis, Isabella Lenzini, Destiny Fore Msw, Anna W. Wright Phd

IRBEH/Spit for Science Publications and Presentations

No abstract provided.


Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri Jan 2024

Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri

Civil Engineering Dissertations - Archive

Infrastructure projects impact a broad range of stakeholders, particularly local communities, whose engagement is critical for successful outcomes. Despite the importance of public engagement in these projects, traditional methods of capturing and analyzing public opinion often fail to fully represent the diverse, genuine perspectives involved. This has led to conflicts between community members and project sponsors. On the other hand, despite advancements in text analytics, including natural language processing (NLP) and its subfields such as topic modeling, sentiment analysis, and neural networks, their functionalities and effectiveness in analyzing public opinion in the domain of infrastructure projects have not been fully …


A Computational Investigation Of English Spelling, John Winstead Jan 2024

A Computational Investigation Of English Spelling, John Winstead

Theses and Dissertations--Linguistics

This thesis examines the predictability and regularity of English orthography through computational methods. The primary objective is to use n-gram models to predict missing letters in English words by exploiting contextual information from adjacent letters. The study evaluates the impact of dataset size, word length, letter position, and vowel presence on the predictive accuracy of these models, uncovering patterns and structures inherent to English spelling.

The research utilizes a range of datasets, including the Carnegie Mellon University Pronouncing Dictionary, the Brown Corpus, the Corpus of Late Modern English Texts, the Lampeter Corpus of Early Modern English Tracts, and the Open …


A Computer-Assisted Approach To Lexical Borrowing In Northeast Caucasian Languages, Bonnie Eleanor Wren-Hardin Jan 2024

A Computer-Assisted Approach To Lexical Borrowing In Northeast Caucasian Languages, Bonnie Eleanor Wren-Hardin

Theses and Dissertations--Linguistics

The disambiguation of loanwords and cognates can be a challenge, especially in areas where there has been intense language contact over an extended period of time, when the contact is between genetically related languages, and when the number of languages involved is large Over the past several decades, more and more computational approaches to automatic cognate and borrowing detection have been created in an attempt to ease the load of examining hundreds to thousands of individual lexemes, as well as determine language family relationships with allegedly greater accuracy. While these methods are not perfect and cannot replace the knowledge or …