Open Access. Powered by Scholars. Published by Universities.®
- Institution
- Keyword
-
- Avatar (2)
- Emotion (2)
- HamNoSys (2)
- ISL (2)
- Language Acquisition (2)
-
- Machine Learning (2)
- NLP (2)
- AI (1)
- AI and linguistic equity (1)
- Academic writing (1)
- Accessibility (1)
- Algorithms (1)
- American Sign Language (1)
- Analysis (1)
- Análisis del discurso (1)
- Aphasia (1)
- Articulators (1)
- Artificial Intelligence (1)
- Automated Deception Detection (1)
- Automated fact-checking (1)
- Bengali NLP (1)
- Brows (1)
- CWE (1)
- Cantonese-English code-switching (1)
- Clickbait (1)
- Clickbait identification (1)
- Communication (1)
- Communication disorders (1)
- Computational linguistics. (1)
- Computer animation (1)
- Publication Year
- Publication
-
- Publications and Research (3)
- Other Resources (2)
- Conference Papers (1)
- Conference papers (1)
- Data and Test Instruments (1)
-
- Dissertations, Theses, and Capstone Projects (1)
- Journal of English and Applied Linguistics (1)
- Purdue Linguistics, Literature, and Second Language Studies Conference (1)
- Senior Honors Theses (1)
- Student Theses (1)
- Student Works (2010-2019) (1)
- Student Works (2020-2029) (1)
- Theses and Dissertations--Linguistics (1)
- Undergraduate Honors Theses (1)
- Visual Communications and Technology Education Faculty Publications (1)
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Computational Linguistics
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
Student Theses
The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …
A Machine Learning Approach To Disentangling Developmental Language Disorder From Typical Development In Russian-Speaking Children, Katsiaryna Aharodnik
A Machine Learning Approach To Disentangling Developmental Language Disorder From Typical Development In Russian-Speaking Children, Katsiaryna Aharodnik
Dissertations, Theses, and Capstone Projects
This study investigated a machine learning (ML) approach to identifying Developmental Language Disorder (DLD) in Russian-speaking children using narrative data. ML methods can capture subtle linguistic patterns that distinguish typical and atypical development, which is especially important in cross-linguistic contexts where morphosyntactic variation affects the manifestation of DLD. Diagnosis remains challenging in less-studied languages due to limited knowledge of language-specific deficits and a lack of validated assessment tools. This study evaluated whether ML algorithms can provide a more efficient alternative to traditional screening methods.
Two binary classification studies were conducted using corpus data: 1) classification of narratives told by Russian …
Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy
Structural Silence: When Ai Infrastructure Fails Speakers Of Underrepresented Languages, Avijit Roy, Proma Roy
Publications and Research
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools—training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures—encodes a set of assumptions that systematically disadvantages speakers of underrepresented languages before a single model is trained. This paper examines those assumptions through the lens of Bengali, one of the world’s most widely spoken languages with roughly 285 million speakers (Ethnologue, 2025; International Communication and Leadership School, 2026), and the structural barriers that emerge when attempting to build AI-assisted educational tools for Bengali-speaking learners in low-connectivity …
Predicting Language Proficiency Using A Multiple Regression Model, Madisen Barrieau
Predicting Language Proficiency Using A Multiple Regression Model, Madisen Barrieau
Senior Honors Theses
Businesses that design language learning products have a common goal to impart language proficiency to users. Many variables play a role in reaching language proficiency, including time spent in study, method of learning, use of technology, motivation, and more. This study involved creating several multiple regression models in R on a dataset featuring many of these variables. This research sought to identify the relative weight and measurability of these variables unto the goal of predicting language proficiency. Several regression models were created, and the best model showed that language similarity and length of residence in target culture, among other factors, …
Retórica Intercultural En El Discurso Académico Universitario: Las Funciones Retóricas De La Citación En Los Trabajos De Fin De Máster Escritos En Español Y En Inglés Por Hablantes Nativos Y No Nativos, David Sanchez-Jimenez
Retórica Intercultural En El Discurso Académico Universitario: Las Funciones Retóricas De La Citación En Los Trabajos De Fin De Máster Escritos En Español Y En Inglés Por Hablantes Nativos Y No Nativos, David Sanchez-Jimenez
Publications and Research
This research derives from the interest in learning the cultural differences in citation practices in the academic genre of Master's thesis of native Spanish (Ee), non-native Filipino writers of Spanish (Fe), native Filipino writers of English (Fi), and American writers of English. A total of thirty-two (32) master´s theses – eight (8) for each group – were analyzed. A quantitative and qualitative methodology was used to study this phenomenon based on the computerized textual analysis of the rhetorical function of citations arranged in typological classification that modified the outline proposed by Petrić in his 2007 article. The results obtained from …
The Sociolinguistics Of Code-Switching In Hong Kong’S Digital Landscape: A Mixed-Methods Exploration Of Cantonese-English Alternation Patterns On Whatsapp, Wilkinson Daniel Wong Gonzales, Yuen Man Tsang
The Sociolinguistics Of Code-Switching In Hong Kong’S Digital Landscape: A Mixed-Methods Exploration Of Cantonese-English Alternation Patterns On Whatsapp, Wilkinson Daniel Wong Gonzales, Yuen Man Tsang
Journal of English and Applied Linguistics
This paper examines the prevalence of Cantonese-English code mixing in Hong Kong through an under-researched digital medium. Prior research on this code-alternation practice has often been limited to exploring either the social or linguistic constraints of code-switching in spoken or written communication. Our study takes a holistic approach to analyzing code-switching in a hybrid medium that exhibits features of both spoken and written discourse. We specifically analyze the code-switching patterns of 24 undergraduates from a Hong Kong university on WhatsApp and examine how both social and linguistic factors potentially constrain these patterns. Utilizing a self-compiled sociolinguistic corpus as well as …
Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin
Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin
Data and Test Instruments
How do we detect, deter, and prevent the spread of mis- and disinformationwith the human eye and AI? How does theory inform the practice, and how do theevidence-based research and best practices in lie-catching and truth-seekingprofessions—inform AI? The book looks into well-established human practicessuch as the routines and processes used in detective work, journalism, and scientificinquiry, and how they contribute toward innovative AI solutions. The book explainsthe principles, inner workings, and recent evolution of five types of state-of-the-artAI technologies suitable for curtailing the spread of mis- and disinformation:automated deception detectors, clickbait detectors, satirical fake detectors, rumordebunkers, and computational fact-checking tools.
A Critical Genre Analysis Of The User Requirement Specification Report In The Information Technology Industry, Shahizan Shaharuddin
A Critical Genre Analysis Of The User Requirement Specification Report In The Information Technology Industry, Shahizan Shaharuddin
Student Works (2020-2029)
While it has been established that discursive practices involving writing are highly situated and context-bound, few local studies have examined the processes involved in the construction of texts and the shaping of the intended text. Recent perspectives on writing as ways of working and acting, and texts as social action suggest that studying the processes behind the construction of a text could provide a clearer understanding of the discursive practices involved in the shaping of the text. This is seen as relevant to professional writing context in the present study as texts are often shaped to facilitate the social action …
Non-Manual Articulators In Irish Sign Language Verbs: An Analysis With Data Mining Association Rules, Robert G. Smith, Markus Hofmann
Non-Manual Articulators In Irish Sign Language Verbs: An Analysis With Data Mining Association Rules, Robert G. Smith, Markus Hofmann
Conference Papers
The Signs of Ireland (SOI) corpus (Leeson et al., 2006) deploys a complex multi-tiered temporal data structure. The process of manually analyzing such data is laborious, cannot eliminate bias and often, important patterns can go completely unnoticed. In addition to this, as a result of the complex nature of grammatical structures contained in the corpus, identifying complex linguistic associations or patterns across tiers is simply too intricate a task for a human to carry out in an acceptable timeframe. This work explores the application of data mining techniques on a set of multi-tiered temporal data from the SOI corpus. Building …
Nevertheless, She Persisted: A Linguistic Analysis Of The Speech Of Elizabeth Warren, 2007-2017, Matthew Jennings
Nevertheless, She Persisted: A Linguistic Analysis Of The Speech Of Elizabeth Warren, 2007-2017, Matthew Jennings
Undergraduate Honors Theses
A breakout star among American progressives in the recent past, Elizabeth Warren has quickly gone from a law professor to a leading figure in Democratic politics. This paper analyzes Warren’s speech from before her time as a political figure to the present using the quantitative textual methodology established by Jones (2016) in order to see if Warren’s speech supports Jones’s assertion that masculine speech is the language of power. Ratios of feminine to masculine markers ultimately indicate that despite her increasing political sway, Warren’s speech becomes increasingly feminine instead. However, despite associations of feminine speech with weakness, Warren’s speech scores …
Towards A Computational Model Of Frame Of Reference Alignment In Swedish Dialogue, Simon Dobnik, Christine Howes, Kim Demaret, John D. Kelleher
Towards A Computational Model Of Frame Of Reference Alignment In Swedish Dialogue, Simon Dobnik, Christine Howes, Kim Demaret, John D. Kelleher
Conference papers
In this paper we examine how people negotiate, interpret and repair the frame of reference (FoR) in online text based dialogues discussing spatial scenes in Swedish. We describe work-in-progress in which participants are given different perspectives of the same scene and asked to locate several objects that are only shown on one of their pictures. This task requires participants to coordinate on FoR in order to identify the missing objects. This study has implications for situated dialogue systems.
Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan
Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan
Purdue Linguistics, Literature, and Second Language Studies Conference
The synthesized paper covers the topics of “corpus linguistics” and “language instruction and pedagogies”. I would like to do a presentation to highlight the key points in my paper.
The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner
The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner
Theses and Dissertations--Linguistics
This study examines the relationship between types of gestures and instances of self-initiated self-repair (SISR) used by persons with non-fluent aphasia (NFA), which is a type of aphasia characterized by stilted speech or signing (Papathanasiou et al., 2013), in interactions with clinicians. Conversation repairs in this study are assessed using the framework of Conversation Analysis (CA), which is an approach for describing, analyzing, and understanding social interaction (Sidnell, 2010). Previous linguistic studies have demonstrated a distinct preference for the use of gesture during a repair by persons with aphasia (Goodwin, 1995; Klippi, 2015; Wilkinson, 2013). This study draws more conclusive …
Emotional Facial Expressions In Synthesised Sign Language Avatars: A Manual Evaluation., Robert G Smith, Brian Nolan
Emotional Facial Expressions In Synthesised Sign Language Avatars: A Manual Evaluation., Robert G Smith, Brian Nolan
Other Resources
This research explores and evaluates the contribution that facial expressions might have regarding improved comprehension and acceptability in sign language avatars. Focusing specifically on Irish sign language (ISL), the Deaf (the uppercase ‘‘D’’ in the word ‘‘Deaf’’ indicates Deaf as a culture as opposed to ‘‘deaf’’ as a medical condition) community’s responsiveness to sign language avatars is examined. The hypothesis of this is as follows: augmenting an existing avatar with the seven widely accepted universal emotions identified by Ekman (Basic emotions: handbook of cognition and emotion. Wiley, London, 2005) to achieve underlying facial expressions will make that avatar more human-like …
The Role Of Emotion And Facial Expression In Synthesised Sign Language Avatars, Robert G Smith
The Role Of Emotion And Facial Expression In Synthesised Sign Language Avatars, Robert G Smith
Other Resources
This thesis explores the role that underlying emotional facial expressions might have in regards to understandability in sign language avatars. Focusing specifically on Irish Sign Language (ISL), we examine the Deaf community’s requirement for a visual-gestural language as well as some linguistic attributes of ISL which we consider fundamental to this research. Unlike spoken language, visual-gestural languages such as ISL have no standard written representation. Given this, we compare current methods of written representation for signed languages as we consider: which, if any, is the most suitable transcription method for the medical receptionist dialogue corpus. A growing body of work …
Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez
Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez
Publications and Research
The aim of this paper is to compare the rhetorical functions gathered from the citations of (14) fourteen master´s theses written by seven Spanish and seven Philippine authors. A typology of nine categories was used in order to identify the cultural rhetorical differences that exist in the use of citation from the contrast between contrasting this element in the Philippine and Spanish cultures. The methodology used is textual analysis of the linguistic context of these citations and its subsequent classification within these nine categories. The results show that there are quantitative and qualitative differences between the cultural conventions of citations …
Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. Mcdonald, Jerry Schnepp
Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. Mcdonald, Jerry Schnepp
Visual Communications and Technology Education Faculty Publications
Computer-generated three-dimensional animation holds great promise for synthesizing utterances in American Sign Language (ASL) that are not only grammatical, but well tolerated by members of the Deaf community. Unfortunately, animation poses several challenges stemming from the necessity of grappling with massive amounts of data. However, the linguistics of ASL can aid in surmounting the challenge by providing structure and rules for organizing animation data. An exploration of the linguistic and extra linguistic behavior of the brows from an animator’s viewpoint yields a new approach for synthesizing nonmanuals that differs from the conventional animation of anatomy and instead offers a different …
A Computer-Aided Error Analysis Of Multi-Word Units In A Malaysian Learner English Corpus., Lee Yit Sim
A Computer-Aided Error Analysis Of Multi-Word Units In A Malaysian Learner English Corpus., Lee Yit Sim
Student Works (2010-2019)
This is a study on multi-word unit (MWU) errors made by Malaysian English learners. The major approaches used in the analysis of learners’ writings are: error analysis and corpus-based approach. The findings from this study reveal that learners have problems with these MWUs: modal structures , infinitive structures , ‘adjective + noun’ collocations , and connectors . An analysis of these structures shows that the common erroneous patterns in these structures can generally be categorized as ‘overgeneralization’, ‘misformation’, ‘misselection’, and ‘distortion’. The causes of such errors are both interlingual as well as intralingual. Besides these factors, this study also highlighted …