Open Access. Powered by Scholars. Published by Universities.®

Large language models

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 109

Full-Text Articles in Artificial Intelligence and Robotics

Leo’S Choices, Katherine G. Schmidt Aug 2026

Leo’S Choices, Katherine G. Schmidt

The Journal of Social Encounters

No abstract provided.


Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo Aug 2026

Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …


Can Machines Testify? Llms And The Boundaries Of Testimonial Epistemology, Michael J. Cummins Jul 2026

Can Machines Testify? Llms And The Boundaries Of Testimonial Epistemology, Michael J. Cummins

Philosophy Summer Fellows

As Large Language Models and AI chatbots become increasingly prevalent, pressing questions are raised about whether beliefs formed through LLM interactions carry the same epistemic weight as beliefs formed through human testimony. How we answer this question depends on whether LLMs can function as testifiers, a role which is typically assumed to require a human or human-like agent. This assumption has gone largely unexamined, yet its consequences are significant: if LLM outputs cannot constitute testimony, then the justificatory tools of testimonial epistemology are unavailable to any beliefs formed through LLM interaction. This paper challenges that assumption. It first argues that …


Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi Jul 2026

Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi

College of Population Health Faculty Papers

BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.

METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence …


Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto Jul 2026

Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto

School of Accountancy Faculty Publications

This study examines whether large language models exhibit systematic country-contingent differential treatment in financial fraud detection. Analyzing 30,000 synthetic transactions with identical statistical properties across three country attributions (United States, Great Britain, and China), we find LLMs assign significantly higher fraud probabilities to Chinese-attributed transactions (36.2%) compared to Western countries (≈30–31%), resulting in accuracy disparities of 67% versus 74%. The gap remains stable across five independent experimental replications and persists when using Chinese language prompts, ruling out linguistic effects. Bias mitigation strategies, such as requiring explanations or explicit country neutrality instructions, reduce but fail to eliminate these disparities. Testing across …


Semantic Context Improvisational Retrieval-Augmented Generation For Empathic Conversational Ai, Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah Jul 2026

Semantic Context Improvisational Retrieval-Augmented Generation For Empathic Conversational Ai, Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah

Research outputs 2022 to 2026

A fundamental limitation of modern conversational AI is its limited capacity to demonstrate sustained empathy in long-form interactions. We propose SCIRAG (Semantic Context Improvisational Retrieval-Augmented Generation), a feedback-driven retrieval framework for adaptive empathic dialogue. It employs a dual-loop retrieval framework, iteratively optimizing a static counseling dataset through user metadata and feedback memory refinement. To enhance contextual alignment, we deploy retrieval adaptation, enabling the model to retain and leverage past conversational cues based on user preferences. When integrated with Mixtral-8x7B, SCIRAG improves human-rated empathic understanding by +1.26 points and empathic response by +1.00 point on the RoPE scale, while increasing acceptability …


Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo Jul 2026

Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo

Research Collection School Of Computing and Information Systems

Learning-based dynamic fault localization techniques play a crucial role in the field of software engineering. These techniques dynamically execute test cases to meticulously extract useful knowledge from the execution information in the program, with the aim of identifying fault locations by leveraging machine learning, deep learning, and large language models. Currently, there is already a flourishing body of research that is intensely focused on learning-based dynamic fault localization. Research literature can be categorized into two main aspects for learning-based dynamic fault localization: data-based enhancements (i.e., the datasets) and model-based enhancements (i.e., the suspiciousness algorithms). Thus, we conduct an extensive literature …


Where Evidence-Based Medicine Meets Ai: Promise, Pitfalls, And Practice, Sangil Lee, Joshua Davis, Ken Milne, Christina Shenvi, Lars K. Beattie, Martin Wegman, Laura Melville, Richard D. Shih, Bryan Kane Jun 2026

Where Evidence-Based Medicine Meets Ai: Promise, Pitfalls, And Practice, Sangil Lee, Joshua Davis, Ken Milne, Christina Shenvi, Lars K. Beattie, Martin Wegman, Laura Melville, Richard D. Shih, Bryan Kane

Department of Emergency Medicine Faculty Papers

No abstract provided.


Capturing Large Language Model Similarity Through Spectral Analysis, Ishan Verma Prasad Jun 2026

Capturing Large Language Model Similarity Through Spectral Analysis, Ishan Verma Prasad

Computer Science Senior Theses

With the rapid development of open-sourced models on Huggingface, there is a strong need for a way to systematically determine the similarity between models. More strongly, for intellectual property and organization, we need a way to determine the "lineage" of models. We borrow principles from Heavy-Tailed Self-Regularization and Random Matrix Theory to provide an inference-free method to accomplish this. We cluster a corpus of several model families by their spectral fingerprints and demonstrate that each model family occupies a distinct region in weight space. This confirms prior ideas of training setups leaving artifacts on model weights and allows us to …


Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez Jun 2026

Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez

Master's Theses

Legislators frequently discuss the same policy issues across multiple hearings and legislative sessions, sometimes maintaining consistent positions and other times modifying or reframing their stance over time. Understanding how these positions evolve is important for analyzing political discourse and democratic accountability, yet identifying such shifts at scale remains difficult.

We introduce TRACE (Temporal Rhetorical Analysis and Consistency Evaluation), a system built on the Digital Democracy Database (DDDB) for detecting rhetorical inconsistency in California legislative hearing testimony. TRACE organizes utterances into speaker-anchored timelines indexed by bill and session, then applies hybrid semantic retrieval — combining dense BGE embeddings with BM25 lexical …


Scorespeak: An Agentic System For Natural Language Control Of Musical Scores, Nathan S. Lim Jun 2026

Scorespeak: An Agentic System For Natural Language Control Of Musical Scores, Nathan S. Lim

Master's Theses

With the recent popularization of large language models (LLMs), natural language has become one of the most accessible and powerful ways for people to interact with creative tools. Although they have become common in mainstream domains like image and audio editing, there is currently no robust AI-based system that can reliably turn free-form language into edits for symbolic musical scores. This gap represents a missed opportunity to improve human workflows for creating and editing sheet music, but it is also a fundamental limitation for other agentic music systems; without a robust mechanism for translating free-form language into structured scores, AI …


An Automated Generation Method For Combat Simulation Scenarios Based On Large Language Models, Zhiming Dong, Zhongqi Hu, Haoran Dai, Jiancheng Gao May 2026

An Automated Generation Method For Combat Simulation Scenarios Based On Large Language Models, Zhiming Dong, Zhongqi Hu, Haoran Dai, Jiancheng Gao

Journal of System Simulation

To address the issue of low efficiency in generating traditional army tactical combat simulation scenarios, an automated generation method based on large language models is proposed. The large language model invokes a semantic segmentation algorithm to parse and restructure the combat scenario, forming semantic modules. Utilizing a multi-agent collaborative framework based on the model contextual protocol, the large language model drives each agent to extract simulation elements from the corresponding semantic modules, constructing a knowledge graph of scenario elements. Using this knowledge graph as a retrieval medium, the method employs a dense retrieval algorithm to achieve precise matching between simulation …


Synthergy: Social Deduction And Deception In Llm-Powered Agents, Lauren Campbell, Andrew Forney May 2026

Synthergy: Social Deduction And Deception In Llm-Powered Agents, Lauren Campbell, Andrew Forney

Honors Thesis

Synthergy is an online social deduction game designed to enable comparative analysis of how large language model-powered agents engage in social deduction and deception under conditions of asymmetric information. Inspired by social deduction games such as Town of Salem, Throne of Lies, and Mafia, the game consists of two factions, Harmony and Discord, to which agents are secretly assigned. Agents must infer others’ affiliations through dialogue, in-game abilities, and voting behavior. To evaluate agent behavior, we conducted 100 simulated games across six agent types: a random baseline agent (RandomSynth), an LLM-based agent (Synth), a chain-of-thought agent (CoT Synth), a Bayesian …


Llms In Compiler Construction, Raffi Khatchadourian May 2026

Llms In Compiler Construction, Raffi Khatchadourian

Open Educational Resources

These lecture slides survey the use of large language models (LLMs) in compiler construction for a graduate compiler course (CSc 81010). They situate LLMs across the compiler pipeline and examine representative work: foundation models trained on LLVM IR and assembly (Meta's LLM Compiler), LLM-driven code optimization, binary decompilation (LLM4Decompile), and LLM-assisted automated refactoring—alongside the challenges of applying probabilistic models to tasks that demand correctness. The slides are a self-contained HTML (W3C Slidy) deck with editable Pandoc Markdown source. Part of a two-session unit on advanced compiler topics; see also "Deep Learning Compilers."


Probing Representational Emergence In Large Language Models, Shawn Ismail May 2026

Probing Representational Emergence In Large Language Models, Shawn Ismail

Master's Theses

This thesis investigates whether abrupt behavioral gains in large language models under scaling are accompanied by systematic changes in internal representations. It combines a behavioral screen of 65 tasks per family with targeted layerwise probing across eight decoder-only, open-weight model families. Behavioral emergence is defined for each family-task trajectory using an empirical jump detector, with segmented regression retained only as a diagnostic. The representational follow-up analyzes 27 selected MMLU subtasks shared across all families, spanning 37 checkpoints and 216 family-task units.

For each follow-up checkpoint, frozen linear probes are trained on every layer's hidden states to measure how much task-relevant …


Study Of Output And Behavior Of Llms Using Confidence Framing In Prompt Engineering, Micah Parrilla Apr 2026

Study Of Output And Behavior Of Llms Using Confidence Framing In Prompt Engineering, Micah Parrilla

Doctoral Dissertations and Master's Theses

While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.

The findings identify a distinct cognitive …


Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang Apr 2026

Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang

Research Collection School Of Computing and Information Systems

Since the launch of ChatGPT, a powerful AI Chatbot developed by OpenAI, large language models (LLMs) have made significant advancements in both academia and industry, bringing about a fundamental engineering paradigm shift in many areas. While LLMs are powerful, it is also crucial to best use their power where “prompt” plays a core role. However, the booming LLMs themselves, including excellent APIs like ChatGPT, have several inherent limitations: (1) temporal lag of training data, and (2) the lack of physical capabilities to perform external actions. Recently, we have observed the trend of utilizing prompt-based tools to better utilize the power …


Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng Apr 2026

Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng

Research Collection Lee Kong Chian School Of Business

Existing research on AI cultural biases predominantly focuses on Western models, overlooking critical gaps in non-Western models. We conduct a comparative analysis of AI models – ChatGPT (U.S. developed) and ErnieBot (China developed) – from different cultures to investigate how corresponding cultural biases manifest in their outputs. Additionally, we examine how cultural alignment between human users and AI models impacts their collaborative creative performance and the underlying psychological mechanisms. In Study 1, multi-choice prompt with zero-shot technique was used to evaluate cultural biases in four widely used AI models – ChatGPT-3.5/4, ErnieBot-3.5/4 – comparing their responses to established cultural psychometric …


When Ai Writes The Doctoral Thesis: Reclaiming The Oral Defence As A Learning Development Intervention, Valerie A. Storey Mar 2026

When Ai Writes The Doctoral Thesis: Reclaiming The Oral Defence As A Learning Development Intervention, Valerie A. Storey

All Faculty and Staff Scholarship

Large language models have fundamentally challenged traditional methods of verifying doctoral competency as AI-generated text becomes increasingly difficult to distinguish from human scholarship. This paper argues that thesis committees and doctoral supervisors must reclaim the oral defence as a critical checkpoint for assessing authentic threshold crossing rather than a ceremonial rite of passage. Drawing on historical examples from medieval oral disputations through to the rise of written theses, this paper asserts the necessity of returning to rigorous oral assessment. Given the limitations of detection technologies and the growing use of AI in thesis writing, oral defences must move from confirmatory …


Why General Ai Inherited The Body: A Structural Account Of Embodied Modulation Under Monolithic Imitation, Griselda Poe Mar 2026

Why General Ai Inherited The Body: A Structural Account Of Embodied Modulation Under Monolithic Imitation, Griselda Poe

Publications and Research

General AI has pursued the replication of human-level intelligence without first decomposing human cognition into structurally distinct components. Human cognition, however, is shaped by embodied constraints such as mortality, survival pressures, finite lifespan, and physiological states. This paper argues that when cognition is treated as a single undifferentiated whole, embodied modulation is not accidentally introduced into AI systems but structurally entailed. Any attempt to imitate “human intelligence” under a monolithic model necessarily incorporates variability shaped by mortal embodiment. The tensions observed in contemporary AI systems are better understood as consequences of copying an undecomposed target rather than isolated implementation errors. …


The Illusion Of Causality In Llms: A Developmentally Grounded Analysis Of Semantic Scaffolding And Benchmark–Capability Mismatches, Daisuke Akiba Mar 2026

The Illusion Of Causality In Llms: A Developmentally Grounded Analysis Of Semantic Scaffolding And Benchmark–Capability Mismatches, Daisuke Akiba

Publications and Research

Recent benchmarks increasingly report that large language models (LLMs) exhibit human-like causal reasoning abilities, including counterfactual inference and intervention planning. However, many such evaluations rely on domains that are heavily represented in training data and embed strong semantic cues, raising the possibility that apparent causal competence may reflect semantic pattern recombination rather than structure-sensitive causal reasoning. Drawing on human developmental theories of causal induction, this perspective argues that genuine causal understanding requires robustness to novelty and reliance on conditional structure rather than semantic familiarity. To illustrate the testability of this claim, the paper includes a pilot demonstration using synthetic causal …


Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui Mar 2026

Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui

Research Collection School Of Computing and Information Systems

Evaluating the alignment of large language models (LLMs) with user-defined coding preferences is a challenging endeavor that requires a deep assessment of LLMs' outputs. Existing methods and benchmarks rely primarily on automated metrics and static analysis tools, which often fail to capture the nuances of user instructions and LLM outputs. To address this gap, we introduce the LLM-as-a-Judge evaluation framework and present CodeUltraFeedback, a comprehensive dataset for assessing and improving LLM alignment with coding preferences. CodeUltraFeedback consists of 10,000 coding instructions, each annotated with four responses generated from a diverse pool of 14 LLMs. These responses are annotated using GPT-3.5 …


Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell Feb 2026

Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell

Proceedings from the Document Academy

Generative Artificial Intelligences (AIs) and current advanced large language models (LLMs) are algorithmically designed to generate text-based conversations as conversational agents (CAs), by replicating human language and conversational communication. Pairing human cognition with generative computationally coded cognition. We have never been here before: cerebral and artificial information collaborations and processing producing expressions that may or may not become visible as second-hand/secondary source documents.

Sensemaking or sense(un)making is a unique autonomous human drive cognitively, our information processing is sensemaking in action and expressions and articulations are evidence of the sensemaking cycle. Documentation [expressed or articulated through various mediums] are a product …


The Gains Do Not Make Up For The Losses: A Comprehensive Evaluation For Safety Alignment Of Large Language Models Via Machine Unlearning, Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che Feb 2026

The Gains Do Not Make Up For The Losses: A Comprehensive Evaluation For Safety Alignment Of Large Language Models Via Machine Unlearning, Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che

Research Collection School Of Computing and Information Systems

Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased, leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark MuBench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful …


Evaluation Of Large Language Models As Decision Support Tools For Head And Neck Cancer Management: A Blinded Multidisciplinary Simulation Study, Sholem Hack, Ron J. Karni, Antonino Maniaci, Christopher E. Fundakowski, Luca Castellani, Fabiola Incandela, Remo Accorona, Miguel Mayo-Yanez, Martina Violati, Lorenzo Giannini, Niccolo' Mevio, Alberto Maria Saibene Jan 2026

Evaluation Of Large Language Models As Decision Support Tools For Head And Neck Cancer Management: A Blinded Multidisciplinary Simulation Study, Sholem Hack, Ron J. Karni, Antonino Maniaci, Christopher E. Fundakowski, Luca Castellani, Fabiola Incandela, Remo Accorona, Miguel Mayo-Yanez, Martina Violati, Lorenzo Giannini, Niccolo' Mevio, Alberto Maria Saibene

Department of Otolaryngology - Head and Neck Surgery Faculty Papers

BACKGROUND: The management of head and neck cancer relies on multidisciplinary expertise; however, access to tumor boards remains variable. Large language models (LLMs) may support guideline-based decision-making, although performance in complex oncologic scenarios is not well defined.

METHODS: Fourteen synthetic cases based on real tumor board encounters were evaluated. Five blinded comparator arms produced recommendations: a human expert, Non-RAG-GPT-4, Non-RAG-GPT-5, RAG-GPT-4, and RAG-GPT-5. Eight head and neck oncologic surgeons scored each recommendation for appropriateness, clarity, specificity, and feasibility using 5-point Likert scales. Paired permutation testing and inter-rater reliability were assessed.

RESULTS: LLM outputs showed close alignment with expert recommendations. RAG-based …


Large Language Model Communication And Data Transfer Across A Simulated Telephone Line, Jared Reyes Jan 2026

Large Language Model Communication And Data Transfer Across A Simulated Telephone Line, Jared Reyes

Dissertations and Theses

This thesis presents a proof-of-concept communication system in which two local large language model endpoints communicate across a simulated analog telephone line using legacy USB modems. The project combines modem voice mode, modem data mode, local speech processing, and structured machine messaging into a single staged session. During the voice phase, one endpoint places a call, the other answers, speech generated by a locally hosted LLaMA-family model is synthesized with Piper, transmitted through the modem voice path, captured on the remote side, and transcribed with Faster-Whisper to drive the next response. After the voice exchange, the system transitions to a …


Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza Jan 2026

Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza

All Graduate Theses, Dissertations, and Other Capstone Projects

With healthcare systems under growing pressure from rising patient volumes and shrinking consultation windows, improving how patients communicate with physicians has become essential to delivering quality care. Yet patients routinely arrive at appointments unable to clearly describe their symptoms, recall their medical history, or articulate concerns, contributing to miscommunication, diagnostic inefficiency, and pre-visit anxiety. This study introduces PreVisit AI, a conversational system designed to address this gap through structured, knowledge-based patient preparation. The system is built on a Retrieval-Augmented Generation (RAG) architecture combining HuggingFace sentence embeddings (all-MiniLM-L6-v2), a Chroma vector store, and Google’s Gemini language model over a curated seven-document …


Ablative Study Of Large Language Model-Based Gesture Inference For Autonomous Navigation, Neil Loftus Jan 2026

Ablative Study Of Large Language Model-Based Gesture Inference For Autonomous Navigation, Neil Loftus

Theses, Dissertations and Capstones

Human gesture inference has broad applications ranging from sign language interpretation to device control. Traditional methods often rely on extensive manually labeled hand datasets for deep learning. Furthermore, they are typically limited to a discrete set of gestures existing in these datasets. Large Language Models (LLMs) created by enterprise companies such as OpenAI have demonstrated positive results in many artificial intelligence tasks, with a notable strength being their adaptability. Existing literature has shown that LLM based systems can not only perform gesture inference but can propose user intent provided with a context and list of possible actions. We propose an …


How Should Ai Talk About Us? Llms And Social Generics, Tiffany A. Zhu Jan 2026

How Should Ai Talk About Us? Llms And Social Generics, Tiffany A. Zhu

Philosophy Faculty Publications

How should AI-generated speech balance epistemic aims, such as precision and accuracy, with ethical and social considerations? This paper examines a subtle yet consequential aspect of LLM-driven communication: the use of generic generalizations that convey information about social groups (e.g., “immigrants work low-wage jobs”). While central to human epistemic and pedagogical practices, generics are theorized to reinforce stereotypes, essentialism, and injustice. Using ChatGPT-3.5 as a case study, I uncover tendencies for AI chatbots to inconsistently hedge and refuse generics, including those that reflect well-documented social structural patterns, such as “women are more likely to get attacked while walking alone at …


Distilling The Complexity Of Agent-Based Simulations Into Textual Explanations Via Large Language Models, Noé Y. Flandre, Philippe J. Giabbanelli Jan 2026

Distilling The Complexity Of Agent-Based Simulations Into Textual Explanations Via Large Language Models, Noé Y. Flandre, Philippe J. Giabbanelli

VMASC Publications

Communicating the design and results of agent-based models (ABMs) to subject matter experts is challenging, which hinders participation and limits trust in simulation-based decision support. Large language models (LLMs) can communicate ABMs as textual summaries, thus complementing traditional disclosure through statistical and visualization techniques. While prior work translated the structure of conceptual models into narratives via LLMs, our extension covers the dynamics of simulation models via an automated simulation-to-text method that extracts contextual information from NetLogo ABMs, performs repeated simulations, and generates narrative descriptions (including the model’s purpose, parameters, and simulation dynamics) using mutimodal LLMs. Furthermore, four summarization algorithms spanning …