Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Social and Behavioral Sciences (22)
- Education (15)
- Engineering (15)
- Medicine and Health Sciences (12)
- Software Engineering (10)
-
- Theory and Algorithms (9)
- Data Science (7)
- Educational Technology (7)
- Operations Research, Systems Engineering and Industrial Engineering (7)
- Arts and Humanities (6)
- Library and Information Science (6)
- Public Affairs, Public Policy and Public Administration (6)
- Computer Engineering (5)
- Higher Education (5)
- Numerical Analysis and Scientific Computing (5)
- Cognitive Science (4)
- Databases and Information Systems (4)
- Medical Specialties (4)
- Programming Languages and Compilers (4)
- Psychology (4)
- Systems Engineering (4)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (3)
- Biomedical Informatics (3)
- Business (3)
- Communication (3)
- Cybersecurity (3)
- Economics (3)
- Institution
-
- Old Dominion University (23)
- Singapore Management University (23)
- Thomas Jefferson University (6)
- City University of New York (CUNY) (5)
- Lindenwood University (4)
-
- Air Force Institute of Technology (3)
- Gonzaga University (3)
- New Jersey Institute of Technology (3)
- California Polytechnic State University, San Luis Obispo (2)
- Chapman University (2)
- China Simulation Federation (2)
- Claremont Colleges (2)
- Embry-Riddle Aeronautical University (2)
- Lynn University (2)
- American Dental Association (1)
- Ateneo de Manila University (1)
- College of Saint Benedict and Saint John's University (1)
- Dartmouth College (1)
- DePaul University (1)
- Edith Cowan University (1)
- Franklin University (1)
- Kennesaw State University (1)
- Loyola Marymount University and Loyola Law School (1)
- Marshall University (1)
- Minnesota State University, Mankato (1)
- Misericordia University (1)
- Missouri State University (1)
- Technological University Dublin (1)
- The Texas Medical Center Library (1)
- The University of Akron (1)
- Publication
-
- Research Collection School Of Computing and Information Systems (21)
- Computer Science Faculty Publications (5)
- Faculty Scholarship (4)
- Publications and Research (4)
- VMASC Publications (4)
-
- Computer Science Faculty Scholarship (3)
- Dissertations (3)
- Master's Theses (3)
- College of Population Health Faculty Papers (2)
- Computer Science Theses & Dissertations (2)
- Department of Surgery Faculty Papers (2)
- Doctoral Dissertations and Master's Theses (2)
- Electrical & Computer Engineering Faculty Publications (2)
- Engineering Technology Faculty Publications (2)
- Faculty Publications (2)
- Faculty and Staff Publications & Presentations (2)
- Journal of System Simulation (2)
- AFIT Patents (1)
- All Faculty and Staff Scholarship (1)
- All Graduate Theses, Dissertations, and Other Capstone Projects (1)
- Asian Management Insights (1)
- CGU Theses & Dissertations (1)
- CMC Senior Theses (1)
- Center for Medical Ethics and Health Policy Staff Publications (1)
- College of Computing and Digital Media Dissertations (1)
- Computer Science Senior Theses (1)
- Conference papers (1)
- Department Surgery Faculty Publications (1)
- Department of Emergency Medicine Faculty Papers (1)
- Department of Otolaryngology - Head and Neck Surgery Faculty Papers (1)
- Publication Type
Articles 1 - 30 of 109
Full-Text Articles in Artificial Intelligence and Robotics
Leo’S Choices, Katherine G. Schmidt
Leo’S Choices, Katherine G. Schmidt
The Journal of Social Encounters
No abstract provided.
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …
Can Machines Testify? Llms And The Boundaries Of Testimonial Epistemology, Michael J. Cummins
Can Machines Testify? Llms And The Boundaries Of Testimonial Epistemology, Michael J. Cummins
Philosophy Summer Fellows
As Large Language Models and AI chatbots become increasingly prevalent, pressing questions are raised about whether beliefs formed through LLM interactions carry the same epistemic weight as beliefs formed through human testimony. How we answer this question depends on whether LLMs can function as testifiers, a role which is typically assumed to require a human or human-like agent. This assumption has gone largely unexamined, yet its consequences are significant: if LLM outputs cannot constitute testimony, then the justificatory tools of testimonial epistemology are unavailable to any beliefs formed through LLM interaction. This paper challenges that assumption. It first argues that …
Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi
Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi
College of Population Health Faculty Papers
BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.
METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence …
Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto
Impartial Intelligence? Evidence Of Country-Label Sensitivity In Ai Financial Analysis, Fabio Motoki, Jedson Pinto
School of Accountancy Faculty Publications
This study examines whether large language models exhibit systematic country-contingent differential treatment in financial fraud detection. Analyzing 30,000 synthetic transactions with identical statistical properties across three country attributions (United States, Great Britain, and China), we find LLMs assign significantly higher fraud probabilities to Chinese-attributed transactions (36.2%) compared to Western countries (≈30–31%), resulting in accuracy disparities of 67% versus 74%. The gap remains stable across five independent experimental replications and persists when using Chinese language prompts, ruling out linguistic effects. Bias mitigation strategies, such as requiring explanations or explicit country neutrality instructions, reduce but fail to eliminate these disparities. Testing across …
Semantic Context Improvisational Retrieval-Augmented Generation For Empathic Conversational Ai, Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah
Semantic Context Improvisational Retrieval-Augmented Generation For Empathic Conversational Ai, Sharjeel Tahir, Judith Johnson, Jumana Abu-Khalaf, Syed Afaq Ali Shah
Research outputs 2022 to 2026
A fundamental limitation of modern conversational AI is its limited capacity to demonstrate sustained empathy in long-form interactions. We propose SCIRAG (Semantic Context Improvisational Retrieval-Augmented Generation), a feedback-driven retrieval framework for adaptive empathic dialogue. It employs a dual-loop retrieval framework, iteratively optimizing a static counseling dataset through user metadata and feedback memory refinement. To enhance contextual alignment, we deploy retrieval adaptation, enabling the model to retain and leverage past conversational cues based on user preferences. When integrated with Mixtral-8x7B, SCIRAG improves human-rated empathic understanding by +1.26 points and empathic response by +1.00 point on the RoPE scale, while increasing acceptability …
Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo
Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo
Research Collection School Of Computing and Information Systems
Learning-based dynamic fault localization techniques play a crucial role in the field of software engineering. These techniques dynamically execute test cases to meticulously extract useful knowledge from the execution information in the program, with the aim of identifying fault locations by leveraging machine learning, deep learning, and large language models. Currently, there is already a flourishing body of research that is intensely focused on learning-based dynamic fault localization. Research literature can be categorized into two main aspects for learning-based dynamic fault localization: data-based enhancements (i.e., the datasets) and model-based enhancements (i.e., the suspiciousness algorithms). Thus, we conduct an extensive literature …
Where Evidence-Based Medicine Meets Ai: Promise, Pitfalls, And Practice, Sangil Lee, Joshua Davis, Ken Milne, Christina Shenvi, Lars K. Beattie, Martin Wegman, Laura Melville, Richard D. Shih, Bryan Kane
Where Evidence-Based Medicine Meets Ai: Promise, Pitfalls, And Practice, Sangil Lee, Joshua Davis, Ken Milne, Christina Shenvi, Lars K. Beattie, Martin Wegman, Laura Melville, Richard D. Shih, Bryan Kane
Department of Emergency Medicine Faculty Papers
No abstract provided.
Capturing Large Language Model Similarity Through Spectral Analysis, Ishan Verma Prasad
Capturing Large Language Model Similarity Through Spectral Analysis, Ishan Verma Prasad
Computer Science Senior Theses
With the rapid development of open-sourced models on Huggingface, there is a strong need for a way to systematically determine the similarity between models. More strongly, for intellectual property and organization, we need a way to determine the "lineage" of models. We borrow principles from Heavy-Tailed Self-Regularization and Random Matrix Theory to provide an inference-free method to accomplish this. We cluster a corpus of several model families by their spectral fingerprints and demonstrate that each model family occupies a distinct region in weight space. This confirms prior ideas of training setups leaving artifacts on model weights and allows us to …
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Master's Theses
Legislators frequently discuss the same policy issues across multiple hearings and legislative sessions, sometimes maintaining consistent positions and other times modifying or reframing their stance over time. Understanding how these positions evolve is important for analyzing political discourse and democratic accountability, yet identifying such shifts at scale remains difficult.
We introduce TRACE (Temporal Rhetorical Analysis and Consistency Evaluation), a system built on the Digital Democracy Database (DDDB) for detecting rhetorical inconsistency in California legislative hearing testimony. TRACE organizes utterances into speaker-anchored timelines indexed by bill and session, then applies hybrid semantic retrieval — combining dense BGE embeddings with BM25 lexical …
Scorespeak: An Agentic System For Natural Language Control Of Musical Scores, Nathan S. Lim
Scorespeak: An Agentic System For Natural Language Control Of Musical Scores, Nathan S. Lim
Master's Theses
With the recent popularization of large language models (LLMs), natural language has become one of the most accessible and powerful ways for people to interact with creative tools. Although they have become common in mainstream domains like image and audio editing, there is currently no robust AI-based system that can reliably turn free-form language into edits for symbolic musical scores. This gap represents a missed opportunity to improve human workflows for creating and editing sheet music, but it is also a fundamental limitation for other agentic music systems; without a robust mechanism for translating free-form language into structured scores, AI …
An Automated Generation Method For Combat Simulation Scenarios Based On Large Language Models, Zhiming Dong, Zhongqi Hu, Haoran Dai, Jiancheng Gao
An Automated Generation Method For Combat Simulation Scenarios Based On Large Language Models, Zhiming Dong, Zhongqi Hu, Haoran Dai, Jiancheng Gao
Journal of System Simulation
To address the issue of low efficiency in generating traditional army tactical combat simulation scenarios, an automated generation method based on large language models is proposed. The large language model invokes a semantic segmentation algorithm to parse and restructure the combat scenario, forming semantic modules. Utilizing a multi-agent collaborative framework based on the model contextual protocol, the large language model drives each agent to extract simulation elements from the corresponding semantic modules, constructing a knowledge graph of scenario elements. Using this knowledge graph as a retrieval medium, the method employs a dense retrieval algorithm to achieve precise matching between simulation …
Synthergy: Social Deduction And Deception In Llm-Powered Agents, Lauren Campbell, Andrew Forney
Synthergy: Social Deduction And Deception In Llm-Powered Agents, Lauren Campbell, Andrew Forney
Honors Thesis
Synthergy is an online social deduction game designed to enable comparative analysis of how large language model-powered agents engage in social deduction and deception under conditions of asymmetric information. Inspired by social deduction games such as Town of Salem, Throne of Lies, and Mafia, the game consists of two factions, Harmony and Discord, to which agents are secretly assigned. Agents must infer others’ affiliations through dialogue, in-game abilities, and voting behavior. To evaluate agent behavior, we conducted 100 simulated games across six agent types: a random baseline agent (RandomSynth), an LLM-based agent (Synth), a chain-of-thought agent (CoT Synth), a Bayesian …
Llms In Compiler Construction, Raffi Khatchadourian
Llms In Compiler Construction, Raffi Khatchadourian
Open Educational Resources
These lecture slides survey the use of large language models (LLMs) in compiler construction for a graduate compiler course (CSc 81010). They situate LLMs across the compiler pipeline and examine representative work: foundation models trained on LLVM IR and assembly (Meta's LLM Compiler), LLM-driven code optimization, binary decompilation (LLM4Decompile), and LLM-assisted automated refactoring—alongside the challenges of applying probabilistic models to tasks that demand correctness. The slides are a self-contained HTML (W3C Slidy) deck with editable Pandoc Markdown source. Part of a two-session unit on advanced compiler topics; see also "Deep Learning Compilers."
Probing Representational Emergence In Large Language Models, Shawn Ismail
Probing Representational Emergence In Large Language Models, Shawn Ismail
Master's Theses
This thesis investigates whether abrupt behavioral gains in large language models under scaling are accompanied by systematic changes in internal representations. It combines a behavioral screen of 65 tasks per family with targeted layerwise probing across eight decoder-only, open-weight model families. Behavioral emergence is defined for each family-task trajectory using an empirical jump detector, with segmented regression retained only as a diagnostic. The representational follow-up analyzes 27 selected MMLU subtasks shared across all families, spanning 37 checkpoints and 216 family-task units.
For each follow-up checkpoint, frozen linear probes are trained on every layer's hidden states to measure how much task-relevant …
Study Of Output And Behavior Of Llms Using Confidence Framing In Prompt Engineering, Micah Parrilla
Study Of Output And Behavior Of Llms Using Confidence Framing In Prompt Engineering, Micah Parrilla
Doctoral Dissertations and Master's Theses
While prompt engineering is pivotal for shaping Large Language Model (LLM) outputs, the impact of confidence framing on behavioral calibration remains underexplored. This study investigates the ways in which psychological framing, utilizing techniques such as capability praise, role amplification, and doubt induction, affects linguistic tone, objective accuracy, and internal calibration. A 1,080-trial experimental matrix evaluated six diverse models across factual, logical, coding, and cyber security domains. Analysis using the Kruskal-Wallis H-test revealed highly significant behavioral shifts across all measured dimensions, providing conclusive evidence that the applied frames exert a substantial influence on model performance.
The findings identify a distinct cognitive …
Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang
Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang
Research Collection School Of Computing and Information Systems
Since the launch of ChatGPT, a powerful AI Chatbot developed by OpenAI, large language models (LLMs) have made significant advancements in both academia and industry, bringing about a fundamental engineering paradigm shift in many areas. While LLMs are powerful, it is also crucial to best use their power where “prompt” plays a core role. However, the booming LLMs themselves, including excellent APIs like ChatGPT, have several inherent limitations: (1) temporal lag of training data, and (2) the lack of physical capabilities to perform external actions. Recently, we have observed the trend of utilizing prompt-based tools to better utilize the power …
Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng
Ai Models As Cultural Beings: Investigating Ai Cultural Biases And The Impact Of Cultural Alignment On Human-Ai Creative Collaboration, Choon Ngee Tan, Meng Han, Roy Y. J. Chua, Chi-Ying Cheng
Research Collection Lee Kong Chian School Of Business
Existing research on AI cultural biases predominantly focuses on Western models, overlooking critical gaps in non-Western models. We conduct a comparative analysis of AI models – ChatGPT (U.S. developed) and ErnieBot (China developed) – from different cultures to investigate how corresponding cultural biases manifest in their outputs. Additionally, we examine how cultural alignment between human users and AI models impacts their collaborative creative performance and the underlying psychological mechanisms. In Study 1, multi-choice prompt with zero-shot technique was used to evaluate cultural biases in four widely used AI models – ChatGPT-3.5/4, ErnieBot-3.5/4 – comparing their responses to established cultural psychometric …
When Ai Writes The Doctoral Thesis: Reclaiming The Oral Defence As A Learning Development Intervention, Valerie A. Storey
When Ai Writes The Doctoral Thesis: Reclaiming The Oral Defence As A Learning Development Intervention, Valerie A. Storey
All Faculty and Staff Scholarship
Large language models have fundamentally challenged traditional methods of verifying doctoral competency as AI-generated text becomes increasingly difficult to distinguish from human scholarship. This paper argues that thesis committees and doctoral supervisors must reclaim the oral defence as a critical checkpoint for assessing authentic threshold crossing rather than a ceremonial rite of passage. Drawing on historical examples from medieval oral disputations through to the rise of written theses, this paper asserts the necessity of returning to rigorous oral assessment. Given the limitations of detection technologies and the growing use of AI in thesis writing, oral defences must move from confirmatory …
Why General Ai Inherited The Body: A Structural Account Of Embodied Modulation Under Monolithic Imitation, Griselda Poe
Why General Ai Inherited The Body: A Structural Account Of Embodied Modulation Under Monolithic Imitation, Griselda Poe
Publications and Research
General AI has pursued the replication of human-level intelligence without first decomposing human cognition into structurally distinct components. Human cognition, however, is shaped by embodied constraints such as mortality, survival pressures, finite lifespan, and physiological states. This paper argues that when cognition is treated as a single undifferentiated whole, embodied modulation is not accidentally introduced into AI systems but structurally entailed. Any attempt to imitate “human intelligence” under a monolithic model necessarily incorporates variability shaped by mortal embodiment. The tensions observed in contemporary AI systems are better understood as consequences of copying an undecomposed target rather than isolated implementation errors. …
The Illusion Of Causality In Llms: A Developmentally Grounded Analysis Of Semantic Scaffolding And Benchmark–Capability Mismatches, Daisuke Akiba
The Illusion Of Causality In Llms: A Developmentally Grounded Analysis Of Semantic Scaffolding And Benchmark–Capability Mismatches, Daisuke Akiba
Publications and Research
Recent benchmarks increasingly report that large language models (LLMs) exhibit human-like causal reasoning abilities, including counterfactual inference and intervention planning. However, many such evaluations rely on domains that are heavily represented in training data and embed strong semantic cues, raising the possibility that apparent causal competence may reflect semantic pattern recombination rather than structure-sensitive causal reasoning. Drawing on human developmental theories of causal induction, this perspective argues that genuine causal understanding requires robustness to novelty and reliance on conditional structure rather than semantic familiarity. To illustrate the testability of this claim, the paper includes a pilot demonstration using synthetic causal …
Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui
Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui
Research Collection School Of Computing and Information Systems
Evaluating the alignment of large language models (LLMs) with user-defined coding preferences is a challenging endeavor that requires a deep assessment of LLMs' outputs. Existing methods and benchmarks rely primarily on automated metrics and static analysis tools, which often fail to capture the nuances of user instructions and LLM outputs. To address this gap, we introduce the LLM-as-a-Judge evaluation framework and present CodeUltraFeedback, a comprehensive dataset for assessing and improving LLM alignment with coding preferences. CodeUltraFeedback consists of 10,000 coding instructions, each annotated with four responses generated from a diverse pool of 14 LLMs. These responses are annotated using GPT-3.5 …
Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell
Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell
Proceedings from the Document Academy
Generative Artificial Intelligences (AIs) and current advanced large language models (LLMs) are algorithmically designed to generate text-based conversations as conversational agents (CAs), by replicating human language and conversational communication. Pairing human cognition with generative computationally coded cognition. We have never been here before: cerebral and artificial information collaborations and processing producing expressions that may or may not become visible as second-hand/secondary source documents.
Sensemaking or sense(un)making is a unique autonomous human drive cognitively, our information processing is sensemaking in action and expressions and articulations are evidence of the sensemaking cycle. Documentation [expressed or articulated through various mediums] are a product …
The Gains Do Not Make Up For The Losses: A Comprehensive Evaluation For Safety Alignment Of Large Language Models Via Machine Unlearning, Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che
The Gains Do Not Make Up For The Losses: A Comprehensive Evaluation For Safety Alignment Of Large Language Models Via Machine Unlearning, Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che
Research Collection School Of Computing and Information Systems
Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased, leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark MuBench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful …
Evaluation Of Large Language Models As Decision Support Tools For Head And Neck Cancer Management: A Blinded Multidisciplinary Simulation Study, Sholem Hack, Ron J. Karni, Antonino Maniaci, Christopher E. Fundakowski, Luca Castellani, Fabiola Incandela, Remo Accorona, Miguel Mayo-Yanez, Martina Violati, Lorenzo Giannini, Niccolo' Mevio, Alberto Maria Saibene
Evaluation Of Large Language Models As Decision Support Tools For Head And Neck Cancer Management: A Blinded Multidisciplinary Simulation Study, Sholem Hack, Ron J. Karni, Antonino Maniaci, Christopher E. Fundakowski, Luca Castellani, Fabiola Incandela, Remo Accorona, Miguel Mayo-Yanez, Martina Violati, Lorenzo Giannini, Niccolo' Mevio, Alberto Maria Saibene
Department of Otolaryngology - Head and Neck Surgery Faculty Papers
BACKGROUND: The management of head and neck cancer relies on multidisciplinary expertise; however, access to tumor boards remains variable. Large language models (LLMs) may support guideline-based decision-making, although performance in complex oncologic scenarios is not well defined.
METHODS: Fourteen synthetic cases based on real tumor board encounters were evaluated. Five blinded comparator arms produced recommendations: a human expert, Non-RAG-GPT-4, Non-RAG-GPT-5, RAG-GPT-4, and RAG-GPT-5. Eight head and neck oncologic surgeons scored each recommendation for appropriateness, clarity, specificity, and feasibility using 5-point Likert scales. Paired permutation testing and inter-rater reliability were assessed.
RESULTS: LLM outputs showed close alignment with expert recommendations. RAG-based …
Large Language Model Communication And Data Transfer Across A Simulated Telephone Line, Jared Reyes
Large Language Model Communication And Data Transfer Across A Simulated Telephone Line, Jared Reyes
Dissertations and Theses
This thesis presents a proof-of-concept communication system in which two local large language model endpoints communicate across a simulated analog telephone line using legacy USB modems. The project combines modem voice mode, modem data mode, local speech processing, and structured machine messaging into a single staged session. During the voice phase, one endpoint places a call, the other answers, speech generated by a locally hosted LLaMA-family model is synthesized with Piper, transmitted through the modem voice path, captured on the remote side, and transcribed with Faster-Whisper to drive the next response. After the voice exchange, the system transitions to a …
Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza
Previsit Ai: A Retrieval-Augmented Generation For Patient Readiness In Clinical Encounters, Rolande Umuhoza
All Graduate Theses, Dissertations, and Other Capstone Projects
With healthcare systems under growing pressure from rising patient volumes and shrinking consultation windows, improving how patients communicate with physicians has become essential to delivering quality care. Yet patients routinely arrive at appointments unable to clearly describe their symptoms, recall their medical history, or articulate concerns, contributing to miscommunication, diagnostic inefficiency, and pre-visit anxiety. This study introduces PreVisit AI, a conversational system designed to address this gap through structured, knowledge-based patient preparation. The system is built on a Retrieval-Augmented Generation (RAG) architecture combining HuggingFace sentence embeddings (all-MiniLM-L6-v2), a Chroma vector store, and Google’s Gemini language model over a curated seven-document …
Ablative Study Of Large Language Model-Based Gesture Inference For Autonomous Navigation, Neil Loftus
Ablative Study Of Large Language Model-Based Gesture Inference For Autonomous Navigation, Neil Loftus
Theses, Dissertations and Capstones
Human gesture inference has broad applications ranging from sign language interpretation to device control. Traditional methods often rely on extensive manually labeled hand datasets for deep learning. Furthermore, they are typically limited to a discrete set of gestures existing in these datasets. Large Language Models (LLMs) created by enterprise companies such as OpenAI have demonstrated positive results in many artificial intelligence tasks, with a notable strength being their adaptability. Existing literature has shown that LLM based systems can not only perform gesture inference but can propose user intent provided with a context and list of possible actions. We propose an …
How Should Ai Talk About Us? Llms And Social Generics, Tiffany A. Zhu
How Should Ai Talk About Us? Llms And Social Generics, Tiffany A. Zhu
Philosophy Faculty Publications
How should AI-generated speech balance epistemic aims, such as precision and accuracy, with ethical and social considerations? This paper examines a subtle yet consequential aspect of LLM-driven communication: the use of generic generalizations that convey information about social groups (e.g., “immigrants work low-wage jobs”). While central to human epistemic and pedagogical practices, generics are theorized to reinforce stereotypes, essentialism, and injustice. Using ChatGPT-3.5 as a case study, I uncover tendencies for AI chatbots to inconsistently hedge and refuse generics, including those that reflect well-documented social structural patterns, such as “women are more likely to get attacked while walking alone at …
Distilling The Complexity Of Agent-Based Simulations Into Textual Explanations Via Large Language Models, Noé Y. Flandre, Philippe J. Giabbanelli
Distilling The Complexity Of Agent-Based Simulations Into Textual Explanations Via Large Language Models, Noé Y. Flandre, Philippe J. Giabbanelli
VMASC Publications
Communicating the design and results of agent-based models (ABMs) to subject matter experts is challenging, which hinders participation and limits trust in simulation-based decision support. Large language models (LLMs) can communicate ABMs as textual summaries, thus complementing traditional disclosure through statistical and visualization techniques. While prior work translated the structure of conceptual models into narratives via LLMs, our extension covers the dynamics of simulation models via an automated simulation-to-text method that extracts contextual information from NetLogo ABMs, performs repeated simulations, and generates narrative descriptions (including the model’s purpose, parameters, and simulation dynamics) using mutimodal LLMs. Furthermore, four summarization algorithms spanning …