Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Social and Behavioral Sciences (11)
- Data Science (8)
- Engineering (8)
- Software Engineering (7)
- Computer Engineering (6)
-
- Medicine and Health Sciences (6)
- Databases and Information Systems (5)
- Communication (4)
- Communication Technology and New Media (4)
- Cybersecurity (4)
- Education (4)
- Educational Technology (4)
- Graphics and Human Computer Interfaces (4)
- Programming Languages and Compilers (4)
- Information Security (3)
- Other Computer Sciences (3)
- Systems Architecture (3)
- Arts and Humanities (2)
- Biomedical Informatics (2)
- Business (2)
- Chemical and Pharmacologic Phenomena (2)
- Computer and Systems Architecture (2)
- Health Information Technology (2)
- Law (2)
- Linguistics (2)
- Medical Sciences (2)
- Psychology (2)
- Institution
-
- Singapore Management University (30)
- Dartmouth College (5)
- Thomas Jefferson University (3)
- University of Texas at Arlington (3)
- City University of New York (CUNY) (2)
-
- Georgia Southern University (2)
- Lindenwood University (2)
- Old Dominion University (2)
- University of Arkansas, Fayetteville (2)
- University of Missouri, St. Louis (2)
- University of South Florida (2)
- California State University, San Bernardino (1)
- Clemson University (1)
- Dakota State University (1)
- Edith Cowan University (1)
- Embry-Riddle Aeronautical University (1)
- Loyola Marymount University and Loyola Law School (1)
- Michigan Technological University (1)
- Mississippi State University (1)
- Portland State University (1)
- Rollins College (1)
- Southern Methodist University (1)
- United Arab Emirates University (1)
- University of Arkansas Little Rock (1)
- University of Central Florida (1)
- University of Kentucky (1)
- University of Michigan Law School (1)
- University of Missouri-Kansas City School of Law (1)
- University of North Florida (1)
- University of South Carolina (1)
- Publication
-
- Research Collection School Of Computing and Information Systems (26)
- Dissertations and Theses Collection (Open Access) (3)
- College of Graduate Studies: Theses & Dissertations (2)
- College of Population Health Faculty Papers (2)
- Dartmouth College Master’s Theses (2)
-
- Dartmouth College Ph.D Dissertations (2)
- Faculty Scholarship (2)
- Publications (2)
- Publications and Research (2)
- Theses (2)
- Theses and Dissertations (2)
- USF Tampa Graduate Theses and Dissertations (2)
- 2024 Fall Honors Capstone Projects - Archive (1)
- All Dissertations (1)
- Articles (1)
- Computer Science Senior Theses (1)
- Computer Science Theses & Dissertations (1)
- Computer Science and Computer Engineering Undergraduate Honors Theses (1)
- Computer Science and Engineering Dissertations (1)
- Computer Science and Engineering Theses - Archive (1)
- Department of Otolaryngology - Head and Neck Surgery Faculty Papers (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Electrical Engineering and Computer Science Undergraduate Honors Theses (1)
- Electronic Theses, Projects, and Dissertations (1)
- Endeavors: Mississippi State Undergraduate Research Journal (1)
- Faculty Works (1)
- Graduate Student Government Association Research Conference (1)
- Honors Program Theses (1)
- Honors Thesis (1)
- Human-Machine Communication (1)
- Publication Type
Articles 1 - 30 of 77
Full-Text Articles in Artificial Intelligence and Robotics
Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi
Llm-As-A-Judge For Infection Prevention And Control And Antimicrobial Resistance Impact: Comparing Three Main Llms Vs. Human Experts' Assessment, Marcello Di Pumpo, Leonardo Villani, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Patrizia Laurenti, Vittorio Maio, Stefania Boccia, Walter Ricciardi
College of Population Health Faculty Papers
BACKGROUND: Large language models (LLMs) are increasingly used to generate health information, yet their reliability as evaluators remains unclear. This study investigated the feasibility of an LLM-as-a-judge methodology in the context of infection prevention and antimicrobial resistance (AMR), comparing automated ratings with human expert benchmarks.
METHODS: We performed a secondary analysis of an expert-annotated dataset of health messages. Three leading LLMs (ChatGPT, Claude, Gemini) independently evaluated the same messages using an adapted DISCERN tool across five domains: information reliability, quality, AMR impact, persuasiveness, and overall score. We utilized descriptive statistics, intra-rater reliability tests, and mixed-effects ordinal regression to analyze divergence …
A Scoping Review Of Sycophancy In Large Language Models: Operational And Theoretical Recognition, Kallen Zhou, Manning Littlejohn, Isabella Garrard
A Scoping Review Of Sycophancy In Large Language Models: Operational And Theoretical Recognition, Kallen Zhou, Manning Littlejohn, Isabella Garrard
Endeavors: Mississippi State Undergraduate Research Journal
As large language models (LLMs) usage grows across different domains, sycophancy, the tendency for output to align with users, is increasingly being recognized as a primary issue arising from applying LLMs into critical areas. Current research has provided a variety of theoretical definitions, mitigation techniques, and quantification for sycophancy. However, there is little to no consistency across different papers. This scoping review seeks to connect different works on LLM sycophancy by identifying themes in theoretical definitions, measurement methods, and inducement techniques of sycophancy. By analyzing 26 papers (preprints, conference proceedings, and journal articles) from arXiv, ACL Anthology, and Scopus, this …
A Novel Hierarchical Multi-Agent System For Payments Using Llms, Donghao Huang, Joon Kiat Chua, Zhaoxia Wang
A Novel Hierarchical Multi-Agent System For Payments Using Llms, Donghao Huang, Joon Kiat Chua, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Large language model (LLM) agents, such as OpenAI’s Operator and Claude’s Computer Use, can automate workflows but unable to handle payment tasks. Existing agentic solutions have gained significant attention; however, even the latest approaches face challenges in implementing end-to-end agentic payment workflows. To address this gap, this research proposes the Hierarchical Multi-Agent System for Payments (HMASP), which provides an end-to-end agentic method for completing payment workflows. The proposed HMASP leverages either open-weight or proprietary LLMs and employs a modular architecture consisting of the Conversational Payment Agent (CPA - first agent level), Supervisor agents (second agent level), Routing agents (third agent …
Automatically Constructed Preference Pairs For Chain-Of-Thought: Consistency Gains With Accuracy Tradeoffs, Cameron Scolari, Lanyu Shang
Automatically Constructed Preference Pairs For Chain-Of-Thought: Consistency Gains With Accuracy Tradeoffs, Cameron Scolari, Lanyu Shang
Honors Thesis
We investigate preference optimization over chain-of-thought (CoT) reasoning using automatically constructed preference signals derived from the accuracy and internal consistency of a model. Our results show that framing reasoning as a preference learning problem improves both the accuracy of the final answer and the structure of the model outputs. We observe a non-monotonic relationship between performance and the Direct Preference Optimization (DPO) scaling parameter β, where moderate values maximize accuracy while lower values improve stability, highlighting a tradeoff between optimization strength and reliable generation. We further identify a tradeoff between reasoning consistency and accuracy. Increasing the consistency weight improves agreement …
Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang
Generative Ai In Enterprises: Optimizing Applications With Large Language Models, Donghao Huang
Dissertations and Theses Collection (Open Access)
This dissertation investigates how to deploy Large Language Models (LLMs) effectively in enterprise settings, where accuracy, reliability, cost, privacy, and operational constraints often matter more than benchmark performance alone. Drawing on seventeen peer-reviewed publications (eleven published and six accepted for publication), the work develops and validates optimization strategies across three connected themes: retrieval-augmented generation (RAG), agentic AI for workflow automation, and deployment guidelines for real-world enterprise environments.
First, we study RAG optimization through systematic evaluation of open and proprietary models, highlighting conditions under which efficient open-weight models can match or exceed proprietary alternatives. To address a pervasive failure mode in …
Ai Interpretability In Healthcare Communication, Ananya Jeyappragash
Ai Interpretability In Healthcare Communication, Ananya Jeyappragash
Dartmouth College Master’s Theses
Artificial intelligence has increasingly been adopted in healthcare, largely for specialized tasks and under significant human oversight. The use of large black-box systems raises important concerns about transparency in high-stakes environments such as clinical decision-making. Clinical communication is fundamentally human-centered, and failures in judgment can have serious consequences for patient care. Overestimating the reasoning abilities of large language models may lead to undue trust in fabricated or “hallucinated” outputs, while rejecting AI-assisted tools altogether may preserve inefficient workflows and contribute to missed or delayed diagnoses. These concerns reflect a broader tradeoff between accuracy and interpretability: although more complex models may …
Llmqua: Practical Backdoor Injection On Large Language Model Quantization, Xiangxiang Chen, Peixin Zhang, Jun Sun, Jin Song Dong, Wenhai Wang, Jingyi Wang
Llmqua: Practical Backdoor Injection On Large Language Model Quantization, Xiangxiang Chen, Peixin Zhang, Jun Sun, Jin Song Dong, Wenhai Wang, Jingyi Wang
Research Collection School Of Computing and Information Systems
Quantization is widely used to enable local deployment of large language models (LLMs) on resource-constrained devices. Recent work (e.g., QuRA) shows quantization can be exploited via rounding manipulation to implant backdoors. However, such an attack has been evaluated only on small models and does not directly apply to LLMs due to three key constraints: (1) limited poisoning data from small, task-agnostic calibration sets; (2) layer-wise quantization restricting adversarial access to global representations; and (3) lack of gradient access in quantization pipelines, blocking gradient-based attacks.We propose LLMQuA, a practical quantization-phase backdoor attack tailored to the LLM setting. LLMQuA (i) injects backdoors …
Evaluation Of Large Language Models As Decision Support Tools For Head And Neck Cancer Management: A Blinded Multidisciplinary Simulation Study, Sholem Hack, Ron J. Karni, Antonino Maniaci, Christopher E. Fundakowski, Luca Castellani, Fabiola Incandela, Remo Accorona, Miguel Mayo-Yanez, Martina Violati, Lorenzo Giannini, Niccolo' Mevio, Alberto Maria Saibene
Evaluation Of Large Language Models As Decision Support Tools For Head And Neck Cancer Management: A Blinded Multidisciplinary Simulation Study, Sholem Hack, Ron J. Karni, Antonino Maniaci, Christopher E. Fundakowski, Luca Castellani, Fabiola Incandela, Remo Accorona, Miguel Mayo-Yanez, Martina Violati, Lorenzo Giannini, Niccolo' Mevio, Alberto Maria Saibene
Department of Otolaryngology - Head and Neck Surgery Faculty Papers
BACKGROUND: The management of head and neck cancer relies on multidisciplinary expertise; however, access to tumor boards remains variable. Large language models (LLMs) may support guideline-based decision-making, although performance in complex oncologic scenarios is not well defined.
METHODS: Fourteen synthetic cases based on real tumor board encounters were evaluated. Five blinded comparator arms produced recommendations: a human expert, Non-RAG-GPT-4, Non-RAG-GPT-5, RAG-GPT-4, and RAG-GPT-5. Eight head and neck oncologic surgeons scored each recommendation for appropriateness, clarity, specificity, and feasibility using 5-point Likert scales. Paired permutation testing and inter-rater reliability were assessed.
RESULTS: LLM outputs showed close alignment with expert recommendations. RAG-based …
Computational Clinical Judgment: Predicting Risk With Large Language Models, Hannah Laqueur, Ryan W. Copus
Computational Clinical Judgment: Predicting Risk With Large Language Models, Hannah Laqueur, Ryan W. Copus
Faculty Works
For seventy years, research has shown actuarial methods outperform clinical judgment. Yet actuarial approaches have limitations: they generally rely on structured data; cannot exploit rare case-specific details; have limited accuracy where outcome data are scarce or incomplete; and cannot offer case-level justifications. Large language models (LLMs) offer a different approach. Like actuarial methods, they aggregate information algorithmically, but like clinicians, they bring general knowledge and can provide case-level justifications. We prompted seven LLMs to assess rearrest risk from 113 parole hearing transcripts and compared their predictions to a machine learning model trained on 4,000 cases with 91 administrative variables. GPT-5 …
Systematic Approaches To Characterizing Vulnerabilities And Enhancing Robustness Of Text And Vision-Language Models, Poojitha Thota
Systematic Approaches To Characterizing Vulnerabilities And Enhancing Robustness Of Text And Vision-Language Models, Poojitha Thota
Computer Science and Engineering Dissertations
The proliferation of artificial intelligence (AI) across critical domains, including news summarization, privacy-policy analysis, and medical decision support, has raised growing concerns about the security and robustness of these systems against adversarial manipulation. This dissertation investigates adversarial robustness in generative AI by addressing three key research goals: (1) characterizing adversarial vulnerabilities across generative models, (2) developing systematic defenses to improve the robustness of generative models, and (3) designing deployment-time safeguards for securing LLM interactions.
Towards the first goal, we characterize adversarial vulnerabilities across text-based and multimodal systems. In abstractive text summarization, we show that inference-time perturbations can exploit lead bias …
Designing Narrative-Based Ai Assistance For Sensemaking In Collaborative Environments: Case Studies In Education And Dementia Care, Dylan Edward Moore
Designing Narrative-Based Ai Assistance For Sensemaking In Collaborative Environments: Case Studies In Education And Dementia Care, Dylan Edward Moore
Dartmouth College Ph.D Dissertations
This thesis addresses a gap in the human-computer interaction literature regarding the design, development, and evaluation of narrative-based AI assistance for collaborative, complex problem solving. I explore this design space through three case studies across the domains of education and dementia care. This work encompasses multi-year industry partnerships and longitudinal fieldwork, user-centered design, dataset curation, model training, and system evaluation.
Specifically, the first case study considers a story-based web platform for teaching AI literacy through peer-generated, personalized narrative scaffolding. Learners on the platform showed significant knowledge gains and other learning-related outcomes. To describe the novel design of this system, I …
Ai-Driven Real-Time Detection Of Zero-Day Browser Exploits Using Webassembly-Based Instrumentation, Temitope Damilola Elijah
Ai-Driven Real-Time Detection Of Zero-Day Browser Exploits Using Webassembly-Based Instrumentation, Temitope Damilola Elijah
College of Graduate Studies: Theses & Dissertations
The rapid evolution of web browsers into fully fledged application execution environments has significantly expanded their attack surface, making them prime targets for sophisticated zero-day exploits that evade traditional signature-based security mechanisms. To address this challenge, this research proposes an AI-driven framework for real-time detection and analysis of zero-day exploits in web browsers by integrating browser-level telemetry monitoring, unsupervised anomaly detection, and large language model–based threat interpretation. The framework introduces a lightweight WebAssembly telemetry agent embedded within the browser runtime to capture low-level execution behaviors, including WASM module instantiation, memory growth patterns, network interactions, and runtime API activity. These telemetry …
Real-Time Isolated Asl Recognition: Evaluating Spatial-Temporal Networks And Multimodal Llms, Raga Mouni Batchu
Real-Time Isolated Asl Recognition: Evaluating Spatial-Temporal Networks And Multimodal Llms, Raga Mouni Batchu
West Chester University Graduate Theses, Dissertations, and Final Projects
This thesis investigates the deployment of high-accuracy Isolated ASL Recognition (ISLR) in resource-constrained edge environments. We train a lightweight Spatio-Temporal Attention Network (SSTAN,∼2.7 M parameters,∼10 MB) on the WLASL-100 benchmark, achieving 75.25% Top-1 and 88.24% Top-5 accuracy with 139 ms CPU-only inference. A systematic comparison against frontier multimodal LLMs (Gemini 3 Flash, Gemini 3.1 Pro, Qwen 3 VL) shows SSTAN outperforms the best LLM baseline by∼1.85×in accuracy while being 22–230×faster and up to 40×cheaper annually. The LLMs’ core limitation is a lack of fine-grained temporal perception; they impose English-language semantic priors rather than learning the articulatory distinctions that define ASL …
A Novel Federated Llm Framework For Distributed Traffic Modelling In Intelligent Transportation Systems, Seerat Kaur
A Novel Federated Llm Framework For Distributed Traffic Modelling In Intelligent Transportation Systems, Seerat Kaur
Theses and Dissertations (Comprehensive)
Intelligent transportation systems (ITS) depend on accurate traffic prediction to support congestion management, infrastructure planning, and real-time operational decisions. Despite substantial progress in data-driven forecasting, several challenges continue to limit practical deployment: traffic data is distributed across independent regional authorities, making centralized aggregation infeasible, standard federated aggregation strategies ignore traffic-specific characteristics that meaningfully affect model quality, and existing models produce only numerical outputs without interpretable reasoning that urban planners can act upon. This thesis addresses these challenges through four contributions that collectively advance privacy-preserving, explainable, and scalable traffic forecasting.
The first contribution provides a systematic review of 129 peer-reviewed publications, …
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Research Collection School Of Computing and Information Systems
The rapid integration of Large Language Models (LLMs) into software engineering (SE) has revolutionized tasks from code generation to program repair, producing a massive volume of software artifacts. This surge in automated creation has exposed a critical bottleneck: the lack of scalable and reliable methods to evaluate the quality of these outputs. Human evaluation, while effective, is very costly and time-consuming. Traditional automated metrics like BLEU rely on high-quality references and struggle to capture nuanced aspects of software quality, such as readability and usefulness. In response, the LLM-as-a-Judge paradigm, which employs LLMs for automated evaluation, has emerged. This approach leverages …
Leveraging Large Language Models For Career Mobility Analysis: A Study Of Gender, Race, And Job Change Using Us Online Resume Profiles, Palakorn Achananuparp, Ye Xu, Yao Lu, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim
Leveraging Large Language Models For Career Mobility Analysis: A Study Of Gender, Race, And Job Change Using Us Online Resume Profiles, Palakorn Achananuparp, Ye Xu, Yao Lu, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
We present a large-scale analysis of career mobility of college-educated U.S. workers using online resume profiles to investigate how gender, race, and job change options are associated with upward mobility. This study addresses key research questions of how the job changes affect their upward career mobility, and how the outcomes of upward career mobility differ by gender and race. We address data challenges – such as missing demographic attributes, missing wage data, and noisy occupation labels – through various data processing and Artificial Intelligence (AI) methods. In particular, we develop a large language models (LLMs) based occupation classification method known …
My First Conversation With Chatgpt (February 22, 2023): Origins Of A Generative Dialogue, David Smith
My First Conversation With Chatgpt (February 22, 2023): Origins Of A Generative Dialogue, David Smith
Publications and Research
This working paper presents the first recorded interaction between the author and the generative AI system ChatGPT, written on February 22, 2023 during the initial weeks of a faculty sabbatical in Boston. The document preserves a complete and unedited transcript of an exploratory conversation conducted without predetermined research aims, marking the author’s first encounter with a large-language-model conversational interface. Although the exchange includes creative experimentation—including musical and poetic prompts—the discussion remains informal and wide-ranging, and no theoretical framework is articulated at this stage. Rather, this transcript is published as primary-source material documenting the moment of discovery and experimentation that precedes …
Understanding Bias And Fairness In Large Language Models: An Empirical Study, Joshua Johnson
Understanding Bias And Fairness In Large Language Models: An Empirical Study, Joshua Johnson
Electrical Engineering and Computer Science Undergraduate Honors Theses
This thesis investigates demographic bias in large language models (LLMs) through the use of evaluating outcome disparities when utilized in decision making tasks as well as underlying associations that could contribute to furthering these disparities. Using profiles from the Adult dataset, we analyze how Gemini 2.0 Flash performs in an income prediction task using zero-shot and few-shot prompting methods. Our findings show that models exhibit measurable differences in demographic parity and false positive rates, with the use of few-shot prompting reducing these disparities. Alongside this line of testing, we tested associational bias in Qwen 2.5 using probability based association tests …
Toward Personalizing Quantum Computing Education: An Evolutionary Llm-Powered Approach, Iizalaarab Elhaimeur
Toward Personalizing Quantum Computing Education: An Evolutionary Llm-Powered Approach, Iizalaarab Elhaimeur
Computer Science Theses & Dissertations
Quantum computing education faces significant challenges due to its complexity and the limitations of current tools. This thesis introduces a novel Intelligent Teaching Assistant for quantum computing education and details its evolutionary design process. The system combines a knowledge-graph-augmented architecture with two specialized LLM agents: a Teaching Agent for dynamic interaction and a Lesson Planning Agent for lesson generation. The system is designed to adapt to individual student needs, with interactions meticulously tracked and stored in a knowledge graph. This graph represents student actions, learning resources, and their relationships, aiming to enable reasoning about effective learning pathways. We describe the …
Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang
Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) have transformed sentiment analysis, yet balancing accuracy, efficiency, and explainability remains a critical challenge. This study presents the first comprehensive evaluation of DeepSeek-R1—an open-source reasoning model—against OpenAI’s GPT-4o and GPT-4o-mini. We test the full 671B model and its distilled variants, systematically documenting few-shot learning curves. Our experiments show DeepSeek-R1 achieves a 91.39% F1 score on 5-class sentiment and 99.31% accuracy on binary tasks with just 5 shots, an eightfold improvement in few-shot efficiency over GPT-4o. Architecture-specific distillation effects emerge, where a 32B Qwen2.5-based model outperforms the 70B Llama-based variant by 6.69 percentage points. While its reasoning …
Enhancing Llm Code Generation: A Systematic Evaluation Of Multi-Agent Collaboration And Runtime Debugging For Improved Accuracy, Reliability, And Latency, Nazmus Ashrafi
Thesis/ Dissertation Defenses
The use of large language models (LLMs) for automated code generation has emerged as a significant focus within AI research. As these pretrained models continue to evolve, their ability to understand and generate complex code structures has opened up new possibilities for automating intricate programming tasks with greater accuracy. Although contemporary foundational models demonstrate promising results, researchers continue to explore optimal post-training strategies to enhance code quality. These include supervised fine-tuning, retrieval-augmented generation (RAG), debugging, and many others. In this thesis, I combine two such widely used post training approaches—namely (1) multi-agent collaboration and (2) runtime execution of information-based debugging—for …
Large Language Models As Information Providers For Appropriate Antimicrobial Use: Computational Text Analysis And Expert-Rated Comparison Of Chatgpt, Claude And Gemini, Marcello Di Pumpo, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Vittorio Maio, Patrizia Laurenti, Walter Ricciardi, Leonardo Villani
Large Language Models As Information Providers For Appropriate Antimicrobial Use: Computational Text Analysis And Expert-Rated Comparison Of Chatgpt, Claude And Gemini, Marcello Di Pumpo, Maria Rosaria Gualano, Danilo Buonsenso, Francesca Raffaelli, Daniele Donà, Vittorio Maio, Patrizia Laurenti, Walter Ricciardi, Leonardo Villani
College of Population Health Faculty Papers
OBJECTIVES: Antimicrobial resistance is a critical public health threat. Large language models (LLMs) show great capability for providing health information. This study evaluates the effectiveness of LLMs in providing information on antibiotic use and infection management.
METHODS: Using a mixed-method approach, responses to healthcare expert-designed scenarios from ChatGPT 3.5, ChatGPT 4.0, Claude 2.0 and Gemini 1.0, in both Italian and English, were analysed. Computational text analysis assessed readability, lexical diversity and sentiment, while content quality was assessed by three experts via DISCERN tool.
RESULTS: 16 scenarios were developed. A total of 101 outputs and 5454 Likert-scale (1-5) scores were obtained …
Rethinking Cognitive Complexity For Unit Tests: Toward A Readability-Aware Metric Grounded In Developer Perception, Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang, Xin Zhou, Anil Koyuncu, Jacques Klein, David Lo, Tegawendé F. Bissyandé
Rethinking Cognitive Complexity For Unit Tests: Toward A Readability-Aware Metric Grounded In Developer Perception, Wendkûuni C. Ouédraogo, Yinghua Li, Xueqi Dang, Xin Zhou, Anil Koyuncu, Jacques Klein, David Lo, Tegawendé F. Bissyandé
Research Collection School Of Computing and Information Systems
Automatically generated unit tests-from searchbased tools like EvoSuite or LLMs-vary significantly in structure and readability. Yet most evaluations rely on metrics like Cyclomatic Complexity and Cognitive Complexity, designed for functional code rather than test code. Recent studies have shown that SonarSource's Cognitive Complexity metric assigns nearzero scores to LLM-generated tests, yet its behavior on EvoSuitegenerated tests and its applicability to test-specific code structures remain unexplored. We introduce CCTR, a Test-Aware Cognitive Complexity metric tailored for unit tests. CCTR integrates structural and semantic features like assertion density, annotation roles, and test composition patterns-dimensions ignored by traditional complexity models but critical for …
Fact-Checker: A Web Application For Leveraging Large Language Models For Fact-Checking Youtube Videos, Andrew R. Craig
Fact-Checker: A Web Application For Leveraging Large Language Models For Fact-Checking Youtube Videos, Andrew R. Craig
Electronic Theses, Projects, and Dissertations
Fact-Checker is a web application that allows users to fact-check YouTube videos. It feeds YouTube’s closed captioning transcript to a large language model (LLM) to extract claims. It then uses multiple LLMs, such as Gemini, Llama, and Claude, to verify these claims. The modular design makes it easy to change to a different LLM or model if needed. The application is built using Python for access to Application Programming Interfaces (APIs) and Streamlit as the front-end framework. The utilization of Docker and Dockerfiles enables easy distribution and deployment. It enables the application to be deployed on almost any hardware platform …
Llm2rec: Large Language Models Are Powerful Embedding Models For Sequential Recommendation, Yingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma, Tat‑Seng Chua
Llm2rec: Large Language Models Are Powerful Embedding Models For Sequential Recommendation, Yingzhi He, Xiaohao Liu, An Zhang, Yunshan Ma, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
Sequential recommendation aims to predict users' future interactions by modeling collaborative filtering (CF) signals from historical behaviors of similar users or items. Traditional sequential recommenders predominantly rely on ID-based embeddings, which capture CF signals through high-order co-occurrence patterns. However, these embeddings depend solely on past interactions, lacking transferable knowledge to generalize to unseen domains. Recent advances in large language models (LLMs) have motivated text-based recommendation approaches that derive item representations from textual descriptions. While these methods enhance generalization, they fail to encode CF signals-i.e., latent item correlations and preference patterns-crucial for effective recommendation. We argue that an ideal embedding model …
A Web Application For Generating Argument Maps For Essays Using Llms, Alexis R. Chalmers
A Web Application For Generating Argument Maps For Essays Using Llms, Alexis R. Chalmers
Theses
Argumentative writing is a critical skill that strengthens students’ reasoning, communication, and analytical abilities. However, maintaining a clear and organized argument structure while writing can be challenging. Argument maps — visual diagrams which explicitly show an argument’s structure — have been shown to improve students’ writing, but are rarely used outside of the planning stage of an essay due to the time and effort required to create them. Automatically generating argument maps from student essays helps students to evaluate the structure of their argument as they write and makes identifying unsupported claims visible. To evaluate whether large language models (LLMs) …
Limitations Of Using Large Language Models For Automated Essay Scoring, Thomas A. Fink
Limitations Of Using Large Language Models For Automated Essay Scoring, Thomas A. Fink
Theses
Background: Automated essay scoring (AES) is a challenging deep learning problem. The two most widely used methods for predicting essay quality scores, supervised learning-based and LLM-based, have their own limitations. Although supervised learning-based methods are more accurate, they only predict a score and do not offer descriptive feedback to students. On the other hand, LLM-based methods can offer rubric-guided feedback but are known to be less accurate.
Methods: This work focuses on improving the accuracy of state-of-the-art LLM-based AES methods. We began by thoroughly investigating why these methods were performing poorly for certain datasets and certain examples. This led us …
"Chatgpt Told Me To Say It": Ai Chatbots And Class Participation Apprehension In University Students, Daisuke Akiba
"Chatgpt Told Me To Say It": Ai Chatbots And Class Participation Apprehension In University Students, Daisuke Akiba
Publications and Research
The growing prevalence of AI chatbots in everyday life has prompted educators to explore their potential applications in promoting student success, including support for classroom engagement and communication. This exploratory study emerged from semester-long observations of class participation apprehensions in an introductory educational psychology course, examining how chatbots might scaffold students toward active and independent classroom contribution. Four students experiencing situational participation anxiety voluntarily participated in a pilot intervention using AI chatbots as virtual peer partners. Following comprehensive training in AI use and prompt design given to the entire class, participants employed systematic consultation frameworks for managing classroom discourse trepidations. …
Finir: The 2nd Workshop On Financial Information Retrieval In The Era Of Generative Ai, Fengbin Zhu, Yunshan Ma, Fuli Feng, Chao Wang, Huanbo Luan, Guangnan Ye, Shuo Zhang, Dhagash Mehta, Pingping Chen, Bing Xiang, Tat‑Seng Chua
Finir: The 2nd Workshop On Financial Information Retrieval In The Era Of Generative Ai, Fengbin Zhu, Yunshan Ma, Fuli Feng, Chao Wang, Huanbo Luan, Guangnan Ye, Shuo Zhang, Dhagash Mehta, Pingping Chen, Bing Xiang, Tat‑Seng Chua
Research Collection School Of Computing and Information Systems
Recent advancements in Generative AI, such as Large Language Models (LLMs), have demonstrated remarkable success across various general tasks. Extensive studies have explored leveraging generative models in finance, but significant challenges persist. This half-day workshop explores potential approaches and research directions to address these challenges by equipping generative models with advanced Information Retrieval (IR) models. Specifically, this workshop seeks to provide a platform for discussing innovative ideas that facilitate the advancement of IR technology to enrich generative models in finance from four key perspectives: (i) financial IR techniques (ii) financial IR benchmarking and evaluation (iii) financial systems and agents/assistants (iv) …
Context-Switch Attacks: Understanding And Mitigating The Threat To Llm Applications, Sydney Holder, Bivin Sadler
Context-Switch Attacks: Understanding And Mitigating The Threat To Llm Applications, Sydney Holder, Bivin Sadler
SMU Data Science Review
Large Language Models (LLMs) are transforming conversational AI, yet their dependence on prompt-supplied context exposes them to context-switch attacks that covertly steer dialogue toward sensitive or malicious ends. A 70 one-sided conversation transcript evaluation set was constructed spanning various fraudulent scenarios. Each transcript embeds adversarial patterns drawn while preserving natural conversational flow. We introduce a hybrid defense that pairs a BERT-based semantic-drift detector (cosine-similarity threshold = 0.70) with a curated keyword and hack-phrase scanner to counter these threats. In aggregate, the system delivered 100 % recall, intercepting every simulated phishing or data-harvesting attempt. The keyword layer achieved perfect precision, generating …