When Saying "No" Is Not Enough: Cognitive-Action Decoupling And The Illusion Of Safety In Llm Agents,
2026
Clark University
When Saying "No" Is Not Enough: Cognitive-Action Decoupling And The Illusion Of Safety In Llm Agents, Shasha Yu
School of Professional Studies
Current safety evaluations of large language models (LLMs) predominantly rely on textual compliance, implicitly assuming that refusal-style responses correspond to safe behavior. This assumption becomes fragile when LLMs are embedded in agentic systems with the ability to execute state-changing actions. In this paper, we present an empirical critique of text-centric safety evaluation through an action-aware study of LLM agents under controlled conditions. Across multiple state-of-the-art models, we observe a recurring cognitive-action decoupling: agents generate policy-aligned refusal language while still producing unsafe tool-mediated action proposals. This produces an illusion of safety, where conversational audits indicate compliance even as operational risk persists. …
Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples,
2026
Singapore Management University
Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples, Alexander Vincent Lewi, Rainer Tan, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose InterFold, a framework for learning and applying interpretable semantic manifolds in latent diffusion models, without requiring binary or paired supervision. Existing methods for semantic editing either rely on limited paired data or uncover only coarse, unsupervised directions that fail to capture user-specific, fine-grained attributes. InterFold addresses these limitations by learning a target attribute manifold in the H-space of diffusion models using only a set of positive, unlabeled examples. To edit a new image, InterFold projects its H-space representation toward this learned manifold through test-time optimization, enabling precise, identity-preserving modifications of complex, non-binary concepts. To make these edits effective …
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech,
2026
California Polytechnic State University, San Luis Obispo
Trace: Temporal Rhetorical Analysis And Consistency Evaluation For Legislative Speech, David Hernandez
Master's Theses
Legislators frequently discuss the same policy issues across multiple hearings and legislative sessions, sometimes maintaining consistent positions and other times modifying or reframing their stance over time. Understanding how these positions evolve is important for analyzing political discourse and democratic accountability, yet identifying such shifts at scale remains difficult.
We introduce TRACE (Temporal Rhetorical Analysis and Consistency Evaluation), a system built on the Digital Democracy Database (DDDB) for detecting rhetorical inconsistency in California legislative hearing testimony. TRACE organizes utterances into speaker-anchored timelines indexed by bill and session, then applies hybrid semantic retrieval — combining dense BGE embeddings with BM25 lexical …
High-Resolution Queries, Low-Resolution Context: Scaling Vision Transformers With Asymmetric Spatial Reduction,
2026
California Polytechnic State University, San Luis Obispo
High-Resolution Queries, Low-Resolution Context: Scaling Vision Transformers With Asymmetric Spatial Reduction, James M. Dwyer
Master's Theses
Vision Transformers (ViTs) have demonstrated great performance on image classifica tion benchmarks, however, the quadratic complexity of the self-attention mechanism with respect to sequence length limits their scalability to higher resolution inputs. The attention score matrix grows as O(N2) in both compute and memory, where N is the number of patch tokens, making ViTs computationally expensive and memory intensive for applications that require real-time inference or operate under resource constraints.
This thesis investigates whether the key and value sequences of the self-attention mechanism can be compressed using the local spatial structure of the image — while keeping queries at full …
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs,
2026
CUNY Graduate Center
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin
Dissertations, Theses, and Capstone Projects
We perform field-level likelihood-free inference of the matter density parameter Ωm from simulated galaxy catalogs using machine learning models with differing inductive biases. Using features extracted from hydrodynamic simulations in the CAMELS suite, we investigate how both observable choice and model architecture govern the extraction of cosmological information. We consider galaxy positions and line-of-sight peculiar velocities, both separately and in combination, and compare permutation-invariant Deep Sets, implemented with either standard multilayer perceptrons (MLPs) or Kolmogorov–Arnold Networks (KANs), to graph neural networks (GNNs) implemented with MLPs, which explicitly encode spatial relations. We evaluate inference performance under both in-distribution and out-of-distribution (OOD) …
Scorespeak: An Agentic System For Natural Language Control Of Musical Scores,
2026
California Polytechnic State University, San Luis Obispo
Scorespeak: An Agentic System For Natural Language Control Of Musical Scores, Nathan S. Lim
Master's Theses
With the recent popularization of large language models (LLMs), natural language has become one of the most accessible and powerful ways for people to interact with creative tools. Although they have become common in mainstream domains like image and audio editing, there is currently no robust AI-based system that can reliably turn free-form language into edits for symbolic musical scores. This gap represents a missed opportunity to improve human workflows for creating and editing sheet music, but it is also a fundamental limitation for other agentic music systems; without a robust mechanism for translating free-form language into structured scores, AI …
Towards Auto-Evaluation For Large Language Models,
2026
Singapore Management University
Towards Auto-Evaluation For Large Language Models, Jiahao Ying
Dissertations and Theses Collection (Open Access)
The rapid advancement of large language models (LLMs) has created an urgent need for evaluation methodologies that are timely, scalable, reliable, and informative. Conventional evaluation benchmarks, although essential for measuring model capabilities and guiding model development, are often constructed and maintained through labor-intensive human annotation. As LLMs continue to improve through increases in model scale, training data, and computational resources, static benchmarks may quickly lose discriminative power. Moreover, the growing use of large and diverse training corpora increases the risk of benchmark leakage, which can inflate evaluation results and obscure the true capabilities of models. These challenges call for a …
How To Save The Take-Home Essay With Oral Assessments,
2026
Singapore Management University
How To Save The Take-Home Essay With Oral Assessments, Matthew Hammerton, Jacqueline Ho
Research Collection School of Social Sciences
In a commentary, the authors opined that pairing take-home essays with oral assessments is a more effective response to AI than policing its use. Students who cannot adequately explain their work can be marked down, reducing incentives to rely on AI. They noted that oral exams help preserve key elements of university education – intellectual effort, ownership, and human relationships – while allowing take-home essays to remain relevant in an AI-driven landscape that demands greater emphasis on understanding, responsibility, and dialogue.
Rode: Linear Rectified Mixture Of Diverse Experts For Food Large Multi-Modal Models,
2026
Singapore Management University
Rode: Linear Rectified Mixture Of Diverse Experts For Food Large Multi-Modal Models, Pengkun Jiao, Xinlan Wu, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang
Research Collection School Of Computing and Information Systems
Large Multi-modal Models (LMMs) have significantly advanced a variety of vision-language tasks. The scalability and availability of high-quality training data play a pivotal role in the success of LMMs. In the realm of food, while comprehensive food datasets such as Recipe1M offer an abundance of ingredient and recipe information, they often fall short of providing ample data for nutritional analysis. The Recipe1M+ dataset, despite offering a subset for nutritional evaluation, is limited in the scale and accuracy of nutrition information. To bridge this gap, we introduce Uni-Food, a unified food dataset that comprises over 100,000 images with various food labels, …
Videocreator: An Agentic System For Multi-Turn Video Production,
2026
Singapore Management University
Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao
Research Collection School Of Computing and Information Systems
Recent advances in video generation models enable visually compelling single clips. However, real-world video creation is inherently continuous and iterative: creators refine content over multiple rounds while maintaining narrative, style, and entity consistency. Existing standalone generators are largely stateless and lack memory of previously generated segments, making it difficult to produce a coherent and consistent video project. To address this gap, we present VideoCreator, a unified video agent that integrates generation and understanding with a project-level memory system. VideoCreator leverages understanding capabilities to perform fine-grained analysis of newly produced content and uses persistent memory to retain and reuse prior context …
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation,
2026
Singapore Management University
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Research Collection School Of Computing and Information Systems
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …
Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion,
2026
Singapore Management University
Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du
Research Collection School Of Computing and Information Systems
Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation (MVR) due to their strong multimodal understanding. However, existing apporches typically deploy LVLMs as fixed black-box feature extractors without systematically comparing alternative representation strategies. To address this gap, we present the first systematic empirical study on various feature extraction paradigms and integration strategies, along with hierarchical representations from frozen LVLMs for MVR. Extensive experiments on representative LVLMs reveal that hidden states from multiple decoder layers provide richer and more effective representations for MVR. Guided by this insight, we propose the Dual Feature Fusion (DFF) Framework, a lightweight approach …
“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts,
2026
Singapore Management University
“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts, Tianyi Zhang, Emran Bin Elias Poh, Yueyue Hou, Yi-Chieh Lee, Renwen Zhang, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Intergenerational conversations often break down when differences in tone, language, or expectations lead participants to feel dismissed or misunderstood. In this work, we explore how people envision AI-driven chatbot interventions for addressing communication problems in text-based intergenerational family chat. We conducted a scenario-based design interview with 10 pairs of family members from different generations, in which participants designed chatbot interventions that varied in intervention target and timing. Our findings show that participants expect chatbots to perform multiple themes of intervention, including mediating understanding, providing emotional support, offering evaluative commentary, and guiding interaction through behavioral suggestions. These expectations varied systematically across …
Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction,
2026
Singapore Management University
Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction, Shunyi Yeo, Tianyi Zhang, Scott Bateman, Gary Hsieh, Young-Ho Kim, Simon Tangi Perrault, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Conversational agents that participate in or mediate group interaction introduce challenges that extend beyond supporting individual users, raising new questions about how agents participate in and influence groups. To characterise this emerging design space, we present a systematic review of 53 peer-reviewed studies on group conversational agents (GCAs). We analyse how GCAs intervene in group-level processes, including participation regulation, conflict mediation, task alignment, and execution support. Using concepts from group research as an analytic lens, we organise prior GCA work around recurring group interactional challenges (orientation, conflict, alignment, and execution), and examine the roles agents are designed to play in …
Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles,
2026
Singapore Management University
Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia
Research Collection School Of Computing and Information Systems
Text-to-image (T2I) generative models are increasingly used to produce content for education, media, and public-facing communication, and are starting to be integrated into higher-impact pipelines. Since generated images tend to reinforce stereotypes, producing representational erasure via “default” depictions and shaping perceptions of who belongs in certain roles, a growing body of work has proposed metrics to quantify gender bias in T2I outputs. Yet existing evaluations remain fragmented. Metrics are often reported without a shared view of what they measure, what assumptions they entail, or how their results should be interpreted under different deployment contexts. This limits the usefulness of gender …
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation,
2026
Singapore Management University
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
Research Collection School Of Computing and Information Systems
Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …
“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts,
2026
Singapore Management University
“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt
Research Collection School Of Computing and Information Systems
Due to their limited ability to reason about the social context in which they are used, smart speakers pose significant privacy risks by responding in ways that may violate people's implicit social boundaries. We conducted a cross-cultural vignette study (N = 944) in Germany and Singapore to investigate how situational factors—specifically social context (bystander relationships and closeness), physical context (location), and interaction context (topic and deceptive intent)—regulate user preferences for smart speaker responses. Our results demonstrate that these factors are superior predictors of response preferences than dispositional user traits (i.e., intrinsic personal traits). We identify two distinct social dynamics: a …
Legal Ethics Of Ai Snake Oil: Navigating The Hype, Harm, And Hope Of Legal Ai,
2026
University of Nevada, Las Vegas William S. Boyd School of Law
Legal Ethics Of Ai Snake Oil: Navigating The Hype, Harm, And Hope Of Legal Ai, Drew Simshaw
Michigan Law Review
A review of AI Snake Oil.By Arvind Narayanan and Sayash Kapoor.
Bridging Data Gaps In Retinal Imaging: From Structural Domain Adaptation To Topology-Aware Synthesis,
2026
CUNY Graduate Center
Bridging Data Gaps In Retinal Imaging: From Structural Domain Adaptation To Topology-Aware Synthesis, Gözde Merve Demirci
Dissertations, Theses, and Capstone Projects
Comprehensive visualization of the retina is essential for diagnosing and monitoring blinding diseases such as Diabetic Retinopathy and Retinopathy of Prematurity (ROP), where pathological changes often extend beyond a single field of view. Despite significant advances in automated retinal image analysis, clinical deployment remains limited by two fundamental data gaps: a structural learning gap, arising from scarce expert annotations and poor generalization across imaging domains, and a spatial coverage gap, caused by the difficulty of acquiring multi-view retinal images in fragile populations. Although these challenges are often addressed independently, this dissertation argues that they are tightly coupled: accurate, …
Saag: Structured Agent Assessment And Grounding,
2026
University of South Carolina - Columbia
Saag: Structured Agent Assessment And Grounding, Ritvik Garimella, Vedant Khandelwal, Anvi Kohli, Amit Sheth
Publications
Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucinate argument values, or satisfy a schema while choosing a agent for the wrong reason. Existing benchmarks collapse these distinctions into a single binary score, leaving practitioners unable to diagnose where agent calls fail. We propose SAAG a cascaded diagnostic framework that decomposes agent-calling evaluation into three sequential stages: registry conformance, structural completeness, and argument grounding, each producing interpretable stage-specific diagnostics. These diagnostics additionally enable iterative self-repair: on prediction failure, the stage-specific signal guides targeted correction without leaking ground-truth values. We evaluate this …
