Open Access. Powered by Scholars. Published by Universities.®

Artificial Intelligence and Robotics Commons™

Open Access. Powered by Scholars. Published by Universities.®

11,188 Full-Text Articles 24,563 Authors 5,758,021 Downloads 274 Institutions

All Articles in Artificial Intelligence and Robotics

Faceted Search

11,188 full-text articles. Page 11 of 542.

Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao TANG, Pengkun JIAO, Bin ZHU, Huiyan QI, Jingjing CHEN, Yu-Gang JIANG 2026 Singapore Management University

Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …


Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing HAN, Bin ZHU, Shiqi HU, Franklin Mingzhe LI, Patrick CARRINGTON, Roger ZIMMERMANN, Jingjing CHEN 2026 Singapore Management University

Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen

Research Collection School Of Computing and Information Systems

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …


Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin WANG, Yuge HUANG, Jianqing XU, Yue YU, Jiangtao YAN, Shouhong DING, Pan ZHOU, Yong LUO 2026 Singapore Management University

Tranx-Adapter: Bridging Artifacts And Semantics Within Mllms For Robust Ai-Generated Image Detection, Wenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu, Jiangtao Yan, Shouhong Ding, Pan Zhou, Yong Luo

Research Collection School Of Computing and Information Systems

Rapid advances in AI-generated image (AIGI) technology enable highly realistic synthesis, threatening public information integrity and security. Recent studies have demonstrated that incorporating texture-level artifact features alongside semantic features into multimodal large language models (MLLMs) can enhance their AIGI detection capability. However, our preliminary analyses reveal that artifact features exhibit high intra-feature similarity, leading to an almost uniform attention map after the softmax operation. This phenomenon causes attention dilution, thereby hindering effective fusion between semantic and artifact features. To overcome this limitation, we propose a lightweight fusion adapter, TranX-Adapter, which integrates a Task-aware Optimal-Transport Fusion that leverages the Jensen-Shannon divergence …


Towards Uniformity And Alignment For Multimodal Representation Learning, Wenzhe YIN, Pan ZHOU, Zehao XIAO, Jie LIU, Shujian YU, Jan-Jakob SONKE, Efstratios GAVVES 2026 Singapore Management University

Towards Uniformity And Alignment For Multimodal Representation Learning, Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves

Research Collection School Of Computing and Information Systems

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify two conflicts in the multimodal regime, both exacerbated as the number of modalities increases: (i) an alignment–uniformity conflict, whereby the repulsion of uniformity undermines pairwise alignment, and (ii) an intra-alignment conflict, where aligning multiple modalities induces competing alignment directions. To address these issues, we propose a principled decoupling of alignment and uniformity for multimodal representations, providing a conflict-free recipe for multimodal learning that …


Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong ZOU, Jianshu LI, Jing HUANG, Pan ZHOU 2026 Singapore Management University

Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou

Research Collection School Of Computing and Information Systems

Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws MCMC samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using …


Air: Improving Agent Safety Through Incident Response, Zibo XIAO, Jun SUN, Junjie CHEN 2026 Singapore Management University

Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen

Research Collection School Of Computing and Information Systems

Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and …


Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan ZHANG, Jun SUN 2026 Singapore Management University

Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun

Research Collection School Of Computing and Information Systems

Large language models (LLMs) are increasingly trained on massive, heterogeneous text corpora, raising serious concerns about the unauthorised use of proprietary or personal data during model training. In this work, we address the problem of data protection against unwanted model learning in a realistic blackbox setting. We propose Disclaimer Injection, a novel data-level defence that renders text unlearnable to LLMs. Rather than relying on model-side controls or explicit data removal, our approach exploits the models’ own alignment mechanisms: injecting carefully designed alignment-triggers to prevent effective learning. Through layer-wise analysis, we find that finetuning on such protected data induces persistent activation …


Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan XIAO, Yuchen CHEN, Jiaming WANG, Wei SONG, Jun SUN, Shiqing MA, Yanzhou MU, Juan ZHAI, Chunrong FANG, Jin Song DONG, Zhenyu CHEN 2026 Singapore Management University

Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen

Research Collection School Of Computing and Information Systems

The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and …


Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan KE, Yi Meng LAU, Siaw Ling LO 2026 Singapore Management University

Knowledge-State Generative Agents For Pre-Assessment Question Evaluation, Ping Fan Ke, Yi Meng Lau, Siaw Ling Lo

Research Collection School Of Computing and Information Systems

This paper introduces a Knowledge‑State Generative Agent framework for evaluating the quality of pre‑assessment questions. The framework employs large language model (LLM)–based agents prompted to adopt a teacher persona to simulate the responses of students with and without mastery of targeted knowledge components. A preliminary empirical study using archival data from 424 students enrolled in an Information Systems Management course indicates that the proposed approach yields interpretable metrics under Classical Test Theory. Results further show that agents instantiated with the relevant mastered knowledge components exhibit systematically higher performance than agents lacking such mastery. In addition, the study suggests that teacher-persona …


Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan HUANG, Yunshan MA, Hongyu ZHANG, Hua MA, Zhu SUN 2026 Singapore Management University

Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun

Research Collection School Of Computing and Information Systems

Addressing itinerary modification is crucial for enhancing the travel experience as it is a frequent requirement during traveling. However, existing research mainly focuses on fixed itinerary planning, leaving modification underexplored due to the scarcity of shape need-to-modify itinerary data. To bridge this gap, we formally define the itinerary modification task and propose a general pipeline to construct the corresponding dataset, namely iTIMO. This pipeline frames the generation of shape need-to-modify itinerary data as an intent-driven perturbation task. It instructs large language models to perturb real-world itineraries using three operations: REPLACE, ADD, and DELETE. Each perturbation is grounded in three intents: …


Dual-Diffusional Generative Fashion Recommendation, Mingzhe YU, Lei WU, Qianru SUN, Yunshan MA 2026 Singapore Management University

Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma

Research Collection School Of Computing and Information Systems

Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user …


Scattered Hypothesis Generation For Open-Ended Event Forecasting, He CHANG, Zhulin TAO, Lifang YANG, Xianglin HUANG, Yunshan MA 2026 Singapore Management University

Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma

Research Collection School Of Computing and Information Systems

Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hypothesis generation. This paradigm aims to generate an inclusive and diverse set of hypotheses that broadly cover the space of plausible future events. To this end, we propose SCATTER, a reinforcement learning framework that jointly optimizes inclusiveness and diversity of the hypothesis. Specifically, we design a novel hybrid reward that consists of …


Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin XIANG, Yunshan MA, Xiaoyu DU, Yibing CHEN, Yanxin ZHANG, Jinhui TANG 2026 Singapore Management University

Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang

Research Collection School Of Computing and Information Systems

Document Question Answering (DQA) involves generating answers from a document based on a user’s query, representing a key task in document understanding. This task requires interpreting visual layouts, which has prompted recent studies to adopt multimodal Retrieval-Augmented Generation (RAG) that processes page images for answer generation. However, in multimodal RAG, visual DQA struggles to utilize a large number of images effectively, as the retrieval stage often retains only a few candidate pages (e.g., Top-4), causing informative but less visually salient content to be overlooked in favor of common yet low-information pages. To address this issue, we propose a Multi-Armed Bandit–based …


Avadclip: Audio-Visual Collaboration For Robust Video Anomaly Detection, Peng WU, Wanshun SU, Guansong PANG, Yujia SUN, Qingsen YAN, Peng WANG, Yanning ZHANG 2026 Singapore Management University

Avadclip: Audio-Visual Collaboration For Robust Video Anomaly Detection, Peng Wu, Wanshun Su, Guansong Pang, Yujia Sun, Qingsen Yan, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

With the increasing adoption of video anomaly detection in intelligent surveillance domains, conventional visual-only detection approaches often struggle with information insufficiency and high false-positive rates in complex environments. To address these limitations, we present a novel weakly supervised framework that leverages audio-visual collaboration for robust video anomaly detection. Capitalizing on the exceptional cross-modal representation learning capabilities of Contrastive Language-Image Pretraining (CLIP) across visual, audio, and textual domains, our framework introduces two major innovations: an efficient audio-visual fusion that enables adaptive cross-modal integration through lightweight parametric adaptation while maintaining the frozen CLIP backbone, and a novel audio-visual prompt that dynamically enhances …


Robust Graph Learning On The Web: Challenges, Methods, And Applications, Ao XIANG, Yang LIU, Guansong PANG, Yuanhao DING, Hezhe QIAO, Dawei CHENG, Qing HE 2026 Singapore Management University

Robust Graph Learning On The Web: Challenges, Methods, And Applications, Ao Xiang, Yang Liu, Guansong Pang, Yuanhao Ding, Hezhe Qiao, Dawei Cheng, Qing He

Research Collection School Of Computing and Information Systems

Graph learning is transforming web intelligence, powering applications from recommender systems to anomaly detection. However, most existing approaches implicitly assume ideal conditions where training and testing data are accurate, complete, and free from manipulation. In reality, web environments rarely exhibit such stability. Dynamic user behavior, incomplete or outdated content, adversarial interference, and sudden distribution shifts can all erode the reliability of even state-of-the-art models, leading to biased or unsafe outcomes. This tutorial provides a comprehensive survey of emerging strategies for robust graph learning on the web. We first present a structured taxonomy of the principal robustness threats specific to web …


Co-Matching: Towards Human–Model Collaborative Legal Case Matching, Chen HUANG, Xinwei YANG, Yang DENG, Wenqiang LEI, Jiancheng LV, Tat-Seng CHUA 2026 Singapore Management University

Co-Matching: Towards Human–Model Collaborative Legal Case Matching, Chen Huang, Xinwei Yang, Yang Deng, Wenqiang Lei, Jiancheng Lv, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Recent efforts have aimed to improve AI models in legal case matching by integrating legal domain knowledge. However, successful legal case matching requires the tacit knowledge of legal practitioners, which is difficult to verbalize and encode into models. This emphasizes the crucial role of involving legal practitioners in high-stakes legal case matching. To address this, we propose a collaborative matching framework called Co-Matching, which encourages both the model and the legal practitioner to participate in the matching process, integrating tacit knowledge. Unlike existing methods that rely solely on the model, Co-Matching allows both the legal practitioner and the model to …


Multicbr: Multi‑View Contrastive Learning For Bundle Recommendation, Yunshan MA, Yingzhi HE, Xiang WANG, Yinwei WEI, Xiaoyu DU, Yuyangzi FU, Tat‑Seng CHUA 2026 Singapore Management University

Multicbr: Multi‑View Contrastive Learning For Bundle Recommendation, Yunshan Ma, Yingzhi He, Xiang Wang, Yinwei Wei, Xiaoyu Du, Yuyangzi Fu, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

Bundle recommendation seeks to recommend a bundle of related items to users to improve both userexperience and the profits of platform. Existing bundle recommendation models have progressed from capturing only user-bundle interactions to the modeling of multiple relations among users, bundles, and items.CrossCBR, in particular, incorporates cross-view contrastive learning into a two-view preference learningframework, significantly improving SOTA performance. It does, however, have two limitations: (1) the twoview formulation does not fully exploit all the heterogeneous relations among users, bundles, and items; and(2) the “early contrast and late fusion” framework is less effective in capturing user preference and difficultto generalize to …


Advancing Social Media Analytics And Personalized Generation Via Transfer Learning, Discourse-Aware Modeling, And Collaborative Modeling, Gibson Nkhata 2026 University of Arkansas-Fayetteville

Advancing Social Media Analytics And Personalized Generation Via Transfer Learning, Discourse-Aware Modeling, And Collaborative Modeling, Gibson Nkhata

Graduate Theses and Dissertations

Social media platforms have become central to information exchange, shaping public opinion across social, political, and economic domains. However, the massive volume of user-generated content, combined with its informal, nuanced, and often noisy nature, presents significant challenges for automated analysis and generation. Tasks such as stance detection, rumor verification, and personalized content generation are further complicated by sarcasm, evolving discourse structures, and diverse user preferences. Addressing these challenges requires models that can effectively leverage linguistic nuance, conversational dynamics, and collaborative user signals. Transfer learning has emerged as a powerful paradigm for improving performance in low-resource and complex language understanding tasks. …


Improved Pbs Algorithm For Multi-Agent Path Planning Based On Conflict Guidance And Punishment Mechanism, Jinbao Zhang, Jianlin Mao, Chengze Qian, Guimi Sun, Kaixin Tong 2026 Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming 650500, China

Improved Pbs Algorithm For Multi-Agent Path Planning Based On Conflict Guidance And Punishment Mechanism, Jinbao Zhang, Jianlin Mao, Chengze Qian, Guimi Sun, Kaixin Tong

Journal of System Simulation

Abstract: To address the bottleneck in which the priority-based search (priority-based search, PBS) algorithm for multi-agent path planning easily falls into conflict loops and generates invalid node expansions in complex scenarios, an improved algorithm based on conflict guidance and a punishment mechanism (improved PBS multi-agent path finding algorithm based on conflict guidance and punishment mechanism, CGP-PBS) was proposed. A conflict-guided node expansion mechanism was constructed; in high-level search, it comprehensively evaluated path cost and the number of conflicts, preferentially expanded child nodes with high potential for conflict resolution, and delayed the expansion of high-conflict nodes, thereby effectively compressing the search …


Operational Agency: A Permeable Legal Fiction For Tracing Culpability In Ai Systems, Anirban MUKHERJEE, Hannah H. CHANG 2026 Singapore Management University

Operational Agency: A Permeable Legal Fiction For Tracing Culpability In Ai Systems, Anirban Mukherjee, Hannah H. Chang

Research Collection Lee Kong Chian School Of Business

Modern artificial intelligence (AI) systems act with a high degree of independence yet lack legal personhood—a paradox that fractures doctrines grounded in human-centric notions of mens rea and actus reus. This Article introduces Operational Agency (OA)—a permeable legal fiction structured as an ex post evidentiary framework—and Operational Agency Graph (OAG)—a tool for mapping causal interactions among human actors, organizations, and AI systems. OA evaluates an AI’s observable operational characteristics: its goal-directedness (as a proxy for intent), predictive processing (as a proxy for foresight), and safety architecture (as a proxy for standard of care). OAG operationalizes that analysis by embedding these …


Digital Commons powered by bepress