Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (344)
- Engineering (291)
- Operations Research, Systems Engineering and Industrial Engineering (255)
- Graphics and Human Computer Interfaces (171)
- Software Engineering (129)
-
- Business (125)
- Numerical Analysis and Scientific Computing (111)
- Social and Behavioral Sciences (101)
- Theory and Algorithms (97)
- Public Affairs, Public Policy and Public Administration (66)
- Transportation (61)
- Programming Languages and Compilers (57)
- Information Security (40)
- OS and Networks (37)
- Medicine and Health Sciences (35)
- Computer Engineering (31)
- Education (24)
- Health Information Technology (24)
- Asian Studies (23)
- International and Area Studies (23)
- Finance and Financial Management (14)
- Technology and Innovation (13)
- Communication (12)
- Operations and Supply Chain Management (12)
- Social Media (12)
- Higher Education (11)
- Computer and Systems Architecture (7)
- Keyword
-
- Artificial intelligence (49)
- Machine learning (46)
- Reinforcement learning (39)
- Deep learning (37)
- Large Language Models (26)
-
- Large Language Model (22)
- Large language models (21)
- Computer vision (18)
- Optimization (18)
- Scheduling (18)
- Anomaly detection (17)
- Generative AI (17)
- Large language model (17)
- Singapore (17)
- ChatGPT (16)
- Deep reinforcement learning (16)
- Reinforcement Learning (16)
- LLMs (15)
- Vehicle routing problem (15)
- Deep Learning (14)
- Natural language processing (14)
- Artificial Intelligence (13)
- Neural networks (13)
- Uncertainty (13)
- Machine Learning (12)
- Graph neural networks (11)
- Metaverse (10)
- Multi-agent systems (10)
- Representation learning (10)
- Semantics (10)
- Publication Year
- File Type
Articles 1 - 30 of 1664
Full-Text Articles in Artificial Intelligence and Robotics
Stprompt++: Prompting Vision-Language Models For Weakly Supervised Video Anomaly Detection And Fine-Grained Localization, Peng Wu, Chengyu Pan, Guansong Pang, Xiangteng He, Zhiwei Yang, Peng Wang, Yanning Zhang
Stprompt++: Prompting Vision-Language Models For Weakly Supervised Video Anomaly Detection And Fine-Grained Localization, Peng Wu, Chengyu Pan, Guansong Pang, Xiangteng He, Zhiwei Yang, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
Traditional weakly supervised video anomaly detection (WSVAD) tasks typically rely on coarse-grained frame-level labels for training. Although this approach reduces annotation costs, it results in weak semantic understanding and spatial localization capabilities due to the absence of fine-grained annotations, hindering precise pixel-level anomaly detection and localization. Thanks to the success of vision-language models (VLMs), e.g., CLIP, recent approaches leveraging large VLMs focus on exploiting their strong semantic understanding capabilities, but they typically feed only keyframes or short video segments into the models, without supplying sufficient prior contextual information (e.g., contextual frames around anomalies, zoomed-in anomaly regions, and detailed anomaly descriptions), …
Metarag: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented Generation, Cheng Tu, Yunshan Ma, Bingyang Guo, Qianyu Li, Yang Li, Min Zhang, Fan Shi, Xiang Wang
Metarag: Identifying Website Owner Using Meta-Path-Guided Dynamic Graph Retrieval-Augmented Generation, Cheng Tu, Yunshan Ma, Bingyang Guo, Qianyu Li, Yang Li, Min Zhang, Fan Shi, Xiang Wang
Research Collection School Of Computing and Information Systems
Website owner identification aims to link websites to their real-world owners, which is crucial for credibility assessment and information provenance in information retrieval and vital for applications in cybersecurity, Internet governance, and digital regulation. Existing approaches for website owner identification primarily rely on querying infrastructure registration records or analyzing webpage content. However, these methods often fail due to incomplete or outdated registration records and sparse webpage content. We observe that inter-website relationships, derived from shared infrastructure data such as primary domains, IP blocks, and geolocations, can provide valuable but underutilized ownership cues. To exploit this insight, we propose MetaRAG, a …
Logupdater: Automated Detection And Repair Of Specific Defects In Logging Statements, Renyi Zhong, Yichen Li, Jinxi Kuang, Wenwei Gu, Yintong Huo, R. Michael Lyu
Logupdater: Automated Detection And Repair Of Specific Defects In Logging Statements, Renyi Zhong, Yichen Li, Jinxi Kuang, Wenwei Gu, Yintong Huo, R. Michael Lyu
Research Collection School Of Computing and Information Systems
Developers write logging statements to monitor software runtime behaviors and system state. However, poorly constructed or misleading log messages can inadvertently obfuscate actual program execution patterns, thereby impeding effective software maintenance. Existing research on analyzing issues within logging statements is limited, primarily focusing on detecting a singular type of defect and relying on manual intervention for fixes rather than automated solutions.To address the limitation, we initiate a systematic study that pinpoints four specific types of defects in logging statements (i.e., statement code inconsistency, static dynamic inconsistency, temporal relation inconsistency, and readability issues) through the analysis of real-world log-centric changes. We …
Generative Ai Adoption And Solvers' Popularity On Supply-Driven Crowdsourcing Platforms: The Dual Role Of Price Signals, Zimeng Zhu, Carol Hsu, Fiona Fui-Hoon Nah, Na Liu
Generative Ai Adoption And Solvers' Popularity On Supply-Driven Crowdsourcing Platforms: The Dual Role Of Price Signals, Zimeng Zhu, Carol Hsu, Fiona Fui-Hoon Nah, Na Liu
Research Collection School Of Computing and Information Systems
Purpose – We investigate the effect of solvers’ adoption of Generative AI (GenAI) on their popularity in a supply-driven crowdsourcing platform. We also examine the impact of price signals as well as their heterogeneous impact based on the solvers’ membership duration on the platform. Design/methodology/approach – Our analysis focuses on solvers who adopt GenAI for design-related gigs on the supply-driven crowdsourcing platform. By combining propensity score matching (PSM) with multi-period difference-in-differences (DID), we examine how GenAI adoption impacts solvers’ popularity and how price signals affect this main effect. Findings – Our findings reveal that solvers who adopt GenAI tend to …
Spatialimaginer: Towards Adaptive Visual Imagination For Spatial Reasoning, Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, Yu-Gang Jiang
Spatialimaginer: Towards Adaptive Visual Imagination For Spatial Reasoning, Yian Li, Yang Jiao, Bin Zhu, Tianwen Qian, Shaoxiang Chen, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Spatial intelligence, which refers to the ability to reason about geometric and physical structure from visual observations, remains a core challenge for multimodal large language models. Despite promising performance, recent multimodal large language models (MLLMs) often exhibit fragile reasoning traces in spatial intelligence tasks that involve consistent spatial state recognition. We argue that these failures stem from a mismatch between the spatial recognition mechanism and the text-only reasoning behavior of these MLLMs. Effective spatial reasoning requires low-level geometric structure to be faithfully preserved and updated throughout the reasoning process, whereas textual representations tend to abstract away precisely these critical details. …
Semantic-Structural Decoupling: Disentangling Semantic Attention From Structural Bias In The Attention Manifold, Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-Gang Jiang
Semantic-Structural Decoupling: Disentangling Semantic Attention From Structural Bias In The Attention Manifold, Pengkun Jiao, Bin Zhu, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifically, MLLMs consistently exhibit disproportionate attention toward certain semantically uninformative visual tokens, a phenomenon termed "register" or "Visual Attention Sinks." While existing inference intervention methods attempt to identify these sink tokens and redistribute their attention weights, such approaches typically treat these tokens in isolation and suffer from computational inefficiency. Instead, we reframe this phenomenon as a generalized textual bias exerted over visual features that extends beyond isolated sink tokens. From this perspective, a pervasive structural bias leads to the dilution of the …
Ddor: Delta Debugging For Explainable Overrefusal Testing And Repair, Qinyan Zhou, Peixin Zhang, Jun Sun, Haonan Zhang, Dongxia Wang
Ddor: Delta Debugging For Explainable Overrefusal Testing And Repair, Qinyan Zhou, Peixin Zhang, Jun Sun, Haonan Zhang, Dongxia Wang
Research Collection School Of Computing and Information Systems
While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that merely appear risky. We present DDOR (Delta Debugging for OverRefusal), a fully automated and explainable framework for overrefusal testing and repair in a black-box setting, where only model inputs and outputs are accessible and internal safety mechanisms remain opaque. DDOR applies delta debugging to localize minimal refusal-triggering fragments (mRTFs) that provide phrase-level, explainable evidence for why a refusal occurs. Conditioned on these mRTFs, DDOR generates diverse, context-rich prompts and performs multi-oracle validation to filter intrinsically …
Mutation-Based Multi-Agent Test Case Update, Dawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang, Jianlei Chi, Jun Sun, Xiaohong Su
Mutation-Based Multi-Agent Test Case Update, Dawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang, Jianlei Chi, Jun Sun, Xiaohong Su
Research Collection School Of Computing and Information Systems
Modern software systems evolve rapidly under CI/CD practices, where tests are critical for quality. However, substantial code changes often render existing test cases obsolete, causing pipeline disruptions, reduced productivity, and compromised quality. Recent automatic test update approaches leverage LLMs to refine test cases via execution feedback and exact-matching context retrieval, prioritizing executability and line coverage but suffering three limitations: (1) neglecting test assertion adequacy, weakening fault detection; (2) relying on coarse line coverage instead of specific uncovered lines/branches; (3) using exact-matching retrieval, which fails for LLM hallucinated queries. To address these, we propose MuMuTestUp, a mutation-guided multi-agent framework with three …
Learning 1-Bit Lidar-Based Localization With Auxiliary Objective, Kaijie Yin, Zhiyuan Zhang, Tian Gao, Wentao Zhu, Cheng-Zhong Xu, Hui Kong
Learning 1-Bit Lidar-Based Localization With Auxiliary Objective, Kaijie Yin, Zhiyuan Zhang, Tian Gao, Wentao Zhu, Cheng-Zhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
6-DoF LiDAR-based localization is a fundamental capability for autonomous systems operating in large-scale outdoor environments. Many deep-learning-based localization methods have achieved promising performance so far. However, as one of the always-on modules competing for limited on-board computational resources, the localization module is expected to consume only a small portion of the overall compute budget. Most existing learning-based methods are still too heavy for this purpose. In contrast, binary neural networks (BNNs) offer an appealing solution, but the 1-bit compression causes severe information loss and performance drop. In this paper, we address this challenge by proposing Binarized LiDAR-based Localization (BiLoc), the …
Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
Restoring Linguistic Grounding In Vla Models Via Train-Free Attention Recalibration, Ninghao Zhang, Bin Zhu, Shijie Zhou, Jingjing Chen
Research Collection School Of Computing and Information Systems
Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-Of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language–action coupling …
Generalized Logit Adjustment: Improved Fine-Tuning By Mitigating Label Bias In Zero-Shot Vision Models, Beier Zhu, Qianru Sun, Xun Yang, Hanwang Zhang
Generalized Logit Adjustment: Improved Fine-Tuning By Mitigating Label Bias In Zero-Shot Vision Models, Beier Zhu, Qianru Sun, Xun Yang, Hanwang Zhang
Research Collection School Of Computing and Information Systems
Foundation models like CLIP allow zero-shot transfer on various tasks without additional training data. Yet, the zero-shot performance is less competitive than a fully supervised one. Thus, fine-tuning and ensembling are also commonly adopted to better fit the downstream tasks. However, we argue that such prior work has overlooked the inherent biases in foundation models. Due to the highly imbalanced Web-scale training set, foundation models are inevitably skewed toward frequent semantics, and thus the subsequent fine-tuning or ensembling is still biased. In this study, we systematically examine the biases in foundation models and demonstrate the efficacy of our proposed Generalized …
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Efficient Test-Time Retrieval Augmented Generation, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Although Large Language Models (LLMs) demonstrate significant capabilities, their reliance on parametric knowledge often leads to inaccuracies. Retrieval Augmented Generation (RAG) mitigates this by incorporating external knowledge, but these methods may introduce irrelevant retrieved documents, leading to inaccurate responses. While the integration methods filter out incorrect answers from multiple responses, but lack external knowledge like RAG methods, and their high costs require balancing overhead with performance gains. To address these issues, we propose an Efficient Test-Time Retrieval-Augmented Generation Framework named ET2RAG to improve the performance of LLMs while maintaining efficiency. Specifically, ET2RAG is a training-free method, that first retrieves the …
Dynamic Spectral Denoising With Global-Context Attention For Multi-Behavior Recommendation, Miaomiao Cai, Yunshan Ma, Fangqi Zhu, Junfeng Fang, Zhijie Zhang, Zhiyong Cheng, Xiang Wang, See-Kiong Ng
Dynamic Spectral Denoising With Global-Context Attention For Multi-Behavior Recommendation, Miaomiao Cai, Yunshan Ma, Fangqi Zhu, Junfeng Fang, Zhijie Zhang, Zhiyong Cheng, Xiang Wang, See-Kiong Ng
Research Collection School Of Computing and Information Systems
Multi-behavior recommendation improves target-behavior predic-tion by exploiting heterogeneous auxiliary feedback (e.g., view,collect, and cart), yet its robustness is often undermined by behavior-dependent noise and inconsistency. We argue that the key bottle-neck is not merely noisy behaviors, but a representation-level failurecaused by two coupled heterogeneities. First, intra-behavior rep-resentation entanglement arises when multi-hop propagationblends incidental signals with true preferences in the embeddingspace. This entanglement renders coarse spatial denoising inef-fective, since it cannot suppress noise without sacrificing weak-but-informative niche signals. Second, inter-behavior reliabilityheterogeneity complicates cross-behavior fusion, as the predic-tive value of auxiliary behaviors varies substantially across usersand contexts. Without reliability calibration, aggregation can …
Task-Aligned Haze Removal With Semantic-Aware Fusion And Contrast Self-Correction, Jinbin Wang, Aiping Yang, Guosong Jiang, Wenlong Yu, Dongwei Ren, Qinghua Hu
Task-Aligned Haze Removal With Semantic-Aware Fusion And Contrast Self-Correction, Jinbin Wang, Aiping Yang, Guosong Jiang, Wenlong Yu, Dongwei Ren, Qinghua Hu
Research Collection School Of Computing and Information Systems
Adverse haze conditions introduce complex degradations that obscure scene details and distort structural cues critical for object detection, posing persistent challenges for vision‐based sensing systems. Although existing haze removal methods have achieved notable improvements in visual clarity, their optimisation objectives are often misaligned with downstream detection requirements, leading to limited detection performance in real‐world scenarios. To address this issue, this work proposes a task‐aligned weakly supervised haze removal framework, termed Dehaze4Detection, which explicitly aligns low‐level restoration with high‐level detection objectives. The framework incorporates a Semantic‐Aware Multi‐Scale Fusion Module (SMFM) that embeds pixel‐level semantic knowledge into the dehazing process, enabling selective …
Left: Learnable Fusion Of Tri-View Tokens For Unsupervised Time Series Anomaly Detection, Dezheng Wang, Tong Chen, Guansong Pang, Congyan Chen, Shihua Li, Hongzhi Yin
Left: Learnable Fusion Of Tri-View Tokens For Unsupervised Time Series Anomaly Detection, Dezheng Wang, Tong Chen, Guansong Pang, Congyan Chen, Shihua Li, Hongzhi Yin
Research Collection School Of Computing and Information Systems
As a fundamental data mining task, unsupervised time series anomaly detection (TSAD) aims to build a model for identifying abnormal timestamps without assuming the availability of annotations. A key challenge in unsupervised TSAD is that many anomalies are too subtle to exhibit detectable deviation in any single view (e.g., time domain), and instead manifest as inconsistencies across multiple views like time, frequency, and a mixture of resolutions. However, most cross-view methods rely on feature or score fusion and do not enforce analysis–synthesis consistency, meaning the frequency branch is not required to reconstruct the time signal through an inverse transform, and …
Hvi-Cidnet+: Beyond Extreme Darkness For Low-Light Image Enhancement, Kangbiao Shi, Xiaowen Ma, Yixu Feng, Tao Hu, Peng Wu, Guansong Pang, Qingsen Yan
Hvi-Cidnet+: Beyond Extreme Darkness For Low-Light Image Enhancement, Kangbiao Shi, Xiaowen Ma, Yixu Feng, Tao Hu, Peng Wu, Guansong Pang, Qingsen Yan
Research Collection School Of Computing and Information Systems
Low-Light Image Enhancement (LLIE) aims to recover visually pleasing content and details from degraded low-light images. However, existing RGB-based methods often suffer from color bias and brightness artifacts due to inherent high color sensitivity. Although the HSV color space can decouple brightness and color, it introduces noticeable red and black noise artifacts. To address these challenges, we adopt the Horizontal/Vertical-Intensity (HVI) color space for LLIE, which is defined by the HV color map and learnable intensity. The former enforces small distances for red coordinates to alleviate red noise artifacts, while the latter adaptively compresses low-light regions to suppress black noise …
Timeradar: A Domain-Rotatable Foundation Model For Time Series Anomaly Detection, Hui He, Hezhe Qiao, Yutong Chen, Kun Yi, Guansong Pang
Timeradar: A Domain-Rotatable Foundation Model For Time Series Anomaly Detection, Hui He, Hezhe Qiao, Yutong Chen, Kun Yi, Guansong Pang
Research Collection School Of Computing and Information Systems
Current time series foundation models (TSFMs) primarily focus on learning prevalent and regular patterns within a predefined time or frequency domain to enable supervised downstream tasks (\eg, forecasting). Consequently, they are often ineffective for inherently unsupervised downstream tasks—such as time series anomaly detection (TSAD), which aims to identify rare, irregular patterns. This limitation arises because such abnormal patterns can closely resemble the regular patterns when presented in the same time/frequency domain. To address this issue, we introduce TimeRadar, an innovative TSFM built in a fractional time–frequency domain to support generalist TSAD across diverse unseen datasets. Our key insight is that …
Approximation And Learning-Based Algorithms For Influence Maximization In Multilayer Social Networks, Xueqin Chang, Ruize Liu, Qing Liu, Baihua Zheng, Yunjun Gao
Approximation And Learning-Based Algorithms For Influence Maximization In Multilayer Social Networks, Xueqin Chang, Ruize Liu, Qing Liu, Baihua Zheng, Yunjun Gao
Research Collection School Of Computing and Information Systems
Motivated by the observation that users in the real world often engage across multiple social networks simultaneously, we study the problem of influence maximization in multilayer social networks (Mlim), aiming to select a small set of nodes that maximizes the total influence spread across all layers. To this end, we introduce a hybrid propagation model that jointly captures layer-specific diffusion dynamics and probabilistic cross-layer propagation. Based on this model, we formally define the Mlim problem and establish its NP-hardness, monotonicity, and submodularity. To address the Mlim problem, we first propose a greedy baseline Mlim-Greedy, which achieves a (1-1/e) approximation. Since …
Audeter: A Large-Scale Dataset For Deepfake Audio Detection In Open Worlds, Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie
Audeter: A Large-Scale Dataset For Deepfake Audio Detection In Open Worlds, Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie
Research Collection School Of Computing and Information Systems
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving distribution shifts between training and test data, driven by the complexity of human speech and the rapid evolution of synthesis systems. Existing datasets suffer from limited real speech diversity, insufficient coverage of recent synthesis systems, and heterogeneous mixtures of deepfake sources, which hinder systematic evaluation and open-world model training. To address these issues, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale and highly diverse deepfake audio dataset comprising over 4,500 …
Efficient And Universal Watermarking For Llm-Generated Code Detection, Boquan Li, Zirui Fu, Mengdi Zhang, Peixin Zhang, Jun Sun, Xingmei Wang
Efficient And Universal Watermarking For Llm-Generated Code Detection, Boquan Li, Zirui Fu, Mengdi Zhang, Peixin Zhang, Jun Sun, Xingmei Wang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) have significantly enhanced the usability of AI-generated code, providing effective assistance to programmers. This advancement also raises ethical and legal concerns, such as academic dishonesty and the generation of malicious code. For accountability, it is imperative to detect whether a piece of code is AI-generated. Watermarking is broadly considered a promising solution and has been successfully applied to identify LLM-generated text. However, existing efforts on code are far from ideal, suffering from limited universality and excessive time and memory consumption. In this work, we propose a plugand- play watermarking approach for AI-generated code detection, named ACW …
Continuous Query For Top-K Maximal Sum Intervals Over Streaming Data, Zhongshuai Zhang, Xiaochun Yang, Baihua Zheng, Rui Zhu, Haomin Li, Bin Wang
Continuous Query For Top-K Maximal Sum Intervals Over Streaming Data, Zhongshuai Zhang, Xiaochun Yang, Baihua Zheng, Rui Zhu, Haomin Li, Bin Wang
Research Collection School Of Computing and Information Systems
The continuous identification of top-k maximal sum intervals using a sliding window over a data stream is a critical operation for applications in IoT and beyond. A maximal sum interval is a non-overlapping, contiguous subsequence with the maximal sum in a sequence of signed values. Existing algorithms are ill-suited for streaming contexts: they either exhaustively enumerate all intervals even for small k values, or depend on indexes that require frequent and costly restructuring. We propose a novel partition-based strategy. Our core insight is a partitioning scheme that guarantees that any maximal sum interval is fully contained within a single partition, …
Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang
Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …
Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun
Rendering Data Unlearnable By Exploiting Llm Alignment Mechanisms, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Large language models (LLMs) are increasingly trained on massive, heterogeneous text corpora, raising serious concerns about the unauthorised use of proprietary or personal data during model training. In this work, we address the problem of data protection against unwanted model learning in a realistic blackbox setting. We propose Disclaimer Injection, a novel data-level defence that renders text unlearnable to LLMs. Rather than relying on model-side controls or explicit data removal, our approach exploits the models’ own alignment mechanisms: injecting carefully designed alignment-triggers to prevent effective learning. Through layer-wise analysis, we find that finetuning on such protected data induces persistent activation …
Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou
Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou
Research Collection School Of Computing and Information Systems
Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws MCMC samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using …
Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun
Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun
Research Collection School Of Computing and Information Systems
Addressing itinerary modification is crucial for enhancing the travel experience as it is a frequent requirement during traveling. However, existing research mainly focuses on fixed itinerary planning, leaving modification underexplored due to the scarcity of shape need-to-modify itinerary data. To bridge this gap, we formally define the itinerary modification task and propose a general pipeline to construct the corresponding dataset, namely iTIMO. This pipeline frames the generation of shape need-to-modify itinerary data as an intent-driven perturbation task. It instructs large language models to perturb real-world itineraries using three operations: REPLACE, ADD, and DELETE. Each perturbation is grounded in three intents: …
Multimodal Contrastive Spatiotemporal Self-Organizing Neural Networks For In-Home Activity Learning Of Mild Cognitive Impairment, Seng Khoon Teh, Ah-Hwee Tan, Kar Way Tan, Iris Rawtaer
Multimodal Contrastive Spatiotemporal Self-Organizing Neural Networks For In-Home Activity Learning Of Mild Cognitive Impairment, Seng Khoon Teh, Ah-Hwee Tan, Kar Way Tan, Iris Rawtaer
Research Collection School Of Computing and Information Systems
In-home spatiotemporal data, such as the movement trajectory data and the spatial time series data, contains potential predictive utility for detection of geriatric conditions including Mild Cognitive Impairment (MCI), frailty, and cognitive frailty. However, few have explored spatiotemporal learning models for learning and fusion of such disparate spatiotemporal data, owing to the lack of a generalized machine learning model that can jointly model these different spatiotemporal data types. This work reports a multimodal spatiotemporal machine learning model based on a class of self-organizing neural networks that can integrate different spatiotemporal data types for MCI detection. Specifically, Episodic Memory Adaptive Resonance …
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen
Research Collection School Of Computing and Information Systems
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …
Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo
Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo
Research Collection School Of Computing and Information Systems
Learning-based dynamic fault localization techniques play a crucial role in the field of software engineering. These techniques dynamically execute test cases to meticulously extract useful knowledge from the execution information in the program, with the aim of identifying fault locations by leveraging machine learning, deep learning, and large language models. Currently, there is already a flourishing body of research that is intensely focused on learning-based dynamic fault localization. Research literature can be categorized into two main aspects for learning-based dynamic fault localization: data-based enhancements (i.e., the datasets) and model-based enhancements (i.e., the suspiciousness algorithms). Thus, we conduct an extensive literature …
Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen
Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen
Research Collection School Of Computing and Information Systems
Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and …
Deep Learning For Video Anomaly Detection: A Review, Peng Wu, Chengyu Pan, Yuting Yan, Guansong Pang, Qingsen Yan, Peng Wang, Yanning Zhang
Deep Learning For Video Anomaly Detection: A Review, Peng Wu, Chengyu Pan, Yuting Yan, Guansong Pang, Qingsen Yan, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
Video anomaly detection (VAD) aims to discover behaviors or events deviating from the normality in videos. As a long-standing task in the field of computer vision, VAD has witnessed much good progress. In the era of deep learning, with the explosion of architectures of continuously growing capability and capacity, a great variety of deep learning-based methods are constantly emerging for the VAD task, greatly improving the generalization ability of detection algorithms and broadening the application scenarios. Therefore, such a multitude of methods and a large body of literature make a comprehensive survey a pressing necessity. In this article, we present …