Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year
File Type

Articles 31 - 60 of 1664

Full-Text Articles in Artificial Intelligence and Robotics

A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes Jul 2026

A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes

Research Collection School Of Computing and Information Systems

Agentic AI coding tools such as Claude Code and OpenAI Codex execute multi-step coding tasks with limited human oversight. To steer these tools, developers create repository-level configuration artifacts (e.g., Markdown files) for configuration mechanisms such as Context Files, Skills, Rules, and Hooks. There is no curated dataset yet that captures these configurations at scale. This dataset, collected from open-source GitHub repositories, fills that gap. We selected 40,585 actively maintained repositories through metadata filtering, classified them using GPT-5.2 to identify 36,710 as belonging to engineered software projects, and systematically detected configuration artifacts in these repositories. The dataset covers 4,738 repositories across …


Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma Jul 2026

Dual-Diffusional Generative Fashion Recommendation, Mingzhe Yu, Lei Wu, Qianru Sun, Yunshan Ma

Research Collection School Of Computing and Information Systems

Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user …


Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang Jul 2026

Spatiotemporal Sycophancy: Negation-Based Gaslighting In Video Large Language Models, Ziyao Tang, Pengkun Jiao, Bin Zhu, Huiyan Qi, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. …


Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen Jul 2026

Oscbench: Benchmarking Object State Change In Text-To-Video Generation, Xianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li, Patrick Carrington, Roger Zimmermann, Jingjing Chen

Research Collection School Of Computing and Information Systems

Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into …


Towards Uniformity And Alignment For Multimodal Representation Learning, Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves Jul 2026

Towards Uniformity And Alignment For Multimodal Representation Learning, Wenzhe Yin, Pan Zhou, Zehao Xiao, Jie Liu, Shujian Yu, Jan-Jakob Sonke, Efstratios Gavves

Research Collection School Of Computing and Information Systems

Multimodal representation learning aims to construct a shared embedding space in which heterogeneous modalities are semantically aligned. Despite strong empirical results, InfoNCE-based objectives introduce inherent conflicts that yield distribution gaps across modalities. In this work, we identify two conflicts in the multimodal regime, both exacerbated as the number of modalities increases: (i) an alignment–uniformity conflict, whereby the repulsion of uniformity undermines pairwise alignment, and (ii) an intra-alignment conflict, where aligning multiple modalities induces competing alignment directions. To address these issues, we propose a principled decoupling of alignment and uniformity for multimodal representations, providing a conflict-free recipe for multimodal learning that …


Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou Jul 2026

Variational Speculative Decoding: Rethinking Draft Training From Token Likelihood To Sequence Acceptance, Xiandong Zou, Jianshu Li, Jing Huang, Pan Zhou

Research Collection School Of Computing and Information Systems

Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize single greedy trajectories, decoding involves verifying and ranking multiple sampled draft paths. We propose Variational Speculative Decoding (VSD), formulating draft training as variational inference over latent proposals (draft paths). VSD maximizes the marginal probability of target-model acceptance, yielding an ELBO that promotes high-quality latent proposals while minimizing divergence from the target distribution. To enhance quality and reduce variance, we incorporate a path-level utility and optimize via an Expectation-Maximization procedure. The E-step draws MCMC samples from an oracle-filtered posterior, while the M-step maximizes weighted likelihood using …


Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen Jul 2026

Air: Improving Agent Safety Through Incident Response, Zibo Xiao, Jun Sun, Junjie Chen

Research Collection School Of Computing and Information Systems

Large Language Model (LLM) agents are increasingly deployed in practice across a wide range of autonomous applications. Yet current safety mechanisms for LLM agents focus almost exclusively on preventing failures in advance, providing limited capabilities for responding to, containing, or recovering from incidents after they inevitably arise. In this work, we introduce AIR, the first incident response framework for LLM agent systems. AIR defines a domain-specific language for managing the incident response lifecycle autonomously in LLM agent systems, and integrates it into the agent's execution loop to (1) detect incidents via semantic checks grounded in the current environment state and …


Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen Jul 2026

Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen

Research Collection School Of Computing and Information Systems

The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and …


Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun Jul 2026

Itimo: An Llm-Empowered Synthesis Dataset For Travel Itinerary Modification, Zhuoxuan Huang, Yunshan Ma, Hongyu Zhang, Hua Ma, Zhu Sun

Research Collection School Of Computing and Information Systems

Addressing itinerary modification is crucial for enhancing the travel experience as it is a frequent requirement during traveling. However, existing research mainly focuses on fixed itinerary planning, leaving modification underexplored due to the scarcity of shape need-to-modify itinerary data. To bridge this gap, we formally define the itinerary modification task and propose a general pipeline to construct the corresponding dataset, namely iTIMO. This pipeline frames the generation of shape need-to-modify itinerary data as an intent-driven perturbation task. It instructs large language models to perturb real-world itineraries using three operations: REPLACE, ADD, and DELETE. Each perturbation is grounded in three intents: …


Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma Jul 2026

Scattered Hypothesis Generation For Open-Ended Event Forecasting, He Chang, Zhulin Tao, Lifang Yang, Xianglin Huang, Yunshan Ma

Research Collection School Of Computing and Information Systems

Despite the importance of open-ended event forecasting for risk management, current LLM-based methods predominantly target only the most probable outcomes, neglecting the intrinsic uncertainty of real-world events. To bridge this gap, we advance open-ended event forecasting from pinpoint forecasting to scatter forecasting by introducing the proxy task of hypothesis generation. This paradigm aims to generate an inclusive and diverse set of hypotheses that broadly cover the space of plausible future events. To this end, we propose SCATTER, a reinforcement learning framework that jointly optimizes inclusiveness and diversity of the hypothesis. Specifically, we design a novel hybrid reward that consists of …


Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang Jul 2026

Mab-Dqa: Addressing Query Aspect Importance In Document Question Answering With Multi-Armed Bandits, Yixin Xiang, Yunshan Ma, Xiaoyu Du, Yibing Chen, Yanxin Zhang, Jinhui Tang

Research Collection School Of Computing and Information Systems

Document Question Answering (DQA) involves generating answers from a document based on a user’s query, representing a key task in document understanding. This task requires interpreting visual layouts, which has prompted recent studies to adopt multimodal Retrieval-Augmented Generation (RAG) that processes page images for answer generation. However, in multimodal RAG, visual DQA struggles to utilize a large number of images effectively, as the retrieval stage often retains only a few candidate pages (e.g., Top-4), causing informative but less visually salient content to be overlooked in favor of common yet low-information pages. To address this issue, we propose a Multi-Armed Bandit–based …


Co-Matching: Towards Human–Model Collaborative Legal Case Matching, Chen Huang, Xinwei Yang, Yang Deng, Wenqiang Lei, Jiancheng Lv, Tat-Seng Chua Jul 2026

Co-Matching: Towards Human–Model Collaborative Legal Case Matching, Chen Huang, Xinwei Yang, Yang Deng, Wenqiang Lei, Jiancheng Lv, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Recent efforts have aimed to improve AI models in legal case matching by integrating legal domain knowledge. However, successful legal case matching requires the tacit knowledge of legal practitioners, which is difficult to verbalize and encode into models. This emphasizes the crucial role of involving legal practitioners in high-stakes legal case matching. To address this, we propose a collaborative matching framework called Co-Matching, which encourages both the model and the legal practitioner to participate in the matching process, integrating tacit knowledge. Unlike existing methods that rely solely on the model, Co-Matching allows both the legal practitioner and the model to …


Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo Jul 2026

Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo

Research Collection School Of Computing and Information Systems

Learning-based dynamic fault localization techniques play a crucial role in the field of software engineering. These techniques dynamically execute test cases to meticulously extract useful knowledge from the execution information in the program, with the aim of identifying fault locations by leveraging machine learning, deep learning, and large language models. Currently, there is already a flourishing body of research that is intensely focused on learning-based dynamic fault localization. Research literature can be categorized into two main aspects for learning-based dynamic fault localization: data-based enhancements (i.e., the datasets) and model-based enhancements (i.e., the suspiciousness algorithms). Thus, we conduct an extensive literature …


Multimodal Contrastive Spatiotemporal Self-Organizing Neural Networks For In-Home Activity Learning Of Mild Cognitive Impairment, Seng Khoon Teh, Ah-Hwee Tan, Kar Way Tan, Iris Rawtaer Jul 2026

Multimodal Contrastive Spatiotemporal Self-Organizing Neural Networks For In-Home Activity Learning Of Mild Cognitive Impairment, Seng Khoon Teh, Ah-Hwee Tan, Kar Way Tan, Iris Rawtaer

Research Collection School Of Computing and Information Systems

In-home spatiotemporal data, such as the movement trajectory data and the spatial time series data, contains potential predictive utility for detection of geriatric conditions including Mild Cognitive Impairment (MCI), frailty, and cognitive frailty. However, few have explored spatiotemporal learning models for learning and fusion of such disparate spatiotemporal data, owing to the lack of a generalized machine learning model that can jointly model these different spatiotemporal data types. This work reports a multimodal spatiotemporal machine learning model based on a class of self-organizing neural networks that can integrate different spatiotemporal data types for MCI detection. Specifically, Episodic Memory Adaptive Resonance …


Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning, Brahmanage Janaka Chathuranga Thilakarathna, Akshat Kumar Jul 2026

Beyond Hard Constraints: Budget-Conditioned Reachability For Safe Offline Reinforcement Learning, Brahmanage Janaka Chathuranga Thilakarathna, Akshat Kumar

Research Collection School Of Computing and Information Systems

Sequential decision-making using Markov Decision Process underpins many real-world applications. Both model-based and model-free methods have achieved strong results in these settings. However, real-world tasks must balance reward maximization with safety constraints, often conflicting objectives, that can lead to unstable min–max, adversarial optimization. A promising alternative is safety reachability analysis, which precomputes a forward-invariant safe state–action set, ensuring that an agent starting inside this set remains safe indefinitely. Yet, most reachability-based methods address only hard safety constraints, and little work extends reachability to cumulative cost constraints. To address this, first, we define a safety-conditioned reachability set that decouples reward maximization …


Accountable Agents In Software Engineering: An Analysis Of Terms Of Service And A Research Roadmap, Christoph Treude Jul 2026

Accountable Agents In Software Engineering: An Analysis Of Terms Of Service And A Research Roadmap, Christoph Treude

Research Collection School Of Computing and Information Systems

AI coding assistants and autonomous agents are becoming integral to software development workflows, reshaping how code is produced, reviewed, and maintained. While recent research has focused mainly on the capabilities and impacts of productivity of these systems, much less attention has been paid to accountability: who is responsible when agents generate, modify, or recommend code? In practice, accountability is defined through the Terms of Service (ToS) and related policy documents that govern the use of AI-powered development tools.In this vision paper, we present a comparative analysis of the Terms of Service for widely used AI coding assistants and agent-enabled development …


Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li Jun 2026

Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li

Research Collection School Of Computing and Information Systems

Personalized outfit recommendation poses a significant challenge in e-commerce and social media platforms, requiring systems that balance user preferences with aesthetic compatibility. Collaborative filtering (CF) provides a traditional solution for this, but it struggles with data-sparse scenarios and complex user-item-outfit relationships. Meanwhile, existing template-based approaches are constrained by rigid pre-designed structures. To bridge these research gaps, we introduce CFALR (Collaborative Filtering-Augmented Large Language Model for Recommendation), a novel framework that synergizes collaborative filtering with large language models for personalized outfit recommendation. Specifically, CFALR describes user-outfit interactions in natural language and leverages LLMs to capture fashion semantics while employing CF-enhanced embeddings …


Scaling Up Multi-Agent Reinforcement Learning For Large Agent Teams And Long-Horizon Tasks: A Survey, Minghong Geng, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan Jun 2026

Scaling Up Multi-Agent Reinforcement Learning For Large Agent Teams And Long-Horizon Tasks: A Survey, Minghong Geng, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Multi-agent reinforcement learning (MARL) empowers multiple autonomous agents to acquire effective policies for collaborative problem-solving. Over the last decade, MARL has seen significant advancements, with numerous algorithms achieving impressive performance across various benchmarks and real-world applications. Nevertheless, the scalability of multi-agent systems, in terms of the number of agents and the length of the task horizon, remains a critical consideration for applying MARL methods to complex problem-solving. Given that a dedicated review of the existing approaches and challenges in scaling up multi-agent systems remains largely absent, this survey aims to bridge this gap by delivering a comprehensive review of MARL …


Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du Jun 2026

Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du

Research Collection School Of Computing and Information Systems

Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation (MVR) due to their strong multimodal understanding. However, existing apporches typically deploy LVLMs as fixed black-box feature extractors without systematically comparing alternative representation strategies. To address this gap, we present the first systematic empirical study on various feature extraction paradigms and integration strategies, along with hierarchical representations from frozen LVLMs for MVR. Extensive experiments on representative LVLMs reveal that hidden states from multiple decoder layers provide richer and more effective representations for MVR. Guided by this insight, we propose the Dual Feature Fusion (DFF) Framework, a lightweight approach …


Language Embeddings Meet Shallow Autoencoders, Rodrigo Alves, Vojtěch Vančura, Pavel Kordík, Antoine Ledent Jun 2026

Language Embeddings Meet Shallow Autoencoders, Rodrigo Alves, Vojtěch Vančura, Pavel Kordík, Antoine Ledent

Research Collection School Of Computing and Information Systems

Shallow autoencoders are appealing recommenders due to their simplicity, scalability, and competitive retrieval quality, but they struggle in strict cold-start settings where new items have no interactions. We propose an inductive shallow autoencoder that leverages item side information (language embeddings) by fixing the decoder to item features and learning only an encoder in the same semantic space. To prevent trivial self-reconstruction without enforcing a hard zero diagonal, we introduce diagonal gating: a leave-one-item-out objective that blocks the self-copy shortcut only for the item being updated while retaining context from the rest of the user history. An alternating-style optimization trains the …


History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu Jun 2026

History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu

Research Collection School Of Computing and Information Systems

Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …


“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt Jun 2026

“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt

Research Collection School Of Computing and Information Systems

Due to their limited ability to reason about the social context in which they are used, smart speakers pose significant privacy risks by responding in ways that may violate people's implicit social boundaries. We conducted a cross-cultural vignette study (N = 944) in Germany and Singapore to investigate how situational factors—specifically social context (bystander relationships and closeness), physical context (location), and interaction context (topic and deceptive intent)—regulate user preferences for smart speaker responses. Our results demonstrate that these factors are superior predictors of response preferences than dispositional user traits (i.e., intrinsic personal traits). We identify two distinct social dynamics: a …


How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández Jun 2026

How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández

Research Collection School Of Computing and Information Systems

The proliferation of Machine Learning (ML) models and their open source implementations has transformed AI research and applications. Platforms like Hugging Face (HF) enable this evolving ecosystem, yet a large-scale longitudinal study of how these models change is lacking. This study addresses this gap by analyzing over 680,000 commits from 100,000 models and 2,251 releases from 202 of these models on HF using repository mining and longitudinal methods. We apply an extended ML change taxonomy to classify commits and use Bayesian networks to model temporal patterns in commit and release activities. Our findings show that commit activities align with established …


Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang Jun 2026

Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang

Research Collection School Of Computing and Information Systems

Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …


Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia Jun 2026

Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia

Research Collection School Of Computing and Information Systems

Text-to-image (T2I) generative models are increasingly used to produce content for education, media, and public-facing communication, and are starting to be integrated into higher-impact pipelines. Since generated images tend to reinforce stereotypes, producing representational erasure via “default” depictions and shaping perceptions of who belongs in certain roles, a growing body of work has proposed metrics to quantify gender bias in T2I outputs. Yet existing evaluations remain fragmented. Metrics are often reported without a shared view of what they measure, what assumptions they entail, or how their results should be interpreted under different deployment contexts. This limits the usefulness of gender …


Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le Jun 2026

Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le

Research Collection School Of Computing and Information Systems

A common collocated group setting in mixed-reality (MR) collaboration is a person wearing a MR headset (HMD user) and presenting MR contents to audiences who are not provided with such specialized devices (Non-HMD users). In this setting, while Non-HMD users can view the MR environment shown on a large physical display, it still remains challenging for the HMD user to interpret their pointing gesture when they spatially refer to objects in the MR environment. To address this, we designed and evaluated two pointing techniques—SCREEN and SCREEN+SPACE—that support Non-HMD users in referring to MR content. Screen pointing allows users to refer …


Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang Jun 2026

Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to operate reliably in dynamic environments and often fail under Out-of-Distribution (OOD) conditions. To address this issue, we propose Robot-Conditioned Normalizing Flow(RC-NF), a real-time monitoring model for robotic anomaly detection and intervention that ensures the robot's state and the object's motion trajectory align with the task. RC-NF decouples the processing of task-aware robot and object states within the normalizing flow. It requires only positive samples for unsupervised training and calculates accurate robotic anomaly scores during inference through the …


Anatomical Domain Shifts: Test-Time Heterogeneous Adaptation For 3d Human Pose Prediction, Qiongjie Cui, Pan Zhou, Jingjing Chen, Na Zhao Jun 2026

Anatomical Domain Shifts: Test-Time Heterogeneous Adaptation For 3d Human Pose Prediction, Qiongjie Cui, Pan Zhou, Jingjing Chen, Na Zhao

Research Collection School Of Computing and Information Systems

The research frontier in human pose prediction (HPP) is advancing toward continual test-time adaptation (TTA), where models must self-adapt to dynamic test distributions. To date, the homeostatic continual TTA remains the sole viable solution, which isolates the model parameters and update domain-sensitive ones. Despite mitigating full-body domain gaps, human anatomical heterogeneity (domain shifts often localize to specific regions) is ignored. This anatomical-agnostic approach forces uniform parameter adaptation across kinematically distinct segments, causing: over-adaptation of stable regions and under-adaptation of shift-prone articulations. To address it, we introduce TT-HA, a novel Test-Time Heterogeneous Adaptation that implicitly estimates domain changes for anatomical segments, …


Adaptive Outlier Detection Over Data Stream, Rui Zhu, Mingyuan Jiang, Xiaochun Yang, Baihua Zheng, Bin Wang, Tao Qiu Jun 2026

Adaptive Outlier Detection Over Data Stream, Rui Zhu, Mingyuan Jiang, Xiaochun Yang, Baihua Zheng, Bin Wang, Tao Qiu

Research Collection School Of Computing and Information Systems

Continuous distance-based outlier detection in streaming data poses significant challenges and has a wide range of practical applications. Traditional threshold-based methods perform well under stable streaming conditions, where fixed parameters remain effective. However, they often struggle with dynamic data distributions and high stream speeds, leading to suboptimal performance, limited control over the number of returned outliers, and failure to meet real-time detection requirements. To address these issues, this paper introduces a novel Recall and Proportion-Aware Outlier Detection (RPA-OD) query. In RPA-OD, ρ defines a distance relaxation that enables real-time outlier detection. Specifically, objects with fewer than k neighbors within the …


Task Complexity Matters: An Empirical Study Of Reasoning In Llms For Sentiment Analysis, Donghao Huang, Zhaoxia Wang Jun 2026

Task Complexity Matters: An Empirical Study Of Reasoning In Llms For Sentiment Analysis, Donghao Huang, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations across seven model families—including adaptive, conditional, and reinforcement learning-based reasoning architectures—on sentiment analysis datasets of varying granularity (binary, five-class, and 27-class emotion). Our findings reveal that reasoning effectiveness is strongly task-dependent, challenging prevailing assumptions: (1) Reasoning shows task-complexity dependence—binary classification degrades up to -19.9 F1% points (pp), while 27-class emotion recognition gains up to  +16.0 pp; (2) Distilled reasoning variants underperform base models by 3–18 pp on simpler tasks, …