Rode: Linear Rectified Mixture Of Diverse Experts For Food Large Multi-Modal Models,
2026
Singapore Management University
Rode: Linear Rectified Mixture Of Diverse Experts For Food Large Multi-Modal Models, Pengkun Jiao, Xinlan Wu, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang
Research Collection School Of Computing and Information Systems
Large Multi-modal Models (LMMs) have significantly advanced a variety of vision-language tasks. The scalability and availability of high-quality training data play a pivotal role in the success of LMMs. In the realm of food, while comprehensive food datasets such as Recipe1M offer an abundance of ingredient and recipe information, they often fall short of providing ample data for nutritional analysis. The Recipe1M+ dataset, despite offering a subset for nutritional evaluation, is limited in the scale and accuracy of nutrition information. To bridge this gap, we introduce Uni-Food, a unified food dataset that comprises over 100,000 images with various food labels, …
Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models,
2026
Singapore Management University
Benchmarking Gaslighting Negation Attacks Against Multimodal Large Language Models, Bin Zhu, Yinxuan Gui, Huiyan Qi, Jingjing Chen, Chong-Wah Ngo, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Multimodal Large Language Models (MLLMs) have exhibited remarkable advancements in integrating different modalities, excelling in complex understanding and generation tasks. Despite their success, MLLMs remain vulnerable to conversational adversarial inputs. In this paper, we systematically study gaslighting negation attacks—a phenomenon where models, despite initially providing correct answers, are persuaded by user-provided negations to reverse their outputs, often fabricating justifications. We conduct extensive evaluations of state-of-the-art MLLMs across diverse benchmarks and observe substantial performance drops when negation is introduced. Notably, we introduce the first benchmark GaslightingBench, specifically designed to evaluate the vulnerability of MLLMs to negation arguments. GaslightingBench consists of multiple-choice …
Videocreator: An Agentic System For Multi-Turn Video Production,
2026
Singapore Management University
Videocreator: An Agentic System For Multi-Turn Video Production, Zhengyang Liang, Yan Shu, Cathal Gurrin, Nicu Sebe, Lizi Liao
Research Collection School Of Computing and Information Systems
Recent advances in video generation models enable visually compelling single clips. However, real-world video creation is inherently continuous and iterative: creators refine content over multiple rounds while maintaining narrative, style, and entity consistency. Existing standalone generators are largely stateless and lack memory of previously generated segments, making it difficult to produce a coherent and consistent video project. To address this gap, we present VideoCreator, a unified video agent that integrates generation and understanding with a project-level memory system. VideoCreator leverages understanding capabilities to perform fine-grained analysis of newly produced content and uses persistent memory to retain and reuse prior context …
Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples,
2026
Singapore Management University
Interfold: Learning Interpretable Diffusion Manifolds Beyond Binary Samples, Alexander Vincent Lewi, Rainer Tan, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose InterFold, a framework for learning and applying interpretable semantic manifolds in latent diffusion models, without requiring binary or paired supervision. Existing methods for semantic editing either rely on limited paired data or uncover only coarse, unsupervised directions that fail to capture user-specific, fine-grained attributes. InterFold addresses these limitations by learning a target attribute manifold in the H-space of diffusion models using only a set of positive, unlabeled examples. To edit a new image, InterFold projects its H-space representation toward this learned manifold through test-time optimization, enabling precise, identity-preserving modifications of complex, non-binary concepts. To make these edits effective …
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation,
2026
Singapore Management University
Sam3-Litetext: An Anatomical Study Of The Sam3 Text Encoder For Efficient Vision-Language Segmentation, Chengxi Zeng, Yuxuan Jiang, Ge Gao, Shuai Wang, Duolikun Danier, Bin Zhu, Stevan Rudinac, David Bull, Fan Zhang
Research Collection School Of Computing and Information Systems
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we …
Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion,
2026
Singapore Management University
Frozen Lvlms For Micro-Video Recommendation: A Systematic Study Of Feature Extraction And Fusion, Huatuan Sun, Yunshan Ma, Changguang Wu, Yanxin Zhang, Pengfei Wang, Xiaoyu Du
Research Collection School Of Computing and Information Systems
Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation (MVR) due to their strong multimodal understanding. However, existing apporches typically deploy LVLMs as fixed black-box feature extractors without systematically comparing alternative representation strategies. To address this gap, we present the first systematic empirical study on various feature extraction paradigms and integration strategies, along with hierarchical representations from frozen LVLMs for MVR. Extensive experiments on representative LVLMs reveal that hidden states from multiple decoder layers provide richer and more effective representations for MVR. Guided by this insight, we propose the Dual Feature Fusion (DFF) Framework, a lightweight approach …
“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts,
2026
Singapore Management University
“Grandpa, Can You Speak Nicer?”: Envisioned Chatbot Roles And Design Tensions In Intergenerational Communication Conflicts, Tianyi Zhang, Emran Bin Elias Poh, Yueyue Hou, Yi-Chieh Lee, Renwen Zhang, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Intergenerational conversations often break down when differences in tone, language, or expectations lead participants to feel dismissed or misunderstood. In this work, we explore how people envision AI-driven chatbot interventions for addressing communication problems in text-based intergenerational family chat. We conducted a scenario-based design interview with 10 pairs of family members from different generations, in which participants designed chatbot interventions that varied in intervention target and timing. Our findings show that participants expect chatbots to perform multiple themes of intervention, including mediating understanding, providing emotional support, offering evaluative commentary, and guiding interaction through behavioral suggestions. These expectations varied systematically across …
Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction,
2026
Singapore Management University
Group Conversational Agents: A Review Of Designs That Support And Shape Group Interaction, Shunyi Yeo, Tianyi Zhang, Scott Bateman, Gary Hsieh, Young-Ho Kim, Simon Tangi Perrault, Jiannan Li, Anthony Tang
Research Collection School Of Computing and Information Systems
Conversational agents that participate in or mediate group interaction introduce challenges that extend beyond supporting individual users, raising new questions about how agents participate in and influence groups. To characterise this emerging design space, we present a systematic review of 53 peer-reviewed studies on group conversational agents (GCAs). We analyse how GCAs intervene in group-level processes, including participation regulation, conflict mediation, task alignment, and execution support. Using concepts from group research as an analytic lens, we organise prior GCA work around recurring group interactional challenges (orientation, conflict, alignment, and execution), and examine the roles agents are designed to play in …
Language Embeddings Meet Shallow Autoencoders,
2026
Singapore Management University
Language Embeddings Meet Shallow Autoencoders, Rodrigo Alves, Vojtěch Vančura, Pavel Kordík, Antoine Ledent
Research Collection School Of Computing and Information Systems
Shallow autoencoders are appealing recommenders due to their simplicity, scalability, and competitive retrieval quality, but they struggle in strict cold-start settings where new items have no interactions. We propose an inductive shallow autoencoder that leverages item side information (language embeddings) by fixing the decoder to item features and learning only an encoder in the same semantic space. To prevent trivial self-reconstruction without enforcing a hard zero diagonal, we introduce diagonal gating: a leave-one-item-out objective that blocks the self-copy shortcut only for the item being updated while retaining context from the rest of the user history. An alternating-style optimization trains the …
Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles,
2026
Singapore Management University
Context Matters: Auditing Gender Bias In T2i Generation Through Risk-Tiered Use-Case Profiles, Jose Luis Luna Campoverde, Yankun Wu, Xiaofei Xie, Noa Garcia
Research Collection School Of Computing and Information Systems
Text-to-image (T2I) generative models are increasingly used to produce content for education, media, and public-facing communication, and are starting to be integrated into higher-impact pipelines. Since generated images tend to reinforce stereotypes, producing representational erasure via “default” depictions and shaping perceptions of who belongs in certain roles, a growing body of work has proposed metrics to quantify gender bias in T2I outputs. Yet existing evaluations remain fragmented. Metrics are often reported without a shared view of what they measure, what assumptions they entail, or how their results should be interpreted under different deployment contexts. This limits the usefulness of gender …
Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation,
2026
Singapore Management University
Rc-Nf: Robot-Conditioned Normalizing Flow For Real-Time Anomaly Detection In Robotic Manipulation, Shijie Zhou, Bin Zhu, Jiarui Yang, Xiangyu Zhao, Jingjing Chen, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Recent advances in Vision-Language-Action (VLA) models have enabled robots to execute increasingly complex tasks. However, VLA models trained through imitation learning struggle to operate reliably in dynamic environments and often fail under Out-of-Distribution (OOD) conditions. To address this issue, we propose Robot-Conditioned Normalizing Flow(RC-NF), a real-time monitoring model for robotic anomaly detection and intervention that ensures the robot's state and the object's motion trajectory align with the task. RC-NF decouples the processing of task-aware robot and object states within the normalizing flow. It requires only positive samples for unsupervised training and calculates accurate robotic anomaly scores during inference through the …
Anatomical Domain Shifts: Test-Time Heterogeneous Adaptation For 3d Human Pose Prediction,
2026
Singapore Management University
Anatomical Domain Shifts: Test-Time Heterogeneous Adaptation For 3d Human Pose Prediction, Qiongjie Cui, Pan Zhou, Jingjing Chen, Na Zhao
Research Collection School Of Computing and Information Systems
The research frontier in human pose prediction (HPP) is advancing toward continual test-time adaptation (TTA), where models must self-adapt to dynamic test distributions. To date, the homeostatic continual TTA remains the sole viable solution, which isolates the model parameters and update domain-sensitive ones. Despite mitigating full-body domain gaps, human anatomical heterogeneity (domain shifts often localize to specific regions) is ignored. This anatomical-agnostic approach forces uniform parameter adaptation across kinematically distinct segments, causing: over-adaptation of stable regions and under-adaptation of shift-prone articulations. To address it, we introduce TT-HA, a novel Test-Time Heterogeneous Adaptation that implicitly estimates domain changes for anatomical segments, …
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation,
2026
Singapore Management University
History To Future: Evolving Agent With Experience And Thought For Zero-Shot Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Zihan Wang, Guo-Sen Xie, Yang Yang, Jinshan Pan, Qianru Sun, Xiangbo Shu
Research Collection School Of Computing and Information Systems
Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve …
Adaptive Outlier Detection Over Data Stream,
2026
Singapore Management University
Adaptive Outlier Detection Over Data Stream, Rui Zhu, Mingyuan Jiang, Xiaochun Yang, Baihua Zheng, Bin Wang, Tao Qiu
Research Collection School Of Computing and Information Systems
Continuous distance-based outlier detection in streaming data poses significant challenges and has a wide range of practical applications. Traditional threshold-based methods perform well under stable streaming conditions, where fixed parameters remain effective. However, they often struggle with dynamic data distributions and high stream speeds, leading to suboptimal performance, limited control over the number of returned outliers, and failure to meet real-time detection requirements. To address these issues, this paper introduces a novel Recall and Proportion-Aware Outlier Detection (RPA-OD) query. In RPA-OD, ρ defines a distance relaxation that enables real-time outlier detection. Specifically, objects with fewer than k neighbors within the …
Task Complexity Matters: An Empirical Study Of Reasoning In Llms For Sentiment Analysis,
2026
Singapore Management University
Task Complexity Matters: An Empirical Study Of Reasoning In Llms For Sentiment Analysis, Donghao Huang, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) with reasoning capabilities have fueled a compelling narrative that reasoning universally improves performance across language tasks. We test this claim through a comprehensive evaluation of 504 configurations across seven model families—including adaptive, conditional, and reinforcement learning-based reasoning architectures—on sentiment analysis datasets of varying granularity (binary, five-class, and 27-class emotion). Our findings reveal that reasoning effectiveness is strongly task-dependent, challenging prevailing assumptions: (1) Reasoning shows task-complexity dependence—binary classification degrades up to -19.9 F1% points (pp), while 27-class emotion recognition gains up to +16.0 pp; (2) Distilled reasoning variants underperform base models by 3–18 pp on simpler tasks, …
A Novel Hierarchical Multi-Agent System For Payments Using Llms,
2026
Singapore Management University
A Novel Hierarchical Multi-Agent System For Payments Using Llms, Donghao Huang, Joon Kiat Chua, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Large language model (LLM) agents, such as OpenAI’s Operator and Claude’s Computer Use, can automate workflows but unable to handle payment tasks. Existing agentic solutions have gained significant attention; however, even the latest approaches face challenges in implementing end-to-end agentic payment workflows. To address this gap, this research proposes the Hierarchical Multi-Agent System for Payments (HMASP), which provides an end-to-end agentic method for completing payment workflows. The proposed HMASP leverages either open-weight or proprietary LLMs and employs a modular architecture consisting of the Conversational Payment Agent (CPA - first agent level), Supervisor agents (second agent level), Routing agents (third agent …
Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration,
2026
Singapore Management University
Enhancing Pointing Gestures Of Non-Hmd Users In Asymmetric Collocated Mixed Reality Collaboration, Nam-Dang Vo, Van-Vinh Thai, Anthony Tang, Khanh-Duy Le
Research Collection School Of Computing and Information Systems
A common collocated group setting in mixed-reality (MR) collaboration is a person wearing a MR headset (HMD user) and presenting MR contents to audiences who are not provided with such specialized devices (Non-HMD users). In this setting, while Non-HMD users can view the MR environment shown on a large physical display, it still remains challenging for the HMD user to interpret their pointing gesture when they spatially refer to objects in the MR environment. To address this, we designed and evaluated two pointing techniques—SCREEN and SCREEN+SPACE—that support Non-HMD users in referring to MR content. Screen pointing allows users to refer …
“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts,
2026
Singapore Management University
“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt
Research Collection School Of Computing and Information Systems
Due to their limited ability to reason about the social context in which they are used, smart speakers pose significant privacy risks by responding in ways that may violate people's implicit social boundaries. We conducted a cross-cultural vignette study (N = 944) in Germany and Singapore to investigate how situational factors—specifically social context (bystander relationships and closeness), physical context (location), and interaction context (topic and deceptive intent)—regulate user preferences for smart speaker responses. Our results demonstrate that these factors are superior predictors of response preferences than dispositional user traits (i.e., intrinsic personal traits). We identify two distinct social dynamics: a …
Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation,
2026
Singapore Management University
Cfalr: Collaborative Filtering-Augmented Large Language Model For Personalized Fashion Outfit Recommendation, Yujuan Ding, Junrong Liao, Yunshan Ma, Yi Bin, Wenqi Fan, Tat-Seng Chua, Qing Li
Research Collection School Of Computing and Information Systems
Personalized outfit recommendation poses a significant challenge in e-commerce and social media platforms, requiring systems that balance user preferences with aesthetic compatibility. Collaborative filtering (CF) provides a traditional solution for this, but it struggles with data-sparse scenarios and complex user-item-outfit relationships. Meanwhile, existing template-based approaches are constrained by rigid pre-designed structures. To bridge these research gaps, we introduce CFALR (Collaborative Filtering-Augmented Large Language Model for Recommendation), a novel framework that synergizes collaborative filtering with large language models for personalized outfit recommendation. Specifically, CFALR describes user-outfit interactions in natural language and leverages LLMs to capture fashion semantics while employing CF-enhanced embeddings …
How Do Machine Learning Models Change?,
2026
Singapore Management University
How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández
Research Collection School Of Computing and Information Systems
The proliferation of Machine Learning (ML) models and their open source implementations has transformed AI research and applications. Platforms like Hugging Face (HF) enable this evolving ecosystem, yet a large-scale longitudinal study of how these models change is lacking. This study addresses this gap by analyzing over 680,000 commits from 100,000 models and 2,251 releases from 202 of these models on HF using repository mining and longitudinal methods. We apply an extended ML change taxonomy to classify commits and use Bayesian networks to model temporal patterns in commit and release activities. Our findings show that commit activities align with established …
