Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection,
2025
Singapore Management University
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Research Collection School Of Computing and Information Systems
Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitation: a lack of diversity in the extracted OOD features from the few-shot ID data.To address this issue, we propose to leverage external datasets as auxiliary outlier data (i.e., pseudo OOD samples) to extract rich, diverse OOD features, with the features from not only background regions but also foreground object regions, thereby supporting more discriminative prompt tuning for OOD detection. We further introduce …
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval,
2025
Singapore Management University
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach,
2025
Singapore Management University
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He
Research Collection School Of Computing and Information Systems
In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …
Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars,
2025
Singapore Management University
Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars, Hyuna Seo, Youngki Lee, Rajesh Krishna Balan, Thivya Kandappu
Research Collection School Of Computing and Information Systems
We present EmoShortcuts1, a novel social Mixed Reality (MR) framework that enhances emotional expression by dynamically augmenting avatar body gestures to reflect users’ emotional states. While social MR enables immersive remote interactions through avatars, conveying emotions remains challenging due to limitations in head-mounted display (HMD) tracking (e.g., missing lower-body movements, such as stomping or defensive postures), and users’ tendency to deprioritize nonverbal expressions during multitasking. EmoShortcuts addresses these challenges by introducing an augmentation framework that generates expressive body gestures even when users’ physical movements are restricted. We conducted a formative study with 12 participants to identify key challenges in emotional …
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired,
2025
Singapore Management University
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu
Research Collection School Of Computing and Information Systems
Tactile graphics on a refreshable display have proven effective in enabling visually impaired people to comprehend pictorial content. To further evaluate the effectiveness of refreshable tactile displays in blind education, we designed tactile data comics, a method that combines step-by-step presentation of tactile graphics with verbal narration. We conducted a user study with sixteen visually impaired students to compare tactile data comics against verbal-only and static tactile graphics. Our findings show that tactile data comics significantly improve participants’ comprehension and engagement during the learning experience. These empirical results suggest that the integration of refreshable tactile displays and tactile data comics …
Stable Score Distillation,
2025
Singapore Management University
Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He
Research Collection School Of Computing and Information Systems
Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieve cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's …
A Comprehensive Review Of Financial Knowledge Graphs,
2025
Singapore Management University
A Comprehensive Review Of Financial Knowledge Graphs, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang
Research Collection School Of Computing and Information Systems
Knowledge Graphs (KGs) are increasingly used in finance to manage complex, interconnected data and support advanced analytics. This survey provides an overview of how KGs are applied across various financial areas, such as fraud detection, credit risk assessment, anti-money laundering, and regulatory compliance. We examine key techniques for building and using KGs in finance, including graph construction, embedding methods, and machine learning models. The survey also discusses challenges specific to finance, like handling private data, ensuring interpretability, and managing real-time data. Additionally, we explore the emerging combination of KGs with large language models and generative AI, which offers new possibilities …
Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation,
2025
Singapore Management University
Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He
Research Collection School Of Computing and Information Systems
Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces cross-image stroke attention, a mechanism embedded within self-attention layers to establish fine-grained semantic correspondences and enable accurate stroke attribute transfer. This allows our method to adaptively integrate reference stroke characteristics into content images while maintaining structural integrity. Additionally, we develop adaptive contrast enhancement and semanticfocused attention to reinforce content preservation and foreground emphasis. Stroke2Sketch effectively synthesizes stylistically faithful sketches that closely …
Cross-Subject Mind Decoding From Inaccurate Representations,
2025
Singapore Management University
Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu
Research Collection School Of Computing and Information Systems
Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings generate partially inaccurate representations that, when fed into diffusion models, accumulate errors and degrade reconstruction fidelity. To address this, we propose the Bidirectional Autoencoder Intertwining framework for accurate decoded representation prediction. Our approach unifies multiple subjects through a Subject Bias Modulation Module while leveraging bidirectional mapping to better capture data distributions for precise representation prediction. To further enhance fidelity when decoding representations into stimulus images, …
Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning,
2025
Singapore Management University
Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang
Research Collection School Of Computing and Information Systems
In this work, we propose a novel continuous graph neural network called FCAD (Feature-Coupled Anisotropic Diffusion) for the task of node classification on graphs. Our approach is motivated by the success of feature-coupled anisotropic diffusion PDEs in multivalued image restoration. Our method introduces a total variation regularization-inspired anisotropic term to control diffusion between nodes and incorporates a learnable parameterization for feature coupling during the diffusion process. Our model performs competitively against several GNN baselines for both heterophilous and homophilous graphs, demonstrating notable benefits for heterophilous graphs due to the learnable feature coupling.
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification,
2025
Singapore Management University
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Research Collection School Of Computing and Information Systems
The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …
Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir,
2025
Singapore Management University
Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang
Research Collection School Of Computing and Information Systems
Traditional ship detection methods primarily rely on single-modal approaches, such as visible or infrared images, which limit their application in complex scenarios involving varying lighting conditions and heavy fog. To address this issue, we explore the advantages of short-wave infrared (SWIR) and long-wave infrared (LWIR) in ship detection and propose a novel single-stage image fusion detection algorithm called LSFDNet. This algorithm leverages feature interaction between the image fusion and object detection subtask networks, achieving remarkable detection performance and generating visually impressive fused images. To further improve the saliency of objects in the fused images and improve the performance of the …
Deep Graph Anomaly Detection: A Survey And New Perspectives,
2025
Singapore Management University
Deep Graph Anomaly Detection: A Survey And New Perspectives, Hezhe Qiao, Hanghang Tong, Nanyang Technological University, Irwin King, Charu Aggarwal, Guansong Pang
Research Collection School Of Computing and Information Systems
Graph anomaly detection (GAD), which aims to identify unusual graph instances (e.g., nodes, edges, subgraphs, or graphs), has attracted increasing attention in recent years due to its significance in a wide range of applications. Deep learning approaches, graph neural networks (GNNs) in particular, have been emerging as a promising paradigm for GAD, owing to its strong capability in capturing complex structure and/or node attributes in graph data. Considering the large number of methods proposed for GNN-based GAD, it is of paramount importance to summarize the methodologies and findings in the existing GAD studies, so that we can pinpoint effective model …
Robface: A Test Suite For Efficient Robustness Evaluation Of Face Recognition Systems,
2025
Singapore Management University
Robface: A Test Suite For Efficient Robustness Evaluation Of Face Recognition Systems, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Face recognition is a widely used authentication technology in practice, where robustness is required. It is thus essential to have an efficient and easy-to-use method for evaluating the robustness of (possibly third-party) trained face recognition systems. Existing approaches to evaluating the robustness of face recognition systems are either based on empirical evaluation (e.g., measuring attacking success rate using state-of-the-art attacking methods) or formal analysis (e.g., measuring the Lipschitz constant). While the former demands significant user efforts and expertise, the latter is extremely time-consuming. In pursuit of a comprehensive, efficient, easy-to-use, and scalable estimation of the robustness of face recognition systems, …
Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook,
2025
South China University of Technology
Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook, Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du, Hongmin Cai, Jing Qin, Shengfeng He
Research Collection School Of Computing and Information Systems
Although pre-trained large-scale generative models StyleGAN series have proven to be effective in various editing and translation tasks, they are limited to pre-defined fixed aspect ratio. To overcome this limitation, we propose StyleGAN-∞, a model that enables pre-trained StyleGAN to perform arbitrary-ratio conditional synthesis. Our key insight is to distill the expressive StyleGAN features into a StyleBook, such that an arbitrary-ratio condition can be translated to other forms by properly assembling pre-defined StyleBook vectors. To learn and leverage the StyleBook, we employ a network with three distinct stages, each corresponding to StyleBook extraction, StyleBook correspondence learning, and arbitrary-ratio synthesis. Extensive …
Gti: Graph-Based Tree Index With Logarithm Updates For Nearest Neighbor Search In High-Dimensional Spaces,
2025
Singapore Management University
Gti: Graph-Based Tree Index With Logarithm Updates For Nearest Neighbor Search In High-Dimensional Spaces, Ruoyao Ma, Yifan Zhu, Baihua Zheng, Lu Chen, Congcong Ge, Yunjun Gao
Research Collection School Of Computing and Information Systems
Nearest neighbor search (NNS) is fundamental for high-dimensional space retrieval and impacts various fields, such as pattern recognition, information retrieval, recommendation systems, and vector database management. Among existing NNS methods, graph-based methods often excel in query accuracy and efficiency. However, these methods face significant challenges, including high construction costs and difficulties with dynamic data updates. Recent efforts have focused on combining graph methods with hashing, quantization, and tree-based approaches to address these issues, but problems with large index sizes and update performance remain unresolved. In response, this paper proposes GTI, a novel, lightweight, and dynamic graph-based tree index for high-dimensional …
Vibemus: Proactive Agentic System For Music Personalization,
2025
Singapore Management University
Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) enable diverse forms of AI-assisted creation, yet they often struggle to bridge the preference-articulation gap: users may provide incomplete or vague intentions or lack the vocabulary to specify what they want, yielding outputs misaligned with true preferences. To address this gap and facilitate music creation in a vibe-centric environment, we introduce VibeMus, a proactive agentic system built on open-source components. The system engages in multi-turn dialogue to progressively determine the music’s emotion, genre, lyrics, and other aspects before generation. Simulated evaluations show that proactive clarification improves alignment with users’ intended nuances. Our approach is training-free, leveraging …
Map As A By-Product: Collective Landmark Mapping From Imu Data And User-Provided Texts In Situated Tasks,
2025
Singapore Management University
Map As A By-Product: Collective Landmark Mapping From Imu Data And User-Provided Texts In Situated Tasks, Ryo Yonetani, Kotaro Hara
Research Collection School Of Computing and Information Systems
This paper presents Collective Landmark Mapper, a novel map-as-a-by-product system for generating semantic landmark maps of indoor environments. Consider users engaged in situated tasks that require them to navigate these environments and regularly take notes on their smartphones. Collective Landmark Mapper exploits the smartphone's IMU data and the user's free text input during these tasks to identify a set of landmarks encountered by the user. The identified landmarks are then aggregated across multiple users to generate a unified map representing the positions and semantic information of all landmarks. In developing the proposed system, we focused specifically on retail applications and …
Towards Multimodal Emotional Support Conversation Systems,
2025
Singapore Management University
Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong
Research Collection School Of Computing and Information Systems
The integration of conversational artificial intelligence (AI) into mental health care promises a new horizon for therapist-client interactions, aiming to closely emulate the depth and nuance of human conversations. Despite the potential, the current landscape of conversational AI is markedly limited by its reliance on single-modal data, constraining the systems’ ability to empathize and provide effective emotional support. This limitation stems from a paucity of resources that encapsulate the multimodal nature of human communication essential for therapeutic counseling. To address this gap, we introduce the Multimodal Emotional Support Conversation (MESC) dataset, a first-of-its-kind resource enriched with comprehensive annotations across text, …
Towards Age-Inclusive Human Computer Interaction: A Study Of Text Message Adoption By Seniors,
2025
Women's University in Africa
Towards Age-Inclusive Human Computer Interaction: A Study Of Text Message Adoption By Seniors, Sam Takavarasha, Varaidzo Mapepa
African Conference on Information Systems and Technology
Since mobile phones are increasingly becoming livelihoods-enablers to people in developing economies, inclusive interaction design are critical. This research on the age inclusivity of feature phones was motivated by some observation that elderly users had challenges with adoption of text messaging. We, therefore, hypothesized that elderly people had challenges with texting ‘cash’, ‘cheque’ or ‘visa’ on a feature phone. The paper investigates if the required agility and interface ergonomics were age discriminatory inhibitors of the adoption of text messaging given the diminishing dexterity and cognitive skills of seniors. Using Sen’s (1999) ‘heterogeneity of capabilities’ theory and Davis et al (1989)’s …
