Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir,
2025
Singapore Management University
Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang
Research Collection School Of Computing and Information Systems
Traditional ship detection methods primarily rely on single-modal approaches, such as visible or infrared images, which limit their application in complex scenarios involving varying lighting conditions and heavy fog. To address this issue, we explore the advantages of short-wave infrared (SWIR) and long-wave infrared (LWIR) in ship detection and propose a novel single-stage image fusion detection algorithm called LSFDNet. This algorithm leverages feature interaction between the image fusion and object detection subtask networks, achieving remarkable detection performance and generating visually impressive fused images. To further improve the saliency of objects in the fused images and improve the performance of the …
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired,
2025
Singapore Management University
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu
Research Collection School Of Computing and Information Systems
Tactile graphics on a refreshable display have proven effective in enabling visually impaired people to comprehend pictorial content. To further evaluate the effectiveness of refreshable tactile displays in blind education, we designed tactile data comics, a method that combines step-by-step presentation of tactile graphics with verbal narration. We conducted a user study with sixteen visually impaired students to compare tactile data comics against verbal-only and static tactile graphics. Our findings show that tactile data comics significantly improve participants’ comprehension and engagement during the learning experience. These empirical results suggest that the integration of refreshable tactile displays and tactile data comics …
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval,
2025
Singapore Management University
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …
Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars,
2025
Singapore Management University
Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars, Hyuna Seo, Youngki Lee, Rajesh Krishna Balan, Thivya Kandappu
Research Collection School Of Computing and Information Systems
We present EmoShortcuts1, a novel social Mixed Reality (MR) framework that enhances emotional expression by dynamically augmenting avatar body gestures to reflect users’ emotional states. While social MR enables immersive remote interactions through avatars, conveying emotions remains challenging due to limitations in head-mounted display (HMD) tracking (e.g., missing lower-body movements, such as stomping or defensive postures), and users’ tendency to deprioritize nonverbal expressions during multitasking. EmoShortcuts addresses these challenges by introducing an augmentation framework that generates expressive body gestures even when users’ physical movements are restricted. We conducted a formative study with 12 participants to identify key challenges in emotional …
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach,
2025
Singapore Management University
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He
Research Collection School Of Computing and Information Systems
In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning,
2025
Singapore Management University
Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo
Research Collection School Of Computing and Information Systems
In this paper, we introduce DiffusionMat, a novel image matting framework that employs a diffusion model for the transition from coarse to refined alpha mattes. Diverging from conventional methods that utilize trimaps merely as loose guidance for alpha matte prediction, our approach treats image matting as a deterministic sequential refinement learning process. This process begins with the addition of noise to trimaps and iteratively denoises them using a pre-trained diffusion model, which incrementally guides the prediction towards a clean alpha matte. The key innovation of our framework is a correction module that adjusts the output at each denoising step, ensuring …
Robface: A Test Suite For Efficient Robustness Evaluation Of Face Recognition Systems,
2025
Singapore Management University
Robface: A Test Suite For Efficient Robustness Evaluation Of Face Recognition Systems, Ruihan Zhang, Jun Sun
Research Collection School Of Computing and Information Systems
Face recognition is a widely used authentication technology in practice, where robustness is required. It is thus essential to have an efficient and easy-to-use method for evaluating the robustness of (possibly third-party) trained face recognition systems. Existing approaches to evaluating the robustness of face recognition systems are either based on empirical evaluation (e.g., measuring attacking success rate using state-of-the-art attacking methods) or formal analysis (e.g., measuring the Lipschitz constant). While the former demands significant user efforts and expertise, the latter is extremely time-consuming. In pursuit of a comprehensive, efficient, easy-to-use, and scalable estimation of the robustness of face recognition systems, …
Vibemus: Proactive Agentic System For Music Personalization,
2025
Singapore Management University
Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) enable diverse forms of AI-assisted creation, yet they often struggle to bridge the preference-articulation gap: users may provide incomplete or vague intentions or lack the vocabulary to specify what they want, yielding outputs misaligned with true preferences. To address this gap and facilitate music creation in a vibe-centric environment, we introduce VibeMus, a proactive agentic system built on open-source components. The system engages in multi-turn dialogue to progressively determine the music’s emotion, genre, lyrics, and other aspects before generation. Simulated evaluations show that proactive clarification improves alignment with users’ intended nuances. Our approach is training-free, leveraging …
Gti: Graph-Based Tree Index With Logarithm Updates For Nearest Neighbor Search In High-Dimensional Spaces,
2025
Singapore Management University
Gti: Graph-Based Tree Index With Logarithm Updates For Nearest Neighbor Search In High-Dimensional Spaces, Ruoyao Ma, Yifan Zhu, Baihua Zheng, Lu Chen, Congcong Ge, Yunjun Gao
Research Collection School Of Computing and Information Systems
Nearest neighbor search (NNS) is fundamental for high-dimensional space retrieval and impacts various fields, such as pattern recognition, information retrieval, recommendation systems, and vector database management. Among existing NNS methods, graph-based methods often excel in query accuracy and efficiency. However, these methods face significant challenges, including high construction costs and difficulties with dynamic data updates. Recent efforts have focused on combining graph methods with hashing, quantization, and tree-based approaches to address these issues, but problems with large index sizes and update performance remain unresolved. In response, this paper proposes GTI, a novel, lightweight, and dynamic graph-based tree index for high-dimensional …
Deep Graph Anomaly Detection: A Survey And New Perspectives,
2025
Singapore Management University
Deep Graph Anomaly Detection: A Survey And New Perspectives, Hezhe Qiao, Hanghang Tong, Nanyang Technological University, Irwin King, Charu Aggarwal, Guansong Pang
Research Collection School Of Computing and Information Systems
Graph anomaly detection (GAD), which aims to identify unusual graph instances (e.g., nodes, edges, subgraphs, or graphs), has attracted increasing attention in recent years due to its significance in a wide range of applications. Deep learning approaches, graph neural networks (GNNs) in particular, have been emerging as a promising paradigm for GAD, owing to its strong capability in capturing complex structure and/or node attributes in graph data. Considering the large number of methods proposed for GNN-based GAD, it is of paramount importance to summarize the methodologies and findings in the existing GAD studies, so that we can pinpoint effective model …
Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook,
2025
South China University of Technology
Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook, Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du, Hongmin Cai, Jing Qin, Shengfeng He
Research Collection School Of Computing and Information Systems
Although pre-trained large-scale generative models StyleGAN series have proven to be effective in various editing and translation tasks, they are limited to pre-defined fixed aspect ratio. To overcome this limitation, we propose StyleGAN-∞, a model that enables pre-trained StyleGAN to perform arbitrary-ratio conditional synthesis. Our key insight is to distill the expressive StyleGAN features into a StyleBook, such that an arbitrary-ratio condition can be translated to other forms by properly assembling pre-defined StyleBook vectors. To learn and leverage the StyleBook, we employ a network with three distinct stages, each corresponding to StyleBook extraction, StyleBook correspondence learning, and arbitrary-ratio synthesis. Extensive …
Map As A By-Product: Collective Landmark Mapping From Imu Data And User-Provided Texts In Situated Tasks,
2025
Singapore Management University
Map As A By-Product: Collective Landmark Mapping From Imu Data And User-Provided Texts In Situated Tasks, Ryo Yonetani, Kotaro Hara
Research Collection School Of Computing and Information Systems
This paper presents Collective Landmark Mapper, a novel map-as-a-by-product system for generating semantic landmark maps of indoor environments. Consider users engaged in situated tasks that require them to navigate these environments and regularly take notes on their smartphones. Collective Landmark Mapper exploits the smartphone's IMU data and the user's free text input during these tasks to identify a set of landmarks encountered by the user. The identified landmarks are then aggregated across multiple users to generate a unified map representing the positions and semantic information of all landmarks. In developing the proposed system, we focused specifically on retail applications and …
Towards Multimodal Emotional Support Conversation Systems,
2025
Hefei University of Technology
Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong
Research Collection School Of Computing and Information Systems
The integration of conversational artificial intelligence (AI) into mental health care promises a new horizon for therapist-client interactions, aiming to closely emulate the depth and nuance of human conversations. Despite the potential, the current landscape of conversational AI is markedly limited by its reliance on single-modal data, constraining the systems’ ability to empathize and provide effective emotional support. This limitation stems from a paucity of resources that encapsulate the multimodal nature of human communication essential for therapeutic counseling. To address this gap, we introduce the Multimodal Emotional Support Conversation (MESC) dataset, a first-of-its-kind resource enriched with comprehensive annotations across text, …
Towards Age-Inclusive Human Computer Interaction: A Study Of Text Message Adoption By Seniors,
2025
Women's University in Africa
Towards Age-Inclusive Human Computer Interaction: A Study Of Text Message Adoption By Seniors, Sam Takavarasha, Varaidzo Mapepa
African Conference on Information Systems and Technology
Since mobile phones are increasingly becoming livelihoods-enablers to people in developing economies, inclusive interaction design are critical. This research on the age inclusivity of feature phones was motivated by some observation that elderly users had challenges with adoption of text messaging. We, therefore, hypothesized that elderly people had challenges with texting ‘cash’, ‘cheque’ or ‘visa’ on a feature phone. The paper investigates if the required agility and interface ergonomics were age discriminatory inhibitors of the adoption of text messaging given the diminishing dexterity and cognitive skills of seniors. Using Sen’s (1999) ‘heterogeneity of capabilities’ theory and Davis et al (1989)’s …
Creepy, Invasive, And Exploitative Algorithms: A Cpm Analysis Of Users' Privacy Breakdowns And Recalibration Practices With Social Media Algorithms,
2025
Central Michigan University
Creepy, Invasive, And Exploitative Algorithms: A Cpm Analysis Of Users' Privacy Breakdowns And Recalibration Practices With Social Media Algorithms, Matthew J. A. Craig, Jeffrey T. Child
Human-Machine Communication
Social media content filtering algorithms can both provide desired personalized content and ads for users. However, sometimes these recommendations can resemble individual private information. How might users navigate these experiences to best manage their private information? The present exploratory study utilizes the rules- and systems-based framework of communication privacy management (CPM) theory to explore social media users’ experiences of privacy breakdowns with social media algorithms and investigates what users do in response to said breakdowns. These responses were refined using content analysis and divided into different categories of privacy breakdowns and recalibration strategies. Implications for future research surrounding human-machine communication …
Machine Learning And Crime Prevention,
2025
CUNY John Jay College
Machine Learning And Crime Prevention, Emily Lizewski
Student Theses
Predictive policing uses machine learning to analyze crime patterns and help law enforcement better efficient use their resources. These tools can improve accuracy by highlighting complex trends in large sets of data. While this technology has its advantages, it also raises important ethical and social questions. Within this paper we looks at how predictive policing works, focusing on the machine learning models often used such as decision trees, random forests, gradient boosting, and models that factor in both time and location. It also explores how these tools might unintentionally reinforce biases already present in historical crime data. In reviewing the …
Evometric: An Interactive Framework For Scalable Visual Analytics Of Time Series Data With Dynamic Changes.,
2025
University of Louisville
Evometric: An Interactive Framework For Scalable Visual Analytics Of Time Series Data With Dynamic Changes., Jiahang Huang
Electronic Theses and Dissertations
In today's data-intensive landscape, rapid advances in digital sensing and recording technologies have enabled the acquisition of high-resolution multimodal time series data, capturing intricate real-world dynamics across various domains such as healthcare, behavioral science, and environmental monitoring. However, the complexity and scale of these datasets present significant analytical challenges, particularly in understanding dynamic changes at both individual and cohort levels. This dissertation introduces EvoMetric, a novel visual analytics framework designed to support scalable exploration and analysis of large-scale multimodal time series data with dynamic changes. EvoMetric seamlessly integrates individual-level temporal dynamics with population-level comparative insights, enabling users to visually …
Revolutionizing Digital Privacy Education For Older Adults: Enhanced Interventions And Ai-Assisted Learning Strategies,
2025
Clemson University
Revolutionizing Digital Privacy Education For Older Adults: Enhanced Interventions And Ai-Assisted Learning Strategies, Heba Aly
All Dissertations
As older adults increasingly engage with digital platforms, they face unique privacy risks stemming from limited digital literacy, reduced trust in AI technologies, and constrained access—especially in rural or underserved communities. While digital tools offer benefits like social connection and information access, current privacy education efforts often neglect the needs of older adults. This dissertation addresses this gap by developing, testing, and refining digital privacy education interventions tailored for older adults, with a focus on trust, personalization, and AI-assisted learning.
Study 1 evaluates multiple instructional modalities across age groups, revealing older adults prefer structured videos and interactive tutorials, while younger …
Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond,
2025
Singapore Management University
Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Text-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character …
Affinitytune: A Prompt-Tuning Framework For Few-Shot Anomaly Detection On Graphs,
2025
Singapore Management University
Affinitytune: A Prompt-Tuning Framework For Few-Shot Anomaly Detection On Graphs, Jingyan Chen, Guanghui Zhu, Guansong Pang, Chunfeng Yuan, Yihua Huang
Research Collection School Of Computing and Information Systems
Graph anomaly detection (GAD) is a critical task with applications in domains such as networking, finance, and bioinformatics. % However, the scarcity of labeled anomalies and the limitations of unsupervised methods hinder effective detection. % While semi-supervised and few-shot learning approaches offer improvements, they struggle with knowledge transfer and rely heavily on labeled data. % Recent advancements in prompt tuning on graphs provide a promising direction, but their application to heterophilous graphs in anomaly detection remains underexplored. % In this work, we propose AffinityTune, a novel framework for few-shot graph anomaly detection based on prompt tuning. % Our approach introduces …
