Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year

Articles 61 - 90 of 912

Full-Text Articles in Graphics and Human Computer Interfaces

Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang Oct 2025

Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

In this work, we propose a novel continuous graph neural network called FCAD (Feature-Coupled Anisotropic Diffusion) for the task of node classification on graphs. Our approach is motivated by the success of feature-coupled anisotropic diffusion PDEs in multivalued image restoration. Our method introduces a total variation regularization-inspired anisotropic term to control diffusion between nodes and incorporates a learnable parameterization for feature coupling during the diffusion process. Our model performs competitively against several GNN baselines for both heterophilous and homophilous graphs, demonstrating notable benefits for heterophilous graphs due to the learnable feature coupling.


Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He Oct 2025

Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He

Research Collection School Of Computing and Information Systems

The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …


Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang Oct 2025

Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

Traditional ship detection methods primarily rely on single-modal approaches, such as visible or infrared images, which limit their application in complex scenarios involving varying lighting conditions and heavy fog. To address this issue, we explore the advantages of short-wave infrared (SWIR) and long-wave infrared (LWIR) in ship detection and propose a novel single-stage image fusion detection algorithm called LSFDNet. This algorithm leverages feature interaction between the image fusion and object detection subtask networks, achieving remarkable detection performance and generating visually impressive fused images. To further improve the saliency of objects in the fused images and improve the performance of the …


Deep Graph Anomaly Detection: A Survey And New Perspectives, Hezhe Qiao, Hanghang Tong, Nanyang Technological University, Irwin King, Charu Aggarwal, Guansong Pang Sep 2025

Deep Graph Anomaly Detection: A Survey And New Perspectives, Hezhe Qiao, Hanghang Tong, Nanyang Technological University, Irwin King, Charu Aggarwal, Guansong Pang

Research Collection School Of Computing and Information Systems

Graph anomaly detection (GAD), which aims to identify unusual graph instances (e.g., nodes, edges, subgraphs, or graphs), has attracted increasing attention in recent years due to its significance in a wide range of applications. Deep learning approaches, graph neural networks (GNNs) in particular, have been emerging as a promising paradigm for GAD, owing to its strong capability in capturing complex structure and/or node attributes in graph data. Considering the large number of methods proposed for GNN-based GAD, it is of paramount importance to summarize the methodologies and findings in the existing GAD studies, so that we can pinpoint effective model …


Robface: A Test Suite For Efficient Robustness Evaluation Of Face Recognition Systems, Ruihan Zhang, Jun Sun Sep 2025

Robface: A Test Suite For Efficient Robustness Evaluation Of Face Recognition Systems, Ruihan Zhang, Jun Sun

Research Collection School Of Computing and Information Systems

Face recognition is a widely used authentication technology in practice, where robustness is required. It is thus essential to have an efficient and easy-to-use method for evaluating the robustness of (possibly third-party) trained face recognition systems. Existing approaches to evaluating the robustness of face recognition systems are either based on empirical evaluation (e.g., measuring attacking success rate using state-of-the-art attacking methods) or formal analysis (e.g., measuring the Lipschitz constant). While the former demands significant user efforts and expertise, the latter is extremely time-consuming. In pursuit of a comprehensive, efficient, easy-to-use, and scalable estimation of the robustness of face recognition systems, …


Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook, Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du, Hongmin Cai, Jing Qin, Shengfeng He Sep 2025

Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook, Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du, Hongmin Cai, Jing Qin, Shengfeng He

Research Collection School Of Computing and Information Systems

Although pre-trained large-scale generative models StyleGAN series have proven to be effective in various editing and translation tasks, they are limited to pre-defined fixed aspect ratio. To overcome this limitation, we propose StyleGAN-∞, a model that enables pre-trained StyleGAN to perform arbitrary-ratio conditional synthesis. Our key insight is to distill the expressive StyleGAN features into a StyleBook, such that an arbitrary-ratio condition can be translated to other forms by properly assembling pre-defined StyleBook vectors. To learn and leverage the StyleBook, we employ a network with three distinct stages, each corresponding to StyleBook extraction, StyleBook correspondence learning, and arbitrary-ratio synthesis. Extensive …


Gti: Graph-Based Tree Index With Logarithm Updates For Nearest Neighbor Search In High-Dimensional Spaces, Ruoyao Ma, Yifan Zhu, Baihua Zheng, Lu Chen, Congcong Ge, Yunjun Gao Sep 2025

Gti: Graph-Based Tree Index With Logarithm Updates For Nearest Neighbor Search In High-Dimensional Spaces, Ruoyao Ma, Yifan Zhu, Baihua Zheng, Lu Chen, Congcong Ge, Yunjun Gao

Research Collection School Of Computing and Information Systems

Nearest neighbor search (NNS) is fundamental for high-dimensional space retrieval and impacts various fields, such as pattern recognition, information retrieval, recommendation systems, and vector database management. Among existing NNS methods, graph-based methods often excel in query accuracy and efficiency. However, these methods face significant challenges, including high construction costs and difficulties with dynamic data updates. Recent efforts have focused on combining graph methods with hashing, quantization, and tree-based approaches to address these issues, but problems with large index sizes and update performance remain unresolved. In response, this paper proposes GTI, a novel, lightweight, and dynamic graph-based tree index for high-dimensional …


Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang Sep 2025

Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) enable diverse forms of AI-assisted creation, yet they often struggle to bridge the preference-articulation gap: users may provide incomplete or vague intentions or lack the vocabulary to specify what they want, yielding outputs misaligned with true preferences. To address this gap and facilitate music creation in a vibe-centric environment, we introduce VibeMus, a proactive agentic system built on open-source components. The system engages in multi-turn dialogue to progressively determine the music’s emotion, genre, lyrics, and other aspects before generation. Simulated evaluations show that proactive clarification improves alignment with users’ intended nuances. Our approach is training-free, leveraging …


Map As A By-Product: Collective Landmark Mapping From Imu Data And User-Provided Texts In Situated Tasks, Ryo Yonetani, Kotaro Hara Sep 2025

Map As A By-Product: Collective Landmark Mapping From Imu Data And User-Provided Texts In Situated Tasks, Ryo Yonetani, Kotaro Hara

Research Collection School Of Computing and Information Systems

This paper presents Collective Landmark Mapper, a novel map-as-a-by-product system for generating semantic landmark maps of indoor environments. Consider users engaged in situated tasks that require them to navigate these environments and regularly take notes on their smartphones. Collective Landmark Mapper exploits the smartphone's IMU data and the user's free text input during these tasks to identify a set of landmarks encountered by the user. The identified landmarks are then aggregated across multiple users to generate a unified map representing the positions and semantic information of all landmarks. In developing the proposed system, we focused specifically on retail applications and …


Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong Sep 2025

Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong

Research Collection School Of Computing and Information Systems

The integration of conversational artificial intelligence (AI) into mental health care promises a new horizon for therapist-client interactions, aiming to closely emulate the depth and nuance of human conversations. Despite the potential, the current landscape of conversational AI is markedly limited by its reliance on single-modal data, constraining the systems’ ability to empathize and provide effective emotional support. This limitation stems from a paucity of resources that encapsulate the multimodal nature of human communication essential for therapeutic counseling. To address this gap, we introduce the Multimodal Emotional Support Conversation (MESC) dataset, a first-of-its-kind resource enriched with comprehensive annotations across text, …


Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He Aug 2025

Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

Text-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character …


Quantizing Text-Attributed Graphs For Semantic-Structural Integration, Jianyuan Bo, Hao Wu, Yuan Fang Aug 2025

Quantizing Text-Attributed Graphs For Semantic-Structural Integration, Jianyuan Bo, Hao Wu, Yuan Fang

Research Collection School Of Computing and Information Systems

Text-attributed graphs (TAGs) have emerged as a powerful representation for modeling complex relationships across diverse domains. With the rise of large language models (LLMs), there is growing interest in leveraging their capabilities for graph learning. However, current approaches face significant challenges in embedding structural information into LLM-compatible formats, requiring either computationally expensive alignment mechanisms or manual graph verbalization techniques that often lose critical structural details. Moreover, these methods typically require labeled data from source domains for effective transfer learning, significantly constraining their adaptability. We propose STAG, a novel self-supervised framework that directly quantizes graph structural information into discrete tokens using …


Graph Positional Autoencoders As Self-Supervised Learners, Yang Liu, Deyu Bo, Wenxuan Cao, Yuan Fang, Yawen Li, Chuan Shi Aug 2025

Graph Positional Autoencoders As Self-Supervised Learners, Yang Liu, Deyu Bo, Wenxuan Cao, Yuan Fang, Yawen Li, Chuan Shi

Research Collection School Of Computing and Information Systems

Graph self-supervised learning seeks to learn effective graph representations without relying on labeled data. Among various approaches, graph autoencoders (GAEs) have gained significant attention for their efficiency and scalability. Typically, GAEs take incomplete graphs as input and predict missing elements, such as masked node features or edges. Although effective, our experimental investigation reveals that traditional feature or edge masking paradigms primarily capture low-frequency signals in the graph and fail to learn expressive structural information. To address these issues, we propose Graph Positional Autoencoders (GraphPAE), which employ a dual-path architecture to reconstruct both node features and positions. Specifically, the feature path …


Bhvit: Binarized Hybrid Vision Transformer, Tian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu, Kaijie Yin, Chengzhong Xu, Hui Kong Aug 2025

Bhvit: Binarized Hybrid Vision Transformer, Tian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu, Kaijie Yin, Chengzhong Xu, Hui Kong

Research Collection School Of Computing and Information Systems

Model binarization has made significant progress in enabling real-time and energy-efficient computation for con-volutional neural networks (CNN), offering a potential solution to the deployment challenges faced by Vision Transformers (ViTs) on edge devices. However, due to the structural differences between CNN and Transformer architectures, simply applying binary CNN strategies to the ViT models will lead to a significant performance drop. To tackle this challenge, we propose BHViT, a binarization-friendly hybrid ViT architecture and its full binarization model with the guidance of three important observations. Initially, BHViT utilizes the local information interaction and hierarchical feature aggregation technique from coarse to fine …


Affinitytune: A Prompt-Tuning Framework For Few-Shot Anomaly Detection On Graphs, Jingyan Chen, Guanghui Zhu, Guansong Pang, Chunfeng Yuan, Yihua Huang Aug 2025

Affinitytune: A Prompt-Tuning Framework For Few-Shot Anomaly Detection On Graphs, Jingyan Chen, Guanghui Zhu, Guansong Pang, Chunfeng Yuan, Yihua Huang

Research Collection School Of Computing and Information Systems

Graph anomaly detection (GAD) is a critical task with applications in domains such as networking, finance, and bioinformatics. % However, the scarcity of labeled anomalies and the limitations of unsupervised methods hinder effective detection. % While semi-supervised and few-shot learning approaches offer improvements, they struggle with knowledge transfer and rely heavily on labeled data. % Recent advancements in prompt tuning on graphs provide a promising direction, but their application to heterophilous graphs in anomaly detection remains underexplored. % In this work, we propose AffinityTune, a novel framework for few-shot graph anomaly detection based on prompt tuning. % Our approach introduces …


Advancing Molecular Graph-Text Pre-Training Via Fine-Grained Alignment, Yibo Li, Yuan Fang, Mengmei Zhang, Chuan Shi Aug 2025

Advancing Molecular Graph-Text Pre-Training Via Fine-Grained Alignment, Yibo Li, Yuan Fang, Mengmei Zhang, Chuan Shi

Research Collection School Of Computing and Information Systems

Understanding molecular structure and related knowledge is crucialfor scientific research. Recent studies integrate molecular graphswith their textual descriptions to enhance molecular representationlearning. However, they focus on the whole molecular graph andneglect frequently occurring subgraphs, known as motifs, whichare essential for determining molecular properties. Without suchfine-grained knowledge, these models struggle to generalize to un-seen molecules and tasks that require motif-level insights. To bridgethis gap, we propose FineMolTex, a novel Fine-grained Moleculargraph-Text pre-training framework to jointly learn coarse-grainedmolecule-level knowledge and fine-grained motif-level knowledge.Specifically, FineMolTex consists of two pre-training tasks: a con-trastive alignment task for coarse-grained matching and a maskedmulti-modal modeling task for …


Gcot: Chain-Of-Thought Prompt Learning For Graphs, Xingtong Yu, Chang Zhou, Zhongwei Kuai, Xinming Zhang, Yuan Fang Aug 2025

Gcot: Chain-Of-Thought Prompt Learning For Graphs, Xingtong Yu, Chang Zhou, Zhongwei Kuai, Xinming Zhang, Yuan Fang

Research Collection School Of Computing and Information Systems

Chain-of-thought (CoT) prompting has achieved remarkable success in natural language processing (NLP). However, its vast potential remains largely unexplored for graphs. This raises an interesting question: How can we design CoT prompting for graphs to guide graph models to learn step by step? On one hand, unlike natural languages, graphs are non-linear and characterized by complex topological structures. On the other hand, many graphs lack textual data, making it difficult to formulate language-based CoT prompting. %Therefore we cannot directly adopt the CoT prompting methods used in the language domain. In this work, we propose the first CoT prompt learning framework …


Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang Aug 2025

Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang

Research Collection School Of Computing and Information Systems

Humans possess a remarkable ability to interpret underspecified ambiguous statements by inferring their meanings from contexts such as visual inputs. This ability, however, may not be as developed in recent pre-trained visionlanguage models (VLMs). In this paper, we introduce a novel probing dataset called FOCUS to evaluate whether state-of-the-art VLMs have this ability. FOCUS consists of underspecified sentences paired with image contexts and carefully designed probing questions. Our experiments reveal that VLMs still fall short in handling underspecification even when visual inputs that can help resolve the ambiguities are available. To further support research in underspecification, FOCUS will be released …


Empowering Weight Loss: A Pragmatic Randomized Controlled Trial Of A Theory-Driven Self-Regulation Mobile App For Young Adults With Excess Body Weight, H. S. J. Chew, J. W. Ngooi, R. C. Du, P. Z. Chan, M. Jansson, B. Zhu, Y. Cao, Chong-Wah Ngo, R. Foo, A. Shabbir, D. Ho, N. Sevdalis, K. Y. Ngiam Jul 2025

Empowering Weight Loss: A Pragmatic Randomized Controlled Trial Of A Theory-Driven Self-Regulation Mobile App For Young Adults With Excess Body Weight, H. S. J. Chew, J. W. Ngooi, R. C. Du, P. Z. Chan, M. Jansson, B. Zhu, Y. Cao, Chong-Wah Ngo, R. Foo, A. Shabbir, D. Ho, N. Sevdalis, K. Y. Ngiam

Research Collection School Of Computing and Information Systems

Background/Introduction: Obesity is projected to affect more than half of the global population by 2035, posing significant health and economic challenges. While lifestyle modification is considered a cornerstone of weight management, its effectiveness often relies on substantial support systems. Purpose: This study aimed to evaluate the effectiveness of a 12-week, standalone Temporal Self-Regulation Theory (TST)-based weight loss mobile application, which integrates self-regulation techniques, food logging, and dietary nudging, in promoting weight loss among young adults with excess body weight. Methods: A two-arm, parallel-group, 1:1 randomized controlled trial was conducted, adhering to the CONSORT-Outcomes 2022 Extension guidelines. Participants completed a face-to-face …


Instruct2see: Learning To Remove Any Obstructions Across Distributions, Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He Jul 2025

Instruct2see: Learning To Remove Any Obstructions Across Distributions, Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He

Research Collection School Of Computing and Information Systems

Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehensive data collection impractical. To overcome these challenges, we propose Instruct2See, a novel zero-shot framework capable of handling both seen and unseen obstacles. The core idea of our approach is to unify obstruction removal by treating it as a soft-hard mask restoration problem, where any obstruction can be represented using multi-modal prompts, such as visual semantics and textual instructions, …


Cracking Aegis: An Adversarial Llm-Based Game For Raising Awareness Of Vulnerabilities In Privacy Protection, Jiaying Fu, Yiyang Lu, Zehua Yang, Fiona Fui-Hoon Nah, Ray Lc Jul 2025

Cracking Aegis: An Adversarial Llm-Based Game For Raising Awareness Of Vulnerabilities In Privacy Protection, Jiaying Fu, Yiyang Lu, Zehua Yang, Fiona Fui-Hoon Nah, Ray Lc

Research Collection School Of Computing and Information Systems

Traditional methods for raising awareness of privacy protection often fail to engage users or provide hands-on insights into how privacy vulnerabilities are exploited. To address this, we incorporate an adversarial mechanic in the design of the dialogue-based serious game Cracking Aegis. Leveraging LLMs to simulate natural interactions, the game challenges players to impersonate characters and extract sensitive information from an AI agent, Aegis. A user study (n=22) revealed that players employed diverse deceptive linguistic strategies, including storytelling and emotional rapport, to manipulate Aegis. After playing, players reported connecting in-game scenarios with real-world privacy vulnerabilities, such as phishing and impersonation, and …


Advancing Food Nutrition Estimation Via Visual-Ingredient Feature Fusion, Huiyan Qi, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Ee-Peng Lim Jul 2025

Advancing Food Nutrition Estimation Via Visual-Ingredient Feature Fusion, Huiyan Qi, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Nutrition estimation is an important component of promoting healthy eating and mitigating diet-related health risks. Despite advances in tasks such as food classification and ingredient recognition, progress in nutrition estimation is limited due to the lack of datasets with nutritional annotations. To address this issue, we introduce FastFood, a dataset with 84,446 images across 908 fast food categories, featuring ingredient and nutritional annotations. In addition, we propose a new model-agnostic Visual-Ingredient Feature Fusion (VIF2 ) method to enhance nutrition estimation by integrating visual and ingredient features. Ingredient robustness is improved through synonym replacement and resampling strategies during training. The ingredient-aware …


Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan Jul 2025

Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan

Research Collection School Of Computing and Information Systems

Graph Transformers (GTs) have demonstrated remarkable performance in graph representation learning over popular graph neural networks (GNNs). However, self-attention, the core module of GTs, preserves only low-frequency signals in graph features, leading to ineffectiveness in capturing other important signals like high-frequency ones. Some recent GT models help alleviate this issue, but their flexibility and expressiveness are still limited since the filters they learn are fixed on predefined graph spectrum or spectral order. To tackle this challenge, we propose a Graph Fourier Kolmogorov-Arnold Transformer (GrokFormer), a novel GT model that learns highly expressive spectral filters with adaptive graph spectrum and spectral …


Unambiguous Granularity Distillation For Asymmetric Image Retrieval, Hongrui Zhang, Yi Xie, Haoquan Zhang, Cheng Xu, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng Ann Heng, Shengfeng He Jul 2025

Unambiguous Granularity Distillation For Asymmetric Image Retrieval, Hongrui Zhang, Yi Xie, Haoquan Zhang, Cheng Xu, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng Ann Heng, Shengfeng He

Research Collection School Of Computing and Information Systems

Previous asymmetric image retrieval methods based on knowledge distillation have primarily focused on aligning the global features of two networks to transfer global semantic information from the gallery network to the query network. However, these methods often fail to effectively transfer local semantic information, limiting the fine-grained alignment of feature representation spaces between the two networks. To overcome this limitation, we propose a novel approach called Layered-Granularity Localized Distillation (GranDist). GranDist constructs layered feature representations that balance the richness of contextual information with the granularity of local features. As we progress through the layers, the contextual information becomes more detailed, …


Hdifftg: A Lightweight Hybrid Diffusion-Transformer-Gcn Architecture For 3d Human Pose Estimation, Yajie Fu, Chaorui Huang, Junwei Li, Hui Kong, Yibin Tian, Huakang Li, Zhiyuan Zhang Jul 2025

Hdifftg: A Lightweight Hybrid Diffusion-Transformer-Gcn Architecture For 3d Human Pose Estimation, Yajie Fu, Chaorui Huang, Junwei Li, Hui Kong, Yibin Tian, Huakang Li, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

We propose HDiffTG, a novel 3D Human Pose Estimation (3DHPE) method that integrates Transformer, Graph Convolutional Network (GCN), and diffusion model into a unified framework. HDiffTG leverages the strengths of these techniques to significantly improve pose estimation accuracy and robustness while maintaining a lightweight design. The Transformer captures global spatiotemporal dependencies, the GCN models local skeletal structures, and the diffusion model provides step-by-step optimization for fine-tuning, achieving a complementary balance between global and local features. This integration enhances the model’s ability to handle pose estimation under occlusions and in complex scenarios. Furthermore, we introduce lightweight optimizations to the integrated model …


Efficient Prompt Tuning For Hierarchical Ingredient Recognition, Yinxuan Gui, Bin Zhu, Jingjing Chen, Chong-Wah Ngo Jul 2025

Efficient Prompt Tuning For Hierarchical Ingredient Recognition, Yinxuan Gui, Bin Zhu, Jingjing Chen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Fine-grained ingredient recognition presents a significant challenge due to the diverse appearances of ingredients, resulting from different cutting and cooking methods. While existing approaches have shown promising results, they still require extensive training costs and focus solely on fine-grained ingredient recognition. In this paper, we address these limitations by introducing an efficient prompt-tuning framework that adapts pretrained visual-language models (VLMs), such as CLIP, to the ingredient recognition task without requiring full model finetuning. Additionally, we introduce three-level ingredient hierarchies to enhance both training performance and evaluation robustness. Specifically, we propose a hierarchical ingredient recognition task, designed to evaluate model performance …


Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He Jul 2025

Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He

Research Collection School Of Computing and Information Systems

We introduce the task of Audible Action Temporal Localization, which aims to identify the spatiotemporal coordinates of audible movements. Unlike conventional tasks such as action recognition and temporal action localization, which broadly analyze video content, our task focuses on the distinct kinematic dynamics of audible actions. It is based on the premise that key actions are driven by inflectional movements; for example, collisions that produce sound often involve abrupt changes in motion. To capture this, we propose T A2Net, a novel architecture that estimates inflectional flow using the second derivative of motion to determine collision timings without relying on audio …


Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo Jul 2025

Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for …


Retrieval Augmented Generation For Dynamic Graph Modeling, Yuxia Wu, Lizi Liao, Yuan Fang Jul 2025

Retrieval Augmented Generation For Dynamic Graph Modeling, Yuxia Wu, Lizi Liao, Yuan Fang

Research Collection School Of Computing and Information Systems

Modeling dynamic graphs, such as those found in social networks, recommendation systems, and e-commerce platforms, is crucial for capturing evolving relationships and delivering relevant insights over time. Traditional approaches primarily rely on graph neural networks with temporal components or sequence generation models, which often focus narrowly on the historical context of target nodes. This limitation restricts the ability to adapt to new and emerging patterns in dynamic graphs. To address this challenge, we propose a novel framework, Retrieval-Augmented Generation for Dy namic Graph modeling (RAG4DyG ), which enhances dynamic graph predictions by incorporating contextually and temporally relevant examples from broader …


Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du Jun 2025

Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du

Research Collection School Of Computing and Information Systems

Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, …