Open Access. Powered by Scholars. Published by Universities.®

Graphics and Human Computer Interfaces

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 61 - 90 of 414

Full-Text Articles in Artificial Intelligence and Robotics

Registration Is A Powerful Rotation-Invariance Learner For 3d Anomaly Detection, Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang, Haoxin Yang, Yongwei Nie, Shengfeng He Dec 2025

Registration Is A Powerful Rotation-Invariance Learner For 3d Anomaly Detection, Yuyang Yu, Zhengwei Chen, Xuemiao Xu, Lei Zhang, Haoxin Yang, Yongwei Nie, Shengfeng He

Research Collection School Of Computing and Information Systems

3D anomaly detection in point-cloud data is critical for industrial quality control, aiming to identify structural defects with high reliability. However, current memory bank-based methods often suffer from inconsistent feature transformations and limited discriminative capacity, particularly in capturing local geometric details and achieving rotation invariance. These limitations become more pronounced when registration fails, leading to unreliable detection results. We argue that point-cloud registration plays an essential role not only in aligning geometric structures but also in guiding feature extraction toward rotation-invariant and locally discriminative representations. To this end, we propose a registration-induced, rotation-invariant feature extraction framework that integrates the objectives …


Jury-And-Judge Chain-Of-Thought For Uncovering Toxic Data In 3d Visual Grounding, Kaixiang Huang, Qifeng Zhang, Jin Wang, Jingru Yang, Yang Zhou, Huan Yu, Guodong Lu, Shengfeng He Dec 2025

Jury-And-Judge Chain-Of-Thought For Uncovering Toxic Data In 3d Visual Grounding, Kaixiang Huang, Qifeng Zhang, Jin Wang, Jingru Yang, Yang Zhou, Huan Yu, Guodong Lu, Shengfeng He

Research Collection School Of Computing and Information Systems

3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framework that harnesses the reasoning capabilities of Multimodal Large Language Models (MLLMs) to identify and mitigate toxic data. At the core of Refer-Judge is a Jury-and-Judge Chain-of-Thought paradigm, inspired by the deliberative process of the judicial system. This framework targets the root causes of annotation noise: jurors collaboratively assess 3DVG samples from diverse perspectives, providing structured, multi-faceted evaluations. Judges then consolidate these insights …


Instance-Level Video Depth In Groups Beyond Occlusions, Yuan Liang, Yang Zhou, Ziming Sun, Tianyi Xiang, Guiqing Li, Shengfeng He Dec 2025

Instance-Level Video Depth In Groups Beyond Occlusions, Yuan Liang, Yang Zhou, Ziming Sun, Tianyi Xiang, Guiqing Li, Shengfeng He

Research Collection School Of Computing and Information Systems

Depth estimation in dynamic, multi-object scenes remains a major challenge, especially under severe occlusions. Existing monocular models, including foundation models, struggle with instance-wise depth consistency due to their reliance on global regression. We tackle this problem from two key aspects: data and methodology. First, we introduce the Group Instance Depth (GID) dataset, the first large-scale video depth dataset with instance-level annotations, featuring 101,500 frames from real-world activity scenes. GID bridges the gap between synthetic and real-world depth data by providing high-fidelity depth supervision for multi-object interactions. Second, we propose InstanceDepth, the first occlusion-aware depth estimation framework for multi-object environments. Our …


Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao Nov 2025

Graph Perturbations For Robust Knowledge Discovery And Retrieval, Hanhua Xiao

Dissertations and Theses Collection (Open Access)

Graph perturbation, rooted in classical perturbation theory, studies how small topology edits, i.e., adding or deleting edges, affects graph properties (e.g., density, centrality). This fundamental problem underpins applications like bioinformatics, privacy preservation and system defense. While much prior work targets perturbations that influence global graph statistics or model outputs, comparatively little addresses robustness for knowledge discovery and information retrieval. In these settings, graphs are attributed: nodes carry real-world semantics (e.g., locations, people) and edges encode interactions or relationships. This thesis proposes new formulations and algorithms that generate and leverage graph perturbations to make knowledge discovery and retrieval more robust. Specifically, …


Instructors’ Strategies In Creating And Implementing Constructivist Llm-Based Learning Activities, Emily Aurelia, Shun Yi Yeo, Michelle Lui, Effie Lai-Chong Law, Anthony Tang Nov 2025

Instructors’ Strategies In Creating And Implementing Constructivist Llm-Based Learning Activities, Emily Aurelia, Shun Yi Yeo, Michelle Lui, Effie Lai-Chong Law, Anthony Tang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) are increasingly being integrated into educational settings, enabling more adoption of constructivist teaching and learning approaches in classrooms. This paper explores the strategies instructors are currently using to incorporate LLMs into learning activities that align with constructivist principles, which emphasize that learners actively construct their own knowledge. Through interviews with nine instructors who have designed eleven distinct LLM-based activities and using reflexive thematic analysis, this study identifies various types of learning activities with respect to four different aspects of the constructivist learning theory. The strategies employed and challenges faced to foster constructivist student-LLM interaction were also …


Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo Nov 2025

Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures.In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual …


Genwardrobe: A Fully Generative System For Travel Fashion Wardrobe Construction, Peng Jin, Yilin Wen, Mingzhe Yu, Yunshan Ma, Rong Zheng, Jin‑Tu Fan, Chong Wah Ngo Oct 2025

Genwardrobe: A Fully Generative System For Travel Fashion Wardrobe Construction, Peng Jin, Yilin Wen, Mingzhe Yu, Yunshan Ma, Rong Zheng, Jin‑Tu Fan, Chong Wah Ngo

Research Collection School Of Computing and Information Systems

With the increasing demand for outfit planning in real-world travel scenarios, the need for constructing a travel fashion wardrobe, a series of outfits tailored to a user's personalization and destination-specific context over a short travel period, has grown significantly. However, existing systems or works often focus on isolated factors and rely on retrieval-based methods, with insufficient utilization of generative models, limiting their adaptability to real-world travel scenarios. To address this issue, this study introduces GenWardrobe, a fully generative system for travel fashion wardrobe construction. GenWardrobe consists of three key modules: user query analysis, fashion knowledge retrieval via retrieval-augmented generation and …


Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang Oct 2025

Cookingdiffusion: Cooking Procedural Image Generation With Stable Diffusion, Yuan Wang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Yi Tan, Xiang Wang

Research Collection School Of Computing and Information Systems

Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation. This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while …


Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He Oct 2025

Stroke2sketch: Harnessing Stroke Attributes For Training-Free Sketch Generation, Rui Yang, Huining Li, Yiyi Long, Xiaojun Wu, Shengfeng He

Research Collection School Of Computing and Information Systems

Generating sketches guided by reference styles requires precise transfer of stroke attributes, such as line thickness, deformation, and texture sparsity, while preserving semantic structure and content fidelity. To this end, we propose Stroke2Sketch, a novel training-free framework that introduces cross-image stroke attention, a mechanism embedded within self-attention layers to establish fine-grained semantic correspondences and enable accurate stroke attribute transfer. This allows our method to adaptively integrate reference stroke characteristics into content images while maintaining structural integrity. Additionally, we develop adaptive contrast enhancement and semanticfocused attention to reinforce content preservation and foreground emphasis. Stroke2Sketch effectively synthesizes stylistically faithful sketches that closely …


Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He Oct 2025

Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He

Research Collection School Of Computing and Information Systems

Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieve cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's …


Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He Oct 2025

Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He

Research Collection School Of Computing and Information Systems

The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …


Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du Oct 2025

Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du

Research Collection School Of Computing and Information Systems

Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve …


Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu Oct 2025

Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu

Research Collection School Of Computing and Information Systems

Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings generate partially inaccurate representations that, when fed into diffusion models, accumulate errors and degrade reconstruction fidelity. To address this, we propose the Bidirectional Autoencoder Intertwining framework for accurate decoded representation prediction. Our approach unifies multiple subjects through a Subject Bias Modulation Module while leveraging bidirectional mapping to better capture data distributions for precise representation prediction. To further enhance fidelity when decoding representations into stimulus images, …


Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang Oct 2025

Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

In this work, we propose a novel continuous graph neural network called FCAD (Feature-Coupled Anisotropic Diffusion) for the task of node classification on graphs. Our approach is motivated by the success of feature-coupled anisotropic diffusion PDEs in multivalued image restoration. Our method introduces a total variation regularization-inspired anisotropic term to control diffusion between nodes and incorporates a learnable parameterization for feature coupling during the diffusion process. Our model performs competitively against several GNN baselines for both heterophilous and homophilous graphs, demonstrating notable benefits for heterophilous graphs due to the learnable feature coupling.


Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim Oct 2025

Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …


Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He Oct 2025

Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He

Research Collection School Of Computing and Information Systems

In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …


Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo Oct 2025

Diffusionmat: Alpha Matting As Deterministic Sequential Refinement Learning, Yangyang Xu, Shengfeng He, Wenqi Shao, Yong Du, Kwan-Yee K. Wong, Yu Qiao, Jun Yu, Ping Luo

Research Collection School Of Computing and Information Systems

In this paper, we introduce DiffusionMat, a novel image matting framework that employs a diffusion model for the transition from coarse to refined alpha mattes. Diverging from conventional methods that utilize trimaps merely as loose guidance for alpha matte prediction, our approach treats image matting as a deterministic sequential refinement learning process. This process begins with the addition of noise to trimaps and iteratively denoises them using a pre-trained diffusion model, which incrementally guides the prediction towards a clean alpha matte. The key innovation of our framework is a correction module that adjusts the output at each denoising step, ensuring …


Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang Sep 2025

Vibemus: Proactive Agentic System For Music Personalization, Zhiliang Guo, Teng Tu, Yunshan Ma, Xun Yang

Research Collection School Of Computing and Information Systems

Large language models (LLMs) enable diverse forms of AI-assisted creation, yet they often struggle to bridge the preference-articulation gap: users may provide incomplete or vague intentions or lack the vocabulary to specify what they want, yielding outputs misaligned with true preferences. To address this gap and facilitate music creation in a vibe-centric environment, we introduce VibeMus, a proactive agentic system built on open-source components. The system engages in multi-turn dialogue to progressively determine the music’s emotion, genre, lyrics, and other aspects before generation. Simulated evaluations show that proactive clarification improves alignment with users’ intended nuances. Our approach is training-free, leveraging …


Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong Sep 2025

Towards Multimodal Emotional Support Conversation Systems, Yuqi Chu, Lizi Liao, Zhiyuan Zhou, Chong-Wah Ngo, Richang Hong

Research Collection School Of Computing and Information Systems

The integration of conversational artificial intelligence (AI) into mental health care promises a new horizon for therapist-client interactions, aiming to closely emulate the depth and nuance of human conversations. Despite the potential, the current landscape of conversational AI is markedly limited by its reliance on single-modal data, constraining the systems’ ability to empathize and provide effective emotional support. This limitation stems from a paucity of resources that encapsulate the multimodal nature of human communication essential for therapeutic counseling. To address this gap, we introduce the Multimodal Emotional Support Conversation (MESC) dataset, a first-of-its-kind resource enriched with comprehensive annotations across text, …


Machine Learning And Crime Prevention, Emily Lizewski Aug 2025

Machine Learning And Crime Prevention, Emily Lizewski

Student Theses

Predictive policing uses machine learning to analyze crime patterns and help law enforcement better efficient use their resources. These tools can improve accuracy by highlighting complex trends in large sets of data. While this technology has its advantages, it also raises important ethical and social questions. Within this paper we looks at how predictive policing works, focusing on the machine learning models often used such as decision trees, random forests, gradient boosting, and models that factor in both time and location. It also explores how these tools might unintentionally reinforce biases already present in historical crime data. In reviewing the …


Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He Aug 2025

Dreamanime: Learning Style-Identity Textual Disentanglement For Anime And Beyond, Chenshu Xu, Yangyang Xu, Huaidong Zhang, Xuemiao Xu, Shengfeng He

Research Collection School Of Computing and Information Systems

Text-to-image generation models have significantly broadened the horizons of creative expression through the power of natural language. However, navigating these models to generate unique concepts, alter their appearance, or reimagine them in unfamiliar roles presents an intricate challenge. For instance, how can we exploit language-guided models to transpose an anime character into a different art style, or envision a beloved character in a radically different setting or role? This paper unveils a novel approach named DreamAnime, designed to provide this level of creative freedom. Using a minimal set of 2-3 images of a user-specified concept such as an anime character …


Bhvit: Binarized Hybrid Vision Transformer, Tian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu, Kaijie Yin, Chengzhong Xu, Hui Kong Aug 2025

Bhvit: Binarized Hybrid Vision Transformer, Tian Gao, Yu Zhang, Zhiyuan Zhang, Huajun Liu, Kaijie Yin, Chengzhong Xu, Hui Kong

Research Collection School Of Computing and Information Systems

Model binarization has made significant progress in enabling real-time and energy-efficient computation for con-volutional neural networks (CNN), offering a potential solution to the deployment challenges faced by Vision Transformers (ViTs) on edge devices. However, due to the structural differences between CNN and Transformer architectures, simply applying binary CNN strategies to the ViT models will lead to a significant performance drop. To tackle this challenge, we propose BHViT, a binarization-friendly hybrid ViT architecture and its full binarization model with the guidance of three important observations. Initially, BHViT utilizes the local information interaction and hierarchical feature aggregation technique from coarse to fine …


Modern Procedural Terrain Generation Techniques And Their Background, Hunter A. Barton Jul 2025

Modern Procedural Terrain Generation Techniques And Their Background, Hunter A. Barton

2025 Symposium

Procedural terrain generation has become a staple in many digital environments, enabling the automated creation of large-scale and realistic landscapes for applications such as video games and movies. This paper provides an in-depth look at smooth noise functions and their use for terrain generation, as well as an overview of some more modern methods of generation. A method utilizing machine learning stlye transfer was reproduced for this paper with some alterations to improve visualization and realism.


Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan Jul 2025

Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan

Research Collection School Of Computing and Information Systems

Graph Transformers (GTs) have demonstrated remarkable performance in graph representation learning over popular graph neural networks (GNNs). However, self-attention, the core module of GTs, preserves only low-frequency signals in graph features, leading to ineffectiveness in capturing other important signals like high-frequency ones. Some recent GT models help alleviate this issue, but their flexibility and expressiveness are still limited since the filters they learn are fixed on predefined graph spectrum or spectral order. To tackle this challenge, we propose a Graph Fourier Kolmogorov-Arnold Transformer (GrokFormer), a novel GT model that learns highly expressive spectral filters with adaptive graph spectrum and spectral …


Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He Jul 2025

Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He

Research Collection School Of Computing and Information Systems

We introduce the task of Audible Action Temporal Localization, which aims to identify the spatiotemporal coordinates of audible movements. Unlike conventional tasks such as action recognition and temporal action localization, which broadly analyze video content, our task focuses on the distinct kinematic dynamics of audible actions. It is based on the premise that key actions are driven by inflectional movements; for example, collisions that produce sound often involve abrupt changes in motion. To capture this, we propose T A2Net, a novel architecture that estimates inflectional flow using the second derivative of motion to determine collision timings without relying on audio …


Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo Jul 2025

Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for …


Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu Jun 2025

Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu

Research Collection School Of Computing and Information Systems

Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations …


Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du Jun 2025

Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du

Research Collection School Of Computing and Information Systems

Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, …


Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He Jun 2025

Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He

Research Collection School Of Computing and Information Systems

Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying sample quality across modalities can lead to the propagation of inaccurate information, resulting in error accumulation. To address this, we propose Modal-Affinity Multimodal Domain Adaptation (MODfinity), a method that dynamically manages multimodal information flow through fine-grained control over teacher model selection, guiding information intertwining at both feature and label levels. By treating labels as an independent modality, MODfinity enables balanced performance assessment across modalities, employing a novel …


Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al. Jun 2025

Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.

Research Collection School Of Computing and Information Systems

No abstract provided.