Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type

Articles 61 - 90 of 179

Full-Text Articles in Graphics and Human Computer Interfaces

Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang Aug 2025

Focus: Evaluating Pre-Trained Vision-Language Models On Underspecification Reasoning, Kankan Zhou, Yibin Lai, Kyriakos Mouratidis, Jing Jiang

Research Collection School Of Computing and Information Systems

Humans possess a remarkable ability to interpret underspecified ambiguous statements by inferring their meanings from contexts such as visual inputs. This ability, however, may not be as developed in recent pre-trained visionlanguage models (VLMs). In this paper, we introduce a novel probing dataset called FOCUS to evaluate whether state-of-the-art VLMs have this ability. FOCUS consists of underspecified sentences paired with image contexts and carefully designed probing questions. Our experiments reveal that VLMs still fall short in handling underspecification even when visual inputs that can help resolve the ambiguities are available. To further support research in underspecification, FOCUS will be released …


Relightable Neural Radiance Fields For Novel View Synthesis, Malena I. Mahoney Jul 2025

Relightable Neural Radiance Fields For Novel View Synthesis, Malena I. Mahoney

Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal

This paper describes relighting neural radiance fields for novel view synthesis. View synthesis is the problem of using input images with corresponding camera angles to produce a photorealistic 3D model of an environment and its objects. Neural radiance fields (NeRFs) were created as a solution to view synthesis. Neural radiance field models work well for generating realistic 3D models from 2D image inputs; how-ever, they do not support changing the lighting or placing the objects from the input images into different environments. The problem comes from the fact that NeRFs rely on a neural network that is essentially overfitted to …


Modern Procedural Terrain Generation Techniques And Their Background, Hunter A. Barton Jul 2025

Modern Procedural Terrain Generation Techniques And Their Background, Hunter A. Barton

2025 Symposium

Procedural terrain generation has become a staple in many digital environments, enabling the automated creation of large-scale and realistic landscapes for applications such as video games and movies. This paper provides an in-depth look at smooth noise functions and their use for terrain generation, as well as an overview of some more modern methods of generation. A method utilizing machine learning stlye transfer was reproduced for this paper with some alterations to improve visualization and realism.


Enhancing Graph Representation Learning Through Self-Supervision: An Augmentation Perspective, Jianyuan Bo Jul 2025

Enhancing Graph Representation Learning Through Self-Supervision: An Augmentation Perspective, Jianyuan Bo

Dissertations and Theses Collection (Open Access)

Graph representation learning has become fundamental in various domains, from social networks to molecular structures, enabling extraction of meaningful patterns from graph-structured data. While deep learning approaches, particularly graph neural networks, have shown promising results, their effectiveness is often limited by the scarcity of labeled data. This challenge is particularly acute in graph domains where annotation requires specialized expertise and is prohibitively expensive. Self-supervised learning has emerged as a promising direction to address this limitation by creating auxiliary tasks from unlabeled data, with augmentation strategies playing a crucial role in their success.

Current graph self-supervised learning methods face several critical …


Experimental Analysis Of Satellite Operator Training Using Game-Based Virtual Reality Simulation, Lana Laskey Jul 2025

Experimental Analysis Of Satellite Operator Training Using Game-Based Virtual Reality Simulation, Lana Laskey

Doctoral Dissertations and Master's Theses

Satellite data plays a vital role in modern global infrastructure by enabling communications, navigation, and weather forecasting. As demand for satellite technology grows, so does the need for highly trained satellite ground operators. Traditional training regimens for satellite operators employ simulation using two-dimensional computer console displays paired with the varied ability of trainees to generate abstract mental imagery of the scenario. However, this development of mental imagery imposes a considerable learning curve and cognitive workload on the trainee, which may negatively impact the user experience and knowledge gained during the training scenario.

This experimental study investigated the effects of game-based …


Empowering Weight Loss: A Pragmatic Randomized Controlled Trial Of A Theory-Driven Self-Regulation Mobile App For Young Adults With Excess Body Weight, H. S. J. Chew, J. W. Ngooi, R. C. Du, P. Z. Chan, M. Jansson, B. Zhu, Y. Cao, Chong-Wah Ngo, R. Foo, A. Shabbir, D. Ho, N. Sevdalis, K. Y. Ngiam Jul 2025

Empowering Weight Loss: A Pragmatic Randomized Controlled Trial Of A Theory-Driven Self-Regulation Mobile App For Young Adults With Excess Body Weight, H. S. J. Chew, J. W. Ngooi, R. C. Du, P. Z. Chan, M. Jansson, B. Zhu, Y. Cao, Chong-Wah Ngo, R. Foo, A. Shabbir, D. Ho, N. Sevdalis, K. Y. Ngiam

Research Collection School Of Computing and Information Systems

Background/Introduction: Obesity is projected to affect more than half of the global population by 2035, posing significant health and economic challenges. While lifestyle modification is considered a cornerstone of weight management, its effectiveness often relies on substantial support systems. Purpose: This study aimed to evaluate the effectiveness of a 12-week, standalone Temporal Self-Regulation Theory (TST)-based weight loss mobile application, which integrates self-regulation techniques, food logging, and dietary nudging, in promoting weight loss among young adults with excess body weight. Methods: A two-arm, parallel-group, 1:1 randomized controlled trial was conducted, adhering to the CONSORT-Outcomes 2022 Extension guidelines. Participants completed a face-to-face …


Instruct2see: Learning To Remove Any Obstructions Across Distributions, Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He Jul 2025

Instruct2see: Learning To Remove Any Obstructions Across Distributions, Junhang Li, Yu Guo, Chuhua Xian, Shengfeng He

Research Collection School Of Computing and Information Systems

Images are often obstructed by various obstacles due to capture limitations, hindering the observation of objects of interest. Most existing methods address occlusions from specific elements like fences or raindrops, but are constrained by the wide range of real-world obstructions, making comprehensive data collection impractical. To overcome these challenges, we propose Instruct2See, a novel zero-shot framework capable of handling both seen and unseen obstacles. The core idea of our approach is to unify obstruction removal by treating it as a soft-hard mask restoration problem, where any obstruction can be represented using multi-modal prompts, such as visual semantics and textual instructions, …


Cracking Aegis: An Adversarial Llm-Based Game For Raising Awareness Of Vulnerabilities In Privacy Protection, Jiaying Fu, Yiyang Lu, Zehua Yang, Fiona Fui-Hoon Nah, Ray Lc Jul 2025

Cracking Aegis: An Adversarial Llm-Based Game For Raising Awareness Of Vulnerabilities In Privacy Protection, Jiaying Fu, Yiyang Lu, Zehua Yang, Fiona Fui-Hoon Nah, Ray Lc

Research Collection School Of Computing and Information Systems

Traditional methods for raising awareness of privacy protection often fail to engage users or provide hands-on insights into how privacy vulnerabilities are exploited. To address this, we incorporate an adversarial mechanic in the design of the dialogue-based serious game Cracking Aegis. Leveraging LLMs to simulate natural interactions, the game challenges players to impersonate characters and extract sensitive information from an AI agent, Aegis. A user study (n=22) revealed that players employed diverse deceptive linguistic strategies, including storytelling and emotional rapport, to manipulate Aegis. After playing, players reported connecting in-game scenarios with real-world privacy vulnerabilities, such as phishing and impersonation, and …


Advancing Food Nutrition Estimation Via Visual-Ingredient Feature Fusion, Huiyan Qi, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Ee-Peng Lim Jul 2025

Advancing Food Nutrition Estimation Via Visual-Ingredient Feature Fusion, Huiyan Qi, Bin Zhu, Chong-Wah Ngo, Jingjing Chen, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Nutrition estimation is an important component of promoting healthy eating and mitigating diet-related health risks. Despite advances in tasks such as food classification and ingredient recognition, progress in nutrition estimation is limited due to the lack of datasets with nutritional annotations. To address this issue, we introduce FastFood, a dataset with 84,446 images across 908 fast food categories, featuring ingredient and nutritional annotations. In addition, we propose a new model-agnostic Visual-Ingredient Feature Fusion (VIF2 ) method to enhance nutrition estimation by integrating visual and ingredient features. Ingredient robustness is improved through synonym replacement and resampling strategies during training. The ingredient-aware …


Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan Jul 2025

Grokformer: Graph Fourier Kolmogorov‑Arnold Transformers, Guoguo Ai, Guansong Pang, Hezhe Qiao, Yuan Gao, Hui Yan

Research Collection School Of Computing and Information Systems

Graph Transformers (GTs) have demonstrated remarkable performance in graph representation learning over popular graph neural networks (GNNs). However, self-attention, the core module of GTs, preserves only low-frequency signals in graph features, leading to ineffectiveness in capturing other important signals like high-frequency ones. Some recent GT models help alleviate this issue, but their flexibility and expressiveness are still limited since the filters they learn are fixed on predefined graph spectrum or spectral order. To tackle this challenge, we propose a Graph Fourier Kolmogorov-Arnold Transformer (GrokFormer), a novel GT model that learns highly expressive spectral filters with adaptive graph spectrum and spectral …


Unambiguous Granularity Distillation For Asymmetric Image Retrieval, Hongrui Zhang, Yi Xie, Haoquan Zhang, Cheng Xu, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng Ann Heng, Shengfeng He Jul 2025

Unambiguous Granularity Distillation For Asymmetric Image Retrieval, Hongrui Zhang, Yi Xie, Haoquan Zhang, Cheng Xu, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng Ann Heng, Shengfeng He

Research Collection School Of Computing and Information Systems

Previous asymmetric image retrieval methods based on knowledge distillation have primarily focused on aligning the global features of two networks to transfer global semantic information from the gallery network to the query network. However, these methods often fail to effectively transfer local semantic information, limiting the fine-grained alignment of feature representation spaces between the two networks. To overcome this limitation, we propose a novel approach called Layered-Granularity Localized Distillation (GranDist). GranDist constructs layered feature representations that balance the richness of contextual information with the granularity of local features. As we progress through the layers, the contextual information becomes more detailed, …


Full-Stack Web Applications: Infrastructure, Development Pipelines & Devsecops, Yassine Chahid, Patrick Slattery Jul 2025

Full-Stack Web Applications: Infrastructure, Development Pipelines & Devsecops, Yassine Chahid, Patrick Slattery

Publications and Research

This research explores emerging development methodologies and technologies which facilitate the deployment and maintenance of software applications. It evaluates architectural styles for the development of software such as monolithic (legacy) and microservice models, with a focus on their key differences such as scalability or project structure through to the development of an application. By examining methodologies such as Agile and continuous integration/continuous development pipelines along with the deployment tools Docker and Git for version/release control, the study analyzes how these innovations speed up development, improve existing practices, and serve as the foundation for development operations. Cloud solutions for tasks such …


Hdifftg: A Lightweight Hybrid Diffusion-Transformer-Gcn Architecture For 3d Human Pose Estimation, Yajie Fu, Chaorui Huang, Junwei Li, Hui Kong, Yibin Tian, Huakang Li, Zhiyuan Zhang Jul 2025

Hdifftg: A Lightweight Hybrid Diffusion-Transformer-Gcn Architecture For 3d Human Pose Estimation, Yajie Fu, Chaorui Huang, Junwei Li, Hui Kong, Yibin Tian, Huakang Li, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

We propose HDiffTG, a novel 3D Human Pose Estimation (3DHPE) method that integrates Transformer, Graph Convolutional Network (GCN), and diffusion model into a unified framework. HDiffTG leverages the strengths of these techniques to significantly improve pose estimation accuracy and robustness while maintaining a lightweight design. The Transformer captures global spatiotemporal dependencies, the GCN models local skeletal structures, and the diffusion model provides step-by-step optimization for fine-tuning, achieving a complementary balance between global and local features. This integration enhances the model’s ability to handle pose estimation under occlusions and in complex scenarios. Furthermore, we introduce lightweight optimizations to the integrated model …


Efficient Prompt Tuning For Hierarchical Ingredient Recognition, Yinxuan Gui, Bin Zhu, Jingjing Chen, Chong-Wah Ngo Jul 2025

Efficient Prompt Tuning For Hierarchical Ingredient Recognition, Yinxuan Gui, Bin Zhu, Jingjing Chen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Fine-grained ingredient recognition presents a significant challenge due to the diverse appearances of ingredients, resulting from different cutting and cooking methods. While existing approaches have shown promising results, they still require extensive training costs and focus solely on fine-grained ingredient recognition. In this paper, we address these limitations by introducing an efficient prompt-tuning framework that adapts pretrained visual-language models (VLMs), such as CLIP, to the ingredient recognition task without requiring full model finetuning. Additionally, we introduce three-level ingredient hierarchies to enhance both training performance and evaluation robustness. Specifically, we propose a hierarchical ingredient recognition task, designed to evaluate model performance …


Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He Jul 2025

Action Dubber: Timing Audible Actions Via Inflectional Flow, Wenlong Wan, Weiying Zheng, Tianyi Xiang, Guiqing Li, Shengfeng He

Research Collection School Of Computing and Information Systems

We introduce the task of Audible Action Temporal Localization, which aims to identify the spatiotemporal coordinates of audible movements. Unlike conventional tasks such as action recognition and temporal action localization, which broadly analyze video content, our task focuses on the distinct kinematic dynamics of audible actions. It is based on the premise that key actions are driven by inflectional movements; for example, collisions that produce sound often involve abrupt changes in motion. To capture this, we propose T A2Net, a novel architecture that estimates inflectional flow using the second derivative of motion to determine collision timings without relying on audio …


Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo Jul 2025

Robust Relevance Feedback For Interactive Known-Item Video Search, Zhixin Ma, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Known-item search (KIS) involves only a single search target, making relevance feedback-typically a powerful technique for efficiently identifying multiple positive examples to infer user intent-inapplicable. PicHunter addresses this issue by asking users to select the top-k most similar examples to the unique search target from a displayed set. Under ideal conditions, when the user's perception aligns closely with the machine's perception of similarity, consistent and precise judgments can elevate the target to the top position within a few iterations. However, in practical scenarios, expecting users to provide consistent judgments is often unrealistic, especially when the underlying embedding features used for …


Retrieval Augmented Generation For Dynamic Graph Modeling, Yuxia Wu, Lizi Liao, Yuan Fang Jul 2025

Retrieval Augmented Generation For Dynamic Graph Modeling, Yuxia Wu, Lizi Liao, Yuan Fang

Research Collection School Of Computing and Information Systems

Modeling dynamic graphs, such as those found in social networks, recommendation systems, and e-commerce platforms, is crucial for capturing evolving relationships and delivering relevant insights over time. Traditional approaches primarily rely on graph neural networks with temporal components or sequence generation models, which often focus narrowly on the historical context of target nodes. This limitation restricts the ability to adapt to new and emerging patterns in dynamic graphs. To address this challenge, we propose a novel framework, Retrieval-Augmented Generation for Dy namic Graph modeling (RAG4DyG ), which enhances dynamic graph predictions by incorporating contextually and temporally relevant examples from broader …


Classification Of Human Trust In Ai Using Brain Activity Data, Danushka Bandara, Ruhuan Liao, Fatima Chowdhury, Leslie Abbott Jun 2025

Classification Of Human Trust In Ai Using Brain Activity Data, Danushka Bandara, Ruhuan Liao, Fatima Chowdhury, Leslie Abbott

Northeast Journal of Complex Systems (NEJCS)

Trust plays a crucial role in human-computer interaction, particularly in scenarios involving artificial intelligence (AI) systems. This study explores the feasibility of using functional near-infrared spectroscopy (fNIRS) data to classify trust levels in human-AI interaction scenarios. A total of 18 participants completed an image classification task with an AI team member while their hemodynamic responses were recorded using fNIRS. Preprocessing of fNIRS data involved motion artifact removal, filtering, and normalization. Exploratory analysis identified significant associations between hemodynamic responses in the prefrontal cortex and trust levels. An across-subject binary trust classification model was developed using machine learning techniques, achieving an F1 …


Multi-Level Differentiable Moving Particles With Partition Of Unity, Jinjin He Jun 2025

Multi-Level Differentiable Moving Particles With Partition Of Unity, Jinjin He

Dartmouth College Master’s Theses

Representing implicit geometry with intricate features has long been a challenge. Recent advances in Implicit Neural Representations (INRs) have shown great promise in applications such as 3D reconstruction, inverse rendering, and dynamic surface evolution. These methods leverage neural networks to model complex shapes continuously, offering advantages in resolution and flexibility over traditional discrete representations. Despite their success, efficiently handling fine geometric details and evolving dynamic scenes remains an open problem.

We introduce a differentiable moving particle representation based on the multi-level partition of unity (MPU) to model dynamic implicit geometries efficiently. Our approach employs two types of particles—feature particles and …


Tellings Of The Pacific Ocean: A Landscape-Based Approach For Multispecies Design And Hci, Maliheh Ghajargar Jun 2025

Tellings Of The Pacific Ocean: A Landscape-Based Approach For Multispecies Design And Hci, Maliheh Ghajargar

Engineering Faculty Articles and Research

Environmental disturbances induced by climate change have caused significant changes in our ecosystems and are threatening the health of our environments. As a response to this issue, a growing body of work has emerged in HCI and design, which seeks to foreground more-than-human stories in support of making more sustainable and just futures. This research contributes to this broad agenda by probing graphic novels as a multispecies storytelling method for design and HCI. Combining ideas from Anna Tsing’s adventures of landscape and from HCI and design’s use of sequential art (e.g., storyboards), we use landscape as the main protagonist of …


The Impact Of Accessibility Features On Player Experience In Video Games, Christine M. Widden Jun 2025

The Impact Of Accessibility Features On Player Experience In Video Games, Christine M. Widden

Master's Theses

While video game accessibility is a growing research topic, few studies investigate how players perceive the presence versus the absence of accessibility features, or how non-disabled players react to the option of accessibility features. This study explores these research gaps, investigating how access to accessibility features affects the experience of both disabled and non-disabled players. For the purposes of this study, a small platformer game was developed with as many accessibility features as feasible for the scope of the project. An A vs.\ B study was conducted in the game, with anonymous participants randomly assigned to version A, with all …


Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du Jun 2025

Nexusgs: Sparse View Synthesis With Epipolar Depth Priors In 3d Gaussian Splatting, Yulong Zheng, Zicheng Jiang, Shengfeng He, Yandu Sun, Junyu Dong, Huaidong Zhang, Yong Du

Research Collection School Of Computing and Information Systems

Neural Radiance Field (NeRF) and 3D Gaussian Splatting (3DGS) have noticeably advanced photo-realistic novel view synthesis using images from densely spaced camera viewpoints. However, these methods struggle in few-shot scenarios due to limited supervision. In this paper, we present NexusGS, a 3DGS-based approach that enhances novel view synthesis from sparse-view images by directly embedding depth information into point clouds, without relying on complex manual regularizations. Exploiting the inherent epipolar geometry of 3DGS, our method introduces a novel point cloud densification strategy that initializes 3DGS with a dense point cloud, reducing randomness in point placement while preventing over-smoothing and overfitting. Specifically, …


Hd-Epic: A Highly-Detailed Egocentric Video Dataset, Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guerrier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, Dima Damen Jun 2025

Hd-Epic: A Highly-Detailed Egocentric Video Dataset, Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Kumar Parida, Kaiting Liu, Prajwal Gatti, Siddhant Bansal, Kevin Flanagan, Jacob Chalk, Zhifan Zhu, Rhodri Guerrier, Fahd Abdelazim, Bin Zhu, Davide Moltisanti, Michael Wray, Hazel Doughty, Dima Damen

Research Collection School Of Computing and Information Systems

We present a validation dataset of newly-collected kitchenbased egocentric videos, manually annotated with highly detailed and interconnected ground-truth labels covering: recipe steps, fine-grained actions, ingredients with nutritional values, moving objects, and audio annotations. Importantly, all annotations are grounded in 3D through digital twinning of the scene, fixtures, object locations, and primed with gaze. Footage is collected from unscripted recordings in diverse home environments, making HDEPIC the first dataset collected in-the-wild but with detailed annotations matching those in controlled lab environments. We show the potential of our highly-detailed annotations through a challenging VQA benchmark of 26K questions assessing the capability to …


Dupin: A Parallel Framework For Densest Subgraph Discovery In Fraud Detection On Massive Graphs, Jiaxin Jiang, Siyuan Yao, Yuchen Li, Qiange Wang, Bingsheng He, Min Chen Jun 2025

Dupin: A Parallel Framework For Densest Subgraph Discovery In Fraud Detection On Massive Graphs, Jiaxin Jiang, Siyuan Yao, Yuchen Li, Qiange Wang, Bingsheng He, Min Chen

Research Collection School Of Computing and Information Systems

Detecting fraudulent activities in financial and e-commerce transaction networks is crucial. One effective method for this is Densest Subgraph Discovery (DSD). However, deploying DSD methods in production systems faces substantial scalability challenges due to the predominantly sequential nature of existing methods, which impedes their ability to handle large-scale transaction networks and results in significant detection delays. To address these challenges, we introduce Dupin, a novel parallel processing framework designed for efficient DSD processing in billion-scale graphs. Dupin is powered by a processing engine that exploits the unique properties of the peeling process, with theoretical guarantees on detection quality and efficiency. …


A Digital Dive: Redesigning The Cabrillo High School Aquarium Website, Jacob V. Cacho Jun 2025

A Digital Dive: Redesigning The Cabrillo High School Aquarium Website, Jacob V. Cacho

Graphic Communication

Tucked away on the Central Coast in Lompoc, you’ll find the Cabrillo High School (CHS) Aquarium. Started in 1986, the CHS Aquarium is the only high school aquarium of its kind in the nation run entirely by high school students. This 10,000+ square foot aquarium serves an underserved community at a Title I school, where students manage all aspects of animal care, nutrition, breeding, educational curriculum development, and visitor tours.

This program is truly one-of-a-kind and deserves the spotlight for just how unique it is. As a CHS graduate, I felt the current website lacked in many areas and could …


Community Detection In Heterogeneous Information Networks Without Materialization, Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu Jun 2025

Community Detection In Heterogeneous Information Networks Without Materialization, Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu

Research Collection School Of Computing and Information Systems

Community detection in heterogeneous information networks (HINs) poses significant challenges due to the diversity of entity types and the complexity of their interrelations. While traditional algorithms may perform adequately in some scenarios, many struggle with the high memory usage and computational demands of large-scale HINs. To address these challenges, we introduce a novel framework, SCAR, which efficiently uncovers community structures in HINs without requiring network materialization. SCAR leverages insights from meta-paths to interpret multi-relational data through compact vertex-based sketches, significantly reducing computational overhead and materialization overhead. We propose a sketch-based technique for estimating changes in modularity, improving both the precision …


Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu Jun 2025

Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu

Research Collection School Of Computing and Information Systems

Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations …


Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang Jun 2025

Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang

Research Collection School Of Computing and Information Systems

Low-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable inten sity. The former enforces …


Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al. Jun 2025

Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.

Research Collection School Of Computing and Information Systems

No abstract provided.


Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He Jun 2025

Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He

Research Collection School Of Computing and Information Systems

Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying sample quality across modalities can lead to the propagation of inaccurate information, resulting in error accumulation. To address this, we propose Modal-Affinity Multimodal Domain Adaptation (MODfinity), a method that dynamically manages multimodal information flow through fine-grained control over teacher model selection, guiding information intertwining at both feature and label levels. By treating labels as an independent modality, MODfinity enables balanced performance assessment across modalities, employing a novel …