Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2821 - 2850 of 63010

Full-Text Articles in Computer Sciences

Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He Oct 2025

Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He

Research Collection School Of Computing and Information Systems

3Dvisual grounding aims to identify and localize objects in a 3Dspacebasedontextualdescriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsisten cies in spatial descriptions caused by perspective variations. To tackle these challenges, we propose ViewSRD, a frame work that formulates 3D visual grounding as a structured multi-view decomposition process. First, the Simple Rela tion Decoupling (SRD) module restructures complex multi anchor queries into a set of targeted single-anchor state ments, generating a structured set of perspective-aware de scriptions that clarify positional relationships. These de composed representations serve as the foundation for the Multi-view …


Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu Oct 2025

Polymind: Parallel Visual Diagramming With Large Language Models To Support Prewriting Through Microtasks, Qian Wan, Jiannan Li, Huanchen Wang, Zhicong Lu

Research Collection School Of Computing and Information Systems

Prewriting is the process of generating and organising ideas before a first draft. It consists of a combination of informal, iterative, and semi-structured strategies such as visual diagramming, which poses a challenge for collaborating with large language models (LLMs) in a turn-taking conversational manner. We present Polymind, a visual diagramming tool that leverages multiple LLM-powered agents to support prewriting. The system features a parallel collaboration workflow in place of the turn-taking conversational interactions. It defines multiple ''microtasks'' to simulate group collaboration scenarios such as collaborative writing and group brainstorming. Instead of repetitively prompting a chatbot for various purposes, Polymind enables …


Filterfl: Knowledge Filtering-Based Data-Free Backdoor Defense For Federated Learning, Yanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao, Pengyu Zhang, Yihao Huang, Mingsong Chen Oct 2025

Filterfl: Knowledge Filtering-Based Data-Free Backdoor Defense For Federated Learning, Yanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao, Pengyu Zhang, Yihao Huang, Mingsong Chen

Research Collection School Of Computing and Information Systems

As a distributed machine learning paradigm, Federated Learning (FL) enables large-scale clients to collaboratively train a model without sharing their raw data. However, due to the lack of data auditing for untrusted clients, FL is vulnerable to poisoning attacks, especially backdoor attacks. By using poisoned data for local training or directly changing the model parameters, attackers can easily inject backdoors into the model, which can trigger the model to make misclassification of targeted patterns in images. To address these issues, we propose a novel data-free trigger-generation-based defense approach based on the two characteristics of backdoor attacks: i) triggers are learned …


Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu Oct 2025

Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu

Research Collection School Of Computing and Information Systems

Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings generate partially inaccurate representations that, when fed into diffusion models, accumulate errors and degrade reconstruction fidelity. To address this, we propose the Bidirectional Autoencoder Intertwining framework for accurate decoded representation prediction. Our approach unifies multiple subjects through a Subject Bias Modulation Module while leveraging bidirectional mapping to better capture data distributions for precise representation prediction. To further enhance fidelity when decoding representations into stimulus images, …


Memad: Structured Memory Of Debates For Enhanced Multi-Agent Reasoning, Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan Oct 2025

Memad: Structured Memory Of Debates For Enhanced Multi-Agent Reasoning, Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) demonstrate remarkable in-context learning capabilities but often struggle with complex, multi-step reasoning. Multi-Agent Debate (MAD) frameworks partially address these limitations by enabling iterative agent interactions. However, they neglect valuable historical insights by treating each new debate independently. In this paper, we propose Memory-Augmented MAD (MeMAD), a parameter-free memory-augmented MAD framework that systematically organizes and reuses past debate transcripts. MeMAD stores structured representations of successful and unsuccessful reasoning attempts enriched with self-reflections and peer feedback. It systematically retrieves them via semantic similarity at inference time to inform new reasoning tasks. Our experiments on challenging mathematical reasoning, scientific …


Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars, Hyuna Seo, Youngki Lee, Rajesh Krishna Balan, Thivya Kandappu Oct 2025

Emoshortcuts: Emotionally Expressive Body Augmentation For Social Mixed Reality Avatars, Hyuna Seo, Youngki Lee, Rajesh Krishna Balan, Thivya Kandappu

Research Collection School Of Computing and Information Systems

We present EmoShortcuts1, a novel social Mixed Reality (MR) framework that enhances emotional expression by dynamically augmenting avatar body gestures to reflect users’ emotional states. While social MR enables immersive remote interactions through avatars, conveying emotions remains challenging due to limitations in head-mounted display (HMD) tracking (e.g., missing lower-body movements, such as stomping or defensive postures), and users’ tendency to deprioritize nonverbal expressions during multitasking. EmoShortcuts addresses these challenges by introducing an augmentation framework that generates expressive body gestures even when users’ physical movements are restricted. We conducted a formative study with 12 participants to identify key challenges in emotional …


Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao Oct 2025

Boosting Chart-To-Code Generation In Mllm Via Dual Preference-Guided Refinement, Zhihan Zhang, Yixin Cao, Lizi Liao

Research Collection School Of Computing and Information Systems

Translating chart images into executable plotting scripts-referred to as the chart-to-code generation task-requires Multimodal Large Language Models (MLLMs) to perform fine-grained visual parsing, precise code synthesis, and robust cross-modal reasoning. However, this task is inherently under-constrained: multiple valid code implementations can produce the same visual chart, and evaluation must consider both code correctness and visual fidelity across diverse dimensions. This makes it difficult to learn accurate and generalizable mappings through standard supervised fine-tuning. To address these challenges, we propose a dual preference-guided refinement framework that combines a feedback-driven, dual-modality reward mechanism with iterative preference learning. Our approach introduces a structured …


Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He Oct 2025

Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He

Research Collection School Of Computing and Information Systems

Handwritten Mathematical Expression Recognition (HMER) remains a challenging task due to the structural complexity of mathematical notation and the ambiguity of handwritten symbols-e.g., ''ρ'' vs. ''p'' or ''B'' vs. ''β''. While stroke-based models offer disambiguation via temporal cues, most existing methods are constrained by coarse modality fusion and a lack of fine-grained cross-modal alignment, further hindered by limited annotated data. We introduce Art for Math (Art4Math), a novel framework that leverages the structural richness of human sketches to enhance HMER through fine-grained, modality-aware learning. Art4Math follows a two-stage training paradigm: Art Grounding (A-Grd) and Math Decoding (M-Dec). In A-Grd, the …


Fine-Grained Abnormality Prompt Learning For Zero-Shot Anomaly Detection, Jiawen Zhu, Yew‑Soon Ong, Chunhua Shen, Guansong Pang Oct 2025

Fine-Grained Abnormality Prompt Learning For Zero-Shot Anomaly Detection, Jiawen Zhu, Yew‑Soon Ong, Chunhua Shen, Guansong Pang

Research Collection School Of Computing and Information Systems

Current zero-shot anomaly detection (ZSAD) methods show remarkable success in prompting large pre-trained visionlanguage models to detect anomalies in a target dataset without using any dataset-specific training or demonstration. However, these methods often focus on crafting/learning prompts that capture only coarse-grained semantics of abnormality, e.g., high-level semantics like ‘damaged’, ‘imperfect’, or ‘defective’ objects. They therefore have limited capability in recognizing diverse abnormality details that deviate from these general abnormal patterns in various ways. To address this limitation, we propose FAPrompt, a novel framework designed to learn Fine-grained Abnormality Prompts for accurate ZSAD. To this end, a novel Compound Abnormality Prompt …


Spd: Shallow Backdoor Protecting Deep Backdoor Against Backdoor Detection, Shunjie Yuan, Xinghua Li, Xuelin Cao, Haiyan Zhang, Mengyao Zhu, Robert H. Deng Oct 2025

Spd: Shallow Backdoor Protecting Deep Backdoor Against Backdoor Detection, Shunjie Yuan, Xinghua Li, Xuelin Cao, Haiyan Zhang, Mengyao Zhu, Robert H. Deng

Research Collection School Of Computing and Information Systems

Backdoor attacks have revealed the vulnerability of deep neural networks (DNNs), which motivates the development of secure deep learning systems. However, existing backdoor attacks often fail to bypass backdoor detection and human visual inspection, resulting in the exposure of the backdoor implanted in DNNs, which can subsequently be significantly mitigated through pruning or fine-tuning on benign data. To address this issue, in this paper, we propose a novel backdoor attack called SPD (Shallow Protecting Deep), which consists of a deep backdoor in the frequency domain and a shallow backdoor in the pixel domain, where the shallow backdoor acts as a …


Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai Oct 2025

Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai

Research Collection School Of Computing and Information Systems

Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitation: a lack of diversity in the extracted OOD features from the few-shot ID data.To address this issue, we propose to leverage external datasets as auxiliary outlier data (i.e., pseudo OOD samples) to extract rich, diverse OOD features, with the features from not only background regions but also foreground object regions, thereby supporting more discriminative prompt tuning for OOD detection. We further introduce …


Impact Of Original Versus Reposted Social Endorsements On Content Consumption: The Moderating Role Of Endorsers’ Network Characteristics, Anqi Zhao, Qian Tang Oct 2025

Impact Of Original Versus Reposted Social Endorsements On Content Consumption: The Moderating Role Of Endorsers’ Network Characteristics, Anqi Zhao, Qian Tang

Research Collection School Of Computing and Information Systems

Social endorsements broadcast endorsers’ positive attitudes toward content or products, especially to their social ties. Original endorsements created by endorsers can be propagated further as reposted endorsements. Both are important marketing tools to increase content consumption, yet their differences are unclear. This study compares the impacts of original and reposted endorsements on content consumption and their contingencies on the endorsers’ network characteristics. Using data on social endorsements of YouTube videos on Twitter, we find that original endorsements (i.e., original tweets) significantly boost content consumption, and the effect is positively moderated by the endorsers’ network size but not their tie strength. …


Juxtaposing Approaches To Risk-Based Ai Governance In Different ‘Rights’ Contexts: A Comparative Analysis Between Singapore And The Eu, Jane Loo, Mark Findlay Oct 2025

Juxtaposing Approaches To Risk-Based Ai Governance In Different ‘Rights’ Contexts: A Comparative Analysis Between Singapore And The Eu, Jane Loo, Mark Findlay

Research Collection Yong Pung How School Of Law

Comparative analysis of European and certain Asian approaches to governance often degenerates into simplistic dichotomies based on universal human rights assumptions. This chapter rejects such dualities, ill-informed by theory and historical reflection. The emerging argument is founded on a historical realist approach to theorising difference. Assisted by Polanyi’s double movement, the detailed substantive comparison is preceded by considerations of how recent trends in governing AI have uniformly adopted a countermovement against the dis-embedding of data and technology from the social leading to a risk/responsibility paradigm. From here, a more nuanced reflection of AI governance approaches in the EU and Singapore …


Multi-Period Risk-Aware Procurement Optimization Under Covid-19 Disruption, Jonathan Chase, Hoong Chuin Lau, Jinfeng Yang, Lu Liu Oct 2025

Multi-Period Risk-Aware Procurement Optimization Under Covid-19 Disruption, Jonathan Chase, Hoong Chuin Lau, Jinfeng Yang, Lu Liu

Research Collection School Of Computing and Information Systems

Supply chain resilience has been a topic of active research in the operations research and AI communities for several years, but the COVID-19 pandemic threw the frailties of global supply chains into sharp relief. Disruptions and delays caused by fresh outbreaks leading to lockdowns, put severe strain on supply chains in many industries. In this work we develop lockdown-resilient procurement capabilities for a global technology company. First, through analysis of lockdown data from China we develop a logarithmic regression-based lockdown prediction method to complement a supplier risk metric for conventional risks. Second, we develop a multi-period stochastic optimization model that …


Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang Oct 2025

Look Before You Decide: Prompting Active Deduction Of Mllms For Assumptive Reasoning, Yian Li, Wentao Tian, Yang Jiao, Jingjing Chen, Tianwen Qian, Bin Zhu, Na Zhao, Yu‑Gang Jiang

Research Collection School Of Computing and Information Systems

Recently, Multimodal Large Language Models (MLLMs) have achieved significant success across multiple disciplines due to their exceptional instruction-following capabilities and extensive world knowledge. However, whether these MLLMs possess human-like compositional reasoning abilities remains an open problem. To unveil their reasoning behaviors, we first curate a Multimodal Assumptive Reasoning Benchmark (MARS-Bench) in this paper. Interestingly, we find that most prevalent MLLMs can be easily fooled by the introduction of a presupposition into the question, whereas such presuppositions appear naive to human reasoning. Besides, we also propose a simple yet effective method, Active Deduction (AD), a novel reinforcement learning paradigm to encourage …


Memory-Efficient 4-Bit Preconditioned Stochastic Optimization, Jingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan Zhou Oct 2025

Memory-Efficient 4-Bit Preconditioned Stochastic Optimization, Jingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan Zhou

Research Collection School Of Computing and Information Systems

Preconditioned stochastic optimization algorithms, exemplified by Shampoo, outperform first-order optimizers by offering theoretical convergence benefits and practical gains in large-scale neural network training. However, they incur substantial memory overhead due to the storage demands of non-diagonal preconditioning matrices. To address this, we introduce 4-bit quantization for Shampoo’s preconditioners. We introduce two key methods: First, we apply Cholesky decomposition followed by quantization of the Cholesky factors, reducing memory usage by leveraging their lower triangular structure while better preserving spectral properties to minimize information loss. To our knowledge, this is the first quantization approach applied to Cholesky factors of preconditioners. Second, we …


What Students Really Think: Unpacking Ai Ethics In Educational Assessments Through A Triadic Framework, Lim Ming Soon Tristan, Gottipati Swapna, Michelle L. F. Cheong Oct 2025

What Students Really Think: Unpacking Ai Ethics In Educational Assessments Through A Triadic Framework, Lim Ming Soon Tristan, Gottipati Swapna, Michelle L. F. Cheong

Research Collection School Of Computing and Information Systems

The rise of AI in educational assessments has significantly enhanced efficiency and accuracy. However, it also introduces critical ethical challenges, including bias in grading, data privacy risks, and accountability gaps. These issues can undermine trust in AI-driven assessments and compromise educational fairness, making a structured ethical framework essential. To address these challenges, this study empirically validates an existing triadic ethical framework for AI-assisted educational assessments, originally proposed by Lim, Gottipati and Cheong (In: Keengwe (ed) Creative AI tools and ethical implications in teaching and learning, IGI Global, 2023), grounded in student perceptions. The framework encompasses three ethical domains—physical, cognitive, and …


Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington Oct 2025

Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington

Research Collection School Of Computing and Information Systems

Cooking plays a vital role in everyday independence and well-being, yet remains challenging for people with vision impairments due to limited support for tracking progress and receiving contextual feedback. Object status — the condition or transformation of ingredients and tools — offers a promising but underexplored foundation for context-aware cooking support. In this paper, we present OSCAR (Object Status Context Awareness for Recipes), a technical pipeline that explores the use of object status recognition to enable recipe progress tracking in non-visual cooking. OSCAR integrates recipe parsing, object status extraction, visual alignment with cooking steps, and time-causal modeling to support real-time …


A Comprehensive Review Of Financial Knowledge Graphs, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang Oct 2025

A Comprehensive Review Of Financial Knowledge Graphs, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang

Research Collection School Of Computing and Information Systems

Knowledge Graphs (KGs) are increasingly used in finance to manage complex, interconnected data and support advanced analytics. This survey provides an overview of how KGs are applied across various financial areas, such as fraud detection, credit risk assessment, anti-money laundering, and regulatory compliance. We examine key techniques for building and using KGs in finance, including graph construction, embedding methods, and machine learning models. The survey also discusses challenges specific to finance, like handling private data, ensuring interpretability, and managing real-time data. Additionally, we explore the emerging combination of KGs with large language models and generative AI, which offers new possibilities …


Rethinking Teaching Evaluation Reports: Designing Ai-Transformed Student Feedback For Instructor Engagement, Ruoxi Shang, Keri Mallari, Au Wei Bin Yeong, Ken Yasuhara, Anthony Tang, Gary Hsieh Oct 2025

Rethinking Teaching Evaluation Reports: Designing Ai-Transformed Student Feedback For Instructor Engagement, Ruoxi Shang, Keri Mallari, Au Wei Bin Yeong, Ken Yasuhara, Anthony Tang, Gary Hsieh

Research Collection School Of Computing and Information Systems

Student feedback is critical for improving teaching, yet instructors often avoid reading evaluations due to emotional burden and information overload. We present a systematic exploration of how language models can distill and transform student evaluations into adaptive, actionable insights. Through a systematic design space exploration combining 4 feedback strategies (removing harmful content, paraphrasing criticism, sandwiching negatives, adding constructive suggestions) with 4 presentation formats (themes, cards, letters, chatbots), we created six AI-augmented prototypes of teaching evaluations. Interviews with 16 post-secondary instructors revealed that effective use of AI in feedback processing should: (1) support action formation through focused views and divergent thinking, …


Conditional Attribute-Based Pre: Definition And Construction From Lwe, Lisha Yao, Jian Weng, Pengfei Wu, Guofeng Tang, Guomin Yang, Haiyang Xue, Robert H. Deng Oct 2025

Conditional Attribute-Based Pre: Definition And Construction From Lwe, Lisha Yao, Jian Weng, Pengfei Wu, Guofeng Tang, Guomin Yang, Haiyang Xue, Robert H. Deng

Research Collection School Of Computing and Information Systems

Attribute-based proxy re-encryption (AB-PRE) is a crucial variant of proxy re-encryption. It allows a proxy with a re-encryption key to transform a delegator’s ciphertext associated with an access policy into another ciphertext associated with a new access policy, enabling delegatees with matching attributes to decrypt the transformed ciphertext. However, a key limitation of AB-PRE is that the delegator cannot control which ciphertexts are transformed. As a result, the proxy, once given the re-encryption key, indiscriminately transforms all ciphertexts, effectively switching their underlying policies—an issue known as the all-or-nothing problem. It limits the system’s flexibility and practicality in real-world use cases.In …


Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He Oct 2025

Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He

Research Collection School Of Computing and Information Systems

Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieve cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's …


Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He Oct 2025

Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He

Research Collection School Of Computing and Information Systems

The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …


Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du Oct 2025

Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du

Research Collection School Of Computing and Information Systems

Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve …


Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang Oct 2025

Fcad: Feature-Coupled Anisotropic Diffusion For Continuous Graph Learning, Amitoz Azad, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

In this work, we propose a novel continuous graph neural network called FCAD (Feature-Coupled Anisotropic Diffusion) for the task of node classification on graphs. Our approach is motivated by the success of feature-coupled anisotropic diffusion PDEs in multivalued image restoration. Our method introduces a total variation regularization-inspired anisotropic term to control diffusion between nodes and incorporates a learnable parameterization for feature coupling during the diffusion process. Our model performs competitively against several GNN baselines for both heterophilous and homophilous graphs, demonstrating notable benefits for heterophilous graphs due to the learnable feature coupling.


Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang Oct 2025

Lsfdnet: A Single-Stage Fusion And Detection Network For Ships Using Swir And Lwir, Yanyin Guo, Runxuan An, Junwei Li, Zhiyuan Zhang

Research Collection School Of Computing and Information Systems

Traditional ship detection methods primarily rely on single-modal approaches, such as visible or infrared images, which limit their application in complex scenarios involving varying lighting conditions and heavy fog. To address this issue, we explore the advantages of short-wave infrared (SWIR) and long-wave infrared (LWIR) in ship detection and propose a novel single-stage image fusion detection algorithm called LSFDNet. This algorithm leverages feature interaction between the image fusion and object detection subtask networks, achieving remarkable detection performance and generating visually impressive fused images. To further improve the saliency of objects in the fused images and improve the performance of the …


Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu Oct 2025

Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu

Research Collection School Of Computing and Information Systems

Tactile graphics on a refreshable display have proven effective in enabling visually impaired people to comprehend pictorial content. To further evaluate the effectiveness of refreshable tactile displays in blind education, we designed tactile data comics, a method that combines step-by-step presentation of tactile graphics with verbal narration. We conducted a user study with sixteen visually impaired students to compare tactile data comics against verbal-only and static tactile graphics. Our findings show that tactile data comics significantly improve participants’ comprehension and engagement during the learning experience. These empirical results suggest that the integration of refreshable tactile displays and tactile data comics …


Search Trajectory Network-Enhanced Multi-Objective Dynamic Algorithm Configuration, Robbert Reijnen, Zaharah Bukhsh, Hoong Chuin Lau, Yaoxin Wu, Yingqian Zhang Oct 2025

Search Trajectory Network-Enhanced Multi-Objective Dynamic Algorithm Configuration, Robbert Reijnen, Zaharah Bukhsh, Hoong Chuin Lau, Yaoxin Wu, Yingqian Zhang

Research Collection School Of Computing and Information Systems

Deep reinforcement learning (DRL) has emerged as an effective technique for dynamic algorithm configuration, particularly in evolutionary computation, enabling adaptive parameter updates during algorithmic execution. DRL-based methods have shown broad applicability across different problem domains and are designed to configure algorithms without problem-specific information, making them highly transferable across problem variants and scalable to different problem sizes. This paper proposes a novel graph neural network-based approach that learns representations of Search Trajectory Networks (STNs) to track the convergence behavior of multiple objectives and dynamically reconfigures multiobjective evolutionary algorithms during execution. By capturing how solutions evolve and interact over time, the …


Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He Oct 2025

Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He

Research Collection School Of Computing and Information Systems

In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …


Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang Oct 2025

Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang

Research Collection School Of Computing and Information Systems

Website owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable …