Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (3436)
- Software Engineering (2136)
- Artificial Intelligence and Robotics (1648)
- Information Security (1051)
- Numerical Analysis and Scientific Computing (1024)
-
- Graphics and Human Computer Interfaces (916)
- Engineering (857)
- Social and Behavioral Sciences (661)
- Business (625)
- Theory and Algorithms (493)
- Computer Engineering (431)
- Operations Research, Systems Engineering and Industrial Engineering (399)
- Programming Languages and Compilers (379)
- OS and Networks (322)
- Communication (297)
- Social Media (240)
- Public Affairs, Public Policy and Public Administration (207)
- Transportation (185)
- Medicine and Health Sciences (177)
- Education (164)
- Management Information Systems (164)
- Data Storage Systems (160)
- E-Commerce (146)
- Health Information Technology (107)
- International and Area Studies (107)
- Asian Studies (106)
- Technology and Innovation (100)
- Digital Communications and Networking (96)
- Keyword
-
- Deep learning (122)
- Machine learning (121)
- Social media (74)
- Artificial intelligence (70)
- Reinforcement learning (69)
-
- Data mining (64)
- Privacy (61)
- Cloud computing (58)
- Deep Learning (55)
- Empirical study (54)
- Security (53)
- Optimization (52)
- Visualization (51)
- Software engineering (49)
- Training (49)
- Neural networks (48)
- Online learning (48)
- Task analysis (48)
- Anomaly detection (47)
- Singapore (47)
- Twitter (46)
- Feature extraction (45)
- Blockchain (44)
- Collaboration (44)
- Semantics (43)
- Large Language Models (42)
- Access control (41)
- Algorithms (40)
- Android (39)
- Machine Learning (38)
- Publication Year
- File Type
Articles 331 - 360 of 8458
Full-Text Articles in Computer Sciences
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu
Tactile Data Comics: Combining Step-By-Step Presentation Of Tactile Graphics With Verbal Narration For The Blind And Visually Impaired, Yang Jiao, Ruoting Sun, Rong Luo, Xiwen Yao, Xinran She, Kotaro Hara, Yuewen Zhang, Xinyi Fu
Research Collection School Of Computing and Information Systems
Tactile graphics on a refreshable display have proven effective in enabling visually impaired people to comprehend pictorial content. To further evaluate the effectiveness of refreshable tactile displays in blind education, we designed tactile data comics, a method that combines step-by-step presentation of tactile graphics with verbal narration. We conducted a user study with sixteen visually impaired students to compare tactile data comics against verbal-only and static tactile graphics. Our findings show that tactile data comics significantly improve participants’ comprehension and engagement during the learning experience. These empirical results suggest that the integration of refreshable tactile displays and tactile data comics …
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He
Language-Driven 3d Human Pose Estimation In Multi-Person Scenarios: A New Dataset And Approach, Tingrui Shen, Bangzhen Liu, Zhirun Fan, Shiting Zhang, Weifeng Pan, Sun Fan, Dan Cao, Shengfeng He
Research Collection School Of Computing and Information Systems
In an NBA game scenario, consider the challenge of locating and analyzing the 3D poses of players performing a user-specified action, such as attempting a shot. Traditional 3D human pose estimation (3DHPE) methods often fall short in such complex, multi-person scenes due to their lack of semantic integration and reliance on isolated pose data. To address these limitations, we introduce Language-Driven 3D Human Pose Estimation (L3DHPE), a novel approach that extends 3DHPE to general multi-person contexts by incorporating detailed language descriptions. We present Panoptic-L3D, the first dataset designed for L3DHPE, featuring 3,838 linguistic annotations for 1,476 individuals across 588 videos, …
Information-Bottleneck Driven Binary Neural Network For Change Detection, Kaijie Yin, Zhiyuan Zhang, Shu Kong, Tian Gao, Cheng-Zhong Xu, Hui Kong
Information-Bottleneck Driven Binary Neural Network For Change Detection, Kaijie Yin, Zhiyuan Zhang, Shu Kong, Tian Gao, Cheng-Zhong Xu, Hui Kong
Research Collection School Of Computing and Information Systems
In this paper, we propose Binarized Change Detection (BiCD), the first binary neural network (BNN) designed specifically for change detection. Conventional network binarization approaches, which directly quantize both weights and activations in change detection models, severely limit the network's ability to represent input data and distinguish between changed and unchanged regions. This results in significantly lower detection accuracy compared to real-valued networks. To overcome these challenges, BiCD enhances both the representational power and feature separability of BNNs, improving detection performance. Specifically, we introduce an auxiliary objective based on the Information Bottleneck (IB) principle, guiding the encoder to retain essential input …
Developing A Strong Cps Defender: An Evolutionary Approach, Qingyuan Hu, Christopher M. Poskitt, Jun Sun, Yuqi Chen
Developing A Strong Cps Defender: An Evolutionary Approach, Qingyuan Hu, Christopher M. Poskitt, Jun Sun, Yuqi Chen
Research Collection School Of Computing and Information Systems
Cyber-physical systems (CPSs) are used extensively in critical infrastructure, underscoring the need for anomaly detection systems that are able to catch even the most motivated attackers. Traditional anomaly detection techniques typically do `one-off' training on datasets crafted by experts or generated by fuzzers, potentially limiting their ability to generalize to unseen and more subtle attack strategies. Stopping at this point misses a key opportunity: a defender can actively challenge the attacker to find more nuanced attacks, which in turn can lead to more effective detection capabilities. Building on this concept, we propose Evo-Defender, an evolutionary framework that iteratively strengthens CPS …
Deep Learning For Hate Speech Detection: A Comparative Study, Jitendra Singh Malik, Hezhe Qiao, Guansong Pang, Anton Van Den Hengel
Deep Learning For Hate Speech Detection: A Comparative Study, Jitendra Singh Malik, Hezhe Qiao, Guansong Pang, Anton Van Den Hengel
Research Collection School Of Computing and Information Systems
Automated hate speech detection is an important tool in combating the spread of hate speech, particularly in social media. Numerous methods have been developed for the task, including a recent proliferation of deep-learning based approaches. A variety of datasets have also been developed, exemplifying various manifestations of the hate-speech detection problem. We present here a largescale empirical comparison of deep and shallow hate-speech detection methods, mediated through the three most commonly used datasets. Our goal is to illuminate progress in the area, and identify strengths and weaknesses in the current state-of-the-art. We particularly focus our analysis on measures of practical …
Multi-Period Risk-Aware Procurement Optimization Under Covid-19 Disruption, Jonathan Chase, Hoong Chuin Lau, Jinfeng Yang, Lu Liu
Multi-Period Risk-Aware Procurement Optimization Under Covid-19 Disruption, Jonathan Chase, Hoong Chuin Lau, Jinfeng Yang, Lu Liu
Research Collection School Of Computing and Information Systems
Supply chain resilience has been a topic of active research in the operations research and AI communities for several years, but the COVID-19 pandemic threw the frailties of global supply chains into sharp relief. Disruptions and delays caused by fresh outbreaks leading to lockdowns, put severe strain on supply chains in many industries. In this work we develop lockdown-resilient procurement capabilities for a global technology company. First, through analysis of lockdown data from China we develop a logarithmic regression-based lockdown prediction method to complement a supplier risk metric for conventional risks. Second, we develop a multi-period stochastic optimization model that …
Lightweight Population-Based Policy Optimization For Pickup And Delivery Problems, Yizhou Liu, Li Li, Yixin Xu, Tang Liu, Rong Cheng, Die Wu, Jilin Yang, Jingwen Li
Lightweight Population-Based Policy Optimization For Pickup And Delivery Problems, Yizhou Liu, Li Li, Yixin Xu, Tang Liu, Rong Cheng, Die Wu, Jilin Yang, Jingwen Li
Research Collection School Of Computing and Information Systems
In recent years, applying deep models to automatically learn construction heuristics for vehicle routing problems has achieved remarkable advancements. However, they are less effective in searching solutions due to two primary limitations: relying on deterministic probability distributions and overlooking the strategic advantage of prioritizing nearby unvisited nodes during the route construction process, resulting in suboptimal policies In this paper, we propose a novel lightweight population-based policy optimization (LPPO) framework that learns a diverse population of solution strategies through the utilization of innovative perturbation factors, in order to facilitate search exploration. Moreover, we design a localized attention synthesis (LAS) network to …
From Holistic To Localized: Local Enhanced Adapters For Efficient Visual Instruction Fine-Tuning, Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yugang Jiang
From Holistic To Localized: Local Enhanced Adapters For Efficient Visual Instruction Fine-Tuning, Pengkun Jiao, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yugang Jiang
Research Collection School Of Computing and Information Systems
Efficient Visual Instruction Fine-Tuning (EVIT) seeks to adapt Multimodal Large Language Models (MLLMs) to downstream tasks with minimal computational overhead. However, as task diversity and complexity increase, EVIT faces significant challenges in resolving data conflicts. To address this limitation, we propose the Dual Low-Rank Adaptation (Dual-LoRA), a holistic-to-local framework that enhances the adapter’s capacity to address data conflict through dual structural optimization. Specifically, we utilize two subspaces: a skill space for stable, holistic knowledge retention, and a rank-rectified task space that locally activates the holistic knowledge. Additionally, we introduce Visual Cue Enhancement (VCE), a multi-level local feature aggregation module designed …
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Exploring Object Status Recognition For Recipe Progress Tracking In Non-Visual Cooking, Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington
Research Collection School Of Computing and Information Systems
Cooking plays a vital role in everyday independence and well-being, yet remains challenging for people with vision impairments due to limited support for tracking progress and receiving contextual feedback. Object status — the condition or transformation of ingredients and tools — offers a promising but underexplored foundation for context-aware cooking support. In this paper, we present OSCAR (Object Status Context Awareness for Recipes), a technical pipeline that explores the use of object status recognition to enable recipe progress tracking in non-visual cooking. OSCAR integrates recipe parsing, object status extraction, visual alignment with cooking steps, and time-causal modeling to support real-time …
Reproducibility Debt In Scientific Software, Zara Hassan, Christoph Treude, Graham Williams, Michael Norrish, Alex Potanin
Reproducibility Debt In Scientific Software, Zara Hassan, Christoph Treude, Graham Williams, Michael Norrish, Alex Potanin
Research Collection School Of Computing and Information Systems
Reproducibility Debt (RpD) refers to accumulated technical and organisational issues in scientific software that hinder the ability to reproduce research results. While reproducibility is essential to scientific integrity, RpD remains poorly defined and under-addressed. This study introduces a formal definition of RpD and investigates its causes, effects, and mitigation strategies using a mixed-methods approach involving a systematic literature review (214 papers), interviews (23 practitioners), and a global survey (59 participants). We identify seven categories of contributing issues, 75 causes, 110 effects, and 61 mitigation strategies. Findings are synthesised into a cause-effect model and supported by taxonomies of team roles and …
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Viewsrd: 3d Visual Grounding Via Structured Multi-View Decomposition, Ronggang Huang, Haoxin Yang, Yan Cai, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
3Dvisual grounding aims to identify and localize objects in a 3Dspacebasedontextualdescriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsisten cies in spatial descriptions caused by perspective variations. To tackle these challenges, we propose ViewSRD, a frame work that formulates 3D visual grounding as a structured multi-view decomposition process. First, the Simple Rela tion Decoupling (SRD) module restructures complex multi anchor queries into a set of targeted single-anchor state ments, generating a structured set of perspective-aware de scriptions that clarify positional relationships. These de composed representations serve as the foundation for the Multi-view …
A Comprehensive Review Of Financial Knowledge Graphs, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang
A Comprehensive Review Of Financial Knowledge Graphs, Jeyaraman Brindha Priyadarshini, Bing Tian Dai, Yuan Fang
Research Collection School Of Computing and Information Systems
Knowledge Graphs (KGs) are increasingly used in finance to manage complex, interconnected data and support advanced analytics. This survey provides an overview of how KGs are applied across various financial areas, such as fraud detection, credit risk assessment, anti-money laundering, and regulatory compliance. We examine key techniques for building and using KGs in finance, including graph construction, embedding methods, and machine learning models. The survey also discusses challenges specific to finance, like handling private data, ensuring interpretability, and managing real-time data. Additionally, we explore the emerging combination of KGs with large language models and generative AI, which offers new possibilities …
Information Provision And Search Frictions: Evidence From The Taxi Industry In Singapore, Sumit Agarwal, Shih-Fen Cheng, Jussi Keppo, Long Wang, Yang Yang
Information Provision And Search Frictions: Evidence From The Taxi Industry In Singapore, Sumit Agarwal, Shih-Fen Cheng, Jussi Keppo, Long Wang, Yang Yang
Research Collection School Of Computing and Information Systems
Search frictions and misallocation are common in decentralized transportation markets. Using novel trip-level data of taxis in Singapore, this paper examines the impactof real-time demand information at airport terminals on search frictions. The information reduces taxi supply misallocation, increasing deadheading speed by 16.3% and decreasing deadheading time by 10.77%, benefiting both passengers and drivers. It raises daily earnings by $3.70 USD and adds 6.2 minutes of operational time per airport-trip taxi. Spatial spillovers are primarily observed among drivers in adjacentdistricts. Taxis from the Budget Terminal and drivers with fewer prior airport pickups benefit more from this information.
Filterfl: Knowledge Filtering-Based Data-Free Backdoor Defense For Federated Learning, Yanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao, Pengyu Zhang, Yihao Huang, Mingsong Chen
Filterfl: Knowledge Filtering-Based Data-Free Backdoor Defense For Federated Learning, Yanxin Yang, Ming Hu, Xiaofei Xie, Yue Cao, Pengyu Zhang, Yihao Huang, Mingsong Chen
Research Collection School Of Computing and Information Systems
As a distributed machine learning paradigm, Federated Learning (FL) enables large-scale clients to collaboratively train a model without sharing their raw data. However, due to the lack of data auditing for untrusted clients, FL is vulnerable to poisoning attacks, especially backdoor attacks. By using poisoned data for local training or directly changing the model parameters, attackers can easily inject backdoors into the model, which can trigger the model to make misclassification of targeted patterns in images. To address these issues, we propose a novel data-free trigger-generation-based defense approach based on the two characteristics of backdoor attacks: i) triggers are learned …
Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He
Stable Score Distillation, Haiming Zhu, Yangyang Xu, Chenshu Xu, Tingrui Shen, Wenxi Liu, Yong Du, Jun Yu, Shengfeng He
Research Collection School Of Computing and Information Systems
Text-guided image and 3D editing have advanced with diffusion-based models, yet methods like Delta Denoising Score often struggle with stability, spatial control, and editing strength. These limitations stem from reliance on complex auxiliary structures, which introduce conflicting optimization signals and restrict precise, localized edits. We introduce Stable Score Distillation (SSD), a streamlined framework that enhances stability and alignment in the editing process by anchoring a single classifier to the source prompt. Specifically, SSD utilizes Classifier-Free Guidance (CFG) equation to achieve cross-prompt alignment, and introduces a constant term null-text branch to stabilize the optimization process. This approach preserves the original content's …
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Seeing 3d Through 2d Lenses: 3d Few-Shot Class-Incremental Learning Via Cross-Modal Geometric Rectification, Tuo Xiang, Xuemiao Xu, Bangzhen Liu, Jinyi Li, Yong Li, Shengfeng He
Research Collection School Of Computing and Information Systems
The rapid growth of 3D digital content necessitates expandable recognition systems for open-world scenarios. However, existing 3D class-incremental learning methods struggle under extreme data scarcity due to geometric misalignment and texture bias. While recent approaches integrate 3D data with 2D foundation models (e.g., CLIP), they suffer from semantic blurring caused by texture-biased projections and indiscriminate fusion of geometric-textural cues, leading to unstable decision prototypes and catastrophic forgetting. To address these issues, we propose Cross-Modal Geometric Rectification (CMGR), a framework that enhances 3D geometric fidelity by leveraging CLIP’s hierarchical spatial semantics. Specifically, we introduce a Structure-Aware Geometric Rectification module that hierarchically …
Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du
Omnivton: Training-Free Universal Virtual Try-On, Zhaotong Yang, Yuhui Li, Shengfeng He, Xinzhe Li, Yangyang Xu, Junyu Dong, Yong Du
Research Collection School Of Computing and Information Systems
Image-based Virtual Try-On (VTON) techniques rely on either supervised in-shop approaches, which ensure high fidelity but struggle with cross-domain generalization, or unsupervised in-the-wild methods, which improve adaptability but remain constrained by data biases and limited universality. A unified, training-free solution that works across both scenarios remains an open challenge. We propose OmniVTON, the first training-free universal VTON framework that decouples garment and pose conditioning to achieve both texture fidelity and pose consistency across diverse settings. To preserve garment details, we introduce a garment prior generation mechanism that aligns clothing with the body, followed by continuous boundary stitching technique to achieve …
Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu
Cross-Subject Mind Decoding From Inaccurate Representations, Yangyang Xu, Bangzhen Liu, Wenqi Shao, Yong Du, Shengfeng He, Tingting Zhu
Research Collection School Of Computing and Information Systems
Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings generate partially inaccurate representations that, when fed into diffusion models, accumulate errors and degrade reconstruction fidelity. To address this, we propose the Bidirectional Autoencoder Intertwining framework for accurate decoded representation prediction. Our approach unifies multiple subjects through a Subject Bias Modulation Module while leveraging bidirectional mapping to better capture data distributions for precise representation prediction. To further enhance fidelity when decoding representations into stimulus images, …
Search Trajectory Network-Enhanced Multi-Objective Dynamic Algorithm Configuration, Robbert Reijnen, Zaharah Bukhsh, Hoong Chuin Lau, Yaoxin Wu, Yingqian Zhang
Search Trajectory Network-Enhanced Multi-Objective Dynamic Algorithm Configuration, Robbert Reijnen, Zaharah Bukhsh, Hoong Chuin Lau, Yaoxin Wu, Yingqian Zhang
Research Collection School Of Computing and Information Systems
Deep reinforcement learning (DRL) has emerged as an effective technique for dynamic algorithm configuration, particularly in evolutionary computation, enabling adaptive parameter updates during algorithmic execution. DRL-based methods have shown broad applicability across different problem domains and are designed to configure algorithms without problem-specific information, making them highly transferable across problem variants and scalable to different problem sizes. This paper proposes a novel graph neural network-based approach that learns representations of Search Trajectory Networks (STNs) to track the convergence behavior of multiple objectives and dynamically reconfigures multiobjective evolutionary algorithms during execution. By capturing how solutions evolve and interact over time, the …
Memad: Structured Memory Of Debates For Enhanced Multi-Agent Reasoning, Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan
Memad: Structured Memory Of Debates For Enhanced Multi-Agent Reasoning, Shuai Ling, Lizi Liao, Dongmei Jiang, Weili Guan
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) demonstrate remarkable in-context learning capabilities but often struggle with complex, multi-step reasoning. Multi-Agent Debate (MAD) frameworks partially address these limitations by enabling iterative agent interactions. However, they neglect valuable historical insights by treating each new debate independently. In this paper, we propose Memory-Augmented MAD (MeMAD), a parameter-free memory-augmented MAD framework that systematically organizes and reuses past debate transcripts. MeMAD stores structured representations of successful and unsuccessful reasoning attempts enriched with self-reflections and peer feedback. It systematically retrieves them via semantic similarity at inference time to inform new reasoning tasks. Our experiments on challenging mathematical reasoning, scientific …
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Mitigating Cross-Modal Representation Bias For Multicultural Image-To-Recipe Retrieval, Qing Wang, Chong-Wah Ngo, Yu Cao, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Existing approaches for image-to-recipe retrieval have the implicit assumption that a food image can fully capture the details textually documented in its recipe. However, a food image only reflects the visual outcome of a cooked dish and not the underlying cooking process. Consequently, learning cross-modal representations to bridge the modality gap between images and recipes tends to ignore subtle, recipe-specific details that are not visually apparent but are crucial for recipe retrieval. Specifically, the representations are biased to capture the dominant visual elements, resulting in difficulty in ranking similar recipes with subtle differences in use of ingredients and cooking methods. …
Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
Art4math: Handwritten Mathematical Expression Recognition Via Multimodal Sketch Grounding, Yang Zhou, Jin Wang, Yuxiao Zhang, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
Research Collection School Of Computing and Information Systems
Handwritten Mathematical Expression Recognition (HMER) remains a challenging task due to the structural complexity of mathematical notation and the ambiguity of handwritten symbols-e.g., ''ρ'' vs. ''p'' or ''B'' vs. ''β''. While stroke-based models offer disambiguation via temporal cues, most existing methods are constrained by coarse modality fusion and a lack of fine-grained cross-modal alignment, further hindered by limited annotated data. We introduce Art for Math (Art4Math), a novel framework that leverages the structural richness of human sketches to enhance HMER through fine-grained, modality-aware learning. Art4Math follows a two-stage training paradigm: Art Grounding (A-Grd) and Math Decoding (M-Dec). In A-Grd, the …
Fine-Grained Abnormality Prompt Learning For Zero-Shot Anomaly Detection, Jiawen Zhu, Yew‑Soon Ong, Chunhua Shen, Guansong Pang
Fine-Grained Abnormality Prompt Learning For Zero-Shot Anomaly Detection, Jiawen Zhu, Yew‑Soon Ong, Chunhua Shen, Guansong Pang
Research Collection School Of Computing and Information Systems
Current zero-shot anomaly detection (ZSAD) methods show remarkable success in prompting large pre-trained visionlanguage models to detect anomalies in a target dataset without using any dataset-specific training or demonstration. However, these methods often focus on crafting/learning prompts that capture only coarse-grained semantics of abnormality, e.g., high-level semantics like ‘damaged’, ‘imperfect’, or ‘defective’ objects. They therefore have limited capability in recognizing diverse abnormality details that deviate from these general abnormal patterns in various ways. To address this limitation, we propose FAPrompt, a novel framework designed to learn Fine-grained Abnormality Prompts for accurate ZSAD. To this end, a novel Compound Abnormality Prompt …
Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang
Website Owner Identification Through Multi-Level Contrastive Representation Learning, Cheng Tu, Yunshan Ma, Yang Li, Min Zhang, Miao Hu, Fan Shi, Xiang Wang
Research Collection School Of Computing and Information Systems
Website owner identification aims to recognize the organization or individual who owns a given website that is served on the web. It is a crucial step for cyberspace surveying and mapping, playing a significant role in cyberspace administration and governance. Existing widely employed solutions for website owner identification mainly fall into two paradigms: (1) querying the public information databases such as WHOIS, which store the Internet resource’s registered users or assignees; and (2) directly extracting the organization or individual name of the website owner from the webpage using the technique of named entity recognition. However, the former is less reliable …
Unsupervised Visual Chain-Of-Thought Reasoning Via Preference Optimization, Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang
Unsupervised Visual Chain-Of-Thought Reasoning Via Preference Optimization, Kesen Zhao, Beier Zhu, Qianru Sun, Hanwang Zhang
Research Collection School Of Computing and Information Systems
Chain-of-thought (CoT) reasoning greatly improves the interpretability and problem-solving abilities of multimodal large language models (MLLMs). However, existing ap proaches focus on text CoT, limiting their ability to lever age visual cues. Visual CoT remains underexplored, and the only work [35] is based on supervised fine-tuning that relies on extensive labeled bounding-box data and is hard to generalize to unseen cases. In this paper, we introduce Unsupervised Visual CoT (UV-CoT), a novel framework for image-level CoT reasoning via preference optimization. UV-CoTperforms preference comparisons between model generated bounding boxes (one is preferred and the other is dis-preferred), eliminating the need for …
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Auxiliary Prompt Tuning Of Vision‑Language Models For Few‑Shot Out‑Of‑Distribution Detection, Wenjun Miao, Guansong Pang, Zihan Wang, Jin Zheng, Xiao Bai
Research Collection School Of Computing and Information Systems
Recent advancements in CLIP-based out-of-distribution (OOD) detection have shown promising results via regularization on prompt tuning, leveraging background features extracted from a few in-distribution (ID) samples as proxies for OOD features.However, these methods suffer from an inherent limitation: a lack of diversity in the extracted OOD features from the few-shot ID data.To address this issue, we propose to leverage external datasets as auxiliary outlier data (i.e., pseudo OOD samples) to extract rich, diverse OOD features, with the features from not only background regions but also foreground object regions, thereby supporting more discriminative prompt tuning for OOD detection. We further introduce …
Morphology-Aware Hrv Estimation From Wrist Ppg In Sedentary Scenarios, Changshuo Hu, Hung Manh Pham, Dong Ma
Morphology-Aware Hrv Estimation From Wrist Ppg In Sedentary Scenarios, Changshuo Hu, Hung Manh Pham, Dong Ma
Research Collection School Of Computing and Information Systems
Photoplethysmography (PPG) is widely used in wearable devices for non-invasive heart rate variability (HRV) monitoring. While most prior work focuses on mitigating motion artifacts, recent studies highlight that even subtle contact pressure variations can distort waveform morphology and lead to inaccurate HRV estimates. In this work, we propose a morphology-aware deep learning framework that conditions HRV estimation on beat-level waveform types. Our model jointly encodes the raw PPG waveform and a sequence of pressure-induced morphology labels using parallel encoders, integrates them via cross-attention, and predicts normal-to-normal (NN) intervals and beat count to support downstream HRV computation. Evaluated on the public …
From Release To Adoption: Challenges In Reusing Pre-Trained Ai Models For Downstream Developers, Peerachai Banyongrakkul, Mansooreh Zahedi, Patanamon Thongtanunam, Christoph Treude, Haoyu Gao
From Release To Adoption: Challenges In Reusing Pre-Trained Ai Models For Downstream Developers, Peerachai Banyongrakkul, Mansooreh Zahedi, Patanamon Thongtanunam, Christoph Treude, Haoyu Gao
Research Collection School Of Computing and Information Systems
Pre-trained models (PTMs) have gained widespread popularity and achieved remarkable success across various fields, driven by their groundbreaking performance and easy accessibility through hosting providers. However, the challenges faced by downstream developers in reusing PTMs in software systems are less explored. To bridge this knowledge gap, we qualitatively created and analyzed a dataset of 840 PTM-related issue reports from 31 OSS GitHub projects. We systematically developed a comprehensive taxonomy of PTM-related challenges that developers face in downstream projects. Our study identifies seven key categories of challenges that downstream developers face in reusing PTMs, such as model usage, model performance, and …
Large Lithium-Ion Battery Model For Secure Shared E-Bike Battery In Smart Cities, Donghui Ding, Zhao Li, Linhao Luo, Ming Jin, Bin Zhu, Yichen Zhong, Junhao Hu, Peng Cai, Huiqi Hu
Large Lithium-Ion Battery Model For Secure Shared E-Bike Battery In Smart Cities, Donghui Ding, Zhao Li, Linhao Luo, Ming Jin, Bin Zhu, Yichen Zhong, Junhao Hu, Peng Cai, Huiqi Hu
Research Collection School Of Computing and Information Systems
Electric bikes powered by lithium-ion batteries are increasingly used in smart cities to promote sustainable mobility and efficient delivery services. However, limited battery range and slow plug-in charging remain key challenges. Shared electric bike battery systems, facilitated by battery swapping stations, offer a promising solution by enabling quick and efficient battery replacements. However, their success hinges on accurate anomaly detection, battery health estimation and remain range prediction. These tasks remain challenging due to data scarcity, battery diversity and environmental variability. Here we show that a large-scale lithium-ion battery model trained on over ten million battery time series data enables robust …
Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook, Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du, Hongmin Cai, Jing Qin, Shengfeng He
Stylegan-∞: Extending Stylegan To Arbitrary-Ratio Translation With Stylebook, Yihua Dai, Tianyi Xiang, Bailin Deng, Yong Du, Hongmin Cai, Jing Qin, Shengfeng He
Research Collection School Of Computing and Information Systems
Although pre-trained large-scale generative models StyleGAN series have proven to be effective in various editing and translation tasks, they are limited to pre-defined fixed aspect ratio. To overcome this limitation, we propose StyleGAN-∞, a model that enables pre-trained StyleGAN to perform arbitrary-ratio conditional synthesis. Our key insight is to distill the expressive StyleGAN features into a StyleBook, such that an arbitrary-ratio condition can be translated to other forms by properly assembling pre-defined StyleBook vectors. To learn and leverage the StyleBook, we employ a network with three distinct stages, each corresponding to StyleBook extraction, StyleBook correspondence learning, and arbitrary-ratio synthesis. Extensive …