Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (344)
- Engineering (291)
- Operations Research, Systems Engineering and Industrial Engineering (255)
- Graphics and Human Computer Interfaces (171)
- Software Engineering (129)
-
- Business (125)
- Numerical Analysis and Scientific Computing (111)
- Social and Behavioral Sciences (101)
- Theory and Algorithms (97)
- Public Affairs, Public Policy and Public Administration (66)
- Transportation (61)
- Programming Languages and Compilers (57)
- Information Security (40)
- OS and Networks (37)
- Medicine and Health Sciences (35)
- Computer Engineering (31)
- Education (24)
- Health Information Technology (24)
- Asian Studies (23)
- International and Area Studies (23)
- Finance and Financial Management (14)
- Technology and Innovation (13)
- Communication (12)
- Operations and Supply Chain Management (12)
- Social Media (12)
- Higher Education (11)
- Computer and Systems Architecture (7)
- Keyword
-
- Artificial intelligence (49)
- Machine learning (46)
- Reinforcement learning (39)
- Deep learning (37)
- Large Language Models (26)
-
- Large Language Model (22)
- Large language models (21)
- Computer vision (18)
- Optimization (18)
- Scheduling (18)
- Anomaly detection (17)
- Generative AI (17)
- Large language model (17)
- Singapore (17)
- ChatGPT (16)
- Deep reinforcement learning (16)
- Reinforcement Learning (16)
- LLMs (15)
- Vehicle routing problem (15)
- Deep Learning (14)
- Natural language processing (14)
- Artificial Intelligence (13)
- Neural networks (13)
- Uncertainty (13)
- Machine Learning (12)
- Graph neural networks (11)
- Metaverse (10)
- Multi-agent systems (10)
- Representation learning (10)
- Semantics (10)
- Publication Year
- File Type
Articles 571 - 600 of 1664
Full-Text Articles in Artificial Intelligence and Robotics
Jigsaw: Edge-Based Streaming Perception Over Spatially Overlapped Multi-Camera Deployments, Ila Gokarn, Yigong Hu, Tarek Abdelzaher, Archan Misra
Jigsaw: Edge-Based Streaming Perception Over Spatially Overlapped Multi-Camera Deployments, Ila Gokarn, Yigong Hu, Tarek Abdelzaher, Archan Misra
Research Collection School Of Computing and Information Systems
We present JIGSAW, a novel system that performs edge-based streaming perception over multiple video streams, while additionally factoring in the redundancy offered by the spatial overlap often exhibited in urban, multi-camera deployments. To assure high streaming throughput, JIGSAW extracts and spatially multiplexes multiple regions-of-interest from different camera frames into a smaller canvas frame. Moreover, to ensure that perception stays abreast of evolving object kinematics, JIGSAW includes a utility-based weighted scheduler to preferentially prioritize and even skip object-specific tiles extracted from an incoming stream of camera frames. Using the CityflowV2 traffic surveillance dataset, we show that JIGSAW can simultaneously process 25 …
Reinforcement Learning For Strategic Airport Slot Scheduling: Analysis Of State Observations And Reward Designs, Anh Nguyen-Duy, Duc-Thinh Pham, Jian-Yi Lye, Nguyen Binh Duong Ta
Reinforcement Learning For Strategic Airport Slot Scheduling: Analysis Of State Observations And Reward Designs, Anh Nguyen-Duy, Duc-Thinh Pham, Jian-Yi Lye, Nguyen Binh Duong Ta
Research Collection School Of Computing and Information Systems
Due to the NP-hard nature, the strategic airport slot scheduling problem is calling for exploring sub-optimal approaches, such as heuristics and learning-based approaches. Moreover, the continuous increase in air traffic demand requires approaches that can work well in new scenarios. While heuristics rely on a fixed set of rules, which limits the ability to explore new solutions, Reinforcement Learning offers a versatile framework to automate the search and generalize to unseen scenarios. Finding a suitable state observation and reward structure design is essential in using Reinforcement Learning. In this paper, we investigate the impact of providing the Reinforcement Learning agent …
Configurable Mirror Descent : Towards A Unification Of Decision Making, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Hau Chan, Bo An
Configurable Mirror Descent : Towards A Unification Of Decision Making, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Hau Chan, Bo An
Research Collection School Of Computing and Information Systems
Decision-making problems, categorized as single-agent, e.g., Atari, cooperative multi-agent, e.g., Hanabi, competitive multi-agent, e.g., Hold’em poker, and mixed cooperative and competitive, e.g., football, are ubiquitous in the real world. Although various methods have been proposed to address the specific decision-making categories, these methods typically evolve independently and cannot generalize to other categories. Therefore, a fundamental question for decision-making is: Can we develop a single algorithm to tackle ALL categories of decision-making problems? There are several main challenges to address this question: i) different decision-making categories involve different numbers of agents and different relationships between agents, ii) different categories have different …
Topic Modeling On Document Networks With Dirichlet Optimal Transport Barycenter (Extended Abstract), Ce Zhang, Hady Wirawan Lauw
Topic Modeling On Document Networks With Dirichlet Optimal Transport Barycenter (Extended Abstract), Ce Zhang, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
Texts are often interconnected in a network structure, e.g., academic papers via citations. On the one hand, though Graph Neural Networks (GNNs) have shown promising ability to derive effective embeddings for networked documents, they do not assume latent topics, resulting in uninterpretahle embeddings. On the other hand, topic models can infer interpretable document representations. However, most topic models focus on plain text and fail to leverage network structure across documents. In this paper, we propose a GNN-based topic model that both captures network connection and derives semantically interpretable text representations. For network modeling, we build our model with Optimal Transport …
Augmenting Decision With Hypothesis In Reinforcement Learning, Minh Quang Nguyen, Hady Wirawan Lauw
Augmenting Decision With Hypothesis In Reinforcement Learning, Minh Quang Nguyen, Hady Wirawan Lauw
Research Collection School Of Computing and Information Systems
Value-based reinforcement learning is the current State-Of-The-Art due to high sampling efficiency. However, our study shows it suffers from low exploitation in early training period and bias sensitiveness. To address these issues, we propose to augment the decision-making process with hypothesis, a weak form of environment description. Our approach relies on prompting the learning agent with accurate hypotheses, and designing a ready-to-adapt policy through incremental learning. We propose the ALH algorithm, showing detailed analyses on a typical learning scheme and a diverse set of Mujoco benchmarks. Our algorithm produces a significant improvement over value-based learning algorithms and other strong baselines. …
Unified Training Of Universal Time Series Forecasting Transformers, Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doyen Sahoo
Unified Training Of Universal Time Series Forecasting Transformers, Gerald Woo, Chenghao Liu, Akshat Kumar, Caiming Xiong, Silvio Savarese, Doyen Sahoo
Research Collection School Of Computing and Information Systems
Deep learning for time series forecasting has traditionally operated within a one-model-per-dataset framework, limiting its potential to leverage the game-changing impact of large pre-trained models. The concept of universal forecasting, emerging from pre-training on a vast collection of time series datasets, envisions a single Large Time Series Model capable of addressing diverse downstream forecasting tasks. However, constructing such a model poses unique challenges specific to time series data: i) cross-frequency learning, ii) accommodating an arbitrary number of variates for multivariate time series, and iii) addressing the varying distributional properties inherent in large-scale data. To address these challenges, we present novel …
The Information Content Of Financial Statement Fraud Risk: An Ensemble Learning Approach, Wei Duan, Nan Hu, Fujing Xue
The Information Content Of Financial Statement Fraud Risk: An Ensemble Learning Approach, Wei Duan, Nan Hu, Fujing Xue
Research Collection School Of Computing and Information Systems
This study aims to assess the financial statement fraud risk ex ante and empirically explore its information content to help improve decision-making and daily operations. We propose an ex-ante fraud risk index by adopting an ensemble learning approach and a theoretically grounded framework. Our ensemble learning model systematically examines the fraud process and deals effectively with the unique challenges in the financial fraud setting, which yields superior prediction performance. More importantly, we empirically examine the information content of our estimated ex-ante fraud risk from the perspective of operational efficiency. Our empirical results find that the estimated ex-ante fraud risk is …
Diffusion Models For Generative Outfit Recommendation, Yiyan Xu, Wenjie Wang, Fuli Feng, Yunshan Ma, Jizhi Zhang, Xiangnan He
Diffusion Models For Generative Outfit Recommendation, Yiyan Xu, Wenjie Wang, Fuli Feng, Yunshan Ma, Jizhi Zhang, Xiangnan He
Research Collection School Of Computing and Information Systems
Outfit Recommendation (OR) in the fashion domain has evolved through two stages: Pre-defined Outfit Recommendation and Personalized Outfit Composition. However, both stages are constrained by existing fashion products, limiting their effectiveness in addressing users' diverse fashion needs. Recently, the advent of AI-generated content provides the opportunity for OR to transcend these limitations, showcasing the potential for personalized outfit generation and recommendation.To this end, we introduce a novel task called Generative Outfit Recommendation (GOR), aiming to generate a set of fashion images and compose them into a visually compatible outfit tailored to specific users. The key objectives of GOR lie in …
Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun
Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun
Research Collection School Of Computing and Information Systems
Existing learning-based methods for solving job shop scheduling problems (JSSP) usually use off-the-shelf GNN models tailored to undirected graphs and neglect the rich and meaningful topological structures of disjunctive graphs (DGs). This paper proposes the topology-aware bidirectional graph attention network (TBGAT), a novel GNN architecture based on the attention mechanism, to embed the DG for solving JSSP in a local search framework. Specifically, TBGAT embeds the DG from a forward and a backward view, respectively, where the messages are propagated by following the different topologies of the views and aggregated via graph attention. Then, we propose a novel operator based …
Adaptive Stabilization Based On Machine Learning For Column Generation, Yunzhuang Shen, Yuan Sun, Xiaodong Li, Zhiguang Cao, Eberhard Andrew, Guangquan Zhang
Adaptive Stabilization Based On Machine Learning For Column Generation, Yunzhuang Shen, Yuan Sun, Xiaodong Li, Zhiguang Cao, Eberhard Andrew, Guangquan Zhang
Research Collection School Of Computing and Information Systems
Column generation (CG) is a well-established method for solving large-scale linear programs. It involves iteratively optimizing a subproblem containing a subset of columns and using its dual solution to generate new columns with negative reduced costs. This process continues until the dual values converge to the optimal dual solution to the original problem. A natural phenomenon in CG is the heavy oscillation of the dual values during iterations, which can lead to a substantial slowdown in the convergence rate. Stabilization techniques are devised to accelerate the convergence of dual values by using information beyond the state of the current subproblem. …
Mvmoe: Multi-Task Vehicle Routing Solver With Mixture-Of-Experts, Jianan Zhou, Zhiguang Cao, Yaoxin Wu, Wen Song, Yining Ma, Jie Zhang, Chi Xu
Mvmoe: Multi-Task Vehicle Routing Solver With Mixture-Of-Experts, Jianan Zhou, Zhiguang Cao, Yaoxin Wu, Wen Song, Yining Ma, Jie Zhang, Chi Xu
Research Collection School Of Computing and Information Systems
Learning to solve vehicle routing problems (VRPs) has garnered much attention. However, most neural solvers are only structured and trained independently on a specific problem, making them less generic and practical. In this paper, we aim to develop a unified neural solver that can cope with a range of VRP variants simultaneously. Specifically, we propose a multi-task vehicle routing solver with mixture-of-experts (MVMoE), which greatly enhances the model capacity without a proportional increase in computation. We further develop a hierarchical gating mechanism for the MVMoE, delivering a good trade-off between empirical performance and computational complexity. Experimentally, our method significantly promotes …
Is There A Space In Landslide Susceptibility Modelling: A Case Study Of Valtellina Valley, Northern Italy, Min Naing Khant, Mei Yi Victoria Grace Ann, Tin Seong Kam
Is There A Space In Landslide Susceptibility Modelling: A Case Study Of Valtellina Valley, Northern Italy, Min Naing Khant, Mei Yi Victoria Grace Ann, Tin Seong Kam
Research Collection School Of Computing and Information Systems
Landslides pose significant and ever-threatening risks to human life and infrastructure worldwide. Landslide susceptibility modelling is an emerging field of research seeking to determine contributing factors of these events. Yet, previous studies rarely explored the spatial variation of different landslide factors. Hence, this study aims to demonstrate the potential contribution of spatial nonstationarity in landslide susceptibility modelling using Global Logistic Regression (GLR) and Geographically Weighted Logistic Regression (GWLR). The second objective of this study is to demonstrate the important role of data preparation, data sampling, variable sensing, and variable selections in landslide susceptibility modelling. Using Valtellina Valley in Northern Italy …
Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang
Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang
Research Collection School Of Computing and Information Systems
With the rapid development of Quantum Machine Learning, quantum neural networks (QNN) have experienced great advancement in the past few years, harnessing the advantages of quantum computing to significantly speed up classical machine learning tasks. Despite their increasing popularity, the quantum neural network is quite counter-intuitive and difficult to understand, due to their unique quantum-specific layers (e.g., data encoding and measurement) in their architecture. It prevents QNN users and researchers from effectively understanding its inner workings and exploring the model training status. To fill the research gap, we propose VIOLET , a novel visual analytics approach to improve the explainability …
Poster: Profiling Event Vision Processing On Edge Devices, Ila Nitin Gokarn, Archan Misra
Poster: Profiling Event Vision Processing On Edge Devices, Ila Nitin Gokarn, Archan Misra
Research Collection School Of Computing and Information Systems
As RGB camera resolutions and frame-rates improve, their increased energy requirements make it challenging to deploy fast, efficient, and low-power applications on edge devices. Newer classes of sensors, such as the biologically inspired neuromorphic event-based camera, capture only changes in light intensity per-pixel to achieve operational superiority in sensing latency (O(μs)), energy consumption (O(mW)), high dynamic range (140dB), and task accuracy such as in object tracking, over traditional RGB camera streams. However, highly dynamic scenes can yield an event rate of up to 12MEvents/second, the processing of which could overwhelm …
Criticality Aware Canvas-Based Visual Perception At The Edge, Ila Gokarn
Criticality Aware Canvas-Based Visual Perception At The Edge, Ila Gokarn
Research Collection School Of Computing and Information Systems
Efficient and effective machine perception remains a formidable challenge in sustaining high fidelity and high throughput of perception tasks on affordable edge devices. This is especially due to the continuing increase in resolution of sensor streams (e.g., video input streams generated by 4K/8K cameras and neuromorphic event cameras that produce ≥ 10 MEvents/second) and computational complexity of Deep Neural Network (DNN) models, which overwhelms edge platforms, adversely impacting machine perception efficiency. Given the insufficiency of the available computation resources, a question then arises on whether selected regions/components of the perception task can be prioritized (and executed preferentially) to achieve highest …
Refining Chatgpt-Generated Code: Characterizing And Mitigating Code Quality Issues, Yue Liu, Thanh Le-Cong, Ratnadira Widyasari, Chakkrit Kla Tantihamthavorn, Li Li, Xuan-Bach Dinh Le, David Lo
Refining Chatgpt-Generated Code: Characterizing And Mitigating Code Quality Issues, Yue Liu, Thanh Le-Cong, Ratnadira Widyasari, Chakkrit Kla Tantihamthavorn, Li Li, Xuan-Bach Dinh Le, David Lo
Research Collection School Of Computing and Information Systems
Since its introduction in November 2022, ChatGPT has rapidly gained popularity due to its remarkable ability in language understanding and human-like responses. ChatGPT, based on GPT-3.5 architecture, has shown great promise for revolutionizing various research fields, including code generation. However, the reliability and quality of code generated by ChatGPT remain unexplored, raising concerns about potential risks associated with the widespread use of ChatGPT-driven code generation.In this article, we systematically study the quality of 4,066 ChatGPT-generated programs of code implemented in two popular programming languages, i.e., Java and Python, for 2,033 programming tasks. The goal of this work is threefold. First, …
Predicting Mild Cognitive Impairment Through Ambient Sensing And Artificial Intelligence, Ah-Hwee Tan, Weng Yan Ying, Budhitama Subagdja, Anni Huang, Shanthoshigaa D, Tony Chin-Ian Tay, Iris Rawtaer
Predicting Mild Cognitive Impairment Through Ambient Sensing And Artificial Intelligence, Ah-Hwee Tan, Weng Yan Ying, Budhitama Subagdja, Anni Huang, Shanthoshigaa D, Tony Chin-Ian Tay, Iris Rawtaer
Research Collection School Of Computing and Information Systems
This paper reports an emerging application leveraging ambient and artificial intelligence techniques for in-home sensing and cognitive health assessment. The application involves a prospective longitudinal study, wherein non-pervasive sensing devices are installed in homes of over 63 real users undergoing clinical cognitive assessment, and digital signals of the users’ activities and behaviour are transmitted to a central cloud-based data server for further processing and analysis. Based on the sensor readings, we identify a set of digital biomarkers covering four key aspects of daily living, namely physical, activity, cognitive, and sleep, and develop a suite of customized feature extraction methods for …
Locality-Aware Tail Node Embeddings On Homogeneous And Heterogeneous Networks, Zemin Liu, Yuan Fang, Wentao Zhang, Xinming Zhang, Steven C. H. Hoi
Locality-Aware Tail Node Embeddings On Homogeneous And Heterogeneous Networks, Zemin Liu, Yuan Fang, Wentao Zhang, Xinming Zhang, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
While the state-of-the-art network embedding approaches often learn high-quality embeddings for high-degree nodes with abundant structural connectivity, the quality of the embeddings for low-degree or nodes is often suboptimal due to their limited structural connectivity. While many real-world networks are long-tailed, to date little effort has been devoted to tail node embeddings. In this article, we formulate the goal of learning tail node embeddings as a problem, given the few links on each tail node. In particular, since each node resides in its own local context, we personalize the regression model for each tail node. To reduce overfitting in the …
Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang
Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang
Research Collection School Of Computing and Information Systems
In recent years, vision Transformers and MLPs have demonstrated remarkable performance in image understanding tasks. However, their inherently dense computational operators, such as self-attention and token-mixing layers, pose significant challenges when applied to spatio-temporal video data. To address this gap, we propose PosMLP-Video, a lightweight yet powerful MLP-like backbone for video recognition. Instead of dense operators, we use efficient relative positional encoding (RPE) to build pairwise token relations, leveraging small-sized parameterized relative position biases to obtain each relation score. Specifically, to enable spatio-temporal modeling, we extend the image PosMLP’s positional gating unit to temporal, spatial, and spatio-temporal variants, namely PoTGU, …
Actively Learn From Llms With Uncertainty Propagation For Generalized Category Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Bobo Li, Jing Jiang
Actively Learn From Llms With Uncertainty Propagation For Generalized Category Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Bobo Li, Jing Jiang
Research Collection School Of Computing and Information Systems
Generalized category discovery faces a key issue: the lack of supervision for new and unseen data categories. Traditional methods typically combine supervised pretraining with self-supervised learning to create models, and then employ clustering for category identification. However, these approaches tend to become overly tailored to known categories, failing to fully resolve the core issue. Hence, we propose to integrate the feedback from LLMs into an active learning paradigm. Specifically, our method innovatively employs uncertainty propagation to select data samples from high-uncertainty regions, which are then labeled using LLMs through a comparison-based prompting scheme. This not only eases the labeling task …
Mix-Initiative Response Generation With Dynamic Prefix Tuning, Yuxiang Nie, Heyan Huang, Xian-Ling Mao, Lizi Liao
Mix-Initiative Response Generation With Dynamic Prefix Tuning, Yuxiang Nie, Heyan Huang, Xian-Ling Mao, Lizi Liao
Research Collection School Of Computing and Information Systems
Mixed initiative serves as one of the key factors in controlling conversation directions. For a speaker, responding passively or leading proactively would result in rather different responses. However, most dialogue systems focus on training a holistic response generation model without any distinction among different initiatives. It leads to the cross-contamination problem, where the model confuses different initiatives and generates inappropriate responses. Moreover, obtaining plenty of human annotations for initiative labels can be expensive. To address this issue, we propose a general mix-Initiative Dynamic Prefix Tuning framework (IDPT) to decouple different initiatives from the generation model, which learns initiative-aware prefixes in …
Sgsh : Stimulate Large Language Models With Skeleton Heuristics For Knowledge Base Question Generation, Shasha Guo, Lizi Liao, Jing Zhang, Yanling Wang, Cuiping Li, Hong Chen
Sgsh : Stimulate Large Language Models With Skeleton Heuristics For Knowledge Base Question Generation, Shasha Guo, Lizi Liao, Jing Zhang, Yanling Wang, Cuiping Li, Hong Chen
Research Collection School Of Computing and Information Systems
Knowledge base question generation (KBQG) aims to generate natural language questions from a set of triplet facts extracted from KB. Existing methods have significantly boosted the performance of KBQG via pre-trained language models (PLMs) thanks to the richly endowed semantic knowledge. With the advance of pre-training techniques, large language models (LLMs) (e.g., GPT-3.5) undoubtedly possess much more semantic knowledge. Therefore, how to effectively organize and exploit the abundant knowledge for KBQG becomes the focus of our study. In this work, we propose SGSH — a simple and effective framework to Stimulate GPT-3.5 with Skeleton Heuristics to enhance KBQG. The framework …
Learning Transferable Negative Prompts For Out-Of-Distribution Detection, Tianqi Li, Guansong Pang, Xiao Bai, Wenjun Miao, Jin Zheng
Learning Transferable Negative Prompts For Out-Of-Distribution Detection, Tianqi Li, Guansong Pang, Xiao Bai, Wenjun Miao, Jin Zheng
Research Collection School Of Computing and Information Systems
Existing prompt learning methods have shown certain capabilities in Out-of-Distribution (OOD) detection, but the lack of OOD images in the target dataset in their training can lead to mismatches between OOD images and In-Distribution (ID) categories, resulting in a high false positive rate. To address this issue, we introduce a novel OOD detection method, named ‘NegPrompt’, to learn a set of negative prompts, each representing a negative connotation of a given class label, for delineating the boundaries between ID and OOD images. It learns such negative prompts with ID data only, without any reliance on external out-lier data. Further, current …
Anomaly Heterogeneity Learning For Open-Set Supervised Anomaly Detection, Jiawen Zhu, Choubo Ding, Yu Tian, Guansong Pang
Anomaly Heterogeneity Learning For Open-Set Supervised Anomaly Detection, Jiawen Zhu, Choubo Ding, Yu Tian, Guansong Pang
Research Collection School Of Computing and Information Systems
Open-set supervised anomaly detection (OSAD) - a recently emerging anomaly detection area - aims at utilizing a few samples of anomaly classes seen during training to detect unseen anomalies (i.e., samples from open-set anomaly classes), while effectively identifying the seen anomalies. Benefiting from the prior knowledge illustrated by the seen anomalies, current OSAD methods can often largely reduce false positive errors. However, these methods are trained in a closed-set setting and treat the anomaly examples as from a homogeneous distribution, rendering them less effective in generalizing to unseen anomalies that can be drawn from any distribution. This paper proposes to …
Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang
Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang
Research Collection School Of Computing and Information Systems
Current video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step …
Toward Generalist Anomaly Detection Via In-Context Residual Learning With Few-Shot Sample Prompts, Jiawen Zhu, Guansong Pang
Toward Generalist Anomaly Detection Via In-Context Residual Learning With Few-Shot Sample Prompts, Jiawen Zhu, Guansong Pang
Research Collection School Of Computing and Information Systems
This paper explores the problem of Generalist Anomaly Detection (GAD), aiming to train one single detection model that can generalize to detect anomalies in diverse datasets from different application domains without any further training on the target data. Some recent studies have showed that large pre-trained Visual-Language Models (VLMs) like CLIP have strong generalization capabilities on detecting industrial defects from various datasets, but their methods rely heavily on handcrafted text prompts about defects, making them difficult to generalize to anomalies in other applications, e.g., medical image anomalies or semantic anomalies in natural images. In this work, we propose to train …
Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He
Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He
Research Collection School Of Computing and Information Systems
Point-based interactive editing serves as an essential tool to complement the controllability of existing generative models. A concurrent work, DragDiffusion, updates the diffusion latent map in response to user inputs, causing global latent map alterations. This results in imprecise preservation of the original content and unsuccessful editing due to gradient vanishing. In contrast, we present DragNoise, offering robust and accelerated editing without retracing the latent map. The core rationale of DragNoise lies in utilizing the predicted noise output of each U-Net as a semantic editor. This approach is grounded in two critical observations: firstly, the bottleneck features of U-Net inherently …
Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
In this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which …
Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He
Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose a voxel-based optimization framework, Re VoRF, for few-shot radiance fields that strategically ad-dress the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the ab-solute color values in disoccluded areas. Consequently, we devise a bilateral geometric consistency loss that carefully navigates the trade-off between color fidelity and geometric accuracy in the context of depth consistency for uncertain regions. Moreover, we present a reliability-guided learning strategy to discern and utilize the variable quality across syn-thesized views, complemented by a reliability-aware voxel smoothing algorithm that smoothens …
D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He
D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He
Research Collection School Of Computing and Information Systems
Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these oneto-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is …