Open Access. Powered by Scholars. Published by Universities.®

2024

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 601 - 630 of 1390

Full-Text Articles in Artificial Intelligence and Robotics

Machines Of The Absurd: Leveraging Generative Ai For Creativity, Humor, And Playfulness, Tyler Sanders Jun 2024

Machines Of The Absurd: Leveraging Generative Ai For Creativity, Humor, And Playfulness, Tyler Sanders

College of Computing and Digital Media Dissertations

Machines of The Absurd is a collection of four projects exploring how generative AI can be leveraged for creativity, humor and playfulness.

1. neverOS — A node-based visual playground for interacting with large language models.

2. Other Calc — An iOS app with a calculator interface, where players can “calculate” text instead of numbers.

3. What Must Burn — An experiment where players type in text that can be dragged into a campfire to produce contextually appropriate sound effects.

4. Jazz vs Waffles — A turn-based comedy game, where players battle anything they type in.

Together, these projects make the …


Back To The Future: A Case For The Resurgence Of Approximation Theory For Enabling Data Driven “Intelligence”, Michael Dominic Ciocco Jun 2024

Back To The Future: A Case For The Resurgence Of Approximation Theory For Enabling Data Driven “Intelligence”, Michael Dominic Ciocco

Theses and Dissertations

Artificial Intelligence (AI) has exploded into mainstream consciousness with commercial investments exceeding $90 billion in the last year alone. Inasmuch as consumer-facing applications such ChatGPT offer astounding access to algorithms that were hitherto restricted to academic research labs, public focus of attention on AI has created an avalanche of misinformation. The nexus of investor-driven hype, “surprising” inaccuracies in the answers provided by AI models – now anthropomorphically labeled as “hallucinations”, and impending legislation by well-meaning and concerned governments has resulted in a crisis of confidence in the science of AI. The primary driver for AI’s recent growth is the convergence …


Perceptions And Aspirations Of Undergraduate Computer Science Students Towards Generative Ai: A Qualitative Inquiry, James Hutson, Theresa Jeevanjee Jun 2024

Perceptions And Aspirations Of Undergraduate Computer Science Students Towards Generative Ai: A Qualitative Inquiry, James Hutson, Theresa Jeevanjee

Faculty Scholarship

This article presents a comprehensive study conducted during the spring semester of 2024, aimed at exploring undergraduate computer science students’ perceptions, awareness, and understanding of generative artificial intelligence (GAI) tools within the context of their Artificial Intelligence (AI) courses. The research methodology employed qualitative techniques, including human-subject research and focus groups, to delve into students’ insights on the evolution of AI as delineated in the seminal textbook by Russell and Norvig. The study-initiated discussions on the historical development of AI, prompting students to reflect on the aspects that intrigued them the most, and to identify which historical concepts and methodologies, …


Predictive Power Of Machine Learning Models On Degree Completion Among Adult Learners, Emily Barnes, James Hutson, Karriem Perry Jun 2024

Predictive Power Of Machine Learning Models On Degree Completion Among Adult Learners, Emily Barnes, James Hutson, Karriem Perry

Faculty Scholarship

The integration of machine learning (ML) into higher education has been recognized as a transformative force for adult learners, a growing demographic facing unique educational challenges. This study evaluates the predictive power of three ML models—Random Forest, Gradient-Boosting Machine, and Decision Trees—in forecasting degree completion among this group. Utilizing a dataset from the academic years 2013-14 to 2021-22, which includes demographic and academic performance metrics, the study employs accuracy, precision, recall, and F1 score to assess the efficacy of these models. The results indicate that the Gradient-Boosting Machine model outperforms others in predicting degree completion, suggesting that ML can significantly …


Confronting Algorithms: Conscience Catching In The Criminal Trial And Beyond, Sherman J. Clark Jun 2024

Confronting Algorithms: Conscience Catching In The Criminal Trial And Beyond, Sherman J. Clark

University of Michigan Journal of Law Reform

Using the question of how to treat algorithmic evidence under the Confrontation Clause as an entry point, I argue that the use of AI in ethically salient situations presents a risk. It may cause us to avoid confronting our own responsibility. This matters because facing up to what we do, including what we delegate, can help us grow and thrive. Bearing responsibility can help us nurture vital capacities, including forms of empathy, honesty, and dignity. In the language of ethics, these are eudaimonist virtues—traits and capacities that can help us live well and fully. We should thus find ways of …


Architectural Elements Contributing To Interpretability Of Deep Neural Networks (Dnns), Emily Barnes, James Hutson Jun 2024

Architectural Elements Contributing To Interpretability Of Deep Neural Networks (Dnns), Emily Barnes, James Hutson

Faculty Scholarship

The interpretability of Deep Neural Networks (DNNs) has become a critical focus in artificial intelligence and machine learning, particularly as DNNs are increasingly used in high-stakes applications like healthcare, finance, and autonomous driving. Interpretability refers to the extent to which humans can understand the reasons behind a model's decisions, which is essential for trust, accountability, and transparency. However, the complexity and depth of DNN architectures often compromise interpretability as these models function as "black boxes." This article reviews key architectural elements of DNNs that affect their interpretability, aiming to guide the design of more transparent and trustworthy models. The primary …


Navigating The Complexities Of Ai: The Critical Role Of Interpretability And Explainability In Ensuring Transparency And Trust, Emily Barnes, James Hutson Jun 2024

Navigating The Complexities Of Ai: The Critical Role Of Interpretability And Explainability In Ensuring Transparency And Trust, Emily Barnes, James Hutson

Faculty Scholarship

The interpretability and explainability of deep neural networks (DNNs) are paramount in artificial intelligence (AI), especially when applied to high-stakes fields such as healthcare, finance, and autonomous driving. The need for this study arises from the growing integration of AI into critical areas where transparency, trust, and ethical decision-making are essential. This paper explores the impact of architectural design choices on DNN interpretability, focusing on how different architectural elements like layer types, network depth, connectivity patterns, and attention mechanisms affect model transparency. Methodologically, the study employs a comprehensive review of case studies and experimental results to analyze the balance between …


Combinatorial Creativity: Knowledge Graphs And Idea Generation In Crowdsourcing Innovation, Zhi Wei Vincent Mack Jun 2024

Combinatorial Creativity: Knowledge Graphs And Idea Generation In Crowdsourcing Innovation, Zhi Wei Vincent Mack

Dissertations and Theses Collection (Open Access)

This dissertation explores the dynamic interplay between combinatorial creativity and technology-driven innovation within various knowledge-intensive fields. It critically examines the role of combinatorial creativity in generating groundbreaking innovations by amalgamating existing ideas and technologies. This research incorporates a detailed examination of how knowledge, whether tacit or explicit, can be transformed into actionable data to foster innovation in crowdsourcing contexts. Chapter 2 provides an overview of the relevant literature on how Artificial Intelligence and Knowledge Management Systems can support combinatorial creativity. The study further delves into the transformative impact of knowledge management systems, particularly focusing on crowdsourcing platforms that leverage collective …


Evaluating Methods For Assessing Interpretability Of Deep Neural Networks (Dnns), Emily Barnes, James Hutson Jun 2024

Evaluating Methods For Assessing Interpretability Of Deep Neural Networks (Dnns), Emily Barnes, James Hutson

Faculty Scholarship

The interpretability of deep neural networks (DNNs) is a critical focus in artificial intelligence (AI) and machine learning (ML), particularly as these models are increasingly deployed in high-stakes applications such as healthcare, finance, and autonomous systems. In the context of these technologies, interpretability refers to the extent to which a human can understand the cause of a decision made by a model. This article evaluates various methods for assessing the interpretability of DNNs, recognizing the significant challenges posed by their complex and opaque nature. The review encompasses both quantitative metrics and qualitative evaluations, aiming to identify effective strategies that enhance …


Design And Implementation Of A Vision-Based Deep-Learning Protocol For Kinematic Feature Extraction With Application To Stroke Rehabilitation, Juan Diego Luna Inga Jun 2024

Design And Implementation Of A Vision-Based Deep-Learning Protocol For Kinematic Feature Extraction With Application To Stroke Rehabilitation, Juan Diego Luna Inga

Master's Theses

Stroke is a leading cause of long-term disability, affecting thousands of individuals annually and significantly impairing their mobility, independence, and quality of life. Traditional methods for assessing motor impairments are often costly and invasive, creating substantial barriers to effective rehabilitation. This thesis explores the use of DeepLabCut (DLC), a deep-learning-based pose estimation tool, to extract clinically meaningful kinematic features from video data of stroke survivors with upper-extremity (UE) impairments.

To conduct this investigation, a specialized protocol was developed to tailor DLC for analyzing movements characteristic of UE impairments in stroke survivors. This protocol was validated through comparative analysis using peak …


Navigating The Ethical Terrain Of Ai In Higher Education: Strategies For Mitigating Bias And Promoting Fairness, Emily Barnes, James Hutson Jun 2024

Navigating The Ethical Terrain Of Ai In Higher Education: Strategies For Mitigating Bias And Promoting Fairness, Emily Barnes, James Hutson

Faculty Scholarship

Artificial intelligence (AI) and machine learning (ML) are transforming higher education by enhancing personalized learning and academic support, yet they pose significant ethical challenges, particularly in terms of inherent biases. This review critically examines the integration of AI in higher education, underscoring the dual aspects of its potential to innovate educational paradigms and the essential need to address ethical implications to avoid perpetuating existing inequalities. The researchers employed a methodological approach that analyzed case studies and literature as primary data collection methods, focusing on strategies to mitigate biases through technical solutions, diverse datasets, and strict adherence to ethical guidelines. Their …


Accessible Real-Time Eye-Gaze Tracking For Neurocognitive Health Assessments, A Multimodal Web-Based Approach, Daniel C. Tisdale Jun 2024

Accessible Real-Time Eye-Gaze Tracking For Neurocognitive Health Assessments, A Multimodal Web-Based Approach, Daniel C. Tisdale

Master's Theses

We introduce a novel integration of real-time, predictive eye-gaze tracking models into a multimodal dialogue system tailored for remote health assessments. This system is designed to be highly accessible requiring only a conventional webcam for video input along with minimal cursor interaction and utilizes engaging gaze-based tasks that can be performed directly in a web browser. We have crafted dynamic subsystems that capture high-quality data efficiently and maintain quality through instances of user attrition and incomplete calls. Additionally, these subsystems are designed with the foresight to allow for future re-analysis using improved predictive models, as well as enable the creation …


Strategic Integration Of Ai In Higher Education And Industry: The Ai8-Point Model, Emily Barnes, James Hutson Jun 2024

Strategic Integration Of Ai In Higher Education And Industry: The Ai8-Point Model, Emily Barnes, James Hutson

Faculty Scholarship

The AI8-Point Model, derived from extensive experience in technology, AI, and higher education administration, addresses the critical need for cost-effective, high-impact strategies tailored to higher education. Despite the transformative potential of AI in enhancing student engagement, optimizing processes, and improving educational outcomes, institutions often struggle with practical implementation. The AI8-Point Model fills this gap by offering strategies that balance cost and impact. Visualized as a circle divided into four quadrants, the model encompasses phases of student engagement and institutional interaction: pre-enrollment beyond institutional control, pre-enrollment within institutional control, post-enrollment within institutional control, and post-enrollment beyond institutional control. Each quadrant contains …


D-Hacking, Emily Black, Talia B. Gillis, Zara Hall Jun 2024

D-Hacking, Emily Black, Talia B. Gillis, Zara Hall

Faculty Scholarship

Recent regulatory efforts, including Executive Order 14110 and the AI Bill of Rights, have focused on mitigating discrimination in AI systems through novel and traditional application of anti-discrimination laws. While these initiatives rightly emphasize fairness testing and mitigation, we argue that they pay insufficient attention to robust bias measurement and mitigation — and that without doing so, the frameworks cannot effectively achieve the goal of reducing discrimination in deployed AI models. This oversight is particularly concerning given the instability and brittleness of current algorithmic bias mitigation and fairness optimization methods, as highlighted by growing evidence in the algorithmic fairness literature. …


Assessing Job Vulnerability And Employment Growth In The Era Of Large Language Models (Llms), Prudence P. Brou Jun 2024

Assessing Job Vulnerability And Employment Growth In The Era Of Large Language Models (Llms), Prudence P. Brou

Dissertations, Theses, and Capstone Projects

This paper explores the impact of Large Language Models (LLMs) and artificial intelligence (AI) on white-collar occupations in the context of job vulnerability and employment growth. Utilizing the Kaggle dataset "Occupation Salary and Likelihood of Automation," the study employs a data-driven approach to analyze trends across states. Through interactive data visualization, the project aims to provide actionable insights for affected workers, businesses, and policymakers navigating the changing dynamics of the workforce amidst technological advancements.


Present Case Studies Highlighting Practical Implications Of Architectural Design Choices, Emily Barnes, James Hutson Jun 2024

Present Case Studies Highlighting Practical Implications Of Architectural Design Choices, Emily Barnes, James Hutson

Faculty Scholarship

The interpretability of deep neural networks (DNNs) has become a crucial focus within artificial intelligence and machine learning, particularly as these models are increasingly used in high-stakes applications such as healthcare, finance, and autonomous driving. This article explores the impact of architectural design choices on the interpretability of DNNs, emphasizing the importance of transparency, trust, and accountability in AI systems. By presenting case studies and experimental results, the article highlights how different architectural elements—such as layer types, network depth, connectivity patterns, and attention mechanisms—affect model interpretability and performance. The discussion is structured into three main sections: real-world applications, architectural trade-offs, …


Does Generative Ai Facilitate Investor Trading? Evidence From Chatgpt Outages, Qiang Cheng, Pengkai Lin, Yue Zhao Jun 2024

Does Generative Ai Facilitate Investor Trading? Evidence From Chatgpt Outages, Qiang Cheng, Pengkai Lin, Yue Zhao

Research Collection School Of Accountancy

In this paper, we use ChatGPT outages to investigate whether investors rely on generative artificial intelligence (GAI) to perform trading-related tasks and the associated impact on stock price informativeness. We first document a significant decline in stock trading volume during ChatGPT outages and find that the effect is stronger for firms with corporate news released immediately before or during the outages. We further document similar declines in the short-run price impact, return variance, and bid-ask spreads, consistent with a reduction in informed trading during the outage periods. Lastly, we use trading volume changes during outages to construct a firm-level measure …


Towards Faster Inference Of Transformers: Strategies For Accelerating Decoding Processes, Cunxiao Du Jun 2024

Towards Faster Inference Of Transformers: Strategies For Accelerating Decoding Processes, Cunxiao Du

Dissertations and Theses Collection (Open Access)

This thesis delves into the acceleration and optimization of Transformer inference, a subject of increasing importance with the emergence of Large Language Models (LLMs). The study primarily addresses the challenges posed by two inherent properties of Transformers during inference: the quadratic complexity of the attention mechanism and the sequential nature of autoregressive inference. The research is structured into three main parts. The first part enhances the learning capabilities of non-autoregressive Transformers, achieving a remarkable 15.0x acceleration on machine translation tasks. The following section focuses on lossless acceleration through speculative decoding, where the proposed algorithm, Glide with CAPE, is shown to …


Locality-Aware Tail Node Embeddings On Homogeneous And Heterogeneous Networks, Zemin Liu, Yuan Fang, Wentao Zhang, Xinming Zhang, Steven C. H. Hoi Jun 2024

Locality-Aware Tail Node Embeddings On Homogeneous And Heterogeneous Networks, Zemin Liu, Yuan Fang, Wentao Zhang, Xinming Zhang, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

While the state-of-the-art network embedding approaches often learn high-quality embeddings for high-degree nodes with abundant structural connectivity, the quality of the embeddings for low-degree or nodes is often suboptimal due to their limited structural connectivity. While many real-world networks are long-tailed, to date little effort has been devoted to tail node embeddings. In this article, we formulate the goal of learning tail node embeddings as a problem, given the few links on each tail node. In particular, since each node resides in its own local context, we personalize the regression model for each tail node. To reduce overfitting in the …


Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang Jun 2024

Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang

Research Collection School Of Computing and Information Systems

In recent years, vision Transformers and MLPs have demonstrated remarkable performance in image understanding tasks. However, their inherently dense computational operators, such as self-attention and token-mixing layers, pose significant challenges when applied to spatio-temporal video data. To address this gap, we propose PosMLP-Video, a lightweight yet powerful MLP-like backbone for video recognition. Instead of dense operators, we use efficient relative positional encoding (RPE) to build pairwise token relations, leveraging small-sized parameterized relative position biases to obtain each relation score. Specifically, to enable spatio-temporal modeling, we extend the image PosMLP’s positional gating unit to temporal, spatial, and spatio-temporal variants, namely PoTGU, …


Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang Jun 2024

Violet: Visual Analytics For Explainable Quantum Neural Networks, Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Robert Griffin, Xiaolin Wen, Yanna Lin, Yong Wang

Research Collection School Of Computing and Information Systems

With the rapid development of Quantum Machine Learning, quantum neural networks (QNN) have experienced great advancement in the past few years, harnessing the advantages of quantum computing to significantly speed up classical machine learning tasks. Despite their increasing popularity, the quantum neural network is quite counter-intuitive and difficult to understand, due to their unique quantum-specific layers (e.g., data encoding and measurement) in their architecture. It prevents QNN users and researchers from effectively understanding its inner workings and exploring the model training status. To fill the research gap, we propose VIOLET , a novel visual analytics approach to improve the explainability …


Poster: Profiling Event Vision Processing On Edge Devices, Ila Nitin Gokarn, Archan Misra Jun 2024

Poster: Profiling Event Vision Processing On Edge Devices, Ila Nitin Gokarn, Archan Misra

Research Collection School Of Computing and Information Systems

As RGB camera resolutions and frame-rates improve, their increased energy requirements make it challenging to deploy fast, efficient, and low-power applications on edge devices. Newer classes of sensors, such as the biologically inspired neuromorphic event-based camera, capture only changes in light intensity per-pixel to achieve operational superiority in sensing latency (O(μs)), energy consumption (O(mW)), high dynamic range (140dB), and task accuracy such as in object tracking, over traditional RGB camera streams. However, highly dynamic scenes can yield an event rate of up to 12MEvents/second, the processing of which could overwhelm …


Criticality Aware Canvas-Based Visual Perception At The Edge, Ila Gokarn Jun 2024

Criticality Aware Canvas-Based Visual Perception At The Edge, Ila Gokarn

Research Collection School Of Computing and Information Systems

Efficient and effective machine perception remains a formidable challenge in sustaining high fidelity and high throughput of perception tasks on affordable edge devices. This is especially due to the continuing increase in resolution of sensor streams (e.g., video input streams generated by 4K/8K cameras and neuromorphic event cameras that produce ≥ 10 MEvents/second) and computational complexity of Deep Neural Network (DNN) models, which overwhelms edge platforms, adversely impacting machine perception efficiency. Given the insufficiency of the available computation resources, a question then arises on whether selected regions/components of the perception task can be prioritized (and executed preferentially) to achieve highest …


Refining Chatgpt-Generated Code: Characterizing And Mitigating Code Quality Issues, Yue Liu, Thanh Le-Cong, Ratnadira Widyasari, Chakkrit Kla Tantihamthavorn, Li Li, Xuan-Bach Dinh Le, David Lo Jun 2024

Refining Chatgpt-Generated Code: Characterizing And Mitigating Code Quality Issues, Yue Liu, Thanh Le-Cong, Ratnadira Widyasari, Chakkrit Kla Tantihamthavorn, Li Li, Xuan-Bach Dinh Le, David Lo

Research Collection School Of Computing and Information Systems

Since its introduction in November 2022, ChatGPT has rapidly gained popularity due to its remarkable ability in language understanding and human-like responses. ChatGPT, based on GPT-3.5 architecture, has shown great promise for revolutionizing various research fields, including code generation. However, the reliability and quality of code generated by ChatGPT remain unexplored, raising concerns about potential risks associated with the widespread use of ChatGPT-driven code generation.In this article, we systematically study the quality of 4,066 ChatGPT-generated programs of code implemented in two popular programming languages, i.e., Java and Python, for 2,033 programming tasks. The goal of this work is threefold. First, …


Predicting Mild Cognitive Impairment Through Ambient Sensing And Artificial Intelligence, Ah-Hwee Tan, Weng Yan Ying, Budhitama Subagdja, Anni Huang, Shanthoshigaa D, Tony Chin-Ian Tay, Iris Rawtaer Jun 2024

Predicting Mild Cognitive Impairment Through Ambient Sensing And Artificial Intelligence, Ah-Hwee Tan, Weng Yan Ying, Budhitama Subagdja, Anni Huang, Shanthoshigaa D, Tony Chin-Ian Tay, Iris Rawtaer

Research Collection School Of Computing and Information Systems

This paper reports an emerging application leveraging ambient and artificial intelligence techniques for in-home sensing and cognitive health assessment. The application involves a prospective longitudinal study, wherein non-pervasive sensing devices are installed in homes of over 63 real users undergoing clinical cognitive assessment, and digital signals of the users’ activities and behaviour are transmitted to a central cloud-based data server for further processing and analysis. Based on the sensor readings, we identify a set of digital biomarkers covering four key aspects of daily living, namely physical, activity, cognitive, and sleep, and develop a suite of customized feature extraction methods for …


Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang Jun 2024

Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

Current video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step …


Anomaly Heterogeneity Learning For Open-Set Supervised Anomaly Detection, Jiawen Zhu, Choubo Ding, Yu Tian, Guansong Pang Jun 2024

Anomaly Heterogeneity Learning For Open-Set Supervised Anomaly Detection, Jiawen Zhu, Choubo Ding, Yu Tian, Guansong Pang

Research Collection School Of Computing and Information Systems

Open-set supervised anomaly detection (OSAD) - a recently emerging anomaly detection area - aims at utilizing a few samples of anomaly classes seen during training to detect unseen anomalies (i.e., samples from open-set anomaly classes), while effectively identifying the seen anomalies. Benefiting from the prior knowledge illustrated by the seen anomalies, current OSAD methods can often largely reduce false positive errors. However, these methods are trained in a closed-set setting and treat the anomaly examples as from a homogeneous distribution, rendering them less effective in generalizing to unseen anomalies that can be drawn from any distribution. This paper proposes to …


Learning Transferable Negative Prompts For Out-Of-Distribution Detection, Tianqi Li, Guansong Pang, Xiao Bai, Wenjun Miao, Jin Zheng Jun 2024

Learning Transferable Negative Prompts For Out-Of-Distribution Detection, Tianqi Li, Guansong Pang, Xiao Bai, Wenjun Miao, Jin Zheng

Research Collection School Of Computing and Information Systems

Existing prompt learning methods have shown certain capabilities in Out-of-Distribution (OOD) detection, but the lack of OOD images in the target dataset in their training can lead to mismatches between OOD images and In-Distribution (ID) categories, resulting in a high false positive rate. To address this issue, we introduce a novel OOD detection method, named ‘NegPrompt’, to learn a set of negative prompts, each representing a negative connotation of a given class label, for delineating the boundaries between ID and OOD images. It learns such negative prompts with ID data only, without any reliance on external out-lier data. Further, current …


Toward Generalist Anomaly Detection Via In-Context Residual Learning With Few-Shot Sample Prompts, Jiawen Zhu, Guansong Pang Jun 2024

Toward Generalist Anomaly Detection Via In-Context Residual Learning With Few-Shot Sample Prompts, Jiawen Zhu, Guansong Pang

Research Collection School Of Computing and Information Systems

This paper explores the problem of Generalist Anomaly Detection (GAD), aiming to train one single detection model that can generalize to detect anomalies in diverse datasets from different application domains without any further training on the target data. Some recent studies have showed that large pre-trained Visual-Language Models (VLMs) like CLIP have strong generalization capabilities on detecting industrial defects from various datasets, but their methods rely heavily on handcrafted text prompts about defects, making them difficult to generalize to anomalies in other applications, e.g., medical image anomalies or semantic anomalies in natural images. In this work, we propose to train …


Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He Jun 2024

Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He

Research Collection School Of Computing and Information Systems

Point-based interactive editing serves as an essential tool to complement the controllability of existing generative models. A concurrent work, DragDiffusion, updates the diffusion latent map in response to user inputs, causing global latent map alterations. This results in imprecise preservation of the original content and unsuccessful editing due to gradient vanishing. In contrast, we present DragNoise, offering robust and accelerated editing without retracing the latent map. The core rationale of DragNoise lies in utilizing the predicted noise output of each U-Net as a semantic editor. This approach is grounded in two critical observations: firstly, the bottleneck features of U-Net inherently …