Open Access. Powered by Scholars. Published by Universities.®

Graphics and Human Computer Interfaces

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 151 - 180 of 414

Full-Text Articles in Artificial Intelligence and Robotics

Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang Jun 2024

Posmlp-Video: Spatial And Temporal Relative Position Encoding For Efficient Video Recognition, Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Xiangnan He, Meng Wang

Research Collection School Of Computing and Information Systems

In recent years, vision Transformers and MLPs have demonstrated remarkable performance in image understanding tasks. However, their inherently dense computational operators, such as self-attention and token-mixing layers, pose significant challenges when applied to spatio-temporal video data. To address this gap, we propose PosMLP-Video, a lightweight yet powerful MLP-like backbone for video recognition. Instead of dense operators, we use efficient relative positional encoding (RPE) to build pairwise token relations, leveraging small-sized parameterized relative position biases to obtain each relation score. Specifically, to enable spatio-temporal modeling, we extend the image PosMLP’s positional gating unit to temporal, spatial, and spatio-temporal variants, namely PoTGU, …


Impact Of Similarities In Gender And Physical Appearance Between User And Embodied Conversational Agents On Trustworthiness, Empathy, And Service Evaluation, Sookyoung Park Jun 2024

Impact Of Similarities In Gender And Physical Appearance Between User And Embodied Conversational Agents On Trustworthiness, Empathy, And Service Evaluation, Sookyoung Park

Dartmouth College Master’s Theses

Embodied conversational agents (ECAs) have significantly enhanced human-machine interactions and show considerable potential in various industries such as customer service, education, healthcare, entertainment, and finance [1, 2]. This study explores the impact of similarities in gender and physical appearance between ECAs and users on the perceptions of trustworthiness, empathy, and service evaluation within the context of counselor ECAs. We conducted a within-subject experiment (n=50), using a 2x2 factorial arrangement, that varied the gender and the physical appearance of four distinct AI avatars. Participants interacted with each avatar, completing a post-experiment survey and participating in semi-structured interviews. Our findings indicate that …


Community Discovery Over Attributed Graphs, Yudong Niu Jun 2024

Community Discovery Over Attributed Graphs, Yudong Niu

Dissertations and Theses Collection (Open Access)

Community discovery, as a fundamental problem in graph mining, finds applications in various domains such as biological analysis, system optimization and fraud detection. Although many efforts have been made to address community discovery based on graph topology, few works have been devoted to community discovery over attributed graphs, where graphs are equipped with attribute information such as node and edge types. Thus, this thesis is devoted to designing innovative solutions that can utilize the attribute information together with graph topology for community discovery. In particular, we study novel problems with efficient algorithms for both homogeneous and heterogeneous attributed graphs and …


Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang Jun 2024

Open-Vocabulary Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Yujia Sun, Jing Liu, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

Current video anomaly detection (VAD) approaches with weak supervisions are inherently limited to a closed-set setting and may struggle in open-world applications where there can be anomaly categories in the test data unseen during training. A few recent studies attempt to tackle a more realistic setting, open-set VAD, which aims to de-tect unseen anomalies given seen anomalies and normal videos. However, such a setting focuses on predicting frame anomaly scores, having no ability to recognize the specific categories of anomalies, despite the fact that this ability is essential for building more informed video surveillance systems. This paper takes a step …


Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He Jun 2024

Drag Your Noise: Interactive Point-Based Editing Via Diffusion Semantic Propagation, Haofeng Liu, Chenshu Xu, Yifei Yang, Lihua Zeng, Shengfeng He

Research Collection School Of Computing and Information Systems

Point-based interactive editing serves as an essential tool to complement the controllability of existing generative models. A concurrent work, DragDiffusion, updates the diffusion latent map in response to user inputs, causing global latent map alterations. This results in imprecise preservation of the original content and unsuccessful editing due to gradient vanishing. In contrast, we present DragNoise, offering robust and accelerated editing without retracing the latent map. The core rationale of DragNoise lies in utilizing the predicted noise output of each U-Net as a semantic editor. This approach is grounded in two critical observations: firstly, the bottleneck features of U-Net inherently …


Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He Jun 2024

Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He

Research Collection School Of Computing and Information Systems

In this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which …


Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He Jun 2024

Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He

Research Collection School Of Computing and Information Systems

We propose a voxel-based optimization framework, Re VoRF, for few-shot radiance fields that strategically ad-dress the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the ab-solute color values in disoccluded areas. Consequently, we devise a bilateral geometric consistency loss that carefully navigates the trade-off between color fidelity and geometric accuracy in the context of depth consistency for uncertain regions. Moreover, we present a reliability-guided learning strategy to discern and utilize the variable quality across syn-thesized views, complemented by a reliability-aware voxel smoothing algorithm that smoothens …


D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He Jun 2024

D3still : Decoupled Differential Distillation For Asymmetric Image Retrieval, Yi Xie, Yihong Lin, Wenjie Cai, Xuemiao Xu, Huaidong Zhang, Yong Du, Shengfeng He

Research Collection School Of Computing and Information Systems

Existing methods for asymmetric image retrieval employ a rigid pairwise similarity constraint between the query network and the larger gallery network. However, these oneto-one constraint approaches often fail to maintain retrieval order consistency, especially when the query network has limited representational capacity. To overcome this problem, we introduce the Decoupled Differential Distillation (D3still) framework. This framework shifts from absolute one-to-one supervision to optimizing the relational differences in pairwise similarities produced by the query and gallery networks, thereby preserving a consistent retrieval order across both networks. Our method involves computing a pairwise similarity differential matrix within the gallery domain, which is …


Rethinking Multi-View Representation Learning Via Distilled Disentangling, Guanzhou Ke, Bo Wang, Xiaoli Wang, Shengfeng He Jun 2024

Rethinking Multi-View Representation Learning Via Distilled Disentangling, Guanzhou Ke, Bo Wang, Xiaoli Wang, Shengfeng He

Research Collection School Of Computing and Information Systems

Multi-view representation learning aims to derive robust representations that are both view-consistent and view-specific from diverse data sources. This paper presents an in-depth analysis of existing approaches in this domain, highlighting a commonly overlooked aspect: the redundancy between view-consistent and view-specific representations. To this end, we propose an innovative framework for multi-view representation learning, which incorporates a technique we term 'distilled disentangling'. Our method introduces the concept of masked cross-view prediction, enabling the extraction of compact, high-quality view-consistent representations from various sources without incurring extra computational overhead. Additionally, we develop a distilled disentangling module that efficiently filters out consistency-related information …


More Human-Likeness, Less Self-Disclosure? Avatars' Form Realism And Job Applicants' Self-Disclosure In Ai Interviews, Yamin Xu, Keng Siau, Fiona Fui-Hoon Nah Jun 2024

More Human-Likeness, Less Self-Disclosure? Avatars' Form Realism And Job Applicants' Self-Disclosure In Ai Interviews, Yamin Xu, Keng Siau, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

The rise of AI in recruitment promises to revolutionize how organizations evaluate job candidates. The quality of AI evaluations is determined by the input data, which depends on job applicants' self-disclosure. However, little is known about how the design elements of AI interview systems, particularly avatar interviewers, influence job applicants' self-disclosure during these interactions. This study aims to address this gap by specifically focusing on how the form realism of avatar interviewers affects job applicants' self-disclosure through their perceptions. In addition, the study will examine the effects of job type as a moderator. Drawing on the Stimulus-Organism-Response (S-O-R) model, this …


Companionship, Romance, And Self-Perception With Conversational Chatbots, Jonathan Windsor May 2024

Companionship, Romance, And Self-Perception With Conversational Chatbots, Jonathan Windsor

Departmental Honors & Graduate Capstone Projects

Serving as a metaphorical gateway transcending the communicative barriers of physical relationships in interpersonal dialogues, artificial imators of human behavior and speech, also known as conversational chatbots; a simulation of human knowledge and existence in a bi-directional conversation, functions as a rhetor of expression. Spanning from contexts of professional to romantic, I serve to dissect and critically analyze the nuances of human-machine relationships based on pre-established literature, inviting ethical considerations and biases in their design and marketing. Corporate influences spark pre-established servitude-esque relationships with conversational agents. Professional applications, both task-oriented and emotionally based alike, paint a mixed picture of …


Diffusion-Based Negative Sampling On Graphs For Link Prediction, Yuan Fang, Yuan Fang May 2024

Diffusion-Based Negative Sampling On Graphs For Link Prediction, Yuan Fang, Yuan Fang

Research Collection School Of Computing and Information Systems

Link prediction is a fundamental task for graph analysis with important applications on the Web, such as social network analysis and recommendation systems, etc. Modern graph link prediction methods often employ a contrastive approach to learn robust node representations, where negative sampling is pivotal. Typical negative sampling methods aim to retrieve hard examples based on either predefined heuristics or automatic adversarial approaches, which might be inflexible or difficult to control. Furthermore, in the context of link prediction, most previous methods sample negative nodes from existing substructures of the graph, missing out on potentially more optimal samples in the latent space. …


Next-Generation Crop Monitoring Technologies: Case Studies About Edge Image Processing For Crop Monitoring And Soil Water Property Modeling Via Above-Ground Sensors, Nipuna Chamara May 2024

Next-Generation Crop Monitoring Technologies: Case Studies About Edge Image Processing For Crop Monitoring And Soil Water Property Modeling Via Above-Ground Sensors, Nipuna Chamara

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

Artificial Intelligence (AI) has advanced rapidly in the past two decades. Internet of Things (IoT) technology has advanced rapidly during the last decade. Merging these two technologies has immense potential in several industries, including agriculture.

We have identified several research gaps in utilizing IoT technology in agriculture. One problem was the digital divide between rural, unconnected, or limited connected areas and urban areas for utilizing images for decision-making, which has advanced with the growth of AI. Another area for improvement was the farmers' demotivation to use in-situ soil moisture sensors for irrigation decision-making due to inherited installation difficulties. As Nebraska …


Automated Cinematographer For Vr Viewing Experiences, Zihan Wu May 2024

Automated Cinematographer For Vr Viewing Experiences, Zihan Wu

Dartmouth College Master’s Theses

As the virtual reality (VR) industry continues to evolve, the question of how to effectively capture VR experiences for an audience remains a challenge. The predominant method of showcasing VR applications through first-person recordings lacks cinematic interest, failing to capture other viewpoints and the essence of the moment. Meanwhile, manually setting up cameras and editing videos requires technical expertise on behalf of the user. In this paper, we propose the use of machine learning (ML) to automatically select the most compelling predefined viewpoint in a VR environment, at any given moment. Our models, trained on actor motion and voice volume, …


Vr Circuit Simulation With Advanced Visualization For Enhancing Comprehension In Electrical Engineering, Elliott Wolbach May 2024

Vr Circuit Simulation With Advanced Visualization For Enhancing Comprehension In Electrical Engineering, Elliott Wolbach

Department of Electrical and Computer Engineering: Dissertations, Theses, and Student Research

As technology advances, the field of electrical and computer engineering continuously demands innovative tools and methodologies to facilitate effective learning and comprehension of fundamental concepts. Through a comprehensive literature review, it was discovered that there was a gap in the current research on using VR technology to effectively visualize and comprehend non-observable electrical characteristics of electronic circuits. This thesis explores the integration of Virtual Reality (VR) technology and real-time electronic circuit simulation with enhanced visualization of non-observable concepts such as voltage distribution and current flow within these circuits. The primary objective is to develop an immersive educational platform that makes …


Privacy Protection In Mobile Photography With Face Cloaking, Rithyka Heng May 2024

Privacy Protection In Mobile Photography With Face Cloaking, Rithyka Heng

Computer Science and Computer Engineering Undergraduate Honors Theses

In a world of increasing connectivity, privacy is becoming ever-more difficult to maintain. People have little control over the capture of their image while in public and have even less control over the online sharing or posting of their image. This leaves many people vulnerable to being tracked or profiled via their image’s presence in other people’s photos. This thesis implements and evaluates an approach to privacy protection that involves the photographers protecting the privacy of bystanders. Because most photographs are now being taken by smartphones, a mobile application is decidedly the technology that would best achieve widespread adoption and …


3-D Reconstruction For Underwater Robots With A Monocular Camera And Lights, Monika Roznere May 2024

3-D Reconstruction For Underwater Robots With A Monocular Camera And Lights, Monika Roznere

Dartmouth College Ph.D Dissertations

Before a robot can act, it must perceive its environment. Though, this is not a simple task when considering the challenges in underwater domains -- poor visibility conditions, limited sensor configurations, and lack of readily accessible localization. Underwater robots have, nevertheless, improved dramatically with more extensive sensor and navigation equipment. Robot and sensor use have enabled us to explore all reaches of our oceans. On the other hand, these same robots are not easily accessible or transferable to many practical tasks, including fishery management, infrastructure maintenance, disaster response, site conservation, and ecological surveys. There is a growing need for robots …


An Empirical Study On The Efficacy Of Llm-Powered Chatbots In Basic Information Retrieval Tasks, Naja Faysal May 2024

An Empirical Study On The Efficacy Of Llm-Powered Chatbots In Basic Information Retrieval Tasks, Naja Faysal

Electronic Theses, Projects, and Dissertations

The rise of conversational user interfaces (CUIs) powered by large language models (LLMs) is transforming human-computer interaction. This study evaluates the efficacy of LLM-powered chatbots, trained on website data, compared to browsing websites for finding information about organizations across diverse sectors. A within-subjects experiment with 165 participants was conducted, involving similar information retrieval (IR) tasks using both websites (GUIs) and chatbots (CUIs). The research questions are: (Q1) Which interface helps users find information faster: LLM chatbots or websites? (Q2) Which interface helps users find more accurate information: LLM chatbots or websites?. The findings are: (Q1) Participants found information significantly faster …


Learning Nighttime Semantic Segmentation The Hard Way, Wenxi Liu, Jiaxin Cai, Qi Li, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu May 2024

Learning Nighttime Semantic Segmentation The Hard Way, Wenxi Liu, Jiaxin Cai, Qi Li, Chenyang Liao, Jingjing Cao, Shengfeng He, Yuanlong Yu

Research Collection School Of Computing and Information Systems

Nighttime semantic segmentation is an important but challenging research problem for autonomous driving. The major challenges lie in the small objects or regions from the under-/over-exposed areas or suffer from motion blur caused by the camera deployed on moving vehicles. To resolve this, we propose a novel hard- class-aware module that bridges the main network for full-class segmentation and the hard-class network for segmenting aforementioned hard-class objects. In specific, it exploits the shared focus of hard-class objects from the dual-stream network, enabling the contextual information flow to guide the model to concentrate on the pixels that are hard to classify. …


Fashionregen: Llm‑Empowered Fashion Report Generation, Yujuan Ding, Yunshan Ma, Wenqi Fan, Yige Yao, Tat‑Seng Chua, Qing Li May 2024

Fashionregen: Llm‑Empowered Fashion Report Generation, Yujuan Ding, Yunshan Ma, Wenqi Fan, Yige Yao, Tat‑Seng Chua, Qing Li

Research Collection School Of Computing and Information Systems

Fashion analysis refers to the process of examining and evaluating trends, styles, and elements within the fashion industry to understand and interpret its current state, generating fashion reports. It is traditionally performed by fashion professionals based on their expertise and experience, which requires high labour cost and may also produce biased results for relying heavily on a small group of people. In this paper, to tackle the Fashion Report Generation (FashionReGen) task, we propose an intelligent Fashion Analyzing and Reporting system based the advanced Large Language Models (LLMs), debbed as GPT-FAR. Specifically, it tries to deliver FashionReGen based on effective …


Immersive Japanese Language Learning Web Application Using Spaced Repetition, Active Recall, And An Artificial Intelligent Conversational Chat Agent Both In Voice And In Text, Marc Butler Apr 2024

Immersive Japanese Language Learning Web Application Using Spaced Repetition, Active Recall, And An Artificial Intelligent Conversational Chat Agent Both In Voice And In Text, Marc Butler

MS in Computer Science Project Reports

In the last two decades various human language learning applications, spaced repetition software, online dictionaries, and artificial intelligent chat agents have been developed. However, there is no solution to cohesively combine these technologies into a comprehensive language learning application including skills such as speaking, typing, listening, and reading. Our contribution is to provide an immersive language learning web application to the end user which combines spaced repetition, a study technique used to review information at systematic intervals, and active recall, the process of purposely retrieving information from memory during a review session, with an artificial intelligent conversational chat agent both …


Visualizing Routes With Ai-Discovered Street-View Patterns, Tsung Heng Wu, Md Amiruzzaman, Ye Zhao, Deepshikha Bhati, Jing Yang Apr 2024

Visualizing Routes With Ai-Discovered Street-View Patterns, Tsung Heng Wu, Md Amiruzzaman, Ye Zhao, Deepshikha Bhati, Jing Yang

Computer Science Faculty Publications

Street-level visual appearances play an important role in studying social systems, such as understanding the built environment, driving routes, and associated social and economic factors. It has not been integrated into a typical geographical visualization interface (e.g., map services) for planning driving routes. In this article, we study this new visualization task with several new contributions. First, we experiment with a set of AI techniques and propose a solution of using semantic latent vectors for quantifying visual appearance features. Second, we calculate image similarities among a large set of street-view images and then discover spatial imagery patterns. Third, we integrate …


Hop‑Based Heterogeneous Graph Transformer, Zixuan Yang, Xiao Wang, Yanhua Yu, Yuling Wang, Kangkang Lu, Zirui Guo, Xiting Qin, Yunshan Ma, Tat‑Seng Chua Apr 2024

Hop‑Based Heterogeneous Graph Transformer, Zixuan Yang, Xiao Wang, Yanhua Yu, Yuling Wang, Kangkang Lu, Zirui Guo, Xiting Qin, Yunshan Ma, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

The Graph Transformer (GT) has shown significant ability in processing graph-structured data, addressing limitations in graph neural networks, such as over-smoothing and over-squashing. However, the implementation of GT in real-world heterogeneous graphs (HGs) with complex topology continues to present numerous challenges. Firstly, a challenge arises in designing a tokenizer that is compatible with heterogeneity. Secondly, the complexity of the transformer hampers the acquisition of high-order neighbor information in HGs. In this paper, we propose a novel Hop-basedHeterogeneous Graph Transformer (H2Gormer) framework, paving a promising path for HGs to benefit from the capabilities of Transformers. We propose a Heterogeneous Hop-based Token …


Test-Time Augmentation For 3d Point Cloud Classification And Segmentation, Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung Mar 2024

Test-Time Augmentation For 3d Point Cloud Classification And Segmentation, Tuan-Anh Vu, Srinjay Sarkar, Zhiyuan Zhang, Binh-Son Hua, Sai-Kit Yeung

Research Collection School Of Computing and Information Systems

Data augmentation is a powerful technique to enhance the performance of a deep learning task but has received less attention in 3D deep learning. It is well known that when 3D shapes are sparsely represented with low point density, the performance of the downstream tasks drops significantly. This work explores test-time augmentation (TTA) for 3D point clouds. We are inspired by the recent revolution of learning implicit representation and point cloud upsampling, which can produce high-quality 3D surface reconstruction and proximity-to-surface, respectively. Our idea is to leverage the implicit field reconstruction or point cloud upsampling techniques as a systematic way …


Delving Into Multimodal Prompting For Fine-Grained Visual Classification, Xin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du, Shengfeng He, Zechao Li Feb 2024

Delving Into Multimodal Prompting For Fine-Grained Visual Classification, Xin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du, Shengfeng He, Zechao Li

Research Collection School Of Computing and Information Systems

Fine-grained visual classification (FGVC) involves categorizing fine subdivisions within a broader category, which poses challenges due to subtle inter-class discrepancies and large intra-class variations. However, prevailing approaches primarily focus on uni-modal visual concepts. Recent advancements in pre-trained vision-language models have demonstrated remarkable performance in various high-level vision tasks, yet the applicability of such models to FGVC tasks remains uncertain. In this paper, we aim to fully exploit the capabilities of cross-modal description to tackle FGVC tasks and propose a novel multimodal prompting solution, denoted as MP-FGVC, based on the contrastive language-image pertaining (CLIP) model. Our MP-FGVC comprises a multimodal prompts …


Leveraging Llms And Generative Models For Interactive Known-Item Video Search, Zhixin Ma, Jiaxin Wu, Chong-Wah Ngo Feb 2024

Leveraging Llms And Generative Models For Interactive Known-Item Video Search, Zhixin Ma, Jiaxin Wu, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

While embedding techniques such as CLIP have considerably boosted search performance, user strategies in interactive video search still largely operate on a trial-and-error basis. Users are often required to manually adjust their queries and carefully inspect the search results, which greatly rely on the users’ capability and proficiency. Recent advancements in large language models (LLMs) and generative models offer promising avenues for enhancing interactivity in video retrieval and reducing the personal bias in query interpretation, particularly in the known-item search. Specifically, LLMs can expand and diversify the semantics of the queries while avoiding grammar mistakes or the language barrier. In …


Out-Of-Distribution Detection In Long-Tailed Recognition With Calibrated Outlier Class Learning, Wenjun Miao, Guansong Pang, Xiao Bai, Tianqi Li, Jin Zheng Feb 2024

Out-Of-Distribution Detection In Long-Tailed Recognition With Calibrated Outlier Class Learning, Wenjun Miao, Guansong Pang, Xiao Bai, Tianqi Li, Jin Zheng

Research Collection School Of Computing and Information Systems

Existing out-of-distribution (OOD) methods have shown great success on balanced datasets but become ineffective in long-tailed recognition (LTR) scenarios where 1) OOD samples are often wrongly classified into head classes and/or 2) tail-class samples are treated as OOD samples. To address these issues, current studies fit a prior distribution of auxiliary/pseudo OOD data to the long-tailed in-distribution (ID) data. However, it is difficult to obtain such an accurate prior distribution given the unknowingness of real OOD samples and heavy class imbalance in LTR. A straightforward solution to avoid the requirement of this prior is to learn an outlier class to …


Vadclip: Adapting Vision-Language Models For Weakly Supervised Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang, Yanning Zhang Feb 2024

Vadclip: Adapting Vision-Language Models For Weakly Supervised Video Anomaly Detection, Peng Wu, Xuerong Zhou, Guansong Pang, Lingru Zhou, Qingsen Yan, Peng Wang, Yanning Zhang

Research Collection School Of Computing and Information Systems

The recent contrastive language-image pre-training (CLIP) model has shown great success in a wide range of image-level tasks, revealing remarkable ability for learning powerful visual representations with rich semantics. An open and worthwhile problem is efficiently adapting such a strong model to the video domain and designing a robust video anomaly detector. In this work, we propose VadCLIP, a new paradigm for weakly supervised video anomaly detection (WSVAD) by leveraging the frozen CLIP model directly without any pre-training and fine-tuning process. Unlike current works that directly feed extracted features into the weakly supervised classifier for frame-level binary classification, VadCLIP makes …


Living Datasets: Towards Data-Centric Ai Explainability And Bias Mitigation, Akib Zaman Jan 2024

Living Datasets: Towards Data-Centric Ai Explainability And Bias Mitigation, Akib Zaman

Computer Science and Engineering Dissertations - Archive

Benchmark datasets are critical to the evolution of AI efforts yet often embed unintended biases that influence the models that drive human-AI interactions. A deeper inspection and awareness of data is needed to understand the biases datasets may contain. In this dissertation, I introduce the Tag-and-Release method, inspired from wildlife research, that treats data as an organism and examines how different environments (i.e., CNNs) select for unique traits or characteristics that ultimately impact data's survival. Using the canonical MNIST handwritten digit dataset as a case study, I describe how the Tag-and-Release method can be used to analyze how dataset imbalance …


A Unified Cross-Modal Interactive System For Assisting Vision Impaired In Human Navigation And Indoor Based Human Robot Interaction, Harish Ram Nambiappan Jan 2024

A Unified Cross-Modal Interactive System For Assisting Vision Impaired In Human Navigation And Indoor Based Human Robot Interaction, Harish Ram Nambiappan

Computer Science and Engineering Dissertations - Archive

People who are blind and vision impaired often require assistance in performing various tasks. With new technologies emerging in the recent years, vision impaired people either require assistance in accessing those technologies or in using those technologies to perform different tasks in real life. Previous works have focused on assisting vision impaired people in different scenarios such as navigation, accessing smartphone interfaces etc. With the recent developments in robotics, a new research has emerged where new systems can be developed for vision impaired people to interact with robots to perform various human robot interactive tasks. But with developing new and …