Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (42)
- Computer Engineering (24)
- Graphics and Human Computer Interfaces (17)
- Electrical and Computer Engineering (15)
- Data Science (14)
-
- Numerical Analysis and Scientific Computing (8)
- Robotics (8)
- Social and Behavioral Sciences (8)
- Theory and Algorithms (8)
- Medicine and Health Sciences (7)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (6)
- Other Computer Sciences (6)
- Software Engineering (6)
- Statistics and Probability (6)
- Systems Architecture (6)
- Aerospace Engineering (5)
- Databases and Information Systems (5)
- Diagnosis (5)
- Geography (5)
- Navigation, Guidance, Control and Dynamics (5)
- Operations Research, Systems Engineering and Industrial Engineering (5)
- Medical Specialties (4)
- Systems Science (4)
- Applied Statistics (3)
- Biomedical Engineering and Bioengineering (3)
- Civil and Environmental Engineering (3)
- Computer and Systems Architecture (3)
- Institution
-
- MBZUAI (22)
- Singapore Management University (19)
- Old Dominion University (16)
- Portland State University (8)
- Loyola University Chicago (7)
-
- New Jersey Institute of Technology (6)
- San Jose State University (6)
- University of Kentucky (5)
- China Simulation Federation (4)
- University of Arkansas, Fayetteville (4)
- University of Texas at Arlington (4)
- California Polytechnic State University, San Luis Obispo (3)
- Edith Cowan University (3)
- Purdue University (3)
- Technological University Dublin (3)
- University of Denver (3)
- Air Force Institute of Technology (2)
- City University of New York (CUNY) (2)
- Dartmouth College (2)
- Kennesaw State University (2)
- Northern Illinois University (2)
- University at Albany, State University of New York (2)
- University of Nebraska - Lincoln (2)
- University of South Florida (2)
- Chapman University (1)
- Clemson University (1)
- Embry-Riddle Aeronautical University (1)
- Florida Institute of Technology (1)
- Georgia Southern University (1)
- Louisiana State University (1)
- Publication Year
- Publication
-
- Computer Vision Faculty Publications (20)
- Research Collection School Of Computing and Information Systems (18)
- Computer Science: Faculty Publications and Other Works (7)
- Dissertations (6)
- Master's Projects (6)
-
- Dissertations and Theses (5)
- Theses and Dissertations--Computer Science (5)
- Electronic Theses and Dissertations (4)
- Journal of System Simulation (4)
- Computer Science Faculty Publications (3)
- Electrical & Computer Engineering Faculty Publications (3)
- Electrical & Computer Engineering Theses & Dissertations (3)
- Master's Theses (3)
- Computer Science Theses & Dissertations (2)
- Computer Science and Computer Engineering Undergraduate Honors Theses (2)
- Computer Science and Engineering Dissertations - Archive (2)
- Conference papers (2)
- Dissertations, Theses, and Capstone Projects (2)
- Graduate Research Theses & Dissertations (2)
- Legacy Theses & Dissertations (2009 - 2024) (2)
- Machine Learning Faculty Publications (2)
- Open Access Dissertations (2)
- Published and Grey Literature from PhD Candidates (2)
- Student Research Symposium (2)
- Theses and Dissertations (2)
- Theses: Doctorates and Masters (2)
- USF Tampa Graduate Theses and Dissertations (2)
- All Theses (1)
- Articles (1)
- Beyond: Undergraduate Research Journal (1)
- Publication Type
Articles 31 - 60 of 157
Full-Text Articles in Artificial Intelligence and Robotics
Etdsuite: A Toolkit To Mine Electronic Theses And Dissertations To Enrich Scholarly Big Data Using Natural Language Processing And Computer Vision, Muntabir Hasan Choudhury
Etdsuite: A Toolkit To Mine Electronic Theses And Dissertations To Enrich Scholarly Big Data Using Natural Language Processing And Computer Vision, Muntabir Hasan Choudhury
Computer Science Theses & Dissertations
In the past decades, there has been a growing interest in mining scientific documents to obtain domain knowledge automatically. One of the understudied types of scientific documents is Electronic Theses and Dissertations (ETDs), as ETDs have distinct features compared with conference proceedings and journal articles. ETDs usually serve as partial requirements of academic degrees for students pursuing higher education. They are book-length documents (i.e., 100 – 400 pages long), and the topics may shift across chapters, exhibit the significant contribution of a student’s research over the entire degree pursuing period, and have unique metadata schema and page layouts. However, the …
Improving Out-Of-Distribution Detection With Disentangled Foreground And Background Features, Choubo Ding, Guansong Pang
Improving Out-Of-Distribution Detection With Disentangled Foreground And Background Features, Choubo Ding, Guansong Pang
Research Collection School Of Computing and Information Systems
Detecting out-of-distribution (OOD) inputs is a principal task for ensuring the safety of deploying deep-neural-network classifiers in open-set scenarios. OOD samples can be drawn from arbitrary distributions and exhibit deviations from in-distribution (ID) data in various dimensions, such as foreground features (e.g., objects in CIFAR100 images vs. those in CIFAR10 images) and background features (e.g., textural images vs. objects in CIFAR10). Existing methods can confound foreground and background features in training, failing to utilize the background features for OOD detection. This paper considers the importance of feature disentanglement in out-of-distribution detection and proposes the simultaneous exploitation of both foreground and …
Robotic Odor Source Localization Using Vision And Olfaction Sensing, Sunzid Hassan
Robotic Odor Source Localization Using Vision And Olfaction Sensing, Sunzid Hassan
Master's Theses
Robotic Odor Source Localization (ROSL) technology allows autonomous agents like robots to find an odor source in unknown environments. A successful odor source location depends crucially on an effective navigation algorithm that directs the robot towards the odor source. This thesis is a combination of three projects. First, we detail development of a versatile multi-modal robotic platform for ROSL real-world ROSL experimentation and discussed real-world validation of a traditional olfactionbased ROSL algorithm. Secondly, we introduced vision in ROSL by proposing a fusion navigation algorithm that integrates deep-learning enabled vision and olfaction-based navigation. This hybrid approach tackles challenges such as turbulent …
Development Of Feature Extraction Models To Improve Image Analysis Applications In Cancer, Yu Shi
Development Of Feature Extraction Models To Improve Image Analysis Applications In Cancer, Yu Shi
Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–
Cancer poses a significant global health challenge. With an estimated 20 million new cases diagnosed worldwide in 2022 and 9.7 million fatalities attributable to the disease, the economic burden of cancer is immense. It impacts healthcare systems and imposes substantial costs for its care on patients and their families. Despite advancements in early detection, prevention, and treatment that have reduced overall cancer mortality rates, the growing prevalence of cancer, particularly among younger individuals, remains a pressing issue.
Recent advancements in medical imaging technology have progressed significantly with the help of emerging computer vision and artificial intelligence (AI) technology. Despite these …
Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Beyond Textual Constraints : Learning Novel Diffusion Conditions With Fewer Examples, Yuyang Yu, Bangzhen Liu, Chenxi Zheng, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
In this paper, we delve into a novel aspect of learning novel diffusion conditions with datasets an order of magnitude smaller. The rationale behind our approach is the elimination of textual constraints during the few-shot learning process. To that end, we implement two optimization strategies. The first, prompt-free conditional learning, utilizes a prompt-free encoder derived from a pre-trained Stable Diffusion model. This strategy is designed to adapt new conditions to the diffusion process by minimizing the textual-visual cor-relation, thereby ensuring a more precise alignment between the generated content and the specified conditions. The second strategy entails condition-specific negative rectification, which …
Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He
Learning With Unreliability : Fast Few-Shot Voxel Radiance Fields With Relative Geometric Consistency, Yingjie Xu, Bangzhen Liu, Hao Tang, Bailin Deng, Shengfeng He
Research Collection School Of Computing and Information Systems
We propose a voxel-based optimization framework, Re VoRF, for few-shot radiance fields that strategically ad-dress the unreliability in pseudo novel view synthesis. Our method pivots on the insight that relative depth relationships within neighboring regions are more reliable than the ab-solute color values in disoccluded areas. Consequently, we devise a bilateral geometric consistency loss that carefully navigates the trade-off between color fidelity and geometric accuracy in the context of depth consistency for uncertain regions. Moreover, we present a reliability-guided learning strategy to discern and utilize the variable quality across syn-thesized views, complemented by a reliability-aware voxel smoothing algorithm that smoothens …
Rethinking Multi-View Representation Learning Via Distilled Disentangling, Guanzhou Ke, Bo Wang, Xiaoli Wang, Shengfeng He
Rethinking Multi-View Representation Learning Via Distilled Disentangling, Guanzhou Ke, Bo Wang, Xiaoli Wang, Shengfeng He
Research Collection School Of Computing and Information Systems
Multi-view representation learning aims to derive robust representations that are both view-consistent and view-specific from diverse data sources. This paper presents an in-depth analysis of existing approaches in this domain, highlighting a commonly overlooked aspect: the redundancy between view-consistent and view-specific representations. To this end, we propose an innovative framework for multi-view representation learning, which incorporates a technique we term 'distilled disentangling'. Our method introduces the concept of masked cross-view prediction, enabling the extraction of compact, high-quality view-consistent representations from various sources without incurring extra computational overhead. Additionally, we develop a distilled disentangling module that efficiently filters out consistency-related information …
Analyzing Swimming Performance Using Drone Captured Aerial Videos, Ngoc Doan Thu Tran, Kenny Tsu Wei Choo, Shaohui Foong, Hitesh Bhardwaj, Shane Kyi Hla Win, Wei Jun Ang, Kenneth T. Goh, Rajesh Krishna Balan
Analyzing Swimming Performance Using Drone Captured Aerial Videos, Ngoc Doan Thu Tran, Kenny Tsu Wei Choo, Shaohui Foong, Hitesh Bhardwaj, Shane Kyi Hla Win, Wei Jun Ang, Kenneth T. Goh, Rajesh Krishna Balan
Research Collection School Of Computing and Information Systems
Monitoring swimmer performance is crucial for improving training and enhancing athletic techniques. Traditional methods for tracking swimmers, such as above-water and underwater cameras, face limitations due to the need for multiple cameras and obstructions from water splashes. This paper presents a novel approach for tracking swimmers using a moving UAV. The proposed system employs a UAV equipped with a high-resolution camera to capture aerial footage of the swimmers. The footage is then processed using computer vision algorithms to extract the swimmers' positions and movements. This approach offers several advantages, including single camera use and comprehensive coverage. The system's accuracy is …
3-D Reconstruction For Underwater Robots With A Monocular Camera And Lights, Monika Roznere
3-D Reconstruction For Underwater Robots With A Monocular Camera And Lights, Monika Roznere
Dartmouth College Ph.D Dissertations
Before a robot can act, it must perceive its environment. Though, this is not a simple task when considering the challenges in underwater domains -- poor visibility conditions, limited sensor configurations, and lack of readily accessible localization. Underwater robots have, nevertheless, improved dramatically with more extensive sensor and navigation equipment. Robot and sensor use have enabled us to explore all reaches of our oceans. On the other hand, these same robots are not easily accessible or transferable to many practical tasks, including fishery management, infrastructure maintenance, disaster response, site conservation, and ecological surveys. There is a growing need for robots …
Context-Aware Affective Behavior Modeling And Analytics, Md Taufeeq Uddin
Context-Aware Affective Behavior Modeling And Analytics, Md Taufeeq Uddin
USF Tampa Graduate Theses and Dissertations
Affective computing (AC) is a sub-domain of AI that has the potential to assist people by assessing mental states and making appropriate recommendations to patients, loved ones, caregivers, and domain experts. Humans usually produce an enormous amount of data (such as face videos) every day. One of the major challenges for affective computer vision is to efficiently deal with high volumes of data to facilitate automated model development. To cope with this challenge, we developed computer vision algorithms that measure the expressivity of the human face from video data. More precisely, the developed algorithms can map complex affect information from …
A Computer Vision Solution To Cross-Cultural Food Image Classification And Nutrition Logging, Rohan Sethi, George K. Thiruvathukal
A Computer Vision Solution To Cross-Cultural Food Image Classification And Nutrition Logging, Rohan Sethi, George K. Thiruvathukal
Computer Science: Faculty Publications and Other Works
The US is a culturally and ethnically diverse country, and with this diversity comes a myriad of cuisines and eating habits that expand well beyond that of western culture. Each of these meals have their own good and bad effects when it comes to the nutritional value and its potential impact on human health. Thus, there is a greater need for people to be able to access the nutritional profile of their diverse daily meals and better manage their health. A revolutionary solution to democratize food image classification and nutritional logging is using deep learning to extract that information from …
An Automated Approach For Improving The Inference Latency And Energy Efficiency Of Pretrained Cnns By Removing Irrelevant Pixels With Focused Convolutions, Caleb Tung, Nick Eliopoulos, Purvish Jajal, Gowri Ramshankar, Chen-Yun Yang, Nicholas Synovic, Xuecen Zhang, Vipin Chaudhary, George K. Thiruvathukal, Yung-Hsiang Lu
An Automated Approach For Improving The Inference Latency And Energy Efficiency Of Pretrained Cnns By Removing Irrelevant Pixels With Focused Convolutions, Caleb Tung, Nick Eliopoulos, Purvish Jajal, Gowri Ramshankar, Chen-Yun Yang, Nicholas Synovic, Xuecen Zhang, Vipin Chaudhary, George K. Thiruvathukal, Yung-Hsiang Lu
Computer Science: Faculty Publications and Other Works
Computer vision often uses highly accurate Convolutional Neural Networks (CNNs), but these deep learning models are associated with ever-increasing energy and computation requirements. Producing more energy-efficient CNNs often requires model training which can be cost-prohibitive. We propose a novel, automated method to make a pretrained CNN more energy-efficient without re-training. Given a pretrained CNN, we insert a threshold layer that filters activations from the preceding layers to identify regions of the image that are irrelevant, i.e. can be ignored by the following layers while maintaining accuracy. Our modified focused convolution operation saves inference latency (by up to 25%) and energy …
Efficient Classification Of Very High Resolution Images, Mohammad I. Nouyed
Efficient Classification Of Very High Resolution Images, Mohammad I. Nouyed
Graduate Theses, Dissertations, and Problem Reports (ETD)
In recent decades, deep learning approaches have shown significant improvement in various image understanding tasks. However, analysis of high-resolution images remains a major challenge. In this work, we address the challenge of very high-resolution histopathological image (VHRHI) classification using a new information-theoretic discriminative patch selection approach. We show results on a high-resolution image dataset, namely, gigapixel whole slide tissue images for cancer tumors. Then we address how to efficiently classify challenging histopathology images, such as gigapixel whole-slide images for cancer diagnostics with image-level annotation. These ``weak labels'' are applied throughout the image but describe tumor regions of variable sizes and …
Automatic Classification Of Activities In Classroom Videos, Jonathan K. Foster, Matthew Korban, Peter Youngs, Ginger S. Watson, Scott T. Acton
Automatic Classification Of Activities In Classroom Videos, Jonathan K. Foster, Matthew Korban, Peter Youngs, Ginger S. Watson, Scott T. Acton
VMASC Publications
Classroom videos are a common source of data for educational researchers studying classroom interactions as well as a resource for teacher education and professional development. Over the last several decades emerging technologies have been applied to classroom videos to record, transcribe, and analyze classroom interactions. With the rise of machine learning, we report on the development and validation of neural networks to classify instructional activities using video signals, without analyzing speech or audio features, from a large corpus of nearly 250 h of classroom videos from elementary mathematics and English language arts instruction. Results indicated that the neural networks performed …
A Survey On Few-Shot Class-Incremental Learning, Songsong Tian, Lusi Li, Weijun Li, Hang Ran, Xin Ning, Prayag Tiwari
A Survey On Few-Shot Class-Incremental Learning, Songsong Tian, Lusi Li, Weijun Li, Hang Ran, Xin Ning, Prayag Tiwari
Computer Science Faculty Publications
Large deep learning models are impressive, but they struggle when real-time data is not available. Few-shot class-incremental learning (FSCIL) poses a significant challenge for deep neural networks to learn new tasks from just a few labeled samples without forgetting the previously learned ones. This setup can easily leads to catastrophic forgetting and overfitting problems, severely affecting model performance. Studying FSCIL helps overcome deep learning model limitations on data volume and acquisition time, while improving practicality and adaptability of machine learning models. This paper provides a comprehensive survey on FSCIL. Unlike previous surveys, we aim to synthesize few-shot learning and incremental …
Cvii: Enhancing Interpretability In Intelligent Sensor Systems Via Computer Vision Interpretability Index, Hossein Mohammadi, Krishnaprasad Thirunarayan, Lingwei Chen
Cvii: Enhancing Interpretability In Intelligent Sensor Systems Via Computer Vision Interpretability Index, Hossein Mohammadi, Krishnaprasad Thirunarayan, Lingwei Chen
Computer Science and Engineering Faculty Publications
In the realm of intelligent sensor systems, the dependence on Artificial Intelligence (AI) applications has heightened the importance of interpretability. This is particularly critical for opaque models such as Deep Neural Networks (DNN), as understanding their decisions is essential, not only for ethical and regulatory compliance, but also for fostering trust in AI-driven outcomes. This paper introduces the novel concept of a Computer Vision Interpretability Index (CVII). The CVII framework is designed to emulate human cognitive processes, specifically in tasks related to vision. It addresses the intricate challenge of quantifying interpretability, a task that is inherently subjective and varies across …
Implementation Of Adas And Autonomy On Unlv Campus, Zillur Rahman
Implementation Of Adas And Autonomy On Unlv Campus, Zillur Rahman
UNLV Theses, Dissertations, Professional Papers, and Capstones
The integration of Advanced Driving Assistance Systems (ADAS) and autonomous driving functionalities into contemporary vehicles has notably surged, driven by the remarkable progress in artificial intelligence (AI). These AI systems, capable of learning from real-world data, now exhibit the capability to perceive their surroundings via a suite of sensors, create optimal routes from source to destination, and execute vehicle control akin to a human driver.
Within the context of this thesis, we undertake a comprehensive exploration of three distinct yet interrelated ADAS and Autonomy projects. Our central objective is the implementation of autonomous driving(AD) technology at UNLV campus, culminating in …
Dynamic Translation Of Drone View To Vehicle View For Autonomous Driving, Grayson Byrd
Dynamic Translation Of Drone View To Vehicle View For Autonomous Driving, Grayson Byrd
All Theses
As the field of computer vision continues to advance, the use of autonomous vehicles in military applications has increased and the tasks associated with these systems have grown in scope and complexity. These vehicles tend to operate in combat situations, presenting significant risk in the form of sensor damage. Since these autonomy algorithms rely on sensor information to navigate their environment, any threat to sensor functionality negatively impacts the reliability of the system. As a potential solution, I propose Dynamic Diffusion-based View Translation (DDVT), a novel computer vision algorithm capable of restoring image sensor function in ground vehicles through aerial …
Smart Street Light Control: A Review On Methods, Innovations, And Extended Applications, Fouad Agramelal, Mohamed Sadik, Youssef Moubarak, Saad Abouzahir
Smart Street Light Control: A Review On Methods, Innovations, And Extended Applications, Fouad Agramelal, Mohamed Sadik, Youssef Moubarak, Saad Abouzahir
Computer Vision Faculty Publications
As urbanization increases, streetlights have become significant consumers of electrical power, making it imperative to develop effective control methods for sustainability. This paper offers a comprehensive review on control methods of smart streetlight systems, setting itself apart by introducing a novel light scheme framework that provides a structured classification of various light control patterns, thus filling an existing gap in the literature. Unlike previous studies, this work dives into the technical specifics of individual research papers and methodologies, ranging from basic to advanced control methods like computer vision and deep learning, while also assessing the energy consumption associated with each …
Object Recognition With Deep Neural Networks In Low-End Systems, Lillian Davis
Object Recognition With Deep Neural Networks In Low-End Systems, Lillian Davis
Mahurin Honors College Capstone Experience/Thesis Projects
Object recognition is an important area in computer vision. Object recognition has been advanced significantly by deep learning that unifies feature extraction and classification. In general, deep neural networks, such as Convolution Neural Networks (CNNs), are trained in high-performance systems. Aiming to extend the reach of deep learning to personal computing, I propose a study of deep learning-based object recognition in low-end systems, such as laptops. This research includes how differing layer configurations and hyperparameter values used in CNNs can either create or resolve the issue of overfitting and affect final accuracy levels of object recognition systems. The main contribution …
3d-Aware Multi-Class Image-To-Image Translation With Nerfs, Senmao Li, Joost Van De Weijer, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu, Jian Yang
3d-Aware Multi-Class Image-To-Image Translation With Nerfs, Senmao Li, Joost Van De Weijer, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu, Jian Yang
Computer Vision Faculty Publications
Recent advances in 3D-aware generative models (3D-aware GANs) combined with Neural Radiance Fields (NeRF) have achieved impressive results. However no prior works investigate 3D-aware GANs for 3D consistent multiclass image-to-image (3D-aware 121) translation. Naively using 2D-121 translation methods suffers from unrealistic shape/identity change. To perform 3D-aware multiclass 121 translation, we decouple this learning process into a multiclass 3D-aware GAN step and a 3D-aware 121 translation step. In the first step, we propose two novel techniques: a new conditional architecture and an effective training strategy. In the second step, based on the well-trained multiclass 3D-aware GAN architecture, that preserves view-consistency, we …
Discriminative Co-Saliency And Background Mining Transformer For Co-Salient Object Detection, Long Li, Junwei Han, Ni Zhang, Nian Liu, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan
Discriminative Co-Saliency And Background Mining Transformer For Co-Salient Object Detection, Long Li, Junwei Han, Ni Zhang, Nian Liu, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan
Computer Vision Faculty Publications
Most previous co-salient object detection works mainly focus on extracting co-salient cues via mining the consistency relations across images while ignore explicit exploration of background regions. In this paper, we propose a Discriminative co-saliency and background Mining Transformer framework (DMT) based on several economical multi-grained correlation modules to explicitly mine both co-saliency and background information and effectively model their discrimination. Specifically, we first propose a region-to-region correlation module for introducing inter-image relations to pixel-wise segmentation features while maintaining computational efficiency. Then, we use two types of pre-defined tokens to mine co-saliency and background information via our proposed contrast-induced pixel-to-token correlation …
Multiclass Confidence And Localization Calibration For Object Detection, Bimsara Pathiraja, Malitha Gunawardhana, Muhammad Haris Khan
Multiclass Confidence And Localization Calibration For Object Detection, Bimsara Pathiraja, Malitha Gunawardhana, Muhammad Haris Khan
Computer Vision Faculty Publications
Albeit achieving high predictive accuracy across many challenging computer vision problems, recent studies suggest that deep neural networks (DNNs) tend to make over-confident predictions, rendering them poorly calibrated. Most of the existing attempts for improving DNN calibration are limited to classification tasks and restricted to calibrating in-domain predictions. Surprisingly, very little to no attempts have been made in studying the calibration of object detection methods, which occupy a pivotal space in vision-based security-sensitive, and safety-critical applications. In this paper, we propose a new train-time technique for calibrating modern object detection methods. It is capable of jointly calibrating multiclass confidence and …
Autonomous Shipwreck Detection & Mapping, William Ard
Autonomous Shipwreck Detection & Mapping, William Ard
LSU Master's Theses
This thesis presents the development and testing of Bruce, a low-cost hybrid Remote Operated Vehicle (ROV) / Autonomous Underwater Vehicle (AUV) system for the optical survey of marine archaeological sites, as well as a novel sonar image augmentation strategy for semantic segmentation of shipwrecks. This approach takes side-scan sonar and bathymetry data collected using an EdgeTech 2205 AUV sensor integrated with an Harris Iver3, and generates augmented image data to be used for the semantic segmentation of shipwrecks. It is shown that, due to the feature enhancement capabilities of the proposed shipwreck detection strategy, correctly identified areas have a 15% …
Fine-Grained Domain Adaptive Crowd Counting Via Point-Derived Segmentation, Yongtuo Liu, Dan Xu, Sucheng Ren, Hanjie Wu, Hongmin Cai, Shengfeng He
Fine-Grained Domain Adaptive Crowd Counting Via Point-Derived Segmentation, Yongtuo Liu, Dan Xu, Sucheng Ren, Hanjie Wu, Hongmin Cai, Shengfeng He
Research Collection School Of Computing and Information Systems
Due to domain shift, a large performance drop is usually observed when a trained crowd counting model is deployed in the wild. While existing domain-adaptive crowd counting methods achieve promising results, they typically regard each crowd image as a whole and reduce domain discrepancies in a holistic manner, thus limiting further improvement of domain adaptation performance. To this end, we propose to untangle domain-invariant crowd and domain-specific background from crowd images and design a fine-grained domain adaption method for crowd counting. Specifically, to disentangle crowd from background, we propose to learn crowd segmentation from point-level crowd counting annotations in a …
Tree-Based Unidirectional Neural Networks For Low-Power Computer Vision, Abhinav Goel, Caleb Tung, Nick Eliopoulos, Amy Wang, Jamie C. Davis, George K. Thiruvathukal, Yung-Hisang Lu
Tree-Based Unidirectional Neural Networks For Low-Power Computer Vision, Abhinav Goel, Caleb Tung, Nick Eliopoulos, Amy Wang, Jamie C. Davis, George K. Thiruvathukal, Yung-Hisang Lu
Computer Science: Faculty Publications and Other Works
This article describes the novel Tree-based Unidirectional Neural Network (TRUNK) architecture. This architecture improves computer vision efficiency by using a hierarchy of multiple shallow Convolutional Neural Networks (CNNs), instead of a single very deep CNN. We demonstrate this architecture’s versatility in performing different computer vision tasks efficiently on embedded devices. Across various computer vision tasks, the TRUNK architecture consumes 65% less energy and requires 50% less memory than representative low-power CNN architectures, e.g., MobileNet v2, when deployed on the NVIDIA Jetson Nano.
Curricular Contrastive Regularization For Physics-Aware Single Image Dehazing, Yu Zheng, Jiahui Zhan, Shengfeng He, Yong Du
Curricular Contrastive Regularization For Physics-Aware Single Image Dehazing, Yu Zheng, Jiahui Zhan, Shengfeng He, Yong Du
Research Collection School Of Computing and Information Systems
Considering the ill-posed nature, contrastive regularization has been developed for single image dehazing, introducing the information from negative images as a lower bound. However, the contrastive samples are non-consensual, as the negatives are usually represented distantly from the clear (i.e., positive) image, leaving the solution space still under-constricted. Moreover, the interpretability of deep dehazing models is underexplored towards the physics of the hazing process. In this paper, we propose a novel curricular contrastive regularization targeted at a consensual contrastive space as opposed to a non-consensual one. Our negatives, which provide better lower-bound constraints, can be assembled from 1) the hazy …
Where Is My Spot? Few-Shot Image Generation Via Latent Subspace Optimization, Chenxi Zheng, Bangzhen Liu, Huaidong Zhang, Xuemiao Xu, Shengfeng He
Where Is My Spot? Few-Shot Image Generation Via Latent Subspace Optimization, Chenxi Zheng, Bangzhen Liu, Huaidong Zhang, Xuemiao Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Image generation relies on massive training data that can hardly produce diverse images of an unseen category according to a few examples. In this paper, we address this dilemma by projecting sparse few-shot samples into a continuous latent space that can potentially generate infinite unseen samples. The rationale behind is that we aim to locate a centroid latent position in a conditional StyleGAN, where the corresponding output image on that centroid can maximize the similarity with the given samples. Although the given samples are unseen for the conditional StyleGAN, we assume the neighboring latent subspace around the centroid belongs to …
Bubbleu: Exploring Augmented Reality Game Design With Uncertain Ai-Based Interaction, Minji Kim, Kyungjin Lee, Rajesh Krishna Balan, Youngki Lee
Bubbleu: Exploring Augmented Reality Game Design With Uncertain Ai-Based Interaction, Minji Kim, Kyungjin Lee, Rajesh Krishna Balan, Youngki Lee
Research Collection School Of Computing and Information Systems
Object detection, while being an attractive interaction method for Augmented Reality (AR), is fundamentally error-prone due to the probabilistic nature of the underlying AI models, resulting in sub-optimal user experiences. In this paper, we explore the effect of three game design concepts, Ambiguity, Transparency, and Controllability, to provide better gameplay experiences in AR games that use error-prone object detection-based interaction modalities. First, we developed a base AR pet breeding game, called Bubbleu that uses object detection as a key interaction method. We then implemented three different variants, each according to the three concepts, to investigate the impact of each design …
Pose- And Attribute-Consistent Person Image Synthesis, Cheng Xu, Zejun Chen, Jiajie Mai, Xuemiao Xu, Shengfeng He
Pose- And Attribute-Consistent Person Image Synthesis, Cheng Xu, Zejun Chen, Jiajie Mai, Xuemiao Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
PersonImageSynthesisaimsattransferringtheappearanceofthesourcepersonimageintoatargetpose. Existingmethods cannot handle largeposevariations and therefore suffer fromtwocritical problems: (1)synthesisdistortionduetotheentanglementofposeandappearanceinformationamongdifferentbody componentsand(2)failureinpreservingoriginalsemantics(e.g.,thesameoutfit).Inthisarticle,weexplicitly addressthesetwoproblemsbyproposingaPose-andAttribute-consistentPersonImageSynthesisNetwork (PAC-GAN).Toreduceposeandappearancematchingambiguity,weproposeacomponent-wisetransferring modelconsistingoftwostages.Theformerstagefocusesonlyonsynthesizingtargetposes,whilethelatter renderstargetappearancesbyexplicitlytransferringtheappearanceinformationfromthesourceimageto thetargetimageinacomponent-wisemanner. Inthisway,source-targetmatchingambiguityiseliminated duetothecomponent-wisedisentanglementofposeandappearancesynthesis.Second,tomaintainattribute consistency,werepresenttheinputimageasanattributevectorandimposeahigh-levelsemanticconstraint usingthisvectortoregularizethetargetsynthesis.ExtensiveexperimentalresultsontheDeepFashiondataset demonstratethesuperiorityofourmethodoverthestateoftheart,especiallyformaintainingposeandattributeconsistenciesunderlargeposevariations.