Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Computer vision

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 151 - 180 of 349

Full-Text Articles in Computer Sciences

Real Time Evaluation Of Boom And Drogue Occlusion With Aar, Xiaoyang Wu Mar 2022

Real Time Evaluation Of Boom And Drogue Occlusion With Aar, Xiaoyang Wu

Theses and Dissertations

In recent years, Unmanned Aerial Vehicles (UAV) have seen a rise in popularity. Various navigational algorithms have been developed as a solution to estimate a UAV’s pose relative to the refueler aircraft. The result can be used to safely automate aerial refueling (AAR) to improve UAVs’ time-on-station and ensure the success of military operations. This research aims to reach real-time performance using a GPU accelerated approach. It also conducts various experiments to quantify the effects of refueling boom/drogue occlusion and image exposure on the pose estimation pipeline in a lab setting.


Cocoa: Context-Conditional Adaptation For Recognizing Unseen Classes In Unseen Domains, Puneet Mangla, Shivam Chandhok, Vineeth N. Balasubramanian, Fahad Shahbaz Khan Feb 2022

Cocoa: Context-Conditional Adaptation For Recognizing Unseen Classes In Unseen Domains, Puneet Mangla, Shivam Chandhok, Vineeth N. Balasubramanian, Fahad Shahbaz Khan

Computer Vision Faculty Publications

Recent progress towards designing models that can generalize to unseen domains (i.e domain generalization) or unseen classes (i.e zero-shot learning) has embarked interest towards building models that can tackle both domain-shift and semantic shift simultaneously (i.e zero-shot domain generalization). For models to generalize to unseen classes in unseen domains, it is crucial to learn feature representation that preserves class-level (domain-invariant) as well as domain-specific information. Motivated from the success of generative zero-shot approaches, we propose a feature generative framework integrated with a COntext COnditional Adaptive (COCOA) Batch-Normalization layer to seamlessly integrate class-level semantic and domain-specific information. The generated visual features …


Low-Power Computer Vision: Improve The Efficiency Of Artificial Intelligence, George K. Thiruvathukal, Yung-Hisang Lu, Jaeyoun Kim, Yiran Chen, Bo Chen Feb 2022

Low-Power Computer Vision: Improve The Efficiency Of Artificial Intelligence, George K. Thiruvathukal, Yung-Hisang Lu, Jaeyoun Kim, Yiran Chen, Bo Chen

Computer Science: Faculty Publications and Other Works

Energy efficiency is critical for running computer vision on battery-powered systems, such as mobile phones or UAVs (unmanned aerial vehicles, or drones). This book collects the methods that have won the annual IEEE Low-Power Computer Vision Challenges since 2015. The winners share their solutions and provide insight on how to improve the efficiency of machine learning systems.


Transformers In Vision: A Survey, Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, Mubarak Shah Jan 2022

Transformers In Vision: A Survey, Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, Mubarak Shah

Computer Vision Faculty Publications

Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks e.g., Long short-term memory (LSTM). Different from convolutional networks, Transformers require minimal inductive biases for their design and are naturally suited as set-functions. Furthermore, the straightforward design of Transformers allows processing multiple modalities (e.g., images, videos, text and speech) using similar processing blocks and demonstrates excellent scalability to very large capacity networks and huge …


Survey Of Ship Detection In Video Surveillance Based On Shallow Machine Learning, Zhenbo Bi, Shiyou Zhang, Yang Hua, Yuanhong Wu Jan 2022

Survey Of Ship Detection In Video Surveillance Based On Shallow Machine Learning, Zhenbo Bi, Shiyou Zhang, Yang Hua, Yuanhong Wu

Journal of System Simulation

Abstract: At present, detection of ship targets in video surveillance based on shallow machine learning methods is still attracting attention in the fields of underwater cultural heritage protection, marine aquaculture, maritime traffic, and port management. This paper provides a review and discussion for this kind of ship detection methods. The ship target detection based on video surveillance is divided into five parts according to the key technologies involved: preprocessing, region of interest extraction, target segmentation, ship feature extraction and ship type recognition. According to different functional modules, the core problems involved in them are pointed out, and the core ideas, …


Explainabilityaudit: An Automated Evaluation Of Local Explainability In Rooftop Image Classification, Duleep Rathgamage Don, Jonathan Boardman, Sudhashree Sayenju, Ramazan Aygun, Yifan Zhang, Bill Franks, Sereres Johnston, George Lee, Dan Sullivan, Girish Modgil Jan 2022

Explainabilityaudit: An Automated Evaluation Of Local Explainability In Rooftop Image Classification, Duleep Rathgamage Don, Jonathan Boardman, Sudhashree Sayenju, Ramazan Aygun, Yifan Zhang, Bill Franks, Sereres Johnston, George Lee, Dan Sullivan, Girish Modgil

Published and Grey Literature from PhD Candidates

Explainable Artificial Intelligence (XAI) is a key concept in building trustworthy machine learning models. Local explainability methods seek to provide explanations for individual predictions. Usually, humans must check these explanations manually. When large numbers of predictions are being made, this approach does not scale. We address this deficiency for a rooftop classification problem specifically with ExplainabilityAudit, a method that automatically evaluates explanations generated by a local explainability toolkit and identifies rooftop images that require further auditing by a human expert. The proposed method utilizes explanations generated by the Local Interpretable Model-Agnostic Explanations (LIME) framework as the most important superpixels of …


Jointly-Learnt Networks For Future Action Anticipation Via Self-Knowledge Distillation And Cycle Consistency, Md Moniruzzaman, Zhaozheng Yin, Zhihai He, Ming-Chuan Leu, Ruwen Qin Jan 2022

Jointly-Learnt Networks For Future Action Anticipation Via Self-Knowledge Distillation And Cycle Consistency, Md Moniruzzaman, Zhaozheng Yin, Zhihai He, Ming-Chuan Leu, Ruwen Qin

Mechanical and Aerospace Engineering Faculty Research & Creative Works

Future action anticipation aims to infer future actions from the observation of a small set of past video frames. In this paper, we propose a novel Jointly learnt Action Anticipation Network (J-AAN) via Self-Knowledge Distillation (Self-KD) and cycle consistency for future action anticipation. In contrast to the current state-of-the-art methods which anticipate the future actions either directly or recursively, our proposed J-AAN anticipates the future actions jointly in both direct and recursive ways. However, when dealing with future action anticipation, one important challenge to address is the future's uncertainty since multiple action sequences may come from or be followed by …


Lpvit: A Transformer Based Model For Pcb Image Classification And Defect Detection, Kang An, Yanping Zhang Jan 2022

Lpvit: A Transformer Based Model For Pcb Image Classification And Defect Detection, Kang An, Yanping Zhang

Computer Science Faculty Scholarship

PCB (printed circuit board) is an extremely important component of all electronic products, which has greatly facilitated human life. Meanwhile, tons of PCBs in the waste streams become a waste of resources, which puts the recycling and reuse of PCBs in urgent need. In the manufacturing and recycling of electronic products, the classification of PCBs, recognition of sub-components, and defect detection have been the key technology. Traditional manual detection and classification are subjective and rely on individuals’ experience. With the development of artificial intelligence, lots of research efforts have been dedicated to the automated detection and recognition of PCBs. In …


Image-Based Malware Classification Hybrid Framework Based On Space-Filling Curves, Stephen O Shaughnessy, Stephen Sheridan Jan 2022

Image-Based Malware Classification Hybrid Framework Based On Space-Filling Curves, Stephen O Shaughnessy, Stephen Sheridan

Articles

There exists a never-ending “arms race” between malware analysts and adversarial malicious code developers as malevolent programs evolve and countermeasures are developed to detect and eradicate them. Malware has become more complex in its intent and capabilities over time, which has prompted the need for constant improvement in detection and defence methods. Of particular concern are the anti-analysis obfuscation techniques, such as packing and encryption, that are employed by malware developers to evade detection and thwart the analysis process. In such cases, malware is generally impervious to basic analysis methods and so analysts must use more invasive techniques to extract …


Faking Sensor Noise Information, Justin Chang Jan 2022

Faking Sensor Noise Information, Justin Chang

Master's Projects

Noise residue detection in digital images has recently been used as a method to classify images based on source camera model type. The meteoric rise in the popularity of using Neural Network models has also been used in conjunction with the concept of noise residuals to classify source camera models. However, many papers gloss over the details on the methods of obtaining noise residuals and instead rely on the self- learning aspect of deep neural networks to implicitly discover this themselves. For this project I propose a method of obtaining noise residuals (“noiseprints”) and denoising an image, as well as …


Construction Of A Repeatable Framework For Prostate Cancer Lesion Binary Semantic Segmentation Using Convolutional Neural Networks, Ian Vincent O. Mirasol, Patricia Angela R. Abu, Rosula Sj Reyes Jan 2022

Construction Of A Repeatable Framework For Prostate Cancer Lesion Binary Semantic Segmentation Using Convolutional Neural Networks, Ian Vincent O. Mirasol, Patricia Angela R. Abu, Rosula Sj Reyes

Department of Information Systems & Computer Science Faculty Publications

Prostate cancer is the 3rd most diagnosed cancer overall. Current screening methods such as the prostate-specific antigen test could result in overdiagonosis and overtreatment while other methods such as a transrectal ultrasonography are invasive. Recent medical advancements have allowed the use of multiparametric MRI — a noninvasive and reliable screening process for prostate cancer. However, assessment would still vary from different professionals introducing subjectivity. While con-volutional neural network has been used in multiple studies to ob-jectively segment prostate lesions, due to the sensitivity of datasets and varying ground-truth established used in these studies, it is not possible to reproduce and …


Building An Understanding Of Human Activities In First Person Video Using Fuzzy Inference, Bradley A. Schneider Jan 2022

Building An Understanding Of Human Activities In First Person Video Using Fuzzy Inference, Bradley A. Schneider

Browse all Theses and Dissertations

Activities of Daily Living (ADL’s) are the activities that people perform every day in their home as part of their typical routine. The in-home, automated monitoring of ADL’s has broad utility for intelligent systems that enable independent living for the elderly and mentally or physically disabled individuals. With rising interest in electronic health (e-Health) and mobile health (m-Health) technology, opportunities abound for the integration of activity monitoring systems into these newer forms of healthcare. In this dissertation we propose a novel system for describing ’s based on video collected from a wearable camera. Most in-home activities are naturally defined by …


Monitoring Plants Growth In Indoor Vertical Farms Using Computer Vision And Ai Techniques, Bhama Krishna Pillutla Jan 2022

Monitoring Plants Growth In Indoor Vertical Farms Using Computer Vision And Ai Techniques, Bhama Krishna Pillutla

Graduate Research Theses & Dissertations

Climatic conditions like temperature, drought, and heavy metals disturb plant cell structures and, ultimately, plant growth that significantly affects crop production. Due to increasing climate change, maize crop yields are projected to decline by 24% by the end of century. With the increase in food demands and decrease in agricultural land and water resources, the space for effective farming is left much desired. Though limited to a few crops at this moment, Indoor Vertical Farming is one technique that requires much less land space, water, soil, and sunlight when compared to traditional farming. Vertical farming allows artificial control of temperature, …


Customer Gaze Estimation In Retail Using Deep Learning, Shashimal Senarath, Primesh Pathirana, Dulani Meedeniya, Sampath Jayarathna Jan 2022

Customer Gaze Estimation In Retail Using Deep Learning, Shashimal Senarath, Primesh Pathirana, Dulani Meedeniya, Sampath Jayarathna

Computer Science Faculty Publications

At present, intelligent computing applications are widely used in different domains, including retail stores. The analysis of customer behaviour has become crucial for the benefit of both customers and retailers. In this regard, the concept of remote gaze estimation using deep learning has shown promising results in analyzing customer behaviour in retail due to its scalability, robustness, low cost, and uninterrupted nature. This study presents a three-stage, three-attention-based deep convolutional neural network for remote gaze estimation in retail using image data. In the first stage, we design a mechanism to estimate the 3D gaze of the subject using image data …


Machine Learning And Computer Vision In Solar Physics, Haodi Jiang Dec 2021

Machine Learning And Computer Vision In Solar Physics, Haodi Jiang

Dissertations

In the recent decades, the difficult task of understanding and predicting violent solar eruptions and their terrestrial impacts has become a strategic national priority, as it affects the life of human beings, including communication, transportation, the power grid, national defense, space travel, and more. This dissertation explores new machine learning and computer vision techniques to tackle this difficult task. Specifically, the dissertation addresses four interrelated problems in solar physics: magnetic flux tracking, fibril tracing, Stokes inversion and vector magnetogram generation.

First, the dissertation presents a new deep learning method, named SolarUnet, to identify and track solar magnetic flux elements in …


Ow-Detr: Open-World Detection Transformer, Akshita Gupta, Sanath Narayan, K.J. Joseph, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah Dec 2021

Ow-Detr: Open-World Detection Transformer, Akshita Gupta, Sanath Narayan, K.J. Joseph, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah

Computer Vision Faculty Publications

Open-world object detection (OWOD) is a challenging computer vision problem, where the task is to detect a known set of object categories while simultaneously identifying unknown objects. Additionally, the model must incrementally learn new classes that become known in the next training episodes. Distinct from standard object detection, the OWOD setting poses significant challenges for generating quality candidate proposals on potentially unknown objects, separating the unknown objects from the background and detecting diverse unknown objects. Here, we introduce a novel end-to-end transformer-based framework, OW-DETR, for open-world object detection. The proposed OW-DETR comprises three dedicated components namely, attention-driven pseudo-labeling, novelty classification …


Situate: An Agent-Based System For Situation Recognition, Max Henry Quinn Nov 2021

Situate: An Agent-Based System For Situation Recognition, Max Henry Quinn

Dissertations and Theses

Computer vision and machine learning systems have improved significantly in recent years, largely based on the development of deep learning systems, leading to impressive performance on object detection tasks. Understanding the content of images is considerably more difficult. Even simple situations, such as "a handshake", "walking the dog", "a game of ping-pong", or "people waiting for a bus", present significant challenges. Each consists of common objects, but are not reliably detectable as a single entity nor through the simple co-occurrence of their parts.

In this dissertation, toward the goal of developing machine learning systems that demonstrate properties associated with understanding, …


Fingerlings Mass Estimation: A Comparison Between Deep And Shallow Learning Algorithms, Adair Da Silva Oliveira Junior, Diego André Sant’Ana, Marcio Carneiro Brito Pache, Vanir Garcia, Vanessa Aparecida De Moares Weber, Gilberto Astolfi, Fabricio De Lima Weber, Geazy Vilharva Menezes, Gabriel Kirsten Menezes, Pedro Lucas França Albuquerque, Celso Soares Costa, Eduardo Quirino Arguelho De Queiroz, João Victor Araújo Rozales, Milena Wolff Ferreira, Marco Hiroshi Naka, Hemerson Pistori Nov 2021

Fingerlings Mass Estimation: A Comparison Between Deep And Shallow Learning Algorithms, Adair Da Silva Oliveira Junior, Diego André Sant’Ana, Marcio Carneiro Brito Pache, Vanir Garcia, Vanessa Aparecida De Moares Weber, Gilberto Astolfi, Fabricio De Lima Weber, Geazy Vilharva Menezes, Gabriel Kirsten Menezes, Pedro Lucas França Albuquerque, Celso Soares Costa, Eduardo Quirino Arguelho De Queiroz, João Victor Araújo Rozales, Milena Wolff Ferreira, Marco Hiroshi Naka, Hemerson Pistori

School of Computing: Faculty Publications

The paper presents some results regarding the automatic mass estimation of Pintado Real fingerlings, using machine learning techniques to support the fish production process. For this purpose, an image dataset called FISHCV1206FSEG, was created which is composed of 1206 images of fingerlings with their respective annotated masses. Through the fish contours, the area and perimeter were extracted, and submitted to the J48, SVM, and KNN classification algorithms and a linear regression algorithm. The images were also submitted to ResNet50, In- ceptionV3, Exception, VGG16, and VGG19 convolutional neural networks. As a result, the classification algorithm J48 reached an accuracy of 58.2% …


Advances In Deep Learning With Applications To Computer Vision And Astronomy, Zhihang Hu Aug 2021

Advances In Deep Learning With Applications To Computer Vision And Astronomy, Zhihang Hu

Dissertations

Deep Learning has spanned a variety of applications in computer vision as well as computational astronomy. These two aspects obtained similar data structure, therefore, their solutions can be transferable between each other. This dissertation look into two video-related tasks in computer vision and propose a novel problem in computational astronomy.

Specifically, acquiring an in-depth understanding of videos has been a cornerstone problem in computer vision. This problem has been studied by various researchers from different perspectives, among which video prediction has attracted much attention. Video prediction aims to generate the pixels of future frames given a sequence of context frames. …


Novel Statistical Modeling Methods For Traffic Video Analysis, Hang Shi Aug 2021

Novel Statistical Modeling Methods For Traffic Video Analysis, Hang Shi

Dissertations

Video analysis is an active and rapidly expanding research area in computer vision and artificial intelligence due to its broad applications in modern society. Many methods have been proposed to analyze the videos, but many challenging factors remain untackled. In this dissertation, four statistical modeling methods are proposed to address some challenging traffic video analysis problems under adverse illumination and weather conditions.

First, a new foreground detection method is presented to detect the foreground objects in videos. A novel Global Foreground Modeling (GFM) method, which estimates a global probability density function for the foreground and applies the Bayes decision rule …


Multi-Vehicle Speed Estimation Algorithm Based On Real-Time Inter-Frame Tracking Technique, Ernest Kisingo, Ndyetabura Hamisi, Hashim U. Iddi, Baraka J. Maiseli Aug 2021

Multi-Vehicle Speed Estimation Algorithm Based On Real-Time Inter-Frame Tracking Technique, Ernest Kisingo, Ndyetabura Hamisi, Hashim U. Iddi, Baraka J. Maiseli

Tanzania Journal of Science

Inappropriate vehicle speeding remains a central factor that causes road accidents claiming millions of lives every year. This challenge has raised concerns for vehicle speed estimation as an attempt to promote speed enforcement methods. Traditionally, radar and lidar systems have widely been used for this purpose, despite their several shortfalls: cosine error effects, need for direct line-of-sight, and inability to simultaneously and accurately measure speed from multiple vehicles. The current work proposes an algorithm and a multi-vehicle speed estimation system in a multi-lane road environment to address multi-vehicle speed estimation shortfalls. The proposed solution exploits image processing and computer vision …


Computer Vision Applications For Autonomous Aerial Vehicles, Burak Kakillioglu Aug 2021

Computer Vision Applications For Autonomous Aerial Vehicles, Burak Kakillioglu

Dissertations - ALL

Undoubtedly, unmanned aerial vehicles (UAVs) have experienced a great leap forward over the last decade. It is not surprising anymore to see a UAV being used to accomplish a certain task, which was previously carried out by humans or a former technology. The proliferation of special vision sensors, such as depth cameras, lidar sensors and thermal cameras, and major breakthroughs in computer vision and machine learning fields accelerated the advance of UAV research and technology. However, due to certain unique challenges imposed by UAVs, such as limited payload capacity, unreliable communication link with the ground stations and data safety, UAVs …


Understanding Complex Human Activities In Videos : The Study Of Concurrent Activity Detection And Group Activity Recognition, Yi Wei Aug 2021

Understanding Complex Human Activities In Videos : The Study Of Concurrent Activity Detection And Group Activity Recognition, Yi Wei

Legacy Theses & Dissertations (2009 - 2024)

Human activity understanding, as one of the most important task in video analysis, has been studied for decades. Great efforts have been made to push the activity recognition models towards effective and efficient representation learning. However, it is difficult to define an explicit semantic organization of activities, even for human. Current activity recognition benchmarks only organize the activity labels with shallow hierarchies, which hinders the development of activity recognition system.


Material Detection With Thermal Imaging And Computer Vision: Potentials And Limitations, Jared Poe Jul 2021

Material Detection With Thermal Imaging And Computer Vision: Potentials And Limitations, Jared Poe

Graduate Theses and Dissertations

The goal of my masters thesis research is to develop an affordable and mobile infraredbased environmental sensoring system for the control of a servo motor based on material identification. While this sensing could be oriented towards different applications, my thesis is particularly interested in material detection due to the wide range of possible applications in mechanical engineering. Material detection using a thermal mobile camera could be used in manufacturing, recycling or autonomous robotics. For my research, the application that will be focused on is using this material detection to control a servo motor by identifying and sending control inputs based …


Signal Processing And Data Analysis For Real-Time Intermodal Freight Classification Through A Multimodal Sensor System., Enrique J. Sanchez Headley Jul 2021

Signal Processing And Data Analysis For Real-Time Intermodal Freight Classification Through A Multimodal Sensor System., Enrique J. Sanchez Headley

Graduate Theses and Dissertations

Identifying freight patterns in transit is a common need among commercial and municipal entities. For example, the allocation of resources among Departments of Transportation is often predicated on an understanding of freight patterns along major highways. There exist multiple sensor systems to detect and count vehicles at areas of interest. Many of these sensors are limited in their ability to detect more specific features of vehicles in traffic or are unable to perform well in adverse weather conditions. Despite this limitation, to date there is little comparative analysis among Laser Imaging and Detection and Ranging (LIDAR) sensors for freight detection …


Methods For Detecting Floodwater On Roadways From Ground Level Images, Cem Sazara Jul 2021

Methods For Detecting Floodwater On Roadways From Ground Level Images, Cem Sazara

Computational Modeling & Simulation Engineering Theses & Dissertations

Recent research and statistics show that the frequency of flooding in the world has been increasing and impacting flood-prone communities severely. This natural disaster causes significant damages to human life and properties, inundates roads, overwhelms drainage systems, and disrupts essential services and economic activities. The focus of this dissertation is to use machine learning methods to automatically detect floodwater in images from ground level in support of the frequently impacted communities. The ground level images can be retrieved from multiple sources, including the ones that are taken by mobile phone cameras as communities record the state of their flooded streets. …


Pedestrian Attribute Recognition Using Trainable Gabor Wavelets, Imran N Junejo, Naveed Ahmed, Mohammad Lataifeh Jun 2021

Pedestrian Attribute Recognition Using Trainable Gabor Wavelets, Imran N Junejo, Naveed Ahmed, Mohammad Lataifeh

All Works

Surveillance cameras are everywhere keeping an eye on pedestrians or people as they navigate through the scene. Within this context, our paper addresses the problem of pedestrian attribute recognition (PAR). This problem entails the extraction of different attributes such as age-group, clothing style, accessories, footwear style etc. This is a multi-label problem with a host of challenges even for human observers. As such, the topic has rightly attracted attention recently. In this work, we integrate trainable Gabor wavelet (TGW) layers inside a convolution neural network (CNN). Whereas other researchers have used fixed Gabor filters with the CNN, the proposed layers …


A Quantitative Validation Of Multi-Modal Image Fusion And Segmentation For Object Detection And Tracking, Nicholas Lahaye, Michael J. Garay, Brian D. Bue, Hesham El-Askary, Erik Linstead Jun 2021

A Quantitative Validation Of Multi-Modal Image Fusion And Segmentation For Object Detection And Tracking, Nicholas Lahaye, Michael J. Garay, Brian D. Bue, Hesham El-Askary, Erik Linstead

Mathematics, Physics, and Computer Science Faculty Articles and Research

In previous works, we have shown the efficacy of using Deep Belief Networks, paired with clustering, to identify distinct classes of objects within remotely sensed data via cluster analysis and qualitative analysis of the output data in comparison with reference data. In this paper, we quantitatively validate the methodology against datasets currently being generated and used within the remote sensing community, as well as show the capabilities and benefits of the data fusion methodologies used. The experiments run take the output of our unsupervised fusion and segmentation methodology and map them to various labeled datasets at different levels of global …


Projecting Your View Attentively: Monocular Road Scene Layout Estimation Via Cross-View Transformation, Weixiang Yang, Qi Li, Wenxi Liu, Yuanlong Yu, Yuexin Ma, Shengfeng He, Jia Pan Jun 2021

Projecting Your View Attentively: Monocular Road Scene Layout Estimation Via Cross-View Transformation, Weixiang Yang, Qi Li, Wenxi Liu, Yuanlong Yu, Yuexin Ma, Shengfeng He, Jia Pan

Research Collection School Of Computing and Information Systems

HD map reconstruction is crucial for autonomous driving. LiDAR-based methods are limited due to the deployed expensive sensors and time-consuming computation. Camera-based methods usually need to separately perform road segmentation and view transformation, which often causes distortion and the absence of content. To push the limits of the technology, we present a novel framework that enables reconstructing a local map formed by road layout and vehicle occupancy in the bird's-eye view given a front-view monocular image only. In particular, we propose a cross-view transformation module, which takes the constraint of cycle consistency between views into account and makes full use …


Counterfactual Zero-Shot And Open-Set Visual Recognition, Zhongqi Yue, Tan Wang, Qianru Sun, Xian-Sheng Hua, Hanwang Zhang Jun 2021

Counterfactual Zero-Shot And Open-Set Visual Recognition, Zhongqi Yue, Tan Wang, Qianru Sun, Xian-Sheng Hua, Hanwang Zhang

Research Collection School Of Computing and Information Systems

We present a novel counterfactual framework for both Zero-Shot Learning (ZSL) and Open-Set Recognition (OSR), whose common challenge is generalizing to the unseen-classes by only training on the seen-classes. Our idea stems from the observation that the generated samples for unseen-classes are often out of the true distribution, which causes severe recognition rate imbalance between the seen-class (high) and unseen-class (low). We show that the key reason is that the generation is not Counterfactual Faithful, and thus we propose a faithful one, whose generation is from the sample-specific counterfactual question: What would the sample look like, if we set its …