Open Access. Powered by Scholars. Published by Universities.®

Articles 61 - 90 of 98

Full-Text Articles in Artificial Intelligence and Robotics

Multimodal Multi-Head Convolutional Attention With Various Kernel Sizes For Medical Image Super-Resolution, Mariana-Iuliana Georgescu, Radu Tudor Ionescu, Andreea-Iuliana Miron, Olivian Savencu, Nicolae Verga, Nicolae-Cătălin Ristea, Fahad Shabaz Khan Apr 2022

Multimodal Multi-Head Convolutional Attention With Various Kernel Sizes For Medical Image Super-Resolution, Mariana-Iuliana Georgescu, Radu Tudor Ionescu, Andreea-Iuliana Miron, Olivian Savencu, Nicolae Verga, Nicolae-Cătălin Ristea, Fahad Shabaz Khan

Computer Vision Faculty Publications

Super-resolving medical images can help physicians in providing more accurate diagnostics. In many situations, computed tomography (CT) or magnetic resonance imaging (MRI) techniques output several scans (modes) during a single investigation, which can jointly be used (in a multimodal fashion) to further boost the quality of super-resolution results. To this end, we propose a novel multimodal multi-head convolutional attention module to super-resolve CT and MRI scans. Our attention module uses the convolution operation to perform joint spatial-channel attention on multiple concatenated input tensors, where the kernel (receptive field) size controls the reduction rate of the spatial attention and the number …


Energy-Based Latent Aligner For Incremental Learning, K.J. Joseph, Salman Khan, Fahad Shahbaz Khan, Rao Muhammad Anwer, Vineeth N. Balasubramanian Mar 2022

Energy-Based Latent Aligner For Incremental Learning, K.J. Joseph, Salman Khan, Fahad Shahbaz Khan, Rao Muhammad Anwer, Vineeth N. Balasubramanian

Computer Vision Faculty Publications

Deep learning models tend to forget their earlier knowledge while incrementally learning new tasks. This behavior emerges because the parameter updates optimized for the new tasks may not align well with the updates suitable for older tasks. The resulting latent representation mismatch causes forgetting. In this work, we propose ELI: Energy-based Latent Aligner for Incremental Learning, which first learns an energy manifold for the latent representations such that previous task latents will have low energy and the current task latents have high energy values. This learned manifold is used to counter the representational shift that happens during incremental learning. The …


Deep Learning Techniques For Diabetic Retinopathy Classification: A Survey, Mohammad Z. Atwany, Abdulwahab H. Sahyoun, Mohammad Yaqub Mar 2022

Deep Learning Techniques For Diabetic Retinopathy Classification: A Survey, Mohammad Z. Atwany, Abdulwahab H. Sahyoun, Mohammad Yaqub

Computer Vision Faculty Publications

Diabetic Retinopathy (DR) is a degenerative disease that impacts the eyes and is a consequence of Diabetes mellitus, where high blood glucose levels induce lesions on the eye retina. Diabetic Retinopathy is regarded as the leading cause of blindness for diabetic patients, especially the working-age population in developing nations. Treatment involves sustaining the patient's current grade of vision since the disease is irreversible. Early detection of Diabetic Retinopathy is crucial in order to sustain the patient's vision effectively. The main issue involved with DR detection is that the manual diagnosis process is very time, money, and effort consuming and involves …


Pseudo-Stereo For Monocular 3d Object Detection In Autonomous Driving, Yi-Nan Chen, Hang Dai, Yong Ding Mar 2022

Pseudo-Stereo For Monocular 3d Object Detection In Autonomous Driving, Yi-Nan Chen, Hang Dai, Yong Ding

Computer Vision Faculty Publications

Pseudo-LiDAR 3D detectors have made remarkable progress in monocular 3D detection by enhancing the capability of perceiving depth with depth estimation networks, and using LiDAR-based 3D detection architectures. The advanced stereo 3D detectors can also accurately localize 3D objects. The gap in image-to-image generation for stereo views is much smaller than that in image-to-LiDAR generation. Motivated by this, we propose a Pseudo-Stereo 3D detection framework with three novel virtual view generation methods, including image-level generation, feature-level generation, and feature-clone, for detecting 3D objects from a single image. Our analysis of depth-aware learning shows that the depth loss is effective in …


An Ensemble Approach For Patient Prognosis Of Head And Neck Tumor Using Multimodal Data, Numan Saeed, Roba Al Majzoub, Ikboljon Sobirov, Mohammad Yaqub Feb 2022

An Ensemble Approach For Patient Prognosis Of Head And Neck Tumor Using Multimodal Data, Numan Saeed, Roba Al Majzoub, Ikboljon Sobirov, Mohammad Yaqub

Computer Vision Faculty Publications

Accurate prognosis of a tumor can help doctors provide a proper course of treatment and, therefore, save the lives of many. Tradi-tional machine learning algorithms have been eminently useful in crafting prognostic models in the last few decades. Recently, deep learning algorithms have shown significant improvement when developing diag-nosis and prognosis solutions to different healthcare problems. However, most of these solutions rely solely on either imaging or clinical data. Utilizing patient tabular data such as demographics and patient med-ical history alongside imaging data in a multimodal approach to solve a prognosis task has started to gain more interest recently and …


Deep-Precognitive Diagnosis: Preventing Future Pandemics By Novel Disease Detection With Biologically-Inspired Conv-Fuzzy Network, Aviral Chharia, Rahul Upadhyay, Vinay Kumar, Chao Cheng, Jing Zhang, Tianyang Wang, Min Xu Feb 2022

Deep-Precognitive Diagnosis: Preventing Future Pandemics By Novel Disease Detection With Biologically-Inspired Conv-Fuzzy Network, Aviral Chharia, Rahul Upadhyay, Vinay Kumar, Chao Cheng, Jing Zhang, Tianyang Wang, Min Xu

Computer Vision Faculty Publications

Deep learning-based Computer-Aided Diagnosis has gained immense attention in recent years due to its capability to enhance diagnostic performance and elucidate complex clinical tasks. However, conventional supervised deep learning models are incapable of recognizing novel diseases that do not exist in the training dataset. Automated early-stage detection of novel infectious diseases can be vital in controlling their rapid spread. Moreover, the development of a conventional CAD model is only possible after disease outbreaks and datasets become available for training (viz. COVID-19 outbreak). Since novel diseases are unknown and cannot be included in training data, it is challenging to recognize them …


Cocoa: Context-Conditional Adaptation For Recognizing Unseen Classes In Unseen Domains, Puneet Mangla, Shivam Chandhok, Vineeth N. Balasubramanian, Fahad Shahbaz Khan Feb 2022

Cocoa: Context-Conditional Adaptation For Recognizing Unseen Classes In Unseen Domains, Puneet Mangla, Shivam Chandhok, Vineeth N. Balasubramanian, Fahad Shahbaz Khan

Computer Vision Faculty Publications

Recent progress towards designing models that can generalize to unseen domains (i.e domain generalization) or unseen classes (i.e zero-shot learning) has embarked interest towards building models that can tackle both domain-shift and semantic shift simultaneously (i.e zero-shot domain generalization). For models to generalize to unseen classes in unseen domains, it is crucial to learn feature representation that preserves class-level (domain-invariant) as well as domain-specific information. Motivated from the success of generative zero-shot approaches, we propose a feature generative framework integrated with a COntext COnditional Adaptive (COCOA) Batch-Normalization layer to seamlessly integrate class-level semantic and domain-specific information. The generated visual features …


Subomiembed: Self-Supervised Representation Learning Of Multi-Omics Data For Cancer Type Classification, Sayed Hashim, Muhammad Ali, Karthik Nandakumar, Mohammad Yaqub Feb 2022

Subomiembed: Self-Supervised Representation Learning Of Multi-Omics Data For Cancer Type Classification, Sayed Hashim, Muhammad Ali, Karthik Nandakumar, Mohammad Yaqub

Computer Vision Faculty Publications

For personalized medicines, very crucial intrinsic information is present in high dimensional omics data which is difficult to capture due to the large number of molecular features and small number of available samples. Different types of omics data show various aspects of samples. Integration and analysis of multi-omics data give us a broad view of tumours, which can improve clinical decision making. Omics data, mainly DNA methylation and gene expression profiles are usually high dimensional data with a lot of molecular features. In recent years, variational autoencoders (VAE) [13] have been extensively used in embedding image and text data into …


Hyperparameter Optimization For Covid-19 Chest X-Ray Classification, Ibraheem Hamdi, Muhammad Ridzuan, Mohammad Yaqub Jan 2022

Hyperparameter Optimization For Covid-19 Chest X-Ray Classification, Ibraheem Hamdi, Muhammad Ridzuan, Mohammad Yaqub

Computer Vision Faculty Publications

Despite the introduction of vaccines, Coronavirus disease (COVID-19) remains a worldwide dilemma, continuously developing new variants such as Delta and the recent Omicron. The current standard for testing is through polymerase chain reaction (PCR). However, PCRs can be expensive, slow, and/or inaccessible to many people. X-rays on the other hand have been readily used since the early 20th century and are relatively cheaper, quicker to obtain, and typically covered by health insurance. With a careful selection of model, hyperparameters, and augmentations, we show that it is possible to develop models with 83% accuracy in binary classification and 64% in multi-class …


Transformers In Medical Imaging: A Survey, Fahad Shamshad, Salman Khan, Syed Waqas Zamir, Muhammad Haris Khan, Munawar Hayat, Fahad Shahbaz Khan, Huazhu Fu Jan 2022

Transformers In Medical Imaging: A Survey, Fahad Shamshad, Salman Khan, Syed Waqas Zamir, Muhammad Haris Khan, Munawar Hayat, Fahad Shahbaz Khan, Huazhu Fu

Computer Vision Faculty Publications

Following unprecedented success on the natural language tasks, Transformers have been successfully applied to several computer vision problems, achieving state-of-the-art results and prompting researchers to reconsider the supremacy of convolutional neural networks (CNNs) as de facto operators. Capitalizing on these advances in computer vision, the medical imaging field has also witnessed growing interest for Transformers that can capture global context compared to CNNs with local receptive fields. Inspired from this transition, in this survey, we attempt to provide a comprehensive review of the applications of Transformers in medical imaging covering various aspects, ranging from recently proposed architectural designs to unsolved …


Transformers In Vision: A Survey, Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, Mubarak Shah Jan 2022

Transformers In Vision: A Survey, Salman Khan, Muzammal Naseer, Munawar Hayat, Syed Waqas Zamir, Fahad Shahbaz Khan, Mubarak Shah

Computer Vision Faculty Publications

Astounding results from Transformer models on natural language tasks have intrigued the vision community to study their application to computer vision problems. Among their salient benefits, Transformers enable modeling long dependencies between input sequence elements and support parallel processing of sequence as compared to recurrent networks e.g., Long short-term memory (LSTM). Different from convolutional networks, Transformers require minimal inductive biases for their design and are naturally suited as set-functions. Furthermore, the straightforward design of Transformers allows processing multiple modalities (e.g., images, videos, text and speech) using similar processing blocks and demonstrates excellent scalability to very large capacity networks and huge …


Deep Learning-Based Quality Assessment Of Clinical Protocol Adherence In Fetal Ultrasound Dating Scans, Sevim Cengiz, Mohammad Yaqub Jan 2022

Deep Learning-Based Quality Assessment Of Clinical Protocol Adherence In Fetal Ultrasound Dating Scans, Sevim Cengiz, Mohammad Yaqub

Computer Vision Faculty Publications

To assess fetal health during pregnancy, doctors use the gestational age (GA) calculation based on the Crown Rump Length (CRL) measurement in order to check for fetal size and growth trajectory. However, GA estimation based on CRL, requires proper positioning of calipers on the fetal crown and rump view, which is not always an easy plane to find, especially for an inexperienced sonographer. Finding a slightly oblique view from the true CRL view could lead to a different CRL value and therefore incorrect estimation of GA. This study presents an AI-based method for a quality assessment of the CRL view …


Automatic Segmentation Of Head And Neck Tumor: How Powerful Transformers Are?, Ikboljon Sobirov, Otabek Nazarov, Hussain Alasmawi, Mohammad Yaqub Jan 2022

Automatic Segmentation Of Head And Neck Tumor: How Powerful Transformers Are?, Ikboljon Sobirov, Otabek Nazarov, Hussain Alasmawi, Mohammad Yaqub

Computer Vision Faculty Publications

Cancer is one of the leading causes of death worldwide, and head and neck (H&N) cancer is amongst the most prevalent types. Positron emission tomography and computed tomography are used to detect and segment the tumor region. Clinically, tumor segmentation is extensively time-consuming and prone to error. Machine learning, and deep learning in particular, can assist to automate this process, yielding results as accurate as the results of a clinician. In this research study, we develop a vision transformers-based method to automatically delineate H&N tumor, and compare its results to leading convolutional neural network (CNN)-based models. We use multi-modal data …


Is Contrastive Learning Suitable For Left Ventricular Segmentation In Echocardiographic Images?, Mohamed Saeed, Rand Muhtaseb, Mohammad Yaqub Jan 2022

Is Contrastive Learning Suitable For Left Ventricular Segmentation In Echocardiographic Images?, Mohamed Saeed, Rand Muhtaseb, Mohammad Yaqub

Computer Vision Faculty Publications

Contrastive learning has proven useful in many applications where access to labelled data is limited. The lack of annotated data is particularly problematic in medical image segmenta-tion as it is difficult to have clinical experts manually annotate large volumes of data. One such task is the segmentation of cardiac structures in ultrasound images of the heart. In this paper, we argue whether or not contrastive pretraining is helpful for the segmentation of the left ventricle in echocardiography images. Furthermore, we study the effect of this on two segmentation networks, DeepLabV3, as well as the commonly used segmentation net-work, UNet. Our …


Is It Possible To Predict Mgmt Promoter Methylation From Brain Tumor Mri Scans Using Deep Learning Models?, Numan Saeed, Shahad Hardan, Kudaibergen Abutalip, Mohammad Yaqub Jan 2022

Is It Possible To Predict Mgmt Promoter Methylation From Brain Tumor Mri Scans Using Deep Learning Models?, Numan Saeed, Shahad Hardan, Kudaibergen Abutalip, Mohammad Yaqub

Computer Vision Faculty Publications

Glioblastoma is a common brain malignancy that tends to occur in older adults and is almost always lethal. The effectiveness of chemotherapy, being the standard treatment for most cancer types, can be improved if a particular genetic sequence in the tumor known as MGMT promoter is methylated. However, to identify the state of the MGMT promoter, the conventional approach is to perform a biopsy for genetic analysis, which is time and effort consuming. A couple of recent publications proposed a connection between the MGMT promoter state and the MRI scans of the tumor and hence suggested the use of deep …


Challenges In Covid-19 Chest X-Ray Classification: Problematic Data Or Ineffective Approaches?, Muhammad Ridzuan, Ameera Ali Bawazir, Ivo Gollini Navarrete, Ibrahim Almakky, Mohammad Yaqub Jan 2022

Challenges In Covid-19 Chest X-Ray Classification: Problematic Data Or Ineffective Approaches?, Muhammad Ridzuan, Ameera Ali Bawazir, Ivo Gollini Navarrete, Ibrahim Almakky, Mohammad Yaqub

Computer Vision Faculty Publications

The value of quick, accurate, and confident diagnoses cannot be undermined to mitigate the effects of COVID-19 infection, particularly for severe cases. Enormous effort has been put towards developing deep learning methods to classify and detect COVID-19 infections from chest radiography images. However, recently some questions have been raised surrounding the clinical viability and effectiveness of such methods. In this work, we carry out extensive experiments on a large COVID-19 chest X-ray dataset to investigate the challenges faced with creating reliable solutions from both the data and machine learning perspectives. Accordingly, we offer an in-depth discussion into the challenges faced …


An Attention-Based Resnet Architecture For Acute Hemorrhage Detection And Classification: Toward A Health 4.0 Digital Twin Study, Aftab Hussain, Muhammad Usman Yaseen, Muhammad Imran, Muhammad Waqar, Adnan Akhunzada, Mohammad Al-Ja'afreh, Abdulmotaleb El Saddik Jan 2022

An Attention-Based Resnet Architecture For Acute Hemorrhage Detection And Classification: Toward A Health 4.0 Digital Twin Study, Aftab Hussain, Muhammad Usman Yaseen, Muhammad Imran, Muhammad Waqar, Adnan Akhunzada, Mohammad Al-Ja'afreh, Abdulmotaleb El Saddik

Computer Vision Faculty Publications

Due to the advancement of digital twin (DT) technology, Health 4.0 applications have become reality and starting to take roots. In this article, we focus on intracranial hemorrhage (ICH) which is a life-threatening emergency that needs immediate diagnosis and treatment. ICH is caused by bleeding inside the skull or brain. Radiologists typically examine computed tomography (CT) scans of the patients to determine the ICH and its subtype. But the manual assessment of the CT scan is a complex and time-consuming task. The existing pre-trained convolutional neural network (CNN) models are state-of-the-art for ICH classification. However, they employ poor feature extraction …


Spatio-Temporal Relation Modeling For Few-Shot Action Recognition, Anirudh Thatipelli, Sanath Narayan, Salman Hameed Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, Bernard Ghanem Dec 2021

Spatio-Temporal Relation Modeling For Few-Shot Action Recognition, Anirudh Thatipelli, Sanath Narayan, Salman Hameed Khan, Rao Muhammad Anwer, Fahad Shahbaz Khan, Bernard Ghanem

Computer Vision Faculty Publications

We propose a novel few-shot action recognition framework, STRM, which enhances class-specific feature discriminability while simultaneously learning higher-order temporal representations. The focus of our approach is a novel spatio-temporal enrichment module that aggregates spatial and temporal contexts with dedicated local patch-level and global frame-level feature enrichment sub-modules. Local patch-level enrichment captures the appearance-based characteristics of actions. On the other hand, global framelevel enrichment explicitly encodes the broad temporal context, thereby capturing the relevant object features over time. The resulting spatio-temporally enriched representations are then utilized to learn the relational matching between query and support action sub-sequences. We further introduce a …


Ow-Detr: Open-World Detection Transformer, Akshita Gupta, Sanath Narayan, K.J. Joseph, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah Dec 2021

Ow-Detr: Open-World Detection Transformer, Akshita Gupta, Sanath Narayan, K.J. Joseph, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah

Computer Vision Faculty Publications

Open-world object detection (OWOD) is a challenging computer vision problem, where the task is to detect a known set of object categories while simultaneously identifying unknown objects. Additionally, the model must incrementally learn new classes that become known in the next training episodes. Distinct from standard object detection, the OWOD setting poses significant challenges for generating quality candidate proposals on potentially unknown objects, separating the unknown objects from the background and detecting diverse unknown objects. Here, we introduce a novel end-to-end transformer-based framework, OW-DETR, for open-world object detection. The proposed OW-DETR comprises three dedicated components namely, attention-driven pseudo-labeling, novelty classification …


Deep Learning Predicts Ebv Status In Gastric Cancer Based On Spatial Patterns Of Lymphocyte Infiltration, Baoyi Zhang, Kevin Yao, Min Xu, Jia Wu, Chao Cheng Nov 2021

Deep Learning Predicts Ebv Status In Gastric Cancer Based On Spatial Patterns Of Lymphocyte Infiltration, Baoyi Zhang, Kevin Yao, Min Xu, Jia Wu, Chao Cheng

Computer Vision Faculty Publications

EBV infection occurs in around 10% of gastric cancer cases and represents a distinct subtype, characterized by a unique mutation profile, hypermethylation, and overexpression of PD-L1. Moreover, EBV positive gastric cancer tends to have higher immune infiltration and a better prognosis. EBV infection status in gastric cancer is most commonly determined using PCR and in situ hybridization, but such a method requires good nucleic acid preservation. Detection of EBV status with histopathology images may complement PCR and in situ hybridization as a first step of EBV infection assessment. Here, we developed a deep learning-based algorithm to directly predict EBV infection …


Multi-Modal Transformers Excel At Class-Agnostic Object Detection, Muhammad Maaz, Hanoona Bangalath Rasheed, Salman Hameed Khan, Fahad Shahbaz Khan, Rao Muhammad Anwer, Ming-Hsuan Yang Nov 2021

Multi-Modal Transformers Excel At Class-Agnostic Object Detection, Muhammad Maaz, Hanoona Bangalath Rasheed, Salman Hameed Khan, Fahad Shahbaz Khan, Rao Muhammad Anwer, Ming-Hsuan Yang

Computer Vision Faculty Publications

What constitutes an object? This has been a longstanding question in computer vision. Towards this goal, numerous learning-free and learning-based approaches have been developed to score objectness. However, they generally do not scale well across new domains and for unseen objects. In this paper, we advocate that existing methods lack a top-down supervision signal governed by human-understandable semantics. To bridge this gap, we explore recent Multi-modal Vision Transformers (MViT) that have been trained with aligned image-text pairs. Our extensive experiments across various domains and novel objects show the state-of-the-art performance of MViTs to localize generic objects in images. Based on …


Restormer: Efficient Transformer For High-Resolution Image Restoration, Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang Nov 2021

Restormer: Efficient Transformer For High-Resolution Image Restoration, Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Ming-Hsuan Yang

Computer Vision Faculty Publications

Since convolutional neural networks (CNNs) perform well at learning generalizable image priors from large-scale data, these models have been extensively applied to image restoration and related tasks. Recently, another class of neural architectures, Transformers, have shown significant performance gains on natural language and high-level vision tasks. While the Transformer model mitigates the shortcomings of CNNs (i.e., limited receptive field and in-adaptability to input content), its computational complexity grows quadratically with the spatial resolution, therefore making it infeasible to apply to most image restoration tasks involving high-resolution images. In this work, we propose an efficient Transformer model by making several key …


Tensor Pooling-Driven Instance Segmentation Framework For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi Sep 2021

Tensor Pooling-Driven Instance Segmentation Framework For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi

Computer Vision Faculty Publications

Automated systems designed for screening contraband items from the X-ray imagery are still facing difficulties with high clutter, concealment, and extreme occlusion. In this paper, we addressed this challenge using a novel multi-scale contour instance segmentation framework that effectively identifies the cluttered contraband data within the baggage X-ray scans. Unlike standard models that employ region-based or keypoint-based techniques to generate multiple boxes around objects, we propose to derive proposals according to the hierarchy of the regions defined by the contours. The proposed framework is rigorously validated on three public datasets, dubbed GDXray, SIXray, and OPIXray, where it outperforms the state-of-the-art …


Discriminative Region-Based Multi-Label Zero-Shot Learning, Sanath Narayan, Akshita Gupta, Salman Khan, Fahad Shahbaz Khan, Ling Shao, Mubarak Shah Aug 2021

Discriminative Region-Based Multi-Label Zero-Shot Learning, Sanath Narayan, Akshita Gupta, Salman Khan, Fahad Shahbaz Khan, Ling Shao, Mubarak Shah

Computer Vision Faculty Publications

Multi-label zero-shot learning (ZSL) is a more realistic counter-part of standard single-label ZSL since several objects can co-exist in a natural image. However, the occurrence of multiple objects complicates the reasoning and requires region-specific processing of visual features to preserve their contextual cues. We note that the best existing multi-label ZSL method takes a shared approach towards attending to region features with a common set of attention maps for all the classes. Such shared maps lead to diffused attention, which does not discriminatively focus on relevant locations when the number of classes are large. Moreover, mapping spatially-pooled visual features to …


Unsupervised Anomaly Instance Segmentation For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi Jul 2021

Unsupervised Anomaly Instance Segmentation For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi

Computer Vision Faculty Publications

Identifying potential threats concealed within the baggage is of prime concern for the security staff. Many researchers have developed frameworks that can automatically detect baggage threats from security X-ray scans. However, to the best of our knowledge, all of these frameworks require extensive training efforts on large-scale and well-annotated datasets, which are hard to procure in the real world, especially for the rarely seen contraband items. This paper presents a novel unsupervised anomaly instance segmentation framework that recognizes baggage threats, in X-ray scans, as anomalies without requiring any ground truth labels. Furthermore, thanks to its stylization capacity, the framework is …


P2v-Rcnn: Point To Voxel Feature Learning For 3d Object Detection From Point Clouds, Jiale Li, Yu Sun, Shujie Luo, Ziqi Zhu, Hang Dai, Andrey S. Krylov, Yong Ding, Ling Shao Jul 2021

P2v-Rcnn: Point To Voxel Feature Learning For 3d Object Detection From Point Clouds, Jiale Li, Yu Sun, Shujie Luo, Ziqi Zhu, Hang Dai, Andrey S. Krylov, Yong Ding, Ling Shao

Computer Vision Faculty Publications

The most recent 3D object detectors for point clouds rely on the coarse voxel-based representation rather than the accurate point-based representation due to a higher box recall in the voxel-based Region Proposal Network (RPN). However, the detection accuracy is severely restricted by the information loss of pose details in the voxels. Different from considering the point cloud as voxel or point representation only, we propose a point-to-voxel feature learning approach to voxelize the point cloud with both the point-wise semantic and local spatial features, which maintains the voxel-wise features to build the high-recall voxel-based RPN and also provides the accurate …


Edge Detail Analysis Of Wear Particles, Mohammad Shakeel Laghari, Ahmed Hassan, Mubashir Noman Jul 2021

Edge Detail Analysis Of Wear Particles, Mohammad Shakeel Laghari, Ahmed Hassan, Mubashir Noman

Computer Vision Faculty Publications

Tribology is the study of wear particles that are generated in all machines with interacting mechanical parts. Particles are separated from the surfaces due to friction and relative motion. These microscopic particles vary in certain characteristics of size, quantity, composition, and morphology. Wear particles or wear debris are categorized by six morphological attributes of shape, edge details, texture, color, size, and thickness ratio. Particles can be identified with the help of some or all of these attributes however, only edge details analysis is considered in this paper. The objective is to classify these particles in a coherent way based on …


Towards Open World Object Detection, K. J. Joseph, Salman Khan, Fahad Shahbaz Khan, Vineeth N. Balasubramanian May 2021

Towards Open World Object Detection, K. J. Joseph, Salman Khan, Fahad Shahbaz Khan, Vineeth N. Balasubramanian

Computer Vision Faculty Publications

Humans have a natural instinct to identify unknown object instances in their environments. The intrinsic curiosity about these unknown instances aids in learning about them, when the corresponding knowledge is eventually available. This motivates us to propose a novel computer vision problem called: 'Open World Object Detection', where a model is tasked to: 1) identify objects that have not been introduced to it as 'unknown', without explicit supervision to do so, and 2) incrementally learn these identified unknown categories without forgetting previously learned classes, when the corresponding labels are progressively received. We formulate the problem, introduce a strong evaluation protocol …


Exploring Complementary Strengths Of Invariant And Equivariant Representations For Few-Shot Learning, Mamshad Nayeem Rizve, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah Apr 2021

Exploring Complementary Strengths Of Invariant And Equivariant Representations For Few-Shot Learning, Mamshad Nayeem Rizve, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah

Computer Vision Faculty Publications

In many real-world problems, collecting a large number of labeled samples is infeasible. Few-shot learning (FSL) is the dominant approach to address this issue, where the objective is to quickly adapt to novel categories in presence of a limited number of samples. FSL tasks have been predominantly solved by leveraging the ideas from gradient-based meta-learning and metric learning approaches. However, recent works have demonstrated the significance of powerful feature representations with a simple embedding network that can outperform existing sophisticated FSL algorithms. In this work, we build on this insight and propose a novel training mechanism that simultaneously enforces equivariance …


Learning To Fuse Asymmetric Feature Maps In Siamese Trackers, Wencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao, Jianbing Shen Mar 2021

Learning To Fuse Asymmetric Feature Maps In Siamese Trackers, Wencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao, Jianbing Shen

Computer Vision Faculty Publications

Recently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCorr has several limitations within Siamese-based tracking: it can easily be fooled by distractors, has fewer activated channels and provides weak discrimination of object boundaries. Further, DW-XCorr is a handcrafted parameter-free module and cannot fully benefit from offline learning on large-scale data. We propose a learnable module, called the asymmetric convolution (ACM), which learns to better capture the semantic correlation information in offline training on …