Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
- Keyword
-
- Deep learning (27)
- Computer Vision and Pattern Recognition (cs.CV) (25)
- Computer vision (22)
- Machine Learning (cs.LG) (21)
- Object detection (14)
-
- Computational linguistics (13)
- Learning systems (12)
- Image and Video Processing (eess.IV) (11)
- Machine learning (11)
- Artificial Intelligence (cs.AI) (10)
- Artificial intelligence (10)
- Semantics (10)
- Benchmarking (8)
- Performance (8)
- COVID-19 (7)
- Convolutional neural networks (7)
- Image segmentation (7)
- Large dataset (7)
- Learn+ (7)
- Medical imaging (7)
- Optimization (7)
- Classification (of information) (6)
- Digital twin (6)
- Reinforcement learning (6)
- Transformers (6)
- Computational modeling (5)
- Computerized tomography (5)
- Convolution (5)
- Deep neural networks (5)
- Object recognition (5)
- Publication Year
Articles 211 - 233 of 233
Full-Text Articles in Artificial Intelligence and Robotics
Notmad: Estimating Bayesian Networks With Sample-Specific Structures And Parameters, Benjamin Lengerich, Caleb Ellington, Bryon Aragam, Eric P. Xing, Manolis Kellis
Notmad: Estimating Bayesian Networks With Sample-Specific Structures And Parameters, Benjamin Lengerich, Caleb Ellington, Bryon Aragam, Eric P. Xing, Manolis Kellis
Machine Learning Faculty Publications
Context-specific Bayesian networks (i.e. directed acyclic graphs, DAGs) identify context-dependent relationships between variables, but the non-convexity induced by the acyclicity requirement makes it difficult to share information between context-specific estimators (e.g. with graph generator functions). For this reason, existing methods for inferring context-specific Bayesian networks have favored breaking datasets into subsamples, limiting statistical power and resolution, and preventing the use of multidimensional and latent contexts. To overcome this challenge, we propose NOTEARS-optimized Mixtures of Archetypal DAGs (NOTMAD). NOTMAD models context-specific Bayesian networks as the output of a function which learns to mix archetypal networks according to sample context. The archetypal …
Tensor Pooling-Driven Instance Segmentation Framework For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi
Tensor Pooling-Driven Instance Segmentation Framework For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi
Computer Vision Faculty Publications
Automated systems designed for screening contraband items from the X-ray imagery are still facing difficulties with high clutter, concealment, and extreme occlusion. In this paper, we addressed this challenge using a novel multi-scale contour instance segmentation framework that effectively identifies the cluttered contraband data within the baggage X-ray scans. Unlike standard models that employ region-based or keypoint-based techniques to generate multiple boxes around objects, we propose to derive proposals according to the hierarchy of the regions defined by the contours. The proposed framework is rigorously validated on three public datasets, dubbed GDXray, SIXray, and OPIXray, where it outperforms the state-of-the-art …
Discriminative Region-Based Multi-Label Zero-Shot Learning, Sanath Narayan, Akshita Gupta, Salman Khan, Fahad Shahbaz Khan, Ling Shao, Mubarak Shah
Discriminative Region-Based Multi-Label Zero-Shot Learning, Sanath Narayan, Akshita Gupta, Salman Khan, Fahad Shahbaz Khan, Ling Shao, Mubarak Shah
Computer Vision Faculty Publications
Multi-label zero-shot learning (ZSL) is a more realistic counter-part of standard single-label ZSL since several objects can co-exist in a natural image. However, the occurrence of multiple objects complicates the reasoning and requires region-specific processing of visual features to preserve their contextual cues. We note that the best existing multi-label ZSL method takes a shared approach towards attending to region features with a common set of attention maps for all the classes. Such shared maps lead to diffused attention, which does not discriminatively focus on relevant locations when the number of classes are large. Moreover, mapping spatially-pooled visual features to …
Panoramic Learning With A Standardized Machine Learning Formalism, Zhiting Hu, Eric P. Xing
Panoramic Learning With A Standardized Machine Learning Formalism, Zhiting Hu, Eric P. Xing
Machine Learning Faculty Publications
Machine Learning (ML) is about computational methods that enable machines to learn concepts from experiences. In handling a wide variety of experiences ranging from data instances, knowledge, constraints, to rewards, adversaries, and lifelong interplay in an ever-growing spectrum of tasks, contemporary ML/AI research has resulted in a multitude of learning paradigms and methodologies. Despite the continual progresses on all different fronts, the disparate narrowly-focused methods also make standardized, composable, and reusable development of learning solutions difficult, and make it costly if possible to build AI agents that panoramically learn from all types of experiences. This paper presents a standardized ML …
Unsupervised Anomaly Instance Segmentation For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi
Unsupervised Anomaly Instance Segmentation For Baggage Threat Recognition, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi
Computer Vision Faculty Publications
Identifying potential threats concealed within the baggage is of prime concern for the security staff. Many researchers have developed frameworks that can automatically detect baggage threats from security X-ray scans. However, to the best of our knowledge, all of these frameworks require extensive training efforts on large-scale and well-annotated datasets, which are hard to procure in the real world, especially for the rarely seen contraband items. This paper presents a novel unsupervised anomaly instance segmentation framework that recognizes baggage threats, in X-ray scans, as anomalies without requiring any ground truth labels. Furthermore, thanks to its stylization capacity, the framework is …
P2v-Rcnn: Point To Voxel Feature Learning For 3d Object Detection From Point Clouds, Jiale Li, Yu Sun, Shujie Luo, Ziqi Zhu, Hang Dai, Andrey S. Krylov, Yong Ding, Ling Shao
P2v-Rcnn: Point To Voxel Feature Learning For 3d Object Detection From Point Clouds, Jiale Li, Yu Sun, Shujie Luo, Ziqi Zhu, Hang Dai, Andrey S. Krylov, Yong Ding, Ling Shao
Computer Vision Faculty Publications
The most recent 3D object detectors for point clouds rely on the coarse voxel-based representation rather than the accurate point-based representation due to a higher box recall in the voxel-based Region Proposal Network (RPN). However, the detection accuracy is severely restricted by the information loss of pose details in the voxels. Different from considering the point cloud as voxel or point representation only, we propose a point-to-voxel feature learning approach to voxelize the point cloud with both the point-wise semantic and local spatial features, which maintains the voxel-wise features to build the high-recall voxel-based RPN and also provides the accurate …
Edge Detail Analysis Of Wear Particles, Mohammad Shakeel Laghari, Ahmed Hassan, Mubashir Noman
Edge Detail Analysis Of Wear Particles, Mohammad Shakeel Laghari, Ahmed Hassan, Mubashir Noman
Computer Vision Faculty Publications
Tribology is the study of wear particles that are generated in all machines with interacting mechanical parts. Particles are separated from the surfaces due to friction and relative motion. These microscopic particles vary in certain characteristics of size, quantity, composition, and morphology. Wear particles or wear debris are categorized by six morphological attributes of shape, edge details, texture, color, size, and thickness ratio. Particles can be identified with the help of some or all of these attributes however, only edge details analysis is considered in this paper. The objective is to classify these particles in a coherent way based on …
Self-Supervised Learning For Fine-Grained Visual Categorization, Muhammad Maaz, Hanoona Abdul Rasheed, Dhanalaxmi Gaddam
Self-Supervised Learning For Fine-Grained Visual Categorization, Muhammad Maaz, Hanoona Abdul Rasheed, Dhanalaxmi Gaddam
Student Publications
Recent research in self-supervised learning (SSL) has shown its capability in learning useful semantic representations from images for classification tasks. Through our work, we study the usefulness of SSL for Fine-Grained Visual Categorization (FGVC). FGVC aims to distinguish objects of visually similar subcategories within a general category. The small inter-class, but large intra-class variations within the dataset makes it a challenging task. The limited availability of annotated labels for such fine-grained data encourages the need for SSL, where additional supervision can boost learning without the cost of extra annotations. Our baseline achieves 86.36% top-1 classification accuracy on CUB-200-2011 dataset by …
Towards Open World Object Detection, K. J. Joseph, Salman Khan, Fahad Shahbaz Khan, Vineeth N. Balasubramanian
Towards Open World Object Detection, K. J. Joseph, Salman Khan, Fahad Shahbaz Khan, Vineeth N. Balasubramanian
Computer Vision Faculty Publications
Humans have a natural instinct to identify unknown object instances in their environments. The intrinsic curiosity about these unknown instances aids in learning about them, when the corresponding knowledge is eventually available. This motivates us to propose a novel computer vision problem called: 'Open World Object Detection', where a model is tasked to: 1) identify objects that have not been introduced to it as 'unknown', without explicit supervision to do so, and 2) incrementally learn these identified unknown categories without forgetting previously learned classes, when the corresponding labels are progressively received. We formulate the problem, introduce a strong evaluation protocol …
Exploring Complementary Strengths Of Invariant And Equivariant Representations For Few-Shot Learning, Mamshad Nayeem Rizve, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah
Exploring Complementary Strengths Of Invariant And Equivariant Representations For Few-Shot Learning, Mamshad Nayeem Rizve, Salman Khan, Fahad Shahbaz Khan, Mubarak Shah
Computer Vision Faculty Publications
In many real-world problems, collecting a large number of labeled samples is infeasible. Few-shot learning (FSL) is the dominant approach to address this issue, where the objective is to quickly adapt to novel categories in presence of a limited number of samples. FSL tasks have been predominantly solved by leveraging the ideas from gradient-based meta-learning and metric learning approaches. However, recent works have demonstrated the significance of powerful feature representations with a simple embedding network that can outperform existing sophisticated FSL algorithms. In this work, we build on this insight and propose a novel training mechanism that simultaneously enforces equivariance …
Learning To Fuse Asymmetric Feature Maps In Siamese Trackers, Wencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao, Jianbing Shen
Learning To Fuse Asymmetric Feature Maps In Siamese Trackers, Wencheng Han, Xingping Dong, Fahad Shahbaz Khan, Ling Shao, Jianbing Shen
Computer Vision Faculty Publications
Recently, Siamese-based trackers have achieved promising performance in visual tracking. Most recent Siamese-based trackers typically employ a depth-wise cross-correlation (DW-XCorr) to obtain multi-channel correlation information from the two feature maps (target and search region). However, DW-XCorr has several limitations within Siamese-based tracking: it can easily be fooled by distractors, has fewer activated channels and provides weak discrimination of object boundaries. Further, DW-XCorr is a handcrafted parameter-free module and cannot fully benefit from offline learning on large-scale data. We propose a learnable module, called the asymmetric convolution (ACM), which learns to better capture the semantic correlation information in offline training on …
Deep Gaussian Processes For Few-Shot Segmentation, Joakim Johnander, Johan Edstedt, Martin Danelljan, Michael Felsberg, Fahad Shahbaz Khan
Deep Gaussian Processes For Few-Shot Segmentation, Joakim Johnander, Johan Edstedt, Martin Danelljan, Michael Felsberg, Fahad Shahbaz Khan
Computer Vision Faculty Publications
Few-shot segmentation is a challenging task, requiring the extraction of a generalizable representation from only a few annotated samples, in order to segment novel query images. A common approach is to model each class with a single prototype. While conceptually simple, these methods suffer when the target appearance distribution is multi-modal or not linearly separable in feature space. To tackle this issue, we propose a few-shot learner formulation based on Gaussian process (GP) regression. Through the expressivity of the GP, our approach is capable of modeling complex appearance distributions in the deep feature space. The GP provides a principled way …
On Generating Transferable Targeted Perturbations, Muzammal Naseer, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Fatih Porikli
On Generating Transferable Targeted Perturbations, Muzammal Naseer, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, Fatih Porikli
Computer Vision Faculty Publications
While the untargeted black-box transferability of adversarial perturbations has been extensively studied before, changing an unseen model's decisions to a specific 'targeted' class remains a challenging feat. In this paper, we propose a new generative approach for highly transferable targeted perturbations (TTP). We note that the existing methods are less suitable for this task due to their reliance on class-boundary information that changes from one model to another, thus reducing transferability. In contrast, our approach matches the perturbed image 'distribution' with that of the target class, leading to high targeted transferability rates. To this end, we propose a new objective …
Orthogonal Projection Loss, Kanchana Ranasinghe, Muzammal Naseer, Munawar Hayat, Salman Khan, Fahad Shahbaz Khan
Orthogonal Projection Loss, Kanchana Ranasinghe, Muzammal Naseer, Munawar Hayat, Salman Khan, Fahad Shahbaz Khan
Computer Vision Faculty Publications
Deep neural networks have achieved remarkable performance on a range of classification tasks, with softmax cross-entropy (CE) loss emerging as the de-facto objective function. The CE loss encourages features of a class to have a higher projection score on the true class-vector compared to the negative classes. However, this is a relative constraint and does not explicitly force different class features to be well-separated. Motivated by the observation that ground-truth class representations in CE loss are orthogonal (one-hot encoded vectors), we develop a novel loss function termed 'Orthogonal Projection Loss' (OPL) which imposes orthogonality in the feature space. OPL augments …
Efficient Cnn Building Blocks For Encrypted Data, Nayna Jain, Karthik Nandakumar, Nalini K. Ratha, Sharath U. Pankanti, Uttam Kumar
Efficient Cnn Building Blocks For Encrypted Data, Nayna Jain, Karthik Nandakumar, Nalini K. Ratha, Sharath U. Pankanti, Uttam Kumar
Computer Vision Faculty Publications
Machine learning on encrypted data can address the concerns related to privacy and legality of sharing sensitive data with untrustworthy service providers, while leveraging their resources to facilitate extraction of valuable insights from otherwise non-shareable data. Fully Homomorphic Encryption (FHE) is a promising technique to enable machine learning and inferencing while providing strict guarantees against information leakage. Since deep convolutional neural networks (CNNs) have become the machine learning tool of choice in several applications, several attempts have been made to harness CNNs to extract insights from encrypted data. However, existing works focus only on ensuring data security and ignore security …
Generative Multi-Label Zero-Shot Learning, Akshita Gupta, Sanath Narayan, Salman Khan, Fahad Shahbaz Khan, Ling Shao, Joost Van De Weijer
Generative Multi-Label Zero-Shot Learning, Akshita Gupta, Sanath Narayan, Salman Khan, Fahad Shahbaz Khan, Ling Shao, Joost Van De Weijer
Computer Vision Faculty Publications
Multi-label zero-shot learning strives to classify images into multiple unseen categories for which no data is available during training. The test samples can additionally contain seen categories in the generalized variant. Existing approaches rely on learning either shared or label-specific attention from the seen classes. Nevertheless, computing reliable attention maps for unseen classes during inference in a multi-label setting is still a challenge. In contrast, state-of-the-art single-label generative adversarial network (GAN) based approaches learn to directly synthesize the class-specific visual features from the corresponding class attribute embeddings. However, synthesizing multi-label features from GANs is still unexplored in the context of …
Molecule Optimization By Explainable Evolution, Binghong Chen, Tianzhe Wang, Chengtao Li, Hanjun Dai, Le Song
Molecule Optimization By Explainable Evolution, Binghong Chen, Tianzhe Wang, Chengtao Li, Hanjun Dai, Le Song
Machine Learning Faculty Publications
Optimizing molecules for desired properties is a fundamental yet challenging task in chemistry, material science, and drug discovery. This paper develops a novel algorithm for optimizing molecular properties via an Expectation-Maximization (EM) like explainable evolutionary process. The algorithm is designed to mimic human experts in the process of searching for desirable molecules and alternate between two stages: the first stage on explainable local search which identifies rationales, i.e., critical subgraph patterns accounting for desired molecular properties, and the second stage on molecule completion which explores the larger space of molecules containing good rationales. We test our approach against various baselines …
Low Light Image Enhancement Via Global And Local Context Modeling, Aditya Arora, Muhammad Haris, Syed Waqas Zamir, Munawar Hayat, Fahad Shahbaz Khan, Ling Shao, Ming-Hsuan Yang
Low Light Image Enhancement Via Global And Local Context Modeling, Aditya Arora, Muhammad Haris, Syed Waqas Zamir, Munawar Hayat, Fahad Shahbaz Khan, Ling Shao, Ming-Hsuan Yang
Computer Vision Faculty Publications
Images captured under low-light conditions manifest poor visibility, lack contrast and color vividness. Compared to conventional approaches, deep convolutional neural networks (CNNs) perform well in enhancing images. However, being solely reliant on confined fixed primitives to model dependencies, existing data-driven deep models do not exploit the contexts at various spatial scales to address low-light image enhancement. These contexts can be crucial towards inferring several image enhancement tasks, e.g., local and global contrast, brightness and color corrections; which requires cues from both local and global spatial extent. To this end, we introduce a context-aware deep network for low-light image enhancement. First, …
Inexact Tensor Methods And Their Application To Stochastic Convex Optimization, Artem Agafonov, Dmitry Kamzolov, Pavel Dvurechensky, Alexander Gasnikov, Martin Takac
Inexact Tensor Methods And Their Application To Stochastic Convex Optimization, Artem Agafonov, Dmitry Kamzolov, Pavel Dvurechensky, Alexander Gasnikov, Martin Takac
Machine Learning Faculty Publications
We propose general non-accelerated and accelerated tensor methods under inexact information on the derivatives of the objective, analyze their convergence rate. Further, we provide conditions for the inexactness in each derivative that is sufficient for each algorithm to achieve a desired accuracy. As a corollary, we propose stochastic tensor methods for convex optimization and obtain sufficient mini-batch sizes for each derivative. © 2020, CC BY.
Human Parsing Based Texture Transfer From Single Image To 3d Human Via Cross-View Consistency, Fang Zhao, Shengcai Liao, Kaihao Zhang, Ling Shao
Human Parsing Based Texture Transfer From Single Image To 3d Human Via Cross-View Consistency, Fang Zhao, Shengcai Liao, Kaihao Zhang, Ling Shao
Machine Learning Faculty Publications
This paper proposes a human parsing based texture transfer model via cross-view consistency learning to generate the texture of 3D human body from a single image. We use the semantic parsing of human body as input for providing both the shape and pose information to reduce the appearance variation of human image and preserve the spatial distribution of semantic parts. Meanwhile, in order to improve the prediction for textures of invisible parts, we explicitly enforce the consistency across different views of the same subject by exchanging the textures predicted by two views to render images during training. The perceptual loss …
Trainable Structure Tensors For Autonomous Baggage Threat Detection Under Extreme Occlusion, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi
Trainable Structure Tensors For Autonomous Baggage Threat Detection Under Extreme Occlusion, Taimur Hassan, Samet Akçay, Mohammed Bennamoun, Salman Khan, Naoufel Werghi
Computer Vision Faculty Publications
Detecting baggage threats is one of the most difficult tasks, even for expert officers. Many researchers have developed computer-aided screening systems to recognize these threats from the baggage X-ray scans. However, all of these frameworks are limited in identifying the contraband items under extreme occlusion. This paper presents a novel instance segmentation framework that utilizes trainable structure tensors to highlight the contours of the occluded and cluttered contraband items (by scanning multiple predominant orientations), while simultaneously suppressing the irrelevant baggage content. The proposed framework has been extensively tested on four publicly available X-ray datasets where it outperforms the state-of-the-art frameworks …
Learning To Learn Kernels With Variational Random Features, Xiantong Zhen, Haoliang Sun, Yingjun Du, Jun Xu, Yilong Yin, Ling Shao, Cees Snoek
Learning To Learn Kernels With Variational Random Features, Xiantong Zhen, Haoliang Sun, Yingjun Du, Jun Xu, Yilong Yin, Ling Shao, Cees Snoek
Machine Learning Faculty Publications
We introduce kernels with random Fourier features in the meta-learning framework for few-shot learning. We propose meta variational random features (MetaVRF) to learn adaptive kernels for the base-learner, which is developed in a latent variable model by treating the random feature basis as the latent variable. We formulate the optimization of MetaVRF as a variational inference problem by deriving an evidence lower bound under the meta-learning framework. To incorporate shared knowledge from related tasks, we propose a context inference of the posterior, which is established by an LSTM architecture. The LSTMbased inference network effectively integrates the context information of previous …
Any-Shot Object Detection, Shafin Rahman, Salman Khan, Nick Barnes, Fahad Shahbaz Khan
Any-Shot Object Detection, Shafin Rahman, Salman Khan, Nick Barnes, Fahad Shahbaz Khan
Computer Vision Faculty Publications
Previous work on novel object detection considers zero or few-shot settings where none or few examples of each category are available for training. In real world scenarios, it is less practical to expect that ‘all’ the novel classes are either unseen or have few-examples. Here, we propose a more realistic setting termed ‘Any-shot detection’, where totally unseen and few-shot categories can simultaneously co-occur during inference. Any-shot detection offers unique challenges compared to conventional novel object detection such as, a high imbalance between unseen, few-shot and seen object classes, susceptibility to forget base-training while learning novel classes and distinguishing novel classes …