Open Access. Powered by Scholars. Published by Universities.®

Articles 121 - 150 of 233

Full-Text Articles in Artificial Intelligence and Robotics

Learning Hierarchical Metrical Structure Beyond Measures, Junyan Jiang, Daniel Chin, Yixiao Zhang, Gus Xia Sep 2022

Learning Hierarchical Metrical Structure Beyond Measures, Junyan Jiang, Daniel Chin, Yixiao Zhang, Gus Xia

Machine Learning Faculty Publications

Music contains hierarchical structures beyond beats and measures. While hierarchical structure annotations are helpful for music information retrieval and computer musicology, such annotations are scarce in current digital music databases. In this paper, we explore a data-driven approach to automatically extract hierarchical metrical structures from scores. We propose a new model with a Temporal Convolutional Network-Conditional Random Field (TCN-CRF) architecture. Given a symbolic music score, our model takes in an arbitrary number of voices in a beat-quantized form, and predicts a 4-level hierarchical metrical structure from downbeat-level to section-level. We also annotate a dataset using RWC-POP MIDI files to facilitate …


Led Down The Rabbit Hole: Exploring The Potential Of Global Attention For Biomedical Multi-Document Summarisation, Yulia Otmakhova, Hung Thinh Truong, Timothy Baldwin, Trevor Cohn, Karin Verspoor, Jey Han Lau Sep 2022

Led Down The Rabbit Hole: Exploring The Potential Of Global Attention For Biomedical Multi-Document Summarisation, Yulia Otmakhova, Hung Thinh Truong, Timothy Baldwin, Trevor Cohn, Karin Verspoor, Jey Han Lau

Natural Language Processing Faculty Publications

In this paper we report on our submission to the Multidocument Summarisation for Literature Review (MSLR) shared task. Specifically, we adapt PRIMERA (Xiao et al., 2022) to the biomedical domain by placing global attention on important biomedical entities in several ways. We analyse the outputs of the 23 resulting models, and report patterns in the results related to the presence of additional global attention, number of training steps, and the input configuration. © 2022, CC BY-SA.


Unsupervised Lexical Substitution With Decontextualised Embeddings, Takashi Wada, Timothy Baldwin, Yuji Matsumoto, Jey Han Lau Sep 2022

Unsupervised Lexical Substitution With Decontextualised Embeddings, Takashi Wada, Timothy Baldwin, Yuji Matsumoto, Jey Han Lau

Natural Language Processing Faculty Publications

We propose a new unsupervised method for lexical substitution using pre-trained language models. Compared to previous approaches that use the generative capability of language models to predict substitutes, our method retrieves substitutes based on the similarity of contextualised and decontextualised word embeddings, i.e. the average contextual representation of a word in multiple contexts. We conduct experiments in English and Italian, and show that our method substantially outperforms strong baselines and establishes a new state-of-the-art without any explicit supervision or fine-tuning. We further show that our method performs particularly well at predicting low-frequency substitutes, and also generates a diverse list of …


Artificial Intelligence-Driven Design Of Fuel Mixtures, Nursulu Kuzhagaliyeva, Samuel Horváth, John Williams, Andre Nicolle, S. Mani Sarathy Sep 2022

Artificial Intelligence-Driven Design Of Fuel Mixtures, Nursulu Kuzhagaliyeva, Samuel Horváth, John Williams, Andre Nicolle, S. Mani Sarathy

Machine Learning Faculty Publications

High-performance fuel design is imperative to achieve cleaner burning and high-efficiency engine systems. We introduce a data-driven artificial intelligence (AI) framework to design liquid fuels exhibiting tailor-made properties for combustion engine applications to improve efficiency and lower carbon emissions. The fuel design approach is a constrained optimization task integrating two parts: (i) a deep learning (DL) model to predict the properties of pure components and mixtures and (ii) search algorithms to efficiently navigate in the chemical space. Our approach presents the mixture-hidden vector as a linear combination of each single component’s vectors in each blend and incorporates it into the …


Beat Transformer: Demixed Beat And Downbeat Tracking With Dilated Self-Attention, Jingwei Zhao, Gus Xia, Ye Wang Sep 2022

Beat Transformer: Demixed Beat And Downbeat Tracking With Dilated Self-Attention, Jingwei Zhao, Gus Xia, Ye Wang

Machine Learning Faculty Publications

We propose Beat Transformer, a novel Transformer encoder architecture for joint beat and downbeat tracking. Different from previous models that track beats solely based on the spectrogram of an audio mixture, our model deals with demixed spectrograms with multiple instrument channels. This is inspired by the fact that humans perceive metrical structures from richer musical contexts, such as chord progression and instrumentation. To this end, we develop a Transformer model with both time-wise attention and instrument-wise attention to capture deep-buried metrical cues. Moreover, our model adopts a novel dilated self-attention mechanism, which achieves powerful hierarchical modelling with only linear complexity. …


Domain Adversarial Training On Conditional Variational Auto-Encoder For Controllable Music Generation, Jingwei Zhao, Gus Xia, Ye Wang Sep 2022

Domain Adversarial Training On Conditional Variational Auto-Encoder For Controllable Music Generation, Jingwei Zhao, Gus Xia, Ye Wang

Machine Learning Faculty Publications

The variational auto-encoder has become a leading framework for symbolic music generation, and a popular research direction is to study how to effectively control the generation process. A straightforward way is to control a model using different conditions during inference. However, in music practice, conditions are usually sequential (rather than simple categorical labels), involving rich information that overlaps with the learned representation. Consequently, the decoder gets confused about whether to “listen to” the latent representation or the condition, and sometimes just ignores the condition. To solve this problem, we leverage domain adversarial training to disentangle the representation from condition cues …


Cmr3d: Contextualized Multi-Stage Refinement For 3d Object Detection, Dhanalaxmi Gaddam, Jean Lahoud, Fahad Shahbaz Khan, Rao Anwer, Hisham Cholakkal Sep 2022

Cmr3d: Contextualized Multi-Stage Refinement For 3d Object Detection, Dhanalaxmi Gaddam, Jean Lahoud, Fahad Shahbaz Khan, Rao Anwer, Hisham Cholakkal

Computer Vision Faculty Publications

Existing deep learning-based 3D object detectors typically rely on the appearance of individual objects and do not explicitly pay attention to the rich contextual information of the scene. In this work, we propose Contextualized Multi-Stage Refinement for 3D Object Detection (CMR3D) framework, which takes a 3D scene as input and strives to explicitly integrate useful contextual information of the scene at multiple levels to predict a set of object bounding-boxes along with their corresponding semantic labels. To this end, we propose to utilize a context enhancement network that captures the contextual information at different levels of granularity followed by a …


Decentralized Personalized Federated Learning: Lower Bounds And Optimal Algorithm For All Personalization Modes, Abdurakhmon Sadiev, Ekaterina Borodich, Aleksandr Beznosikov, Darina Dvinskikh, Saveliy Chezhegov, Rachael Tappenden, Martin Takac, Alexander Gasnikov Sep 2022

Decentralized Personalized Federated Learning: Lower Bounds And Optimal Algorithm For All Personalization Modes, Abdurakhmon Sadiev, Ekaterina Borodich, Aleksandr Beznosikov, Darina Dvinskikh, Saveliy Chezhegov, Rachael Tappenden, Martin Takac, Alexander Gasnikov

Machine Learning Faculty Publications

This paper considers the problem of decentralized, personalized federated learning. For centralized personalized federated learning, a penalty that measures the deviation from the local model and its average, is often added to the objective function. However, in a decentralized setting this penalty is expensive in terms of communication costs, so here, a different penalty — one that is built to respect the structure of the underlying computational network — is used instead. We present lower bounds on the communication and local computation costs for this problem formulation and we also present provably optimal methods for decentralized personalized federated learning. Numerical …


Transformers In Remote Sensing: A Survey, Abdulaziz Amer Aleissaee, Amandeep Kumar, Rao Anwer, Salman Khan, Hisham Cholakkal, Gui-Song Xia, Fahad Shahbaz Khan Sep 2022

Transformers In Remote Sensing: A Survey, Abdulaziz Amer Aleissaee, Amandeep Kumar, Rao Anwer, Salman Khan, Hisham Cholakkal, Gui-Song Xia, Fahad Shahbaz Khan

Computer Vision Faculty Publications

Deep learning-based algorithms have seen a massive popularity in different areas of remote sensing image analysis over the past decade. Recently, transformers-based architectures, originally introduced in natural language processing, have pervaded computer vision field where the self-attention mechanism has been utilized as a replacement to the popular convolution operator for capturing long-range dependencies. Inspired by recent advances in computer vision, remote sensing community has also witnessed an increased exploration of vision transformers for a diverse set of tasks. Although a number of surveys have focused on transformers in computer vision in general, to the best of our knowledge we are …


Accomontage2: A Complete Harmonization And Accompaniment Arrangement System, Li Yi, Haochen Hu, Jingwei Zhao, Gus Xia Sep 2022

Accomontage2: A Complete Harmonization And Accompaniment Arrangement System, Li Yi, Haochen Hu, Jingwei Zhao, Gus Xia

Machine Learning Faculty Publications

We propose AccoMontage2, a system capable of doing full-length song harmonization and accompaniment arrangement based on a lead melody. Following AccoMontage, this study focuses on generating piano arrangements for popular/folk songs and it carries on the generalized template-based retrieval method. The novelties of this study are twofold. First, we invent a harmonization module (which AccoMontage does not have). This module generates structured and coherent full-length chord progression by optimizing and balancing three loss terms: a micro-level loss for note-wise dissonance, a meso-level loss for phrase-template matching, and a macro-level loss for full piece coherency. Second, we develop a graphical user …


Overview Of The Clef-2022 Checkthat! Lab Task 2 On Detecting Previously Fact-Checked Claims, Preslav Nakov, Giovanni Da San Martino, Firoj Alam, Shaden Shaar, Hamdy Mubarak, Nikolay Babulkov Sep 2022

Overview Of The Clef-2022 Checkthat! Lab Task 2 On Detecting Previously Fact-Checked Claims, Preslav Nakov, Giovanni Da San Martino, Firoj Alam, Shaden Shaar, Hamdy Mubarak, Nikolay Babulkov

Natural Language Processing Faculty Publications

We describe the fourth edition of the CheckThat! Lab, part of the 2022 Conference and Labs of the Evaluation Forum (CLEF). The lab evaluates technology supporting three tasks related to factuality, and it covers seven languages such as Arabic, Bulgarian, Dutch, English, German, Spanish, and Turkish. Here, we present the task 2, which asks to detect previously fact-checked claims (in two languages). A total of six teams participated in this task, submitted a total of 37 runs, and most submissions managed to achieve sizable improvements over the baselines using transformer based models such as BERT, RoBERTa. In this paper, we …


Negational Symmetry Of Quantum Neural Networks For Binary Pattern Classification, Nanqing Dong, Michael Kampffmeyer, Irina Voiculescu, Eric P. Xing Sep 2022

Negational Symmetry Of Quantum Neural Networks For Binary Pattern Classification, Nanqing Dong, Michael Kampffmeyer, Irina Voiculescu, Eric P. Xing

Machine Learning Faculty Publications

Although quantum neural networks (QNNs) have shown promising results in solving simple machine learning tasks recently, the behavior of QNNs in binary pattern classification is still underexplored. In this work, we find that QNNs have an Achilles’ heel in binary pattern classification. To illustrate this point, we provide a theoretical insight into the properties of QNNs by presenting and analyzing a new form of symmetry embedded in a family of QNNs with full entanglement, which we term negational symmetry. Due to negational symmetry, QNNs can not differentiate between a quantum binary signal and its negational counterpart. We empirically evaluate the …


Overview Of The Clef-2022 Checkthat! Lab Task 1 On Identifying Relevant Claims In Tweets, Preslav Nakov, Alberto Barrón-Cedeño, Giovanni Da San Martino, Firoj Alam, Mucahid Kutlu, Wajdi Zaghouani, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Hamdy Mubarak, Alex Nikolov Sep 2022

Overview Of The Clef-2022 Checkthat! Lab Task 1 On Identifying Relevant Claims In Tweets, Preslav Nakov, Alberto Barrón-Cedeño, Giovanni Da San Martino, Firoj Alam, Mucahid Kutlu, Wajdi Zaghouani, Mucahid Kutlu, Wajdi Zaghouani, Chengkai Li, Shaden Shaar, Hamdy Mubarak, Alex Nikolov

Natural Language Processing Faculty Publications

We present an overview of CheckThat! lab 2022 Task 1, part of the 2022 Conference and Labs of the Evaluation Forum (CLEF). Task 1 asked to predict which posts in a Twitter stream are worth fact-checking, focusing on COVID-19 and politics in six languages: Arabic, Bulgarian, Dutch, English, Spanish, and Turkish. A total of 19 teams participated and most submissions managed to achieve sizable improvements over the baselines using Transformer-based models such as BERT and GPT-3. Across the four subtasks, approaches that targetted multiple languages (be it individually or in conjunction, in general obtained the best performance. We describe the …


Truncated Matrix Power Iteration For Differentiable Dag Learning, Zhen Zhang, Ignavier Ng, Dong Gong, Yuhang Liu, Ehsan M. Abbasnejad, Mingming Gong, Kun Zhang, Javen Qinfeng Shi Aug 2022

Truncated Matrix Power Iteration For Differentiable Dag Learning, Zhen Zhang, Ignavier Ng, Dong Gong, Yuhang Liu, Ehsan M. Abbasnejad, Mingming Gong, Kun Zhang, Javen Qinfeng Shi

Machine Learning Faculty Publications

Recovering underlying Directed Acyclic Graph structures (DAG) from observational data is highly challenging due to the combinatorial nature of the DAG-constrained optimization problem. Recently, DAG learning has been cast as a continuous optimization problem by characterizing the DAG constraint as a smooth equality one, generally based on polynomials over adjacency matrices. Existing methods place very small coefficients on high-order polynomial terms for stabilization, since they argue that large coefficients on the higher-order terms are harmful due to numeric exploding. On the contrary, we discover that large coefficients on higher-order terms are beneficial for DAG learning, when the spectral radiuses of …


Exploiting Higher-Order Derivatives In Convex Optimization Methods, Dmitry Kamzolov, Alexander Gasnikov, Pavel Dvurechensky, Artem Agafonov, Martin Takac Aug 2022

Exploiting Higher-Order Derivatives In Convex Optimization Methods, Dmitry Kamzolov, Alexander Gasnikov, Pavel Dvurechensky, Artem Agafonov, Martin Takac

Machine Learning Faculty Publications

Exploiting higher-order derivatives in convex optimization is known at least since 1970’s. In each iteration higher-order (also called tensor) methods minimize a regularized Taylor expansion of the objective function, which leads to faster convergence rates if the corresponding higher-order derivative is Lipschitz-continuous. Recently a series of lower iteration complexity bounds for such methods were proved, and a gap between upper an lower complexity bounds was revealed. Moreover, it was shown that such methods can be implementable since the appropriately regularized Taylor expansion of a convex function is also convex and, thus, can be minimized in polynomial time. Only very recently …


Sr-Dcsk Cooperative Communication System With Code Index Modulation: A New Design For 6g New Radios, Yi Fang, Wang Chen, Pingping Chen, Yiwei Tao, Mohsen Guizani Aug 2022

Sr-Dcsk Cooperative Communication System With Code Index Modulation: A New Design For 6g New Radios, Yi Fang, Wang Chen, Pingping Chen, Yiwei Tao, Mohsen Guizani

Machine Learning Faculty Publications

This paper proposes a high-throughput short reference differential chaos shift keying cooperative communication system with the aid of code index modulation, referred to as CIM-SR-DCSK-CC system. In the proposed CIM-SR-DCSK-CC system, the source transmits information bits to both the relay and destination in the first time slot, while the relay not only forwards the source information bits but also sends new information bits to the destination in the second time slot. To be specific, the relay employs an N-order Walsh code to carry additional log2N information bits, which are superimposed onto the SR-DCSK signal carrying the decoded source information bits. …


Interpreting Song Lyrics With An Audio-Informed Pre-Trained Language Model, Yixiao Zhang, Junyan Jiang, Gus Xia, Simon Dixon Aug 2022

Interpreting Song Lyrics With An Audio-Informed Pre-Trained Language Model, Yixiao Zhang, Junyan Jiang, Gus Xia, Simon Dixon

Machine Learning Faculty Publications

Lyric interpretations can help people understand songs and their lyrics quickly, and can also make it easier to manage, retrieve and discover songs efficiently from the growing mass of music archives. In this paper we propose BART-fusion, a novel model for generating lyric interpretations from lyrics and music audio that combines a large-scale pre-trained language model with an audio encoder. We employ a cross-modal attention module to incorporate the audio representation into the lyrics representation to help the pre-trained language model understand the song from an audio perspective, while preserving the language model’s original generative performance. We also release the …


Fdrl Approach For Association And Resource Allocation In Multi-Uav Air-To-Ground Iomt Network, Abegaz Mohammed, Aiman Erbad, Hayla Nahom, Abdullatif Albaseer, Mohammed Abdallah, Mohsen Guizani Aug 2022

Fdrl Approach For Association And Resource Allocation In Multi-Uav Air-To-Ground Iomt Network, Abegaz Mohammed, Aiman Erbad, Hayla Nahom, Abdullatif Albaseer, Mohammed Abdallah, Mohsen Guizani

Machine Learning Faculty Publications

In 6G networks, unmanned aerial vehicles (UAVs) can serve as aerial flying base stations (AFBS) with aerial mobile edge computing (AMEC) server capabilities. AFBS is an increasingly popular solution for delivering time-sensitive applications, extending network coverage, and assisting ground base stations in the healthcare systems for remote areas with limited infrastructure. Furthermore, the UAVs are deployed in the healthcare system to support the Internet of medical things (IoMT) devices in data collection, medical equipment distribution, and providing smart services. However, ensuring the privacy and security of patients’ data with the limited UAV resources is a major challenge. In this paper, …


Reconfigurable Intelligent Surfaces And Capacity Optimization: A Large System Analysis, Aris L. Moustakas, George C. Alexandropoulos, Mérouane Debbah Aug 2022

Reconfigurable Intelligent Surfaces And Capacity Optimization: A Large System Analysis, Aris L. Moustakas, George C. Alexandropoulos, Mérouane Debbah

Machine Learning Faculty Publications

Reconfigurable Intelligent Surfaces (RISs), comprising large numbers of low-cost and almost passive metamaterials with tunable reflection properties, have been recently proposed as an enabling technology for programmable wireless propagation environments. In this paper, we present asymptotic closed-form expressions for the mean and variance of the mutual information metric for a multi-antenna transmitter-receiver pair in the presence of multiple RISs, using methods from statistical physics. While nominally valid in the large system limit, we show that the derived Gaussian approximation for the mutual information can be quite accurate, even for modest-sized antenna arrays and metasurfaces. The above results are particularly useful …


Transformnet: Self-Supervised Representation Learning Through Predicting Geometric Transformations, Muhammad Ali, Sayed Hashim Aug 2022

Transformnet: Self-Supervised Representation Learning Through Predicting Geometric Transformations, Muhammad Ali, Sayed Hashim

Computer Vision Faculty Publications

Deep neural networks need a big amount of training data, while in the real world there is a scarcity of data available for training purposes. To resolve this issue unsupervised methods are used for training with limited data. In this report, we describe the unsupervised semantic feature learning approach for recognition of the geometric transformation applied to the input data. The basic concept of our approach is that if someone is unaware of the objects in the images, he/she would not be able to quantitatively predict the geometric transformation that was applied to them. This self supervised scheme is based …


Avist: A Benchmark For Visual Object Tracking In Adverse Visibility, Mubashir Noman, Wafa Al Ghallabi, Daniya Najiha, Christoph Mayer, Hisham Cholakkal, Salman Khan, Luc Van Gool, Fahad Shahbaz Khan Aug 2022

Avist: A Benchmark For Visual Object Tracking In Adverse Visibility, Mubashir Noman, Wafa Al Ghallabi, Daniya Najiha, Christoph Mayer, Hisham Cholakkal, Salman Khan, Luc Van Gool, Fahad Shahbaz Khan

Computer Vision Faculty Publications

One of the key factors behind the recent success in visual tracking is the availability of dedicated benchmarks. While being greatly benefiting to the tracking research, existing benchmarks do not pose the same difficulty as before with recent trackers achieving higher performance mainly due to (i) the introduction of more sophisticated transformers-based methods and (ii) the lack of diverse scenarios with adverse visibility such as, severe weather conditions, camouflage and imaging effects. We introduce AVisT, a dedicated benchmark for visual tracking in diverse scenarios with adverse visibility. AVisT comprises 120 challenging sequences with 80k annotated frames, spanning 18 diverse scenarios …


3d Vision With Transformers: A Survey, Jean Lahoud, Jiale Cao, Fahad Shahbaz Khan, Hisham Cholakkal, Rao Anwer, Salman Khan, Ming-Hsuan Yang Aug 2022

3d Vision With Transformers: A Survey, Jean Lahoud, Jiale Cao, Fahad Shahbaz Khan, Hisham Cholakkal, Rao Anwer, Salman Khan, Ming-Hsuan Yang

Computer Vision Faculty Publications

The success of the transformer architecture in natural language processing has recently triggered attention in the computer vision field. The transformer has been used as a replacement for the widely used convolution operators, due to its ability to learn long-range dependencies. This replacement was proven to be successful in numerous tasks, in which several state-of-the-art methods rely on transformers for better learning. In computer vision, the 3D field has also witnessed an increase in employing the transformer for 3D convolution neural networks and multi-layer perceptron networks. Although a number of surveys have focused on transformers in vision in general, 3D …


A Multi-Dimensional Matrix Pencil-Based Channel Prediction Method For Massive Mimo With Mobility, Weidong Li, Haifan Yin, Ziao Qin, Yandi Cao, Mérouane Debbah Aug 2022

A Multi-Dimensional Matrix Pencil-Based Channel Prediction Method For Massive Mimo With Mobility, Weidong Li, Haifan Yin, Ziao Qin, Yandi Cao, Mérouane Debbah

Machine Learning Faculty Publications

This paper addresses the mobility problem in massive multiple-input multiple-output systems, which leads to significant performance losses in the practical deployment of the fifth generation mobile communication networks. We propose a novel channel prediction method based on multi-dimensional matrix pencil (MDMP), which estimates the path parameters by exploiting the angular-frequency-domain and angular-timedomain structures of the wideband channel. The MDMP method also entails a novel path pairing scheme to pair the delay and Doppler, based on the super-resolution property of the angle estimation. Our method is able to deal with the realistic constraint of time-varying path delays introduced by user movements, …


Meta-Detr: Image-Level Few-Shot Detection With Inter-Class Correlation Exploitation, Gongjie Zhang, Zhipeng Luo, Kaiwen Cui, Shijian Lu, Eric P. Xing Jul 2022

Meta-Detr: Image-Level Few-Shot Detection With Inter-Class Correlation Exploitation, Gongjie Zhang, Zhipeng Luo, Kaiwen Cui, Shijian Lu, Eric P. Xing

Machine Learning Faculty Publications

Few-shot object detection has been extensively investigated by incorporating meta-learning into region-based detection frameworks. Despite its success, the said paradigm is still constrained by several factors, such as (i) low-quality region proposals for novel classes and (ii) negligence of the inter-class correlation among different classes. Such limitations hinder the generalization of base-class knowledge for the detection of novel-class objects. In this work, we design Meta-DETR, which (i) is the first image-level few-shot detector, and (ii) introduces a novel inter-class correlational meta-learning strategy to capture and leverage the correlation among different classes for robust and accurate few-shot object detection. Meta-DETR works …


Semantic-Aligned Matching For Enhanced Detr Convergence And Multi-Scale Feature Fusion, Gongjie Zhang, Zhipeng Luo, Yingchen Yu, Jiaxing Huang, Kaiwen Cui, Shijian Lu, Eric Xing Jul 2022

Semantic-Aligned Matching For Enhanced Detr Convergence And Multi-Scale Feature Fusion, Gongjie Zhang, Zhipeng Luo, Yingchen Yu, Jiaxing Huang, Kaiwen Cui, Shijian Lu, Eric Xing

Machine Learning Faculty Publications

The recently proposed DEtection TRansformer (DETR) has established a fully end-to-end paradigm for object detection. However, DETR suffers from slow training convergence, which hinders its applicability to various detection tasks. We observe that DETR's slow convergence is largely attributed to the difficulty in matching object queries to relevant regions due to the unaligned semantics between object queries and encoded image features. With this observation, we design Semantic-Aligned-Matching DETR++ (SAM-DETR++) to accelerate DETR's convergence and improve detection performance. The core of SAM-DETR++ is a plug-andplay module that projects object queries and encoded image features into the same feature embedding space, where …


Towards Smart City Security: Violence And Weaponized Violence Detection Using Dcnn, Toluwani Aremu, Li Zhiyuan, Reem Alameeri, Moayad Aloqaily, Mohsen Guizani Jul 2022

Towards Smart City Security: Violence And Weaponized Violence Detection Using Dcnn, Toluwani Aremu, Li Zhiyuan, Reem Alameeri, Moayad Aloqaily, Mohsen Guizani

Machine Learning Faculty Publications

In this ever connected society, CCTVs have had a pivotal role in enforcing safety and security of the citizens by recording unlawful activities for the authorities to take actions. In a smart city context, using Deep Convolutional Neural Networks (DCNN) to detection violence and weaponized violence from CCTV videos will provide an additional layer of security by ensuring real-time detection around the clock. In this work, we introduced a new specialised dataset by gathering real CCTV footage of both weaponized and non-weaponized violence as well as non-violence videos from YouTube. We also proposed a novel approach in merging consecutive video …


Self-Distilled Vision Transformer For Domain Generalization, Maryam Sultana, Muzammal Naseer, Muhammad Haris Khan, Salman Khan, Fahad Shahbaz Khan Jul 2022

Self-Distilled Vision Transformer For Domain Generalization, Maryam Sultana, Muzammal Naseer, Muhammad Haris Khan, Salman Khan, Fahad Shahbaz Khan

Computer Vision Faculty Publications

In recent past, several domain generalization (DG) methods have been proposed, showing encouraging performance, however, almost all of them build on convolutional neural networks (CNNs). There is little to no progress on studying the DG performance of vision transformers (ViTs), which are challenging the supremacy of CNNs on standard benchmarks, often built on i.i.d assumption. This renders the real-world deployment of ViTs doubtful. In this paper, we attempt to explore ViTs towards addressing the DG problem. Similar to CNNs, ViTs also struggle in out-of-distribution scenarios and the main culprit is overfitting to source domains. Inspired by the modular architecture of …


Green, Quantized Federated Learning Over Wireless Networks: An Energy-Efficient Design, Minsu Kim, Walid Saad, Mohammad Mozaffari, Mérouane Debbah Jul 2022

Green, Quantized Federated Learning Over Wireless Networks: An Energy-Efficient Design, Minsu Kim, Walid Saad, Mohammad Mozaffari, Mérouane Debbah

Machine Learning Faculty Publications

The practical deployment of federated learning (FL) over wireless networks requires balancing energy efficiency and convergence time due to the limited available resources of devices. Prior art on FL often trains deep neural networks (DNNs) to achieve high accuracy and fast convergence using 32 bits of precision level. However, such scenarios will be impractical for resource-constrained devices since DNNs typically have high computational complexity and memory requirements. Thus, there is a need to reduce the precision level in DNNs to reduce the energy expenditure. In this paper, a green-quantized FL framework, which represents data with a finite precision level in …


Adversarial Pixel Restoration As A Pretext Task For Transferable Perturbations, Hashmat Shadab Malik, Shahina K. Kunhimon, Muzammal Nasser, Salman Khan, Fahad Shahbaz Khan Jul 2022

Adversarial Pixel Restoration As A Pretext Task For Transferable Perturbations, Hashmat Shadab Malik, Shahina K. Kunhimon, Muzammal Nasser, Salman Khan, Fahad Shahbaz Khan

Computer Vision Faculty Publications

Transferable adversarial attacks optimize adversaries from a pretrained surrogate model and known label space to fool the unknown black-box models. Therefore, these attacks are restricted by the availability of an effective surrogate model. In this work, we relax this assumption and propose Adversarial Pixel Restoration as a self-supervised alternative to train an effective surrogate model from scratch under the condition of no labels and few data samples. Our training approach is based on a min-max objective which reduces overfitting via an adversarial objective and thus optimizes for a more generalizable surrogate model. Our proposed attack is complimentary to our adversarial …


Robustar: Interactive Toolbox Supporting Precise Data Annotation For Robust Vision Learning, Chonghan Chen, Haohan Wang, Leyang Hu, Yuhao Zhang, Shuguang Lyu, Jingcheng Wu, Xinnuo Li, Linjing Sun, Eric Xing Jul 2022

Robustar: Interactive Toolbox Supporting Precise Data Annotation For Robust Vision Learning, Chonghan Chen, Haohan Wang, Leyang Hu, Yuhao Zhang, Shuguang Lyu, Jingcheng Wu, Xinnuo Li, Linjing Sun, Eric Xing

Machine Learning Faculty Publications

We introduce the initial release of our software Robustar, which aims to improve the robustness of vision classification machine learning models through a data-driven perspective. Building upon the recent understanding that the lack of machine learning model’s robustness is the tendency of the model’s learning of spurious features, we aim to solve this problem from its root at the data perspective by removing the spurious features from the data before training. In particular, we introduce a software that helps the users to better prepare the data for training image classification models by allowing the users to annotate the spurious features …