Open Access. Powered by Scholars. Published by Universities.®

Articles 31 - 60 of 233

Full-Text Articles in Artificial Intelligence and Robotics

Multiclass Confidence And Localization Calibration For Object Detection, Bimsara Pathiraja, Malitha Gunawardhana, Muhammad Haris Khan Aug 2023

Multiclass Confidence And Localization Calibration For Object Detection, Bimsara Pathiraja, Malitha Gunawardhana, Muhammad Haris Khan

Computer Vision Faculty Publications

Albeit achieving high predictive accuracy across many challenging computer vision problems, recent studies suggest that deep neural networks (DNNs) tend to make over-confident predictions, rendering them poorly calibrated. Most of the existing attempts for improving DNN calibration are limited to classification tasks and restricted to calibrating in-domain predictions. Surprisingly, very little to no attempts have been made in studying the calibration of object detection methods, which occupy a pivotal space in vision-based security-sensitive, and safety-critical applications. In this paper, we propose a new train-time technique for calibrating modern object detection methods. It is capable of jointly calibrating multiclass confidence and …


N-Shot Benchmarking Of Whisper On Diverse Arabic Speech Recognition, Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed Aug 2023

N-Shot Benchmarking Of Whisper On Diverse Arabic Speech Recognition, Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed

Natural Language Processing Faculty Publications

Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under diverse conditions even on languages it was evaluated on such as Arabic. In this work, we address this gap by comprehensively evaluating Whisper on several varieties of Arabic speech for the ASR task. Our evaluation covers most publicly available Arabic speech data and is performed under n-shot (zero-, few-, and full) finetuning. We also investigate the robustness of Whisper under completely novel conditions, such as in …


Vision Language Navigation With Knowledge-Driven Environmental Dreamer, Fengda Zhu, Vincent C.S. Lee, Xiaojun Chang, Xiaodan Liang Aug 2023

Vision Language Navigation With Knowledge-Driven Environmental Dreamer, Fengda Zhu, Vincent C.S. Lee, Xiaojun Chang, Xiaodan Liang

Computer Vision Faculty Publications

Vision-language navigation (VLN) requires an agent to perceive visual observation in a house scene and navigate step-by-step following natural language instruction. Due to the high cost of data annotation and data collection, current VLN datasets provide limited instruction-trajectory data samples. Learning vision-language alignment for VLN from limited data is challenging since visual observation and language instruction are both complex and diverse. Previous works only generate augmented data based on original scenes while failing to generate data samples from unseen scenes, which limits the generalization ability of the navigation agent. In this paper, we introduce the Knowledge-driven Environmental Dreamer (KED), a …


Reinforcement Learning Approach To Stochastic Vehicle Routing Problem With Correlated Demands, Zangir Iklassov, Ikboljon Sobirov, Ruben Solozabal, Martin Takac Aug 2023

Reinforcement Learning Approach To Stochastic Vehicle Routing Problem With Correlated Demands, Zangir Iklassov, Ikboljon Sobirov, Ruben Solozabal, Martin Takac

Machine Learning Faculty Publications

We present a novel end-to-end framework for solving the Vehicle Routing Problem with stochastic demands (VRPSD) using Reinforcement Learning (RL). Our formulation incorporates the correlation between stochastic demands through other observable stochastic variables, thereby offering an experimental demonstration of the theoretical premise that non-i.i.d. stochastic demands provide opportunities for improved routing solutions. Our approach bridges the gap in the application of RL to VRPSD and consists of a parameterized stochastic policy optimized using a policy gradient algorithm to generate a sequence of actions that form the solution. Our model outperforms previous state-of-the-art metaheuristics and demonstrates robustness to changes in the …


A Multi-Layer Information Dissemination Model And Interference Optimization Strategy For Communication Networks In Disaster Areas, Yuexia Zhang, Yang Hong, Mohsen Guizani, Sheng Wu, Peiying Zhang, Ruiqi Liu Aug 2023

A Multi-Layer Information Dissemination Model And Interference Optimization Strategy For Communication Networks In Disaster Areas, Yuexia Zhang, Yang Hong, Mohsen Guizani, Sheng Wu, Peiying Zhang, Ruiqi Liu

Machine Learning Faculty Publications

The communication network in disaster areas (CNDA) can disseminate the key disaster information in time and provide basic information support for decision-making and rescuing. Therefore, it is of great significance to study the information dissemination mechanism of CNDA. However, a CNDA is vulnerable to interference, which affects information dissemination and rescuing. To solve this problem, this paper established a multi-layer information dissemination model of CNDA (MMND) which models the CNDA from the perspective of degree distribution of nodes. The information dissemination process and equilibrium state in CNDA is analyzed by an improved dynamic dissemination method. Then, the effects of the …


Arabic Dysarthric Speech Recognition Using Adversarial And Signal-Based Augmentation, Massa Baali, Ibrahim Almakky, Shady Shehata, Fakhri Karray Aug 2023

Arabic Dysarthric Speech Recognition Using Adversarial And Signal-Based Augmentation, Massa Baali, Ibrahim Almakky, Shady Shehata, Fakhri Karray

Machine Learning Faculty Publications

Despite major advancements in Automatic Speech Recognition (ASR), the state-of-the-art ASR systems struggle to deal with impaired speech even with high-resource languages. In Arabic, this challenge gets amplified, with added complexities in collecting data from dysarthric speakers. In this paper, we aim to improve the performance of Arabic dysarthric automatic speech recognition through a multi-stage augmentation approach. To this effect, we first propose a signal-based approach to generate dysarthric Arabic speech from healthy Arabic speech by modifying its speed and tempo. We also propose a second stage Parallel Wave Generative (PWG) adversarial model that is trained on an English dysarthric …


S2cd: Self-Heuristic Speaker Content Disentanglement For Any-To-Any Voice Conversion, Pengfei Wei, Xiang Yin, Chunfeng Wang, Zhonghao Li, Xinghua Qu, Zhiqiang Xu, Zejun Ma Aug 2023

S2cd: Self-Heuristic Speaker Content Disentanglement For Any-To-Any Voice Conversion, Pengfei Wei, Xiang Yin, Chunfeng Wang, Zhonghao Li, Xinghua Qu, Zhiqiang Xu, Zejun Ma

Machine Learning Faculty Publications

In this paper, we propose a Self-heuristic Speaker Content Disentanglement (S2CD) model for any to any voice conversion without using any external resources, e.g., speaker labels or vectors, linguistic models, and transcriptions. S2CD is built on the disentanglement sequential variational autoencoder (DSVAE), but improves DSVAE structure at the model architecture level from three perspectives. Specifically, we develop different structures for speaker and content encoders based on their underlying static/dynamic property. We further propose a generative graph, modelled by S2CD, so as to make S2CD well mimic the multi-speaker speech generation process. Finally, we propose a self-heuristic way to introduce bias …


Fooctts: Generating Arabic Speech With Acoustic Environment For Football Commentator, Massa Baali, Ahmed Ali Aug 2023

Fooctts: Generating Arabic Speech With Acoustic Environment For Football Commentator, Massa Baali, Ahmed Ali

Machine Learning Faculty Publications

This paper presents FOOCTTS, an automatic pipeline for a football commentator that generates speech with background crowd noise. The application gets the text from the user, applies text pre-processing such as vowelization, followed by the commentator's speech synthesizer. Our pipeline included Arabic automatic speech recognition for data labeling, CTC segmentation, transcription vowelization to match speech, and fine-tuning the TTS. Our system is capable of generating speech with its acoustic environment within limited 15 minutes of football commentator recording. Our prototype is generalizable and can be easily applied to different domains and languages.


Prompt-Based Tuning Of Transformer Models For Multi-Center Medical Image Segmentation Of Head And Neck Cancer, Numan Saeed, Muhammad Ridzuan, Roba Al Majzoub, Mohammad Yaqub Jul 2023

Prompt-Based Tuning Of Transformer Models For Multi-Center Medical Image Segmentation Of Head And Neck Cancer, Numan Saeed, Muhammad Ridzuan, Roba Al Majzoub, Mohammad Yaqub

Computer Vision Faculty Publications

Medical image segmentation is a vital healthcare endeavor requiring precise and efficient models for appropriate diagnosis and treatment. Vision transformer (ViT)-based segmentation models have shown great performance in accomplishing this task. However, to build a powerful backbone, the self-attention block of ViT requires large-scale pre-training data. The present method of modifying pre-trained models entails updating all or some of the backbone parameters. This paper proposes a novel fine-tuning strategy for adapting a pretrained transformer-based segmentation model on data from a new medical center. This method introduces a small number of learnable parameters, termed prompts, into the input space (less than …


Understanding Political Polarization Using Language Models: A Dataset And Method, Samiran Gode, Supreeth Bare, Bhiksha Raj, Hyungon Yoo Jul 2023

Understanding Political Polarization Using Language Models: A Dataset And Method, Samiran Gode, Supreeth Bare, Bhiksha Raj, Hyungon Yoo

Natural Language Processing Faculty Publications

Our paper aims to analyze political polarization in US political system using language models, and thereby help candidates make an informed decision. The availability of this information will help voters understand their candidates' views on the economy, healthcare, education, and other social issues. Our main contributions are a dataset extracted from Wikipedia that spans the past 120 years and a language model-based method that helps analyze how polarized a candidate is. Our data are divided into two parts, background information and political information about a candidate, since our hypothesis is that the political views of a candidate should be based …


Enhancing Video-Based Learning Using Knowledge Tracing: Personalizing Students’ Learning Experience With Orbits, Shady Shehata, David Santandreu, Philip Purnell, Mark Thompson Jul 2023

Enhancing Video-Based Learning Using Knowledge Tracing: Personalizing Students’ Learning Experience With Orbits, Shady Shehata, David Santandreu, Philip Purnell, Mark Thompson

Natural Language Processing Faculty Publications

As the world regains its footing following the COVID-19 pandemic, academia is striving to consolidate the gains made in students’ education experience. New technologies such as video-based learning have shown some early improvement in student learning and engagement. In this paper, we present ORBITS predictive engine at YOURIKA company, a video-based student support platform powered by knowledge tracing. In an exploratory case study of one master’s level Speech Processing course at the Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI) in Abu Dhabi, half the students used the system while the other half did not. Student qualitative feedback was universally …


Towards Enabling Haptic Communications Over 6g: Issues And Challenges, Muhammad Awais, Fasih Ullah Khan, Muhammad Zafar, Muhammad Mudassar, Muhammad Zaigham Zaheer, Khalid Mehmood Cheema, Muhammad Kamran, Woo Sung Jung Jul 2023

Towards Enabling Haptic Communications Over 6g: Issues And Challenges, Muhammad Awais, Fasih Ullah Khan, Muhammad Zafar, Muhammad Mudassar, Muhammad Zaigham Zaheer, Khalid Mehmood Cheema, Muhammad Kamran, Woo Sung Jung

Computer Vision Faculty Publications

This research paper provides a comprehensive overview of the challenges and potential solutions related to enabling haptic communication over the Tactile Internet in the context of 6G networks. The increasing demand for multimedia services and device proliferation has resulted in limited radio resources, posing challenges in their efficient allocation for Device-to-Device (D2D)-assisted haptic communications. Achieving ultra-low latency, security, and energy efficiency are crucial requirements for enabling haptic communication over TI. The paper explores various methodologies, technologies, and frameworks that can facilitate haptic communication, including backscatter communications (BsC), non-orthogonal multiple access (NOMA), and software-defined networks. Additionally, it discusses the potential of …


Target-Based Offensive Language Identification, Marcos Zampieri, Skye Morgan, Kai North, Tharindu Ranasinghe, Austin Simmons, Paridhi Khandelwal, Sara Rosenthal, Preslav Nakov Jul 2023

Target-Based Offensive Language Identification, Marcos Zampieri, Skye Morgan, Kai North, Tharindu Ranasinghe, Austin Simmons, Paridhi Khandelwal, Sara Rosenthal, Preslav Nakov

Natural Language Processing Faculty Publications

We present TBO, a new dataset for Target-based Offensive language identification. TBO contains post-level annotations regarding the harmfulness of an offensive post and token-level annotations comprising of the target and the offensive argument expression. Popular offensive language identification datasets for social media focus on annotation taxonomies only at the post level and more recently, some datasets have been released that feature only token-level annotations. TBO is an important resource that bridges the gap between post-level and token-level annotation datasets by introducing a single comprehensive unified annotation taxonomy. We use the TBO taxonomy to annotate post-level and token-level offensive language on …


Analysis Of Predictive Performance And Reliability Of Classifiers For Quality Assessment Of Medical Evidence Revealed Important Variation By Medical Area, Simon Šuster, Timothy Baldwin, Karin Verspoor Jul 2023

Analysis Of Predictive Performance And Reliability Of Classifiers For Quality Assessment Of Medical Evidence Revealed Important Variation By Medical Area, Simon Šuster, Timothy Baldwin, Karin Verspoor

Natural Language Processing Faculty Publications

Objectives: A major obstacle in deployment of models for automated quality assessment is their reliability. To analyze their calibration and selective classification performance. Study Design and Setting: We examine two systems for assessing the quality of medical evidence, EvidenceGRADEr and RobotReviewer, both developed from Cochrane Database of Systematic Reviews (CDSR) to measure strength of bodies of evidence and risk of bias (RoB) of individual studies, respectively. We report their calibration error and Brier scores, present their reliability diagrams, and analyze the risk–coverage trade-off in selective classification. Results: The models are reasonably well calibrated on most quality criteria (expected calibration error …


Bertastic At Semeval-2023 Task 3: Fine-Tuning Pretrained Multilingual Transformers – Does Order Matter?, Tarek Mahmoud, Preslav Nakov Jul 2023

Bertastic At Semeval-2023 Task 3: Fine-Tuning Pretrained Multilingual Transformers – Does Order Matter?, Tarek Mahmoud, Preslav Nakov

Natural Language Processing Faculty Publications

The naïve approach for fine-tuning pretrained deep learning models on downstream tasks involves feeding them mini-batches of randomly sampled data. In this paper, we propose a more elaborate method for fine-tuning Pretrained Multilingual Transformers (PMTs) on multilingual data. Inspired by the success of curriculum learning approaches, we investigate the significance of fine-tuning PMTs on multilingual data in a sequential fashion language by language. Unlike the curriculum learning paradigm where the model is presented with increasingly complex examples, we do not adopt a notion of “easy” and “hard” samples. Instead, our experiments draw insight from psychological findings on how the human …


Linear Classifier: An Often-Forgotten Baseline For Text Classification, Yu Chen Lin, Si An Chen, Jie Jyun Liu, Chih Jen Lin Jul 2023

Linear Classifier: An Often-Forgotten Baseline For Text Classification, Yu Chen Lin, Si An Chen, Jie Jyun Liu, Chih Jen Lin

Machine Learning Faculty Publications

Large-scale pre-trained language models such as BERT are popular solutions for text classification. Due to the superior performance of these advanced methods, nowadays, people often directly train them for a few epochs and deploy the obtained model. In this opinion paper, we point out that this way may only sometimes get satisfactory results. We argue the importance of running a simple baseline like linear classifiers on bag-of-words features along with advanced methods. First, for many text data, linear methods show competitive performance, high efficiency, and robustness. Second, advanced models such as BERT may only achieve the best results if properly …


Team Thesyllogist At Semeval-2023 Task 3: Language-Agnostic Framing Detection In Multi-Lingual Online News: A Zero-Shot Transfer Approach, Osama Mohammed Afzal, Preslav Nakov Jul 2023

Team Thesyllogist At Semeval-2023 Task 3: Language-Agnostic Framing Detection In Multi-Lingual Online News: A Zero-Shot Transfer Approach, Osama Mohammed Afzal, Preslav Nakov

Natural Language Processing Faculty Publications

We describe our system for SemEval-2022 Task 3 subtask 2 which on detecting the frames used in a news article in a multi-lingual setup. We propose a multi-lingual approach based on machine translation of the input, followed by an English prediction model. Our system demonstrated good zero-shot transfer capability, achieving micro-F1 scores of 53% for Greek (4th on the leaderboard) and 56.1% for Georgian (3rd on the leaderboard), without any prior training on translated data for these languages. Moreover, our system achieved comparable performance on seven other languages, including German, English, French, Russian, Italian, Polish, and Spanish. Our results demonstrate …


Semeval-2023 Task 3: Detecting The Category, The Framing, And The Persuasion Techniques In Online News In A Multi-Lingual Setup, Jakub Piskorski, Nicolas Stefanovitch, Giovanni Da San Martino, Preslav Nakov Jul 2023

Semeval-2023 Task 3: Detecting The Category, The Framing, And The Persuasion Techniques In Online News In A Multi-Lingual Setup, Jakub Piskorski, Nicolas Stefanovitch, Giovanni Da San Martino, Preslav Nakov

Natural Language Processing Faculty Publications

We describe SemEval-2023 task 3 on Detecting the Category, the Framing, and the Persuasion Techniques in Online News in a Multilingual Setup: the dataset, the task organization process, the evaluation setup, the results, and the participating systems. The task focused on news articles in nine languages (six known to the participants upfront: English, French, German, Italian, Polish, and Russian), and three additional ones revealed to the participants at the testing phase: Spanish, Greek, and Georgian). The task featured three subtasks: (1) determining the genre of the article (opinion, reporting, or satire), (2) identifying one or more frames used in an …


Multilingual Multifaceted Understanding Of Online News In Terms Of Genre, Framing And Persuasion Techniques, Jakub Piskorski, Nicolas Stefanovitch, Nikolaos Nikolaidis, Giovanni Da San Martino, Preslav Nakov Jul 2023

Multilingual Multifaceted Understanding Of Online News In Terms Of Genre, Framing And Persuasion Techniques, Jakub Piskorski, Nicolas Stefanovitch, Nikolaos Nikolaidis, Giovanni Da San Martino, Preslav Nakov

Natural Language Processing Faculty Publications

We present a new multilingual multifacet dataset of news articles, each annotated for genre (objective news reporting vs. opinion vs. satire), framing (what key aspects are highlighted), and persuasion techniques (logical fallacies, emotional appeals, ad hominem attacks, etc.). The persuasion techniques are annotated at the span level, using a taxonomy of 23 fine-grained techniques grouped into 6 coarse categories. The dataset contains 1,612 news articles covering recent news on current topics of public interest in six European languages (English, French, German, Italian, Polish, and Russian), with more than 37k annotated spans of persuasion techniques. We describe the dataset and the …


Can You Answer This? - Exploring Zero-Shot Qa Generalization Capabilities In Large Language Models, Saptarshi Sengupta, Shreya Ghosh, Preslav Nakov, Prasenjit Mitra Jun 2023

Can You Answer This? - Exploring Zero-Shot Qa Generalization Capabilities In Large Language Models, Saptarshi Sengupta, Shreya Ghosh, Preslav Nakov, Prasenjit Mitra

Natural Language Processing Faculty Publications

The buzz around Transformer-based Language Models (TLMs) such as BERT, RoBERTa, etc. is well-founded owing to their impressive results on an array of tasks. However, when applied to areas needing specialized knowledge (closed-domain), such as medical, finance, etc. their performance takes drastic hits, sometimes more than their older recurrent/convolutional counterparts. In this paper, we explore zero-shot capabilities of large language models for extractive Question Answering. Our objective is to examine the performance change in the face of domain drift, i.e., when the target domain data is vastly different in semantic and statistical properties from the source domain, in an attempt …


Adversarial Alignment For Source Free Object Detection, Qiaosong Chu, Shuyan Li, Guangyi Chen, Kai Li, Xiu Li Jun 2023

Adversarial Alignment For Source Free Object Detection, Qiaosong Chu, Shuyan Li, Guangyi Chen, Kai Li, Xiu Li

Machine Learning Faculty Publications

Source-free object detection (SFOD) aims to transfer a detector pre-trained on a label-rich source domain to an unlabeled target domain without seeing source data. While most existing SFOD methods generate pseudo labels via a source-pretrained model to guide training, these pseudo labels usually contain high noises due to heavy domain discrepancy. In order to obtain better pseudo supervisions, we divide the target domain into source-similar and source-dissimilar parts and align them in the feature space by adversarial learning. Specifically, we design a detection variance-based criterion to divide the target domain. This criterion is motivated by a finding that larger detection …


Corruption-Tolerant Algorithms For Generalized Linear Models, Bhaskar Mukhoty, Debojyoti Dey, Purushottam Kar Jun 2023

Corruption-Tolerant Algorithms For Generalized Linear Models, Bhaskar Mukhoty, Debojyoti Dey, Purushottam Kar

Machine Learning Faculty Publications

This paper presents SVAM (Sequential Variance-Altered MLE), a unified framework for learning generalized linear models under adversarial label corruption in training data. SVAM extends to tasks such as least squares regression, logistic regression, and gamma regression, whereas many existing works on learning with label corruptions focus only on least squares regression. SVAM is based on a novel variance reduction technique that may be of independent interest and works by iteratively solving weighted MLEs over variance-altered versions of the GLM objective. SVAM offers provable model recovery guarantees superior to the state-of-the-art for robust regression even when a constant fraction of training …


Graphprompt: Graph-Based Prompt Templates For Biomedical Synonym Prediction, Hanwen Xu, Jiayou Zhang, Zhirui Wang, Shizhuo Zhang, Megh Bhalerao, Yucong Liu, Dawei Zhu, Sheng Wang Jun 2023

Graphprompt: Graph-Based Prompt Templates For Biomedical Synonym Prediction, Hanwen Xu, Jiayou Zhang, Zhirui Wang, Shizhuo Zhang, Megh Bhalerao, Yucong Liu, Dawei Zhu, Sheng Wang

Computer Vision Faculty Publications

In the expansion of biomedical dataset, the same category may be labeled with different terms, thus being tedious and onerous to curate these terms. Therefore, automatically mapping synonymous terms onto the ontologies is desirable, which we name as biomedical synonym prediction task. Unlike biomedical concept normalization (BCN), no clues from context can be used to enhance synonym prediction, making it essential to extract graph features from ontology. We introduce an expert-curated dataset OBO-syn encompassing 70 different types of concepts and 2 million curated concept-term pairs for evaluating synonym prediction methods. We find BCN methods perform weakly on this task for …


Stability-Based Generalization Analysis For Mixtures Of Pointwise And Pairwise Learning, Jiahuan Wang, Jun Chen, Hong Chen, Bin Gu, Weifu Li, Xin Tang Jun 2023

Stability-Based Generalization Analysis For Mixtures Of Pointwise And Pairwise Learning, Jiahuan Wang, Jun Chen, Hong Chen, Bin Gu, Weifu Li, Xin Tang

Machine Learning Faculty Publications

Recently, some mixture algorithms of pointwise and pairwise learning (PPL) have been formulated by employing the hybrid error metric of “pointwise loss + pairwise loss” and have shown empirical effectiveness on feature selection, ranking and recommendation tasks. However, to the best of our knowledge, the learning theory foundation of PPL has not been touched in the existing works. In this paper, we try to fill this theoretical gap by investigating the generalization properties of PPL. After extending the definitions of algorithmic stability to the PPL setting, we establish the high-probability generalization bounds for uniformly stable PPL algorithms. Moreover, explicit convergence …


Class-Independent Regularization For Learning With Noisy Labels, Rumeng Yi, Dayan Guan, Yaping Huang, Shijian Lu Jun 2023

Class-Independent Regularization For Learning With Noisy Labels, Rumeng Yi, Dayan Guan, Yaping Huang, Shijian Lu

Computer Vision Faculty Publications

Training deep neural networks (DNNs) with noisy labels often leads to poorly generalized models as DNNs tend to memorize the noisy labels in training. Various strategies have been developed for improving sample selection precision and mitigating the noisy label memorization issue. However, most existing works adopt a class-dependent softmax classifier that is vulnerable to noisy labels by entangling the classification of multi-class features. This paper presents a class-independent regularization (CIR) method that can effectively alleviate the negative impact of noisy labels in DNN training. CIR regularizes the class-dependent softmax classifier by introducing multi-binary classifiers each of which takes care of …


Fine-Tuned Clip Models Are Efficient Video Learners, Hanoona Rasheed, Muhammad Uzair Khattak, Muhammad Maaz, Salman Khan, Fahad Shahbaz Khan Jun 2023

Fine-Tuned Clip Models Are Efficient Video Learners, Hanoona Rasheed, Muhammad Uzair Khattak, Muhammad Maaz, Salman Khan, Fahad Shahbaz Khan

Computer Vision Faculty Publications

Large-scale multi-modal training with image-text pairs imparts strong generalization to CLIP model. Since training on a similar scale for videos is infeasible, recent approaches focus on the effective transfer of image-based CLIP to the video domain. In this pursuit, new parametric modules are added to learn temporal information and inter-frame relationships which require meticulous design efforts. Furthermore, when the resulting models are learned on videos, they tend to overfit on the given task distribution and lack in generalization aspect. This begs the following question: How to effectively transfer image-level CLIP representations to videos? In this work, we show that a …


Person Image Synthesis Via Denoising Diffusion Model, Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, Fahad Shahbaz Khan Jun 2023

Person Image Synthesis Via Denoising Diffusion Model, Ankan Kumar Bhunia, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, Fahad Shahbaz Khan

Computer Vision Faculty Publications

The pose-guided person image generation task requires synthesizing photorealistic images of humans in arbitrary poses. The existing approaches use generative adversarial networks that do not necessarily maintain realistic textures or need dense correspondences that struggle to handle complex deformations and severe occlusions. In this work, we show how denoising diffusion models can be applied for high-fidelity person image synthesis with strong sample diversity and enhanced mode coverage of the learnt data distribution. Our proposed Person Image Diffusion Model (PIDM) disintegrates the complex transfer problem into a series of simpler forward-backward denoising steps. This helps in learning plausible source-to-target transformation trajectories …


Digital Twin Haptic Robotic Arms: Towards Handshakes In The Metaverse, Mohd Faisal, Fedwa Laamarti, Abdulmotaleb El Saddik Jun 2023

Digital Twin Haptic Robotic Arms: Towards Handshakes In The Metaverse, Mohd Faisal, Fedwa Laamarti, Abdulmotaleb El Saddik

Computer Vision Faculty Publications

More daily interactions are happening in the digital world of the metaverse. Providing individuals with means to perform a handshake during these interactions can enhance the overall user experience. In this paper, we put forward the design and implementation of two right-handed underactuated Digital Twin robotic arms to mediate the physical handshake interaction between two individuals. This allows them to perform a handshake while they are in separate locations. The experimental findings are very promising as our evaluation shows that the participants were highly interested in using our system to shake hands with their loved ones when they are physically …


Joint Flood Risks In The Grand River Watershed, Poornima Unnikrishnan, Kumaraswamy Ponnambalam, Nirupama Agrawal, Fakhri Karray Jun 2023

Joint Flood Risks In The Grand River Watershed, Poornima Unnikrishnan, Kumaraswamy Ponnambalam, Nirupama Agrawal, Fakhri Karray

Machine Learning Faculty Publications

According to the World Meteorological Organization, since 2000, there has been an increase in global flood-related disasters by 134 percent compared to the previous decades. Efficient flood risk management strategies necessitate a holistic approach to evaluating flood vulnerabilities and risks. Catastrophic losses can occur when the peak flow values in the rivers in a basin coincide. Therefore, estimating the joint flood risks in a region is vital, especially when frequent occurrences of extreme events are experienced. This study focuses on estimating the joint flood risks due to river flow extremes in the Grand River watershed in Canada. For this purpose, …


Suitability Of Sdn And Mec To Facilitate Digital Twin Communication Over Lte-A, Hikmat Adhami, Mohammad Alja'afreh, Mohamed Hoda, Jiaqi Zhao, Yong Zhou, Abdulmotaleb Elsaddik Jun 2023

Suitability Of Sdn And Mec To Facilitate Digital Twin Communication Over Lte-A, Hikmat Adhami, Mohammad Alja'afreh, Mohamed Hoda, Jiaqi Zhao, Yong Zhou, Abdulmotaleb Elsaddik

Computer Vision Faculty Publications

Haptic is the modality that complements traditional multimedia, i.e., audiovisual, to evolve the next wave of innovation at which the Internet data stream can be exchanged to enable remote skills and control applications. This will require ultra-low latency and ultra-high reliability to evolve the mobile experience into the era of Digital Twin and Tactile Internet. While the 5th generation of mobile networks is not yet widely deployed, Long-Term Evolution (LTE-A) latency remains much higher than the 1 ms requirement for the Tactile Internet and therefore the Digital Twin. This work investigates an interesting solution based on the incorporation of Software-defined …