Open Access. Powered by Scholars. Published by Universities.®

Articles 61 - 90 of 233

Full-Text Articles in Artificial Intelligence and Robotics

Handling Realistic Label Noise In Bert Text Classification, Maha Tufail Agro, Hanan Al Darmaki May 2023

Handling Realistic Label Noise In Bert Text Classification, Maha Tufail Agro, Hanan Al Darmaki

Natural Language Processing Faculty Publications

Label noise refers to errors in training labels caused by cheap data annotation methods, such as web scraping or crowd-sourcing, which can be detrimental to the performance of supervised classifiers. Several methods have been proposed to counteract the effect of random label noise in supervised classification, and some studies have shown that BERT is already robust against high rates of randomly injected label noise. However, real label noise is not random; rather, it is often correlated with input features or other annotator-specific factors. In this paper, we evaluate BERT in the presence of two types of realistic label noise: feature-dependent …


Transformer-Based Feature Fusion Approach For Multimodal Visual Sentiment Recognition Using Tweets In The Wild, Fatimah Alzamzami, Abdulmotaleb El Saddik May 2023

Transformer-Based Feature Fusion Approach For Multimodal Visual Sentiment Recognition Using Tweets In The Wild, Fatimah Alzamzami, Abdulmotaleb El Saddik

Computer Vision Faculty Publications

We present an image-based real-time sentiment analysis system that can be used to recognize in-the-wild sentiment expressions on online social networks. The system deploys the newly proposed transformer architecture on online social networks (OSN) big data to extract emotion and sentiment features using three types of images: images containing faces, images containing text, and images containing no faces/text. We build three separate models, one for each type of image, and then fuse all the models to learn the online sentiment behavior. Our proposed methodology combines a supervised two-stage training approach and threshold-moving method, which is crucial for the data imbalance …


Fair Enough: Standardizing Evaluation And Model Selection For Fairness Research In Nlp, Xudong Han, Timothy Baldwin, Trevor Cohn May 2023

Fair Enough: Standardizing Evaluation And Model Selection For Fairness Research In Nlp, Xudong Han, Timothy Baldwin, Trevor Cohn

Natural Language Processing Faculty Publications

Modern NLP systems exhibit a range of biases, which a growing literature on model debiasing attempts to correct. However, current progress is hampered by a plurality of definitions of bias, means of quantification, and oftentimes vague relation between debiasing algorithms and theoretical measures of bias. This paper seeks to clarify the current situation and plot a course for meaningful progress in fair learning, with two key contributions: (1) making clear inter-relations among the current gamut of methods, and their relation to fairness theory; and (2) addressing the practical problem of model selection, which involves a trade-off between fairness and accuracy …


Tc-Net: A Modest & Lightweight Emotion Recognition System Using Temporal Convolution Network, Muhammad Ishaq, Mustaqeem Khan, Soonil Kwon Apr 2023

Tc-Net: A Modest & Lightweight Emotion Recognition System Using Temporal Convolution Network, Muhammad Ishaq, Mustaqeem Khan, Soonil Kwon

Computer Vision Faculty Publications

Speech signals play an essential role in communication and provide an efficient way to exchange information between humans and machines. Speech Emotion Recognition (SER) is one of the critical sources for human evaluation, which is applicable in many real-world applications such as healthcare, call centers, robotics, safety, and virtual reality. This work developed a novel TCN-based emotion recognition system using speech signals through a spatial-temporal convolution network to recognize the speaker's emotional state. The authors designed a Temporal Convolutional Network (TCN) core block to recognize long-term dependencies in speech signals and then feed these temporal cues to a dense network …


On The Accelerated Noise-Tolerant Power Method, Zhiqiang Xu Apr 2023

On The Accelerated Noise-Tolerant Power Method, Zhiqiang Xu

Machine Learning Faculty Publications

We revisit the acceleration of the noise-tolerant power method for which, despite previous studies, the results remain unsatisfactory as they are either wrong or suboptimal, also lacking generality. In this work, we present a simple yet general and optimal analysis via noise-corrupted Chebyshev polynomials, which allows a larger iteration rank p than the target rank k, requires less noise conditions in a new form, and achieves the optimal iteration complexity (Equation presented) for some q satisfying k ≤ q ≤ p in a certain regime of the momentum parameter. Interestingly, it shows dynamic dependence of the noise tolerance on the …


Ten Years After Imagenet: A 360° Perspective On Artificial Intelligence, Sanjay Chawla, Preslav Nakov, Ahmed Ali, Wendy Hall, Issa Khalil, Xiaosong Ma, Husrev Taha Sencar, Ingmar Weber, Michael Wooldridge, Ting Yu Mar 2023

Ten Years After Imagenet: A 360° Perspective On Artificial Intelligence, Sanjay Chawla, Preslav Nakov, Ahmed Ali, Wendy Hall, Issa Khalil, Xiaosong Ma, Husrev Taha Sencar, Ingmar Weber, Michael Wooldridge, Ting Yu

Natural Language Processing Faculty Publications

It is 10 years since neural networks made their spectacular comeback. Prompted by this anniversary, we take a holistic perspective on artificial intelligence (AI). Supervised learning for cognitive tasks is effectively solved - provided we have enough high-quality labelled data. However, deep neural network models are not easily interpretable, and thus the debate between blackbox and whitebox modelling has come to the fore. The rise of attention networks, self-supervised learning, generative modelling and graph neural networks has widened the application space of AI. Deep learning has also propelled the return of reinforcement learning as a core building block of autonomous …


Towards Carbon Neutrality: Prediction Of Wave Energy Based On Improved Gru In Maritime Transportation, Zhihan Lv, Nana Wang, Ranran Lou, Yajun Tian, Mohsen Guizani Feb 2023

Towards Carbon Neutrality: Prediction Of Wave Energy Based On Improved Gru In Maritime Transportation, Zhihan Lv, Nana Wang, Ranran Lou, Yajun Tian, Mohsen Guizani

Machine Learning Faculty Publications

Efficient use of renewable energy is one of the critical measures to achieve carbon neutrality. Countries have introduced policies to put carbon neutrality on the agenda to achieve relatively zero emissions of greenhouse gases and to cope with the crisis brought about by global warming. This work analyzes the wave energy with high energy density and wide distribution based on understanding of various renewable energy sources. This study provides a wave energy prediction model for energy harvesting. At the same time, the Gated Recurrent Unit network (GRU), Bayesian optimization algorithm, and attention mechanism are introduced to improve the model's performance. …


Uncertaintyfusenet: Robust Uncertainty-Aware Hierarchical Feature Fusion Model With Ensemble Monte Carlo Dropout For Covid-19 Detection, Moloud Abdar, Soorena Salari, Sina Qahremani, Hak-Keung Lam, Fakhreddine (Fakhri) Karray, Sadiq Hussain, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, Saeid Nahavandi Feb 2023

Uncertaintyfusenet: Robust Uncertainty-Aware Hierarchical Feature Fusion Model With Ensemble Monte Carlo Dropout For Covid-19 Detection, Moloud Abdar, Soorena Salari, Sina Qahremani, Hak-Keung Lam, Fakhreddine (Fakhri) Karray, Sadiq Hussain, Abbas Khosravi, U. Rajendra Acharya, Vladimir Makarenkov, Saeid Nahavandi

Machine Learning Faculty Publications

The COVID-19 (Coronavirus disease 2019) pandemic has become a major global threat to human health and well-being Thus, the development of computer-aided detection (CAD) systems that are capable to accurately distinguish COVID-19 from other diseases using chest computed tomography (CT) and X-ray data is of immediate priority Such automatic systems are usually based on traditional machine learning or deep learning methods Differently from most of existing studies, which used either CT scan or X-ray images in COVID-19-case classification, we present a simple but efficient deep learning feature fusion model, called UncertaintyFuseNet, which is able to classify accurately large datasets of …


Arl-Wavelet-Bpf Optimization Using Pso Algorithm For Bearing Fault Diagnosis, Muhammad Ahsan, Dariusz Bismor, Muhammad Arslan Manzoor Jan 2023

Arl-Wavelet-Bpf Optimization Using Pso Algorithm For Bearing Fault Diagnosis, Muhammad Ahsan, Dariusz Bismor, Muhammad Arslan Manzoor

Computer Vision Faculty Publications

Rotating element bearings are the backbone of every rotating machine. Vibration signals measured from these bearings are used to diagnose the health of the machine, but when the signal-to-noise ratio is low, it is challenging to diagnose the fault frequency. In this paper, a new method is proposed to enhance the signal-to-noise ratio by applying the Asymmetric Real Laplace wavelet Bandpass Filter (ARL-wavelet-BPF). The Gaussian function of the ARL-wavelet represents an excellent BPF with smooth edges which helps to minimize the ripple effects. The bandwidth and center frequency of the ARL-wavelet-BPF are optimized using the Particle Swarm Optimization (PSO) algorithm. …


Digital Twin Of Atmospheric Environment: Sensory Data Fusion For High-Resolution Pm2.5 Estimation And Action Policies Recommendation, Kudaibergen Abutalip, Anas Al-Lahham, Abdulmotaleb Elsaddik Jan 2023

Digital Twin Of Atmospheric Environment: Sensory Data Fusion For High-Resolution Pm2.5 Estimation And Action Policies Recommendation, Kudaibergen Abutalip, Anas Al-Lahham, Abdulmotaleb Elsaddik

Computer Vision Faculty Publications

Particulate matter smaller than 2.5 microns (PM2.5) is one of the main pollutants that has considerable detrimental effects on human health. Estimating its concentration levels with ground monitors is inefficient for several reasons. In this study, we build a digital twin (DT) of an atmospheric environment by fusing remote sensing and observational data. Integral part of DT pipeline is a presence of feedback that can influence future input data. Estimated values of PM2.5 obtained from an ensemble of Random Forest and Gradient Boosting are used to provide recommendations for decreasing the agglomeration levels. A simple optimization problem is formulated for …


Channel-Resilient Deep-Learning-Driven Device Fingerprinting Through Multiple Data Streams, Nora Basha, Bechir Hamdaoui, Kathiravetpillai Sivanesan, Mohsen Guizani Jan 2023

Channel-Resilient Deep-Learning-Driven Device Fingerprinting Through Multiple Data Streams, Nora Basha, Bechir Hamdaoui, Kathiravetpillai Sivanesan, Mohsen Guizani

Machine Learning Faculty Publications

Enabling accurate and automated identification of wireless devices is critical for allowing network access monitoring and ensuring data authentication for large-scale IoT networks. RF fingerprinting has emerged as a solution for device identification by leveraging the transmitters' inevitable hardware impairments that occur during manufacturing. Although deep learning is proven efficient in classifying devices based on hardware impairments, the performance of deep learning models suffers greatly from variations of the wireless channel conditions, across time and space. To the best of our knowledge, we are the first to propose leveraging MIMO capabilities to mitigate the channel effect and provide a channel-resilient …


Self-Omics: A Self-Supervised Learning Framework For Multi-Omics Cancer Data, Sayed Hashim, Karthik Nandakumar, Mohammad Yaqub Jan 2023

Self-Omics: A Self-Supervised Learning Framework For Multi-Omics Cancer Data, Sayed Hashim, Karthik Nandakumar, Mohammad Yaqub

Computer Vision Faculty Publications

We have gained access to vast amounts of multi-omics data thanks to Next Generation Sequencing. However, it is challenging to analyse this data due to its high dimensionality and much of it not being annotated. Lack of annotated data is a significant problem in machine learning, and Self-Supervised Learning (SSL) methods are typically used to deal with limited labelled data. However, there is a lack of studies that use SSL methods to exploit inter-omics relationships on unlabelled multi-omics data. In this work, we develop a novel and efficient pre-training paradigm that consists of various SSL components, including but not limited …


Digital Twin For Railway: A Comprehensive Survey, Sara Ghaboura, Rahatara Ferdousi, Fedwa Laamarti, Chunsheng Yang, Abdulmotaleb El Saddik Jan 2023

Digital Twin For Railway: A Comprehensive Survey, Sara Ghaboura, Rahatara Ferdousi, Fedwa Laamarti, Chunsheng Yang, Abdulmotaleb El Saddik

Computer Vision Faculty Publications

Digital transformation has been prioritized in the railway industry to bring automation to railway operations. Digital Twin (DT) technology has recently gained attention in the railway industry to fulfill this goal. Contemporary researchers argue that DT can be advantageous in Railway manufacturing logistics to planning and scheduling. Although underlying technologies of DT, e.g., modelling, computer vision, and the Internet of Things, have been studied for various railway industry applications, the DT has been least explored in the context of railways. Thus, in this paper, we aim to understand the state-of-the-art of DT for railway (DTR), for advanced railway systems. Besides, …


Maple: Multi-Modal Prompt Learning, Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, Fahad Shahbaz Khan Jan 2023

Maple: Multi-Modal Prompt Learning, Muhammad Uzair Khattak, Hanoona Rasheed, Muhammad Maaz, Salman Khan, Fahad Shahbaz Khan

Computer Vision Faculty Publications

Pre-trained vision-language (V-L) models such as CLIP have shown excellent generalization ability to downstream tasks. However, they are sensitive to the choice of input text prompts and require careful selection of prompt templates to perform well. Inspired by the Natural Language Processing (NLP) literature, recent CLIP adaptation approaches learn prompts as the textual inputs to fine-tune CLIP for downstream tasks. We note that using prompting to adapt representations in a single branch of CLIP (language or vision) is sub-optimal since it does not allow the flexibility to dynamically adjust both representation spaces on a downstream task. In this work, we …


Differentially Private Stochastic Convex Optimization In (Non)-Euclidean Space Revisited, Jinyan Su, Changhong Zhao, Di Wang Jan 2023

Differentially Private Stochastic Convex Optimization In (Non)-Euclidean Space Revisited, Jinyan Su, Changhong Zhao, Di Wang

Machine Learning Faculty Publications

In this paper, we revisit the problem of Differentially Private Stochastic Convex Optimization (DP-SCO) in Euclidean and general `dp spaces. Specifically, we focus on three settings that are still far from well understood: (1) DP-SCO over a constrained and bounded (convex) set in Euclidean space; (2) unconstrained DP-SCO in `dp space; (3) DP-SCO with heavy-tailed data over a constrained and bounded set in `dp space. For problem (1), for both convex and strongly convex loss functions, we propose methods whose outputs could achieve (expected) excess population risks that are only dependent on the Gaussian width of the constraint set, rather …


A Hybrid Artificial Intelligence Model For Detecting Keratoconus, Zaid Abdi Alkareem Alyasseri, Ali H. Al-Timemy, Ammar Kamal Abasi, Alexandru Lavric, Husam Jasim Mohammed, Hidenori Takahashi, Jose Arthur Milhomens Filho, Mauro Campos, Rossen M. Hazarbassanov, Siamak Yousefi Dec 2022

A Hybrid Artificial Intelligence Model For Detecting Keratoconus, Zaid Abdi Alkareem Alyasseri, Ali H. Al-Timemy, Ammar Kamal Abasi, Alexandru Lavric, Husam Jasim Mohammed, Hidenori Takahashi, Jose Arthur Milhomens Filho, Mauro Campos, Rossen M. Hazarbassanov, Siamak Yousefi

Machine Learning Faculty Publications

Machine learning models have recently provided great promise in diagnosis of several ophthalmic disorders, including keratoconus (KCN). Keratoconus, a noninflammatory ectatic corneal disorder characterized by progressive cornea thinning, is challenging to detect as signs may be subtle. Several machine learning models have been proposed to detect KCN, however most of the models are supervised and thus require large well-annotated data. This paper proposes a new unsupervised model to detect KCN, based on adapted flower pollination algorithm (FPA) and the k-means algorithm. We will evaluate the proposed models using corneal data collected from 5430 eyes at different stages of KCN severity …


Towards A Machine Learning-Based Digital Twin For Non-Invasive Human Bio-Signal Fusion, Izaldein Al-Zyoud, Fedwa Laamarti, Xiaocong Ma, Diana Tobón, Abdulmotaleb Elsaddik Dec 2022

Towards A Machine Learning-Based Digital Twin For Non-Invasive Human Bio-Signal Fusion, Izaldein Al-Zyoud, Fedwa Laamarti, Xiaocong Ma, Diana Tobón, Abdulmotaleb Elsaddik

Computer Vision Faculty Publications

Human bio-signal fusion is considered a critical technological solution that needs to be advanced to enable modern and secure digital health and well-being applications in the metaverse. To support such efforts, we propose a new data-driven digital twin (DT) system to fuse three human physiological bio-signals: heart rate (HR), breathing rate (BR), and blood oxygen saturation level (SpO2). To accomplish this goal, we design a computer vision technology based on the non-invasive photoplethysmography (PPG) technique to extract raw time-series bio-signal data from facial video frames. Then, we implement machine learning (ML) technology to model and measure the bio-signals. We accurately …


Polarmix: A General Data Augmentation Technique For Lidar Point Clouds, Aoran Xiao, Jiaxing Huang, Dayan Guan, Kaiwen Cui, Shijian Lu, Ling Shao Dec 2022

Polarmix: A General Data Augmentation Technique For Lidar Point Clouds, Aoran Xiao, Jiaxing Huang, Dayan Guan, Kaiwen Cui, Shijian Lu, Ling Shao

Computer Vision Faculty Publications

LiDAR point clouds, which are usually scanned by rotating LiDAR sensors continuously, capture precise geometry of the surrounding environment and are crucial to many autonomous detection and navigation tasks. Though many 3D deep architectures have been developed, efficient collection and annotation of large amounts of point clouds remain one major challenge in the analytics and understanding of point cloud data. This paper presents PolarMix, a point cloud augmentation technique that is simple and generic but can mitigate the data constraint effectively across different perception tasks and scenarios. PolarMix enriches point cloud distributions and preserves point cloud fidelity via two cross-scan …


Towards Improving Calibration In Object Detection Under Domain Shift, Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali Dec 2022

Towards Improving Calibration In Object Detection Under Domain Shift, Muhammad Akhtar Munir, Muhammad Haris Khan, M. Saquib Sarfraz, Mohsen Ali

Computer Vision Faculty Publications

With deep neural network based solution more readily being incorporated in real-world applications, it has been pressing requirement that predictions by such models, especially in safety-critical environments, be highly accurate and well-calibrated. Although some techniques addressing DNN calibration have been proposed, they are only limited to visual classification applications and in-domain predictions. Unfortunately, very little to no attention is paid towards addressing calibration of DNN-based visual object detectors, that occupy similar space and importance in many decision making systems as their visual classification counterparts. In this work, we study the calibration of DNN-based object detection models, particularly under domain shift. …


Intermediate Prototype Mining Transformer For Few-Shot Semantic Segmentation, Yuanwei Liu, Nian Liu, Xiwen Yao, Junwei Han Dec 2022

Intermediate Prototype Mining Transformer For Few-Shot Semantic Segmentation, Yuanwei Liu, Nian Liu, Xiwen Yao, Junwei Han

Computer Vision Faculty Publications

Few-shot semantic segmentation aims to segment the target objects in query under the condition of a few annotated support images. Most previous works strive to mine more effective category information from the support to match with the corresponding objects in query. However, they all ignored the category information gap between query and support images. If the objects in them show large intra-class diversity, forcibly migrating the category information from the support to the query is ineffective. To solve this problem, we are the first to introduce an intermediate prototype for mining both deterministic category information from the support and adaptive …


Iitd At The Wanlp 2022 Shared Task: Multilingual Multi-Granularity Network For Propaganda Detection, Shubham Mittal, Preslav Nakov Dec 2022

Iitd At The Wanlp 2022 Shared Task: Multilingual Multi-Granularity Network For Propaganda Detection, Shubham Mittal, Preslav Nakov

Natural Language Processing Faculty Publications

We present our system for the two subtasks of the shared task on propaganda detection in Arabic, part of WANLP'2022. Subtask 1 is a multi-label classification problem to find the propaganda techniques used in a given tweet. Our system for this task uses XLM-R to predict probabilities for the target tweet to use each of the techniques. In addition to finding the techniques, Subtask 2 further asks to identify the textual span for each instance of each technique that is present in the tweet; the task can be modeled as a sequence tagging problem. We use a multi-granularity network with …


Hate-Clipper: Multimodal Hateful Meme Classification Based On Cross-Modal Interaction Of Clip Features, Gokul Karthik Kumar, Karthik Nandakumar Dec 2022

Hate-Clipper: Multimodal Hateful Meme Classification Based On Cross-Modal Interaction Of Clip Features, Gokul Karthik Kumar, Karthik Nandakumar

Computer Vision Faculty Publications

Hateful memes are a growing menace on social media. While the image and its corresponding text in a meme are related, they do not necessarily convey the same meaning when viewed individually. Hence, detecting hateful memes requires careful consideration of both visual and textual information. Multimodal pretraining can be beneficial for this task because it effectively captures the relationship between the image and the text by representing them in a similar feature space. Furthermore, it is essential to model the interactions between the image and text features through intermediate fusion. Most existing methods either employ multimodal pre-training or intermediate fusion, …


Predicting Publication Of Clinical Trials Using Structured And Unstructured Data: Model Development And Validation Study, Siyang Wang, Simon Šuster, Timothy Baldwin, Karin Verspoor Dec 2022

Predicting Publication Of Clinical Trials Using Structured And Unstructured Data: Model Development And Validation Study, Siyang Wang, Simon Šuster, Timothy Baldwin, Karin Verspoor

Natural Language Processing Faculty Publications

Background: Publication of registered clinical trials is a critical step in the timely dissemination of trial findings. However, a significant proportion of completed clinical trials are never published, motivating the need to analyze the factors behind success or failure to publish. This could inform study design, help regulatory decision-making, and improve resource allocation. It could also enhance our understanding of bias in the publication of trials and publication trends based on the research direction or strength of the findings. Although the publication of clinical trials has been addressed in several descriptive studies at an aggregate level, there is a lack …


Assisting The Human Fact-Checkers: Detecting All Previously Fact-Checked Claims In A Document, Shaden Shaar, Nikola Georgiev, Firoj Alam, Giovanni Da San Martino, Aisha Mohamed, Preslav Nakov Dec 2022

Assisting The Human Fact-Checkers: Detecting All Previously Fact-Checked Claims In A Document, Shaden Shaar, Nikola Georgiev, Firoj Alam, Giovanni Da San Martino, Aisha Mohamed, Preslav Nakov

Natural Language Processing Faculty Publications

Given the recent proliferation of false claims online, there has been a lot of manual fact-checking effort. As this is very time-consuming, human fact-checkers can benefit from tools that can support them and make them more efficient. Here, we focus on building a system that could provide such support. Given an input document, it aims to detect all sentences that contain a claim that can be verified by some previously fact-checked claims (from a given database). The output is a re-ranked list of the document sentences, so that those that can be verified are ranked as high as possible, together …


Overview Of The Wanlp 2022 Shared Task On Propaganda Detection In Arabic, Firoj Alam, Hamdy Mubarak, Wajdi Zaghouani, Giovanni Da San Martino, Preslav Nakov Dec 2022

Overview Of The Wanlp 2022 Shared Task On Propaganda Detection In Arabic, Firoj Alam, Hamdy Mubarak, Wajdi Zaghouani, Giovanni Da San Martino, Preslav Nakov

Natural Language Processing Faculty Publications

Propaganda is the expression of an opinion or an action by an individual or a group deliberately designed to influence the opinions or the actions of other individuals or groups with reference to predetermined ends, which is achieved by means of well-defined rhetorical and psychological devices. Propaganda techniques are commonly used in social media to manipulate or to mislead users. Thus, there has been a lot of recent research on automatic detection of propaganda techniques in text as well as in memes. However, so far the focus has been primarily on English. With the aim to bridge this language gap, …


Supervised Acoustic Embeddings And Their Transferability Across Languages, Sreepratha Ram, Hanan Aldarmaki Dec 2022

Supervised Acoustic Embeddings And Their Transferability Across Languages, Sreepratha Ram, Hanan Aldarmaki

Natural Language Processing Faculty Publications

In speech recognition, it is essential to model the phonetic content of the input signal while discarding irrelevant factors such as speaker variations and noise, which is challenging in low-resource settings. Self-supervised pretraining has been proposed as a way to improve both supervised and unsupervised speech recognition, including frame-level feature representations and Acoustic Word Embeddings (AWE) for variable-length segments. However, self-supervised models alone cannot learn perfect separation of the linguistic content as they are trained to optimize indirect objectives. In this work, we experiment with different pre-trained self-supervised features as input to AWE models and show that they work best …


Camelira: An Arabic Multi-Dialect Morphological Disambiguator, Ossama Obeid, Go Inoue, Nizar Habash Dec 2022

Camelira: An Arabic Multi-Dialect Morphological Disambiguator, Ossama Obeid, Go Inoue, Nizar Habash

Computer Vision Faculty Publications

We present Camelira, a web-based Arabic multi-dialect morphological disambiguation tool that covers four major variants of Arabic: Modern Standard Arabic, Egyptian, Gulf, and Levantine. Camelira offers a user-friendly web interface that allows researchers and language learners to explore various linguistic information, such as part-of-speech, morphological features, and lemmas. Our system also provides an option to automatically choose an appropriate dialect-specific disambiguator based on the prediction of a dialect identification component. Camelira is publicly accessible at http://camelira.camel-lab.com.


Pasta: Table-Operations Aware Fact Verification Via Sentence-Table Cloze Pre-Training, Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, Xiaoyong Du Dec 2022

Pasta: Table-Operations Aware Fact Verification Via Sentence-Table Cloze Pre-Training, Zihui Gu, Ju Fan, Nan Tang, Preslav Nakov, Xiaoman Zhao, Xiaoyong Du

Natural Language Processing Faculty Publications

Fact verification has attracted a lot of research attention recently, e.g., in journalism, marketing, and policymaking, as misinformation and disinformation online can sway one's opinion and affect one's actions. While fact-checking is a hard task in general, in many cases, false statements can be easily debunked based on analytics over tables with reliable information. Hence, table-based fact verification has recently emerged as an important and growing research area. Yet, progress has been limited due to the lack of datasets that can be used to pre-train language models (LMs) to be aware of common table operations, such as aggregating a column …


Greener: Graph Neural Networks For News Media Profiling, Panayot Panayotov, Utsav Shukla, Husrev T. Sencar, Mohamed Nabeel, Preslav Nakov Dec 2022

Greener: Graph Neural Networks For News Media Profiling, Panayot Panayotov, Utsav Shukla, Husrev T. Sencar, Mohamed Nabeel, Preslav Nakov

Natural Language Processing Faculty Publications

We study the problem of profiling news media on the Web with respect to their factuality of reporting and bias. This is an important but under-studied problem related to disinformation and “fake news” detection, but it addresses the issue at a coarser granularity compared to looking at an individual article or an individual claim. This is useful as it allows to profile entire media outlets in advance. Unlike previous work, which has focused primarily on text (e.g., on the articles published by the target website, or on the textual description in their social media profiles or in Wikipedia), here we …


Asdot: Any-Shot Data-To-Text Generation With Pretrained Language Models, Jiannan Xiang, Zhengzhong Liu, Yucheng Zhou, Eric P. Xing, Zhiting Hu Dec 2022

Asdot: Any-Shot Data-To-Text Generation With Pretrained Language Models, Jiannan Xiang, Zhengzhong Liu, Yucheng Zhou, Eric P. Xing, Zhiting Hu

Machine Learning Faculty Publications

Data-to-text generation is challenging due to the great variety of the input data in terms of domains (e.g., finance vs sports) or schemata (e.g., diverse predicates). Recent end-to-end neural methods thus require substantial training examples to learn to disambiguate and describe the data. Yet, real-world data-to-text problems often suffer from various data-scarce issues: one may have access to only a handful of or no training examples, and/or have to rely on examples in a different domain or schema. To fill this gap, we propose Any-Shot Data-to-Text (ASDOT), a new approach flexibly applicable to diverse settings by making efficient use of …