Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (5391)
- Computer Engineering (4370)
- Operations Research, Systems Engineering and Industrial Engineering (4244)
- Numerical Analysis and Scientific Computing (4163)
- Systems Science (3895)
-
- Social and Behavioral Sciences (982)
- Databases and Information Systems (629)
- Medicine and Health Sciences (605)
- Data Science (527)
- Theory and Algorithms (485)
- Business (447)
- Graphics and Human Computer Interfaces (414)
- Electrical and Computer Engineering (402)
- Education (382)
- Software Engineering (373)
- Arts and Humanities (337)
- Public Affairs, Public Policy and Public Administration (297)
- Information Security (293)
- Other Computer Sciences (274)
- Life Sciences (270)
- Law (241)
- Statistics and Probability (203)
- Medical Specialties (179)
- Library and Information Science (165)
- Psychology (157)
- Robotics (157)
- Programming Languages and Compilers (154)
- Institution
-
- China Simulation Federation (3880)
- Singapore Management University (1897)
- Old Dominion University (641)
- San Jose State University (277)
- MBZUAI (233)
-
- City University of New York (CUNY) (184)
- Technological University Dublin (157)
- Air Force Institute of Technology (137)
- Chapman University (125)
- California Polytechnic State University, San Luis Obispo (116)
- Chinese Academy of Sciences (113)
- University of Arkansas, Fayetteville (102)
- Lindenwood University (97)
- Edith Cowan University (92)
- Embry-Riddle Aeronautical University (92)
- University of Nebraska - Lincoln (78)
- University of Kentucky (76)
- University of South Florida (71)
- University of Nevada, Las Vegas (63)
- Dartmouth College (62)
- Clemson University (60)
- University of Denver (59)
- University of Michigan Law School (57)
- Utah State University (57)
- The Texas Medical Center Library (54)
- Thomas Jefferson University (54)
- New Jersey Institute of Technology (53)
- University of Malaya (50)
- Purdue University (48)
- Missouri University of Science and Technology (47)
- Keyword
-
- Artificial intelligence (778)
- Machine learning (685)
- Deep learning (435)
- Artificial Intelligence (359)
- Machine Learning (359)
-
- AI (239)
- Deep Learning (201)
- Simulation (160)
- Computer vision (157)
- Reinforcement learning (140)
- Generative AI (134)
- Neural networks (128)
- Large language models (109)
- Natural language processing (108)
- Robotics (97)
- Natural Language Processing (90)
- ChatGPT (89)
- Path planning (89)
- Optimization (82)
- Large Language Models (77)
- Computer Vision (76)
- Classification (71)
- Neural network (67)
- Neural Networks (65)
- Virtual reality (64)
- Reinforcement Learning (63)
- Computer Science (59)
- Cybersecurity (59)
- Genetic algorithm (58)
- Algorithms (57)
- Publication Year
- Publication
-
- Journal of System Simulation (3880)
- Research Collection School Of Computing and Information Systems (1664)
- Master's Projects (248)
- Theses and Dissertations (183)
- Computer Science Faculty Publications (125)
-
- Bulletin of Chinese Academy of Sciences (Chinese Version) (113)
- Faculty Scholarship (108)
- Publications and Research (99)
- Computer Vision Faculty Publications (98)
- Master's Theses (96)
- Conference papers (92)
- Electrical & Computer Engineering Faculty Publications (90)
- Machine Learning Faculty Publications (86)
- Electronic Theses and Dissertations (85)
- Faculty Publications (77)
- Dissertations (70)
- Research outputs 2022 to 2026 (64)
- USF Tampa Graduate Theses and Dissertations (59)
- Dissertations and Theses Collection (Open Access) (57)
- Articles (54)
- Dissertations, Theses, and Capstone Projects (53)
- Theses and Dissertations--Computer Science (48)
- Natural Language Processing Faculty Publications (46)
- Teaching and Generative AI: Pedagogical Possibilities and Productive Tensions (46)
- Graduate Theses and Dissertations (44)
- Open Access Theses & Dissertations (42)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (40)
- Theses (40)
- Electrical & Computer Engineering Theses & Dissertations (39)
- Publications (39)
- Publication Type
- File Type
Articles 4111 - 4140 of 11180
Full-Text Articles in Artificial Intelligence and Robotics
Machine Learning-Based Classification Of Chronic Traumatic Brain Injury Using Hybrid Diffusion Imaging, Jennifer Muller, Ruixuan Wang, Devon Middleton, Mahdi Alizadeh, Kichang Kang, Ryan Hryczyk, George Zabrecky, Chloe Hriso, Emily Navarreto, Nancy Wintering, Anthony J. Bazzan, Chengyuan Wu, Daniel A. Monti, Xun Jiao, Qianhong Wu, Andrew B. Newberg, Feroze Mohamed
Machine Learning-Based Classification Of Chronic Traumatic Brain Injury Using Hybrid Diffusion Imaging, Jennifer Muller, Ruixuan Wang, Devon Middleton, Mahdi Alizadeh, Kichang Kang, Ryan Hryczyk, George Zabrecky, Chloe Hriso, Emily Navarreto, Nancy Wintering, Anthony J. Bazzan, Chengyuan Wu, Daniel A. Monti, Xun Jiao, Qianhong Wu, Andrew B. Newberg, Feroze Mohamed
Marcus Institute of Integrative Health Faculty Papers
BACKGROUND AND PURPOSE: Traumatic brain injury (TBI) can cause progressive neuropathology that leads to chronic impairments, creating a need for biomarkers to detect and monitor this condition to improve outcomes. This study aimed to analyze the ability of data-driven analysis of diffusion tensor imaging (DTI) and neurite orientation dispersion imaging (NODDI) to develop biomarkers to infer symptom severity and determine whether they outperform conventional T1-weighted imaging.
MATERIALS AND METHODS: A machine learning-based model was developed using a dataset of hybrid diffusion imaging of patients with chronic traumatic brain injury. We first extracted the useful features from the hybrid diffusion imaging …
Dynamic Graph Enhanced Contrastive Learning For Chest X-Ray Report Generation, Mingjie Li, Bingqian Lin, Zicong Chen, Haokun Lin, Xiaodan Liang, Xiaojun Chang
Dynamic Graph Enhanced Contrastive Learning For Chest X-Ray Report Generation, Mingjie Li, Bingqian Lin, Zicong Chen, Haokun Lin, Xiaodan Liang, Xiaojun Chang
Computer Vision Faculty Publications
Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to eliminate the severe visual and textual bias in this task. The structures of such graphs are exploited by using the clinical dependencies formed by the disease topic tags via general knowledge and usually do not update during the training process. Consequently, the fixed graphs can not guarantee the most appropriate scope of knowledge and limit the effectiveness. To address the limitation, we propose a knowledge graph with Dynamic structure and nodes …
3d Semantic Segmentation In The Wild: Learning Generalized Models For Adverse-Condition Point Clouds, Aoran Xiao, Jiaxing Huang, Weihao Xuan, Ruijie Ren, Kangcheng Liu, Dayan Guan, Abdulmotaleb El Saddik, Shijian Lu, Eric Xing
3d Semantic Segmentation In The Wild: Learning Generalized Models For Adverse-Condition Point Clouds, Aoran Xiao, Jiaxing Huang, Weihao Xuan, Ruijie Ren, Kangcheng Liu, Dayan Guan, Abdulmotaleb El Saddik, Shijian Lu, Eric Xing
Computer Vision Faculty Publications
Robust point cloud parsing under all-weather conditions is crucial to level-5 autonomy in autonomous driving. However, how to learn a universal 3D semantic segmentation (3DSS) model is largely neglected as most existing benchmarks are dominated by point clouds captured under normal weather. We introduce SemanticSTF, an adverse-weather point cloud dataset that provides dense point-level annotations and allows to study 3DSS under various adverse weather conditions. We study all-weather 3DSS modeling under two setups: 1) domain adaptive 3DSS that adapts from normal-weather data to adverse-weather data; 2) domain generalizable 3DSS that learns all-weather 3DSS models from normal-weather data. Our studies reveal …
3d-Aware Multi-Class Image-To-Image Translation With Nerfs, Senmao Li, Joost Van De Weijer, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu, Jian Yang
3d-Aware Multi-Class Image-To-Image Translation With Nerfs, Senmao Li, Joost Van De Weijer, Yaxing Wang, Fahad Shahbaz Khan, Meiqin Liu, Jian Yang
Computer Vision Faculty Publications
Recent advances in 3D-aware generative models (3D-aware GANs) combined with Neural Radiance Fields (NeRF) have achieved impressive results. However no prior works investigate 3D-aware GANs for 3D consistent multiclass image-to-image (3D-aware 121) translation. Naively using 2D-121 translation methods suffers from unrealistic shape/identity change. To perform 3D-aware multiclass 121 translation, we decouple this learning process into a multiclass 3D-aware GAN step and a 3D-aware 121 translation step. In the first step, we propose two novel techniques: a new conditional architecture and an effective training strategy. In the second step, based on the well-trained multiclass 3D-aware GAN architecture, that preserves view-consistency, we …
Discriminative Co-Saliency And Background Mining Transformer For Co-Salient Object Detection, Long Li, Junwei Han, Ni Zhang, Nian Liu, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan
Discriminative Co-Saliency And Background Mining Transformer For Co-Salient Object Detection, Long Li, Junwei Han, Ni Zhang, Nian Liu, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan
Computer Vision Faculty Publications
Most previous co-salient object detection works mainly focus on extracting co-salient cues via mining the consistency relations across images while ignore explicit exploration of background regions. In this paper, we propose a Discriminative co-saliency and background Mining Transformer framework (DMT) based on several economical multi-grained correlation modules to explicitly mine both co-saliency and background information and effectively model their discrimination. Specifically, we first propose a region-to-region correlation module for introducing inter-image relations to pixel-wise segmentation features while maintaining computational efficiency. Then, we use two types of pre-defined tokens to mine co-saliency and background information via our proposed contrast-induced pixel-to-token correlation …
Burstormer: Burst Image Restoration And Enhancement Transformer, Akshay Dudhane, Syed Waqas Zamir, Salman Khan, Fahad Shahbaz Khan, Ming Hsuan Yang
Burstormer: Burst Image Restoration And Enhancement Transformer, Akshay Dudhane, Syed Waqas Zamir, Salman Khan, Fahad Shahbaz Khan, Ming Hsuan Yang
Computer Vision Faculty Publications
On a shutter press, modern handheld cameras capture multiple images in rapid succession and merge them to generate a single image. However, individual frames in a burst are misaligned due to inevitable motions and contain multiple degradations. The challenge is to properly align the successive image shots and merge their complementary information to achieve high-quality outputs. Towards this direction, we propose Burstormer: a novel transformer-based architecture for burst image restoration and enhancement. In comparison to existing works, our approach exploits multi-scale local and non-local features to achieve improved alignment and feature fusion. Our key idea is to enable inter-frame communication …
Clip2protect: Protecting Facial Privacy Using Text-Guided Makeup Via Adversarial Latent Search, Fahad Shamshad, Muzammal Naseer, Karthik Nandakumar
Clip2protect: Protecting Facial Privacy Using Text-Guided Makeup Via Adversarial Latent Search, Fahad Shamshad, Muzammal Naseer, Karthik Nandakumar
Computer Vision Faculty Publications
The success of deep learning based face recognition systems has given rise to serious privacy concerns due to their ability to enable unauthorized tracking of users in the digital world. Existing methods for enhancing privacy fail to generate 'naturalistic' images that can protect facial privacy without compromising user experience. We propose a novel two-step approach for facial privacy protection that relies on finding adversarial latent codes in the low- dimensional manifold of a pretrained generative model. The first step inverts the given face image into the latent space and finetunes the generative model to achieve an accurate reconstruction of the …
Multiclass Confidence And Localization Calibration For Object Detection, Bimsara Pathiraja, Malitha Gunawardhana, Muhammad Haris Khan
Multiclass Confidence And Localization Calibration For Object Detection, Bimsara Pathiraja, Malitha Gunawardhana, Muhammad Haris Khan
Computer Vision Faculty Publications
Albeit achieving high predictive accuracy across many challenging computer vision problems, recent studies suggest that deep neural networks (DNNs) tend to make over-confident predictions, rendering them poorly calibrated. Most of the existing attempts for improving DNN calibration are limited to classification tasks and restricted to calibrating in-domain predictions. Surprisingly, very little to no attempts have been made in studying the calibration of object detection methods, which occupy a pivotal space in vision-based security-sensitive, and safety-critical applications. In this paper, we propose a new train-time technique for calibrating modern object detection methods. It is capable of jointly calibrating multiclass confidence and …
N-Shot Benchmarking Of Whisper On Diverse Arabic Speech Recognition, Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed
N-Shot Benchmarking Of Whisper On Diverse Arabic Speech Recognition, Bashar Talafha, Abdul Waheed, Muhammad Abdul-Mageed
Natural Language Processing Faculty Publications
Whisper, the recently developed multilingual weakly supervised model, is reported to perform well on multiple speech recognition benchmarks in both monolingual and multilingual settings. However, it is not clear how Whisper would fare under diverse conditions even on languages it was evaluated on such as Arabic. In this work, we address this gap by comprehensively evaluating Whisper on several varieties of Arabic speech for the ASR task. Our evaluation covers most publicly available Arabic speech data and is performed under n-shot (zero-, few-, and full) finetuning. We also investigate the robustness of Whisper under completely novel conditions, such as in …
Impact Analysis Of Gpt Technology Revolution On Fundamental Scientific Research, Mengge Sun, Tao Han, Yanpeng Wang, Yuxin Huang, Xiwen Liu
Impact Analysis Of Gpt Technology Revolution On Fundamental Scientific Research, Mengge Sun, Tao Han, Yanpeng Wang, Yuxin Huang, Xiwen Liu
Bulletin of Chinese Academy of Sciences (Chinese Version)
The generative large model GPT represented by ChatGPT is developing rapidly, which has aroused extensive discussion in academic circle and the industry and has an incalculable impact on foundational scientific research development. The study first sorts out the development of the GPT technological revolution, and discusses the new changes brought about by this technology in scientific research. Then, based on the three aspects of application status, core principles and innovation subjects, the impact of the GPT technological revolution on basic scientific research and its development suggestions for China are discussed. The study believes that GPT technology can certainly play a …
Vision Language Navigation With Knowledge-Driven Environmental Dreamer, Fengda Zhu, Vincent C.S. Lee, Xiaojun Chang, Xiaodan Liang
Vision Language Navigation With Knowledge-Driven Environmental Dreamer, Fengda Zhu, Vincent C.S. Lee, Xiaojun Chang, Xiaodan Liang
Computer Vision Faculty Publications
Vision-language navigation (VLN) requires an agent to perceive visual observation in a house scene and navigate step-by-step following natural language instruction. Due to the high cost of data annotation and data collection, current VLN datasets provide limited instruction-trajectory data samples. Learning vision-language alignment for VLN from limited data is challenging since visual observation and language instruction are both complex and diverse. Previous works only generate augmented data based on original scenes while failing to generate data samples from unseen scenes, which limits the generalization ability of the navigation agent. In this paper, we introduce the Knowledge-driven Environmental Dreamer (KED), a …
Threads, Buckets, And Impact: A Framework For Tool Accelerated Machine Learning Courses, Jonathan Adam Niemirowski
Threads, Buckets, And Impact: A Framework For Tool Accelerated Machine Learning Courses, Jonathan Adam Niemirowski
Doctoral Dissertations
Artificial intelligence and machine learning (ML) have exploded in use, accessibility, and awareness in the past few years, particularly with the release of ChatGPT in late 2022. Advances in end-user ML tools are accelerating the development of ML applications, lowering the technical barrier of entry for users outside of the computer science (CS) community. Access to ML education within STEM is mostly limited to upper-level computer science courses that have deep pre-requisite requirements or to introductory workshops that yield limited ML skills. Despite the critical need for ML education, there is a lack of guidance in instructional design for applied …
Don't Fear The Artificial Intelligence: A Systematic Review Of Machine Learning For Prostate Cancer Detection In Pathology, Aaryn Frewing, Alexander B. Gibson, Richard Robertson, Paul Urie, Dennis Della Corte
Don't Fear The Artificial Intelligence: A Systematic Review Of Machine Learning For Prostate Cancer Detection In Pathology, Aaryn Frewing, Alexander B. Gibson, Richard Robertson, Paul Urie, Dennis Della Corte
Faculty Publications
The adoption of whole slide image (WSI) scanners in clinical practice was accelerated by US Food and Drug Administration approval in 2017, which allowed primary pathologic diagnoses to be made on scanned images. Images in the digital domain allow the application of pathology artificial intelligence (AI), including clinical decision support with algorithms performing specific diagnoses.1,2 These algorithms, if trained properly, could go beyond the ability of human observation to detect and quantify features that are not recognizable by human perception.1,3,4
Reinforcement Learning Approach To Stochastic Vehicle Routing Problem With Correlated Demands, Zangir Iklassov, Ikboljon Sobirov, Ruben Solozabal, Martin Takac
Reinforcement Learning Approach To Stochastic Vehicle Routing Problem With Correlated Demands, Zangir Iklassov, Ikboljon Sobirov, Ruben Solozabal, Martin Takac
Machine Learning Faculty Publications
We present a novel end-to-end framework for solving the Vehicle Routing Problem with stochastic demands (VRPSD) using Reinforcement Learning (RL). Our formulation incorporates the correlation between stochastic demands through other observable stochastic variables, thereby offering an experimental demonstration of the theoretical premise that non-i.i.d. stochastic demands provide opportunities for improved routing solutions. Our approach bridges the gap in the application of RL to VRPSD and consists of a parameterized stochastic policy optimized using a policy gradient algorithm to generate a sequence of actions that form the solution. Our model outperforms previous state-of-the-art metaheuristics and demonstrates robustness to changes in the …
Ocr Post-Processing Using Large Language Models, Mahdi Hajiali
Ocr Post-Processing Using Large Language Models, Mahdi Hajiali
UNLV Theses, Dissertations, Professional Papers, and Capstones
Optical Character Recognition (OCR) technology transforms textual visuals into an electronically readable, non-graphical format of the text. This allows the editing and other text manipulation of the content by language technology software such as machine translation, text comprehension, query-answering systems, and search engines. While Optical Character Recognition (OCR) systems continually progress towards greater precision, several complications persist when dealing with low-resolution source images or those with multicolored backgrounds. Consequently, the text derived from OCR necessitates additional refinement to optimize accuracy, beneficial for various subsequent applications. It is recognized that the character accuracy of OCR-generated text may influence certain natural language …
A Multi-Layer Information Dissemination Model And Interference Optimization Strategy For Communication Networks In Disaster Areas, Yuexia Zhang, Yang Hong, Mohsen Guizani, Sheng Wu, Peiying Zhang, Ruiqi Liu
A Multi-Layer Information Dissemination Model And Interference Optimization Strategy For Communication Networks In Disaster Areas, Yuexia Zhang, Yang Hong, Mohsen Guizani, Sheng Wu, Peiying Zhang, Ruiqi Liu
Machine Learning Faculty Publications
The communication network in disaster areas (CNDA) can disseminate the key disaster information in time and provide basic information support for decision-making and rescuing. Therefore, it is of great significance to study the information dissemination mechanism of CNDA. However, a CNDA is vulnerable to interference, which affects information dissemination and rescuing. To solve this problem, this paper established a multi-layer information dissemination model of CNDA (MMND) which models the CNDA from the perspective of degree distribution of nodes. The information dissemination process and equilibrium state in CNDA is analyzed by an improved dynamic dissemination method. Then, the effects of the …
Self-Supervised Pretraining And Transfer Learning On Fmri Data With Transformers, Sean Paulsen
Self-Supervised Pretraining And Transfer Learning On Fmri Data With Transformers, Sean Paulsen
Dartmouth College Ph.D Dissertations
Transfer learning is a machine learning technique founded on the idea that knowledge acquired by a model during “pretraining” on a source task can be transferred to the learning of a target task. Successful transfer learning can result in improved performance, faster convergence, and reduced demand for data. This technique is particularly desirable for the task of brain decoding in the domain of functional magnetic resonance imaging (fMRI), wherein even the most modern machine learning methods can struggle to decode labelled features of brain images. This challenge is due to the highly complex underlying signal, physical and neurological differences between …
Autonomous Shipwreck Detection & Mapping, William Ard
Autonomous Shipwreck Detection & Mapping, William Ard
LSU Master's Theses
This thesis presents the development and testing of Bruce, a low-cost hybrid Remote Operated Vehicle (ROV) / Autonomous Underwater Vehicle (AUV) system for the optical survey of marine archaeological sites, as well as a novel sonar image augmentation strategy for semantic segmentation of shipwrecks. This approach takes side-scan sonar and bathymetry data collected using an EdgeTech 2205 AUV sensor integrated with an Harris Iver3, and generates augmented image data to be used for the semantic segmentation of shipwrecks. It is shown that, due to the feature enhancement capabilities of the proposed shipwreck detection strategy, correctly identified areas have a 15% …
Vertical Federated Learning Using Autoencoders With Applications In Electrocardiograms, Wesley William Chorney
Vertical Federated Learning Using Autoencoders With Applications In Electrocardiograms, Wesley William Chorney
Theses and Dissertations
Federated learning is a framework in machine learning that allows for training a model while maintaining data privacy. Moreover, it allows clients with their own data to collaborate in order to build a stronger, shared model. Federated learning is of particular interest to healthcare data, since it is of the utmost importance to respect patient privacy while still building useful diagnostic tools. However, healthcare data can be complicated — data format might differ across providers, leading to unexpected inputs and incompatibility between different providers. For example, electrocardiograms might differ in sampling rate or number of leads used, meaning that a …
Ide-Based Learning Analytics For Assessing Introductory Programming Skill, Phyllis J. Beck
Ide-Based Learning Analytics For Assessing Introductory Programming Skill, Phyllis J. Beck
Theses and Dissertations
Providing a sufficient level of personalized feedback on students' current level of strategic knowledge within the context of the natural programming environment through IDE-based learning analytics would transform learning outcomes for introductory programming students. However, providing sufficient insight into the programming process was previously inaccessible due to the need for more complex and scalable data collection methods and metrics with a wider variety for understanding programming metacognition and the full programming process.
This research developed a custom-built web-based IDE and event compression system to investigate two of the five components of a five-dimensional model of cognition for programming skill estimation …
Proposing A Measure Of Ethicality For Humans And Ai, Alejandro Jorge Napolitano Jawerbaum
Proposing A Measure Of Ethicality For Humans And Ai, Alejandro Jorge Napolitano Jawerbaum
Electronic Theses and Dissertations
Smarter people or intelligent machines are able to make more accurate inferences about their environment and other agents more efficiently than less intelligent agents. Formally: ‘Intelligence measures an agent’s ability to achieve goals in a wide range of environments.’ (Legg, 2008)
In this dissertation we extend this definition to include ethical behaviour and we will offer a mathematical formalism and a way to estimate how ethical an action is or will be, both for a human and for a computer, by calculating the expected values of random variables. Formally, we propose the following measure of ethicality, which is computable, or …
Arabic Dysarthric Speech Recognition Using Adversarial And Signal-Based Augmentation, Massa Baali, Ibrahim Almakky, Shady Shehata, Fakhri Karray
Arabic Dysarthric Speech Recognition Using Adversarial And Signal-Based Augmentation, Massa Baali, Ibrahim Almakky, Shady Shehata, Fakhri Karray
Machine Learning Faculty Publications
Despite major advancements in Automatic Speech Recognition (ASR), the state-of-the-art ASR systems struggle to deal with impaired speech even with high-resource languages. In Arabic, this challenge gets amplified, with added complexities in collecting data from dysarthric speakers. In this paper, we aim to improve the performance of Arabic dysarthric automatic speech recognition through a multi-stage augmentation approach. To this effect, we first propose a signal-based approach to generate dysarthric Arabic speech from healthy Arabic speech by modifying its speed and tempo. We also propose a second stage Parallel Wave Generative (PWG) adversarial model that is trained on an English dysarthric …
S2cd: Self-Heuristic Speaker Content Disentanglement For Any-To-Any Voice Conversion, Pengfei Wei, Xiang Yin, Chunfeng Wang, Zhonghao Li, Xinghua Qu, Zhiqiang Xu, Zejun Ma
S2cd: Self-Heuristic Speaker Content Disentanglement For Any-To-Any Voice Conversion, Pengfei Wei, Xiang Yin, Chunfeng Wang, Zhonghao Li, Xinghua Qu, Zhiqiang Xu, Zejun Ma
Machine Learning Faculty Publications
In this paper, we propose a Self-heuristic Speaker Content Disentanglement (S2CD) model for any to any voice conversion without using any external resources, e.g., speaker labels or vectors, linguistic models, and transcriptions. S2CD is built on the disentanglement sequential variational autoencoder (DSVAE), but improves DSVAE structure at the model architecture level from three perspectives. Specifically, we develop different structures for speaker and content encoders based on their underlying static/dynamic property. We further propose a generative graph, modelled by S2CD, so as to make S2CD well mimic the multi-speaker speech generation process. Finally, we propose a self-heuristic way to introduce bias …
Fooctts: Generating Arabic Speech With Acoustic Environment For Football Commentator, Massa Baali, Ahmed Ali
Fooctts: Generating Arabic Speech With Acoustic Environment For Football Commentator, Massa Baali, Ahmed Ali
Machine Learning Faculty Publications
This paper presents FOOCTTS, an automatic pipeline for a football commentator that generates speech with background crowd noise. The application gets the text from the user, applies text pre-processing such as vowelization, followed by the commentator's speech synthesizer. Our pipeline included Arabic automatic speech recognition for data labeling, CTC segmentation, transcription vowelization to match speech, and fine-tuning the TTS. Our system is capable of generating speech with its acoustic environment within limited 15 minutes of football commentator recording. Our prototype is generalizable and can be easily applied to different domains and languages.
Controllable Language Generation Using Deep Learning, Rohola Zandie
Controllable Language Generation Using Deep Learning, Rohola Zandie
Electronic Theses and Dissertations
The advent of deep neural networks has sparked a revolution in Artificial Intelligence (AI), notably with the creation of Transformer models like GPT-X and ChatGPT. These models have surpassed previous methods in various Natural Language Processing (NLP) tasks. As the NLP field evolves, there is a need to further understand and question the capabilities of these models. Text generation, a crucial part of NLP, remains an area where our comprehension is limited while being critical in research.
This dissertation focuses on the challenging problem of controlling the general behaviors of language models such as sentiment, topical focus, and logical reasoning. …
Terrain And Adversary-Aware Autonomous Robot Navigation, Aniekan Ufot Inyang
Terrain And Adversary-Aware Autonomous Robot Navigation, Aniekan Ufot Inyang
Electronic Theses and Dissertations
In autonomous robot navigation, the robot is able to understand the environment around it for intelligent navigation. From its world model of this environment, it generates a global plan for navigation from a position to a goal based on different factors. This research aims to implement autonomous robot navigation by learning terrain affordances: traversability (moving quickly) and concealment (staying hidden from an adversary) using the Preference-based Inverse Reward Learning (PbIRL) methodology. The PbIRL methodology reduces the barrier of generating initial demonstration data to learn the terrain affordances by using a human expert’s preferences to learn individual weights over the terrain …
Topology Optimization For Artificial Neural Networks, Justin Mills
Topology Optimization For Artificial Neural Networks, Justin Mills
Masters Theses & Specialist Projects
This thesis examines the feasibility of implementing two simple optimization methods, namely the Weights Power method (Hagiwara, 1994) and the Tabu Search method (Gupta & Raza, 2020), within an existing framework. The study centers around the generation of artificial neural networks using these methods, assessing their performance in terms of both accuracy and the capacity to reduce components within the Artificial Neural Network’s (ANN) topology.
The evaluation is conducted on three classification datasets: Air Quality (Shahane, 2021), Diabetes (Soni, 2021), and MNIST (Deng, 2012). The main performance metric used is accuracy, which measures the network's predictive capability for the classification …
Evaluating Chatgpt For Recommendation: How Does The Ability To Converse Impact Recommendation?, Kyle Spurlock
Evaluating Chatgpt For Recommendation: How Does The Ability To Converse Impact Recommendation?, Kyle Spurlock
Electronic Theses and Dissertations
Recommendation algorithms have become an absolute necessity in the modern world to avoid information overload. However, the interaction between the human and the system is largely superficial and without any real contact. If you are given poor recommendations, you have no choice but to sift through mountains of content on your own until the model learns to accommodate your tastes more. This is bad for business as well as the consumer. Recently, large language models like ChatGPT have seen a significant rise in popularity due to their ease of use and wide range of knowledge. It has now become nearly …
Application Of Machine Learning Algorithms For Elucidation Of Biological Networks From Time Series Gene Expression Data, Krupa Nagori
Application Of Machine Learning Algorithms For Elucidation Of Biological Networks From Time Series Gene Expression Data, Krupa Nagori
Computational and Data Sciences (PhD) Dissertations
This dissertation provides a deep dive into understanding gene expression, interaction, regulation, and the intricate mechanisms behind heliotropism and phototropism. Additionally, the research accentuates the significance of machine learning techniques, specifically for gene regulatory networks (GRNs).
Chapter 1 offers an exhaustive benchmarking of GRN methodologies, furthering our comprehension of machine-learning models relevant to GRNs. The evaluation revealed that GRNTE, SWING, and BiXGBoost emerged as top-performing methods in GRN inference. The suitability of these models varies depending on specific research criteria such as computational needs, dataset dimensions, and performance metric emphasis. An innovation of this chapter was the introduction of Colab …
Generalizable Deep-Learning-Based Wireless Indoor Localization, Ali Owfi
Generalizable Deep-Learning-Based Wireless Indoor Localization, Ali Owfi
All Theses
The growing interest in indoor localization has been driven by its wide range of applications in areas such as smart homes, industrial automation, and healthcare. With the increasing reliance on wireless devices for location-based services, accurate estimation of device positions within indoor environments has become crucial. Deep learning approaches have shown promise in leveraging wireless parameters like Channel State Information (CSI) and Received Signal Strength Indicator (RSSI) to achieve precise localization. However, despite their success in achieving high accuracy, these deep learning models suffer from limited generalizability, making them unsuitable for deployment in new or dynamic environments without retraining. To …