Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,235 Full-Text Articles 9,310 Authors 1,358,113 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,235 full-text articles. Page 61 of 155.

Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger 2024 The University of Western Ontario

Federated Learning For Sentiment Analysis In Presence Of Non-Iid Data: Sensitivity Of Deep Learning Models, Davoud Gholamiangonabadi, Katarina Grolinger

Electrical and Computer Engineering Publications

In sentiment analysis, data are commonly distributed across many devices, and traditional machine learning requires transferring these data to a central location exposing data to security and privacy risks. Federated Learning (FL) avoids this transfer by training a model without requiring the clients/devices to share their local data; however, FL performance drops when data are not Independent and Identically Distributed (non-IID), such as when label distribution or data size vary across clients. Although techniques for non-IID data have been proposed primarily in the image domain, the sensitivity of various deep learning models to non-IID data needs to be examined. Consequently, …


Data Driven Trade-Off Analysis For Cybersecurity, Goskel Kucukkaya, Murat Ozer, Murat Balci, Emrah Ugurlu 2024 University of Cincinnati

Data Driven Trade-Off Analysis For Cybersecurity, Goskel Kucukkaya, Murat Ozer, Murat Balci, Emrah Ugurlu

Engineering Management & Systems Engineering Faculty Publications

Trade-off analysis, a specialization of systems engineering, addresses design criteria like security, cost, performance, and compliance. Monte Carlo simulations are commonly employed to generate impact scenarios for trade-off analysis combined with solution alternatives that accommodate industry-specific considerations and uncertainties. In the cyber domain, this paper proposes a methodology for data-driven trade-off analysis in cybersecurity, leveraging industry reports as primary data sources using confidentiality, integrity, and availability as trade-off analysis objectives. Distribution functions are derived to manage and model uncertainties for various industries. The approach given in this study aims to facilitate informed choices and to enhance cybersecurity decision making and …


Bootstrap Regression For Investigating Macroeconomics Factors Affecting Usa Home Prices, Benedict Kongyir, Emil Agbemade 2024 Oklahoma State University

Bootstrap Regression For Investigating Macroeconomics Factors Affecting Usa Home Prices, Benedict Kongyir, Emil Agbemade

Data Science and Data Mining

This study investigates the impact of macroeconomic indicators on US home prices, underscoring the importance of understanding these dynamics due to their signifcant socioeconomic consequences. Utilizing a dataset from Kaggle, originally collected by FRED, the research examines variables like the Consumer Price Index, Population, Unemployment, GDP, Stock Prices, Income, and Mortgage Rate to discern their efect on housing market fuctuations. The analysis identifes multicollinearity among predictors, necessitating a shift from traditional multiple linear regression to a more robust bootstrap regression method due to violations of parametric assumptions. Key fndings reveal that Real Disposable Income is a signifcant predictor of home …


Combating Cyberbullying On Social Media: A Machine Learning Approach With Text Analysis On Twitter, Amir Alipour Yengejeh 2024 University of Central Florida

Combating Cyberbullying On Social Media: A Machine Learning Approach With Text Analysis On Twitter, Amir Alipour Yengejeh

Data Science and Data Mining

The popularity of the electronic mobile devices along with social media as well as networking websites have been tremendously increased in the recent year. Most people around the world daily engage in the variety of cyberspace additives. Even though the users can take most advantages of these system such as exchange the idea and information, being sociable, and enjoyments, they might be faced with such adverse behaviors such as toxicity, bullying, extremism, and cruelty. The recent statistics reports that such mentioned behaviors has been noticeably grown on the cyberspace such that can threaten the individuals and even any community. Thus, …


Predicting Road Accident Injury Severity For Drivers In Automobile Crashes In United States Using Machine Learning Models And Ai, Emil Agbemade, Benedict Kongyir 2024 University of Central Florida

Predicting Road Accident Injury Severity For Drivers In Automobile Crashes In United States Using Machine Learning Models And Ai, Emil Agbemade, Benedict Kongyir

Data Science and Data Mining

This study analyzes data from the National Highway Trafc Safety Administration’s 2021 Crash Report Sampling System to identify key factors contributing to the severity of injuries in car accidents. By utilizing various machine learning algorithms and cross-validation techniques, we assessed metrics such as accuracy, sensitivity, precision, specifcity, and the area under the curve (AUC) to evaluate the efectiveness of predictive models. All data preprocessing and model building was done using KNIME Analytical software [9]. Our fndings reveal signifcant correlations between certain variables such as airbag injection, weather conditions, intoxication, vehicle state, driver distractions, and injury severity. These insights underscore the …


Diagnostic In Neuroimaging: A Comparative Study Of Deep Learning And Traditional Approaches, Amina Issoufou Anaroua 2024 University of Central Florida

Diagnostic In Neuroimaging: A Comparative Study Of Deep Learning And Traditional Approaches, Amina Issoufou Anaroua

Data Science and Data Mining

In the realm of medical diagnostics, precise classification of brain tumors is pivotal. This study conducts a comprehensive comparative analysis of a Convolutional Neural Network (CNN) against traditional machine learning models, Logistic Regression (LR) and Support Vector Machines (SVM) on a dataset of MRI scans for multi-class brain tumor classification. The CNN, tailored for image recognition, is evaluated alongside LR and SVM, which have established benchmarks in classification tasks. The investigation reveals that the traditional models hold their ground in terms of precision and interpretability, with the SVM, in particular, achieving remarkable accuracy. However, the CNN distinguishes itself by demonstrating …


Understanding Social Dynamics In Toxic Conversations And Public Health Intervention Acceptance On Social Media, Ana Aleksandric 2024 University of Texas at Arlington

Understanding Social Dynamics In Toxic Conversations And Public Health Intervention Acceptance On Social Media, Ana Aleksandric

Computer Science and Engineering Dissertations - Archive

Social media is now central to daily life, offering users a space to share content and opinions. However, these platforms also facilitate the spread of hate speech and misinformation, which can negatively impact public health. This dissertation develops methodologies to analyze social media data for insights that could inform health interventions. The research first examines user responses to toxic content, focusing on behavioral and emotional reactions, as well as group dynamics and bystander effects in toxic interactions. Another key focus is public opinion toward health interventions, particularly COVID-19 vaccination, using geolocated posts and analyzing factors such as race, ethnicity, and …


Advancing Deep Learning With Graph-Based Structural Insights: From Graph Classification To Semantic Segmentation, Xin Ma 2024 Department of Computer Science and Engineering

Advancing Deep Learning With Graph-Based Structural Insights: From Graph Classification To Semantic Segmentation, Xin Ma

Computer Science and Engineering Dissertations - Archive

Deep learning has profoundly transformed machine learning by offering sophisticated data representations, yet effectively incorporating structural information remains a challenge. Structural data, whether explicit or implicit, has the potential to significantly enhance the performance of deep learning tasks. This research investigates the benefits of structural information across three crucial tasks: classification, clustering, and segmentation. For explicit structural data, where inputs are directly represented as graphs, we investigate graph-level classification in brain connectivity networks. We introduce the Multi-resolution Edge Network (MENET), a novel framework designed to identify disease-specific connectomic benchmarks with high discriminatory power across diagnostic categories. MENET leverages graph-level representations …


Optimizing Ai With Advanced Data Structuring: A Comparative Analysis Of K-Means And Gmm Clustering Techniques, Amir Alipour Yengejeh 2024 University of Central Florida

Optimizing Ai With Advanced Data Structuring: A Comparative Analysis Of K-Means And Gmm Clustering Techniques, Amir Alipour Yengejeh

Data Science and Data Mining

This study presents a detailed comparison of Kmeans and Gaussian Mixture Model (GMM) clustering algorithms, illustrating their unique capabilities and limitations across various synthetic datasets. By utilizing metrics such as the Adjusted Rand Index (ARI) and Normalized Mutual Information (NMI), the research provides nuanced insights into how these algorithms handle datasets with varying structures and complexities. For instance, while both K-means and GMM show robust performance on well-separated clusters, GMM demonstrates a distinct advantage in scenarios with overlapping clusters or unbalanced data distributions. Conversely, K-means excels in identifying clear, distinct groupings, highlighting its utility in simpler clustering contexts. This study …


Adaptive Multi-Label Classification On Drifting Data Streams, Martha Roseberry 2024 Virginia Commonwealth University

Adaptive Multi-Label Classification On Drifting Data Streams, Martha Roseberry

Theses and Dissertations

Drifting data streams and multi-label data are both challenging problems. When multi-label data arrives as a stream, the challenges of both problems must be addressed along with additional challenges unique to the combined problem. Algorithms must be fast and flexible, able to match both the speed and evolving nature of the stream. We propose four methods for learning from multi-label drifting data streams. First, a multi-label k Nearest Neighbors with Self Adjusting Memory (ML-SAM-kNN) exploits short- and long-term memories to predict the current and evolving states of the data stream. Second, a punitive k nearest neighbors algorithm with a self-adjusting …


Advancing Cancer Classifcation Through Machine Learning Analysis Of Rna-Seq Gene Expression Data, Emil Agbemade, Amina Issoufou Anaroua, Dimitri Bamba 2024 University of Central Florida

Advancing Cancer Classifcation Through Machine Learning Analysis Of Rna-Seq Gene Expression Data, Emil Agbemade, Amina Issoufou Anaroua, Dimitri Bamba

Data Science and Data Mining

This study delves into the classifcation of various cancer types using the RNA-Seq (HiSeq) PANCAN dataset from the UCI Machine Learning Repository, which encompasses a rich collection of gene expression data across multiple tumor samples. To improve cancer diagnosis and treatment, our methodology confronts the challenges inherent in high-dimensional datasets, such as the Hughes Effect and the Curse of Dimensionality, through innovative feature selection methods and machine learning approaches. A key component of our strategy includes the use of tree-based algorithms, particularly Random Forest, to refine the dataset to seventy genes of utmost relevance for tumor classifcation, and the application …


Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe 2024 University of Central Florida

Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe

Data Science and Data Mining

This project estimates a regression model to predict the superconducting critical temperature based on variables extracted from the superconductor’s chemical formula. The regression model along with the stepwise variable selection gives a reasonable and good predictive model with a lower prediction error (MSE). Variables extracted based on atomic radius, valence, atomic mass and thermal conductivity appeared to have the most contribution to the predictive model.


Modeling Health Insurance Premium Using Bayesian Hierarchical Models, Bennedict Kongyir, Emil Agbemade 2024 Oklahoma State University - Main Campus

Modeling Health Insurance Premium Using Bayesian Hierarchical Models, Bennedict Kongyir, Emil Agbemade

Data Science and Data Mining

Insurance pricing requires pragmatism and creativity due to the unpredictable nature of risk [3]. This paper explores Bayesian hierarchical models to model health insurance premiums using individual and group predictors like demographics, health status, and geography. Data from Kaggle on health insurance policyholders was utilized, with prior distributions enhanc­ing model interpretability and credibility. Bayesian models improve predictive accuracy and provide valuable insights for actuaries and policymakers, highlighting the signifcant impact of factors such as age and BMI on premium pricing.


Predicting Telecommunication Customer Attrition Using The Hopfeld Neural Network Model., Benedict Kongyir, Emil Agbemade, Kelvin Njuki 2024 Oklahoma State University - Main Campus

Predicting Telecommunication Customer Attrition Using The Hopfeld Neural Network Model., Benedict Kongyir, Emil Agbemade, Kelvin Njuki

Data Science and Data Mining

Customer churn prediction has become one of the crucial steps for customer retention. Telecommunication companies rely on loyal customers to make their proft. It is often very easy for customers to switch from one service provider to the other. To prevent or reduce the rate of customer attrition, there needs to be a model that can identify customers who are at risk of churning in the future in advance. Previous literature has shown that predictive models are efective in predicting customer churn. In this work, four tentative machine-learning models are built using data obtained from Kaggle on telecommunication customer attrition …


A 3-Step, Open-Data, Ride-Hailing Ridership Model With Pricing Applications, Richard A. Mucci 2024 University of Kentucky

A 3-Step, Open-Data, Ride-Hailing Ridership Model With Pricing Applications, Richard A. Mucci

Theses and Dissertations--Civil Engineering

Researchers and practitioners studied the effects ride-hailing had in cities before the covid-19 pandemic. Previous research found ride-hailing to produce negative externalities, such as reducing transit ridership and increasing congestion in various cities. Since the pandemic, ride-hailing ridership has nearly recovered to pre-pandemic levels in Chicago. Ride-hailing ridership has grown steadily since the pandemic while a rider’s willingness to share their trip stagnated. Ride-hailing ridership nearly recovering to pre-covid levels in Chicago suggests that transportation planners, and policy makers, will need to continue assessing the impacts ride-hailing trips have in their cities.

Pickup and drop off locations in the Chicago …


Manifold Learning In Robotics: A Tutorial And Survey, marcus hawkins 2024 University of Texas at Arlington

Manifold Learning In Robotics: A Tutorial And Survey, Marcus Hawkins

Computer Science and Engineering Theses - Archive

In this article, we hope to represent the current state of the art of manifold learning in an understandable and approachable way. The authors will present a general overview core algorithms associated with linear and nonlinear dimensionality reduction techniques, give rudimentary definitions from differential geometry, and tenets of robotic perception, manipulation and path planning. Some of the historical applications of these algorithms will be presented, as well as conjectures about future uses, through examples from peer-reviewed journals.


When Brain Meets Artificial Intelligence, Lu Zhang 2024 The University of Texas at Arlington

When Brain Meets Artificial Intelligence, Lu Zhang

Computer Science and Engineering Dissertations - Archive

When we review the history of development of artificial intelligence (AI), we will find that brain science plays a pivotal role in fostering breakthroughs in AI, such as artificial neural networks (ANNs). Today, AI has made remarkable strides, particularly with the emergence of large language models (LLMs), surpassing expectations and achieving human-level performance in certain tasks. Nonetheless, an insurmountable gap remains between AI and human intelligence. It is urgent to establish a bridge between brain science and AI, promoting their mutual enhancement and collaborations. This involve establishing connections from brain science to AI (brain-inspired AI), and reversely, from AI to …


Content Moderation On Social Media: Social And Computational Standards And Implications, Mohit Singhal 2024 University of Texas at Arlington

Content Moderation On Social Media: Social And Computational Standards And Implications, Mohit Singhal

Computer Science and Engineering Dissertations - Archive

Social media has become a powerful tool that reflects human communication's best and worst aspects. They allow individuals to freely express opinions, communicate with others, and learn about new stories. On the other hand, they have become fertile grounds for several forms of abuse, harassment, and the dissemination of misinformation. Social media platforms have established and employed content moderation to counteract the spread of abuse and misinformation.

Some critical challenges hinder the understanding of the social media content moderation ecosystem. This dissertation investigates various aspects of content moderation, including their coverage, fairness, and effectiveness. Firstly, it investigates how, in practice, …


Natural Language Generation From Large-Scale Open-Domain Knowledge Graphs, Xiao Shi 2024 University of Texas at Arlington

Natural Language Generation From Large-Scale Open-Domain Knowledge Graphs, Xiao Shi

Computer Science and Engineering Dissertations - Archive

This dissertation delves into the realm of natural language generation (NLG) from expansive open-domain knowledge graphs, aiming to bridge the gap between existing methods primarily tested on limited datasets and the demands of real-world large-scale, diverse graph structures. Prior works in NLG often relied on small-scale or restricted datasets, neglecting the complexities of broader knowledge graphs. To address this, we introduce a new dataset called GraphNarrative, designed to encompass a wide range of graph structures and enhance the realism of NLG tasks.

The core contribution of this research lies in devising a novel approach to mitigating information hallucination, a common …


Claim Sensing: A Study Linking Factual Claims To Human Behaviors On Social Media, Zeyu Zhang 2024 University of Texas at Arlington

Claim Sensing: A Study Linking Factual Claims To Human Behaviors On Social Media, Zeyu Zhang

Computer Science and Engineering Dissertations - Archive

The ubiquity of social media has transformed it into a rich source for reflecting people's opinions, behaviors, and interactions. Users frequently encounter factual claims in news, stories, and political statements, which can be either true or false. These claims significantly shape people's minds and behaviors, influencing not only individual perspectives but also broader public discourse. This study explores individuals' behaviors and perceptions toward factual claims by leveraging the concept of "check-worthiness" to analyze the relationship between such claims and user behaviors across datasets containing tens of millions of social media posts, particularly tweets from the platform X (formerly Twitter). It …


Digital Commons powered by bepress