Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1201 - 1230 of 3232

Full-Text Articles in Data Science

Combating Cyberbullying On Social Media: A Machine Learning Approach With Text Analysis On Twitter, Amir Alipour Yengejeh Jan 2024

Combating Cyberbullying On Social Media: A Machine Learning Approach With Text Analysis On Twitter, Amir Alipour Yengejeh

Data Science and Data Mining

The popularity of the electronic mobile devices along with social media as well as networking websites have been tremendously increased in the recent year. Most people around the world daily engage in the variety of cyberspace additives. Even though the users can take most advantages of these system such as exchange the idea and information, being sociable, and enjoyments, they might be faced with such adverse behaviors such as toxicity, bullying, extremism, and cruelty. The recent statistics reports that such mentioned behaviors has been noticeably grown on the cyberspace such that can threaten the individuals and even any community. Thus, …


Predicting Road Accident Injury Severity For Drivers In Automobile Crashes In United States Using Machine Learning Models And Ai, Emil Agbemade, Benedict Kongyir Jan 2024

Predicting Road Accident Injury Severity For Drivers In Automobile Crashes In United States Using Machine Learning Models And Ai, Emil Agbemade, Benedict Kongyir

Data Science and Data Mining

This study analyzes data from the National Highway Trafc Safety Administration’s 2021 Crash Report Sampling System to identify key factors contributing to the severity of injuries in car accidents. By utilizing various machine learning algorithms and cross-validation techniques, we assessed metrics such as accuracy, sensitivity, precision, specifcity, and the area under the curve (AUC) to evaluate the efectiveness of predictive models. All data preprocessing and model building was done using KNIME Analytical software [9]. Our fndings reveal signifcant correlations between certain variables such as airbag injection, weather conditions, intoxication, vehicle state, driver distractions, and injury severity. These insights underscore the …


Diagnostic In Neuroimaging: A Comparative Study Of Deep Learning And Traditional Approaches, Amina Issoufou Anaroua Jan 2024

Diagnostic In Neuroimaging: A Comparative Study Of Deep Learning And Traditional Approaches, Amina Issoufou Anaroua

Data Science and Data Mining

In the realm of medical diagnostics, precise classification of brain tumors is pivotal. This study conducts a comprehensive comparative analysis of a Convolutional Neural Network (CNN) against traditional machine learning models, Logistic Regression (LR) and Support Vector Machines (SVM) on a dataset of MRI scans for multi-class brain tumor classification. The CNN, tailored for image recognition, is evaluated alongside LR and SVM, which have established benchmarks in classification tasks. The investigation reveals that the traditional models hold their ground in terms of precision and interpretability, with the SVM, in particular, achieving remarkable accuracy. However, the CNN distinguishes itself by demonstrating …


Understanding Social Dynamics In Toxic Conversations And Public Health Intervention Acceptance On Social Media, Ana Aleksandric Jan 2024

Understanding Social Dynamics In Toxic Conversations And Public Health Intervention Acceptance On Social Media, Ana Aleksandric

Computer Science and Engineering Dissertations - Archive

Social media is now central to daily life, offering users a space to share content and opinions. However, these platforms also facilitate the spread of hate speech and misinformation, which can negatively impact public health. This dissertation develops methodologies to analyze social media data for insights that could inform health interventions. The research first examines user responses to toxic content, focusing on behavioral and emotional reactions, as well as group dynamics and bystander effects in toxic interactions. Another key focus is public opinion toward health interventions, particularly COVID-19 vaccination, using geolocated posts and analyzing factors such as race, ethnicity, and …


Advancing Deep Learning With Graph-Based Structural Insights: From Graph Classification To Semantic Segmentation, Xin Ma Jan 2024

Advancing Deep Learning With Graph-Based Structural Insights: From Graph Classification To Semantic Segmentation, Xin Ma

Computer Science and Engineering Dissertations - Archive

Deep learning has profoundly transformed machine learning by offering sophisticated data representations, yet effectively incorporating structural information remains a challenge. Structural data, whether explicit or implicit, has the potential to significantly enhance the performance of deep learning tasks. This research investigates the benefits of structural information across three crucial tasks: classification, clustering, and segmentation. For explicit structural data, where inputs are directly represented as graphs, we investigate graph-level classification in brain connectivity networks. We introduce the Multi-resolution Edge Network (MENET), a novel framework designed to identify disease-specific connectomic benchmarks with high discriminatory power across diagnostic categories. MENET leverages graph-level representations …


Optimizing Ai With Advanced Data Structuring: A Comparative Analysis Of K-Means And Gmm Clustering Techniques, Amir Alipour Yengejeh Jan 2024

Optimizing Ai With Advanced Data Structuring: A Comparative Analysis Of K-Means And Gmm Clustering Techniques, Amir Alipour Yengejeh

Data Science and Data Mining

This study presents a detailed comparison of Kmeans and Gaussian Mixture Model (GMM) clustering algorithms, illustrating their unique capabilities and limitations across various synthetic datasets. By utilizing metrics such as the Adjusted Rand Index (ARI) and Normalized Mutual Information (NMI), the research provides nuanced insights into how these algorithms handle datasets with varying structures and complexities. For instance, while both K-means and GMM show robust performance on well-separated clusters, GMM demonstrates a distinct advantage in scenarios with overlapping clusters or unbalanced data distributions. Conversely, K-means excels in identifying clear, distinct groupings, highlighting its utility in simpler clustering contexts. This study …


Adaptive Multi-Label Classification On Drifting Data Streams, Martha Roseberry Jan 2024

Adaptive Multi-Label Classification On Drifting Data Streams, Martha Roseberry

Theses and Dissertations

Drifting data streams and multi-label data are both challenging problems. When multi-label data arrives as a stream, the challenges of both problems must be addressed along with additional challenges unique to the combined problem. Algorithms must be fast and flexible, able to match both the speed and evolving nature of the stream. We propose four methods for learning from multi-label drifting data streams. First, a multi-label k Nearest Neighbors with Self Adjusting Memory (ML-SAM-kNN) exploits short- and long-term memories to predict the current and evolving states of the data stream. Second, a punitive k nearest neighbors algorithm with a self-adjusting …


Advancing Cancer Classifcation Through Machine Learning Analysis Of Rna-Seq Gene Expression Data, Emil Agbemade, Amina Issoufou Anaroua, Dimitri Bamba Jan 2024

Advancing Cancer Classifcation Through Machine Learning Analysis Of Rna-Seq Gene Expression Data, Emil Agbemade, Amina Issoufou Anaroua, Dimitri Bamba

Data Science and Data Mining

This study delves into the classifcation of various cancer types using the RNA-Seq (HiSeq) PANCAN dataset from the UCI Machine Learning Repository, which encompasses a rich collection of gene expression data across multiple tumor samples. To improve cancer diagnosis and treatment, our methodology confronts the challenges inherent in high-dimensional datasets, such as the Hughes Effect and the Curse of Dimensionality, through innovative feature selection methods and machine learning approaches. A key component of our strategy includes the use of tree-based algorithms, particularly Random Forest, to refine the dataset to seventy genes of utmost relevance for tumor classifcation, and the application …


Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe Jan 2024

Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe

Data Science and Data Mining

This project estimates a regression model to predict the superconducting critical temperature based on variables extracted from the superconductor’s chemical formula. The regression model along with the stepwise variable selection gives a reasonable and good predictive model with a lower prediction error (MSE). Variables extracted based on atomic radius, valence, atomic mass and thermal conductivity appeared to have the most contribution to the predictive model.


Modeling Health Insurance Premium Using Bayesian Hierarchical Models, Bennedict Kongyir, Emil Agbemade Jan 2024

Modeling Health Insurance Premium Using Bayesian Hierarchical Models, Bennedict Kongyir, Emil Agbemade

Data Science and Data Mining

Insurance pricing requires pragmatism and creativity due to the unpredictable nature of risk [3]. This paper explores Bayesian hierarchical models to model health insurance premiums using individual and group predictors like demographics, health status, and geography. Data from Kaggle on health insurance policyholders was utilized, with prior distributions enhanc­ing model interpretability and credibility. Bayesian models improve predictive accuracy and provide valuable insights for actuaries and policymakers, highlighting the signifcant impact of factors such as age and BMI on premium pricing.


Predicting Telecommunication Customer Attrition Using The Hopfeld Neural Network Model., Benedict Kongyir, Emil Agbemade, Kelvin Njuki Jan 2024

Predicting Telecommunication Customer Attrition Using The Hopfeld Neural Network Model., Benedict Kongyir, Emil Agbemade, Kelvin Njuki

Data Science and Data Mining

Customer churn prediction has become one of the crucial steps for customer retention. Telecommunication companies rely on loyal customers to make their proft. It is often very easy for customers to switch from one service provider to the other. To prevent or reduce the rate of customer attrition, there needs to be a model that can identify customers who are at risk of churning in the future in advance. Previous literature has shown that predictive models are efective in predicting customer churn. In this work, four tentative machine-learning models are built using data obtained from Kaggle on telecommunication customer attrition …


A 3-Step, Open-Data, Ride-Hailing Ridership Model With Pricing Applications, Richard A. Mucci Jan 2024

A 3-Step, Open-Data, Ride-Hailing Ridership Model With Pricing Applications, Richard A. Mucci

Theses and Dissertations--Civil Engineering

Researchers and practitioners studied the effects ride-hailing had in cities before the covid-19 pandemic. Previous research found ride-hailing to produce negative externalities, such as reducing transit ridership and increasing congestion in various cities. Since the pandemic, ride-hailing ridership has nearly recovered to pre-pandemic levels in Chicago. Ride-hailing ridership has grown steadily since the pandemic while a rider’s willingness to share their trip stagnated. Ride-hailing ridership nearly recovering to pre-covid levels in Chicago suggests that transportation planners, and policy makers, will need to continue assessing the impacts ride-hailing trips have in their cities.

Pickup and drop off locations in the Chicago …


Manifold Learning In Robotics: A Tutorial And Survey, Marcus Hawkins Jan 2024

Manifold Learning In Robotics: A Tutorial And Survey, Marcus Hawkins

Computer Science and Engineering Theses - Archive

In this article, we hope to represent the current state of the art of manifold learning in an understandable and approachable way. The authors will present a general overview core algorithms associated with linear and nonlinear dimensionality reduction techniques, give rudimentary definitions from differential geometry, and tenets of robotic perception, manipulation and path planning. Some of the historical applications of these algorithms will be presented, as well as conjectures about future uses, through examples from peer-reviewed journals.


When Brain Meets Artificial Intelligence, Lu Zhang Jan 2024

When Brain Meets Artificial Intelligence, Lu Zhang

Computer Science and Engineering Dissertations - Archive

When we review the history of development of artificial intelligence (AI), we will find that brain science plays a pivotal role in fostering breakthroughs in AI, such as artificial neural networks (ANNs). Today, AI has made remarkable strides, particularly with the emergence of large language models (LLMs), surpassing expectations and achieving human-level performance in certain tasks. Nonetheless, an insurmountable gap remains between AI and human intelligence. It is urgent to establish a bridge between brain science and AI, promoting their mutual enhancement and collaborations. This involve establishing connections from brain science to AI (brain-inspired AI), and reversely, from AI to …


Content Moderation On Social Media: Social And Computational Standards And Implications, Mohit Singhal Jan 2024

Content Moderation On Social Media: Social And Computational Standards And Implications, Mohit Singhal

Computer Science and Engineering Dissertations - Archive

Social media has become a powerful tool that reflects human communication's best and worst aspects. They allow individuals to freely express opinions, communicate with others, and learn about new stories. On the other hand, they have become fertile grounds for several forms of abuse, harassment, and the dissemination of misinformation. Social media platforms have established and employed content moderation to counteract the spread of abuse and misinformation.

Some critical challenges hinder the understanding of the social media content moderation ecosystem. This dissertation investigates various aspects of content moderation, including their coverage, fairness, and effectiveness. Firstly, it investigates how, in practice, …


Natural Language Generation From Large-Scale Open-Domain Knowledge Graphs, Xiao Shi Jan 2024

Natural Language Generation From Large-Scale Open-Domain Knowledge Graphs, Xiao Shi

Computer Science and Engineering Dissertations - Archive

This dissertation delves into the realm of natural language generation (NLG) from expansive open-domain knowledge graphs, aiming to bridge the gap between existing methods primarily tested on limited datasets and the demands of real-world large-scale, diverse graph structures. Prior works in NLG often relied on small-scale or restricted datasets, neglecting the complexities of broader knowledge graphs. To address this, we introduce a new dataset called GraphNarrative, designed to encompass a wide range of graph structures and enhance the realism of NLG tasks.

The core contribution of this research lies in devising a novel approach to mitigating information hallucination, a common …


Claim Sensing: A Study Linking Factual Claims To Human Behaviors On Social Media, Zeyu Zhang Jan 2024

Claim Sensing: A Study Linking Factual Claims To Human Behaviors On Social Media, Zeyu Zhang

Computer Science and Engineering Dissertations - Archive

The ubiquity of social media has transformed it into a rich source for reflecting people's opinions, behaviors, and interactions. Users frequently encounter factual claims in news, stories, and political statements, which can be either true or false. These claims significantly shape people's minds and behaviors, influencing not only individual perspectives but also broader public discourse. This study explores individuals' behaviors and perceptions toward factual claims by leveraging the concept of "check-worthiness" to analyze the relationship between such claims and user behaviors across datasets containing tens of millions of social media posts, particularly tweets from the platform X (formerly Twitter). It …


Leveraging Machine Learning & Deep Learning Methodologies To Detect Deepfakes, Aniruddha Tiwari Jan 2024

Leveraging Machine Learning & Deep Learning Methodologies To Detect Deepfakes, Aniruddha Tiwari

All Graduate Theses, Dissertations, and Other Capstone Projects

The rapid evolution of deep learning (DL) and machine learning (ML) techniques has facilitated the rise of highly convincing synthetic media, commonly referred to as deepfakes. These manipulative media artifacts, generated through advanced artificial intelligence algorithms, pose significant challenges in distinguishing them from authentic content. Given their potential to be disseminated widely across various online platforms, the imperative for robust detection methodologies becomes apparent. Accordingly, this study explores the efficacy of existing ML/DL-based approaches and aims to compare which type of methodology performs better in identifying deepfake content. In response to the escalating threat posed by deepfakes, previous research efforts …


A Comprehensive Study Of Patent Litigation In The Pharmaceutical Sector: Employing Network Theories, Graph Neural Networks, Agent Based Modeling, Bayesian Network Autocorrelation Models, Sreehas Gopinathan Jan 2024

A Comprehensive Study Of Patent Litigation In The Pharmaceutical Sector: Employing Network Theories, Graph Neural Networks, Agent Based Modeling, Bayesian Network Autocorrelation Models, Sreehas Gopinathan

Information Systems & Operations Management Dissertations - Archive

Understanding the dynamics and predictors of patent litigation is crucial in intellectual property management, especially given the competitive edge patents offer companies. Also, patents serve as both legal tools and repositories of innovation. This research delves into the complex world of patent litigation within the pharmaceutical industry, focusing on creating and applying advanced computational models to study litigation propensities. Techniques such as Graph Neural Networks (GNN), Agent-Based Modeling (ABM), and Bayesian Analysis of Network Autocorrelation Models (BANAM) are employed to explore the litigation phenomenon


Performing Holt-Winters Time Series Forecasting Using Neural Network Based Models, Kazeem Olanrewaju Bankole Jan 2024

Performing Holt-Winters Time Series Forecasting Using Neural Network Based Models, Kazeem Olanrewaju Bankole

College of Graduate Studies: Theses & Dissertations

We show how to create Artificial Neural Network based models for performing the well- known Holt-Winters time series analysis. Our work fares well compared to the well-known Holt-Winter time series prediction method while avoiding the burden of searching for the parameters of the model. We present the theoretical justification of the connection between the two models and experimental results showing the similarities of these models


Comparison Of Classification Methodologies Using Convolutional Neural Networks In A Dataset Of Plant Leaf Diseases., Ruairi O’Donohoe Jan 2024

Comparison Of Classification Methodologies Using Convolutional Neural Networks In A Dataset Of Plant Leaf Diseases., Ruairi O’Donohoe

ICT

This project investigates the impact of classification methodology selection on the performance of four Convolutional Neural Network (CNN) models applied to a multi-label image dataset. The dataset consists of plant leaf images with one or more diseases. Two classification methodologies—multi-label and multi-class—are compared based on their model performance metrics. It was hypothesised that multi-label classification would perform better, but the results show that although multi-label models performed better for Loss and Accuracy metrics, they underperformed in terms of the F1 score, which is considered a more appropriate metric for this task. This surprising result refutes the initial hypothesis. Transfer learning …


Supply Chain Optimisation With Machine Learning And Neural Networks: Applications To Demand Planning, Supply Planning, And Inventory Planning., Laurence Cully Jan 2024

Supply Chain Optimisation With Machine Learning And Neural Networks: Applications To Demand Planning, Supply Planning, And Inventory Planning., Laurence Cully

ICT

This thesis explores the impact of machine learning (ML) on supply chain planning, particularly in demand forecasting, supply planning, and inventory optimisation. By analysing literature on supply chain management, data flow, and the intersection of ML and competitive advantage, the author contextualises the research within a globalised market's demands. Case studies, interviews with industry professionals, and raw data collection provide empirical support for evaluating the research objectives and documenting the integration of ML in supply chain processes.

The findings reveal that optimised ML models, particularly those using model stacking (autoregressors, GRUs, and Random Forests), significantly outperform traditional demand forecasting methods, …


Developing Machine Learning And Time-Series Analysis Methods With Applications In Diverse Fields, Muhammed Aljifri Jan 2024

Developing Machine Learning And Time-Series Analysis Methods With Applications In Diverse Fields, Muhammed Aljifri

Theses and Dissertations

This dissertation introduces methodologies that combine machine learning models with time-series analysis to tackle data analysis challenges in varied fields. The first study enhances the traditional cumulative sum control charts with machine learning models to leverage their predictive power for better detection of process shifts, applying this advanced control chart to monitor hospital readmission rates. The second project develops multi-layer models for predicting chemical concentrations from ultraviolet-visible spectroscopy data, specifically addressing the challenge of analyzing chemicals with a wide range of concentrations. The third study presents a new method for detecting multiple changepoints in autocorrelated ordinal time series, using the …


Large Language Models, Prompting, And Synthetic Data Generation For Continual Named Entity Recognition, Charles I. Cutler Jan 2024

Large Language Models, Prompting, And Synthetic Data Generation For Continual Named Entity Recognition, Charles I. Cutler

Theses and Dissertations

With the ever-growing amount of textual data, the task of Named Entity Recognition (NER) is vital to Natural Language Processing (NLP), a field which focuses on enabling computers to understand and manipulate human language. NER enables the extraction of information from unstructured text. Accurate information extraction is crucial for applications ranging from information retrieval to systems for question-answering. To ensure that NER models are robust to changes in data distributions and capable of recognizing new entity types, one may consider expanding the capabilities of an existing model. Continual learning is a paradigm within machine learning. It studies the objective of …


Evaluation And Implementation Of Machine Learning Models To Predict Customer Churn In The Telecommunications Sector., Stephen Hasson Jan 2024

Evaluation And Implementation Of Machine Learning Models To Predict Customer Churn In The Telecommunications Sector., Stephen Hasson

ICT

This research addresses customer churn in the Telecom industry by utilizing Machine Learning (ML) models to predict customers at risk of leaving and provide data-driven retention strategies. The study highlights the effectiveness of ML, particularly in churn prediction, while noting the need for further exploration into the ethical implications of AI, such as potential biases towards vulnerable groups. Using the CRISP-DM framework, the study develops and compares three Supervised Learning (SL) models: Random Forests (RF), LightGBM (LGBM), and XGBoost (XGB), incorporating class resampling techniques to manage data imbalance.

The findings identified five key features as the most significant predictors of …


Statistical And Machine Learning Techniques For Predicting Solar Power Generation In A Microgrid., Conor Dillon Jan 2024

Statistical And Machine Learning Techniques For Predicting Solar Power Generation In A Microgrid., Conor Dillon

ICT

This study investigates statistical and machine learning models for forecasting solar power generation in microgrids, focusing on the solar installation at Powell-Focht Bioengineering Hall, UC San Diego. Accurate predictions are critical due to the variability of solar energy, aiming to optimise microgrid operations and solar power efficiency. The research compares the performance of SARIMAX, LSTM, Random Forest, and ANN models using meteorological and solar power time series data. It finds that current meteorological inputs, especially solar radiation, enhance short-term forecasting accuracy over reliance on historical patterns.

The Random Forest Auto Regressor (RFAR) outperformed other models in 10-day-ahead solar power forecasting, …


Responsible Natural Language Processing To Aid Employee Performance Reviews., Grace Rubinger Jan 2024

Responsible Natural Language Processing To Aid Employee Performance Reviews., Grace Rubinger

ICT

This research explores the use of Natural Language Processing (NLP) techniques in assessing evaluators' written appraisals during Employee Performance Reviews (EPRs), aiming to address biases inherent in traditional methods. By integrating Responsible Artificial Intelligence (AI) and foundational Large Language Models (LLMs), the study seeks to enhance the objectivity, fairness, and ethical transparency of performance evaluations. It highlights the potential of AI systems to ensure comprehensive assessments while promoting trust, ethical standards, and employee retention.

The research also aims to advance the field of AI Ethics in practical Human Resources Management (HRM) applications, particularly through NLP-driven tools. These tools are designed …


Using Machine Learning To Identify Hate Speech And Offending Language On Twitter., Mayara Lorens, Thayene Lorens Jan 2024

Using Machine Learning To Identify Hate Speech And Offending Language On Twitter., Mayara Lorens, Thayene Lorens

ICT

This project focuses on applying Machine Learning (ML) techniques to detect hate speech and offensive language on Twitter, addressing ethical concerns like cyberbullying and fostering a safer online environment. The topic is chosen for its societal significance and business relevance, as hostile online behaviour negatively impacts user experiences and platform credibility.

To achieve this, the study implements four distinct ML models to develop an automated system capable of identifying and categorising content as offensive, non-offensive, or neutral. The system aims to contribute to mitigating harmful interactions on social media and improving user safety by effectively classifying potentially problematic content.

The …


Using Unsupervised Learning Methods In Extracting Features For Classifying Rice Varieties From Rice Grains Images., Kevin Anthony Martinez Jan 2024

Using Unsupervised Learning Methods In Extracting Features For Classifying Rice Varieties From Rice Grains Images., Kevin Anthony Martinez

ICT

Rice, a staple food for nearly half of the global population, requires accurate classification of its varieties to ensure food quality, support agricultural trade, and enhance yield optimisation. Traditional manual classification methods are time-intensive and error-prone, prompting this study's exploration of unsupervised learning for feature extraction from rice grain images. The research tested classifiers on 75,000 rice samples across five classes, with 15,000 samples per class.

The study's DCGAN-CNN model achieved the highest classification accuracy of 99.67%. However, the PCA-CNN model underperformed, with only 20% accuracy, due to implementation errors. Recommendations for improvement include optimising model parameters such as learning …


Judging Our New Judges: Why We Must Remove Artificial Intelligence From Our Courtrooms Now, Kieran Duffy Newcomb Jan 2024

Judging Our New Judges: Why We Must Remove Artificial Intelligence From Our Courtrooms Now, Kieran Duffy Newcomb

Honors Theses and Capstones

In this paper, I explore some of the ways in which artificial intelligence might enhance the sentencing process through recidivism prediction technology. Notably, this technology can increase the accuracy of risk predictions and the speed with which sentencing decisions are reached. I then show, however, that the recidivism prediction technology is likely to turn into what data scientist Cathy O’Neil calls a Weapon of Math Destruction. The potential harmfulness of this technology is due not to the inherent nature of the technology, but the symbiotic relationship it will have with our already harmful criminal justice system. I argue that the …