Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,232 Full-Text Articles 9,306 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,232 full-text articles. Page 5 of 155.

Reimagining Less-Than-Truckload Pricing Development In Competitive Bid Environments With Artificial Intelligence, Lawson C. Levin 2026 University of Arkansas, Fayetteville

Reimagining Less-Than-Truckload Pricing Development In Competitive Bid Environments With Artificial Intelligence, Lawson C. Levin

Data Science Undergraduate Honors Theses

This undergraduate thesis explores how data analytics and engineering judgment are used to support pricing decisions in the less-than-truckload (LTL) freight market. It’s based on an internship with ArcBest Corporation. It explains the company’s background, its role in the LTL market, and the responsibilities of a Pricing and Supply Chain Engineer within the Yield department.

Most of the internship was spent evaluating requests for proposals (RFPs), in which a negotiating third party provides a customer’s shipment data that must be cleaned, analyzed, and translated into a comprehensive pricing offer. Using the Data Science Analytics Process as a framework, this thesis …


Prescribing Company Action Through Machine Learning And Ai, Breck T. Husong 2026 University of Arkansas, Fayetteville

Prescribing Company Action Through Machine Learning And Ai, Breck T. Husong

Data Science Undergraduate Honors Theses

The purpose of this research is to implement an OpenAI Reinforced Learning prescription-giving model for improving sales on a week-by-week basis. The data used comes from a segment of High Impact Analytics’s sales data that has been anonymized for proprietary reasons. The features among the data include inventory numbers, shipments in transit, total quantity and dollars of products sold each week for the past 2 years, all aggregated at the store-item-week level. In order to build this model, Tigramite, a causal discovery model combined with prediction models XGBoost, Linear Regression, Ridge Regression, Lasso Regression, Scikit-learn’s MLP, and Keras’s Neural Model …


A Comparative Machine Learning Framework For Identifying Ai-Generated Versus Real Celebrity Faces, Sidney Gehring 2026 University of Arkansas, Fayetteville

A Comparative Machine Learning Framework For Identifying Ai-Generated Versus Real Celebrity Faces, Sidney Gehring

Data Science Undergraduate Honors Theses

The rapid advancements in the world of generative artificial intelligence has enabled the creation of highly realistic fictitious facial images, raising concerns about authenticity and bias in computer vision systems. This study investigates the capabilities of machine learning models to distinguish between real and artificially generated facial images across gender and race focusing on celebrity imagery. Four datasets were used against the classification model, each trained on images of a single celebrity within distinct demographic groups: White women, White men, Black women, and Black men. For each group, real images are paired with AI-generated counterparts designed to closely replicate the …


A Forecasting Framework For Distribution Center Capacity Utilization: An Applied Industry Study, Jordan J. Shortt 2026 University of Arkansas, Fayetteville

A Forecasting Framework For Distribution Center Capacity Utilization: An Applied Industry Study, Jordan J. Shortt

Data Science Undergraduate Honors Theses

This project develops and evaluates a predictive modeling framework for forecasting distribution center capacity utilization at Company Y, with monthly forecast horizons up to one year. Motivated by the operational challenges of seasonal demand volatility, promotional cycles, and the absence of a formally defined capacity metric, the study first constructs a historical capacity utilization measure from raw warehouse management system data — reconciling item volumes, location dimensions, and utilization factors across all DCs — which serves as the target variable for all modeling work. Four models are developed and evaluated against a naïve seasonal baseline: SARIMA, LightGBM, LSTM, and a …


Developing Tracking Compliance Standards For Inbound Freight: A Data-Driven Industry Application At O’Reilly Automotive, Jackson Endacott 2026 University of Arkansas, Fayetteville

Developing Tracking Compliance Standards For Inbound Freight: A Data-Driven Industry Application At O’Reilly Automotive, Jackson Endacott

Data Science Undergraduate Honors Theses

Visibility of inbound freight is critical for managing operational efficiency, yet many organizations lack standardized compliance metrics for third-party carriers to uphold, preventing them from utilizing tracking data to make data-driven decisions. During a summer internship with the Transportation Department at O’Reilly Automotive, data inconsistencies were addressed in the Transportation Management System (TMS), and that data was utilized to create tracking compliance standards for third-party carriers. Data populated from various sources within O’Reilly’s TMS was cleaned, validated, and utilized to create a Tracking Scorecard that evaluates message transmission rates, timeliness, and errors. This tool provides actionable insights to improve tracking …


Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder 2026 University of Arkansas, Fayetteville

Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder

Electrical Engineering and Computer Science Undergraduate Honors Theses

Splines are used for representing complex functions. In statistics, splines can be used for distributional shapes that are difficult to model by traditional parametric approaches. Ramsay (1) uses M-Spline bases to estimate continuous distributions. Semi-Parametric Quantile Regression (SPQR), developed by Xu and Reich (2), models conditional distributions where a neural network is used to estimate the basis function weights that depend on covariates. (3) implements a package for SPQR in R. We build on this by implementing a version of SPQR in Python with PyTorch. By using PyTorch, we can use more sophisticated deep learning architectures than those available in …


High Throughput Phenomics Pipeline For Pulse Crop Nutritional Breeding, Amod Udayanga Madurapperumage 2026 Clemson University

High Throughput Phenomics Pipeline For Pulse Crop Nutritional Breeding, Amod Udayanga Madurapperumage

All Dissertations

Dry pea (Pisum sativum L.), lentil (Lens culinaris Medik.), and chickpea (Cicer arietinum L.) are major pulse crops valued for their high nutritional composition and importance to global food systems. Pulses are rich in carbohydrates, protein, and essential minerals, making them ideal whole foods and critical contributors to food and nutrition security. Due to these advantages, pulse breeding programs are increasingly focusing on enhancing nutritional traits, such as protein quality, amino acid balance, and micronutrient density, through the process of biofortification. However, improvement of agronomic traits remains equally essential. Characteristics such as plant height, standability, stress tolerance, …


Modular Category Optimization For Substitutability: An Item-Level Approach, Medhansh A. Sankaran 2026 University of Arkansas, Fayetteville

Modular Category Optimization For Substitutability: An Item-Level Approach, Medhansh A. Sankaran

Data Science Undergraduate Honors Theses

This thesis examines substitutability within Walmart apparel as a foundation for modular category optimization. Using large-scale item-level data, I develop an attribute-based framework that aggregates products to the fineline level, constructs a structured feature space, and identifies candidate substitute relationships through similarity-based matching within relevant merchandise groupings. The results show that Walmart item master data contains sufficient structure to support scalable substitute generation across a high-variety assortment. However, substitutability is not uniform: many item pairs exhibit high similarity but low observed demand transfer, indicating that structural similarity alone does not guarantee substitution. To address this, the framework is positioned within …


Recursion, Regurgitation, And Regeneration: Testing Limits And Revealing Biases Of Generative Ai Models Through Multimodal Feedback Loops, William Donnell-Lonon 2026 University of Arkansas, Fayetteville

Recursion, Regurgitation, And Regeneration: Testing Limits And Revealing Biases Of Generative Ai Models Through Multimodal Feedback Loops, William Donnell-Lonon

Data Science Undergraduate Honors Theses

Contemporary generative AI systems such as OpenAI's GPT-4o and DALL-E models embed complex priors about society, reality, and history shaped by training data distributions, social alignment procedures, legal constraints, and safety regulations. This study uses a "telephone game" methodology to investigate how embedded social, political, and visual biases propagate and reveal themselves through iterative multimodal generation loops, where image captioning and text-to-image models are chained in successive feedback cycles.

Using CLIP similarity metrics, facial recognition algorithms, semantic drift analysis, and qualitative content observations, I tested how image subject matter affects the rate and quality of semantic and visual shift, identity …


Escaping The Promotion Trap: A Machine Learning Framework For Brand Equity Preservation In Beverage Cpg, Lucas P. Jones 2026 University of Arkansas, Fayetteville

Escaping The Promotion Trap: A Machine Learning Framework For Brand Equity Preservation In Beverage Cpg, Lucas P. Jones

Data Science Undergraduate Honors Theses

When companies acquire beverage brands, they typically value them based on total sales revenue. This traditional approach treats all sales equally over time, whether they are driven by genuine consumer demand or temporary discounts. This is important because while promotions can boost short-term sales, they tend to erode brand value over long periods of time. The measurement problem extends to acquisitions, where buyers lack the tools to distinguish real consumer demand from artificial promotional inflation.

This thesis develops a framework to separate genuine baseline demand from promotional dependence using Nielsen scanner data covering 189 beverage brands across 188,304 weekly observations …


A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout 2026 University of Arkansas, Fayetteville

A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout

Data Science Undergraduate Honors Theses

The purpose of this paper is to analyze patterns between public safety and streetlighting for the City of Sugar Land, TX so that they may better protect their citizens.  The data involved come from the City of Sugar Land’s public works division and include type and location for all the attributes. The method of doing so involved visualizing the patterns of streetlights and their closest light readings to visualize which streetlights are underperforming using the Shiny package in R. Statistical tests were also used to quantify the association between lighting, crime occurrence, and crosswalks. From this, and the literature review, …


A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari 2026 Florida Institute of Technology

A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari

Theses and Dissertations

Student retention and degree completion remain central challenges for higher-education institutions, with significant implications for student success, institutional effectiveness, and public accountability. While advances in predictive analytics have enabled earlier identification of students at risk of withdrawal, many commonly used machine learning approaches suffer from limited interpretability, constraining their practical usefulness for advising, intervention, and policy decision making. This dissertation addresses the problem of predicting student persistence by developing and evaluating optimization based, interpretable classification models within the Logical Analysis of Data (LAD) framework. Building on existing LAD formulations, this research introduces two novel pattern generation models, the Best Term …


Utilizing Physics Informed Neural Networks For Disrupted Signal Dynamics, Nicholas J. Joyner 2026 East Tennessee State University

Utilizing Physics Informed Neural Networks For Disrupted Signal Dynamics, Nicholas J. Joyner

Electronic Theses and Dissertations

Physics-informed neural networks (PINNs) have been used in many applications including engineering and physical sciences. PINNs allow the incorporation of a priori understanding of a process’ structure into the modeling. We attempt to leverage the PINN structure toward the evaluation of disruptions to classical dynamical models by combining elements of ordinary differential equations into our loss function with sigmoidal gating to balance the penalties for deviations from the data with those for structural deviations. This enables the identification of the signal structure and the limits of disruption influence. As a use case, we consider stock value from 2019-2021, which expresses …


Analyzing The Writing Style Of Generative Ai When Prompted With Writing Samples, Samuel McDowell 2026 Liberty University

Analyzing The Writing Style Of Generative Ai When Prompted With Writing Samples, Samuel Mcdowell

Senior Honors Theses

Authorship attribution is an important topic in today’s world of Large Language Models (LLMs). It is the technology that helps to verify the author of a written work. This study explores whether LLMs can successfully mimic an individual’s writing style if they are given a text sample. A dataset of human-written texts was collected and used to prompt several LLMs to generate new texts that attempt to replicate the original author’s stylistic characteristics. The generated texts were then tested with modern authorship attribution models to determine whether they would be identified as being written by the original author. The results …


Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg 2026 Macalester College

Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg

Mathematics, Statistics, and Computer Science Honors Projects

Emergency department (ED) wait times remain a persistent bottleneck in the United States healthcare system, impacting patient outcomes, hospital efficiency, and equitable access to care. This study analyzes nationally representative data from the National Hospital Ambulatory Medical Care Survey (NHAMCS), a complex, multi-stage probability sample. Using survey-weighted analyses and predictive modeling, we examine the effects of patient characteristics, triage acuity, and visit timing. Results indicate that operational and system-level factors, including hospital capacity, geographic region, and temporal variation, are among the most influential predictors of ED wait times


Visualizing Probabilistic Model Checking: An Interactive Framework For Exploring Ctmc Models, Ishara Mawelle Kankanamge 2026 Utah State University

Visualizing Probabilistic Model Checking: An Interactive Framework For Exploring Ctmc Models, Ishara Mawelle Kankanamge

All Graduate Theses and Dissertations, Fall 2023 to Present

Probabilistic model checking is a critical method for analyzing systems characterized by uncertainty, such as communication protocols, randomized algorithms, and biochemical networks. While formal verification tools provide precise numerical data about these systems, interpreting these results is often limited by large state space, high-dimensional state space and time dependent evolution. Current tools typically output raw numerical data, offering limited support for intuitively understanding the time-dependent behavior of a model. This research presents an interactive visualization framework designed to bridge the gap between complex numerical analysis and human intuition. The framework integrates coordinated visual interfaces, including lower-dimensional state-space projections and synchronized …


From Vibration To Visualization: Building An Real-Time Audio Visualization System For Learning And Exploration Using Pyqt5, Aidan Roach 2026 Grand Prairie Fine Arts Academy

From Vibration To Visualization: Building An Real-Time Audio Visualization System For Learning And Exploration Using Pyqt5, Aidan Roach

The Transdisciplinary STEAM+ Journal

This paper presents the design, development, and analysis of my real-time audio visualization system created entirely in Python using PyQt5, called WaveCatcher. The system captures live audio input from a microphone and provides simultaneous visual feedback through multiple signal representations: including a time-domain waveform, a frequency-domain FFT spectrum, a scrolling spectrogram, harmonic peak visualization, dynamic range, amplitude envelope, spectral centroid, and spectral bandwidth. These features offer insight not just into the raw structure of sound, but into how humans perceive its qualities–like timbre! This terminology may seem intimidating—it certainly was when I first began learning it—but I’ll explain all of …


Artificial Intelligence In Medicine: Barriers, Solutions, And Strategies, Anil Harrison, Melissa Stradley Moreno, Caroline E. Williams, Munevver Mine Subasi, Ersoy Subasi 2026 Midwestern University College of Medicine

Artificial Intelligence In Medicine: Barriers, Solutions, And Strategies, Anil Harrison, Melissa Stradley Moreno, Caroline E. Williams, Munevver Mine Subasi, Ersoy Subasi

HCA Healthcare Journal of Medicine

The integration of artificial intelligence (AI) and machine learning (ML) into health care holds the potential to revolutionize patient care by enhancing clinical decision-making, improving diagnostic accuracy, and reducing costs. Despite this promise, adoption remains limited due to a range of technical, regulatory, educational, and cultural barriers. This paper examines these challenges and proposes strategies to support safe and effective implementation of AI in clinical practice.

Key barriers include the lack of model interpretability, often referred to as the "black box" problem, which undermines clinician trust and accountability in clinical settings, evolving regulatory frameworks and unresolved questions surrounding liability, and …


A Deep Learning-Based Approach For Bot Detection In Trending Hashtags On X, Mehboob Hussain, Muhammad Rizwan Rashid Rana, Muhammad Imran, Muhammad Shoaib, Muhammad Hasaan Mujtaba 2026 Department of Robotics & Artificial Intelligence, University is Shaheed Zulfikar Ali Bhutto Institute of Science and Technology, Islamabad 44090, Pakistan

A Deep Learning-Based Approach For Bot Detection In Trending Hashtags On X, Mehboob Hussain, Muhammad Rizwan Rashid Rana, Muhammad Imran, Muhammad Shoaib, Muhammad Hasaan Mujtaba

Makara Journal of Technology

The widespread presence of bots on social media platforms, such as X (formerly Twitter), poses a significant threat to the integrity of online information by facilitating the dissemination of misinformation and manipulating public discourse. This study proposes a robust deep learning-based framework, DeepBot, to detect bot participation in trending hashtags and discussions on X. The approach uses a dataset sourced from Kaggle, comprising user profile metadata, including follower count, tweet frequency, account verification status, and engagement metrics. The data were subjected to comprehensive preprocessing, including noise removal, part-of-speech (POS) tagging, and word embedding using the pre-trained GloVe model. RoBERTa is …


Repurposing An Old Pc Into A Self-Hosted Cloud And Media Server, Riley Biggins 2026 Western Michigan University

Repurposing An Old Pc Into A Self-Hosted Cloud And Media Server, Riley Biggins

Honors Theses

Commercial cloud platforms rely heavily on subscription models, so everyday users fall into a "cloud trap" summarized by perpetual fees, forfeited privacy, and a lack of true asset ownership. At the same time, functioning consumer electronics get discarded as e-waste when they fail to meet the hardware requirements of modern operating systems. This project presents a sustainable and cost-effective alternative to both issues by repurposing a 2017 HP Pavilion laptop into a personal cloud and media server.

Despite severe hardware constraints, my software architecture rivals the utility of paid services. By replacing Windows with a headless installation of Ubuntu Server, …


Digital Commons powered by bepress