Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

University of Arkansas, Fayetteville

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 45

Full-Text Articles in Data Science

Augmented Reality In Fashion Retail: A Walmart Unlimited Study, Chloe A. Mcpherson May 2026

Augmented Reality In Fashion Retail: A Walmart Unlimited Study, Chloe A. Mcpherson

Apparel Merchandising and Product Development Undergraduate Honors Theses

As technology continues to evolve, augmented reality (AR) has become increasingly common within the retail and fashion industries. This study explored Gen Z consumers’ perceptions of immersive AR shopping experiences through Walmart Unlimited, an interactive digital shopping platform. The purpose of this research was to better understand how younger consumers respond to AR-enhanced shopping environments and whether these technologies influence attitudes toward convenience, engagement, and sustainability in retail.

A quantitative research design was used for this study. Participants completed the Walmart Unlimited shopping experience and then responded to a Qualtrics survey measuring areas such as immersion, satisfaction, ease of use, …


Strengthening Cyber Resilience In Critical Infrastructure: Lessons From Major Attack Case Studies, Jodi Barnes May 2026

Strengthening Cyber Resilience In Critical Infrastructure: Lessons From Major Attack Case Studies, Jodi Barnes

Data Science Undergraduate Honors Theses

This comparative case study research paper analyzes the Colonial Pipeline attack, the Oldsmar Water Treatment Plant attack, and related case studies to identify past and current gaps in cyber resilience in critical infrastructure.  It provides insights into the importance of cybersecurity and opportunities to enhance protection in an increasingly digital world.  Findings include unsecure practices, limited communication between sectors, outdated technology, and weaknesses in security processes and employee training. These vulnerabilities are interconnected and are largely driven by limited funding within critical infrastructure systems, which restricts the ability to address them effectively.


Reimagining Less-Than-Truckload Pricing Development In Competitive Bid Environments With Artificial Intelligence, Lawson C. Levin May 2026

Reimagining Less-Than-Truckload Pricing Development In Competitive Bid Environments With Artificial Intelligence, Lawson C. Levin

Data Science Undergraduate Honors Theses

This undergraduate thesis explores how data analytics and engineering judgment are used to support pricing decisions in the less-than-truckload (LTL) freight market. It’s based on an internship with ArcBest Corporation. It explains the company’s background, its role in the LTL market, and the responsibilities of a Pricing and Supply Chain Engineer within the Yield department.

Most of the internship was spent evaluating requests for proposals (RFPs), in which a negotiating third party provides a customer’s shipment data that must be cleaned, analyzed, and translated into a comprehensive pricing offer. Using the Data Science Analytics Process as a framework, this thesis …


Prescribing Company Action Through Machine Learning And Ai, Breck T. Husong May 2026

Prescribing Company Action Through Machine Learning And Ai, Breck T. Husong

Data Science Undergraduate Honors Theses

The purpose of this research is to implement an OpenAI Reinforced Learning prescription-giving model for improving sales on a week-by-week basis. The data used comes from a segment of High Impact Analytics’s sales data that has been anonymized for proprietary reasons. The features among the data include inventory numbers, shipments in transit, total quantity and dollars of products sold each week for the past 2 years, all aggregated at the store-item-week level. In order to build this model, Tigramite, a causal discovery model combined with prediction models XGBoost, Linear Regression, Ridge Regression, Lasso Regression, Scikit-learn’s MLP, and Keras’s Neural Model …


A Comparative Machine Learning Framework For Identifying Ai-Generated Versus Real Celebrity Faces, Sidney Gehring May 2026

A Comparative Machine Learning Framework For Identifying Ai-Generated Versus Real Celebrity Faces, Sidney Gehring

Data Science Undergraduate Honors Theses

The rapid advancements in the world of generative artificial intelligence has enabled the creation of highly realistic fictitious facial images, raising concerns about authenticity and bias in computer vision systems. This study investigates the capabilities of machine learning models to distinguish between real and artificially generated facial images across gender and race focusing on celebrity imagery. Four datasets were used against the classification model, each trained on images of a single celebrity within distinct demographic groups: White women, White men, Black women, and Black men. For each group, real images are paired with AI-generated counterparts designed to closely replicate the …


A Forecasting Framework For Distribution Center Capacity Utilization: An Applied Industry Study, Jordan J. Shortt May 2026

A Forecasting Framework For Distribution Center Capacity Utilization: An Applied Industry Study, Jordan J. Shortt

Data Science Undergraduate Honors Theses

This project develops and evaluates a predictive modeling framework for forecasting distribution center capacity utilization at Company Y, with monthly forecast horizons up to one year. Motivated by the operational challenges of seasonal demand volatility, promotional cycles, and the absence of a formally defined capacity metric, the study first constructs a historical capacity utilization measure from raw warehouse management system data — reconciling item volumes, location dimensions, and utilization factors across all DCs — which serves as the target variable for all modeling work. Four models are developed and evaluated against a naïve seasonal baseline: SARIMA, LightGBM, LSTM, and a …


Developing Tracking Compliance Standards For Inbound Freight: A Data-Driven Industry Application At O’Reilly Automotive, Jackson Endacott May 2026

Developing Tracking Compliance Standards For Inbound Freight: A Data-Driven Industry Application At O’Reilly Automotive, Jackson Endacott

Data Science Undergraduate Honors Theses

Visibility of inbound freight is critical for managing operational efficiency, yet many organizations lack standardized compliance metrics for third-party carriers to uphold, preventing them from utilizing tracking data to make data-driven decisions. During a summer internship with the Transportation Department at O’Reilly Automotive, data inconsistencies were addressed in the Transportation Management System (TMS), and that data was utilized to create tracking compliance standards for third-party carriers. Data populated from various sources within O’Reilly’s TMS was cleaned, validated, and utilized to create a Tracking Scorecard that evaluates message transmission rates, timeliness, and errors. This tool provides actionable insights to improve tracking …


Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder May 2026

Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder

Electrical Engineering and Computer Science Undergraduate Honors Theses

Splines are used for representing complex functions. In statistics, splines can be used for distributional shapes that are difficult to model by traditional parametric approaches. Ramsay (1) uses M-Spline bases to estimate continuous distributions. Semi-Parametric Quantile Regression (SPQR), developed by Xu and Reich (2), models conditional distributions where a neural network is used to estimate the basis function weights that depend on covariates. (3) implements a package for SPQR in R. We build on this by implementing a version of SPQR in Python with PyTorch. By using PyTorch, we can use more sophisticated deep learning architectures than those available in …


Modular Category Optimization For Substitutability: An Item-Level Approach, Medhansh A. Sankaran May 2026

Modular Category Optimization For Substitutability: An Item-Level Approach, Medhansh A. Sankaran

Data Science Undergraduate Honors Theses

This thesis examines substitutability within Walmart apparel as a foundation for modular category optimization. Using large-scale item-level data, I develop an attribute-based framework that aggregates products to the fineline level, constructs a structured feature space, and identifies candidate substitute relationships through similarity-based matching within relevant merchandise groupings. The results show that Walmart item master data contains sufficient structure to support scalable substitute generation across a high-variety assortment. However, substitutability is not uniform: many item pairs exhibit high similarity but low observed demand transfer, indicating that structural similarity alone does not guarantee substitution. To address this, the framework is positioned within …


Recursion, Regurgitation, And Regeneration: Testing Limits And Revealing Biases Of Generative Ai Models Through Multimodal Feedback Loops, William Donnell-Lonon May 2026

Recursion, Regurgitation, And Regeneration: Testing Limits And Revealing Biases Of Generative Ai Models Through Multimodal Feedback Loops, William Donnell-Lonon

Data Science Undergraduate Honors Theses

Contemporary generative AI systems such as OpenAI's GPT-4o and DALL-E models embed complex priors about society, reality, and history shaped by training data distributions, social alignment procedures, legal constraints, and safety regulations. This study uses a "telephone game" methodology to investigate how embedded social, political, and visual biases propagate and reveal themselves through iterative multimodal generation loops, where image captioning and text-to-image models are chained in successive feedback cycles.

Using CLIP similarity metrics, facial recognition algorithms, semantic drift analysis, and qualitative content observations, I tested how image subject matter affects the rate and quality of semantic and visual shift, identity …


Escaping The Promotion Trap: A Machine Learning Framework For Brand Equity Preservation In Beverage Cpg, Lucas P. Jones May 2026

Escaping The Promotion Trap: A Machine Learning Framework For Brand Equity Preservation In Beverage Cpg, Lucas P. Jones

Data Science Undergraduate Honors Theses

When companies acquire beverage brands, they typically value them based on total sales revenue. This traditional approach treats all sales equally over time, whether they are driven by genuine consumer demand or temporary discounts. This is important because while promotions can boost short-term sales, they tend to erode brand value over long periods of time. The measurement problem extends to acquisitions, where buyers lack the tools to distinguish real consumer demand from artificial promotional inflation.

This thesis develops a framework to separate genuine baseline demand from promotional dependence using Nielsen scanner data covering 189 beverage brands across 188,304 weekly observations …


A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout May 2026

A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout

Data Science Undergraduate Honors Theses

The purpose of this paper is to analyze patterns between public safety and streetlighting for the City of Sugar Land, TX so that they may better protect their citizens.  The data involved come from the City of Sugar Land’s public works division and include type and location for all the attributes. The method of doing so involved visualizing the patterns of streetlights and their closest light readings to visualize which streetlights are underperforming using the Shiny package in R. Statistical tests were also used to quantify the association between lighting, crime occurrence, and crosswalks. From this, and the literature review, …


Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran May 2025

Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran

Data Science Undergraduate Honors Theses

Early and accurate wildfire detection is critical for minimizing environmental damage and ensuring a timely response. However, existing satellite-based wildfire datasets suffer from limitations such as coarse ground truth, poor spectral coverage, and class imbalance, which hinder progress in developing robust segmentation models. In this paper, we introduce Land8Fire, a new large-scale wildfire segmentation dataset composed of over 20,000 multispectral image patches derived from Landsat 8 and manually annotated for high-quality fire masks. Building on the ActiveFire dataset, Land8Fire improves ground truth reliability and offers predefined splits for consistent benchmarking. We evaluate a range of state-of-the-art convolutional and transformer-based models, …


Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar May 2025

Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar

Data Science Undergraduate Honors Theses

As universities navigate financial constraints and resource allocation challenges, data driven financial analysis has become increasingly important. Universities employ various methods to assess financial efficiency, predict future expenditures, and optimize student credit hour distribution. However, the approaches to financial analysis vary widely, with some institutions leveraging advanced predictive modeling and business intelligence tools, while others rely on traditional budgeting techniques and manual forecasting.

This thesis examines how the University of Arkansas' (“Uark”) financial analysis methods compare to those of other institutions and alternative data-driven approaches. Using four years of financial and student credit hour data, this study evaluates cost trends …


Optimizing Fire Station Placement In Sugar Land, Tx: A Socioeconomic Risk-Based Approach, Alicia Gallemore May 2025

Optimizing Fire Station Placement In Sugar Land, Tx: A Socioeconomic Risk-Based Approach, Alicia Gallemore

Data Science Undergraduate Honors Theses

Fire station placement has a critical role in emergency response efficiency and community safety. Traditional optimization models focus on mainly the minimization of response times and the maximization of coverage. However, this approach may overlook potential socioeconomic disparities that can influence emergency demand. This study seeks to expand upon the existing project of zoning a fire station in Sugar Land, TX, by integrating spatial road network analysis and publicly available census data—including population density, median household income, and age-based vulnerability—into a Maximal Coverage Location Problem (MCLP) framework. Using a road network-based travel time with realistic constraints, the goal is to …


Attribute Based Assortment Using Machine-Learning, Hector Negron May 2025

Attribute Based Assortment Using Machine-Learning, Hector Negron

Data Science Undergraduate Honors Theses

Retail success is influenced by a store's demographic and environmental context, both of which impact item-level sales performance. This study applies machine learning techniques to optimize item allocation based on club attributes at Sam’s Club locations. By analyzing store- specific factors such as proximity to universities, income levels, and regional preferences, the research identifies patterns that contribute to product demand. The results offer insights into how clubs can enhance inventory decisions, improving sales outcomes while reducing inefficiencies. This study reinforces the value of data-driven retail strategies and presents a practical framework for implementing predictive models in a real-world business context.


Shortage To Surge - Studying The Post-Covid-19 Guitar Retail Market, Jed H. Kim May 2025

Shortage To Surge - Studying The Post-Covid-19 Guitar Retail Market, Jed H. Kim

Data Science Undergraduate Honors Theses

The COVID-19 pandemic was one of the most catalyzing events of the 21st century, leading to supply chain disruptions, lifestyle changes, and a massive shift towards digital technologies. During the COVID-19 lockdown, many people had more free time, and over 16 million individuals learned to play guitar in the first 2 years of the pandemic. According to a study by Fender, 62% of these new guitar learners cited the pandemic as their primary reason for learning the instrument. However, pandemic policies and supply chain disruptions meant that many guitar retailers were unable to satisfy demand, and backorders accumulated. After the …


Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer May 2025

Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer

Data Science Undergraduate Honors Theses

Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …


Icylib: A Scalable Solution For Reproducible Image Classification Workflows, Leo Williams May 2025

Icylib: A Scalable Solution For Reproducible Image Classification Workflows, Leo Williams

Data Science Undergraduate Honors Theses

With the rapid expansion of e-commerce over time, ensuring the diversity and quality of product images has become a critical challenge infeasible for human completion. In conjunction with Walmart Global Tech for the Team 1 Data Science Practicum Project, image classification models were trained to assess product image sets, but training and deploying such models often involves repetitive code and inefficient processes. This thesis presents a reusable modeling library, named IcyLib, designed to streamline the training, validation, and testing of image classification models as well as dataset importation using PyTorch. IcyLib provides a structured yet flexible approach for model implementation, …


Extending Simulation-Enhanced Bayesian Optimization Of System Designs: A Computational Study, Luke Kim May 2025

Extending Simulation-Enhanced Bayesian Optimization Of System Designs: A Computational Study, Luke Kim

Data Science Undergraduate Honors Theses

This honors thesis builds off work initially accepted for publication in the Proceedings of the 2025 IISE Annual Conference & Expo, which introduced “Simulation-Enhanced Bayesian Optimization” (SEBO)—a hybrid testing optimization approach that combined the usage of unbiased but costly physical experiments with the usage of cheaper but potentially biased computer experiments to optimize engineered systems. The original study established the SEBO methodology and demonstrated its effectiveness on a multimodal, two-dimensional benchmark function. Expanding on the work performed, we conduct a broader evaluation of the SEBO framework through parameter testing and experimentation under a variety of additional benchmark functions. This investigation …


Enhancing Product Image Classification: Utilizing Machine Learning Models For Retail Applications, Avery A. Thompson May 2025

Enhancing Product Image Classification: Utilizing Machine Learning Models For Retail Applications, Avery A. Thompson

Data Science Undergraduate Honors Theses

The expansion of e-commerce has continued at a blinding pace since the COVID-19 pandemic, and retailers are constantly looking for new ways to retain customers. Ensuring that diverse and well-classified images are on product pages has been a paramount method for retailers to ensure retention as they increase product engagement and sales and enhance user experience. Managing and labeling these vast catalogs of images by hand is becoming increasingly infeasible, so some online retailers have started to turn to automated classification models to assist them. Accuracy in these classification models is integral, as a good image classification model can improve …


The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi May 2024

The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi

Computer Science and Computer Engineering Undergraduate Honors Theses

The strategic planning of offensive passing plays in the NFL incorporates numerous variables, including defensive coverages, player positioning, historical data, etc. This project develops an application using an analytical framework and an interactive model to simulate and visualize an NFL offense's passing strategy under varying conditions. Using R-programming and data management, the model dynamically represents potential passing routes in response to different defensive schemes. The system architecture integrates data from historical NFL league years to generate quantified route scores through designed mathematical equations. This allows for the prediction of potential passing routes for offensive skill players in response to the …


Sequential Optimization For Stressor-Informed Test Planning Through Integration Of Experimental And Simulated Data, Jacob Brecheisen May 2024

Sequential Optimization For Stressor-Informed Test Planning Through Integration Of Experimental And Simulated Data, Jacob Brecheisen

Data Science Undergraduate Honors Theses

This technical report details an innovative approach in reliability engineering aimed at maximizing system durability through a synergistic use of physical experimentation and computer-based modeling. Our methodology explores the efficient design and analysis of computer experiments and physical tests to facilitate accelerated reliability growth, while leveraging a sequential integration of data from these two distinct sources: costly physical experiments, characterized by random errors, and inexpensive computer simulations, marked by inherent systematic errors. The key innovation lies in the adoption of a closed-loop design and analysis method. This method begins by identifying a viable subset of important environmental stressors—such as temperature, …


A Comprehensive Analysis Of Training Induced Heat-Related Injuries At Fort Moore, Anthony Beger May 2024

A Comprehensive Analysis Of Training Induced Heat-Related Injuries At Fort Moore, Anthony Beger

Data Science Undergraduate Honors Theses

Heat related injuries are a significant problem for the United States Armed Forces. There were over 11,000 confirmed cases of heat-related illnesses that were diagnosed at more than 230 military installations from 2018-2022. These injuries are primarily due to hyperthermia (i.e., abnormally high body temperature) resulting from extreme environmental temperatures, high humidity, medications, or excessive physical work or exercise. Fort Moore has the most heat related injuries of any installation in the U.S. Department of Defense since it is home to one of the largest U. S. Army training posts with most training involving intensive outdoor activity in high heat …


Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald May 2024

Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald

Data Science Undergraduate Honors Theses

Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …


The Importance Of Data Preparation In A Data Science Problem, Sophia Beard May 2024

The Importance Of Data Preparation In A Data Science Problem, Sophia Beard

Data Science Undergraduate Honors Theses

This study is going to be based on an inventory outlier automation data science problem that is being solved to identify and prescribe inventory level outliers to help keep shelves stocked in terms of beverages. The objective of this paper will address why it is so important to understand the data that is involved in a particular data science problem and how planning ahead ensures a successful outcome in the data science world. In this data science project, Spatiotemporal Outlier Analysis for Inventory Intervention Automation, it was crucial for the team to understand, research, and visualize the data we were …


Employing Natural Language Processing To Link Customer Survey Feedback With Net Promoter Scores, Gerardo Moreno May 2024

Employing Natural Language Processing To Link Customer Survey Feedback With Net Promoter Scores, Gerardo Moreno

Data Science Undergraduate Honors Theses

This project leverages Natural Language Processing (NLP) to analyze customer feedback from Sam’s Club, aiming to pinpoint key factors influencing Net Promoter Score (NPS). Using sentiment analysis, bigram, and trigram techniques, the project analyses textual data to identify underlying themes and patterns that affect customer satisfaction. These analyses reveal actionable insights into customer preferences and pain points, facilitating a deeper understanding of what drives customer satisfaction in retail environments. By correlating these findings with NPS, this paper details strategies to enhance customer experiences at Sam’s Club, ultimately aiming to improve both satisfaction levels and NPS.


A Spatiotemporal Analysis Of Violent Crime In Little Rock, Arkansas From 1999-2022, Nicole Rogers May 2024

A Spatiotemporal Analysis Of Violent Crime In Little Rock, Arkansas From 1999-2022, Nicole Rogers

Data Science Undergraduate Honors Theses

Little Rock, Arkansas is not only the capital and largest city in Arkansas, but it has one of the highest crime rates amongst cities with over 100,000 people in the country. According to the US Census in 2020, Little Rock had a population of 202,591. In the same year, Little Rock Police Department recorded 3,567 cases of violent crime, leading to a violent crime rate of 1,805 violent crime occurrences per 100,000 people. For perspective, Chicago’s violent crime rate was approximately half of that of Little Rock during the same time period. Crime, like other social phenomena is unevenly distributed …


Predicting True Attributes Of Retailer Data, Abby Willard May 2024

Predicting True Attributes Of Retailer Data, Abby Willard

Data Science Undergraduate Honors Theses

In the rapidly evolving landscape of consumer-packaged goods (CPG) retail, understanding the true values of various factors influencing sales performance is paramount for strategic decision-making and effective resource allocation. In ensuring accuracy of data points, the CatBoost model is utilized, a state-of-the-art gradient boosting technique, to predict the true attribution values of datasets sourced from CPG industry retailers.

By leveraging CatBoost’s inherent capabilities to handle categorical data and its robustness against overfitting, the models are optimized to accurately predict the true attribution values for various items. The performance of the CatBoost models is evaluated through rigorous cross-validation techniques and compared …


Concurrent Processing Of Retail Data In Python To Optimize Runtime, Bobby Slavin May 2024

Concurrent Processing Of Retail Data In Python To Optimize Runtime, Bobby Slavin

Data Science Undergraduate Honors Theses

This thesis explores the application of multiprocessing and multithreading techniques in Python to optimize runtime efficiency on the analysis of retail data. As the retail data processed by a program increases, so does the runtime of the program. If you are performing this processing using only a single core, even a gigabyte of data can potentially take upwards to half an hour to finish processing, while larger datasets of 100 GB or more could take days, heavily limiting the amount of retail data that can be processed in a reasonable amount of time. By employing multithreading and multiprocessing architectures in …