Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Machine Learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 101

Full-Text Articles in Statistics and Probability

Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava Sep 2026

Mapping The Water Quality Of Jamaica Bay, New York (1996-2024): Principal Component Analysis And K-Means Clustering, Sneha Srivastava

Dissertations, Theses, and Capstone Projects

Jamaica Bay, located along the southeastern coast of New York City, acts as a biodiverse estuary of wetlands, meadows, and salt marsh islands. The purpose of this study is to analyze the water quality conditions of the region over time, comparing locations around the bay to identify hyperlocal features that influence larger trends in the hydrological system. Ten variables were used as water quality indicators, including total Kjeldahl nitrogen, salinity, pH, Secchi disk depth, and total phosphorus, among others, across five stations in the bay, between 1994 and 2024. After data cleaning and standardization methods were applied, principal component analysis …


Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen Jul 2026

Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen

Dissertations, Theses, and Projects

The increasing adoption of the Internet of Medical Things (IoMT) has improved healthcare delivery through connected medical devices while simultaneously expanding the cybersecurity risks facing healthcare organizations. Although machine learning based intrusion detection systems have demonstrated high detection accuracy, their ability to respond reliably to previously unseen cyberattacks remains uncertain. This study investigated how a Neural Network model and a Logistic Regression model classified novel cyberattacks within the IoMT environment. The Neural Network and Logistic Regression models were both trained and tested using a subset of the CICIoMT2024 benchmark dataset. The Neural Network achieved 99.82% test accuracy and a 0.94 …


Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D. Apr 2026

Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.

SPARK Symposium Presentations

Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression,  Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …


Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii Jan 2026

Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii

Williams Honors College, Honors Research Projects

Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …


Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci Jan 2026

Challenges And Applications Of Fine-Grained Temporal Action Understanding: Modeling Temporal Granularity And Data Efficiency, Halil I. Helvaci

Theses and Dissertations--Electrical and Computer Engineering

Fine-grained Temporal Action Segmentation (TAS) has become a cornerstone of video understanding, offering dense frame-level predictions essential for clinical assessment, surgical skill evaluation, and human-computer interaction. While TAS methods have delivered strong results on coarse-grained benchmarks, two fundamental challenges persist: (1) global attention mechanisms dilute boundary information critical for subsecond precision, a phenomenon we term the temporal granularity bottleneck, and (2) dense frame-level annotation remains prohibitively expensive, with most datasets requiring exhaustive labeling of lengthy untrimmed videos. These challenges are particularly pronounced in medical domains, where sub-second primitives define clinical outcomes while expert annotation remains scarce. In this dissertation, we …


A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha Jan 2026

A Comparative Evaluation Of Data Imbalance Handling Techniques In Machine Learning Models For One-Year Mortality Prediction In Liver Cirrhosis, Sumiya Hasan Trisha

UNF Graduate Theses and Dissertations

Liver cirrhosis is associated with substantial morbidity and mortality, making one-year mortality prediction a clinically relevant problem. Using a liver cirrhosis dataset as the motivating application, this thesis evaluates five machine learning classifiers—Logistic Regression, Random Forest, XGBoost, LightGBM, and CatBoost—under five class-imbalance handling strategies: Baseline learning, Random Oversampling, SMOTE-NC, ADASYN, and Cost-Sensitive Learning. Hyperparameter tuning was conducted using randomized search, and predictive performance was assessed over 200 iterations of Monte Carlo Cross-Validation using Accuracy, Precision, Recall, Fl-score, and ROC-AUC.

The results suggest that imbalance-handling strategies can materially affect predictive performance, particularly recall. Because the outcome of interest is death within …


Performance Of Numerical Methods Applied To The Black–Scholes Model, Scott Cameron Williams Jan 2026

Performance Of Numerical Methods Applied To The Black–Scholes Model, Scott Cameron Williams

UNF Graduate Theses and Dissertations

We compare five numerical approaches for approximating solutions to the Black–Scholes partial differential equation for pricing European call options: FTCS, BTCS, Crank– Nicolson, Monte Carlo simulation, and a physics–informed neural network (PINN). These methods span finite difference techniques, probabilistic simulation, and machine learning. Performance is evaluated based on computational efficiency and accuracy relative to the analytical Black–Scholes solution.

Among the methods, Crank–Nicolson and the PINN demonstrated the strongest overall performance. Crank–Nicolson achieved the highest accuracy but exhibited increased runtime as the number of underlying stock price grid points grew. In contrast, the PINN produced slightly less accurate results but with …


Explainable Ai For Liver Transplant Survival Prediction: Integrating Immunological Mismatch Features, Sourab Shaik Dec 2025

Explainable Ai For Liver Transplant Survival Prediction: Integrating Immunological Mismatch Features, Sourab Shaik

Honors Projects

Liver Transplantations are crucial treatment for end-stage liver disease. However, a persistent deficit of donor organs necessitates maximizing the utility of each available graft to minimize failure rates. We evaluated whether donor–recipient molecular immunogenicity metrics - Electrostatic and Hydrophobic Mismatch Scores (HMS/EMS) and eplet-based counts - improve post–liver-transplant survival prediction. The analytic cohort comprised adult, first time, single-organ deceased-donor transplants drawn from Scientific Registry of Transplant Recipients; follow-up was truncated at five years, and the endpoint was all-cause graft failure (earliest of graft failure or death; otherwise, censored). HLA variables were derived via high- resolution conversion and molecular mismatch computations …


Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel Aug 2025

Crop Yield Prediction At Multiple Spatial Scales With Statistical Machine Learning, Vaibhav Charan, Pratishtha Poudel

Discovery Undergraduate Interdisciplinary Research Internship

Understanding and accurately predicting crop yield is becoming increasingly important today in the face of global food security challenges, and thus, the availability of standardized data and scalable models is the need of the hour. To support this, researchers have developed CY-Bench (Crop Yield Benchmark), a comprehensive dataset that helps forecast maize and wheat yields on a global scale. This research project primarily involved working with the CY-Bench dataset aiming to improve crop yield prediction through machine learning. Initially, papers explaining the CY-Bench dataset and other papers for agriculture modeling were studied and analyzed in detail. The research then progressed …


Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May Aug 2025

Empirical Evaluation Of Bayes Error Rate Bounds In Binary Classification, Riley May

All Graduate Theses and Dissertations, Fall 2023 to Present

Classification tasks are fundamental in statistical machine learning. In classification tasks, a general goal is to build or select a model that can correctly classify data with as few errors as possible. However, for a particular dataset, the minimal number of errors achievable is seldom zero since overlap in the data makes errors unavoidable. As a result, it is often difficult for machine learning practitioners and data scientists to know whether classification errors can be reduced through further refinement. A potential solution to this lies in the Bayes error rate (BER). The BER is the lowest error rate achievable for …


Three-Stage Latent Dynamics Forecasting (T-Ldf) Framework For Shenzhen Metro Passenger Flow Prediction, Tianze Zhang Jul 2025

Three-Stage Latent Dynamics Forecasting (T-Ldf) Framework For Shenzhen Metro Passenger Flow Prediction, Tianze Zhang

Lingnan Theses (MPhil & PhD)

Accurate forecasting of metro passenger flow is vital for efficient urban transportation management and optimal resource allocation in modern cities. Traditional ARIMA-based models effectively capture regular, cyclical patterns but struggle with sudden, nonlinear fluctuations caused by random events such as weather disruptions, special events, or service interruptions. Moreover, existing research predominantly focuses on individual stations, overlooking the complex cross-station interactions inherent in networked metro systems where passenger flows are interconnected across the entire network.

To address these critical limitations, we propose the Three-Stage Latent Dynamics Forecasting (T-LDF) Framework, a novel approach that systematically integrates temporal decomposition, latent dynamics extraction, and …


A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai Jun 2025

A Comparative Study Of Neural Networks And Xgboost Models For Flight Time Prediction, Ioannis Paraschos, Taryn E. Trimble, Eshna Bhargava, Jake Klingler, Benjamin R. Nicolai

Beyond: Undergraduate Research Journal

Flight time prediction plays a crucial role in modern air travel, benefiting airlines and passengers alike. Accurate predictions enable airlines to optimize schedules, allocate resources effectively, and ensure passenger safety and satisfaction. In recent years, machine learning models, such as neural networks and XGBoost, have gained popularity for predicting flight times. This study aims to compare the performance of neural network and XGBoost models in predicting flight times, considering factors such as weather conditions, air traffic control, and aircraft performance. The results indicate that both models are effective, with XGBoost achieving slightly higher accuracy. However, neural networks offer advantages in …


Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh Jun 2025

Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh

University Honors Theses

This study evaluates ChatGPT's ability to forecast influenza rates, such as the number of flu cases, hospitalizations, and death during peak season periods using CDC data, and comparing forecasts against actual results to calculate statistical accuracy and consistency. Influenza forecasting is essential for public health planning, but traditional methods may not always provide timely or accurate predictions. In this research study, ChatGPT was utilized to predict the influenza rates for the following week based on the previous week's data obtained from the FluView surveillance system. The predicted rates were compared to the actual influenza rates to assess the model's overall …


Ai-Driven Personalized Radiotherapy Planning, Nithin Venkatesh, Marco Pota, Maged Shaban Jun 2025

Ai-Driven Personalized Radiotherapy Planning, Nithin Venkatesh, Marco Pota, Maged Shaban

SAML-25 Workshop on Statistical and Machine Learning

The planning of radiation oncology treatment is made more dynamic and individualized by Artificial Intelligence (AI). Routine radiotherapy practice applies normative procedures indifferent to patient-specific parameters such as tumor volume, patient anatomy, and heterogeneity in the delineation of treatment response. Inadequate and over-radiation treatment is the most prevalent outcome. Further, with the inclusion of AI, it can facilitate enhancing the healthcare industry through optimizing radiotherapy using an array of patient information such as molecular profiles and imaging data. The product offers an end-to-end AI-driven solution to all aspects of radiotherapy, from initial consultation (diagnosis) to adaptive treatment planning. All the …


Interpretable Ai In Education: A Comparison Of Glass-Box Models For Predicting Student Success, Jan Glazenborg Jun 2025

Interpretable Ai In Education: A Comparison Of Glass-Box Models For Predicting Student Success, Jan Glazenborg

SAML-25 Workshop on Statistical and Machine Learning

This Master’s thesis addresses early identification of first-year Computer Science students at risk of underperformance by comparing inherently interpretable (“glass-box”) predictive models with the existing Naïve Bayes–based PreSS tool. The PreSS dataset was originally compiled by Quille & Bergin from 692 first-year CS1 students across eleven institutions in Ireland and Denmark, who completed surveys on programming and mathematics backgrounds, gaming habits and a short programming test four to six hours into the course. Seventeen normalized features capturing demographic, academic and behavioural factors were extracted. In this thesis, four machine learning models are evaluated: Naïve Bayes, explainable boosting machines, automatic piecewise …


Benchmarking Energy And Performance Of Parallel Machine Learning Models Using Hardware And Software Power Meters, Urooj Asgher, Tania Malik Jun 2025

Benchmarking Energy And Performance Of Parallel Machine Learning Models Using Hardware And Software Power Meters, Urooj Asgher, Tania Malik

SAML-25 Workshop on Statistical and Machine Learning

The growing reliance on machine learning algorithms across domains such as healthcare, transportation, and finance has led to their increased deployment on high-performance computing platforms. While performance optimization remains a central concern, energy efficiency is emerging as a critical design consideration, particularly in light of global sustainability goals. This study presents a comparative analysis of the energy consumption and performance of serial and parallel implementations of four machine learning algorithms, K-means clustering, Ant Colony Optimization, Logistic Regression, and Random Search. Experiments were conducted on an HPC testbed using both hardware-based and software-based power meters to measure energy consumption. The results …


Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac Jun 2025

Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac

Dartmouth College Ph.D Dissertations

In this dissertation, we take a step towards addressing the major problem of a lack of standardized and rigorous approaches to testing and evaluation of AI systems. Taking inspiration from both the fields of Property Testing and Property Based Testing (for programs), we develop a novel taxonomy of partially overlapping classes of properties of AI systems, including simple properties, compound properties, higher order properties, data relation properties, and architecture-utility properties. We argue that this taxonomy categorizes a diverse set of AI traits -- including accuracy, fairness, robustness, monotonicity, point-wise and global privacy properties, sensitivity, and more -- according to the …


Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim Jun 2025

Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim

Master's Theses

The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …


Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow May 2025

Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow

Capstone Projects

This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …


Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer May 2025

Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer

Data Science Undergraduate Honors Theses

Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …


Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta May 2025

Clinicogenomic Insights For Prostate Cancer Progression, Kelvin Ofori-Minta

Open Access Theses & Dissertations

Prostate cancer (PrCa) remains a critical challenge in precision oncology due to several reasons including its apparent heterogenous condition, recurrence following treatment and rapid progressive forms. Therefore, identifying patients at risk of progression is essential to fast-track therapeutic decisions and improve outcomes. Despite recent advances in genomic and molecular profiling, conventional PrCa risk assessment tools heavily rely on a few clinical parameters, neglecting the prognostic potential of genomic biomarkers in the presence of clinical biomarkers. This study presents a computational pipeline to harmonize and evaluate the prognostic value of clinicogenomic profiles of patients in modelling progression free survival (PFS). PFS, …


Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan May 2025

Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan

Theses and Dissertations

The ability to characterize how information diffuses online is of paramount importance to stakeholders that are interested in tasks such as proposing solutions for mitigating and countering dis/misinformation, predicting user engagement of content in social media, planning marketing campaigns to roll-out products and planning dissemination of political campaign messaging among others. One such facet of learning the dynamics of information diffusion is the ability to predict user engagement or the popularity of a single piece of information as it spreads through an online medium. Existing works in this regard mainly either obfuscate user level information or utilize frameworks that are …


Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih May 2025

Car Price Prediction Using Machine Learning: Analyzing The Dvm-Car Dataset, Yaman Abu Ghareebaih

Electronic Theses and Dissertations

The objective of this study is to predict car prices using machine learning models and the DVM-CAR dataset, which includes over 1.4 million images and car specifi- cations from 899 car models. Key factors such as mileage, engine power, and year of registration were analyzed for their correlation with car prices. Extensive data cleaning was performed, including filling missing values, identifying outliers, and normalizing numerical variables. Discrete variables like car make and body type were encoded using one-hot encoding. Linear relationships were analyzed with Multiple Logistic Regression, and Random Forest models were used for nonlinear patterns. Model performance was evaluated …


Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins Apr 2025

Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins

Honors College Theses

The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …


Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja Jan 2025

Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja

College of Graduate Studies: Theses & Dissertations

Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.

We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …


Impact Of Urban Development On Uv Exposure: A Clustering And Machine Learning Assessment, Taufik Roni Sahroni Mr., Verdi Yasin, Lulut Alfaris, Reza Ariefka, Ruben Cornelius Siagian, Mohammad Alfin Karim, Nana Rahdiana, Ade Suhara Dec 2024

Impact Of Urban Development On Uv Exposure: A Clustering And Machine Learning Assessment, Taufik Roni Sahroni Mr., Verdi Yasin, Lulut Alfaris, Reza Ariefka, Ruben Cornelius Siagian, Mohammad Alfin Karim, Nana Rahdiana, Ade Suhara

Journal of Environmental Science and Sustainable Development

The relocation of Indonesia's capital city is anticipated to promote inclusive economic growth while embracing cultural diversity. However, this transition may affect ultraviolet (UV) radiation exposure patterns. The study investigated variations in UV exposure in the IKN region, focusing on urban development factors such as land use and population density that affect public health, sun protection, and skin cancer prevention. The research hypothesized that UV radiation is significantly correlated with these factors. UV Index data from 2010-2023, a hierarchical clustering method, identifies complex data patterns without determining the number of clusters. XGBoost, a machine learning model, was used for handling …


Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar Dec 2024

Unlocking The Power Of Data: Enhancing Public Policy Through Advanced Data Infrastructure And Language Model Analysis, Zahid Asghar

CBER Conference

Data is the fundamental building block for advancements in artificial intelligence (AI), general AI (GAI), machine learning (ML), and large language models (LLMs). This study emphasizes the critical need for robust data infrastructure, arguing that without it, countries cannot fully benefit from technological advancements in various economic sectors. Governments possess vast repositories of both structured and unstructured data across multiple domains such as the judiciary, parliaments, and civil bureaucracy. However, these potential goldmines remain untapped due to inadequate data management capabilities and a lack of appreciation for the necessity of high-quality data. The research identifies key issues in public data …


A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang Aug 2024

A Machine Learning Based Approach For The Identification Of Fake Bills, Tianyang Lu, Hongyang Pang

Rose-Hulman Undergraduate Mathematics Journal

Fake or counterfeiting currency, which has been around as long as money has existed, is a major economic problem. Since the US dollar is the most popular form of currency globally, it is the most popular currency to counterfeit. The United States Department of Treasury estimates that between $70 million and $200 million in fake bills are in circulation. The Federal Reserve Bank uses special banknote processing systems to count each bill deposited by the bank and examine them for the possibility of counterfeits. These machines have sensors designed to detect general quality of the bills, including paper type, quality …


Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah Aug 2024

Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah

All Dissertations

The intricate interplay of genetic predisposition, environmental influences, and lifestyle acts as the multifactorial landscape of diseases. Understanding this complexity presents a significant challenge. Molecular insights into disease mechanisms, particularly the interactions of DNA, RNA, and proteins with environmental and lifestyle factors, have revolutionized disease diagnosis, prognosis, and treatment. High-throughput technologies, such as next-generation sequencing, generate large amounts of molecular data, holding a wealth of knowledge. These datasets unveil the roles of genes and their interactions with various factors through analysis, shedding light on previously unknown molecular mechanisms underlying disease pathogenesis. Furthermore, they facilitate the discovery of biomarkers crucial for …


Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald May 2024

Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald

Data Science Undergraduate Honors Theses

Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …