Open Access. Powered by Scholars. Published by Universities.®

Statistical Models Commons

Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 31 - 60 of 84

Full-Text Articles in Statistical Models

Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh Jun 2025

Forecasting Influenza Rates Using Machine Learning: A Study Of Chatgpt's Predictive Accuracy, Sara Saleh

University Honors Theses

This study evaluates ChatGPT's ability to forecast influenza rates, such as the number of flu cases, hospitalizations, and death during peak season periods using CDC data, and comparing forecasts against actual results to calculate statistical accuracy and consistency. Influenza forecasting is essential for public health planning, but traditional methods may not always provide timely or accurate predictions. In this research study, ChatGPT was utilized to predict the influenza rates for the following week based on the previous week's data obtained from the FluView surveillance system. The predicted rates were compared to the actual influenza rates to assess the model's overall …


Bandwagon Behavior In Major League Baseball, Daniel E. Erro Jun 2025

Bandwagon Behavior In Major League Baseball, Daniel E. Erro

Master's Theses

This study investigates “bandwagon” behavior among Major League Baseball (MLB) fans by analyzing Google search interest data from 2004 to 2019. Drawing on publicly available information from Google Trends, the analysis explores how fluctuations in search activity align with team performance during both the regular season and postseason. Hierarchical linear models are used to estimate expected levels of fan interest based on team performance and market characteristics. Deviations from these expectations during the regular season are interpreted as evidence of bandwagon or anti-bandwagon behavior. A drop-off in interest following playoff elimination is also examined to capture shifts in fan attention …


How Argentina Won The 2022 Fifa Men's World Cup: A Data Story, Aniruddha Parthasarathy Jun 2025

How Argentina Won The 2022 Fifa Men's World Cup: A Data Story, Aniruddha Parthasarathy

Dissertations, Theses, and Capstone Projects

This project analyzes Argentina’s 2022 FIFA Men’s World Cup win using open-source football (soccer) data. The project evaluates the team’s performance at a micro-level across three domains: without possessing the ball, possessing the ball and the team’s in-game management tactics. A statistical framework, i.e., multiple linear regression modeling, was used to identify the five key defensive actions influencing the Argentinian team’s intensity of pressure applied, and visualized by heatmaps and time-segmented plots. More specifically, an Expected Threat (xT) analysis quantified the threat or danger from passes and progressive carries (moving the ball at least 10 meters), revealing that Lionel Messi’s …


Welfare Implication Of Alternative Tax Rates Adjustment Policy In Nigeria: A Dsge Analysis, Umar B. Ibrahim, Isah F. Abubakar Jun 2025

Welfare Implication Of Alternative Tax Rates Adjustment Policy In Nigeria: A Dsge Analysis, Umar B. Ibrahim, Isah F. Abubakar

CBN Journal of Applied Statistics (JAS)

This study sets out to determine the desirable policy adjustment in the tax rate for Nigeria that ensures the least welfare cost. A calibrated small open-economy New Keynesian Dynamic Stochastic General Equilibrium (NKDSGE) model of the Nigerian economy is applied to achieve this objective. Within this framework, we examined the impact of an increase in value-added tax (VAT) rate from 7.5 to 15 percent on key macroeconomic variables relative to the impact of an increase in company income tax (CIT) rate from 30 to 35 percent on macroeconomic variables. Furthermore, we examined the welfare costs of the increases in the …


Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac Jun 2025

Property Testing Ai: An Efficient Frontier, Paul Sopher Lintilhac

Dartmouth College Ph.D Dissertations

In this dissertation, we take a step towards addressing the major problem of a lack of standardized and rigorous approaches to testing and evaluation of AI systems. Taking inspiration from both the fields of Property Testing and Property Based Testing (for programs), we develop a novel taxonomy of partially overlapping classes of properties of AI systems, including simple properties, compound properties, higher order properties, data relation properties, and architecture-utility properties. We argue that this taxonomy categorizes a diverse set of AI traits -- including accuracy, fairness, robustness, monotonicity, point-wise and global privacy properties, sensitivity, and more -- according to the …


Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim Jun 2025

Opening The Black Box With Regal: A Novel Explainable Ai Approach To Uncover Key Predictors In Search And Rescue Success, Brandon Hyunjun Kim

Master's Theses

The outcome of a search and rescue (SAR) operation is influenced by a complex, non-linear interplay among numerous factors, including geographic context, subject-specific characteristics, and environmental conditions. The high dimensionality and intricate dependencies among these variables pose significant challenges to traditional exploratory modeling approaches, limiting their ability to uncover meaningful patterns and relationships associated with mission success. This study introduces Rules Based Explanations for Generated neighborhoods Around Localized cases (REGAL), a novel adaptation of the Local Interpretable Model-agnostic Explanations (LIME) framework to explain deep multimodal neural networks and what key features it assesses to determine search and rescue success. REGAL …


Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow May 2025

Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow

Capstone Projects

This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …


Contribution Of Various Factors On The Rate Of Traffic Accidents In The Us, Martin Mnatsakanyan May 2025

Contribution Of Various Factors On The Rate Of Traffic Accidents In The Us, Martin Mnatsakanyan

Undergraduate Research Symposium Lightning Talks

Background:

Until 2020, the number of traffic accidents has been steadily decreasing. After 2020, the number started increasing until 2022, then started s lowly decreasing again. Most drivers aren’t fully aware of the reason behind all of these accidents.


Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer May 2025

Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer

Data Science Undergraduate Honors Theses

Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …


Aleci: An R Package For Non-Parametric Confidence Intervals On Accumulated Local Effects Plots, Matthew R. Lister May 2025

Aleci: An R Package For Non-Parametric Confidence Intervals On Accumulated Local Effects Plots, Matthew R. Lister

All Graduate Reports and Creative Projects, Fall 2023 to Present

Machine learning models can take a collection of inputs and craft an output. The mathematical formulas these models use to calculate their outputs easily become too complex or time consuming for a human to analyze. Collectively, we refer to these as black box models. Accumulated local effects plots (ALE) are a method for adding interpretability and visibility into the effects that individual variables contribute to the predictions made by black box models. The method designed by D.W. Apley calculates equally spaced point estimates of the response value to construct a graph across the range of the variable of interest. AleCI …


Nonlinear Power Function Model Changepoint Detection., Jacob Steven Townson May 2025

Nonlinear Power Function Model Changepoint Detection., Jacob Steven Townson

Electronic Theses and Dissertations

Most work surrounding changepoint analysis focuses on linear models. This dissertation explores changepoint detection in nonlinear power function models, specifically focusing on models where the constant multiplier and power are the parameters to be estimated in addition to the changepoint parameter. The study assumes an asymptotic framework as the number of observations approaches infinity. The study explores various model fitting algorithms, and decides to employ the Newton-Raphson method for parameter estimation, with a custom implementation developed to optimize the process. The research first establishes the strong consistency of estimators for the model without a changepoint. Building on this result, consistency …


Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan May 2025

Expressive And Interpretable User Engagement Prediction Using Multivariate Survival Processes, Akshay Aravamudan

Theses and Dissertations

The ability to characterize how information diffuses online is of paramount importance to stakeholders that are interested in tasks such as proposing solutions for mitigating and countering dis/misinformation, predicting user engagement of content in social media, planning marketing campaigns to roll-out products and planning dissemination of political campaign messaging among others. One such facet of learning the dynamics of information diffusion is the ability to predict user engagement or the popularity of a single piece of information as it spreads through an online medium. Existing works in this regard mainly either obfuscate user level information or utilize frameworks that are …


Innovative Methods For The Design And Analysis Of Phase Ii Clinical Trials, Feng Tian May 2025

Innovative Methods For The Design And Analysis Of Phase Ii Clinical Trials, Feng Tian

Dissertations and Theses (Open Access)

Drug development has become increasingly time-consuming, costly, and risky in recent years. There is significant potential for improving clinical trial designs, particularly for phase II trials, which play a critical role in the drug development process. Innovative methods are especially necessary for addressing key challenges in phase II trials in terms of dose-ranging study, patient population selection, and decentralized clinical trials (DCTs). This dissertation presents a comprehensive set of methodologies that address these critical issues with three projects. The first project introduces a Bayesian adaptive dose-ranging design that integrates both efficacy and toxicity data to evaluate each dose comprehensively. The …


Application Of Ordinal Regression Models To Acquired Stress Resistance In Wild Strains Of Saccharomyces Cerevisiae, Carson Stacy May 2025

Application Of Ordinal Regression Models To Acquired Stress Resistance In Wild Strains Of Saccharomyces Cerevisiae, Carson Stacy

Graduate Theses and Dissertations

This thesis explores the application of ordinal regression to the analysis of semi-quantitative growth assays often used when comparing fitness for different strains of the model yeast Saccharomyces cerevisiae. For stress survival assays, yeast stress resistance is measured using an ordered survival score that ranges from 0 (no growth) to 4 (confluent growth). Traditional approaches to analyze this type of data either treats data as a nominal categorical variable or as a continuous numerical variable. These approaches risk loss of information or violation of testing assumptions. In contrast, cumulative logit ordinal regression uses the information contained in the order …


Nonparametric Methods For Bayesian Community Detection In Complex Networks, Kedran Young May 2025

Nonparametric Methods For Bayesian Community Detection In Complex Networks, Kedran Young

Graduate Theses and Dissertations

Network analysis is becoming an increasingly popular interdisciplinary area of study, with emerging interest in fields like sociology, biology, economics, and ecology. Within the niche of network analysis, capturing the community structure of a network is one important achievement that many statisticians have been working toward over recent decades. The most popular modeling technique for latent community detection is the Stochastic Block Model (SBM), which falls into the category of latent variable models and will serve as the baseline model throughout this thesis. SBM is widely regarded as the most effective community detection method as it detects latent community membership …


Sabrina Vs Steph: The Battle Between The Wnba And Nba, Naysha Mcgriff Apr 2025

Sabrina Vs Steph: The Battle Between The Wnba And Nba, Naysha Mcgriff

Symposium of Student Scholars

The average salary of a Women’s National Basketball Association (WNBA) player is 110 times less than a National Basketball Association (NBA) player’s. Despite growing WNBA viewership, gender inequality in sports remains high, with critics claiming female athletes are less skilled. Gender bias in sports is severely understudied, making direct comparisons to men’s leagues unfair due to long-term lack of investment in women’s sports. This study investigates whether the perceived disparity in skill levels between WNBA and NBA players' is genuine or influenced more by external factors by developing an unbiased measure of player efficiency to compare athletic performance. This dataset …


Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins Apr 2025

Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins

Honors College Theses

The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …


Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun Apr 2025

Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun

SMU Data Science Review

Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …


Profiting On The Kentucky Derby, Bailey Korfhage Apr 2025

Profiting On The Kentucky Derby, Bailey Korfhage

Undergraduate Theses

This paper analyzes the quantitative data of horses that ran in the Kentucky Derby to recognize statistically significant variables to predict the horse that comes in first or in-the-money. This analysis is specific to the post-implementation of the points system that began for the 2013 Kentucky Derby. Churchill Downs, the host of the Kentucky Derby, changed the methodology of qualification for a horse to enter the race; instead of qualifying with highest earnings in lifetime starts, the institution implemented a points system that awarded different proportions of points depending on the value of various prep races leading up to the …


Leveraging Benford’S Law And Machine Learning For Financial Fraud Detection, Benjamin R. Fu Apr 2025

Leveraging Benford’S Law And Machine Learning For Financial Fraud Detection, Benjamin R. Fu

Cybersecurity Undergraduate Research Showcase

Financial fraud, particularly credit card fraud, continues to pose substantial challenges to financial institutions due to its increasing frequency and impact on consumer trust. While traditional rule-based methods have provided foundational defenses, their limitations in scalability and adaptability have accelerated the adoption of machine learning (ML) techniques. Concurrently, Benford’s Law—a statistical principle often used in forensic accounting—has demonstrated efficacy in detecting anomalies within naturally occurring numerical datasets. This study explores a hybrid fraud detection approach that integrates Benford’s Law with supervised machine learning algorithms, including Logistic Regression, Random Forest, and k-Nearest Neighbors. Using the publicly available European credit card fraud …


Irreversible K-Threshold Number Ck(G) And Saturation Probability P[G] For Corona Product And Double Corona Product Graphs, Eric J. Moon, Soumya Bhoumik, Paul Flesher Apr 2025

Irreversible K-Threshold Number Ck(G) And Saturation Probability P[G] For Corona Product And Double Corona Product Graphs, Eric J. Moon, Soumya Bhoumik, Paul Flesher

SACAD: Scholarly Activities

We discuss the Irreversible k-conversion process for graphs, where a vertex becomes saturated and remains saturated indefinitely if at least k of its neighbors are saturated. We investigate sets S0, which when initially saturated, lead to complete graph saturation. We are interested in the minimum |S0| = Ck(G), called the k-threshold number. We consider the construction of the Corona Product Graphs (of Cn and Kp). Additionally, we extend our analysis by defining and exploring Double Corona Product Graphs (of Cn and Kp). Then we incorporate …


Nonparametric Finite Mixture Of Ising Graphical Models, Manal Hamadi Alloqmani Apr 2025

Nonparametric Finite Mixture Of Ising Graphical Models, Manal Hamadi Alloqmani

Dissertations

Statistical applications in fields such as bioinformatics, genomics, speech processing, image processing, and communications often involve large-scale models in which thousands or millions of random variables are linked in complex ways. Graphical models provide a general methodology for approaching these problems, and indeed many of the models developed by researchers in these applied fields are instances of the general graphical model formalism. This formalism gives a nice framework for capturing complex dependencies among the random variables and building a large-scale model for high-dimensional data. Recently, high-dimensional data are more assumed to come from one population and follow a parametric or …


Statistical Inference For Noisy Matrix Completion Incorporating Auxiliary Information, Shujie Ma, Po-Yao Niu, Yichong Zhang, Yinchu Zhu Apr 2025

Statistical Inference For Noisy Matrix Completion Incorporating Auxiliary Information, Shujie Ma, Po-Yao Niu, Yichong Zhang, Yinchu Zhu

Research Collection School Of Economics

This article investigates statistical inference for noisy matrix completion in a semi-supervised model when auxiliary covariates are available. The model consists of two parts. One part is a low-rank matrix induced by unobserved latent factors; the other part models the effects of the observed covariates through a coefficient matrix which is composed of high-dimensional column vectors. We model the observational pattern of the responses through a logistic regression of the covariates, and allow its probability to go to zero as the sample size increases. We apply an iterative least squares (LS) estimation approach in our considered context. The iterative LS …


Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns Mar 2025

Linking Water Quality And Climate Change To Long-Term Trends In Species Abundance In Norwalk Harbor, Viktoria Savatorova, Aidan Kieft, Nicole C. Spiller, Kasey Burns

Spora: A Journal of Biomathematics

This study examines the effects of environmental changes on fish populations in Norwalk Harbor, focusing on winter flounder (Pseudopleuronectes americanus), cunner (Tautogolabrus adspersus), northern pipefish (Syngnathus fuscus), and naked goby (Gobiosoma bosci) as examples of species responding to climate-related shifts. We analyze how water temperature, salinity, and dissolved oxygen correlate with fish abundance. To assess statistically significant differences in catch per unit effort (CPUE) across harbor regions, we applied the Kruskal-Wallis test followed by Dunn's post-hoc test. Seasonal variations in CPUE were examined by comparing monthly catch data for each species. K-means …


Discounting Effect Size When Borrowing External Data In Clinical Studies, Zhuanzhuan Ma, Chul Ahn, Bin Wang, Xuefeng Li Mar 2025

Discounting Effect Size When Borrowing External Data In Clinical Studies, Zhuanzhuan Ma, Chul Ahn, Bin Wang, Xuefeng Li

Research Symposium

Background: When borrowing information from external data to augment a current trial, many available methods discount the sample size but retain the effect size from previous studies. Discounting the sample size is just one way to discount the prior information. It may not be appropriate if the underlying assumption of unbiased treatment effect does not hold, for example, when the treatment effect in the historical study is likely higher than the one expected in the current trial.

Methods: To tackle this potential issue, we study some methods to shrink the effect size from previous studies assuming that the prior effect …


Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma Mar 2025

Sparse Bayesian Variable Selection Using Global-Local Shrinkage Priors For The Analysis Of Cancer Datasets, Zhuanzhuan Ma

Research Symposium

Background: With a rapid development of data collection technology, high dimensional data, whose model dimension k may be growing or much larger than the sample size n, is becoming increasingly prevalent in different fields of study, such as ecology, genetics, among others. This data deluge is introducing new challenges to traditional statistical procedures and theories and is thus generating a renewed interest in the problems of variable selection and classification in high dimensional regression models. In large k, small n settings, variable selection is usually the first step for dimension reduction to uncover significant covariates, which contribute to …


Effect Of Pre-Adsorbed Species On High-Pressure Adsorption Of Methane In Zeolite 5a Using Grand Canonical Monte Carlo (Gcmc) Simulations, Kanhamardi Lao, Brooks D. Rabideau Mar 2025

Effect Of Pre-Adsorbed Species On High-Pressure Adsorption Of Methane In Zeolite 5a Using Grand Canonical Monte Carlo (Gcmc) Simulations, Kanhamardi Lao, Brooks D. Rabideau

Shelby Hall Graduate Research Forum Posters

Natural gas upgrading, which removes impurities from methane (CH4), is essential for industrial applications, including liquefied natural gas (LNG) production and power generation, as well as for residential use. Removing non-hydrocarbon impurities such as carbon dioxide (CO2), nitrogen (N2), and water vapor (H2O), among others, along with separating heavier hydrocarbon gases from raw natural gas, is required to achieve high- purity methane and prevent pipeline corrosion. Zeolite 5A is a microporous aluminosilicate material with a pore size of approximately 5 Å, containing sodium and calcium cations that balance the framework’s negative charge. Its structure offers high thermal stability and a …


Filters For Forecasting Crop Health: Analyzing And Projecting The Temporal Evolution Of Landsat Ndvi Data Using Dynamic Linear Models And The Kalman Filter, Kamal Albousafi, Hossein Moradi, Jung-Han Kimn Feb 2025

Filters For Forecasting Crop Health: Analyzing And Projecting The Temporal Evolution Of Landsat Ndvi Data Using Dynamic Linear Models And The Kalman Filter, Kamal Albousafi, Hossein Moradi, Jung-Han Kimn

SDSU Data Science Symposium

Accurately forecasting food availability is a critical task. One approach involves utilizing remote sensing data, such as satellite images, to observe the health of crop fields using different Vegetation Indices (VI). The Normalized Difference Vegetation Index (NDVI) provides a sound metric to track the “greenness” of crops over time. In this research, we develop statistical models that capture the dynamics of NDVI time series data to make better predictions of its future values. The median NDVI of the pixels of a farm located in Edmunds County, South Dakota, is obtained using imagery from the Landsat 5 and Landsat 8 satellites, …


Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari Jan 2025

Kroger Post-Pandemic Customer Segmentation, Mario Mata, Joey Truitt, Renn Spigelmyer, Dhanuja Kasturiratna, Lisa Holden, Nitish Baidya, Hanna Tafari

Posters-at-the-Capitol

The grocery retail industry landscape has changed greatly in the wake of the pandemic. Specifically, delivery and pickup services have become more popular and customer buying habits have evolved. At the same time, improvements in data collection and analysis have allowed grocery marketing strategies to become highly individualized.

We worked with 84.51, an analytics firm, to identify customer segments for the Kroger Company based on data from 2023. Using clustering techniques, we organized customers into groups, or segments, based on similar characteristics. We identified and profiled four distinct groups of customers. Three segments were characterized by high frequency and spending …


Estimation And Model Misspecification For Recurrent Event Data With Covariates Under Measurement Errors, Ravinath Alahakoon, Gideon K.D. Zamba, Xuerong Meggie Wen, Akim Adekpedjou Jan 2025

Estimation And Model Misspecification For Recurrent Event Data With Covariates Under Measurement Errors, Ravinath Alahakoon, Gideon K.D. Zamba, Xuerong Meggie Wen, Akim Adekpedjou

Mathematics and Statistics Faculty Research & Creative Works

For subject i, we monitor an event that can occur multiple times over a random observation window [0, (Formula presented.)). At each recurrence, p concomitant variables, (Formula presented.), associated to the event recurrence are recorded—a subset ((Formula presented.)) of which is measured with errors. To circumvent the problem of bias and consistency associated with parameter estimation in the presence of measurement errors, we propose inference for corrected estimating equations with well-behaved roots under an additive measurement errors model. We show that estimation is essentially unbiased under the corrected profile likelihood for recurrent events, in comparison to biased estimations under a …