Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

University of Arkansas, Fayetteville

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 113

Full-Text Articles in Statistics and Probability

Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder May 2026

Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder

Electrical Engineering and Computer Science Undergraduate Honors Theses

Splines are used for representing complex functions. In statistics, splines can be used for distributional shapes that are difficult to model by traditional parametric approaches. Ramsay (1) uses M-Spline bases to estimate continuous distributions. Semi-Parametric Quantile Regression (SPQR), developed by Xu and Reich (2), models conditional distributions where a neural network is used to estimate the basis function weights that depend on covariates. (3) implements a package for SPQR in R. We build on this by implementing a version of SPQR in Python with PyTorch. By using PyTorch, we can use more sophisticated deep learning architectures than those available in …


Do I Have To Know Someone? The Impact Of Factors Beyond The Playing Field That Affect Mlb Hall Of Fame Voting., Jackson Wollscheid May 2026

Do I Have To Know Someone? The Impact Of Factors Beyond The Playing Field That Affect Mlb Hall Of Fame Voting., Jackson Wollscheid

Economics Undergraduate Honors Theses

With the creation of WAR (Wins above Replacement), it has created an index to measure the player’s value using a variety of statistics relevant to that player’s position. The higher a player’s career WAR then the greater the impact that player had on the game during his career. However, players are not simply inducted into the Hall of Fame on this measure of value but on voting from members of the BBWAA (Baseball Writers’ Association of America). This research investigates external factors outside of standard playing statistics’ influence on Hall of Fame voting, including team affiliation, career length, and other …


Survival Patterns Among Adult And Pediatric Bone Cancer Patients, Ethan Estes May 2026

Survival Patterns Among Adult And Pediatric Bone Cancer Patients, Ethan Estes

Mathematical Sciences Undergraduate Honors Theses

Recently noted, Huang et al. (2023), machine learning (ML) models, while offering great advantages over traditional statistical predictive modeling methods, are less explored in the analysis of survival and other similar time-to-event predictive data modeling. ML methods such as neural networks offer a great deal of promise but need to be further explored to investigate their comparative power in predicting survival outcomes. Focusing specifically on survival analysis in adult and pediatric bone cancer patients, traditional methods, like shown in Emmert-Streib and Dehmer (2019), will be shown with machine learning models using methods in Hothorn, Hornik, and Zeileis (2006). In this …


Explainability In Deep Learning For Density Regression, Dalton James Oxford May 2026

Explainability In Deep Learning For Density Regression, Dalton James Oxford

Graduate Theses and Dissertations

Classical statistical methods focus on explainability and inferential power. Machine learning and deep learning can handle non-linear, high-dimensional data better than traditional methods. In modeling, a clear understanding and interpretation are essential to decision-making. Recent work in quantile regression and extreme modeling has begun to use deep learning due to its performance on high-dimensional, non-linear data. Semi-Parametric Quantile Regression (SPQR) is a nonparametric spline-based approach to quantile regression that estimates the conditional PDF and CDF of the response. Semi-Parametric Quantile Regression for Extremes (SPQRx) is a recent extension of SPQR that provides two features: out-of-sample estimation and accurate extreme-tailed estimation. …


Two Topics In Survival Analysis: Restricted Distance Covariance Test For Non-Proportional Hazard And A New Estimation For Dropout Rate, Ruizhe Yin Dec 2025

Two Topics In Survival Analysis: Restricted Distance Covariance Test For Non-Proportional Hazard And A New Estimation For Dropout Rate, Ruizhe Yin

Graduate Theses and Dissertations

When treatment effects change over time, standard statistical methods, such as the log-rank test and the Cox proportional hazards model, may give misleading results. This dissertation presents the restricted distance covariance (rdcov) test, a nonparametric method that compares survival curves between groups within a chosen study period [0, τ] using right censored data. The statistic measures the dependence between pre-specified group labels and survival times using pairwise distances from Kaplan-Meier estimates. Our method does not rely on the proportional hazards assumption, and it equals zero only when survival functions are identical across groups. Thus, this test can be applied to …


Tree-Based Differential Item Functioning Detection Methods: Exploring Their Performance In Diverse Measurement Scenarios, Nana Amma Berko Asamoah Aug 2025

Tree-Based Differential Item Functioning Detection Methods: Exploring Their Performance In Diverse Measurement Scenarios, Nana Amma Berko Asamoah

Graduate Theses and Dissertations

Despite the availability of numerous methods for detecting differential item functioning (DIF), the continued development and evaluation of innovative, data-driven approaches remains essential. Tree-based methods, in particular, represent a significant advancement in DIF detection. Unlike some traditional techniques, they can simultaneously screen multiple variables for DIF without discretizing continuous variables, and do not require the pre-specification of focal and reference groups; capabilities that are especially valuable in today’s diverse and multifaceted assessment contexts. However, research systematically examining the performance of these methods under realistic measurement conditions is limited. This dissertation, in three simulation studies, critically examines the robustness and practical …


Using Gaussian Process Regression To Learn Thermodynamic Equations Of State With Uncertainty Quantification, Austen T. Lee May 2025

Using Gaussian Process Regression To Learn Thermodynamic Equations Of State With Uncertainty Quantification, Austen T. Lee

Chemical Engineering Undergraduate Honors Theses

This study investigates the use of derivative-informed Gaussian Process (GP) models to estimate thermodynamic behavior across temperature and density by building a Helmholtz-based equation of state. Argon, a stable monatomic gas, was chosen as a case study within the vapor region. The GP model was trained using values of experimentally measurable properties found by taking first and second derivatives of the original potential function. Results show that while the GP model offered uncertainty quantification and informed thermodynamic behavior, it predicted values that deviated from the ground truth depending on the property. The model exhibited high confidence in regions with substantial …


Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer May 2025

Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer

Data Science Undergraduate Honors Theses

Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …


Exploring Latent Mediation Through Bayesian Regularization Methods Of Lasso, Ridge, Horseshoe, Spike-And-Slab, Ethan Harris May 2025

Exploring Latent Mediation Through Bayesian Regularization Methods Of Lasso, Ridge, Horseshoe, Spike-And-Slab, Ethan Harris

Graduate Theses and Dissertations

Regularization is a powerful tool to combat overfitting and drive sparsity in complex models. Regularization was initially applied in regression modeling but has been increasingly utilized in structural equation modeling where its utility in identifying the essential components has helped improve modeling. As structural equation models have increased in complexity both in the number of indicators but also the number of latent factors, researchers have begun to investigate how applying Bayesian regularization to these systems can further push the limits on modeling complex models with limited sample sizes. One area where research is limited is the application of Bayesian regularizations …


Application Of Ordinal Regression Models To Acquired Stress Resistance In Wild Strains Of Saccharomyces Cerevisiae, Carson Stacy May 2025

Application Of Ordinal Regression Models To Acquired Stress Resistance In Wild Strains Of Saccharomyces Cerevisiae, Carson Stacy

Graduate Theses and Dissertations

This thesis explores the application of ordinal regression to the analysis of semi-quantitative growth assays often used when comparing fitness for different strains of the model yeast Saccharomyces cerevisiae. For stress survival assays, yeast stress resistance is measured using an ordered survival score that ranges from 0 (no growth) to 4 (confluent growth). Traditional approaches to analyze this type of data either treats data as a nominal categorical variable or as a continuous numerical variable. These approaches risk loss of information or violation of testing assumptions. In contrast, cumulative logit ordinal regression uses the information contained in the order …


Nonparametric Methods For Bayesian Community Detection In Complex Networks, Kedran Young May 2025

Nonparametric Methods For Bayesian Community Detection In Complex Networks, Kedran Young

Graduate Theses and Dissertations

Network analysis is becoming an increasingly popular interdisciplinary area of study, with emerging interest in fields like sociology, biology, economics, and ecology. Within the niche of network analysis, capturing the community structure of a network is one important achievement that many statisticians have been working toward over recent decades. The most popular modeling technique for latent community detection is the Stochastic Block Model (SBM), which falls into the category of latent variable models and will serve as the baseline model throughout this thesis. SBM is widely regarded as the most effective community detection method as it detects latent community membership …


Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo May 2025

Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo

Graduate Theses and Dissertations

This thesis explores the use of latent factor models to uncover hidden structures in pair wise outcomes derived from Over/Under betting markets in sports betting. Specifically, we implement and evaluate the Eigen model, a latent space model that represents dyadic data using node-specific vectors whose inner product govern edge probabilities. By modeling relationships between teams as adjacency matrices of binary outcomes, we investigate the extent to which the Eigen model captures both homophily, the tendency of similar teams to yield consistent betting results, and stochastic equivalence, where different teams exhibit indistinguishable patterns of Over/Under outcomes. A Bayesian formulation of the …


Comparison Of Statistical And Machine Learning Genomic Prediction Methods In Plant Breeding: Case Studies In Maize And Soybean, Igor Kuivjogi Fernandes Dec 2024

Comparison Of Statistical And Machine Learning Genomic Prediction Methods In Plant Breeding: Case Studies In Maize And Soybean, Igor Kuivjogi Fernandes

Graduate Theses and Dissertations

Plant breeding is essential to increase genetic gain and food production worldwide. This study was conducted to evaluate new ways to use machine learning (ML) to tackle plant breeding challenges, where two ideas were tested — the first chapter focuses on how to combine genetic and environmental data using ML to improve the prediction of maize grain yield in multi-environment trials, while the second chapter centers on how to couple feature selection of molecular markers with ML to enhance prediction of yield in soybean, and, in both cases, ML approaches were compared to well-established statistical methods greatly adopted by the …


Gan With Skip Patch Discriminator For Biological Electron Microscopy Image Generation, Nishith Ranjon Roy Aug 2024

Gan With Skip Patch Discriminator For Biological Electron Microscopy Image Generation, Nishith Ranjon Roy

Graduate Theses and Dissertations

GAN models have been successfully used for image generation in various sections such as real-life objects like human faces, cars, animal faces, landscapes, etc. This work focuses on biological electron microscopy (EM) image generation. Unlike other real-life objects, biological EM images are obtained through electron microscopy techniques to study biological specimens. Electron microscopy offers high resolution and magnification capabilities, making it a powerful tool for visualizing biological structures at the nanoscale. However, using GAN models for biological EM image generation poses challenges due to the complex and unique arrangements of biological structures and the sparse and asymmetrical patterns in EM …


Sparse Neural Network To Enhance Performance Under Limited Parameter Constraints., Nailah Rawnaq Aug 2024

Sparse Neural Network To Enhance Performance Under Limited Parameter Constraints., Nailah Rawnaq

Graduate Theses and Dissertations

Over the past decade, the widespread adoption of deep neural networks has been a breakthrough driven by significant computational advancements. Additionally, the number of parameters of those models is exponentially increasing for performing complex tasks and achieving better performance. However, in most practical cases, often there are constraints in the number of parameters due to limited resources in storage size and computational cost. Network pruning can lead to an optimal solution to this problem. In this thesis, I present supporting evidence to the hypothesis that higher sparsity leads to better performance for a convolution-based neural network. I perform performance studies …


Effects Of Measurement Error In Student Pre-Post Test Score On The Recovery Of The Estimates Of Teachers Value-Added Scores, Merlin J. Kamgue Aug 2024

Effects Of Measurement Error In Student Pre-Post Test Score On The Recovery Of The Estimates Of Teachers Value-Added Scores, Merlin J. Kamgue

Graduate Theses and Dissertations

Abstract Background: Value-added models (VAMs) are statistical tools used to gauge a teacher’s impact on student performance by analyzing standardized test scores. These models project students’ future performance based on past scores and compare the projection to actual outcomes, accounting for differences in student backgrounds. However, the standard error of measurement (SEM) inherent in all measurement tools is often overlooked in VAMs. Aims and Objectives: This study aims to investigate the impact of test reliability on teacher and school score estimates within a Bayesian framework. We will precisely manipulate the reliability of standardized tests by adjusting the standard error of …


Hierarchical Spatial Abundance Models For Migratory Shorebirds, Md Shahbaz Alam Aug 2024

Hierarchical Spatial Abundance Models For Migratory Shorebirds, Md Shahbaz Alam

Graduate Theses and Dissertations

Predicting the distribution and abundance of migratory shorebirds is crucial for effective conservation planning. This research applies hierarchical spatial models to predict counts and spatial variations of three shorebird species: Semipalmated sandpiper (sesa), Ruddy turnstone (rutu), and Whimbrel (whim). Different versions of the Poisson, Negative Binomial, and Hurdle regression models are employed to tackle specific data characteristics, such as overdispersion and excess zeros. Model comparisons are performed in terms of likelihood measures and cross-validation. The Hurdle model for sesa and rutu and the Negative Binomial model for whim effectively captured spatial patterns, highlighting potential hotspots. Mean predictive count further emphasized …


A Two-Part Validation Study Of The Sexual Minority Identity Emotion Scale, Henrietta Kadi Tettey-Tawiah Aug 2024

A Two-Part Validation Study Of The Sexual Minority Identity Emotion Scale, Henrietta Kadi Tettey-Tawiah

Graduate Theses and Dissertations

A number of existing studies indicate that there is some correlation between pride and shame and various behavioral health outcomes like anxiety, depression, and self-harm in the general population. Research also confirms that this relationship is true among sexual minorities as well. This study sets itself up to interrogate this assertion vis-à-vis sexual minority adolescents (SMAs). Consequently, confirmatory factor analysis (CFA) and differential item functioning (DIF) detection analysis were performed on the 35-item Sexual Minority Identity Emotion Scale (SMIES) for SMAs. A tree-based modeling technique, specifically conditional inference trees (CIT), was then employed to explore and make predictions about the …


The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi May 2024

The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi

Computer Science and Computer Engineering Undergraduate Honors Theses

The strategic planning of offensive passing plays in the NFL incorporates numerous variables, including defensive coverages, player positioning, historical data, etc. This project develops an application using an analytical framework and an interactive model to simulate and visualize an NFL offense's passing strategy under varying conditions. Using R-programming and data management, the model dynamically represents potential passing routes in response to different defensive schemes. The system architecture integrates data from historical NFL league years to generate quantified route scores through designed mathematical equations. This allows for the prediction of potential passing routes for offensive skill players in response to the …


Comparing North American Professional Sports League Season Formats Using Monte Carlo Simulation, Lathan Gregg May 2024

Comparing North American Professional Sports League Season Formats Using Monte Carlo Simulation, Lathan Gregg

Industrial Engineering Undergraduate Honors Theses

Each NFL, NBA, and MLB season consists of a regular season, in which teams play a set number of scheduled games and a playoff, in which qualifying teams compete for a championship. At the conclusion of each season, teams are ranked based on their performance throughout the season. This study aims to investigate the ability of each league's season format to accurately rank teams using Monte Carlo simulation. Matches between two teams are simulated by using the team’s assigned strength ranks to calculate a winning probability for each team. The winning probabilities are simulated with different skill values, dictating how …


Automatic Appraisals Of Houses, Sloan Scroggin May 2024

Automatic Appraisals Of Houses, Sloan Scroggin

Graduate Theses and Dissertations

Multiple hedonic models and an automatic appraiser model were used to create a residential house’s estimated sales price. The goal is to use the limited data available to a REALTOR® to estimate the future sales price of a residential home without the aid of pictures of the property or viewing the physical property. The first model automates some of the actions of an appraiser by finding comparable sales based on proximity, based both on distance between houses and characteristics of the houses, and then calculating a weighted average price for an estimated sales price of future sales. If the model …


Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald May 2024

Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald

Data Science Undergraduate Honors Theses

Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …


Examining Award Compliance To Inform Resource Allocation, Jacob Haarala May 2024

Examining Award Compliance To Inform Resource Allocation, Jacob Haarala

Data Science Undergraduate Honors Theses

This project focuses on JB Hunt Transport Inc's intermodal business unit (JBI) by focusing on the challenges associated with Published Pricing and Contractual Pricing. The primary issue revolves around the variance between the awarded freight volumes in Requests for Pricing (RFPs) and the actual volumes realized when the freight is shipped. This discrepancy poses challenges for effective sales planning, revenue goals, and optimal freight network management within JBI. Reporting tools, such as PowerBI, are currently used by JBI to provide insights into award compliance on a weekly basis. However, our goal with this project was to provide a deeper understanding …


A Nonparametric Test For Comparing Survival Functions Based On Restricted Distance Correlation, Quinyang Zhang Dec 2023

A Nonparametric Test For Comparing Survival Functions Based On Restricted Distance Correlation, Quinyang Zhang

Mathematical Sciences Faculty Publications and Presentations

In this article, we propose an omnibus test for comparing two survival functions under non-proportional hazards. The test statistic is based on a product-limit estimate of the restricted distance correlation, which is closely related to the L2 distance between survival curves. The strong consistency is established under mild regularity conditions. Our simulation studies show that the new test has satisfactory power under proportional hazard and various non-proportional hazards settings including delayed treatment effect, diminishing effect, and crossing survival curves; therefore, it can be a competitive alternative to the existing omnibus tests such as Kolmogorov-Smirnov test, Cramer-von Mises test, two-stage …


Bayesian Learning Of Spatiotemporal Source Distribution For Beached Microplastic In The Gulf Of Mexico, David Pojunas Dec 2023

Bayesian Learning Of Spatiotemporal Source Distribution For Beached Microplastic In The Gulf Of Mexico, David Pojunas

Graduate Theses and Dissertations

Over the last several decades, plastic waste has gradually accumulated while slowly degrading in terrestrial and oceanic environments. Recently, there has been an increased effort to identify the possible sources of plastic to understand how they affect vulnerable beaches. This issue is of particular concern in the Gulf of Mexico due to the presence of oil, natural gas, and plastic production. In this thesis, we expand upon existing Bayesian plastic attribution models and develop a rigorous statistical framework to map observed beached microplastics to their sources. Within this framework, we combine Lagrangian backtracking simulations of floating particles using nurdle beaching …


Comparative Analysis Of Teacher Effects Parameters In Models Used For Assessing School Effectiveness: Value-Added Models & Persistence, Merlin J. Kamgue Dec 2023

Comparative Analysis Of Teacher Effects Parameters In Models Used For Assessing School Effectiveness: Value-Added Models & Persistence, Merlin J. Kamgue

Graduate Theses and Dissertations

Longitudinal measures for students have become increasingly popular to estimate the effects of individual teachers and schools. Value-added models are one of the approaches using longitudinal data to evaluate teachers and schools. In the value-added model (VAM) literature, many statistical approaches have been developed and used to estimate teacher or school effects on student learning. This study opted to use a Bayesian multivariate model for evaluating teacher effects. The generalized persistence models can handle longitudinal data, not vertically scaled, allowing for a below-par teacher’s effects correlation across test administrations. This study first generated longitudinal students’ test score data and used …


Analyses Of Effect Indices Across Single-Case Research Designs In Counseling, Cian L. Brown Dec 2023

Analyses Of Effect Indices Across Single-Case Research Designs In Counseling, Cian L. Brown

Graduate Theses and Dissertations

Single case research design (SCRD) is a common methodology used across clinical disciplines to determine treatments effectiveness by comparing treatment conditions to baseline conditions in individual cases, usually among researchers working with smaller samples. Although popular within behavioral disciplines such as special education and behavioral analysis, studies have begun to emerge in counseling. However, guidance and current understanding of the use of SCRD in counseling is limited. A content analysis of counseling journals from 2003 to 2014 yielded only 7 studies using SCRD. In 2015, the flagship counseling journal, Journal of Counseling and Development, published a special issue on the …


Comparing Predictive Performance Of Garch And Stochastic Volatility Models, Swapnaneel Nath Aug 2023

Comparing Predictive Performance Of Garch And Stochastic Volatility Models, Swapnaneel Nath

Graduate Theses and Dissertations

This paper compares the predictive performance of two commonly used financial models, the Generalized Auto-Regressive Conditional Heteroskedasticity (GARCH) model, and the Stochastic Volatility model. Both techniques are used in the finance literature to model returns on an asset; the main difference between the two is that the former holds volatility as deterministic, whereas the latter treats it as a stochastic component. Three 10-year periods (2006-15, 2008-17, and 2010-19) of returns of the S&P-500 Index are used to train the two models. The parameter estimation is done using Hamiltonian Monte Carlo. Then, using Sequential Monte Carlo updates, returns for 2016, 2018, …


A Comparative Study Of Techniques For Non-Monotonic Dependence With Emphasis On Sensitivity To Sample Size, Noise Level And Computational Attributes, Fariha Tasnim Aug 2023

A Comparative Study Of Techniques For Non-Monotonic Dependence With Emphasis On Sensitivity To Sample Size, Noise Level And Computational Attributes, Fariha Tasnim

Graduate Theses and Dissertations

Evaluating association between variables is often of interest by many researchers. To serve this purpose, different association measures have been developed. However, type of relation between variables affects the degree of relationship. Hence, detection of the rela- tionship between variables is germane to measuring the correlation coefficient. With that mindset, here we explored six non-monotonic measure of association techniques and com- pared them with three classical approaches. Due to inconsistency in definition and range of different techniques, it is not feasible to compare the correlation estimates as their nature of variability differ. Therefore, we used permutation test based on Monte …


The Appropriateness Of Outlier Exclusion Approaches Depends On The Expected Contamination: Commentary On André (2022), Daniel Villanova Jul 2023

The Appropriateness Of Outlier Exclusion Approaches Depends On The Expected Contamination: Commentary On André (2022), Daniel Villanova

Marketing Faculty Publications and Presentations

In a recent article, André (2022) addressed the decision to exclude outliers using a threshold across conditions or within conditions and offered a clear recommendation to avoid within-conditions exclusions because of the possibility for large false-positive inflation. In this commentary, I note that André’s simulations did not include the situation for which within-conditions exclusion has previously been recommended—when across-conditions exclusion would exacerbate selection bias. Examining test performance in this situation confirms the recommendation for within-conditions exclusion in such a circumstance. Critically, the suitability of exclusion criteria must be considered in relationship to assumptions about data-generating mechanisms.