Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Electronic Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 1 - 30 of 58

Full-Text Articles in Applied Statistics

Modeling Mean And Variability Of Anxiety In Ecological Momentary Assessment Data Using Mixed-Effects Location–Scale Models, Trenzy Odero Aug 2026

Modeling Mean And Variability Of Anxiety In Ecological Momentary Assessment Data Using Mixed-Effects Location–Scale Models, Trenzy Odero

Electronic Theses and Dissertations

Ecological Momentary Assessment is a method of collecting repeated measures of people in real time within natural environments. This results in hierarchical data that has a significant amount of variation at the person level. The traditional linear mixedeffects models assume that the residual variance is constant, which might not be true when the residual variance varies among individuals as well as in time. This thesis uses mixed-effects location-scale (MELS) models to model the mean and variance of an EMA outcome together. By introducing the possibility of variability in residual variance within and across individuals and with covariates, the MELS framework …


A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings Dec 2025

A Leslie System For A Demographic Simulation: From An Actuarial Point Of View, David Kings

Electronic Theses and Dissertations

This thesis develops a discrete stochastic linear systems interpretation of age–stage demographic evolution grounded in Leslie operators and realized in a discrete-event simulation implemented with salabim. The central claim is that one annual cycle of the simulation constitutes a cone-preserving, stochastic affine transformation on a high- dimensional population state vector indexed by age, sex, marital status, household type, employment, and education, and that the composition of yearly operators yields a random matrix product whose top Lyapunov exponent is the stochastic counterpart of the Perron–Frobenius growth rate (Caswell, 2001; Tuljapurkar, 1997)[1, 2]. The actuarial bridge is constructed by mapping simulated survival …


Performance Of The Two Sample Likelihood Ratio Test Under A Nested Dirichlet: A Simulation Study, Edwina Agyeman Aug 2025

Performance Of The Two Sample Likelihood Ratio Test Under A Nested Dirichlet: A Simulation Study, Edwina Agyeman

Electronic Theses and Dissertations

Compositional data analysis (CoDA) addresses multivariate data constrained to a constant sum, such as proportions or percentages. Originating from early warnings regarding misinterpretation by Pearson (1897), the field was formalized by John Aitchison in 1986, whose foundational work remains highly influential. Over time, new modeling techniques and visualization tools have advanced the field, as noted by Greenacre et al. More recently, Turner et al. proposed an approach based on the Nested Dirichlet Distribution (NDD), which accommodates more flexible dependence structures than the standard Dirichlet model. This thesis builds on the methodology of Turner et al. Chapter 1 introduces the nature …


Evaluating Alpha Spending Functions Applied To Observational Time-To-Event Analysis, Moses Torgbenu Aug 2025

Evaluating Alpha Spending Functions Applied To Observational Time-To-Event Analysis, Moses Torgbenu

Electronic Theses and Dissertations

This thesis explores the theoretical foundation of the alpha spending approach and extends its application beyond the conventional setting of randomized controlled trials (RCTs) to observational studies with time to event analyses. In these less structured environments, key design parameters such as the total number of events are often unknown, posing challenges for the standard implementation of sequential analysis methods.

Through simulation studies, this research delivers several important contributions. First, it presents a modified approach that uses calendar time to define the timing of interim analyses while relying on event-based information to estimate the correlation among test statistics. This adjustment …


Simultaneous Application Of Multiple Process Control Rules, Tran B. Ngo Aug 2025

Simultaneous Application Of Multiple Process Control Rules, Tran B. Ngo

Electronic Theses and Dissertations

Statistical Process Control (SPC) charts are tools used in quality control to monitor and analyze the stability of a process over time. This study evaluates the effectiveness of eight individual Western Electric rules, also known as WECO rules, and the various combinations of these rules with Shewhart rule (or WECO rule 1) to SPC charts. As more rules are added to a process control scheme with Rule 1, there is a trade-off: a higher false out-of-control signal rate but an increase in sensitivity, that is the ability of a specified process control scheme to capture a true out-of-control signal. This …


Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares Aug 2025

Estimation Methods For Bayesian Exponential Random Graph Models Under The Horseshoe Prior., Pamela Linares

Electronic Theses and Dissertations

Networks are powerful tools for modeling the complexity of social interactions, biological systems, and information spread. A leading statistical frameworks for analyzing network data are Exponential Random Graph Models (ERGMs), which provide a principled approach to capturing structural dependencies. However, ERGMs remain challenging to estimate, especially in sparse or high-dimensional settings where models suffer from degeneracy and unstable parameter inference. This paper proposes a penalized Bayesian approach to ERGMs that utilizes the horseshoe prior, a sparsity-inducing global-local shrinkage prior. This prior offers robust regularization while preserving important signals, improving estimation by shrinking irrelevant parameters and reducing the impact of extreme …


Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson May 2025

Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson

Electronic Theses and Dissertations

The beef industry plays a vital role in global agriculture, with carcass quality and consumer preference being key determinants of market success. This thesis examines predictive modeling techniques for estimating the Total Score of beef carcasses, a composite measure representing yield and quality, primarily used by the Nebraska Cattlemen Association. Using data from the Nebraska Cattlemen’s Foundation Retail Value Steer Challenge (2000–2023), the study compares the performance of First Order Multiple Linear Regression (MLR) with three machine learning techniques: K-Nearest Neighbors (KNN), Random Forest, and Gradient Boosting Machine (GBM).

The analysis focuses on six key predictors: Hot Carcass Weight, Back …


Application Of Deep Learning On Gage R&R For Anomaly Detection, Oluwatope Richard Ojo May 2025

Application Of Deep Learning On Gage R&R For Anomaly Detection, Oluwatope Richard Ojo

Electronic Theses and Dissertations

This thesis explores the application of deep learning techniques, specifically autoencoder based models, to enhance anomaly detection within Gage Repeatability and Reproducibility (Gage R&R) studies—an essential component of Measurement System Analysis (MSA) in quality engineering. Traditional Gage R&R methodologies, while effective for linear and low-dimensional data, exhibit limitations in detecting subtle, nonlinear variations in complex measurement systems. To address this challenge, an unsupervised autoencoder was developed and trained on a synthetically generated dataset comprising 2,500 voltage measurements (5V and 33V) derived using Generative Adversarial Networks (GANs) based on real-world manufacturing data measurements.

The proposed autoencoder model achieved a 95th percentile-based …


Value-Based Healthcare Reimagined: A Mixed-Methods Study On Behavioral Health Clinicians' Perspectives, Amanda L. Strickland Mar 2025

Value-Based Healthcare Reimagined: A Mixed-Methods Study On Behavioral Health Clinicians' Perspectives, Amanda L. Strickland

Electronic Theses and Dissertations

This study explores how behavioral health clinicians perceive Value-Based Healthcare (VBHC), a model designed by Porter and Teisberg (2006) to improve outcomes relative to costs. While widely promoted in healthcare reform, VBHC poses unique challenges when applied to behavioral health settings. Using an explanatory mixed-methods design, this study first assessed clinicians’ awareness of VBHC through a survey of 23 licensed clinicians at a Community Mental Health Center (CMHC) in Colorado. Quantitative findings revealed that one-third of participants were aware of VBHC with awareness differing by role prompting further exploration in a qualitative phase. Semi-structured interviews with eight clinicians provided deeper …


Financialization And Price Volatility: An Empirical Analysis On Speculation In The Oil Market, Audry F. Oliveira Carnivale Nov 2024

Financialization And Price Volatility: An Empirical Analysis On Speculation In The Oil Market, Audry F. Oliveira Carnivale

Electronic Theses and Dissertations

Financialization has facilitated the trade of futures contracts because of deregulatory policies increasing speculation. Speculation has created a more fragile market inducing riskier investments and aggravating price volatility. Three post-Keynesian theories, the financial instability hypothesis, money manager capitalism and markup, explain how policy altering the banking structure has developed financialization from lax regulation. Previous research has emphasized supply and demand as the main determinants of oil price changes, but it’s important to consider how structural changes from policy stimulating a more financialized economy has impacted volatility. With Brent Crude oil price data and West Texas Intermediate (WTI) open interest and …


Investigating Servant Leadership Measurement: A Mixed-Methods Study Integrating Content Analysis And Meta-Analysis, Kai Torsten Schramm Aug 2024

Investigating Servant Leadership Measurement: A Mixed-Methods Study Integrating Content Analysis And Meta-Analysis, Kai Torsten Schramm

Electronic Theses and Dissertations

This study aimed to understand the similarities and differences among servant leadership measures and the variations in their effect sizes on job performance and job satisfaction. This paper explores how the items in servant leadership measures portrayed the servant leadership construct and how these relate to the outcomes. The researcher used an exploratory sequential mixed methods design. Which involved a qualitative content analysis of the measurement items and a meta-analysis of outcomes, considering the findings from the content analysis. Six key categories determined the three main themes: selfless generosity, inspiring influence, adaptive humility, integrity, empowering, and harmonious engagement. The three …


The Impact Of “Multiple Looks” When Performing Survival Analysis, Quentin Eloise Aug 2024

The Impact Of “Multiple Looks” When Performing Survival Analysis, Quentin Eloise

Electronic Theses and Dissertations

Survival analysis is a critical statistical method in healthcare to assess patient treatment effects and disease progression. Another critical area of statistical methodology in health care is the practice of adaptive designs. Adaptive designs allow for interim analyses to take place during a study and various decisions and actions can take place more ethically. This is beneficial for studies that take multiple years to complete and allows administrators and healthcare providers to make sound decisions as early as possible. A challenging aspect of adaptive designs is that the number of interim analyses is known in advance which is applicable in …


A Uniformly Most Powerful Test For The Mean Of A Beta Distribution, Richard Ntiamoah Kyei Aug 2024

A Uniformly Most Powerful Test For The Mean Of A Beta Distribution, Richard Ntiamoah Kyei

Electronic Theses and Dissertations

The beta distribution is used in numerous real-world applications, including areas such as manufacturing (quality control) and analyzing patient outcomes in health care. It also plays a key role in statistical theory, including multivariate analysis of variance (MANOVA) and Bayesian statistics. It is a flexible distribution that can account for many different characteristics of real data. To our surprise, there has been very little work or discussion on performing statistical hypothesis testing for the mean when it is reasonable to assume that the population is beta distributed. Many analysts conduct traditional analyses using a t-test or nonparametric approach, try transformations, …


Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh Aug 2024

Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh

Electronic Theses and Dissertations

This study explores innovative approaches to constructing confidence intervals for the population standard deviation, σ, in non-normal data scenarios. While the sample standard deviation, s, is widely used, its reliability is compromised when dealing with skewed or heavy-tailed distributions and exhibits sensitivity to outliers. Our research addresses these limitations by investigating alternative estimation methods that offer greater robustness and accuracy.


Capturing Latent Abilities And Latent Capacities Of Professional Golfers Using Nonlinear Mixed Effects Growth Modeling, Mac Wetherbee Jun 2024

Capturing Latent Abilities And Latent Capacities Of Professional Golfers Using Nonlinear Mixed Effects Growth Modeling, Mac Wetherbee

Electronic Theses and Dissertations

This study demonstrates an effective and innovative approach to measuring the latent athletic abilities and capacities of professional golfers. I used nonlinear mixed effects growth modeling (e.g., Dynamic Measurement Modeling) to measure professional golfers’ ability levels and capacities for improvement. I accomplished this using a two-stage modeling approach. First, a crossed linear mixed effects model estimated each player’s ability level in each year. In the second stage, I used the results from the first stage to estimate several candidate nonlinear growth trajectories for players’ abilities over time. The quadratic growth trajectory was the best-fitting of these trajectories and was used …


Advancement Of Iterative Optimization Technology Algorithms Toward Calibration-Free Process Analytical Technology Applications, Adam Rish May 2024

Advancement Of Iterative Optimization Technology Algorithms Toward Calibration-Free Process Analytical Technology Applications, Adam Rish

Electronic Theses and Dissertations

The expansion of spectroscopic process analytical technology (PAT) tools within the pharmaceutical industry has the potential to elevate the current state-of-the-art of pharmaceutical manufacturing by offering opportunities for reduced quality testing times, enhanced process control, and greater production flexibility. Spectroscopic PAT tools are dependent on multivariate models to extract the relevant information from the spectral outputs. However, there is a substantial calibration burden for developing and maintaining these multivariate models that discourages the application of PAT, despite the encouragement from regulators. This has led to an interest in calibration-free methods such as iterative optimization technology (IOT) for spectroscopic PAT that …


Proteomics And Machine Learning For Pulmonary Embolism Risk With Protein Markers, Yaa Amankwah Awuah Dec 2023

Proteomics And Machine Learning For Pulmonary Embolism Risk With Protein Markers, Yaa Amankwah Awuah

Electronic Theses and Dissertations

This thesis investigates protein markers linked to pulmonary embolism risk using proteomics and statistical methods, employing unsupervised and supervised machine learning techniques. The research analyzes existing datasets, identifies significant features, and observes gender differences through MANOVA. Principal Component Analysis reduces variables from 378 to 59, and Random Forest achieves 70% accuracy. These findings contribute to our understanding of pulmonary embolism and may lead to diagnostic biomarkers. MANOVA reveals significant gender differences, and applying proteomics holds promise for clinical practice and research.


The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña Nov 2023

The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña

Electronic Theses and Dissertations

Since the late 1970s, multiple linear regression has been the preferred method for identifying discrimination in pay. An empirical study on this topic was conducted using quantitative critical methods. A literature review first examined conflicting views on using multiple linear regression in pay equity studies. The review found that multiple linear regression is used so prevalently in pay equity studies because the courts and practitioners have widely accepted it and because of its simplicity and ability to parse multiple sources of variance simultaneously. Commentaries in the literature cautioned about errors in model specification, the use of tainted variables, and the …


Cannabidiol Tweet Miner: A Framework For Identifying Misinformation In Cbd Tweets., Jason Turner Aug 2023

Cannabidiol Tweet Miner: A Framework For Identifying Misinformation In Cbd Tweets., Jason Turner

Electronic Theses and Dissertations

As regulations surrounding cannabis continue to develop, the demand for cannabis-based products is on the rise. Despite not producing the psychoactive effects commonly associated with THC, products containing cannabidiol (CBD) have gained immense popularity in recent years as a potential treatment option for a range of conditions, particularly those associated with pain or sleep disorders. However, due to current federal policies, these products have yet to undergo comprehensive safety and efficacy testing. Fortunately, utilizing advanced natural language processing (NLP) techniques, data harvested from social networks have been employed to investigate various social trends within healthcare, such as disease tracking and …


Statistical Inference On Lung Cancer Screening Using The National Lung Screening Trial Data., Farhin Rahman Aug 2023

Statistical Inference On Lung Cancer Screening Using The National Lung Screening Trial Data., Farhin Rahman

Electronic Theses and Dissertations

This dissertation consists of three research projects on cancer screening probability modeling. In these projects, the three key modeling parameters (sensitivity, sojourn time, transition density) for cancer screening were estimated, along with the long-term outcomes (including overdiagnosis as one outcome), the optimal screening time/age, the lead time distribution, and the probability of overdiagnosis at the future screening time were simulated to provide a statistical perspective on the effectiveness of cancer screening programs. In the first part of this dissertation, a statistical inference was conducted for male and female smokers using the National Lung Screening Trial (NLST) chest X-ray data. A …


An Analysis Of All-Cause Mortality On Patients With Sickle Cell Disease And Kidney Disease Using Propensity Score Matching, Adam Garrison May 2023

An Analysis Of All-Cause Mortality On Patients With Sickle Cell Disease And Kidney Disease Using Propensity Score Matching, Adam Garrison

Electronic Theses and Dissertations

In this work, we provide an overview of the Cox proportional hazards model for time to event or survival analysis and the notion of propensity score matching to deal with confounding factors. A full analysis is reported in Chapter 2 concerning mortality for in-center dialysis patients with sickle cell disease to demonstrate the application of a general analysis strategy that has some logistical benefits over more traditional approaches to accounting for confounding variables. We also provide some insight and discussions on the challenges and future research questions that will emerge when trying to implement this strategy as a monitoring tool …


Problems With Machine Learning, High-Dimensional Data And Forecasting Stock Returns, Erik Mekelburg Jan 2023

Problems With Machine Learning, High-Dimensional Data And Forecasting Stock Returns, Erik Mekelburg

Electronic Theses and Dissertations

Using a multi-level ensemble design, we forecast international stock market returns with a novel high-dimensional data set of aggregated cross sectional firm-level predictors. The method includes considerations of model uncertainty, parameter instability, model density and non-linearities with machine learning, shrinkage and model averaging. We provide evidence that it is important to systematically focus on all four sources of forecast failure, shed light on the sparsity/density debate in the stock return forecasting dialogue and contribute interesting findings on the efficacy dimensionality reduction with principal components analysis and partial least squares. The robustness of the approach is demonstrated through applications in four …


Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury Dec 2022

Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury

Electronic Theses and Dissertations

Graphical models determine associations between variables through the notion of conditional independence. Gaussian graphical models are a widely used class of such models, where the relationships are formalized by non-null entries of the precision matrix. However, in high-dimensional cases, covariance estimates are typically unstable. Moreover, it is natural to expect only a few significant associations to be present in many realistic applications. This necessitates the injection of sparsity techniques into the estimation method. Classical frequentist methods, like GLASSO, use penalization techniques for this purpose. Fully Bayesian methods, on the contrary, are slow because they require iteratively sampling over a quadratic …


Statistical Methods For Personalized Treatment Selection And Survival Data Analysis Based On Observational Data With High-Dimensional Covariates., Don Ramesh Dinendra Sudaraka Tholkage Aug 2022

Statistical Methods For Personalized Treatment Selection And Survival Data Analysis Based On Observational Data With High-Dimensional Covariates., Don Ramesh Dinendra Sudaraka Tholkage

Electronic Theses and Dissertations

Due to the wide availability of functional data from multiple disciplines, the studies of functional data analysis have become popular in the recent literature. However, the related development in censored survival data has been relatively sparse. In Chapter 2, we consider the problem of analyzing time-to-event data in the presence of functional predictors. We develop a conditional generalized Kaplan Meier (KM) estimator that incorporates functional predictors using kernel weights and rigorously establishes its asymptotic properties. In addition, we propose to select the optimal bandwidth based on a time-dependent Brier score. We then carry out extensive numerical studies to examine the …


Finding A Representative Distribution For The Tail Index Alpha, Α, For Stock Return Data From The New York Stock Exchange, Jett Burns May 2022

Finding A Representative Distribution For The Tail Index Alpha, Α, For Stock Return Data From The New York Stock Exchange, Jett Burns

Electronic Theses and Dissertations

Statistical inference is a tool for creating models that can accurately display real-world events. Special importance is given to the financial methods that model risk and large price movements. A parameter that describes tail heaviness, and risk overall, is α. This research finds a representative distribution that models α. The absolute value of standardized stock returns from the Center for Research on Security Prices are used in this research. The inference is performed using R. Approximations for α are found using the ptsuite package. The GAMLSS package employs maximum likelihood estimation to estimate distribution parameters using the CRSP data. The …


Investigaion Of The Gamma Hurdle Model For A Single Population Mean, Alissa Jacobs Jan 2022

Investigaion Of The Gamma Hurdle Model For A Single Population Mean, Alissa Jacobs

Electronic Theses and Dissertations

A common issue in some statistical inference problems is dealing with a high frequency of zeroes in a sample of data. For many distributions such as the gamma, optimal inference procedures do not allow for zeroes to be present. In practice, however, it is natural to observe real data sets where nonnegative distributions would make sense to model but naturally zeroes will occur. One example of this is in the analysis of cost in insurance claim studies. One common approach to deal with the presence of zeroes is using a hurdle model. Most literary work on hurdle models will focus …


Confidence Interval For The Mean Of A Beta Distribution, Sean Rangel Dec 2021

Confidence Interval For The Mean Of A Beta Distribution, Sean Rangel

Electronic Theses and Dissertations

Statistical inference for the mean of a beta distribution has become increasingly popular in various fields of academic research. In this study, we developed a novel statistical model from likelihood-based techniques to evaluate various confidence interval techniques for the mean of a beta distribution. Simulation studies will be implemented to compare the performance of the confidence intervals. In addition to the development and study involving confidence intervals, we will also apply the confidence intervals to real biological data that was gathered by the Department of Biology at Stephen F. Austin State University and provide recommendations on the best practice.


Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin Aug 2021

Bayesian Variable Selection Strategies In Longitudinal Mixture Models And Categorical Regression Problems., Md Nazir Uddin

Electronic Theses and Dissertations

In this work, we seek to develop a variable screening and selection method for Bayesian mixture models with longitudinal data. To develop this method, we consider data from the Health and Retirement Survey (HRS) conducted by University of Michigan. Considering yearly out-of-pocket expenditures as the longitudinal response variable, we consider a Bayesian mixture model with $K$ components. The data consist of a large collection of demographic, financial, and health-related baseline characteristics, and we wish to find a subset of these that impact cluster membership. An initial mixture model without any cluster-level predictors is fit to the data through an MCMC …


Performance Comparison Of Imputation Methods For Mixed Data Missing At Random With Small And Large Sample Data Set With Different Variability, Kyei Afari Aug 2021

Performance Comparison Of Imputation Methods For Mixed Data Missing At Random With Small And Large Sample Data Set With Different Variability, Kyei Afari

Electronic Theses and Dissertations

One of the concerns in the field of statistics is the presence of missing data, which leads to bias in parameter estimation and inaccurate results. However, the multiple imputation procedure is a remedy for handling missing data. This study looked at the best multiple imputation methods used to handle mixed variable datasets with different sample sizes and variability along with different levels of missingness. The study employed the predictive mean matching, classification and regression trees, and the random forest imputation methods. For each dataset, the multiple regression parameter estimates for the complete datasets were compared to the multiple regression parameter …


Statistical Modeling Of Positive Peer Support On Longitudinal Adolescent Substance Use, Kady Rost Jan 2021

Statistical Modeling Of Positive Peer Support On Longitudinal Adolescent Substance Use, Kady Rost

Electronic Theses and Dissertations

To evaluate this study’s research question of ”Does the latent construct of Positive Peer Support (PPS) relate to the construct of Adolescent Substance Use (ASU) over time, controlling for neighborhood safety, race, and sex?”, Structural Equation (SEM) and Latent Growth Curve Modeling (LGCM) were used to investigate trajectories. Secondary longitudinal data from Zimmerman (2014) of 604 students enrolled for four consecutive years in public schools located in Flint, Michigan. In the secondary data resource, students who participated were declared “at risk” by GPA. Significant relationships were found in SEM: Positive Peer Support to Adolescent Substance Use, All Control Variables to …