Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 31 - 60 of 1633

Full-Text Articles in Statistical Theory

Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad Jan 2025

Theoretical Foundations And Applied Performance Of Periodicity-Aware Imputation: Variable Bandpass Block Bootstrap Methods For Incomplete Time Series, Asmaa Ahmad

Electronic Theses & Dissertations (2024 - present)

Time series data are prevalent across a wide range of disciplines, including health surveillance, public policy, and environmental monitoring. In the presence of underlying cyclical patterns, the integrity of time series analysis depends critically on the ability to detect, model, and impute structured missing data without compromising the temporal structure. This dissertation introduces and validates a novel imputation framework that integrates the Variable Bandpass Periodic Block Bootstrap (VBPBB) into multiple imputation procedures, improving the accuracy, robustness, and interpretability of time series models under high rates of missingness and noise. The overarching goal of this dissertation was to develop and evaluate …


Predictive Modeling For Healthcare Data Using Nonlinear Bayesian Methods, Prince Kofi Asare Jan 2025

Predictive Modeling For Healthcare Data Using Nonlinear Bayesian Methods, Prince Kofi Asare

Theses and Dissertations

Unplanned hospital readmissions represent a significant challenge for healthcare systems, contributing to substantial financial burdens and highlighting gaps in patient care coordination. In the U.S., approximately 20% of Medicare beneficiaries are readmitted within 30 days, costing billions annually. Social determinants of health, such as income, housing stability, and social support, account for up to 80% of health outcomes, yet their integration into predictive models remains underexplored. This study introduces a novel Bayesian framework for predicting 30-day readmission risk, combining Gaussian Process models with spike-and-slab priors and Bayesian Lasso regression with Laplace priors. Utilizing Markov Chain Monte Carlo methods, the approach …


Learning Problems Related To Stochastic Differential Equations, Jinpu Zhou Nov 2024

Learning Problems Related To Stochastic Differential Equations, Jinpu Zhou

LSU Doctoral Dissertations

Stochastic differential equations (SDEs) are essential for modeling systems influenced by both deterministic dynamics and random fluctuations, with applications in a wide variety of disciplines. This thesis develops a Bayesian framework for nonparametric learning in SDEs, addressing key challenges in inference, particularly when dealing with complex systems and incomplete data. The thesis begins by establishing a theoretical foundation in optimization over Hilbert spaces, including a generalized representer theorem to address infinite-dimensional optimization problems encountered in nonparametric inference. Building on this, we introduce a Bayesian framework with shrinkage priors to learn drift functions from high-frequency data. Bayesian approach incorporates low-cost sparse …


A Uniformly Most Powerful Test For The Mean Of A Beta Distribution, Richard Ntiamoah Kyei Aug 2024

A Uniformly Most Powerful Test For The Mean Of A Beta Distribution, Richard Ntiamoah Kyei

Electronic Theses and Dissertations

The beta distribution is used in numerous real-world applications, including areas such as manufacturing (quality control) and analyzing patient outcomes in health care. It also plays a key role in statistical theory, including multivariate analysis of variance (MANOVA) and Bayesian statistics. It is a flexible distribution that can account for many different characteristics of real data. To our surprise, there has been very little work or discussion on performing statistical hypothesis testing for the mean when it is reasonable to assume that the population is beta distributed. Many analysts conduct traditional analyses using a t-test or nonparametric approach, try transformations, …


Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh Aug 2024

Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh

Electronic Theses and Dissertations

This study explores innovative approaches to constructing confidence intervals for the population standard deviation, σ, in non-normal data scenarios. While the sample standard deviation, s, is widely used, its reliability is compromised when dealing with skewed or heavy-tailed distributions and exhibits sensitivity to outliers. Our research addresses these limitations by investigating alternative estimation methods that offer greater robustness and accuracy.


Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han Aug 2024

Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han

Electronic Theses and Dissertations

This dissertation consists of two projects. The first one involves nonparametric methods on Continuous Time Markov Chains (CTMCs). The second one is centered around Bayesian shrinkage models for detecting prognostic and predictive biomarkers in high-dimensional clinical data. Both these projects build on methods from across the frequentist and Bayesian paradigm to offer novel solutions. In the first project, we aim to model the nonlinear effects of continuous variables within multistate framework in a non-parametrically by appealing to the rich mathematical framework of Reproducing Kernel Hilbert Spaces (RKHS). Then we adapted the classical Representer Theorem to penalized (squared norm) log-likelihood which …


Contributions To Nonparametric Testing In Clustered Data, Hasika Kalani Wickrama Senevirathne Jul 2024

Contributions To Nonparametric Testing In Clustered Data, Hasika Kalani Wickrama Senevirathne

Mathematics & Statistics Theses & Dissertations

Clustered data refers to a specific kind of correlated data where units within the same cluster are correlated while units from different clusters are independent. The number of units in each cluster, known as the cluster size, can be associated with the cluster’s outcome. This is known as the informative cluster size (ICS) and affects the inference drawn from clustered data. Recently, a hypothesis testing method has been developed to detect the presence of ICS. However, considering ICS alone may not be sufficient when comparing outcomes across multiple groups of units within clustered data. The size of a group within …


Comparison Of Value At Risk Using Historical And Monte Carlo Methods On Pt Xyz Stock Portofolio, Eka Fitriani, Yulial Hikmah, Ira Rosianal Hikmah Jun 2024

Comparison Of Value At Risk Using Historical And Monte Carlo Methods On Pt Xyz Stock Portofolio, Eka Fitriani, Yulial Hikmah, Ira Rosianal Hikmah

Jurnal Administrasi Bisnis Terapan

One way to achieve profits in a company is through investment activities. However, everything has risks. Investing can also be risky. Therefore, the relationship between risk and investment is important because it will influence the determination of investment selection. The problem faced by investors is choosing an efficient portfolio, or a portfolio that provides the smallest risk. This risk can be done by measuring risk, one of which is using the Value at Risk (VaR) measure. Measurement using Value at Risk has several methods that are quite popular, namely the Historical Method, Variance-Covariance, and Monte Carlo. In this research, the …


Comparing The Preparation Of Youth Services Librarians To Their On-The-Ground Experiences: A Grounded Theory Study Incorporating Criticism And Connoisseurship, Anne Holland Jun 2024

Comparing The Preparation Of Youth Services Librarians To Their On-The-Ground Experiences: A Grounded Theory Study Incorporating Criticism And Connoisseurship, Anne Holland

Electronic Theses and Dissertations

The purpose of this study was to better understand the on-the-ground preparation of youth services librarians, in contrast to their professional training in Master’s of Library Science (MLIS) programs. Classic Grounded Theory was the predominant methodology for this qualitative study, and elements of Criticism and Connoisseurship were also utilized. Document review, interviews, and journaling activities with ten participants were the primary methods of data collection. Key findings from this dissertation include a grounded theory explaining the current state of preparation for youth services librarianship, and multiple avenues for further study.


Trade Liberalization, Non-Oil Export And Economic Growth In Nigeria, Jerome T. Andohol, Terhemen Tarzoor, Dennis T. Nomor Jun 2024

Trade Liberalization, Non-Oil Export And Economic Growth In Nigeria, Jerome T. Andohol, Terhemen Tarzoor, Dennis T. Nomor

CBN Journal of Applied Statistics (JAS)

The study examines the impact of trade liberalization and non-oil exports on economic growth in Nigeria from 1986 to 2021. The study utilizes an autoregressive distributed lag model and found the combined effect of trade liberalization and non-oil exports to be positive and statistical significant. While trade liberalization alone may have negative consequences, its synergy with a robust non-oil export can drive sustainable economic growth. The study recommends that strategies to enhance non-oil exports should be encouraged to support the effectiveness of trade liberalization in promoting growth.


Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov Jun 2024

Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov

CBN Journal of Applied Statistics (JAS)

This paper investigates the time it would take for the FTSE-100 index to reach its post-COVID-19 peak. The paper utilises an exponential generalised autoregressive conditional heteroscedasticity (EGARCH) model that accounts for leverage effect and asymmetries. The preferred models amongst competing variants was the Autoregressive Moving Average (ARMA)-EGARCH(2,1) specification and was used to predict daily FTSE-100 data from 5th January 2000 to 21st June 2024. The empirical exercise showed that the COVID-19-induced financial crisis negatively affected the United Kingdom’s stock market performance. The results show that the FTSE100 index could reach its post-pandemic peak around 27th August, 2024 (two months after …


"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson Apr 2024

"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson

Senior Honors Theses

The authorship of Hebrews has been a point of contention for scholars for the past two millennia. While the epistle is traditionally attributed to Paul, many scholars assert that it carries thematic, structural, and stylistic differences from the remainder of his extant epistles; therefore, many other possible authors have been proposed. Of these, only Luke has other New Testament writings. Therefore, this project conducts a statistical comparison of Hebrews to the Pauline and Lukan corpora using stylometric authorial analysis methods. This analysis demonstrates that Hebrews is stylistically closer to Lukan literature than Pauline (but not to a significant degree), and …


Uconn Baseball Reliever Lane Optimization Tool, Jason Bartholomew Apr 2024

Uconn Baseball Reliever Lane Optimization Tool, Jason Bartholomew

Honors Scholar Theses

The building of a tool to be utilized by UConn’s Division I baseball team that will generate a game plan for when different relievers should be used against different parts of the opponent’s lineup to achieve the lowest total expected value of runs allowed for the remainder of the game based on game situations and matchup probabilities. The tool will also examine and determine situations that may be vital enough to the outcome of the game to bring in a better reliever normally saved for later in the game.


Latent Classification Of Multi-Model Non-Homogeneous Continuous-Time Markov Chains, Joonha Chang Apr 2024

Latent Classification Of Multi-Model Non-Homogeneous Continuous-Time Markov Chains, Joonha Chang

Dissertations and Theses (Open Access)

The continuous-time Markov chain (CTMC) model and latent clustering models are commonly used to study longitudinal measures of categorical outcomes. Because of its simple but powerful Markovian property, CTMC models have been widely used in medical and public health researches. Due to limitations in the standard CTMC model, there have been some studies on non-homogeneous continuous-time Markov chain (NH-CTMC) that utilized time-dependent rates, but the progresses have been limited. NH-CTMC can be more powerful than CTMC by its default nature of time-dependent rate that can be fitted to a wider range of applications in medical studies. In this study, we …


Model Selection Through Cross-Validation For Supervised Learning Tasks With Manifold Data, Derek Brown Jan 2024

Model Selection Through Cross-Validation For Supervised Learning Tasks With Manifold Data, Derek Brown

The Journal of Purdue Undergraduate Research

No abstract provided.


Sensitivity Analysis Of Prior Distributions In Regression Model Estimation, Ayoade I Adewole, Oluwatoyin K. Bodunwa Jan 2024

Sensitivity Analysis Of Prior Distributions In Regression Model Estimation, Ayoade I Adewole, Oluwatoyin K. Bodunwa

Al-Bahir

Bayesian inferences depend solely on specification and accuracy of likelihoods and prior distributions of the observed data. The research delved into Bayesian estimation method of regression models to reduce the impact of some of the problems, posed by convectional method of estimating regression models, such as handling complex models, availability of small sample sizes and inclusion of background information in the estimation procedure. Posterior distributions are based on prior distributions and the data accuracy, which is the fundamental principles of Bayesian statistics to produce accurate final model estimates. Sensitivity analysis is an essential part of mathematical model validation in obtaining …


Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe Jan 2024

Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe

Data Science and Data Mining

This project estimates a regression model to predict the superconducting critical temperature based on variables extracted from the superconductor’s chemical formula. The regression model along with the stepwise variable selection gives a reasonable and good predictive model with a lower prediction error (MSE). Variables extracted based on atomic radius, valence, atomic mass and thermal conductivity appeared to have the most contribution to the predictive model.


Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe Jan 2024

Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe

Data Science and Data Mining

Cyberbullying refers to the act of bullying using electronic means and the internet. In recent years, this act has been identifed to be a major problem among young people and even adults. It can negatively impact one’s emotions and lead to adverse outcomes like depression, anxiety, harassment, and suicide, among others. This has led to the need to employ machine learning techniques to automatically detect cyberbullying and prevent them on various social media platforms. In this study, we want to analyze the combination of some Natural Language Processing (NLP) algorithms (such as Bag-of-Words and TFIDF) with some popular machine learning …


Accounting For Variability Due To Resampling Using Bootstrapping, Dipendra Phuyal Jan 2024

Accounting For Variability Due To Resampling Using Bootstrapping, Dipendra Phuyal

College of Graduate Studies: Theses & Dissertations

Bradley Efron (1979) introduced bootrapping. Typically a researcher is interested in studying a process which generates individuals. The collection of individuals the process has(actual) or could have (conceptual) generated is the population. The collection of conceptual members of the population is an uncountable collection. Hence, the population is anuncountable collection of individuals. The collection of individuals the process has generated (actual individuals) is representative of what the process can generate and will bereferred to as the representative sample. The size of this sample is a nonnegative integervalued random variable N which may be a constant random variable such as in …


The Distribution Of The Significance Level, Paul O. Monnu Jan 2024

The Distribution Of The Significance Level, Paul O. Monnu

College of Graduate Studies: Theses & Dissertations

Reporting the p-value is customary when conducting a test of hypothesis or significance. The likelihood of getting a fictitious second sample and presuming the null hypothesis is correct is the p-value. The significance level is a statistic that interests us to investigate. Being a statistic, it has a distribution. For the F-test in a one-way ANOVA and the t-tests for population means, we define the significance level, its observed value, and the observed significance level. It is possible to derive the significance level distribution. The t-test and the F-test are not without controversy. Specifically, we demonstrate that as sample size …


Microplate-Like Metal Pyrophosphate Engineered On Ni-Foam Towards Multifunctional Electrode Material For Energy Conversion And Storage, Rishabh Srivastava Dec 2023

Microplate-Like Metal Pyrophosphate Engineered On Ni-Foam Towards Multifunctional Electrode Material For Energy Conversion And Storage, Rishabh Srivastava

Electronic Theses & Dissertations

High clean energy demand, dire need for sustainable development, and low carbon footprints are the few intuitive challenges, leading researchers to aim for research and development for high-performance energy devices. The development of materials used in energy devices is currently focused on enhancing the performance, electronic properties, and durability of devices. Tunning the attributes of transition metals using pyrophosphate (P2O7) ligand moieties can be a promising approach to meet the requirements of energy devices such as water electrolyzers and supercapacitors, although such a material’s configuration is rarely exposed for this purpose of study.

Herein, we grow …


Exploration And Statistical Modeling Of Profit, Caleb Gibson Dec 2023

Exploration And Statistical Modeling Of Profit, Caleb Gibson

Undergraduate Honors Theses

For any company involved in sales, maximization of profit is the driving force that guides all decision-making. Many factors can influence how profitable a company can be, including external factors like changes in inflation or consumer demand or internal factors like pricing and product cost. Understanding specific trends in one's own internal data, a company can readily identify problem areas or potential growth opportunities to help increase profitability.

In this discussion, we use an extensive data set to examine how a company might analyze their own data to identify potential changes the company might investigate to drive better performance. Based …


Generalized Ratio-Product Cum Regression Variance Estimator In Two-Phase Sampling, Isah Muhammad Dec 2023

Generalized Ratio-Product Cum Regression Variance Estimator In Two-Phase Sampling, Isah Muhammad

CBN Journal of Applied Statistics (JAS)

This study develops a flexible and efficient generalized ratio-product cum regression type estimator of population variance utilizing auxiliary variable in two-phase sampling that incorporates the properties of ratio-type and product-type estimators. The properties of the estimator were derived using first order approximation. The theoretical conditions under which the precision and the flexibility of the estimator is better than some classical estimators are also provided. Empirical evidence from five real datasets suggests that the proposed estimator outperforms the classical variance, ratio variance, product, and exponential ratio type estimators in terms of precision and efficiency. The estimator can be utilized to provide …


Modelling The Naira Exchange Rate Dependence Using Static And Time-Varying Copula, Kabir Katata Dec 2023

Modelling The Naira Exchange Rate Dependence Using Static And Time-Varying Copula, Kabir Katata

CBN Journal of Applied Statistics (JAS)

This paper examines the dependence structure of different currencies versus the Nigerian Naira using constant and time-varying copula. Daily Naira/USD, Naira/Yuan, Naira/Pound, and Naira/Euro exchange rates from 23 December 2011 to 12 May 2020 were utilised. We fitted eight constant and time-varying copula families using the exchange rate standardised residuals. The study finds that the Naira exchange rate may be estimated with student t-copula, Symmetrized Joe-Clayton (SJC), or Rotated Gumbel copula models and Autoregressive (AR)– Glosten Jagannathan RunkleGeneralized Autoregressive Conditional Heteroscedastic (GJR-GARCH) (1,1) models with skewed t residuals for margins. The Naira exchange rate returns is timevarying, tail-dependent, and asymmetric. …


The Private Pilot Check Ride: Applying The Spacing Effect Theory To Predict Time To Proficiency For The Practical Test, Michael Scott Harwin Dec 2023

The Private Pilot Check Ride: Applying The Spacing Effect Theory To Predict Time To Proficiency For The Practical Test, Michael Scott Harwin

Theses and Dissertations

This study examined the relationship between a set of targeted factors and the total flight time students needed to become ready to take the private pilot check ride. The study was grounded in Ebbinghaus’s (1885/1913/2013) forgetting curve theory and spacing effect, and Ausubel’s (1963) theory of meaningful learning. The research factors included (a) training time to proficiency, which represented the number of training days needed to become check-ride ready; (b) flight training program (Part 61 vs. Part 141); (c) organization offering the training program (2- or 4-year college/university vs. FBO); (d) scheduling policy (mandated vs. student-driven); and demographical variables, which …


Statistical Inference On Lung Cancer Screening Using The National Lung Screening Trial Data., Farhin Rahman Aug 2023

Statistical Inference On Lung Cancer Screening Using The National Lung Screening Trial Data., Farhin Rahman

Electronic Theses and Dissertations

This dissertation consists of three research projects on cancer screening probability modeling. In these projects, the three key modeling parameters (sensitivity, sojourn time, transition density) for cancer screening were estimated, along with the long-term outcomes (including overdiagnosis as one outcome), the optimal screening time/age, the lead time distribution, and the probability of overdiagnosis at the future screening time were simulated to provide a statistical perspective on the effectiveness of cancer screening programs. In the first part of this dissertation, a statistical inference was conducted for male and female smokers using the National Lung Screening Trial (NLST) chest X-ray data. A …


A Comparison Of Confidence Intervals In State Space Models, Jinyu Du Jul 2023

A Comparison Of Confidence Intervals In State Space Models, Jinyu Du

Statistical Science Theses and Dissertations

This thesis develops general procedures for constructing confidence intervals (CIs) of the error disturbance parameters (standard deviations) and transformations of the error disturbance parameters in time-invariant state space models (ssm). With only a set of observations, estimating individual error disturbance parameters accurately in the presence of other unknown parameters in ssm is a very challenging problem. We attempted to construct four different types of confidence intervals, Wald, likelihood ratio, score, and higher-order asymptotic intervals for both the simple local level model and the general time-invariant state space models (ssm). We show that for a simple local level model, both the …


Improving The Efficiency Of Exponential Ratio-Type Estimator For Population Median: A Calibration Weight Adjustment Approach, Mathew J. Iseh, Kufre J. Bassey Jun 2023

Improving The Efficiency Of Exponential Ratio-Type Estimator For Population Median: A Calibration Weight Adjustment Approach, Mathew J. Iseh, Kufre J. Bassey

CBN Journal of Applied Statistics (JAS)

This paper modifies the Bahl and Tuteja exponential ratio-type estimator for population median under simple random and stratified sampling schemes using calibration weight adjustment technique with supplementary information to vary the stratum weights. The bias and mean square error of the modified estimator were obtained up to the second-order approximation, which satisfies the necessary conditions for efficiency. The findings show that the new estimator surpasses existing estimators in efficiency gain. This suggests the appropriateness of calibration weight modification in boosting the efficiency of a population parameter estimator under stratified random sampling especially where the population parameter of the auxiliary variable …


Testing For Dice Control Based On Observations Of The Length Of The Shooter's Hand, Stewart N. Ethier, Hokwon Cho May 2023

Testing For Dice Control Based On Observations Of The Length Of The Shooter's Hand, Stewart N. Ethier, Hokwon Cho

International Conference on Gambling & Risk Taking

uploaded


Uconn Baseball Batting Order Optimization, Gavin Rublewski, Gavin Rublewski May 2023

Uconn Baseball Batting Order Optimization, Gavin Rublewski, Gavin Rublewski

Honors Scholar Theses

Challenging conventional wisdom is at the very core of baseball analytics. Using data and statistical analysis, the sets of rules by which coaches make decisions can be justified, or possibly refuted. One of those sets of rules relates to the construction of a batting order. Through data collection, data adjustment, the construction of a baseball simulator, and the use of a Monte Carlo Simulation, I have assessed thousands of possible batting orders to determine the roster-specific strategies that lead to optimal run production for the 2023 UConn baseball team. This paper details a repeatable process in which basic player statistics …