Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Bootstrap

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 59

Full-Text Articles in Statistics and Probability

Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh Aug 2024

Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh

Electronic Theses and Dissertations

This study explores innovative approaches to constructing confidence intervals for the population standard deviation, σ, in non-normal data scenarios. While the sample standard deviation, s, is widely used, its reliability is compromised when dealing with skewed or heavy-tailed distributions and exhibits sensitivity to outliers. Our research addresses these limitations by investigating alternative estimation methods that offer greater robustness and accuracy.


Finite Mixtures Of Mean-Parameterized Conway-Maxwell-Poisson Models, Dongying Zhan Jan 2023

Finite Mixtures Of Mean-Parameterized Conway-Maxwell-Poisson Models, Dongying Zhan

Theses and Dissertations--Statistics

For modeling count data, the Conway-Maxwell-Poisson (CMP) distribution is a popular generalization of the Poisson distribution due to its ability to characterize data over- or under-dispersion. While the classic parameterization of the CMP has been well-studied, its main drawback is that it is does not directly model the mean of the counts. This is mitigated by using a mean-parameterized version of the CMP distribution. In this work, we are concerned with the setting where count data may be comprised of subpopulations, each possibly having varying degrees of data dispersion. Thus, we propose a finite mixture of mean-parameterized CMP distributions. An …


A Bootstrap Method For A Multiple-Imputation Variance Estimator In Survey Sampling, Lili Yu, Yichuan Zhao Nov 2022

A Bootstrap Method For A Multiple-Imputation Variance Estimator In Survey Sampling, Lili Yu, Yichuan Zhao

Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications

Rubin’s variance estimator of the multiple imputation estimator for a domain mean is not asymptotically unbiased. Kim et al. derived the closed-form bias for Rubin’s variance estimator. In addition, they proposed an asymptotically unbiased variance estimator for the multiple imputation estimator when the imputed values can be written as a linear function of the observed values. However, this needs the assumption that the covariance of the imputed values in the same imputed dataset is twice that in the different imputed datasets. In this study, we proposed a bootstrap variance estimator that does not need this assumption. Both theoretical argument and …


Non-Inferiority Testing: Kernel Estimation And Overlap Measure, Larie C. Ward Jan 2022

Non-Inferiority Testing: Kernel Estimation And Overlap Measure, Larie C. Ward

College of Graduate Studies: Theses & Dissertations

In non-inferiority testing, the decision of whether a proposed treatment is non-inferior to a reference treatment depends on model assumptions and choices of acceptable tolerance limits. Here, we consider a method that employs kernels to estimate the probability density functions of both the experimental and reference populations from two independent samples. Based on these densities, we introduce a quantity called the overlap coefficient or overlap measure. A bootstrap technique is helpful in exploring the distribution and variance empirically. We derive the distribution of this measure and define a hypothesis test that can be applied to the non-inferiority setting under some …


Bayesian Analysis Of Extended Cox Model With Time-Varying Covariates Using Bootstrap Prior, Oyebayo R. Olaniran, Mohd Asrul A. Abdullah Jul 2020

Bayesian Analysis Of Extended Cox Model With Time-Varying Covariates Using Bootstrap Prior, Oyebayo R. Olaniran, Mohd Asrul A. Abdullah

Journal of Modern Applied Statistical Methods

A new Bayesian estimation procedure for extended cox model with time varying covariate was presented. The prior was determined using bootstrapping technique within the framework of parametric empirical Bayes. The efficiency of the proposed method was observed using Monte Carlo simulation of extended Cox model with time varying covariates under varying scenarios. Validity of the proposed method was also ascertained using real life data set of Stanford heart transplant. Comparison of the proposed method with its competitor established appreciable supremacy of the method.


Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris Jul 2020

Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris

Journal of Modern Applied Statistical Methods

Re-sampling based statistical tests are known to be computationally heavy, but reliable when small sample sizes are available. Despite their nice theoretical properties not much effort has been put to make them efficient. Computationally efficient method for calculating permutation-based p-values for the Pearson correlation coefficient and two independent samples t-test are proposed. The method is general and can be applied to other similar two sample mean or two mean vectors cases.


Quasi-Likelihood Ratio Tests For Homoscedasticity In Linear Regression, Lili Yu, Varadan Sevilimedu, Robert Vogel, Hani Samawi Apr 2020

Quasi-Likelihood Ratio Tests For Homoscedasticity In Linear Regression, Lili Yu, Varadan Sevilimedu, Robert Vogel, Hani Samawi

Journal of Modern Applied Statistical Methods

Two quasi-likelihood ratio tests are proposed for the homoscedasticity assumption in the linear regression models. They require few assumptions than the existing tests. The properties of the tests are investigated through simulation studies. An example is provided to illustrate the usefulness of the new proposed tests.


Using Hac Estimators For Intervention Analysis, Ashok K. Singh, Rohan J. Dalpatadu Jan 2020

Using Hac Estimators For Intervention Analysis, Ashok K. Singh, Rohan J. Dalpatadu

Hospitality Faculty Research

The purpose of this article is to present an alternative method for intervention analysis of time series data that is simpler to use than the traditional method of fitting an explanatory Autoregressive Integrated Moving Average (ARIMA) model. Time series regression analysis is commonly used to test the effect of an event on a time series. An econometric modeling method, which uses a heteroskedasticity and autocorrelation consistent (HAC) estimator of the covariance matrix instead of fitting an ARIMA model, is proposed as an alternative. The method of parametric bootstrap is used to compare the two approaches for intervention analysis. The results …


A Flexible Zero-Inflated Poisson Regression Model, Eric S. Roemmele Jan 2019

A Flexible Zero-Inflated Poisson Regression Model, Eric S. Roemmele

Theses and Dissertations--Statistics

A practical problem often encountered with observed count data is the presence of excess zeros. Zero-inflation in count data can easily be handled by zero-inflated models, which is a two-component mixture of a point mass at zero and a discrete distribution for the count data. In the presence of predictors, zero-inflated Poisson (ZIP) regression models are, perhaps, the most commonly used. However, the fully parametric ZIP regression model could sometimes be restrictive, especially with respect to the mixing proportions. Taking inspiration from some of the recent literature on semiparametric mixtures of regressions models for flexible mixture modeling, we propose a …


Evaluation Of Using The Bootstrap Procedure To Estimate The Population Variance, Nghia Trong Nguyen May 2018

Evaluation Of Using The Bootstrap Procedure To Estimate The Population Variance, Nghia Trong Nguyen

Electronic Theses and Dissertations

The bootstrap procedure is widely used in nonparametric statistics to generate an empirical sampling distribution from a given sample data set for a statistic of interest. Generally, the results are good for location parameters such as population mean, median, and even for estimating a population correlation. However, the results for a population variance, which is a spread parameter, are not as good due to the resampling nature of the bootstrap method. Bootstrap samples are constructed using sampling with replacement; consequently, groups of observations with zero variance manifest in these samples. As a result, a bootstrap variance estimator will carry a …


Multiple Ratio Imputation By The Emb Algorithm: Theory And Simulation, Masayoshi Takahashi May 2017

Multiple Ratio Imputation By The Emb Algorithm: Theory And Simulation, Masayoshi Takahashi

Journal of Modern Applied Statistical Methods

Although multiple imputation is the gold standard of treating missing data, single ratio imputation is often used in practice. Based on Monte Carlo simulation, the Expectation-Maximization with Bootstrapping (EMB) algorithm to create multiple ratio imputation is used to fill in the gap between theory and practice.


Jmasm44: Implementing Multiple Ratio Imputation By The Emb Algorithm (R), Masayoshi Takahashi May 2017

Jmasm44: Implementing Multiple Ratio Imputation By The Emb Algorithm (R), Masayoshi Takahashi

Journal of Modern Applied Statistical Methods

Although single ratio imputation is often used to deal with missing values in practice, there is a paucity of discussion regarding multiple ratio imputation. Code in the R statistical environment is presented to execute multiple ratio imputation by the Expectation-Maximization with Bootstrapping (EMB) algorithm.


Statistical Methodology For Data With Multiple Limits Of Detection, Robert M. Flikkema Jun 2016

Statistical Methodology For Data With Multiple Limits Of Detection, Robert M. Flikkema

Dissertations

Limitations of instruments used to collect continuous data sometimes lead to obtaining observations lower than a limit of detection. These observations are known as nondetects. They could be zeroes, or positive numbers, but they are too small to be recorded by a measuring device. Nondetects frequently occur in environmental data. Trace amounts of chemicals can exist in soil or groundwater and are undetectable by a machine reading. These observations pose a problem to researchers since the true values are unknown.

Simulations in the literature have led to inconsistent conclusions regarding what estimation technique to use with nondetect data when estimating …


Causal Inference In Observational Studies With Clustered Data, Meng Wu Jan 2016

Causal Inference In Observational Studies With Clustered Data, Meng Wu

Legacy Theses & Dissertations (2009 - 2024)

In this thesis, we study causal inference in observational studies with clustered data.


Combating Anti-Statistical Thinking Using Simulation-Based Methods Throughout The Undergraduate Curriculum, Nathan L. Tintle, Beth Chance, George Cobb, Soma Roy, Todd Swanson, Jill Vanderstoep Dec 2015

Combating Anti-Statistical Thinking Using Simulation-Based Methods Throughout The Undergraduate Curriculum, Nathan L. Tintle, Beth Chance, George Cobb, Soma Roy, Todd Swanson, Jill Vanderstoep

Faculty Work Comprehensive List

The use of simulation-based methods for introducing inference is growing in popularity for the Stat 101 course, due in part to increasing evidence of the methods ability to improve students’ statistical thinking. This impact comes from simulation-based methods (a) clearly presenting the overarching logic of inference, (b) strengthening ties between statistics and probability/mathematical concepts, (c) encouraging a focus on the entire research process, (d) facilitating student thinking about advanced statistical concepts, (e) allowing more time to explore, do, and talk about real research and messy data, and (f) acting as a firmer foundation on which to build statistical intuition. Thus, …


Per-Contact Infectivity Of Hcv Associated With Injection Exposures In A Prospective Cohort Of Young Injection Drug Users In San Francisco, Ca (Ufo Study), Yuridia Leyva Sep 2015

Per-Contact Infectivity Of Hcv Associated With Injection Exposures In A Prospective Cohort Of Young Injection Drug Users In San Francisco, Ca (Ufo Study), Yuridia Leyva

Mathematics & Statistics ETDs

Sharing needles and ancillary injection drug equipment places injection drug users (IDU) at risk for Hepatitis C Virus (HCV), a highly infectious blood-borne virus. A limited number of studies have analyzed the per-contact infectivity of HCV associated with the use of previously-used needles, but per-contact infectivity of ancillary injecting equipment has not been previously investigated. Our goal is to estimate the per-contact infectivity of HCV associated with (1) injecting with another person's previously-used needle, classified as receptive needle sharing (RNS), and (2) using another person's previously-used ancillary injecting equipment, such as cookers to melt drugs and cottons to strain impurities …


Testing Equality Of Locally Stationary Covariances With Application To Mortality Rate Modeling, Zahra Teimouri, Ali R. Taheriyoun Jun 2015

Testing Equality Of Locally Stationary Covariances With Application To Mortality Rate Modeling, Zahra Teimouri, Ali R. Taheriyoun

Applications and Applied Mathematics: An International Journal (AAM)

No abstract provided.


Constructing Confidence Intervals For Effect Sizes In Anova Designs, Li-Ting Chen, Chao-Ying Joanne Peng Nov 2013

Constructing Confidence Intervals For Effect Sizes In Anova Designs, Li-Ting Chen, Chao-Ying Joanne Peng

Journal of Modern Applied Statistical Methods

A confidence interval for effect sizes provides a range of plausible population effect sizes (ES) that are consistent with data. This article defines an ES as a standardized linear contrast of means. The noncentral method, Bonett’s method, and the bias-corrected and accelerated bootstrap method are illustrated for constructing the confidence interval for such an effect size. Results obtained from the three methods are discussed and interpretations of results are offered.


Bootstrap Interval Estimation Of Reliability Via Coefficient Omega, Miguel A. Padilla, Jasmin Divers May 2013

Bootstrap Interval Estimation Of Reliability Via Coefficient Omega, Miguel A. Padilla, Jasmin Divers

Journal of Modern Applied Statistical Methods

Three different bootstrap confidence intervals (CIs) for coefficient omega were investigated. The CIs were assessed through a simulation study with conditions not previously investigated. All methods performed well; however, the normal theory bootstrap (NTB) CI had the best performance because it had more consistent acceptable coverage under the simulation conditions investigated.


The Length-Biased Lognormal Distribution And Its Application In The Analysis Of Data From Oil Field Exploration Studies, Makarand V. Ratnaparkhi, Uttara V. Naik-Nimbalkar May 2012

The Length-Biased Lognormal Distribution And Its Application In The Analysis Of Data From Oil Field Exploration Studies, Makarand V. Ratnaparkhi, Uttara V. Naik-Nimbalkar

Journal of Modern Applied Statistical Methods

The length-biased version of the lognormal distribution and related estimation problems are considered and sized-biased data arising in the exploration of oil fields is analyzed. The properties of the estimators are studied using simulations and the use of sample mode as an estimate of the lognormal parameter is discussed.


Empirical Sampling From Permutation Space With Unique Patterns, Justice I. Odiase May 2012

Empirical Sampling From Permutation Space With Unique Patterns, Justice I. Odiase

Journal of Modern Applied Statistical Methods

The exact distribution of a test statistic ultimately guarantees that the probability of a Type I error is exactly α. Several methods for estimating the exact distribution of a test statistic have evolved over the years with inherent computational problems and varying degrees of accuracy. The unique pattern of permutations resulting from using experimental data to sample within the permutation space without the risk of repeating permutations is identified. The method presented circumvents the theoretical requirements of asymptotic procedures and the computational difficulties associated with an exhaustive enumeration of permutations. Results show that time and space complexities are drastically reduced …


A Systematic Selection Method For The Development Of Cancer Staging Systems, Yunzhi Lin, Richard Chappell, Mithat Gonen Jan 2012

A Systematic Selection Method For The Development Of Cancer Staging Systems, Yunzhi Lin, Richard Chappell, Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

The tumor-node-metastasis (TNM) staging system has been the anchor of cancer diagnosis, treatment, and prognosis for many years. For meaningful clinical use, an orderly, progressive condensation of the T and N categories into an overall staging system needs to be defined, usually with respect to a time-to-event outcome. This can be considered as a cutpoint selection problem for a censored response partitioned with respect to two ordered categorical covariates and their interaction. The aim is to select the best grouping of the TN categories. A novel bootstrap cutpoint/model selection method is proposed for this task by maximizing bootstrap estimates of …


A Stochastic Version Of The Em Algorithm To Analyze Multivariate Skew-Normal Data With Missing Responses, M. Khounsiavash, M. Ganjali, T. Baghfalaki Dec 2011

A Stochastic Version Of The Em Algorithm To Analyze Multivariate Skew-Normal Data With Missing Responses, M. Khounsiavash, M. Ganjali, T. Baghfalaki

Applications and Applied Mathematics: An International Journal (AAM)

In this paper an algorithm called SEM, which is a stochastic version of the EM algorithm, is used to analyze multivariate skew-normal data with intermittent missing values. Also, a multivariate selection model framework for modeling of both missing and response mechanisms is formulated. By the SEM algorithm missing values of responses are inputed by the conditional distribution of missing values given observed data and then the log-likelihood of the pseudocomplete data is maximized. The algorithm is iterated until convergence of parameter estimates. Results of an application are also reported where a Bootstrap approach is used to compute the standard error …


Modeling Repairable System Failures With Interval Failure Data And Time Dependent Covariate, Jayanthi Arasan, Samira Ehsani Nov 2011

Modeling Repairable System Failures With Interval Failure Data And Time Dependent Covariate, Jayanthi Arasan, Samira Ehsani

Journal of Modern Applied Statistical Methods

An application of a repairable system model for interval failure data with a time dependent covariate is examined. The performance of several models based on the NHPP when applied to real data on ball bearing failures is also explored. The best model for the data was selected based on results of the likelihood ratio test. The bootstrapping technique was applied to obtain the variance estimate for the estimated expected number of failures. Results demonstrate that the proposed model works well and is easy to implement, in addition the bootstrap variance estimate provides a simple substitute for the traditional estimate.


Weighting Large Datasets With Complex Sampling Designs: Choosing The Appropriate Variance Estimation Method, Sara Mann, James Chowhan May 2011

Weighting Large Datasets With Complex Sampling Designs: Choosing The Appropriate Variance Estimation Method, Sara Mann, James Chowhan

Journal of Modern Applied Statistical Methods

Using the Canadian Workplace and Employee Survey (WES), three variance estimation methods for weighting large datasets with complex sampling designs are compared: simple final weighting, standard bootstrapping and mean bootstrapping. Using a logit analysis, it is shown - depending on which weighting method is used - different predictor variables are significant. The potential lack of independence inherent in a multi-stage cluster sample design, as in the WES, results in a downward bias in the variance when conducting statistical inference (using the simple final weight), which in turn results in increased Type I errors. Bootstrap methods can account for the survey’s …


Generalized Variances Ratio Test For Comparing K Covariance Matrices From Dependent Normal Populations, Marcelo Angelo Cirillo, Daniel Furtado Ferreira, Thelma Sáfadi, Eric Batista Ferreira Nov 2010

Generalized Variances Ratio Test For Comparing K Covariance Matrices From Dependent Normal Populations, Marcelo Angelo Cirillo, Daniel Furtado Ferreira, Thelma Sáfadi, Eric Batista Ferreira

Journal of Modern Applied Statistical Methods

New tests based on the ratio of generalized variances are presented to compare covariance matrices from dependent normal populations. Monte Carlo simulation concluded that the tests considered controlled the Type I error, providing empirical probabilities that were consistent with the nominal level stipulated.


Another Look At Resampling: Replenishing Small Samples With Virtual Data Through S-Smart, Haiyan Bai, Wei Pan, Leigh Lihshing Wang, Phillip Neal Ritchey May 2010

Another Look At Resampling: Replenishing Small Samples With Virtual Data Through S-Smart, Haiyan Bai, Wei Pan, Leigh Lihshing Wang, Phillip Neal Ritchey

Journal of Modern Applied Statistical Methods

A new resampling method is introduced to generate virtual data through a smoothing technique for replenishing small samples. The replenished analyzable sample retains the statistical properties of the original small sample, has small standard errors and possesses adequate statistical power.


Level Robust Methods Based On The Least Squares Regression Estimator, Marie Ng, Rand R. Wilcox Nov 2009

Level Robust Methods Based On The Least Squares Regression Estimator, Marie Ng, Rand R. Wilcox

Journal of Modern Applied Statistical Methods

Heteroscedastic consistent covariance matrix (HCCM) estimators provide ways for testing hypotheses about regression coefficients under heteroscedasticity. Recent studies have found that methods combining the HCCM-based test statistic with the wild bootstrap consistently perform better than non-bootstrap HCCM-based methods (Davidson & Flachaire, 2008; Flachaire, 2005; Godfrey, 2006). This finding is more closely examined by considering a broader range of situations which were not included in any of the previous studies. In addition, the latest version of HCCM, HC5 (Cribari-Neto, et al., 2007), is evaluated.


Comparing Bootstrap And Jackknife Variance Estimation Methods For Area Under The Roc Curve Using One-Stage Cluster Survey Data, Allison Dunning Jun 2009

Comparing Bootstrap And Jackknife Variance Estimation Methods For Area Under The Roc Curve Using One-Stage Cluster Survey Data, Allison Dunning

Theses and Dissertations

The purpose of this research is to examine the bootstrap and jackknife as methods for estimating the variance of the AUC from a study using a complex sampling design and to determine which characteristics of the sampling design effects this estimation. Data from a one-stage cluster sampling design of 10 clusters was examined. Factors included three true AUCs (.60, .75, and .90), three prevalence levels (50/50, 70/30, 90/10) (non-disease/disease), and finally three number of clusters sampled (2, 5, or 7). A simulated sample was constructed for each of the 27 combinations of AUC, prevalence and number of clusters. Estimates of …


The Bootstrap Method For The Selection Of A Shrinkage Factor In Two-Stage Estimation Of The Reliability Function Of An Exponential Distribution, Makarand V. Ratnaparkhi, Vasant B. Waikar, Fredrick J. Schuurmann May 2009

The Bootstrap Method For The Selection Of A Shrinkage Factor In Two-Stage Estimation Of The Reliability Function Of An Exponential Distribution, Makarand V. Ratnaparkhi, Vasant B. Waikar, Fredrick J. Schuurmann

Journal of Modern Applied Statistical Methods

An application of a bootstrap method for selecting a suitable shrinkage factor for the two-stage shrinkage estimator of a reliability function for the exponential distribution is discussed. The estimator obtained here has higher efficiency as compared to the one where the shrinkage factor is not subjected to bootstrapping.