Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 781 - 810 of 1633

Full-Text Articles in Statistical Theory

Bias In Monte Carlo Simulations Due To Pseudo-Random Number Generator Initial Seed Selection, Jack C. Hill, Shlomo S. Sawilowsky May 2011

Bias In Monte Carlo Simulations Due To Pseudo-Random Number Generator Initial Seed Selection, Jack C. Hill, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

Pseudo-random number generators can bias Monte Carlo simulations of the standard normal probability distribution function with initial seeds selection. Five generator designs were initial-seeded with values from 10000HEX to 1FFFFHEX, estimates of the mean were calculated for each seed, the distribution of mean estimates was determined for each generator and simulation histories were graphed for selected seeds.


New Perspectives In Applying The Regression-Discontinuity Design For Program Evaluation: A Simulation Analysis, Sally A. Lesik May 2011

New Perspectives In Applying The Regression-Discontinuity Design For Program Evaluation: A Simulation Analysis, Sally A. Lesik

Journal of Modern Applied Statistical Methods

Evaluating educational programs is a core component of assessment. One challenge occurs because participants often enter into programs with diverse skills and backgrounds. The regression-discontinuity design has been used to evaluate programs amongst a diverse group, but noncompliance is a limitation. A simulation analysis illustrates the impact of noncompliance.


Information Technology For Increasing Qualitative Information Processing Efficiency, S. N. Martyshenko, E. A. Egorov May 2011

Information Technology For Increasing Qualitative Information Processing Efficiency, S. N. Martyshenko, E. A. Egorov

Journal of Modern Applied Statistical Methods

The problem of qualitative information processing in questionnaires is considered and a solution for this problem is offered. The computer technology developed by the authors to automate the offered decision is described.


General Piecewise Growth Mixture Model: Word Recognition Development For Different Learners In Different Phases, Amery D. Wu, Bruno D. Zumbo, Linda S. Siegel May 2011

General Piecewise Growth Mixture Model: Word Recognition Development For Different Learners In Different Phases, Amery D. Wu, Bruno D. Zumbo, Linda S. Siegel

Journal of Modern Applied Statistical Methods

The General Piecewise Growth Mixture Model (GPGMM), without losing generality to other fields of study, can answer six crucial research questions regarding children’s word recognition development. Using child word recognition data as an example, this study demonstrates the flexibility and versatility of the GPGMM in investigating growth trajectories that are potentially phasic and heterogeneous. The strengths and limitations of the GPGMM and lessons learned from this hands-on experience are discussed.


A Simulation Study Of The Relative Efficiency Of The Minimized Integrated Square Error Estimator (L2e) For Phase I Control Charting, John N. Dyer May 2011

A Simulation Study Of The Relative Efficiency Of The Minimized Integrated Square Error Estimator (L2e) For Phase I Control Charting, John N. Dyer

Journal of Modern Applied Statistical Methods

Parameter estimates used in control charting, the sample mean and variance, are based on maximum likelihood estimation (MLE). Unfortunately, MLEs are not robust to contaminated data and can lead to improper conclusions regarding parameter values. This article proposes a more robust estimation technique; the minimized integrated square error estimator (L2E).


Maximum Likelihood Solution For The Linear Structural Relationship With Three Parameters Known, Androulla Michaeloudis May 2011

Maximum Likelihood Solution For The Linear Structural Relationship With Three Parameters Known, Androulla Michaeloudis

Journal of Modern Applied Statistical Methods

A maximum likelihood solution is obtained for the simple linear structural relation model where the underlying incidental distribution and one error variance are assumed known. Expressions for the asymptotic standard errors of the maximum likelihood estimates are obtained and these are verified using a simulation study.


Logistic Regression Models For Higher Order Transition Probabilities Of Markov Chain For Analyzing The Occurrences Of Daily Rainfall Data, Narayan Chanra Sinha, M. Ataharul Islam, Kazi Saleh Ahamed May 2011

Logistic Regression Models For Higher Order Transition Probabilities Of Markov Chain For Analyzing The Occurrences Of Daily Rainfall Data, Narayan Chanra Sinha, M. Ataharul Islam, Kazi Saleh Ahamed

Journal of Modern Applied Statistical Methods

Logistic regression models for transition probabilities of higher order Markov models are developed for the sequence of chain dependent repeated observations. To identify the significance of these models and their parameters a test procedure for a likelihood ratio criterion is developed. A method of model selection is suggested on the basis of AIC and BIC procedures. The proposed models and test procedures are applied to analyze the occurrences of daily rainfall data for selected stations in Bangladesh. Based on results from these models, the transition probabilities of first order Markov model for temperature and humidity provided the most suitable option …


Number Of Replications Required In Monte Carlo Simulation Studies: A Synthesis Of Four Studies, Daniel J. Mundform, Jay Schaffer, Myoung-Jin Kim, Dale Shaw, Ampai Thongteeraparp, Pornsin Supawan May 2011

Number Of Replications Required In Monte Carlo Simulation Studies: A Synthesis Of Four Studies, Daniel J. Mundform, Jay Schaffer, Myoung-Jin Kim, Dale Shaw, Ampai Thongteeraparp, Pornsin Supawan

Journal of Modern Applied Statistical Methods

Monte Carlo simulations are used extensively to study the performance of statistical tests and control charts. Researchers have used various numbers of replications, but rarely provide justification for their choice. Currently, no empirically-based recommendations regarding the required number of replications exist. Twenty-two studies were re-analyzed to determine empirically-based recommendations.


Matched-Pair Studies With Misclassified Ordinal Data, Tze-San Lee May 2011

Matched-Pair Studies With Misclassified Ordinal Data, Tze-San Lee

Journal of Modern Applied Statistical Methods

The problem of matched-pair studies with misclassified ordinal data is considered. Misclassification is assumed to occur only between the adjacent columns/rows. Bias-adjusted generalized odds ratio and a test for marginal homogeneity are presented to account for misclassification bias. Data from lambing records of 227 Merino ewes are used to illustrate how to calculate these bias-adjusted estimators and – because validation data are not available – a sensitivity analysis is conducted.


A Robust Root Mean Square Standardized Effect Size In One-Way Fixed-Effects Anova, Guili Zhang, James Algina May 2011

A Robust Root Mean Square Standardized Effect Size In One-Way Fixed-Effects Anova, Guili Zhang, James Algina

Journal of Modern Applied Statistical Methods

A robust Root Mean Square Standardized Effect Size (RMSSER) was developed to address the unsatisfactory performance of the Root Mean Square Standardized Effect Size. The coverage performances of the confidence intervals (CI) for RMSSER were investigated. The coverage probabilities of the non-central F distribution-based CI for RMSSER were adequate.


The Overall F-Tests For Seasonal Unit Roots Under Nonstationary Alternatives: Some Theoretical Results And A Monte Carlo Investigation, Ghassen El Montasser May 2011

The Overall F-Tests For Seasonal Unit Roots Under Nonstationary Alternatives: Some Theoretical Results And A Monte Carlo Investigation, Ghassen El Montasser

Journal of Modern Applied Statistical Methods

In many empirical studies concerning seasonal time series, it has been shown that the whole set of unit roots associated with seasonal random walks are not present. This article focuses on the overall F-tests for seasonal unit roots under some nonstationary alternatives different from the seasonal random walk. The asymptotic theory of these tests is established for these cases using a new approach based on circulant matrix concepts. The simulation results joined to this theoretic analysis showed that the overall F-tests, as well as their augmented versions, maintained high power against the nonstationary alternatives.


Weighting Large Datasets With Complex Sampling Designs: Choosing The Appropriate Variance Estimation Method, Sara Mann, James Chowhan May 2011

Weighting Large Datasets With Complex Sampling Designs: Choosing The Appropriate Variance Estimation Method, Sara Mann, James Chowhan

Journal of Modern Applied Statistical Methods

Using the Canadian Workplace and Employee Survey (WES), three variance estimation methods for weighting large datasets with complex sampling designs are compared: simple final weighting, standard bootstrapping and mean bootstrapping. Using a logit analysis, it is shown - depending on which weighting method is used - different predictor variables are significant. The potential lack of independence inherent in a multi-stage cluster sample design, as in the WES, results in a downward bias in the variance when conducting statistical inference (using the simple final weight), which in turn results in increased Type I errors. Bootstrap methods can account for the survey’s …


Using Finite Mixture Modeling To Deal With Systematic Measurement Error: A Case Study, Min Liu, Gregory R. Hancock, Jeffrey R. Harring May 2011

Using Finite Mixture Modeling To Deal With Systematic Measurement Error: A Case Study, Min Liu, Gregory R. Hancock, Jeffrey R. Harring

Journal of Modern Applied Statistical Methods

Conventional methods and analyses view measurement error as random. A scenario is presented where a variable was measured with systematic error. Mixture models with systematic parameter constraints were used to test hypotheses in the context of general linear models; this accommodated the heterogeneity arising due to systematic measurement error.


Estimating Internal Consistency Using Bayesian Methods, Miguel A. Padilla, Guili Zhang May 2011

Estimating Internal Consistency Using Bayesian Methods, Miguel A. Padilla, Guili Zhang

Journal of Modern Applied Statistical Methods

Bayesian internal consistency and its Bayesian credible interval (BCI) are developed and Bayesian internal consistency and its percentile and normal theory based BCIs were investigated in a simulation study. Results indicate that the Bayesian internal consistency is relatively unbiased under all investigated conditions and the percentile based BCIs yielded better coverage performance.


A Broad Symmetry Criterion For Nonparametric Validity Of Parametrically-Based Tests In Randomized Trials, Russell T. Shinohara, Constantine E. Frangakis, Constantine G.. Lyketos Apr 2011

A Broad Symmetry Criterion For Nonparametric Validity Of Parametrically-Based Tests In Randomized Trials, Russell T. Shinohara, Constantine E. Frangakis, Constantine G.. Lyketos

Johns Hopkins University, Dept. of Biostatistics Working Papers

Summary. Pilot phases of a randomized clinical trial often suggest that a parametric model may be an accurate description of the trial's longitudinal trajectories. However, parametric models are often not used for fear that they may invalidate tests of null hypotheses of equality between the experimental groups. Existing work has shown that when, for some types of data, certain parametric models are used, the validity for testing the null is preserved even if the parametric models are incorrect. Here, we provide a broader and easier to check characterization of parametric models that can be used to (a) preserve nonparametric validity …


Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan Apr 2011

Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This article is devoted to the construction and asymptotic study of adaptive group sequential covariate-adjusted randomized clinical trials analyzed through the prism of the semiparametric methodology of targeted maximum likelihood estimation (TMLE). We show how to build, as the data accrue group-sequentially, a sampling design which targets a user-supplied optimal design. We also show how to carry out a sound TMLE statistical inference based on such an adaptive sampling scheme (therefore extending some results known in the i.i.d setting only so far), and how group-sequential testing applies on top of it. The procedure is robust (i.e., consistent even if the …


Targeted Maximum Likelihood Estimation For Dynamic Treatment Regimes In Sequential Randomized Controlled Trials, Paul Chaffee, Mark J. Van Der Laan Mar 2011

Targeted Maximum Likelihood Estimation For Dynamic Treatment Regimes In Sequential Randomized Controlled Trials, Paul Chaffee, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Sequential Randomized Controlled Trials (SRCTs) are rapidly becoming essential tools in the search for optimized treatment regimes in ongoing treatment settings. Analyzing data for multiple time-point treatments with a view toward optimal treatment regimes is of interest in many types of afflictions: HIV infection, Attention Deficit Hyperactivity Disorder in children, leukemia, prostate cancer, renal failure, and many others. Methods for analyzing data from SRCTs exist but they are either inefficient or suffer from the drawbacks of estimating equation methodology. We describe an estimation procedure, targeted maximum likelihood estimation (TMLE), which has been fully developed and implemented in point treatment settings, …


Estimating Subject-Specific Treatment Differences For Risk-Benefit Assessment With Competing Risk Event-Time Data, Brian Claggett, Lihui Zhao, Lu Tian, Davide Castagno, L. J. Wei Mar 2011

Estimating Subject-Specific Treatment Differences For Risk-Benefit Assessment With Competing Risk Event-Time Data, Brian Claggett, Lihui Zhao, Lu Tian, Davide Castagno, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan Mar 2011

Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan

Johns Hopkins University, Dept. of Biostatistics Working Papers

We present a brief overview of targeted maximum likelihood for estimating the causal effect of a single time point treatment and of a two time point treatment. We focus on simple examples demonstrating how to apply the methodology developed in (van der Laan and Rubin, 2006; Moore and van der Laan, 2007; van der Laan, 2010a,b). We include R code for the single time point case.


Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu Jan 2011

Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We establish a fundamental equivalence between singular value decomposition (SVD) and functional principal components analysis (FPCA) models. The constructive relationship allows to deploy the numerical efficiency of SVD to fully estimate the components of FPCA, even for extremely high-dimensional functional objects, such as brain images. As an example, a functional mixed effect model is fitted to high-resolution morphometric (RAVENS) images. The main directions of morphometric variation in brain volumes are identified and discussed.


A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos Jan 2011

A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos

U.C. Berkeley Division of Biostatistics Working Paper Series

In many analyses, one has data on one level but desires to draw inference on another level. For example, in genetic association studies, one observes units of DNA referred to as SNPs, but wants to determine whether genes that are comprised of SNPs are associated with disease. While there are some available approaches for addressing this issue, they usually involve making parametric assumptions and are not easily generalizable. A statistical test is proposed for testing the association of a set of variables with an outcome of interest. No assumptions are made about the functional form relating the variables to the …


Flipping The Winner Of A Poset Game, Adam O. Kalinich '12 Jan 2011

Flipping The Winner Of A Poset Game, Adam O. Kalinich '12

Student Publications & Research

Partially-ordered set games, also called poset games, are a class of two-player combinatorial games. The playing field consists of a set of elements, some of which are greater than other elements. Two players take turns removing an element and all elements greater than it, and whoever takes the last element wins. Examples of poset games include Nim and Chomp. We investigate the complexity of computing which player of a poset game has a winning strategy. We give an inductive procedure that modifies poset games to change the nim-value which informally captures the winning strategies in the game. For a generic …


Parametric Estimation In Competing Risks And Multi-State Models, Yushun Lin Jan 2011

Parametric Estimation In Competing Risks And Multi-State Models, Yushun Lin

Theses and Dissertations--Statistics

The typical research of Alzheimer's disease includes a series of cognitive states. Multi-state models are often used to describe the history of disease evolvement. Competing risks models are a sub-category of multi-state models with one starting state and several absorbing states.

Analyses for competing risks data in medical papers frequently assume independent risks and evaluate covariate effects on these events by modeling distinct proportional hazards regression models for each event. Jeong and Fine (2007) proposed a parametric proportional sub-distribution hazard (SH) model for cumulative incidence functions (CIF) without assumptions about the dependence among the risks. We modified their model to …


Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan Dec 2010

Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan

UW Biostatistics Working Paper Series

In the presence of missing response, reweighting the complete case subsample by the inverse of nonmissing probability is both intuitive and easy to implement. However, inverse probability weighting is not efficient in general and is not robust against misspecification of the missing probability model. Calibration was developed by survey statisticians for improving efficiency of inverse probability weighting estimators when population totals of auxiliary variables are known and when inclusion probability is known by design. In missing data problem we can calibrate auxiliary variables in the complete case subsample to the full sample. However, the inclusion probability is unknown in general …


Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan Dec 2010

Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan

UW Biostatistics Working Paper Series

An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …


Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan Dec 2010

Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan

UW Biostatistics Working Paper Series

An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …


Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel Dec 2010

Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel

COBRA Preprint Series

In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …


Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley Dec 2010

Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley

UW Biostatistics Working Paper Series

Association studies in environmental statistics often involve exposure and outcome data that are misaligned in space. A common strategy is to employ a spatial model such as universal kriging to predict exposures at locations with outcome data and then estimate a regression parameter of interest using the predicted exposures. This results in measurement error because the predicted exposures do not correspond exactly to the true values. We characterize the measurement error by decomposing it into Berkson-like and classical-like components. One correction approach is the parametric bootstrap, which is effective but computationally intensive since it requires solving a nonlinear optimization problem …


Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel Nov 2010

Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel

COBRA Preprint Series

The goal of determining which of hundreds of thousands of SNPs are associated with disease poses one of the most challenging multiple testing problems. Using the empirical Bayes approach, the local false discovery rate (LFDR) estimated using popular semiparametric models has enjoyed success in simultaneous inference. However, the estimated LFDR can be biased because the semiparametric approach tends to overestimate the proportion of the non-associated single nucleotide polymorphisms (SNPs). One of the negative consequences is that, like conventional p-values, such LFDR estimates cannot quantify the amount of information in the data that favors the null hypothesis of no disease-association.

We …


Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano Nov 2010

Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano

Harvard University Biostatistics Working Paper Series

No abstract provided.