Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1531 - 1560 of 1633

Full-Text Articles in Statistical Theory

A Parametric Bootstrap Version Of Hedges’ Homogeneity Test, Wim Van Den Noortgate, Patrick Onghena May 2003

A Parametric Bootstrap Version Of Hedges’ Homogeneity Test, Wim Van Den Noortgate, Patrick Onghena

Journal of Modern Applied Statistical Methods

Hedges’ Q-test is frequently used in meta-analyses to evaluate the homogeneity of effect sizes, but for several kinds of effect size measures it does not always appropriately control the Type 1 error probability. Therefore we propose a parametric bootstrap version, which shows Type 1 error control under a broad set of circumstances. This is confirmed in a small simulation study.


Fast Permutation Tests That Maximize Power Under Conventional Monte Carlo Sampling For Pairwise And Multiple Comparisons, J. D. Opdyke May 2003

Fast Permutation Tests That Maximize Power Under Conventional Monte Carlo Sampling For Pairwise And Multiple Comparisons, J. D. Opdyke

Journal of Modern Applied Statistical Methods

While the distribution-free nature of permutation tests makes them the most appropriate method for hypothesis testing under a wide range of conditions, their computational demands can be runtime prohibitive, especially if samples are not very small and/or many tests must be conducted (e.g. all pairwise comparisons). This paper presents statistical code that performs continuous-data permutation tests under such conditions very quickly – often more than an order of magnitude faster than widely available commercial alternatives when many tests must be performed and some of the sample pairs contain a large sample. Also presented is an efficient method for obtaining a …


Screening Properties And Design Selection Of Certain Two-Level Designs, H. Evangelaras, Christos Koukouvinos May 2003

Screening Properties And Design Selection Of Certain Two-Level Designs, H. Evangelaras, Christos Koukouvinos

Journal of Modern Applied Statistical Methods

Screening designs are useful for situations where a large number of factors (q) is examined but only few (k) of these are expected to be important. It is of practical interest for a given k to know all the inequivalent projections of the design into the k dimensions. In this paper we give all the inequivalent projections of inequivalent Hadamard matrices of order 28 into k=3 and 4 dimensions and furthermore, we give partial results for k=5. Then, we sort these projections according to their generalized resolution and their generalized aberration.


Analyzing Group By Time Effects In Longitudinal Two-Group Randomized Trial Designs With Missing Data, James Algina, H. J. Keselman, Abdul R. Othman May 2003

Analyzing Group By Time Effects In Longitudinal Two-Group Randomized Trial Designs With Missing Data, James Algina, H. J. Keselman, Abdul R. Othman

Journal of Modern Applied Statistical Methods

We investigated bias, sampling variability, Type I error and power of nine approaches for testing the group by time interaction in a repeated measures design under three types of missing data mechanisms. One procedure due to Overall, Ahn, Shivakumar, and Kalburgi (1999) performed reasonably well over a range of conditions.


A More Efficient Way Of Obtaining A Unique Median Estimate For Circular Data, B. Sango Otieno, Christine M. Anderson-Cook May 2003

A More Efficient Way Of Obtaining A Unique Median Estimate For Circular Data, B. Sango Otieno, Christine M. Anderson-Cook

Journal of Modern Applied Statistical Methods

The procedure for computing the sample circular median occasionally leads to a non-unique estimate of the population circular median, since there can sometimes be two or more diameters that divide data equally and have the same circular mean deviation. A modification in the computation of the sample median is suggested, which not only eliminates this non-uniqueness problem, but is computationally easier and faster to work with than the existing alternative.


You Think You’Ve Got Trivials?, Shlomo S. Sawilowsky May 2003

You Think You’Ve Got Trivials?, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

Effect sizes are important for power analysis and meta-analysis. This has led to a debate on reporting effect sizes for studies that are not statistically significant. Contrary and supportive evidence has been offered on the basis of Monte Carlo methods. In this article, clarifications are given regarding what should be simulated to determine the possible effects of piecemeal publishing trivial effect sizes.


A Semiparametric Regression Model For Oligonucleotide Arrays, Jianhua Hu, Guosheng Yin May 2003

A Semiparametric Regression Model For Oligonucleotide Arrays, Jianhua Hu, Guosheng Yin

Journal of Modern Applied Statistical Methods

A semiparametric model incorporating the spline smoothing technique is proposed to study oligonucleotide gene expression data. No specific parametric functional form is assumed for mismatch probe intensities, which allows much more flexibility in the fitted model. The new approach improves the model fitting, hence the estimation of expression indexes. The method is applied to a data set of 18 HuGeneFL arrays.


Jmasm6: An Algorithm For Generating Exact Critical Values For The Kruskal-Wallis One-Way Anova, Todd C. Headrick May 2003

Jmasm6: An Algorithm For Generating Exact Critical Values For The Kruskal-Wallis One-Way Anova, Todd C. Headrick

Journal of Modern Applied Statistical Methods

A Fortran 77 subroutine is provided for computing exact critical values for the Kruskal-Wallis test on k independent groups with equal or unequal samples sizes. The subroutine requires the user to provide sorting and ranking routines and a uniform pseudo-random number generator. The program is available from the author on request.


Randomization Technique, Allocation Concealment, Masking, And Susceptibility Of Trials To Selection Bias, Vance W. Berger, Costas A. Christophi May 2003

Randomization Technique, Allocation Concealment, Masking, And Susceptibility Of Trials To Selection Bias, Vance W. Berger, Costas A. Christophi

Journal of Modern Applied Statistical Methods

It is widely believed that baseline imbalances in randomized clinical trials must necessarily be random. Yet even among masked randomized trials conducted with allocation concealment, there are mechanisms by which patients with specific covariates may be selected for inclusion into a particular treatment group. This selection bias would force imbalance in those covariates, measured or unmeasured, that are used for the patient selection. Unfortunately, few trials provide adequate information to determine even if there was allocation concealment, how the randomization was conducted, and how successful the masking may have been, let alone if selection bias was adequately controlle d. In …


Improved Multiple Comparisons With The Best In Response Surface Methodology, Laura K. Miller, Ping Sa May 2003

Improved Multiple Comparisons With The Best In Response Surface Methodology, Laura K. Miller, Ping Sa

Journal of Modern Applied Statistical Methods

A method to construct simultaneous confidence intervals about the difference in mean responses at the stationary point and at x for all x within a sphere with radius I R is proposed. Results of an efficiency study to compare the new method and the existing method by Moore and Sa (1999) are provided.


Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little Mar 2003

Penalized Spline Nonparametric Mixed Models For Inference About A Finite Population Mean From Two-Stage Samples, Hui Zheng, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Samplers often distrust model-based approaches to survey inference due to concerns about model misspecification when applied to large samples from complex populations. We suggest that the model-based paradigm can work very successfully in survey settings, provided models are chosen that take into account the sample design and avoid strong parametric assumptions. The Horvitz-Thompson (HT) estimator is a simple design-unbiased estimator of the finite population total in probability sampling designs. From a modeling perspective, the HT estimator performs well when the ratios of the outcome values and the inclusion probabilities are exchangeable. When this assumption is not met, the HT estimator …


A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan Mar 2003

A Semiparametric Model Selection Criterion With Applications To The Marginal Structural Model, M. Alan Brookhart, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Estimators for the parameter of interest in semiparametric models often depend on a guessed model for the nuisance parameter. The choice of the model for the nuisance parameter can affect both the finite sample bias and efficiency of the resulting estimator of the parameter of interest. In this paper we propose a finite sample criterion based on cross validation that can be used to select a nuisance parameter model from a list of candidate models. We show that expected value of this criterion is minimized by the nuisance parameter model that yields the estimator of the parameter of interest with …


Ibd Configuration Transition Matrices And Linkage Score Tests For Unilineal Relative Pairs, Sandrine Dudoit Feb 2003

Ibd Configuration Transition Matrices And Linkage Score Tests For Unilineal Relative Pairs, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Properties of transition matrices between IBD configurations are derived for four general classes of unilineal relative pairs obtained from the grand-parent/ grand-child, half-sib, avuncular, and cousin relationships. In this setting, IBD configurations are defined as orbits of groups acting on a set of inheritance vectors. Properties of the transition matrix between IBD configurations at two linked loci are derived by relating its infinitesimal generator to the adjacency matrix of a quotient graph. The second largest eigenvalue of the infinitesimal generator and its multiplicity are key in determining the form of the transition matrix and of likelihood-based linkage tests such as …


Asymptotic Optimality Of Likelihood Based Cross-Validation, Mark J. Van Der Laan, Sandrine Dudoit, Sunduz Keles Feb 2003

Asymptotic Optimality Of Likelihood Based Cross-Validation, Mark J. Van Der Laan, Sandrine Dudoit, Sunduz Keles

U.C. Berkeley Division of Biostatistics Working Paper Series

Likelihood-based cross-validation is a statistical tool for selecting a density estimate based on n i.i.d. observations from the true density among a collection of candidate density estimators. General examples are the selection of a model indexing a maximum likelihood estimator, and the selection of a bandwidth indexing a nonparametric (e.g. kernel) density estimator. In this article, we establish asymptotic optimality of a general class of likelihood based cross-validation procedures (as indexed by the type of sample splitting used, e.g. V-fold cross-validation), in the sense that the cross-validation selector performs asymptotically as well (w.r.t. to the Kullback-Leibler distance to the true …


Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan Feb 2003

Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Risk estimation is an important statistical question for the purposes of selecting a good estimator (i.e., model selection) and assessing its performance (i.e., estimating generalization error). This article introduces a general framework for cross-validation and derives distributional properties of cross-validated risk estimators in the context of estimator selection and performance assessment. Arbitrary classes of estimators are considered, including density estimators and predictors for both continuous and polychotomous outcomes. Results are provided for general full data loss functions (e.g., absolute and squared error, indicator, negative log density). A broad definition of cross-validation is used in order to cover leave-one-out cross-validation, V-fold …


Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe Jan 2003

Semiparametric Receiver Operating Characteristic Analysis To Evaluate Biomarkers For Disease, Tianxi Cai, Margaret S. Pepe

UW Biostatistics Working Paper Series

The receiver operating characteristic (ROC) curve is a popular method for characterizing the accuracy of diagnostic tests when test results are not binary. Various methodologies for estimating and comparing ROC curves have been developed. One approach, due to Pepe, uses a parametric regression model with the baseline function specified up to a finite-dimensional parameter. In this article we extend the regression models by allowing arbitrary nonparametric baseline functions. We also provide asymptotic distribution theory and procedures for making statistical inference. We illustrate our approach with dataset from a prostate cancer biomarker study. Simulation studies suggest that the extra flexibility inherent …


Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe Jan 2003

Semi-Parametric Regression For The Area Under The Receiver Operating Characteristic Curve, Lori E. Dodd, Margaret S. Pepe

UW Biostatistics Working Paper Series

Medical advances continue to provide new and potentially better means for detecting disease. Such is true in cancer, for example, where biomarkers are sought for early detection and where improvements in imaging methods may pick up the initial functional and molecular changes associated with cancer development. In other binary classification tasks, computational algorithms such as Neural Networks, Support Vector Machines and Evolutionary Algorithms have been applied to areas as diverse as credit scoring, object recognition, and peptide-binding prediction. Before a classifier becomes an accepted technology, it must undergo rigorous evaluation to determine its ability to discriminate between states. Characterization of …


Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger Jan 2003

Checking Assumptions In Latent Class Regression Models Via A Markov Chain Monte Carlo Estimation Approach: An Application To Depression And Socio-Economic Status, Elizabeth Garrett, Richard Miech, Pamela Owens, William W. Eaton, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

Latent class regression models are useful tools for assessing associations between covariates and latent variables. However, evaluation of key model assumptions cannot be performed using methods from standard regression models due to the unobserved nature of latent outcome variables. This paper presents graphical diagnostic tools to evaluate whether or not latent class regression models adhere to standard assumptions of the model: conditional independence and non-differential measurement. An integral part of these methods is the use of a Markov Chain Monte Carlo estimation procedure. Unlike standard maximum likelihood implementations for latent class regression model estimation, the MCMC approach allows us to …


Estimation Of Cumulative Incidence Functions In Competing Risks Studies Under An Order Restriction, Hammou El Barmi, Subhash C. Kochar, Hari Mukerjee, Francisco J. Samaniego Jan 2003

Estimation Of Cumulative Incidence Functions In Competing Risks Studies Under An Order Restriction, Hammou El Barmi, Subhash C. Kochar, Hari Mukerjee, Francisco J. Samaniego

Mathematics and Statistics Faculty Publications and Presentations

In the competing risks problem an important role is played by the cumulative incidence function (CIF), whose value at time t is the probability of failure by time t for a particular type of failure in the presence of other risks. Its estimation and asymptotic distribution theory have been studied by many. In some cases there are reasons to believe that the CIFs due to two types of failure are order restricted. Several procedures have appeared in the literature for testing for such orders. In this paper we initiate the study of estimation of two CIFs subject to a type …


Recurrent Events Analysis In The Presence Of Time Dependent Covariates And Dependent Censoring, Maja Miloslavsky, Sunduz Keles, Mark J. Van Der Laan, Steve Butler Dec 2002

Recurrent Events Analysis In The Presence Of Time Dependent Covariates And Dependent Censoring, Maja Miloslavsky, Sunduz Keles, Mark J. Van Der Laan, Steve Butler

U.C. Berkeley Division of Biostatistics Working Paper Series

Recurrent events models have lately received a lot of attention in the literature. The majority of approaches discussed show the consistency of parameter estimates under the assumption that censoring is independent of the recurrent events process of interest conditional on the covariates included into the model. We provide an overview of available recurrent events analysis methods, and present an inverse probability of censoring weighted estimator for the regression parameters in the Andersen-Gill model that is commonly used for recurrent event analysis. This estimator remains consistent under informative censoring if the censoring mechanism is estimated consistently, and generally improves on the …


Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan Dec 2002

Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Robins' causal inference theory assumes existence of treatment specific counterfactual variables so that the observed data augmented by the counterfactual data will satisfy a consistency and a randomization assumption. Gill and Robins [2001] show that the consistency and randomization assumptions do not add any restrictions to the observed data distribution. In particular, they provide a construction of counterfactuals as a function of the observed data distribution. In this paper we provide a construction of counterfactuals as a function of the observed data itself. Our construction provides a new statistical tool for estimation of counterfactual distributions. Robins [1987b] shows that the …


On The Misuse Of Confidence Intervals For Two Means In Testing For The Significance Of The Difference Between The Means, George W. Ryan, Steven D. Leadbetter Nov 2002

On The Misuse Of Confidence Intervals For Two Means In Testing For The Significance Of The Difference Between The Means, George W. Ryan, Steven D. Leadbetter

Journal of Modern Applied Statistical Methods

Comparing individual confidence intervals of two population means is an incorrect procedure for determining the statistical significance of the difference between the means. We show conditions where confidence intervals for the means from two independent samples overlap and the difference between the means is in fact significant.


Constructive Criticism, Ronald C. Serlin Nov 2002

Constructive Criticism, Ronald C. Serlin

Journal of Modern Applied Statistical Methods

Attempts to attain knowledge as certified true belief have failed to circumvent Hume’s injunction against induction. Theories must be viewed as unprovable, improbable, and undisprovable. The empirical basis is fallible, and yet the method of conjectures and refutations is untouched by Hume’s insights. The implications for statistical methodology is that the requisite severity of testing is achieved through the use of robust procedures, whose assumptions have not been shown to be substantially violated, to test predesignated range null hypotheses. Nonparametric range null hypothesis tests need to be developed to examine whether or not effect sizes or measures of association, as …


Extensions Of The Concept Of Exchangeability And Their Applications, Phillip I. Good Nov 2002

Extensions Of The Concept Of Exchangeability And Their Applications, Phillip I. Good

Journal of Modern Applied Statistical Methods

Permutation tests provide exact p-values in a wide variety of practical testing situations. But permutation tests rely on the assumption of exchangeability, that is, under the hypothesis, the joint distribution of the observations is invariant under permutations of the subscripts. Observations are exchangeable if they are independent, identically distributed (i.i.d.), or if they are jointly normal with identical covariances. The range of applications of these exact, powerful, distribution-free tests can be enlarged through exchangeability- preserving transforms, asymptotic exchangeability, partial exchangeability, and weak exchangeability. Original exact tests for comparing the slopes of two regression lines and for the analysis of …


A Test Of Symmetry, Abdul R. Othman, H. J. Keselman, Rand R. Wilcox, Katherine Fradette, A. R. Padmanabhan Nov 2002

A Test Of Symmetry, Abdul R. Othman, H. J. Keselman, Rand R. Wilcox, Katherine Fradette, A. R. Padmanabhan

Journal of Modern Applied Statistical Methods

When data are nonnormal in form classical procedures for assessing treatment group equality are prone to distortions in rates of Type I error and power to detect effects. Replacing the usual means with trimmed means reduces rates of Type I error and increases sensitivity to detect effects. If data are skewed, say to the right, then it has been postulated that asymmetric trimming, to the right, should be better at controlling rates of Type I error and power to detect effects than symmetric trimming from both tails of the data distribution. Keselman, Wilcox, Othman and Fradette (2002) found that Babu, …


Best Regression Model Using Information Criteria, Phill Gagné, C. Mitchell Dayton Nov 2002

Best Regression Model Using Information Criteria, Phill Gagné, C. Mitchell Dayton

Journal of Modern Applied Statistical Methods

The accuracy of AIC and BIC is evaluated under simulated multiple regression conditions, varying number of total and valid predictors, R2, and n. AIC and BIC were increasingly accurate as n increased and as total predictors decreased. Interactions of the ratio of valid/total predictors affected accuracy.


The Statistical Modeling Of The Fertility Of Chinese Women, Dudley L. Poston Jr. Nov 2002

The Statistical Modeling Of The Fertility Of Chinese Women, Dudley L. Poston Jr.

Journal of Modern Applied Statistical Methods

This article is concerned with the statistical modeling of children ever born (CEB) fertility data. It is shown that in a low fertility population, such as China, the use of linear regression approaches to model CEB is statistically inappropriate because the distribution of the CEB variable is often heavily skewed with a long right tail. For five sub-groups of Chinese women, their fertility is modeled using Poisson, negative binomial, and ordinary least squares (OLS) regression models. It is shown that in almost all instances there would have been major errors of statistical inference had the interpretations of the results been …


Fermat, Schubert, Einstein, And Behrens-Fisher: The Probable Difference Between Two Means When Σ_1^2≠Σ_2^2, Shlomo S. Sawilowsky Nov 2002

Fermat, Schubert, Einstein, And Behrens-Fisher: The Probable Difference Between Two Means When Σ_1^2≠Σ_2^2, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

The history of the Behrens-Fisher problem and some approximate solutions are reviewed. In outlining relevant statistical hypotheses on the probable difference between two means, the importance of the Behrens- Fisher problem from a theoretical perspective is acknowledged, but it is concluded that this problem is irrelevant for applied research in psychology, education, and related disciplines. The focus is better placed on “shift in location” and, more importantly, “shift in location and change in scale” treatment alternatives.


Double Median Ranked Set Sample: Comparing To Other Double Ranked Samples For Mean And Ratio Estimators, Hani M. Samawi, Eman M. Tawalbeh Nov 2002

Double Median Ranked Set Sample: Comparing To Other Double Ranked Samples For Mean And Ratio Estimators, Hani M. Samawi, Eman M. Tawalbeh

Journal of Modern Applied Statistical Methods

Double median ranked set sample (DMRSS) and its properties for estimating the population mean, when the underlying distribution is assumed to be symmetric about its mean, are introduced. Also, the performance of DMRSS with respect to other ranked set samples and double ranked set samples, for estimating the population mean and ratio, is considered. Real data that consist of heights and diameters of 399 trees are used to illustrate the procedure. The analysis and simulation indicate that using DMRSS for estimating the population mean is more efficient than using the other ranked samples and double ranked samples schemes except in …


Robust Estimation Of Multivariate Failure Data With Time-Modulated Frailty, Pingfu Fu, J. Sunil Rao, Jiming Jiang Nov 2002

Robust Estimation Of Multivariate Failure Data With Time-Modulated Frailty, Pingfu Fu, J. Sunil Rao, Jiming Jiang

Journal of Modern Applied Statistical Methods

A time-modulated frailty model is proposed for analyzing multivariate failure data. The effect of frailties, which may not be constant over time, is discussed. We assume a parametric model for the baseline hazard, but avoid the parametric assumption for the frailty distribution. The well-known connection between survival times and Poisson regression model is used. The parameters of interest are estimated by generalized estimating equations (GEE) or by penalized GEE. Simulation studies show that the procedure is successful to detect the effect of time-modulated frailty. The method is also applied to a placebo controlled randomized clinical trial of gamma interferon, a …