Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 31 - 60 of 72

Full-Text Articles in Statistical Theory

Jmasm4: Critical Values For Four Nonparametric And/Or Distribution-Free Tests Of Location For Two Independent Samples, Bruce R. Fay Nov 2002

Jmasm4: Critical Values For Four Nonparametric And/Or Distribution-Free Tests Of Location For Two Independent Samples, Bruce R. Fay

Journal of Modern Applied Statistical Methods

Researchers engaged in computer-intensive studies may need exact critical values, especially for sample sizes and alpha levels not normally found in published tables, as well as the ability to control ‘best-fit’ criteria. They may also benefit from the ability to directly generate these values rather than having to create lookup tables. Fortran 90 programs generate ‘best-conservative’ (bc) and ‘best-fit’ (bf) critical values with associated probabilities for the Kolmogorov-Smirnov test of general differences (bc), Rosenbaum’s test of location (bc), Tukey’s quick test (bc and bf)) and the Wilcoxon rank-sum test (bc).


Locally Efficient Estimation With Bivariate Right Censored Data , Christopher M. Quale, Mark J. Van Der Laan, James M. Robins Oct 2002

Locally Efficient Estimation With Bivariate Right Censored Data , Christopher M. Quale, Mark J. Van Der Laan, James M. Robins

U.C. Berkeley Division of Biostatistics Working Paper Series

Estimation for bivariate right censored data is a problem that has had much study over the past 15 years. In this paper we propose a new class of estimators for the bivariate survivor function based on locally efficient estimation. The locally efficient estimator takes bivariate estimators Fn and Gn of the distributions of the time variables T1,T2 and the censoring variables C1,C2, respectively, and maps them to the resulting estimator. If Fn and Gn are consistent estimators of F and G, respectively, then the resulting estimator will be nonparametrically efficient (thus the term ``locally efficient''). However, if either Fn or …


The Analysis Of Placement Values For Evaluating Discriminatory Measures, Margaret S. Pepe, Tianxi Cai Sep 2002

The Analysis Of Placement Values For Evaluating Discriminatory Measures, Margaret S. Pepe, Tianxi Cai

UW Biostatistics Working Paper Series

The idea of using measurements such as biomarkers, clinical data, or molecular biology assays for classification and prediction is popular in modern medicine. The scientific evaluation of such measures includes assessing the accuracy with which they predict the outcome of interest. Receiver operating characteristic curves are commonly used for evaluating the accuracy of diagnostic tests. They can be applied more broadly, indeed to any problem involving classification to two states or populations (D = 0 or D = 1). We show that the ROC curve can be interpreted as a cumulative distribution function for the discriminatory measure Y in the …


Accelerated Hazards Model: Method, Theory And Applications, Ying Qing Chen, Nicholas P. Jewell, Jingrong Yang Sep 2002

Accelerated Hazards Model: Method, Theory And Applications, Ying Qing Chen, Nicholas P. Jewell, Jingrong Yang

U.C. Berkeley Division of Biostatistics Working Paper Series

In an accelerated hazards model, the hazard functions of a failure time are related through the time scale-change, which is often a function of covariates and associated parameters. When the hazard functions have special properties, such as monotonicity in time, the parameters may be clinically meaningful in measuring a treatment effect. This paper reviews methodological and theoretical development of this model. Applications of the accelerated hazards model including sample size calculation in clinical trials, are also explored.


Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins Sep 2002

Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins

U.C. Berkeley Division of Biostatistics Working Paper Series

In biostatistics applications interest often focuses on the estimation of the distribution of a time-variable T. If one only observes whether or not T exceeds an observed monitoring time C, then the data structure is called current status data, also known as interval censored data, case I. We consider this data structure extended to allow the presence of both time-independent covariates and time-dependent covariate processes that are observed until the monitoring time. We assume that the monitoring process satisfies coarsening at random.

Our goal is to estimate the regression parameter beta of the regression model T = Z*beta+epsilon where the …


Case-Control Current Status Data, Nicholas P. Jewell, Mark J. Van Der Laan Sep 2002

Case-Control Current Status Data, Nicholas P. Jewell, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Current status observation on survival times has recently been widely studied. An extreme form of interval censoring, this data structure refers to situations where the only available information on a survival random variable, T, is whether or not T exceeds a random independent monitoring time C, a binary random variable, Y. To date, nonparametric analyses of current status data have assumed the availability of i.i.d. random samples of the random variable (Y, C), or a similar random sample at each of a set of fixed monitoring times. In many situations, it is useful to consider a case-control sampling scheme. Here, …


Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan Sep 2002

Why Prefer Double Robust Estimates? Illustration With Causal Point Treatment Studies, Romain Neugebauer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In point treatment marginal structural models with treatment A, outcome Y and covariates W, causal parameters can be estimated under the assumption of no unobserved confounders. Three estimates can be used: the G-computation, Inverse Probability of Treatment Weighted (IPTW) or Double Robust (DR) estimates. The properties of the IPTW and DR estimates are known under an assumption on the treatment mechanism that we name "Experimental Treatment Assignment" (ETA) assumption. We show that the DR estimating function is unbiased when the ETA assumption is violated if the model used to regress Y on A and W is correctly specified. The practical …


Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell Sep 2002

Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

In many applications, it is often of interest to estimate a bivariate distribution of two survival random variables. Complete observation of such random variables is often incomplete. If one only observes whether or not each of the individual survival times exceeds a common observed monitoring time C, then the data structure is referred to as bivariate current status data (Wang and Ding, 2000). For such data, we show that the identifiable part of the joint distribution is represented by three univariate cumulative distribution functions, namely the two marginal cumulative distribution functions, and the bivariate cumulative distribution function evaluated on the …


Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan Sep 2002

Current Status Data: Review, Recent Developments And Open Problems, Nicholas P. Jewell, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Researchers working with survival data are by now adept at handling issues associated with incomplete data, particular those associated with various forms of censoring. An extreme form of interval censoring, known as current status observation, refers to situations where the only available information on a survival random variable T is whether or not T exceeds a random independent monitoring time C. This article contains a brief review of the extensive literature on the analysis of current status data, discussing the implications of response-based sampling on these methods. The majority of the paper introduces some recent extensions of these ideas to …


Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang Aug 2002

Semiparametric Regression Analysis On Longitudinal Pattern Of Recurrent Gap Times, Ying Qing Chen, Mei-Cheng Wang, Yijian Huang

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal studies, individual subjects may experience recurrent events of the same type over a relatively long period of time. The longitudinal pattern of the gaps between the successive recurrent events is often of great research interest. In this article, the probability structure of the recurrent gap times is first explored in the presence of censoring. According to the discovered structure, we introduce the proportional reverse-time hazards models with unspecified baseline functions to accommodate heterogeneous individual underlying distributions, when the ongitudinal pattern parameter is of main interest. Inference procedures are proposed and studied by way of proper riskset construction. The …


Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick Aug 2002

Multiple Hypothesis Testing In Microarray Experiments, Sandrine Dudoit, Juliet Popper Shaffer, Jennifer C. Boldrick

U.C. Berkeley Division of Biostatistics Working Paper Series

DNA microarrays are a new and promising biotechnology which allows the monitoring of expression levels in cells for thousands of genes simultaneously. An important and common question in microarray experiments is the identification of differentially expressed genes, i.e., genes whose expression levels are associated with a response or covariate of interest. The biological question of differential expression can be restated as a problem in multiple hypothesis testing: the simultaneous test for each gene of the null hypothesis of no association between the expression levels and the responses or covariates. As a typical microarray experiment measures expression levels for thousands of …


Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins Aug 2002

Estimation Of The Bivariate Survival Function With Generalized Bivariate Right Censored Data Structures, Sunduz Keles, Mark J. Van Der Laan, James M. Robins

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a bivariate survival function estimator for a general right censored data structure that includes a time dependent covariate process. Firstly, an initial estimator that generalizes Dabrowska's (1988) estimator is introduced. We obtain this estimator by a general methodology of constructing estimating functions in censored data models. The initial estimator is guaranteed to improve on Dabrowska's estimator and remains consistent and asymptotically linear under informative censoring schemes if the censoring mechanism is estimated consistently. We then construct an orthogonalized estimating function which results in a more robust and efficient estimator than our initial estimator. A simulation study demonstrates the …


Nonlinear Regression Based On Ranks, Asheber Abebe Jun 2002

Nonlinear Regression Based On Ranks, Asheber Abebe

Dissertations

This study presents robust methods for estimating parameters of nonlinear regression models. The proposed methods obtain estimates by minimizing rankbased dispersions instead of the Euclidean norm. We focus on the Wilcoxon and generalized signed-rank dispersion functions. Asymptotic properties of the estimators are established under mild regularity conditions similar to those used in least squares and least absolute deviations estimation. The study also shows that by considering the generalized signed-rank dispersion we obtain a class of estimators that encompasses most of the existing popular nonlinear regression estimators. As in linear models, these rank-based procedures provide estimators that are highly efficient. This …


Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell May 2002

Inference For Proportional Mean Residual Life Model In The Presence Of Censoring, Ying Q. Chen, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

As a function of time t, mean residual life is defined as remaining life expectancy of a subject given its survival to t. It plays an important role in many research areas to characterise stochastic behavior of survival over time. Similar to the Cox proportional hazard model, the proportional mean residual life model were proposed in statistical literature to study association between the mean residual life and individual subject's explanatory covariates. In this article, we will study this model and develop appropriate inference procedures in presence of censoring. Numerical studies including simulation and real data analysis are presented as well.


Power Analyses When Comparing Trimmed Means, Rand R. Wilcox, H. J. Keselman May 2002

Power Analyses When Comparing Trimmed Means, Rand R. Wilcox, H. J. Keselman

Journal of Modern Applied Statistical Methods

Given a random sample from each of two independent groups, this article takes up the problem of estimating power, as well as a power curve, when comparing 20% trimmed means with a percentile bootstrap method. Many methods were considered, but only one was found to be satisfactory in terms of obtaining both a point estimate of power as well as a (one-sided) confidence interval. The method is illustrated with data from a reading study where theory suggests two groups should differ but nonsignificant results were obtained.


Some Locally Most Powerful Rank Tests For Correlation, W. J. Conover May 2002

Some Locally Most Powerful Rank Tests For Correlation, W. J. Conover

Journal of Modern Applied Statistical Methods

Four examples are given to illustrate the ease and practicality of the procedure for finding locally most powerful rank tests for correlation. The first two examples deal with bivariate exponential models. The third example uses the bivariate normal distribution, and the fourth example analyzes the Morgenstem’s general correlation model.


Exact Level And Power Of Permutation, Bootstrap, And Asymptotic Tests Of Trend, Christopher D. Corcoran, Cyrus R. Mehta May 2002

Exact Level And Power Of Permutation, Bootstrap, And Asymptotic Tests Of Trend, Christopher D. Corcoran, Cyrus R. Mehta

Journal of Modern Applied Statistical Methods

We develop computational tools that can evaluate the exact size and power of three tests of trend (e.g., permutation, bootstrap and asymptotic) without resorting to large-sample theory or simulations. We then use these tools to compare the operating characteristics of the three tests. It is seen that the bootstrap test is ultra-conservative relative to the other two tests and as a result suffers from a severe deterioration in power. The power of the asymptotic test is uniformly larger than that of the other two tests, but it fails to preserve the Type I error for most of the range of …


Six Modifications Of The Aligned Rank Transform Test For Interaction, Kathleen Peterson May 2002

Six Modifications Of The Aligned Rank Transform Test For Interaction, Kathleen Peterson

Journal of Modern Applied Statistical Methods

Testing for interactions in multivariate experiments is an important function. Studies indicate that much data from social studies research is not normally distributed, thus violating that assumption of the AN OVA procedure. The aligned rank transformation test (ART), aligning using the means of columns and rows, has been found, in limited situations, to be robust to Type I error rates and to have greater power than the ANOVA. This study explored a variety of alignments, including the median, Winsorized trimmed means (10%) and (20%), the Huber1.28 M-estimator, and the Harrell-Davis estimator of the median. Results are reported for Type …


Two Methods To Estimate Homogenous Markov Processes, Ricardo Ocaña-Rilola May 2002

Two Methods To Estimate Homogenous Markov Processes, Ricardo Ocaña-Rilola

Journal of Modern Applied Statistical Methods

Multi-state Markov processes have been introduced recently in Health Sciences in order to study disease history events. This sort of model have some advantages respect to traditional survival analysis, therefore they are an important line of research into stochastic processes applied to Epidemiology. However these types of models increase the complexity of analysis, even for simpler processes, and standard software is limited. In this paper, two methods for fitting homogeneous Markov models are proposed and compared.


The Trouble With Trivials (P > .05), Shlomo S. Sawilowsky, Jina S. Yoon May 2002

The Trouble With Trivials (P > .05), Shlomo S. Sawilowsky, Jina S. Yoon

Journal of Modern Applied Statistical Methods

Trivials are effect sizes associated with statistically non-significant results. Trivials are like Tribbles in the Star Trek television show. They are cute and loveable. They proliferate without limit. They probably growl at Bayesians. But they are troublesome. This brief report discusses the trouble with trivials.


An Error In Statistical Logic In The Application Of Genetic Paternity Testing, Ernest P. Chiodo, Joseph L. Musial, J. Sia Robinson May 2002

An Error In Statistical Logic In The Application Of Genetic Paternity Testing, Ernest P. Chiodo, Joseph L. Musial, J. Sia Robinson

Journal of Modern Applied Statistical Methods

A Bayes probability computer program was written in Fortran to examine issues related to genetic paternity testing. An application was given to demonstrate the effects improper assumptions of prior probability of culpability. The seriousness of such errors include the potential of assigning paternity to wrongly accused men, or wrongly refuting paternity.


An Adaptive Inference Strategy: The Case Of Auditory Data, Bruno D. Zumbo May 2002

An Adaptive Inference Strategy: The Case Of Auditory Data, Bruno D. Zumbo

Journal of Modern Applied Statistical Methods

By way of an example some of the basic features in the derivation and use of adaptive inferential methods are demonstrated. The focus of this paper is dyadic (coupled) data in auditory and perceptual research. We present: (a) why one should not use the conventional methods, (b) a derivation of an adaptive method, and (c) how the new adaptive method works with the example data. In the concluding remarks we draw attention to the work of Professor George Barnard who provided the adaptive inference strategy in the context of the Behrens-Fisher problem -- testing the equality of means when one …


Parametric Analyses In Randomized Clinical Trials, Vance W. Berger, Clifford E. Lunneborg, Michael D. Ernst, Jonathan G. Levine May 2002

Parametric Analyses In Randomized Clinical Trials, Vance W. Berger, Clifford E. Lunneborg, Michael D. Ernst, Jonathan G. Levine

Journal of Modern Applied Statistical Methods

One salient feature of randomized clinical trials is that patients are randomly allocated to treatment groups, but not randomly sampled from any target population. Without random sampling parametric analyses are inexact, yet they are still often used in clinical trials. Given the availability of an exact test, it would still be conceivable to argue convincingly that for technical reasons (upon which we elaborate) a parametric test might be preferable in some situations. Having acknowledged this possibility, we point out that such an argument cannot be convincing without supporting facts concerning the specifics of the problem at hand. Moreover, we have …


Alternatives To SW In The Bracketed Interval Of The Trimmed Mean, Jennifer Bunner, Shlomo S. Sawilowsky May 2002

Alternatives To SW In The Bracketed Interval Of The Trimmed Mean, Jennifer Bunner, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

The aim of this Monte Carlo study is to examine alternatives to estimated variability in building bracketed intervals about the trimmed mean.


Generation Of Combinations Using Excel, Constantine Stamatopoulos May 2002

Generation Of Combinations Using Excel, Constantine Stamatopoulos

Journal of Modern Applied Statistical Methods

Theoretical development of combinations via enumeration methods are considered. An Excel macro is provided.


Jmasm1: Rangen 2.0 (Fortran 90/95), Gail F. Fahoome May 2002

Jmasm1: Rangen 2.0 (Fortran 90/95), Gail F. Fahoome

Journal of Modern Applied Statistical Methods

Rangen 2.0 is Fortran 90 module of subroutines used to generate uniform and nonuniform pseudo-random deviates. It includes uni1, an uniform pseudo-random number generator, and non-uniform generators based on unil. The subroutines in Rangen 2.0 were written using Essential Lahey Fortran 90, a proper subset of Fortran 90. It includes both source code for the subroutines and a short description of each subroutine, its purpose, and the arguments including data type and usage.


Asymptotic And Exact Tests In 2 X C Ordered Categorical Contingency Tables With Statxact 2.0 - 4.0, Margaret Posch May 2002

Asymptotic And Exact Tests In 2 X C Ordered Categorical Contingency Tables With Statxact 2.0 - 4.0, Margaret Posch

Journal of Modern Applied Statistical Methods

The purpose of this study was to compare the statistical power of a variety of exact tests in the 2 x C ordered categorical contingency table using StatXact software. The Wilcoxon Rank Sum, Expected Nonnal Scores, Savage Scores (or its Log Rank equivalent), and Permutation tests were studied. Results indicated that the procedures were nearly the same in terms of comparative statistical power.


Quantifying Bimodality Part I: An Easily Implemented Method Using Spss, B. W. Frankland, Bruno D. Zumbo May 2002

Quantifying Bimodality Part I: An Easily Implemented Method Using Spss, B. W. Frankland, Bruno D. Zumbo

Journal of Modern Applied Statistical Methods

Scientists in a variety of fields are faced with the question of whether or not a particular sample of data are best described as unimodal or bimodal. We provide a simple and convenient method for assessing bimodality. The use of the non-linear algorithms in SPSS for modeling complex mixture distributions is demonstrated on a unimodal normal distribution (with 2 free parameters) and on bimodal mixture of two normal distributions (with 5 free parameters).


A Measure Of Relative Efficiency For Location Of A Single Sample, Shlomo S. Sawilowsky May 2002

A Measure Of Relative Efficiency For Location Of A Single Sample, Shlomo S. Sawilowsky

Journal of Modern Applied Statistical Methods

The question of how much to trim or which weighting constant to use are practical considerations in applying robust methods such as trimmed means (L-estimators) and Huber statistics (M-estimators). An index oflocation relative efficiency (LRE), which is a ratio of the narrowness of resulting confidence intervals, was applied to various trimmed means and Huber M-estimators calculated on seven representative data sets from applied education and psychology research. On the basis of LREs, lightly trimmed means were found to be more efficient than heavily trimmed means, but Huber M-estimators systematically produced narrower confidence intervals. The weighting constant of ψ = 1.28 …


Two-Sided Equivalence Testing Of The Difference Between Two Means, R. Clifford Blair, Stephen R. Cole May 2002

Two-Sided Equivalence Testing Of The Difference Between Two Means, R. Clifford Blair, Stephen R. Cole

Journal of Modern Applied Statistical Methods

Studies designed to examine the equivalence of treatments are increasingly common in social and biomedical research. Herein, we outline the rationale and some nuances underlying equivalence testing of the difference between two means. Specifically, we note the odd relation between tests of hypothesis and confidence intervals in the equivalence setting.