Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons™

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type

Articles 1291 - 1320 of 1633

Full-Text Articles in Statistical Theory

Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin Sep 2005

Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin

Harvard University Biostatistics Working Paper Series

There is an emerging interest in modeling spatially correlated survival data in biomedical and epidemiological studies. In this paper, we propose a new class of semiparametric normal transformation models for right censored spatially correlated survival data. This class of models assumes that survival outcomes marginally follow a Cox proportional hazard model with unspecified baseline hazard, and their joint distribution is obtained by transforming survival outcomes to normal random variables, whose joint distribution is assumed to be multivariate normal with a spatial correlation structure. A key feature of the class of semiparametric normal transformation models is that it provides a rich …


Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan Sep 2005

Inference On Survival Data With Covariate Measurement Error - An Imputation-Based Approach, Yi Li, Louise Ryan

Harvard University Biostatistics Working Paper Series

We propose a new method for fitting proportional hazards models with error-prone covariates. Regression coefficients are estimated by solving an estimating equation that is the average of the partial likelihood scores based on imputed true covariates. For the purpose of imputation, a linear spline model is assumed on the baseline hazard. We discuss consistency and asymptotic normality of the resulting estimators, and propose a stochastic approximation scheme to obtain the estimates. The algorithm is easy to implement, and reduces to the ordinary Cox partial likelihood approach when the measurement error has a degenerative distribution. Simulations indicate high efficiency and robustness. …


The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek Sep 2005

The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek

UW Biostatistics Working Paper Series

As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the Optimal Discovery Procedure (ODP), which has recently been introduced and theoretically shown …


Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen Aug 2005

Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

The causal effect of a treatment on an outcome is generally mediated by several intermediate variables. Estimation of the component of the causal effect of a treatment that is mediated by a given intermediate variable (the indirect effect of the treatment), and the component that is not mediated by that intermediate variable (the direct effect of the treatment) is often relevant to mechanistic understanding and to the design of clinical and public health interventions. Under the assumption of no-unmeasured confounders for treatment and the intermediate variable, Robins & Greenland (1992) define an individual direct effect as the counterfactual effect of …


Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan Aug 2005

Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …


Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan Aug 2005

Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We present a cross-validated bagging scheme in the context of partitioning algorithms. To explore the benefits of the various bagging scheme, we compare via simulations the predictive ability of single Classification and Regression (CART) Tree with several previously suggested bagging schemes and with our proposed approach. Additionally, a variable importance measure is explained and illustrated.


Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit Jul 2005

Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …


Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang Jul 2005

Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang

UW Biostatistics Working Paper Series

Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.


A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan Jun 2005

A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Robins' causal inference theory assumes existence of treatment specific counterfactual variables so that the observed data augmented by the counterfactual data will satisfy a consistency and a randomization assumption. In this paper we provide an explicit function that maps the observed data into a counterfactual variable which satisfies the consistency and randomization assumptions. This offers a practically useful imputation method for counterfactuals. Gill & Robins [2001]'s construction of counterfactuals can be used as an imputation method in principle, but it is very hard to implement in practice. Robins [1987] shows that the counterfactual distribution can be identified from the observed …


Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen Jun 2005

Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

Many applications aim to learn a high dimensional parameter of a data generating distribution based on a sample of independent and identically distributed observations. For example, the goal might be to estimate the conditional mean of an outcome given a list of input variables. In this prediction context, Breiman (1996a) introduced bootstrap aggregating (bagging) as a method to reduce the variance of a given estimator at little cost to bias. Bagging involves applying the estimator to multiple bootstrap samples, and averaging the result across bootstrap samples. In order to deal with the curse of dimensionality, typical practice has been to …


On Additive Regression Of Expectancy, Ying Qing Chen Jun 2005

On Additive Regression Of Expectancy, Ying Qing Chen

UW Biostatistics Working Paper Series

Regression models have been important tools to study the association between outcome variables and their covariates. The traditional linear regression models usually specify such an association by the expectations of the outcome variables as function of the covariates and some parameters. In reality, however, interests often focus on their expectancies characterized by the conditional means. In this article, a new class of additive regression models is proposed to model the expectancies. The model parameters carry practical implication, which may allow the models to be useful in applications such as treatment assessment, resource planning or short-term forecasting. Moreover, the new model …


An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley Jun 2005

An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley

UW Biostatistics Working Paper Series

We consider data that are dependent, but where most small sets of observations are independent. By extending Bernstein's inequality we prove a strong law of law numbers and an empirical process central limit theorem under bracketing entropy conditions.


New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski May 2005

New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski

COBRA Preprint Series

As the field of functional genetics and genomics is beginning to mature, we become confronted with new challenges. The constant drop in price for sequencing and gene expression profiling as well as the increasing number of genetic and genomic variables that can be measured makes it feasible to address more complex questions. The success with rare diseases caused by single loci or genes has provided us with a proof-of-concept that new therapies can be developed based on functional genomics and genetics.

Common diseases, however, typically involve genetic epistasis, genomic pathways, and proteomic pattern. Moreover, to better understand the underlying biologi-cal …


Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou May 2005

Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In the case in which all subjects are screened using a common test, and only a subset of these subjects are tested using a golden standard test, it is well documented that there is a risk for bias, called verification bias. When the test has only two levels (e.g. positive and negative) and we are trying to estimate the sensitivity and specificity of the test, one is actually constructing a confidence interval for a binomial proportion. Since it is well documented that this estimation is not trivial even with complete data, we adopt Multiple imputation (MI) framework for verification bias …


On The Power Function Of Bayesian Tests With Application To Design Of Clinical Trials: The Fixed-Sample Case, Lyle Broemeling, Dongfeng Wu May 2005

On The Power Function Of Bayesian Tests With Application To Design Of Clinical Trials: The Fixed-Sample Case, Lyle Broemeling, Dongfeng Wu

Journal of Modern Applied Statistical Methods

Using a Bayesian approach to clinical trial design is becoming more common. For example, at the MD Anderson Cancer Center, Bayesian techniques are routinely employed in the design and analysis of Phase I and II trials. It is important that the operating characteristics of these procedures be determined as part of the process when establishing a stopping rule for a clinical trial. This study determines the power function for some common fixed-sample procedures in hypothesis testing, namely the one and two-sample tests involving the binomial and normal distributions. Also considered is a Bayesian test for multi-response (response and toxicity) in …


Two Sides Of The Same Coin: Bootstrapping The Restricted Vs. Unrestricted Model, Panagiotis Mantalos May 2005

Two Sides Of The Same Coin: Bootstrapping The Restricted Vs. Unrestricted Model, Panagiotis Mantalos

Journal of Modern Applied Statistical Methods

The properties of the bootstrap test for restrictions are studied in two versions: 1) bootstrapping under the null hypothesis, restricted, and 2) bootstrapping under the alternative hypothesis, unrestricted. This article demonstrates the equivalence of these two methods, and illustrates the small sample properties of the Wald test for testing Granger-Causality in a stable stationary VAR system by Monte Carlo methods. The analysis regarding the size of the test reveals that, as expected, both bootstrap tests have actual sizes that lie close to the nominal size. Regarding the power of the test, the Wald and bootstrap tests share the same power …


Manifestation Of Differences In Item-Level Characteristics In Scale-Level Measurement Invariance Tests Of Multi-Group Confirmatory Factor Analyses, Bruno D. Zumbo, Kim H. Koh May 2005

Manifestation Of Differences In Item-Level Characteristics In Scale-Level Measurement Invariance Tests Of Multi-Group Confirmatory Factor Analyses, Bruno D. Zumbo, Kim H. Koh

Journal of Modern Applied Statistical Methods

If a researcher applies the conventional tests of scale-level measurement invariance through multi-group confirmatory factor analysis of a PC matrix and MLE to test hypotheses of strong and full measurement invariance when the researcher has a rating scale response format wherein the item characteristics are different for the two groups of respondents, do these scale-level analyses reflect (or ignore) differences in item threshold characteristics? Results of the current study demonstrate the inadequacy of judging the suitability of a measurement instrument across groups by only investigating the factor structure of the measure for the different groups with a PC matrix and …


Right-Tailed Testing Of Variance For Non-Normal Distributions, Michael C. Long, Ping Sa May 2005

Right-Tailed Testing Of Variance For Non-Normal Distributions, Michael C. Long, Ping Sa

Journal of Modern Applied Statistical Methods

A new test of variance for non-normal distribution with fewer restrictions than the current tests is proposed. Simulation study shows that the new test controls the Type I error rate well, and has power performance comparable to the competitors. In addition, it can be used without restrictions.


Multiple Imputation For Missing Ordinal Data, Ling Chen, Marian Toma-Drane, Robert F. Valois, J. Wanzer Drane May 2005

Multiple Imputation For Missing Ordinal Data, Ling Chen, Marian Toma-Drane, Robert F. Valois, J. Wanzer Drane

Journal of Modern Applied Statistical Methods

Simulations were used to compare complete case analysis of ordinal data with including multivariate normal imputations. MVN methods of imputation were not as good as using only complete cases. Bias and standard errors were measured against coefficients estimated from logistic regression and a standard data set.


Coverage Properties Of Optimized Confidence Intervals For Proportions, John P. Wendell, Sharon P. Cox May 2005

Coverage Properties Of Optimized Confidence Intervals For Proportions, John P. Wendell, Sharon P. Cox

Journal of Modern Applied Statistical Methods

Wardell (1997) provided a method for constructing confidence intervals on a proportion that modifies the Clopper-Pearson (1934) interval by allowing for the upper and lower binomial tail probabilities to be set in a way that minimizes the interval width. This article investigates the coverage properties of these optimized intervals. It is found that the optimized intervals fail to provide coverage at or above the nominal rate over some portions of the binomial parameter space but may be useful as an approximate method.


Testing The Goodness Of Fit Of Multivariate Multiplicative-Intercept Risk Models Based On Case-Control Data, Biao Zhang May 2005

Testing The Goodness Of Fit Of Multivariate Multiplicative-Intercept Risk Models Based On Case-Control Data, Biao Zhang

Journal of Modern Applied Statistical Methods

The validity of the multivariate multiplicative-intercept risk model with I +1 categories based on casecontrol data is tested. After reparametrization, the assumed risk model is equivalent to an (I +1) -sample semiparametric model in which the I ratios of two unspecified density functions have known parametric forms. By identifying this (I +1) -sample semiparametric model, which is of intrinsic interest in general (I +1) -sample problems, with an (I +1) -sample semiparametric selection bias model, we propose a weighted Kolmogorov-Smirnov-type statistic to test the validity of the multivariate multiplicativeintercept risk model. Established are some asymptotic results …


Within By Within Anova Based On Medians, Rand R. Wilcox May 2005

Within By Within Anova Based On Medians, Rand R. Wilcox

Journal of Modern Applied Statistical Methods

This article considers a J by K ANOVA design where all JK groups are dependent and where groups are to be compared based on medians. Two general approaches are considered. The first is based on an omnibus test for no main effects and no interactions and the other tests each member of a collection of relevant linear contrasts. Based on an earlier paper dealing with multiple comparisons, an obvious speculation is that a particular bootstrap method should be used. One of the main points here is that, in general, this is not the case for the problem at hand. The …


Testing The Casual Relation Between Sunspots And Temperature Using Wavelets Analysis, Abdullah Almasri, Ghazi Shukur May 2005

Testing The Casual Relation Between Sunspots And Temperature Using Wavelets Analysis, Abdullah Almasri, Ghazi Shukur

Journal of Modern Applied Statistical Methods

Investigated and tested in this article are the causal nexus between sunspots and temperature by using statistical methodology and causality tests. Because this kind of relationship cannot be properly captured in the short run (daily, monthly or yearly data), the relationship is investigated in the long run using a very low frequency Wavelets-based decomposed data such as D8 (128 - 256 months). Results indicate that during the period 1854-1989, the causality nexus between these two series is as expected of onedirectional form, i.e., from sunspots to temperature.


Model-Selection-Based Monitoring Of Structural Change, Kosei Fukuda May 2005

Model-Selection-Based Monitoring Of Structural Change, Kosei Fukuda

Journal of Modern Applied Statistical Methods

Monitoring structural change is performed not by hypothesis testing but by model selection using a modified Bayesian information criterion. It is found that concerning detection accuracy and detection speed, the proposed method shows better performance than the hypothesis-testing method. Two advantages of the proposed method are also discussed.


Using Scale Mixtures Of Normals To Model Continuously Compounded Returns, Hasan Hamdan, John Nolan, Melanie Wilson, Kristen Dardia May 2005

Using Scale Mixtures Of Normals To Model Continuously Compounded Returns, Hasan Hamdan, John Nolan, Melanie Wilson, Kristen Dardia

Journal of Modern Applied Statistical Methods

A new method for estimating the parameters of scale mixtures of normals (SMN) is introduced and evaluated. The new method is called UNMIX and is based on minimizing the weighted square distance between exact values of the density of the scale mixture and estimated values using kernel smoothing techniques over a specified grid of x-values and a grid of potential scale values. Applications of the method are made in modeling the continuously compounded return, CCR, of stock prices. Modeling this ratio with UNMIX proves promising in comparison with other existing techniques that use only one normal component, or those that …


Bayesian Reliability Modeling Using Monte Carlo Integration, Vincent A. R. Camara, Chris P. Tsokos May 2005

Bayesian Reliability Modeling Using Monte Carlo Integration, Vincent A. R. Camara, Chris P. Tsokos

Journal of Modern Applied Statistical Methods

Bayesian Reliability Modeling Using Monte Carlo IntegrationThe aim of this article is to introduce the concept of Monte Carlo Integration in Bayesian estimation and Bayesian reliability analysis. Using the subject concept, approximate estimates of parameters and reliability functions are obtained for the three-parameter Weibull and the gamma failure models. Four different loss functions are used: square error, Higgins-Tsokos, Harris, and a logarithmic loss function proposed in this article. Relative efficiency is used to compare results obtained under the above mentioned loss functions.


Exploratory Factor Analysis In Two Measurement Journals: Hegemony By Default, J. Thomas Kellow May 2005

Exploratory Factor Analysis In Two Measurement Journals: Hegemony By Default, J. Thomas Kellow

Journal of Modern Applied Statistical Methods

Exploratory factor analysis studies in two prominent measurement journals were explored. Issues addressed were: (a) factor extraction methods, (b) factor retention rules, (c) factor rotation strategies, and (d) saliency criteria for including variables. Many authors continue to use principal components extraction, orthogonal (varimax) rotation, and retain factors with eigenvalues greater than 1.0.


An Algorithm For Generating Unconditional Exact Permutation Distribution For A Two-Sample Experiment, Justice I. Odiase, Sunday M. Ogbonmwan May 2005

An Algorithm For Generating Unconditional Exact Permutation Distribution For A Two-Sample Experiment, Justice I. Odiase, Sunday M. Ogbonmwan

Journal of Modern Applied Statistical Methods

An Algorithm that generates the unconditional exact permutation distribution of a 2 x n experiment is presented. The algorithm is able to handle ranks as well as actual observations. It makes it possible to obtain exact p-values for several statistics, especially when sample sizes are small and the application of large sample approximation is unreliable. An illustrative implementation is achieved and leads to the computation of exact p-values for the Mood test when the sample size is small.


A Comparison Of Nonlinear Regression Codes, Paul Fredrick Mondragon, Brian Borchers May 2005

A Comparison Of Nonlinear Regression Codes, Paul Fredrick Mondragon, Brian Borchers

Journal of Modern Applied Statistical Methods

Five readily available software packages were tested on nonlinear regression test problems from the NIST Statistical Reference Datasets. None of the packages was consistently able to obtain solutions accurate to at least three digits. However, two of the packages were somewhat more reliable than the others.


Bias Of The Cox Model Hazard Ratio, Inger Persson, Harry Khamis May 2005

Bias Of The Cox Model Hazard Ratio, Inger Persson, Harry Khamis

Journal of Modern Applied Statistical Methods

The hazard ratio estimated with the Cox model is investigated under proportional and five forms of nonproportional hazards. Results indicate that the highest bias occurs for diverging hazards with early censoring, and for increasing and crossing hazards under a high censoring rate.