Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 841 - 870 of 1108

Full-Text Articles in Statistics and Probability

The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek Sep 2005

The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek

UW Biostatistics Working Paper Series

As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the Optimal Discovery Procedure (ODP), which has recently been introduced and theoretically shown …


Comparison Of Affymetrix Genechip Expression Measures, Rafael A. Irizarry, Zhijin Wu, Harris A. Jaffee Sep 2005

Comparison Of Affymetrix Genechip Expression Measures, Rafael A. Irizarry, Zhijin Wu, Harris A. Jaffee

Johns Hopkins University, Dept. of Biostatistics Working Papers

Affymetrix GeneChip expression array technology has become a standard tool in medical science and basic biology research. In this system, preprocessing occurs before one obtains expression level measurements. Because the number of competing preprocessing methods was large and growing, in the summer of 2003 we developed a benchmark to help users of the technology identify the best method for their application. In conjunction with the release of a Bioconductor R package (affycomp), a webtool was made available for developers of preprocessing methods to submit them to a benchmark for comparison. There have now been over 30 methods compared via the …


Sample Size And Power Calculations For Body Weight In Beef Cattle, Claudia Cristina Paro Paz, Alfredo Ribeiro De Freitas, Irineu Umberto Packer, Daniela Tambasco-Talhari, Luciana Correa De Almeida Regitano, Mauricio Mello Alencar Aug 2005

Sample Size And Power Calculations For Body Weight In Beef Cattle, Claudia Cristina Paro Paz, Alfredo Ribeiro De Freitas, Irineu Umberto Packer, Daniela Tambasco-Talhari, Luciana Correa De Almeida Regitano, Mauricio Mello Alencar

COBRA Preprint Series

Estimates of minimum sample sizes are calculated in order to test differences in rates of changes over time for longitudinal designs. In this study, body weight of crossbred beef cattle, considering 14 measurements on individuals, taken at birth, weaning (7 months of age) and monthly from 8 to 19 months of age, were analyzed by an usual mixed model for repeated measures. The number of individuals n required to detect significant differences (delta) between any two consecutive measurements on the individual, was obtained by a SAS program considering a t-variate normal distribution (t = 14), sample variance–covariance matrix among the …


Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen Aug 2005

Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

The causal effect of a treatment on an outcome is generally mediated by several intermediate variables. Estimation of the component of the causal effect of a treatment that is mediated by a given intermediate variable (the indirect effect of the treatment), and the component that is not mediated by that intermediate variable (the direct effect of the treatment) is often relevant to mechanistic understanding and to the design of clinical and public health interventions. Under the assumption of no-unmeasured confounders for treatment and the intermediate variable, Robins & Greenland (1992) define an individual direct effect as the counterfactual effect of …


Statistical Inference For Variable Importance, Mark J. Van Der Laan Aug 2005

Statistical Inference For Variable Importance, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Many statistical problems involve the learning of an importance/effect of a variable for predicting an outcome of interest based on observing a sample of n independent and identically distributed observations on a list of input variables and an outcome. For example, though prediction/machine learning is, in principle, concerned with learning the optimal unknown mapping from input variables to an outcome from the data, the typical reported output is a list of importance measures for each input variable. The typical approach in prediction has been to learn the unknown optimal predictor from the data and derive, for each of the input …


Computing The Total Sample Size When Group Sizes Are Not Fixed, Mithat Gonen Aug 2005

Computing The Total Sample Size When Group Sizes Are Not Fixed, Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

This article is concerned with computing the total sample size required for a two-sample comparison when the sizes of the two groups to be compared cannot be fixed in advance. This is frequently encountered when group membership depends on a variable which is observable only after the subject is enrolled to the study, such as a genetic or a biological marker. The most common way of circumventing this problem is assuming a fixed number for the prevalence of the condition that will determine the group membership and compute the required sample size conditionally. In this article this practice is formalized …


Survival Point Estimate Prediction In Matched And Non-Matched Case-Control Subsample Designed Studies, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore, Karla Kerlikowske Aug 2005

Survival Point Estimate Prediction In Matched And Non-Matched Case-Control Subsample Designed Studies, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore, Karla Kerlikowske

U.C. Berkeley Division of Biostatistics Working Paper Series

Providing information about the risk of disease and clinical factors that may increase or decrease a patient's risk of disease is standard medical practice. Although case-control studies can provide evidence of strong associations between diseases and risk factors, clinicians need to be able to communicate to patients the age-specific risks of disease over a defined time interval for a set of risk factors.

An estimate of absolute risk cannot be determined from case-control studies because cases are generally chosen from a population whose size is not known (necessary for calculation of absolute risk) and where duration of follow-up is not …


A User-Friendly Introduction To Link-Probit-Normal Models, Brian S. Caffo, Michael Griswold Aug 2005

A User-Friendly Introduction To Link-Probit-Normal Models, Brian S. Caffo, Michael Griswold

Johns Hopkins University, Dept. of Biostatistics Working Papers

Probit-normal models have attractive properties compared to logit-normal models. In particular, they allow for easy specification of marginal links of interest while permitting a conditional random effects structure. Moreover, programming fitting algorithms for probit-normal models can be trivial with the use of well-developed algorithms for approximating multivariate normal quantiles. In typical settings, the data cannot distinguish between probit and logit conditional link functions. Therefore, if marginal interpretations are desired, the default conditional link should be the most convenient one. We refer to models with a probit conditional link an arbitrary marginal link and a normal random effect distribution as link-probit-normal …


Semiparametric Inferences For Association With Semi-Competing Risks Data, Debashis Ghosh Aug 2005

Semiparametric Inferences For Association With Semi-Competing Risks Data, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

In many biomedical studies, it is of interest to assess dependence between bivariate failure time data. We focus here on a special type of such data, referred to as semi-competing risks data. In this article, we develop methods for making inferences regarding dependence of semi-competing risks data across strata of a discrete covariate Z. A class of rank statistics for testing constancy of association across strata are proposed; its asymptotic properties are also derived. We develop a novel resampling-based technique for calculating the variances of the proposed test statistics. In addition, we develop methods for combining test statistics for assessing …


Simultaneous Estimation Procedures And Multiple Testing: A Decision-Theoretic Framework, Debashis Ghosh Aug 2005

Simultaneous Estimation Procedures And Multiple Testing: A Decision-Theoretic Framework, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

There is recent tremendous interest in statistical methods regarding the false discovery rate (FDR). Two classes of literature on this topic exist. In the first, authors have proposed sequential testing procedures that control the false discovery rate. For the second, authors have studied the procedures involving FDR in a univariate mixture model setting. We consider a decision-theoretic approach to the assessment of FDR-based methods. In particular, we attempt to reconcile the current literature on false discovery rate procedures with more classical simultaneous estimation procedures. Formulation of the link will allow us to apply results from decision theory; we can then …


Shrunken P-Values For Assessing Differential Expression, With Applications To Genomic Data Analysis, Debashis Ghosh Aug 2005

Shrunken P-Values For Assessing Differential Expression, With Applications To Genomic Data Analysis, Debashis Ghosh

The University of Michigan Department of Biostatistics Working Paper Series

n many scientific problems involving high-throughput technology, inference must be made involving several hundreds or thousands of hypotheses. Recent attention has focused on how to address the multiple testing issue; much focus has been devoted towards use of the false discovery rate. In this article, we consider an alternative estimation procedure titled shrunken p-values for assessing differential expression (SPADE). The estimators are motivated by risk considerations from decision theory and lead to a completely new method for adjustment in the multiple testing problem. Some theoretical results are outlined. The proposed methodology is illustrated using simulation studies and with application to …


Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan Aug 2005

Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …


Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan Aug 2005

Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We present a cross-validated bagging scheme in the context of partitioning algorithms. To explore the benefits of the various bagging scheme, we compare via simulations the predictive ability of single Classification and Regression (CART) Tree with several previously suggested bagging schemes and with our proposed approach. Additionally, a variable importance measure is explained and illustrated.


Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit Jul 2005

Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …


When Should One Substract Background Fluorescence In Two Color Microarrays?, Robert B. Scharpf, Christine A. Iacobuzio-Donahue, Julie B. Sneddon, Giovanni Parmigiani Jul 2005

When Should One Substract Background Fluorescence In Two Color Microarrays?, Robert B. Scharpf, Christine A. Iacobuzio-Donahue, Julie B. Sneddon, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

Two color microarrays are a powerful tool for genomic analysis, but have noise components that make inferences regarding gene expression inefficient and potentially misleading. Background fluorescence,whether attributable to non-specific binding or other sources,is an important component of noise. The decision to subtract fluorescence surrounding spots of hybridization from spot fluorescence has been controversial, with no clear criteria for determining circumstances that may favor, or disfavor, background subtraction. While it is generally accepted that subtracting background reduces bias but increases variance in the estimates of the ratios of interest, no formal analysis of the bias-variance trade off of background subtraction has …


Does The Effect Of Micronutrient Supplementation On Neonatal Survival Vary With Respect To The Percentiles Of The Birth Weight Distribution?, Francesca Dominici, Scott L. Zeger, Giovanni Parmigiani, Joanne Katz, Parul Christian Jul 2005

Does The Effect Of Micronutrient Supplementation On Neonatal Survival Vary With Respect To The Percentiles Of The Birth Weight Distribution?, Francesca Dominici, Scott L. Zeger, Giovanni Parmigiani, Joanne Katz, Parul Christian

Johns Hopkins University, Dept. of Biostatistics Working Papers

Scientific Background: In developing countries, higher infant mortality is partially caused by poor maternal and fetal nutrition. Clinical trials of micronutrient supplementation are aimed at reducing the risk of infant mortality by increasing birth weight. Because infant mortality is greatest among the low birth weight infants (LBW) (less than or equal to 2500 grams), an effective intervention might be needed to increase birth weight among the smallest babies. Although it has been demonstrated that supplementation increases the birth weight in a trial conducted in Nepal, there is inconclusive evidence that the supplementation improves their survival. It has been hypothesized that …


Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang Jul 2005

Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang

UW Biostatistics Working Paper Series

Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.


G-Computation Estimation Of Nonparametric Causal Effects On Time-Dependent Mean Outcomes In Longitudinal Studies, Romain Neugebauer, Mark J. Van Der Laan Jul 2005

G-Computation Estimation Of Nonparametric Causal Effects On Time-Dependent Mean Outcomes In Longitudinal Studies, Romain Neugebauer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Two approaches to Causal Inference based on Marginal Structural Models (MSM) have been proposed. They provide different representations of causal effects with distinct causal parameters. Initially, a parametric MSM approach to Causal Inference was developed: it relies on correct specification of a parametric MSM. Recently, a new approach based on nonparametric MSM was introduced. This later approach does not require the assumption of a correctly specified MSM and thus is more realistic if one believes that correct specification of a parametric MSM is unlikely in practice. However, this approach was described only for investigating causal effects on mean outcomes collected …


A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan Jun 2005

A Note On The Construction Of Counterfactuals And The G-Computation Formula, Zhuo Yu, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Robins' causal inference theory assumes existence of treatment specific counterfactual variables so that the observed data augmented by the counterfactual data will satisfy a consistency and a randomization assumption. In this paper we provide an explicit function that maps the observed data into a counterfactual variable which satisfies the consistency and randomization assumptions. This offers a practically useful imputation method for counterfactuals. Gill & Robins [2001]'s construction of counterfactuals can be used as an imputation method in principle, but it is very hard to implement in practice. Robins [1987] shows that the counterfactual distribution can be identified from the observed …


Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen Jun 2005

Cross-Validated Bagged Learning, Mark J. Van Der Laan, Sandra E. Sinisi, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

Many applications aim to learn a high dimensional parameter of a data generating distribution based on a sample of independent and identically distributed observations. For example, the goal might be to estimate the conditional mean of an outcome given a list of input variables. In this prediction context, Breiman (1996a) introduced bootstrap aggregating (bagging) as a method to reduce the variance of a given estimator at little cost to bias. Bagging involves applying the estimator to multiple bootstrap samples, and averaging the result across bootstrap samples. In order to deal with the curse of dimensionality, typical practice has been to …


On Additive Regression Of Expectancy, Ying Qing Chen Jun 2005

On Additive Regression Of Expectancy, Ying Qing Chen

UW Biostatistics Working Paper Series

Regression models have been important tools to study the association between outcome variables and their covariates. The traditional linear regression models usually specify such an association by the expectations of the outcome variables as function of the covariates and some parameters. In reality, however, interests often focus on their expectancies characterized by the conditional means. In this article, a new class of additive regression models is proposed to model the expectancies. The model parameters carry practical implication, which may allow the models to be useful in applications such as treatment assessment, resource planning or short-term forecasting. Moreover, the new model …


Spatio-Temporal Point Processes: Methods And Applications, Peter J. Diggle Jun 2005

Spatio-Temporal Point Processes: Methods And Applications, Peter J. Diggle

Johns Hopkins University, Dept. of Biostatistics Working Papers

No abstract provided.


A Partial Likelihood For Spatio-Temporal Point Processes, Peter J. Diggle Jun 2005

A Partial Likelihood For Spatio-Temporal Point Processes, Peter J. Diggle

Johns Hopkins University, Dept. of Biostatistics Working Papers

Spatio-temporal point process data arise in many fields of application. An intuitively natural way to specify a model for a spatio-temporal point process is through its conditional intensity at location x and time t, given the history of the process up to time t. Typically, this results in an analytically intractable likelihood. Likelihood-based inference therefore relies on Monte Carlo methods which are computationally intensive and require careful tuning to each application. We propose a partial likelihood alternative which is computationally straightforward and can be applied routinely. We apply the method to data from the 2001 foot-and-mouth epidemic in the UK, …


Polydesigns And Causal Inference, Fan Li, Constantine E. Frangakis Jun 2005

Polydesigns And Causal Inference, Fan Li, Constantine E. Frangakis

Johns Hopkins University, Dept. of Biostatistics Working Papers

In an increasingly common class of studies, the goal is to evaluate causal effects of treatments that are only partially controlled by the investigator. In such studies there are two conflicting features: (1) a model on the full cohort design and data can identify the causal effects of interest, but can be sensitive to extreme regions of that design's data, where model specification can have more impact; and (2) models on a reduced design (i.e., a subset of the full data), e.g., conditional likelihood on matched subsets of data, can avoid such sensitivity, but do not generally identify the causal …


An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley Jun 2005

An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley

UW Biostatistics Working Paper Series

We consider data that are dependent, but where most small sets of observations are independent. By extending Bernstein's inequality we prove a strong law of law numbers and an empirical process central limit theorem under bracketing entropy conditions.


Model Choice In Time Series Studies Of Air Pollution And Mortality, Roger D. Peng, Francesca Dominici, Thomas A. Louis Jun 2005

Model Choice In Time Series Studies Of Air Pollution And Mortality, Roger D. Peng, Francesca Dominici, Thomas A. Louis

Johns Hopkins University, Dept. of Biostatistics Working Papers

Multi-city time series studies of particulate matter (PM) and mortality and morbidity have provided evidence that daily variation in air pollution levels is associated with daily variation in mortality counts. These findings served as key epidemiological evidence for the recent review of the United States National Ambient Air Quality Standards (NAAQS) for PM. As a result, methodological issues concerning time series analysis of the relation between air pollution and health have attracted the attention of the scientific community and critics have raised concerns about the adequacy of current model formulations. Time series data on pollution and mortality are generally analyzed …


A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe May 2005

A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe

UW Biostatistics Working Paper Series

In the field of medical diagnostic testing, the receiver operating characteristics(ROC) curve has long been used as a standard statistical tool to assess the accuracy of tests that yield continuous results. Although previous research in this area focused mostly on estimating the ROC curve, recently it has been recognized that the accuracy of a given test may fluctuate depending on certain factors, which motivates modelling covariate effects on the ROC curve. Comparing the corresponding ROC curves between two or more tests is a special case of covariate effect modelling. In this manuscript, we introduce a linear regression framework to model …


Attributable Risk Function In The Proportional Hazards Model, Ying Qing Chen, Chengcheng Hu, Yan Wang May 2005

Attributable Risk Function In The Proportional Hazards Model, Ying Qing Chen, Chengcheng Hu, Yan Wang

UW Biostatistics Working Paper Series

As an epidemiological parameter, the population attributable fraction is an important measure to quantify the public health attributable risk of an exposure to morbidity and mortality. In this article, we extend this parameter to the attributable fraction function in survival analysis of time-to-event outcomes, and further establish its estimation and inference procedures based on the widely used proportional hazards models. Numerical examples and simulations studies are presented to validate and demonstrate the proposed methods.


New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski May 2005

New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski

COBRA Preprint Series

As the field of functional genetics and genomics is beginning to mature, we become confronted with new challenges. The constant drop in price for sequencing and gene expression profiling as well as the increasing number of genetic and genomic variables that can be measured makes it feasible to address more complex questions. The success with rare diseases caused by single loci or genes has provided us with a proof-of-concept that new therapies can be developed based on functional genomics and genetics.

Common diseases, however, typically involve genetic epistasis, genomic pathways, and proteomic pattern. Moreover, to better understand the underlying biologi-cal …


Estimating Function Based Cross-Validation And Learning, Mark J. Van Der Laan, Daniel Rubin May 2005

Estimating Function Based Cross-Validation And Learning, Mark J. Van Der Laan, Daniel Rubin

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose that we observe a sample of independent and identically distributed realizations of a random variable. Given a model for the data generating distribution, assume that the parameter of interest can be characterized as the parameter value which makes the population mean of a possibly infinite dimensional estimating function equal to zero. Given a collection of candidate estimators of this parameter, and specification of the vector estimating function, we propose cross-validation criteria for selecting among these estimators. This cross-validation criteria is defined as the Euclidean norm of the empirical mean over the validation sample of the estimating function at the …