Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 481 - 510 of 1108

Full-Text Articles in Statistics and Probability

Causal Inference In Epidemiological Studies With Strong Confounding, Kelly L. Moore, Romain S. Neugebauer, Mark J. Van Der Laan, Ira B. Tager Oct 2009

Causal Inference In Epidemiological Studies With Strong Confounding, Kelly L. Moore, Romain S. Neugebauer, Mark J. Van Der Laan, Ira B. Tager

U.C. Berkeley Division of Biostatistics Working Paper Series

One of the identifiabilty assumptions of causal effects defined by marginal structural model (MSM) parameters is the experimental treatment assignment (ETA) assumption. Practical violations of this assumption frequently occur in data analysis, when certain exposures are rarely observed within some strata of the population. The inverse probability of treatment weighted (IPTW) estimator is particularly sensitive to violations of this assumption, however, we demonstrate that this is a problem for all estimators of causal effects. This is due to the fact that the ETA assumption is about information (or lack thereof) in the data. A new class of causal models, causal …


Lasagna Plots: A Saucy Alternative To Spaghetti Plots, Bruce Swihart, Brian Caffo, Bryan D. James, Matthew Strand, Brian S. Schwartz, Naresh M. Punjabi Oct 2009

Lasagna Plots: A Saucy Alternative To Spaghetti Plots, Bruce Swihart, Brian Caffo, Bryan D. James, Matthew Strand, Brian S. Schwartz, Naresh M. Punjabi

Johns Hopkins University, Dept. of Biostatistics Working Papers

Longitudinal repeated measures data has often been visualized with spaghetti plots for continuous out- comes. For large datasets, this often leads to over-plotting and consequential obscuring of trends in the data. This is primarily due to overlapping of trajectories. Here, we suggest a framework called lasagna plot ting that constrains the subject-specific trajectories to prevent overlapping and utilizes gradients of color to depict the outcome. Dynamic sorting and visualization is demonstrated as an exploratory data analysis tool. Supplemental material in the form of sample R code additional illustrated examples are available online.


Modeling Multilevel Sleep Transitional Data Via Poisson Log-Linear Multilevel Models, Bruce J. Swihart Oct 2009

Modeling Multilevel Sleep Transitional Data Via Poisson Log-Linear Multilevel Models, Bruce J. Swihart

COBRA Preprint Series

This paper proposes Poisson log-linear multilevel models to investigate population variability in sleep state transition rates. We specifically propose a Bayesian Poisson regression model that is more flexible, scalable to larger studies, and easily fit than other attempts in the literature. We further use hierarchical random effects to account for pairings of individuals and repeated measures within those individuals, as comparing diseased to non-diseased subjects while minimizing bias is of epidemiologic importance. We estimate essentially non-parametric piecewise constant hazards and smooth them, and allow for time varying covariates and segment of the night comparisons. The Bayesian Poisson regression is justified …


Quasi-Least Squares With Mixed Linear Correlation Structures, Jichun Xie, Justine Shults, Jon Peet, Dwight Stambolian, Mary F. Cotch Oct 2009

Quasi-Least Squares With Mixed Linear Correlation Structures, Jichun Xie, Justine Shults, Jon Peet, Dwight Stambolian, Mary F. Cotch

UPenn Biostatistics Working Papers

Quasi-least squares (QLS) is a two-stage computational approach for estimation of the correlation parameters in the framework of generalized estimating equations (GEE). We prove two general results for the class of mixed linear correlation structures: namely, that the stage one QLS estimate of the correlation parameter always exists and is feasible (yields a positive definite estimated correlation matrix) for any correlation structure, while the stage two estimator exists and is unique (and therefore consistent) with probability one, for the class of mixed linear correlation structures. Our general results justify the implementation of QLS for particular members of the class of …


Composite Likelihood Em Algorithm With Applications To Multivariate Hidden Markov Model , Xin Gao, Peter Xuekun Song Sep 2009

Composite Likelihood Em Algorithm With Applications To Multivariate Hidden Markov Model , Xin Gao, Peter Xuekun Song

COBRA Preprint Series

The method of composite likelihood is useful to deal with estimation and inference in parametric models with high-dimensional data, where the full likelihood approach renders to intractable computational complexity. We develop an extension of the EM algorithm in the framework of composite likelihood estimation in the presence of missing data or latent variables. We establish three key theoretical properties of the composite likelihood EM (CLEM) algorithm, including the ascent property, the algorithmic convergence and the convergence rate. The proposed method is applied to estimate the transition probabilities in multivariate hidden Markov model. Simulation studies are presented to demonstrate the empirical …


Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber Sep 2009

Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber

U.C. Berkeley Division of Biostatistics Working Paper Series

This is a compilation of current and past work on targeted maximum likelihood estimation. It features the original targeted maximum likelihood learning paper as well as chapters on super (machine) learning using cross validation, randomized controlled trials, realistic individualized treatment rules in observational studies, biomarker discovery, case-control studies, and time-to-event outcomes with censored data, among others. We hope this collection is helpful to the interested reader and stimulates additional research in this important area.


Robustness Of Semiparametric Efficiency In Nearly-Correct Models For Two-Phase Samples, Thomas Lumley Sep 2009

Robustness Of Semiparametric Efficiency In Nearly-Correct Models For Two-Phase Samples, Thomas Lumley

UW Biostatistics Working Paper Series

Augmented inverse-probability weighted (AIPW) estimators for incomplete-data models typically do not have full semiparametric efficiency, but do have model-robustness properties not shared by the efficient estimator. We examine the performance of efficient and AIPW estimators when the complete-data model is nearly correctly specified, in the sense that the misspecification is not reliably detectable from the data by any possible diagnostic or test. Asymptotic results for these nearly true models are obtained by representing them as sequences of misspecified models that are mutually contiguous with a correctly specified model. For some least favorable direction of model misspecification the bias in the …


Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan Sep 2009

Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

A nested case-control study is conducted within a well-defined cohort arising out of a population of interest. This design is often used in epidemiology to reduce the costs associated with collecting data on the full cohort; however, the case control sample within the cohort is a biased sample. Methods for analyzing case-control studies have largely focused on logistic regression models that provide conditional and not marginal causal estimates of the odds ratio. We previously developed a Case-Control Weighted Targeted Maximum Likelihood Estimation (TMLE) procedure for case-control study designs, which relies on the prevalence probability q0. We propose the use of …


Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry Sep 2009

Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

The DNA of most vertebrates is depleted in CpG dinucleotides; C followed by a G in the 5’ to 3’ direction. CpGs are the target for DNA methylation, a chemical modification of cytosine (C) heritable during cell division and the most well characterized epigenetic mechanism. The remaining CpGs tend to cluster in regions referred to as CpG islands (CGI). Knowing CGI locations is important because they mark functionally relevant epigenetic loci in development and disease. For various mammals, including human, a readily available and widely used list of CGI is available from the UCSC Genome Browser. This list was derived …


Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan Aug 2009

Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This paper provides a concise introduction to targeted maximum likelihood estimation (TMLE) of causal effect parameters. The interested analyst should gain sufficient understanding of TMLE from this introductory tutorial to be able to apply the method in practice. A program written in R is provided. This program implements a basic version of TMLE that can be used to estimate the effect of a binary point treatment on a continuous or binary outcome.


Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei Aug 2009

Comparing Risk Scoring Systems Beyond The Roc Paradigm In Survival Analysis, Hajime Uno, Lu Tian, Tianxi Cai, Isaac S. Kohane, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani Aug 2009

Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce combinatorial mixtures - a flexible class of models for inference on mixture distributions whose component have multidimensional parameters. The key idea is to allow each element of the component-specific parameter vectors to be shared by a subset of other components. This approach allows for mixtures that range from very flexible to very parsimonious, and unifies inference on component-specific parameters with inference on the number of components. We develop Bayesian inference and computation approaches for this class of distributions, and illustrate them in an application. This work was originally motivated by the analysis of cancer subtypes: in terms of …


Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel Aug 2009

Shrinkage Estimation Of Expression Fold Change As An Alternative To Testing Hypotheses Of Equivalent Expression, Zahra Montazeri, Corey M. Yanofsky, David R. Bickel

COBRA Preprint Series

Research on analyzing microarray data has focused on the problem of identifying differentially expressed genes to the neglect of the problem of how to integrate evidence that a gene is differentially expressed with information on the extent of its differential expression. Consequently, researchers currently prioritize genes for further study either on the basis of volcano plots or, more commonly, according to simple estimates of the fold change after filtering the genes with an arbitrary statistical significance threshold. While the subjective and informal nature of the former practice precludes quantification of its reliability, the latter practice is equivalent to using a …


Reliability Of The Model For Clustering Of Longitudinal Datasets Of Infant Mortality Rate In India, Ajay Kumar Bansal, S D. Sharma Jul 2009

Reliability Of The Model For Clustering Of Longitudinal Datasets Of Infant Mortality Rate In India, Ajay Kumar Bansal, S D. Sharma

COBRA Preprint Series

Because of the natural tendency of human beings and heavenly bodies to form groups, the technique of cluster analysis or segmentation analysis find its importance and applications in many fields of study. A model for clustering of time trends was proposed by authors whose beauty is that 2-way dimensions that is the horizontal flow of the trend and vertical distance of the trend from a common base are considered to obtain the natural clusters. In the present paper, the reliability of this model is studied in two steps namely (i) by repeating the analysis but using different interval distance measures …


The Effect Of Correlation In False Discovery Rate Estimation, Armin Schwartzman, Xihong Lin Jul 2009

The Effect Of Correlation In False Discovery Rate Estimation, Armin Schwartzman, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan Jun 2009

Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

For estimating regressions for repeated measures outcome data, a popular choice is the population average models estimated by generalized estimating equations (GEE). We review in this report the derivation of the robust inference (sandwich-type estimator of the standard error). In addition, we present formally how the approximation of a misspecified working population average model relates to the true model and in turn how to interpret the results of such a misspecified model.


A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry Jun 2009

A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …


A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry Jun 2009

A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …


A Spatio-Temporal Approach For Estimating Chronic Effects Of Air Pollution, Sonja Greven, Francesca Dominici, Scott L. Zeger Jun 2009

A Spatio-Temporal Approach For Estimating Chronic Effects Of Air Pollution, Sonja Greven, Francesca Dominici, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

Estimating the health risks associated with air pollution exposure is of great importance in public health. In air pollution epidemiology, two study designs have been used mainly. Time series studies estimate acute risk associated with short-term exposure. They compare day-to-day variation of pollution concentrations and mortality rates, and have been criticized for potential confounding by time-varying covariates. Cohort studies estimate chronic effects associated with long-term exposure. They compare long-term average pollution concentrations and time-to-death across cities, and have been criticized for potential confounding by individual risk factors or city-level characteristics.

We propose a new study design and a statistical model, …


Spatial Cluster Detection For Repeatedly Measured Outcomes While Accounting For Residential History, Andrea J. Cook, Diane Gold, Yi Li Jun 2009

Spatial Cluster Detection For Repeatedly Measured Outcomes While Accounting For Residential History, Andrea J. Cook, Diane Gold, Yi Li

Harvard University Biostatistics Working Paper Series

No abstract provided.


Marginalized Frailty Models For Multivariate Survival Data, Megan Othus, Yi Li Jun 2009

Marginalized Frailty Models For Multivariate Survival Data, Megan Othus, Yi Li

Harvard University Biostatistics Working Paper Series

No abstract provided.


Spatial Cluster Detection For Weighted Outcomes Using Cumulative Geographic Residuals, Andrea J. Cook, Yi Li, David Arterburn, Ram C. Tiwari Jun 2009

Spatial Cluster Detection For Weighted Outcomes Using Cumulative Geographic Residuals, Andrea J. Cook, Yi Li, David Arterburn, Ram C. Tiwari

Harvard University Biostatistics Working Paper Series

No abstract provided.


"Implementation Of Quasi-Least Squares With The R Package Qlspack", Jichun Xie, Justine Shults Jun 2009

"Implementation Of Quasi-Least Squares With The R Package Qlspack", Jichun Xie, Justine Shults

UPenn Biostatistics Working Papers

Quasi-least squares (QLS) is an alternative method for estimating the correlation parameters within the framework of generalized estimating equations (GEE) that has two main advantages over the moment estimates that are typically applied for GEE: (1) It guarantees a consistent estimate of the correlation parameter and a positive definite estimated correlation matrix, for several correlation structures; and (2) It allows for easier implementation of some correlation structures that have not yet been implemented in the framework of GEE. Furthermore, because QLS is a method in the framework of GEE, existing software can be employed within the QLS algorithm for estimation …


Simple, Defensible Sample Sizes Based On Cost Efficiency -- With Discussion And Rejoinder, Peter Bacchetti, Charles E. Mcculloch, Mark R. Segal, Richard Simon, Peter Muller, Gary L. Rosner, James A. Hanley, Stan Shapiro Jun 2009

Simple, Defensible Sample Sizes Based On Cost Efficiency -- With Discussion And Rejoinder, Peter Bacchetti, Charles E. Mcculloch, Mark R. Segal, Richard Simon, Peter Muller, Gary L. Rosner, James A. Hanley, Stan Shapiro

COBRA Preprint Series

The conventional approach of choosing sample size to provide 80% or greater power ignores the cost implications of different sample size choices. Costs, however, are often impossible for investigators and funders to ignore in actual practice. Here, we propose and justify a new approach for choosing sample size based on cost efficiency, the ratio of a study’s projected scientific and/or practical value to its total cost. By showing that a study’s projected value exhibits diminishing marginal returns as a function of increasing sample size for a wide variety of definitions of study value, we are able to develop two simple …


Nonparametric And Semiparametric Estimation Of The Three Way Receiver Operating Characteristic Surface, Jialiang Li, Xiao-Hua Zhou Jun 2009

Nonparametric And Semiparametric Estimation Of The Three Way Receiver Operating Characteristic Surface, Jialiang Li, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In many situations the diagnostic decision is not limited to a binary choice. Binary statistical tools such as receiver operating characteristic (ROC) curve and area under the ROC curve (AUC) need to be expanded to address three-category classification problem. Previous authors have suggest various ways to model the extension of AUC but not the ROC surface. Only simple parametric approaches are proposed for modeling the ROC measure under the assumption that test results all follow normal distributions. We study the estimation methods of three dimensional ROC surfaces with nonparametric and semiparametric estimators. Asymptotical results are provided as a basis for …


Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou Jun 2009

Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

For many medical conditions several treatment options may be available for treating patients. We consider evaluating markers based on a simple treatment selection policy that incorporates information on the patient's marker value exceeding a threshold. For example, colon cancer patients may be treated by surgery alone or surgery plus chemotherapy. The c-myc gene expression level may be used as a biomarker for treatment selection. Although traditional regression methods may assess the effect of the marker and treatment on outcomes, it is appealing to quantify more directly the potential impact on the population of using the marker to select treatment. A …


On The C-Statistics For Evaluating Overall Adequacy Of Risk Prediction Procedures With Censored Survival Data, Hajime Uno, Tianxi Cai, Michael J. Pencina, Ralph B. D'Agostino, L. J. Wei Jun 2009

On The C-Statistics For Evaluating Overall Adequacy Of Risk Prediction Procedures With Censored Survival Data, Hajime Uno, Tianxi Cai, Michael J. Pencina, Ralph B. D'Agostino, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell Jun 2009

A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

No abstract provided.


A Bayesian Shrinkage Model For Incomplete Longitudinal Binary Data With Application To The Breast Cancer Prevention Trial, C. Wang, M.J. Daniels, Daniel O. Scharfstein, S. Land May 2009

A Bayesian Shrinkage Model For Incomplete Longitudinal Binary Data With Application To The Breast Cancer Prevention Trial, C. Wang, M.J. Daniels, Daniel O. Scharfstein, S. Land

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider inference in randomized studies, in which repeatedly measured outcomes may be informatively missing due to drop out. In this setting, it is well known that full data estimands are not identified unless unverified assumptions are imposed. We assume a non-future dependence model for the drop-out mechanism and posit an exponential tilt model that links non-identifiable and identifiable distributions. This model is indexed by non-identified parameters, which are assumed to have an informative prior distribution, elicited from subject-matter experts. Under this model, full data estimands are shown to be expressed as functionals of the distribution of the observed data. …


Estimating Subject-Specific Dependent Competing Risk Profile With Censored Event Time Observations, Yi Li, Lu Tian, L. J. Wei May 2009

Estimating Subject-Specific Dependent Competing Risk Profile With Censored Event Time Observations, Yi Li, Lu Tian, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.