Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 391 - 420 of 1108

Full-Text Articles in Statistics and Probability

Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan Mar 2011

Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan

Johns Hopkins University, Dept. of Biostatistics Working Papers

We present a brief overview of targeted maximum likelihood for estimating the causal effect of a single time point treatment and of a two time point treatment. We focus on simple examples demonstrating how to apply the methodology developed in (van der Laan and Rubin, 2006; Moore and van der Laan, 2007; van der Laan, 2010a,b). We include R code for the single time point case.


Bate Curve In Assessment Of Clinical Utility Of Predictive Biomarkers, Xiao-Hua Zhou, Yunbei Ma Feb 2011

Bate Curve In Assessment Of Clinical Utility Of Predictive Biomarkers, Xiao-Hua Zhou, Yunbei Ma

UW Biostatistics Working Paper Series

In this paper, for time-to-event data, we propose a new statistical framework for casual inference in evaluating clinical utility of predictive biomarkers and in selecting an optimal treatment for a particular patient. This new casual framework is based on a new concept, called Biomarker Adjusted Treatment Effect (BATE) curve, which can be used to represent the clinical utility of a predictive biomarker and select an optimal treatment for one particular patient. We then propose semi-parametric methods for estimating the BATE curves of biomarkers and establish asymptotic results of the proposed estimators for the BATE curves. We also conduct extensive simulation …


Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan Feb 2011

Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation (TMLE) presents an approach for construction of an efficient double-robust semi-parametric substitution estimator of a target feature of the data generating distribution, such as a statistical association measure or a causal effect parameter. tmle is a recently developed R package that implements TMLE for estimation of the effect of a binary treatment at a single point in time on an outcome of interest, controlling for user supplied covariates: the additive treatment effect, the relative risk, the odds ratio. The package allows outcome data with missingness, and experimental units that contribute repeated records of the point-treatment data …


Causal Inference Under Multiple Versions Of Treatment, Tyler J. Vanderweele, Miguel A. Hernan Feb 2011

Causal Inference Under Multiple Versions Of Treatment, Tyler J. Vanderweele, Miguel A. Hernan

COBRA Preprint Series

In this article we discuss the no-multiple-versions-of-treatment assumption and extend the potential outcomes framework to accommodate causal inference under violations of this assumption. A variety of examples are discussed in which the assumption may be violated. Identification results are provided for the overall treatment effect and the effect of treatment on the treated when multiple versions of treatment are present and also for the causal effect comparing a version of one treatment to some other version of the same or a different treatment. Further identification and interpretative results are given for cases in which a treatment variable is dichotomized to …


Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett Feb 2011

Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett

Johns Hopkins University, Dept. of Biostatistics Working Papers

In this manuscript, we use a two-stage decomposition for the analysis of func- tional magnetic resonance imaging (fMRI). In the first stage, spatial independent component analysis is applied to the group fMRI data to obtain common brain networks (spatial maps) and subject-specific mixing matrices (time courses). In the second stage, functional principal component analysis is utilized to decompose the mixing matrices into population- level eigenvectors and subject-specific loadings. Inference is performed using permutation-based exact conditional logistic regression for matched pairs data. Simulation studies suggest the ability of the decomposition methods to recover population brain networks and the major direction of …


Non-Homogeneous Markov Process Models With Incomplete Observations: Application To A Dementia Disease Study, Xiao-Hua Zhou, Baojiang Chen Jan 2011

Non-Homogeneous Markov Process Models With Incomplete Observations: Application To A Dementia Disease Study, Xiao-Hua Zhou, Baojiang Chen

UW Biostatistics Working Paper Series

Identifying risk factors for transition rates among normal cognition, mildly cognitive impairment, dementia and death in an Alzheimer's disease study is very important. It is known that transition rates among these states are strongly time dependent. While Markov process models are often used to describe these disease progressions, the literature mainly focuses on time homogeneous processes, and limited tools are available for dealing with non-homogeneity. Further, patients may choose when they want to visit the clinics, which creates informative observations. In this paper, we develop methods to deal with non-homogeneous Markov processes through time scale transformation when observation times are …


Doubly Robust Estimates For Binary Longitudinal Data Analysis With Missing Response And Missing Covariates, Baojiang Chen, Xiao-Hua Zhou Jan 2011

Doubly Robust Estimates For Binary Longitudinal Data Analysis With Missing Response And Missing Covariates, Baojiang Chen, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Longitudinal studies often feature incomplete response and covariate data. Likelihood-based methods such as the EM algorithm give consistent estimators for model parameters when data are missing at random provided that the response model and the missing covariate model are correctly specified; but we do not need to specify the missing data mechanism. An alternative method is the weighted estimating equation which gives consistent estimators if the missing data and response models are correctly specified; but we do not need to specify the distribution of the covariates that have missing values. In this paper we develop a doubly robust estimation method …


Semiparametric Estimation Of The Covariate-Specific Roc Curve In Presence Of Ignorable Verification Bias, Danping Liu, Xiao-Hua Zhou Jan 2011

Semiparametric Estimation Of The Covariate-Specific Roc Curve In Presence Of Ignorable Verification Bias, Danping Liu, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Covariate-specific ROC curves are often used to evaluate the classification accuracy of a medical diagnostic test or a biomarker, when the accuracy of the test is associated with certain covariates. In many large-scale screening tests, the gold standard is subject to missingness due to high cost or harmfulness to the patient. In this paper, we propose a semiparametric estimation method for the covariate-specific ROC curves with a partial missing gold standard. A location-scale model is constructed for the test result to model the covariates' effect, but the residual distributions are left unspecified. Thus the baseline and link functions of the …


Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou Jan 2011

Evaluating Markers For Treatment Selection Based On Survival Time, Xiao Song, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

For many medical conditions there are several treatment options available to patients. We consider evaluating markers based on a simple treatment selection policy that incorporates information on the patient's marker value exceeding a threshold. Although traditional regression methods may assess the effect of the marker and treatment on outcomes, it is appealing to quantify more directly the potential impact on the population of using the marker to select treatment. A useful tool is the selection impact (SI) curve proposed by Song and Pepe (2004, \textit{Biometrics} \textbf{60}, 874--883) for binary outcomes. However, this approach does not deal with continuous outcomes, nor …


A Flexible Spatio-Temporal Model For Air Pollution: Allowing For Spatio-Temporal Covariates, Johan Lindstrom, Adam A. Szpiro, Paul D. Sampson, Lianne Sheppard, Assaf Oron, Mark Richards, Tim Larson Jan 2011

A Flexible Spatio-Temporal Model For Air Pollution: Allowing For Spatio-Temporal Covariates, Johan Lindstrom, Adam A. Szpiro, Paul D. Sampson, Lianne Sheppard, Assaf Oron, Mark Richards, Tim Larson

UW Biostatistics Working Paper Series

Given the increasing interest in the association between exposure to air pollution and adverse health outcomes, the development of models that provide accurate spatio-temporal predictions of air pollution concentrations at small spatial scales is of great importance when assessing potential health effects of air pollution. The methodology presented here has been developed as part of the Multi-Ethnic Study of Atherosclerosis and Air Pollution (MESA Air), a prospective cohort study funded by the US EPA to investigate the relationship between chronic exposure to air pollution and cardiovascular disease. We present a spatio-temporal framework that models and predicts ambient air pollution by …


Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu Jan 2011

Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We establish a fundamental equivalence between singular value decomposition (SVD) and functional principal components analysis (FPCA) models. The constructive relationship allows to deploy the numerical efficiency of SVD to fully estimate the components of FPCA, even for extremely high-dimensional functional objects, such as brain images. As an example, a functional mixed effect model is fitted to high-resolution morphometric (RAVENS) images. The main directions of morphometric variation in brain volumes are identified and discussed.


A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos Jan 2011

A Generalized Approach For Testing The Association Of A Set Of Predictors With An Outcome: A Gene Based Test, Benjamin A. Goldstein, Alan E. Hubbard, Lisa F. Barcellos

U.C. Berkeley Division of Biostatistics Working Paper Series

In many analyses, one has data on one level but desires to draw inference on another level. For example, in genetic association studies, one observes units of DNA referred to as SNPs, but wants to determine whether genes that are comprised of SNPs are associated with disease. While there are some available approaches for addressing this issue, they usually involve making parametric assumptions and are not easily generalizable. A statistical test is proposed for testing the association of a set of variables with an outcome of interest. No assumptions are made about the functional form relating the variables to the …


Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan Dec 2010

Oracle And Multiple Robustness Properties Of Survey Calibration Estimator In Missing Response Problem, Kwun Chuen Gary Chan

UW Biostatistics Working Paper Series

In the presence of missing response, reweighting the complete case subsample by the inverse of nonmissing probability is both intuitive and easy to implement. However, inverse probability weighting is not efficient in general and is not robust against misspecification of the missing probability model. Calibration was developed by survey statisticians for improving efficiency of inverse probability weighting estimators when population totals of auxiliary variables are known and when inclusion probability is known by design. In missing data problem we can calibrate auxiliary variables in the complete case subsample to the full sample. However, the inclusion probability is unknown in general …


Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan Dec 2010

Modification And Improvement Of Empirical Likelihood For Missing Response Problem, Kwun Chuen Gary Chan

UW Biostatistics Working Paper Series

An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …


Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan Dec 2010

Modification And Improvement Of Empirical Liklihood For Missing Response Problem, Gary Chan

UW Biostatistics Working Paper Series

An empirical likelihood (EL) estimator was proposed by Qin and Zhang (2007) for a missing response problem under a missing at random assumption. They showed by simulation studies that the finite sample performance of EL estimator is better than some existing estimators. However, the empirical likelihood estimator does not have a uniformly smaller asymptotic variance than other estimators in general. We consider several modifications to the empirical likelihood estimator and show that the proposed estimator dominates the empirical likelihood estimator and several other existing estimators in terms of asymptotic efficiencies. The proposed estimator also attains the minimum asymptotic variance among …


Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel Dec 2010

Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel

COBRA Preprint Series

In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …


Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley Dec 2010

Efficient Measurement Error Correction With Spatially Misaligned Data, Adam A. Szpiro, Lianne Sheppard, Thomas Lumley

UW Biostatistics Working Paper Series

Association studies in environmental statistics often involve exposure and outcome data that are misaligned in space. A common strategy is to employ a spatial model such as universal kriging to predict exposures at locations with outcome data and then estimate a regression parameter of interest using the predicted exposures. This results in measurement error because the predicted exposures do not correspond exactly to the true values. We characterize the measurement error by decomposing it into Berkson-like and classical-like components. One correction approach is the parametric bootstrap, which is effective but computationally intensive since it requires solving a nonlinear optimization problem …


Predicting Treatment Efficacy Via Quantitative Mri: A Bayesian Joint Model, Jincao Wu, Tim Johnson Dec 2010

Predicting Treatment Efficacy Via Quantitative Mri: A Bayesian Joint Model, Jincao Wu, Tim Johnson

The University of Michigan Department of Biostatistics Working Paper Series

The prognosis for patients with high-grade gliomas is poor, with a median survival of one year. Treatment efficacy assessment is typically unavailable until 5{6 months post diagnosis. Investigators hypothesize that quantitative MRI (qMRI) can assess treatment efficacy three weeks after therapy starts, thereby allowing salvage treatments to begin earlier. The purpose of this work is to build a predictive model of treatment efficacy using qMRI data and to assess its performance. The outcome is one-year survival status. We propose a joint, two-stage Bayesian model. In stage I, we smooth the image data with a multivariate spatio-temporal pairwise dierence prior. We …


Asymptotic Theory For Cross-Validated Targeted Maximum Likelihood Estimation, Wenjing Zheng, Mark J. Van Der Laan Nov 2010

Asymptotic Theory For Cross-Validated Targeted Maximum Likelihood Estimation, Wenjing Zheng, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider a targeted maximum likelihood estimator of a path-wise differentiable parameter of the data generating distribution in a semi-parametric model based on observing n independent and identically distributed observations. The targeted maximum likelihood estimator (TMLE) uses V-fold sample splitting for the initial estimator in order to make the TMLE maximally robust in its bias reduction step. We prove a general theorem that states asymptotic efficiency (and thereby regularity) of the targeted maximum likelihood estimator when the initial estimator is consistent and a second order term converges to zero in probability at a rate faster than the square root of …


A Bayesian Shared Component Model For Genetic Association Studies, Juan J. Abellan, Carlos Abellan, Juan R. Gonzalez Nov 2010

A Bayesian Shared Component Model For Genetic Association Studies, Juan J. Abellan, Carlos Abellan, Juan R. Gonzalez

COBRA Preprint Series

We present a novel approach to address genome association studies between single nucleotide polymorphisms (SNPs) and disease. We propose a Bayesian shared component model to tease out the genotype information that is common to cases and controls from the one that is specific to cases only. This allows to detect the SNPs that show the strongest association with the disease. The model can be applied to case-control studies with more than one disease. In fact, we illustrate the use of this model with a dataset of 23,418 SNPs from a case-control study by The Welcome Trust Case Control Consortium (2007) …


Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel Nov 2010

Minimum Description Length And Empirical Bayes Methods Of Identifying Snps Associated With Disease, Ye Yang, David R. Bickel

COBRA Preprint Series

The goal of determining which of hundreds of thousands of SNPs are associated with disease poses one of the most challenging multiple testing problems. Using the empirical Bayes approach, the local false discovery rate (LFDR) estimated using popular semiparametric models has enjoyed success in simultaneous inference. However, the estimated LFDR can be biased because the semiparametric approach tends to overestimate the proportion of the non-associated single nucleotide polymorphisms (SNPs). One of the negative consequences is that, like conventional p-values, such LFDR estimates cannot quantify the amount of information in the data that favors the null hypothesis of no disease-association.

We …


Observational Study And Individualized Antiretroviral Therapy Initiation Rules For Reducing Cancer Incidence In Hiv-Infected Patients, Romain Neugebauer, Michael J. Silverberg, Mark J. Van Der Laan Nov 2010

Observational Study And Individualized Antiretroviral Therapy Initiation Rules For Reducing Cancer Incidence In Hiv-Infected Patients, Romain Neugebauer, Michael J. Silverberg, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted Maximum Likelihood Learning (TMLL) has been proposed as a general estimation methodology that can, in particular, be applied to draw causal inferences based on marginal structural modeling with observational data using either a point treatment approach (all confounders are assumed not to be affected by the exposure(s) of interest) or a longitudinal data approach (some confounders may be affected by one of the exposures of interest). While formal development of TMLL has included road maps for applications in longitudinal data approaches, real-life implementations have been restricted to studies based on a point treatment approach. In this article, we illustrate …


Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano Nov 2010

Improving The Power Of Chronic Disease Surveillance By Incorporating Residential History, Justin Manjourides, Marcello Pagano

Harvard University Biostatistics Working Paper Series

No abstract provided.


Gains In Power From Structured Two-Sample Tests Of Means On Graphs, Laurent Jacob, Pierre Neuvial, Sandrine Dudoit Oct 2010

Gains In Power From Structured Two-Sample Tests Of Means On Graphs, Laurent Jacob, Pierre Neuvial, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider multivariate two-sample tests of means, where the location shift between the two populations is expected to be related to a known graph structure. An important application of such tests is the detection of differentially expressed genes between two patient populations, as shifts in expression levels are expected to be coherent with the structure of graphs reflecting gene properties such as biological process, molecular function, regulation, or metabolism. For a fixed graph of interest, we demonstrate that accounting for graph structure can yield more powerful tests under the assumption of smooth distribution shift on the graph. We also investigate …


Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert Oct 2010

Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose a novel class of models for functional data exhibiting skewness or other shape characteristics that vary with spatial or temporal location. We use copulas so that the marginal distributions and the dependence structure can be modeled independently. Dependence is modeled with a Gaussian or t-copula, so that there is an underlying latent Gaussian process. We model the marginal distributions using the skew t family. The mean, variance, and shape parameters are modeled nonparametrically as functions of location. A computationally tractable inferential framework for estimating heterogeneous asymmetric or heavy-tailed marginal distributions is introduced. This framework provides a new set …


Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov Oct 2010

Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov

Johns Hopkins University, Dept. of Biostatistics Working Papers

Images, often stored in multidimensional arrays are fast becoming ubiquitous in medical and public health research. Analyzing populations of images is a statistical problem that raises a host of daunting challenges. The most severe challenge is that data sets incorporating images recorded for hundreds or thousands of subjects at multiple visits are massive. We introduce the population value decomposition (PVD), a general method for simultaneous dimensionality reduction of large populations of massive images. We show how PVD can seamlessly be incorporated into statistical modeling and lead to a new, transparent and fast inferential framework. Our methodology was motivated by and …


Multilevel Functional Principal Component Analysis For High-Dimensional Data, Vadim Zipunnikov, Brian Caffo, Ciprian Crainiceanu, David M. Yousem, Christos Davatzikos, Brian S. Schwartz Oct 2010

Multilevel Functional Principal Component Analysis For High-Dimensional Data, Vadim Zipunnikov, Brian Caffo, Ciprian Crainiceanu, David M. Yousem, Christos Davatzikos, Brian S. Schwartz

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose fast and scalable statistical methods for the analysis of hundreds or thousands of high dimensional vectors observed at multiple visits. The proposed inferential methods avoid the difficult task of loading the entire data set at once in the computer memory and use sequential access to data. This allows deployment of our methodology on low-resource computers where computations can be done in minutes on extremely large data sets. Our methods are motivated by and applied to a study where hundreds of subjects were scanned using Magnetic Resonance Imaging (MRI) at two visits roughly five years apart. The original data …


Targeted Bayesian Learning, Ivan Diaz Munoz, Alan E. Hubbard, Mark J. Van Der Laan Oct 2010

Targeted Bayesian Learning, Ivan Diaz Munoz, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation (van der Laan & Rubin 2006) is a loss-based semi-parametric estimation method that yields a substitution estimator of a target parameter of the probability distribution of the data that solves the efficient influence curve estimating equation, and thereby yields a double robust locally efficient estimator of the parameter of interest, under regularity conditions. The Bayesian paradigm is concerned with including the researcher’s prior uncertainty about the parameter through a prior distribution, which combined with the likelihood yields a posterior distribution for the parameter that reflects the researcher’s posterior uncertainty. In this paper, we present a way …


Estimating Temporal Associations In Electrocorticographic (Ecog) Time Series With First Order Pruning, Haley Hedlin, Dana Boatman, Brian Caffo Sep 2010

Estimating Temporal Associations In Electrocorticographic (Ecog) Time Series With First Order Pruning, Haley Hedlin, Dana Boatman, Brian Caffo

Johns Hopkins University, Dept. of Biostatistics Working Papers

Granger causality (GC) is a statistical technique used to estimate temporal associations in multivariate time series. Many applications and extensions of GC have been proposed since its formulation by Granger in 1969. Here we control for potentially mediating or confounding associations between time series in the context of event-related electrocorticographic (ECoG) time series. A pruning approach to remove spurious connections and simultaneously reduce the required number of estimations to fit the effective connectivity graph is proposed. Additionally, we consider the potential of adjusted GC applied to independent components as a method to explore temporal relationships between underlying source signals. Both …


Landmark Prediction Of Survival, Layla Parast, Tianxi Cai Sep 2010

Landmark Prediction Of Survival, Layla Parast, Tianxi Cai

Harvard University Biostatistics Working Paper Series

No abstract provided.