Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

U.C. Berkeley Division of Biostatistics Working Paper Series

Discipline
Keyword
Publication Year

Articles 61 - 90 of 118

Full-Text Articles in Biostatistics

Population Intervention Causal Effects Based On Stochastic Interventions, Ivan Diaz Munoz, Mark J. Van Der Laan Aug 2011

Population Intervention Causal Effects Based On Stochastic Interventions, Ivan Diaz Munoz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Estimating the causal effect of an intervention on a population typically involves defining parameters in a nonparametric structural equation model (Pearl, 2000, NPSEM) in which the treatment or exposure is deter- ministically assigned in a static or dynamic way. We define a new causal parameter that takes into account the fact that intervention policies can result in stochastically assigned exposures. The statistical parameter that identifies the causal parameter of interest is established. Inverse probability of treatment weighting (IPTW), augmented IPTW (A-IPTW), and targeted maximum likelihood estimators (TMLE) are developed. A simulation study is performed to demonstrate the properties of these …


Targeted Maximum Likelihood Estimation Of Natural Direct Effect, Wenjing Zheng, Mark J. Van Der Laan Jul 2011

Targeted Maximum Likelihood Estimation Of Natural Direct Effect, Wenjing Zheng, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In many causal inference problems, one is interested in the direct causal effect of an exposure on an outcome of interest that is not mediated by certain intermediate variables. Robins and Greenland (1992) and Pearl (2000) formalized the definition of two types of direct effects (natural and controlled) under the counterfactual framework. Since then, identifiability conditions for these effects have been studied extensively. By contrast, considerably fewer efforts have been invested in the estimation problem of the natural direct effect. In this article, we propose a semiparametric efficient, multiply robust estimator for the natural direct effect of a binary treatment …


Targeted Minimum Loss Based Estimation Based On Directly Solving The Efficient Influence Curve Equation, Paul Chaffee, Mark J. Van Der Laan Jul 2011

Targeted Minimum Loss Based Estimation Based On Directly Solving The Efficient Influence Curve Equation, Paul Chaffee, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Applying targeted maximum likelihood estimation to longitudinal data can be computationally intensive. As the number of time points and/or number of intermediate factors grows, the computation resources consumed by these algorithms likewise increases. Different TMLE algorithms have different computational speeds and implementation challenges; there may also be efficiency differences of the corresponding estimators. The algorithm we describe here proceeds by solving the empirical efficient influence curve equation directly using numerical computation methods, rather than indirectly (by solving a score equation), which is the usual route. We believe that this estimator is the simplest of the TMLE procedures to implement in …


Targeted Methods For Finding Quantitative Trait Loci, Hui Wang, Sherri Rose, Mark J. Van Der Laan Jul 2011

Targeted Methods For Finding Quantitative Trait Loci, Hui Wang, Sherri Rose, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Conventional genetic mapping methods typically assume parametric models with Gaussian errors, and obtain parameter estimates through maximum likelihood estimation. We propose a general semiparametric model to map quantitative trait loci (QTL) in experimental crosses. In contrast with widely-used interval mapping (IM) derived methods, our model requires fewer assumptions and also accommodates various machine learning algorithms. Estimation using both targeted maximum likelihood and collaborative targeted maximum likelihood methods is compared to a composite interval mapping (CIM) approach. We demonstrate with simulations and real data analyses that, on average, our semiparametric targeted learning approach produces less biased QTL effect estimates than those …


Targeted Maximum Likelihood Estimation Of Conditional Relative Risk In A Semi-Parametric Regression Model, Cathy Tuglus, Kristin E. Porter, Mark J. Van Der Laan Jun 2011

Targeted Maximum Likelihood Estimation Of Conditional Relative Risk In A Semi-Parametric Regression Model, Cathy Tuglus, Kristin E. Porter, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The conditional relative risk is an important measure in medical and epidemiological studies when the outcome of interest is binary (i.e. disease vs. no disease). When the outcome is common, estimation of conditional relative risk and related parameters can be problematic, especially when the exposure or covariates are continuous. We propose a new estimation procedure based on targeted maximum likelihood methodology that targets the parameters relating to the conditional relative risk for common outcomes under a log-linear, or multiplicative, semi-parametric model. In this paper, we present three possible targeted maximum likelihood estimators for relative risk parameters implied by such a …


Super Learner Based Conditional Density Estimation With Application To Marginal Structural Models, Ivan Diaz Munoz, Mark J. Van Der Laan Jun 2011

Super Learner Based Conditional Density Estimation With Application To Marginal Structural Models, Ivan Diaz Munoz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In this paper we present a histogram-like estimator of a conditional density that uses super learner crossvalidation to estimate the histogram probabilities, as well as the optimal number and position of the bins. This estimator is an alternative to kernel density estimators when the dimension of the problem is large. We demonstrate its applicability to estimation of Marginal Structural Model (MSM) parameters in which an initial estimator of the treatment %mechanism is needed. MSM estimation based on the proposed density estimator results in less biased estimates, when compared to estimates based on a misspecified parametric model.


A General Implementation Of Tmle For Longitudinal Data Applied To Causal Inference In Survival Analysis, Ori M. Stitelman, Victor De Gruttola, Mark J. Van Der Laan Apr 2011

A General Implementation Of Tmle For Longitudinal Data Applied To Causal Inference In Survival Analysis, Ori M. Stitelman, Victor De Gruttola, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In many randomized controlled trials the outcome of interest is a time to event, and one measures on each subject baseline covariates and time-dependent covariates until the subject either drops-out, the time to event is observed, or the end of study is reached. The goal of such a study is to assess the causal effect of the treatment on the survival curve. Standard methods (e.g., Kaplan-Meier estimator, Cox-proportional hazards) ignore the available baseline and time-dependent covariates, and are therefore biased if the drop-out is affected by these covariates, and are always inefficient. We present a targeted maximum likelihood estimator of …


Targeted Minimum Loss Based Estimator That Outperforms A Given Estimator, Susan Gruber, Mark J. Van Der Laan Apr 2011

Targeted Minimum Loss Based Estimator That Outperforms A Given Estimator, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted minimum loss based estimation (TMLE) provides a template for the construction of semiparametric locally efficient double robust substitution estimators of the target parameter of the data generating distribution in a semiparametric censored data or causal inference model (van der Laan and Rubin (2006),van der Laan (2008), van der Laan and Rose (2011)). In this article we demonstrate how to construct a TMLE that also satisfies the property that it is at least as efficient as a user supplied asymptotically linear estimator. For the sake of illustration we focus on estimation of the additive average causal effect of a point …


The Relative Performance Of Targeted Maximum Likelihood Estimators, Kristin E. Porter, Susan Gruber, Mark J. Van Der Laan, Jasjeet S. Sekhon Apr 2011

The Relative Performance Of Targeted Maximum Likelihood Estimators, Kristin E. Porter, Susan Gruber, Mark J. Van Der Laan, Jasjeet S. Sekhon

U.C. Berkeley Division of Biostatistics Working Paper Series

There is an active debate in the literature on censored data about the relative performance of model based maximum likelihood estimators, IPCW-estimators, and a variety of double robust semiparametric efficient estimators. Kang and Schafer (2007) demonstrate the fragility of double robust and IPCW-estimators in a simulation study with positivity violations. They focus on a simple missing data problem with covariates where one desires to estimate the mean of an outcome that is subject to missingness. Responses by Robins et al. (2007), Tsiatis and Davidian (2007), Tan (2007a) and Ridgeway and McCaffrey (2007) further explore the challenges faced by double robust …


Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan Apr 2011

Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This article is devoted to the construction and asymptotic study of adaptive group sequential covariate-adjusted randomized clinical trials analyzed through the prism of the semiparametric methodology of targeted maximum likelihood estimation (TMLE). We show how to build, as the data accrue group-sequentially, a sampling design which targets a user-supplied optimal design. We also show how to carry out a sound TMLE statistical inference based on such an adaptive sampling scheme (therefore extending some results known in the i.i.d setting only so far), and how group-sequential testing applies on top of it. The procedure is robust (i.e., consistent even if the …


Threshold Regression Models Adapted To Case-Control Studies, And The Risk Of Lung Cancer Due To Occupational Exposure To Asbestos In France, Antoine Chambaz, Dominique Choudat, Catherine Huber, Jean-Claude Pairon, Mark J. Van Der Laan Mar 2011

Threshold Regression Models Adapted To Case-Control Studies, And The Risk Of Lung Cancer Due To Occupational Exposure To Asbestos In France, Antoine Chambaz, Dominique Choudat, Catherine Huber, Jean-Claude Pairon, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Asbestos has been known for many years as a powerful carcinogen. Our purpose is quantify the relationship between an occupational exposure to asbestos and an increase of the risk of lung cancer. Furthermore, we wish to tackle the very delicate question of the evaluation, in subjects suffering from a lung cancer, of how much the amount of exposure to asbestos explains the occurrence of the cancer. For this purpose, we rely on a recent French case-control study. We build a large collection of threshold regression models, data-adaptively select a better model in it by multi-fold likelihood-based cross-validation, then fit the …


Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan Feb 2011

Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation (TMLE) presents an approach for construction of an efficient double-robust semi-parametric substitution estimator of a target feature of the data generating distribution, such as a statistical association measure or a causal effect parameter. tmle is a recently developed R package that implements TMLE for estimation of the effect of a binary treatment at a single point in time on an outcome of interest, controlling for user supplied covariates: the additive treatment effect, the relative risk, the odds ratio. The package allows outcome data with missingness, and experimental units that contribute repeated records of the point-treatment data …


Asymptotic Theory For Cross-Validated Targeted Maximum Likelihood Estimation, Wenjing Zheng, Mark J. Van Der Laan Nov 2010

Asymptotic Theory For Cross-Validated Targeted Maximum Likelihood Estimation, Wenjing Zheng, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We consider a targeted maximum likelihood estimator of a path-wise differentiable parameter of the data generating distribution in a semi-parametric model based on observing n independent and identically distributed observations. The targeted maximum likelihood estimator (TMLE) uses V-fold sample splitting for the initial estimator in order to make the TMLE maximally robust in its bias reduction step. We prove a general theorem that states asymptotic efficiency (and thereby regularity) of the targeted maximum likelihood estimator when the initial estimator is consistent and a second order term converges to zero in probability at a rate faster than the square root of …


Observational Study And Individualized Antiretroviral Therapy Initiation Rules For Reducing Cancer Incidence In Hiv-Infected Patients, Romain Neugebauer, Michael J. Silverberg, Mark J. Van Der Laan Nov 2010

Observational Study And Individualized Antiretroviral Therapy Initiation Rules For Reducing Cancer Incidence In Hiv-Infected Patients, Romain Neugebauer, Michael J. Silverberg, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted Maximum Likelihood Learning (TMLL) has been proposed as a general estimation methodology that can, in particular, be applied to draw causal inferences based on marginal structural modeling with observational data using either a point treatment approach (all confounders are assumed not to be affected by the exposure(s) of interest) or a longitudinal data approach (some confounders may be affected by one of the exposures of interest). While formal development of TMLL has included road maps for applications in longitudinal data approaches, real-life implementations have been restricted to studies based on a point treatment approach. In this article, we illustrate …


Targeted Bayesian Learning, Ivan Diaz Munoz, Alan E. Hubbard, Mark J. Van Der Laan Oct 2010

Targeted Bayesian Learning, Ivan Diaz Munoz, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation (van der Laan & Rubin 2006) is a loss-based semi-parametric estimation method that yields a substitution estimator of a target parameter of the probability distribution of the data that solves the efficient influence curve estimating equation, and thereby yields a double robust locally efficient estimator of the parameter of interest, under regularity conditions. The Bayesian paradigm is concerned with including the researcher’s prior uncertainty about the parameter through a prior distribution, which combined with the likelihood yields a posterior distribution for the parameter that reflects the researcher’s posterior uncertainty. In this paper, we present a way …


Diagnosing And Responding To Violations In The Positivity Assumption, Maya L. Petersen, Kristin Porter, Susan Gruber, Yue Wang, Mark J. Van Der Laan Sep 2010

Diagnosing And Responding To Violations In The Positivity Assumption, Maya L. Petersen, Kristin Porter, Susan Gruber, Yue Wang, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The assumption of positivity or experimental treatment assignment requires that observed treatment levels vary within confounder strata. This article discusses the positivity assumption in the context of assessing model and parameter-specific identifiability of causal effects. Positivity violations occur when certain subgroups in a sample rarely or never receive some treatments of interest. The resulting sparsity in the data may increase bias with or without an increase in variance and can threaten valid inference. The parametric bootstrap is presented as a tool to assess the severity of such threats and its utility as a diagnostic is explored using simulated data. Several …


Estimation Of Causal Effects Of Community Based Interventions, Mark J. Van Der Laan Jun 2010

Estimation Of Causal Effects Of Community Based Interventions, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose one assigns two interventions to a small number K of different populations or communities, and one measures covariates and outcomes on a random sample of independent individuals from each of the K populations. We investigate the problem of identification and estimation of the causal effect of the choice of intervention assigned at the community level, and, if the intervention is time-dependent, the causal effect of the changes in the intervention at time t, on the outcome. The challenge one is confronted with is that different populations have different environmental factors and that the intervention and environment are assigned to …


A Targeted Maximum Likelihood Estimator Of A Causal Effect On A Bounded Continuous Outcome, Susan Gruber, Mark J. Van Der Laan May 2010

A Targeted Maximum Likelihood Estimator Of A Causal Effect On A Bounded Continuous Outcome, Susan Gruber, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation of a parameter of a data generating distribution, known to be an element of a semiparametric model, involves constructing a parametric model through an initial density estimator with parameter epsilon representing an amount of fluctuation of the initial density estimator, where the score of this fluctuation model at epsilon=0 equals the efficient influence curve/canonical gradient. The latter constraint can be satisfied by many parametric fluctuation models, since it represents only a local constraint of its behavior at zero fluctuation. However, it is very important that the fluctuations stay within the semiparametric model for the observed data …


Super Learner In Prediction, Eric C. Polley, Mark J. Van Der Laan May 2010

Super Learner In Prediction, Eric C. Polley, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Super learning is a general loss based learning method that has been proposed and analyzed theoretically in van der Laan et al. (2007). In this article we consider super learning for prediction. The super learner is a prediction method designed to find the optimal combination of a collection of prediction algorithms. The super learner algorithm finds the combination of algorithms minimizing the cross-validated risk. The super learner framework is built on the theory of cross-validation and allows for a general class of prediction algorithms to be considered for the ensemble. Due to the previously established oracle results for the cross-validation …


Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan Mar 2010

Targeted Maximum Likelihood Method For Repeated Measures Semiparametric Regression: Discovery For Transcription Factor Activity, Catherine Tuglus, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In longitudinal and repeated measures data analysis, often the goal is to determine the effect of a treatment or aspect on a particular outcome (e.g. disease progression). We consider semiparametric repeated measures regression model, where the parametric component models effect of the variable of interest and any modification by other covariates. The expectation of this parametric component over the other covariates is a measure of variable importance. Here we present a targeted maximum likelihood estimator of the finite dimensional regression parameter, which is easily estimated using standard software for generalized estimating equations. The targeted maximum likelihood method provides double robust …


Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan Feb 2010

Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

This article is devoted to the asymptotic study of adaptive group sequential designs in the case of randomized clinical trials with binary treatment, binary outcome and no covariate. By adaptive design, we mean in this setting a clinical trial design that allows the investigator to dynamically modify its course through data-driven adjustment of the randomization probability based on data accrued so far, without negatively impacting on the statistical integrity of the trial. By adaptive group sequential design, we refer to the fact that group sequential testing methods can be equally well applied on top of adaptive designs. Prior to collection …


Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan Feb 2010

Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Given causal graph assumptions, intervention-specific counterfactual distributions of the data can be defined by the so called G-computation formula, which is obtained by carrying out these interventions on the likelihood of the data factorized according to the causal graph. The obtained G-computation formula represents the counterfactual distribution the data would have had if this intervention would have been enforced on the system generating the data. A causal effect of interest can now be defined as some difference between these counterfactual distributions indexed by different interventions. For example, the interventions can represent static treatment regimens or individualized treatment rules that assign …


Simple, Efficient Estimators Of Treatment Effects In Randomized Trials Using Generalized Linear Models To Leverage Baseline Variables, Michael Rosenblum, Mark J. Van Der Laan Jan 2010

Simple, Efficient Estimators Of Treatment Effects In Randomized Trials Using Generalized Linear Models To Leverage Baseline Variables, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Models, such as logistic regression and Poisson regression models, are often used to estimate treatment effects in randomized trials. These models leverage information in variables collected before randomization, in order to obtain more precise estimates of treatment effects. However, there is the danger that model misspecification will lead to bias. We show that certain easy to compute, model-based estimators are asymptotically unbiased even when the working model used is arbitrarily misspecified. Furthermore, these estimators are locally efficient. As a special case of our main result, we consider a simple Poisson working model containing only main terms; in this case, we …


Targeted Maximum Likelihood Estimation Of The Parameter Of A Marginal Structural Model, Michael Rosenblum, Mark J. Van Der Laan Jan 2010

Targeted Maximum Likelihood Estimation Of The Parameter Of A Marginal Structural Model, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Targeted maximum likelihood estimation is a versatile tool for estimating parameters in semiparametric and nonparametric models. We work through an example applying targeted maximum likelihood methodology to estimate the parameter of a marginal structural model. In the case we consider, we show how this can be easily done by clever use of standard statistical software. We point out differences between targeted maximum likelihood estimation and other approaches (including estimating function based methods). The application we consider is to estimate the effect of adherence to antiretroviral medications on virologic failure in HIV positive individuals.


Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber Sep 2009

Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber

U.C. Berkeley Division of Biostatistics Working Paper Series

This is a compilation of current and past work on targeted maximum likelihood estimation. It features the original targeted maximum likelihood learning paper as well as chapters on super (machine) learning using cross validation, randomized controlled trials, realistic individualized treatment rules in observational studies, biomarker discovery, case-control studies, and time-to-event outcomes with censored data, among others. We hope this collection is helpful to the interested reader and stimulates additional research in this important area.


Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan Sep 2009

Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

A nested case-control study is conducted within a well-defined cohort arising out of a population of interest. This design is often used in epidemiology to reduce the costs associated with collecting data on the full cohort; however, the case control sample within the cohort is a biased sample. Methods for analyzing case-control studies have largely focused on logistic regression models that provide conditional and not marginal causal estimates of the odds ratio. We previously developed a Case-Control Weighted Targeted Maximum Likelihood Estimation (TMLE) procedure for case-control study designs, which relies on the prevalence probability q0. We propose the use of …


Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan Jun 2009

Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

For estimating regressions for repeated measures outcome data, a popular choice is the population average models estimated by generalized estimating equations (GEE). We review in this report the derivation of the robust inference (sandwich-type estimator of the standard error). In addition, we present formally how the approximation of a misspecified working population average model relates to the true model and in turn how to interpret the results of such a misspecified model.


A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell Jun 2009

A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

No abstract provided.


Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit Apr 2009

Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Gaussian graphical models have become popular tools for identifying relationships between genes when analyzing microarray expression data. In the classical undirected Gaussian graphical model setting, conditional independence relationships can be inferred from partial correlations obtained from the concentration matrix (= inverse covariance matrix) when the sample size n exceeds the number of parameters p which need to estimated. In situations where n < p, another approach to graphical model estimation may rely on calculating unconditional (zero-order) and first-order partial correlations. In these settings, the goal is to identify a lower-order conditional independence graph, sometimes referred to as a ‘0-1 graphs’. For either choice of graph, model selection may involve a multiple testing problem, in which edges in a graph are drawn only after rejecting hypotheses involving (saturated or lower-order) partial correlation parameters. Most multiple testing procedures applied in previously proposed graphical model selection algorithms rely on standard, marginal testing methods which do not take into account the joint distribution of the test statistics derived from (partial) correlations. We propose and implement a multiple testing framework useful when testing for edge inclusion during graphical model selection. Two features of our methodology include (i) a computationally efficient and asymptotically valid test statistics joint null distribution derived from influence curves for correlation-based parameters, and (ii) the application of empirical Bayes joint multiple testing procedures which can effectively control a variety of popular Type I error rates by incorpo- rating joint null distributions such as those described here (Dudoit and van der Laan, 2008). Using a dataset from Arabidopsis thaliana, we observe that the use of more sophisticated, modular approaches to multiple testing allows one to identify greater numbers of edges when approximating an undirected graphical model using a 0-1 graph. Our framework may also be extended to edge testing algorithms for other types of graphical models (e.g., for classical undirected, bidirected, and directed acyclic graphs).


A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin Dec 2008

A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin

U.C. Berkeley Division of Biostatistics Working Paper Series

The attributable risk, often called the population attributable risk, is in many epidemiological contexts a more relevant measure of exposure-disease association than the excess risk, relative risk, or odds ratio. When estimating attributable risk with case-control data and a rare disease, we present a simple correction to the standard approach making it essentially unbiased, and also less noisy. As with analogous corrections given in Jewell (1986) for other measures of association, the adjustment often won't make a substantial difference unless the sample size is very small or point estimates are desired within fine strata, but we discuss the possible utility …