Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (118)
- Statistical Theory (116)
- Statistical Methodology (114)
- Statistical Models (60)
- Survival Analysis (48)
-
- Medicine and Health Sciences (31)
- Epidemiology (26)
- Public Health (26)
- Life Sciences (20)
- Multivariate Analysis (20)
- Longitudinal Data Analysis and Time Series (19)
- Genetics and Genomics (18)
- Applied Mathematics (17)
- Numerical Analysis and Computation (17)
- Genetics (15)
- Clinical Trials (13)
- Microarrays (12)
- Design of Experiments and Sample Surveys (7)
- Categorical Data Analysis (5)
- Bioinformatics (4)
- Computational Biology (4)
- Disease Modeling (4)
- Diseases (4)
- Laboratory and Basic Science Research (4)
- Applied Statistics (3)
- Medical Specialties (3)
- Other Statistics and Probability (1)
- Vital and Health Statistics (1)
- Keyword
-
- Cross-validation (23)
- Causal inference (20)
- Influence curve (14)
- Prediction (12)
- Efficient influence curve (11)
-
- Model selection (11)
- Targeted maximum likelihood estimation (11)
- Loss function (10)
- Bootstrap (9)
- Causal effect (9)
- Counterfactual (9)
- Adjusted p-value (8)
- Asymptotic linearity (8)
- Multiple testing (8)
- Super-learning (8)
- Type I error rate (8)
- Confounding (7)
- Empirical process (7)
- One-step estimator (7)
- Survival analysis (7)
- Counting process (6)
- Estimating equation (6)
- False discovery rate (6)
- Gene expression (6)
- Null distribution (6)
- Pathwise differentiable parameter (6)
- Semiparametric statistical model (6)
- Asymptotic control (5)
- Censored data (5)
- Censoring (5)
Articles 91 - 120 of 242
Full-Text Articles in Statistics and Probability
Collaborative Targeted Maximum Likelihood For Time To Event Data, Ori M. Stitelman, Mark J. Van Der Laan
Collaborative Targeted Maximum Likelihood For Time To Event Data, Ori M. Stitelman, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Current methods used to analyze time to event data either, rely on highly parametric assumptions which result in biased estimates of parameters which are purely chosen out of convenience, or are highly unstable because they ignore the global constraints of the true model. By using Targeted Maximum Likelihood Estimation one may consistently estimate parameters which directly answer the statistical question of interest. Targeted Maximum Likelihood Estimators are substitution estimators, which rely on estimating the underlying distribution. However, unlike other substitution estimators, the underlying distribution is estimated specifically to reduce bias in the estimate of the parameter of interest. We will …
Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan
Targeting The Optimal Design In Randomized Clinical Trials With Binary Outcomes And No Covariate, Antoine Chambaz, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
This article is devoted to the asymptotic study of adaptive group sequential designs in the case of randomized clinical trials with binary treatment, binary outcome and no covariate. By adaptive design, we mean in this setting a clinical trial design that allows the investigator to dynamically modify its course through data-driven adjustment of the randomization probability based on data accrued so far, without negatively impacting on the statistical integrity of the trial. By adaptive group sequential design, we refer to the fact that group sequential testing methods can be equally well applied on top of adaptive designs. Prior to collection …
Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan
Targeted Maximum Likelihood Based Causal Inference, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Given causal graph assumptions, intervention-specific counterfactual distributions of the data can be defined by the so called G-computation formula, which is obtained by carrying out these interventions on the likelihood of the data factorized according to the causal graph. The obtained G-computation formula represents the counterfactual distribution the data would have had if this intervention would have been enforced on the system generating the data. A causal effect of interest can now be defined as some difference between these counterfactual distributions indexed by different interventions. For example, the interventions can represent static treatment regimens or individualized treatment rules that assign …
Simple, Efficient Estimators Of Treatment Effects In Randomized Trials Using Generalized Linear Models To Leverage Baseline Variables, Michael Rosenblum, Mark J. Van Der Laan
Simple, Efficient Estimators Of Treatment Effects In Randomized Trials Using Generalized Linear Models To Leverage Baseline Variables, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Models, such as logistic regression and Poisson regression models, are often used to estimate treatment effects in randomized trials. These models leverage information in variables collected before randomization, in order to obtain more precise estimates of treatment effects. However, there is the danger that model misspecification will lead to bias. We show that certain easy to compute, model-based estimators are asymptotically unbiased even when the working model used is arbitrarily misspecified. Furthermore, these estimators are locally efficient. As a special case of our main result, we consider a simple Poisson working model containing only main terms; in this case, we …
Targeted Maximum Likelihood Estimation Of The Parameter Of A Marginal Structural Model, Michael Rosenblum, Mark J. Van Der Laan
Targeted Maximum Likelihood Estimation Of The Parameter Of A Marginal Structural Model, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Targeted maximum likelihood estimation is a versatile tool for estimating parameters in semiparametric and nonparametric models. We work through an example applying targeted maximum likelihood methodology to estimate the parameter of a marginal structural model. In the case we consider, we show how this can be easily done by clever use of standard statistical software. We point out differences between targeted maximum likelihood estimation and other approaches (including estimating function based methods). The application we consider is to estimate the effect of adherence to antiretroviral medications on virologic failure in HIV positive individuals.
Causal Inference In Epidemiological Studies With Strong Confounding, Kelly L. Moore, Romain S. Neugebauer, Mark J. Van Der Laan, Ira B. Tager
Causal Inference In Epidemiological Studies With Strong Confounding, Kelly L. Moore, Romain S. Neugebauer, Mark J. Van Der Laan, Ira B. Tager
U.C. Berkeley Division of Biostatistics Working Paper Series
One of the identifiabilty assumptions of causal effects defined by marginal structural model (MSM) parameters is the experimental treatment assignment (ETA) assumption. Practical violations of this assumption frequently occur in data analysis, when certain exposures are rarely observed within some strata of the population. The inverse probability of treatment weighted (IPTW) estimator is particularly sensitive to violations of this assumption, however, we demonstrate that this is a problem for all estimators of causal effects. This is due to the fact that the ETA assumption is about information (or lack thereof) in the data. A new class of causal models, causal …
Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber
Readings In Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Sherri Rose, Susan Gruber
U.C. Berkeley Division of Biostatistics Working Paper Series
This is a compilation of current and past work on targeted maximum likelihood estimation. It features the original targeted maximum likelihood learning paper as well as chapters on super (machine) learning using cross validation, randomized controlled trials, realistic individualized treatment rules in observational studies, biomarker discovery, case-control studies, and time-to-event outcomes with censored data, among others. We hope this collection is helpful to the interested reader and stimulates additional research in this important area.
Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan
Causal Inference For Nested Case-Control Studies Using Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
A nested case-control study is conducted within a well-defined cohort arising out of a population of interest. This design is often used in epidemiology to reduce the costs associated with collecting data on the full cohort; however, the case control sample within the cohort is a biased sample. Methods for analyzing case-control studies have largely focused on logistic regression models that provide conditional and not marginal causal estimates of the odds ratio. We previously developed a Case-Control Weighted Targeted Maximum Likelihood Estimation (TMLE) procedure for case-control study designs, which relies on the prevalence probability q0. We propose the use of …
Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan
Targeted Maximum Likelihood Estimation: A Gentle Introduction, Susan Gruber, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper provides a concise introduction to targeted maximum likelihood estimation (TMLE) of causal effect parameters. The interested analyst should gain sufficient understanding of TMLE from this introductory tutorial to be able to apply the method in practice. A program written in R is provided. This program implements a basic version of TMLE that can be used to estimate the effect of a binary point treatment on a continuous or binary outcome.
Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan
Nonparametric Population Average Models: Deriving The Form Of Approximate Population Average Models Estimated Using Generalized Estimating Equations, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
For estimating regressions for repeated measures outcome data, a popular choice is the population average models estimated by generalized estimating equations (GEE). We review in this report the derivation of the robust inference (sandwich-type estimator of the standard error). In addition, we present formally how the approximation of a misspecified working population average model relates to the true model and in turn how to interpret the results of such a misspecified model.
A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell
A Machine-Learning Algorithm For Estimating And Ranking The Impact Of Environmental Risk Factors In Exploratory Epidemiological Studies, Jessica G. Young, Alan E. Hubbard, B Eskenazi, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
No abstract provided.
Resampling-Based Multiple Hypothesis Testing With Applications To Genomics: New Developments In The R/Bioconductor Package Multtest, Houston N. Gilbert, Katherine S. Pollard, Mark J. Van Der Laan, Sandrine Dudoit
Resampling-Based Multiple Hypothesis Testing With Applications To Genomics: New Developments In The R/Bioconductor Package Multtest, Houston N. Gilbert, Katherine S. Pollard, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
The multtest package is a standard Bioconductor package containing a suite of functions useful for executing, summarizing, and displaying the results from a wide variety of multiple testing procedures (MTPs). In addition to many popular MTPs, the central methodological focus of the multtest package is the implementation of powerful joint multiple testing procedures. Joint MTPs are able to account for the dependencies between test statistics by effectively making use of (estimates of) the test statistics joint null distribution. To this end, two additional bootstrap-based estimates of the test statistics joint null distribution have been developed for use in the …
Application Of Time-To-Event Methods In The Assessment Of Safety In Clinical Trials, Kelly L. Moore, Mark J. Van Der Laan
Application Of Time-To-Event Methods In The Assessment Of Safety In Clinical Trials, Kelly L. Moore, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Since randomized controlled trials (RCT) are typically designed and powered for efficacy rather than safety, power is an important concern in the analysis of the effect of treatment on the occurrence of adverse events (AE). These outcomes are often time-to-event outcomes which will naturally be subject to right-censoring due to early patient withdrawals. In the analysis of the treatment effect on such an outcome, gains in efficiency, and thus power, can be achieved by exploiting covariate information. We apply the targeted maximum likelihood methodology to the estimation of treatment specific survival at a fixed end point for right-censored survival outcomes. …
Collaborative Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Susan Gruber
Collaborative Targeted Maximum Likelihood Estimation, Mark J. Van Der Laan, Susan Gruber
U.C. Berkeley Division of Biostatistics Working Paper Series
Collaborative double robust targeted maximum likelihood estimators represent a fundamental further advance over standard targeted maximum likelihood estimators of causal inference and variable importance parameters. The targeted maximum likelihood approach involves fluctuating an initial density estimate, (Q), in order to make a bias/variance tradeoff targeted towards a specific parameter in a semi-parametric model. The fluctuation involves estimation of a nuisance parameter portion of the likelihood, g. TMLE and other double robust estimators have been shown to be consistent and asymptotically normally distributed (CAN) under regularity conditions, when either one of these two factors of the likelihood of the data is …
Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit
Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Gaussian graphical models have become popular tools for identifying relationships between genes when analyzing microarray expression data. In the classical undirected Gaussian graphical model setting, conditional independence relationships can be inferred from partial correlations obtained from the concentration matrix (= inverse covariance matrix) when the sample size n exceeds the number of parameters p which need to estimated. In situations where n < p, another approach to graphical model estimation may rely on calculating unconditional (zero-order) and first-order partial correlations. In these settings, the goal is to identify a lower-order conditional independence graph, sometimes referred to as a ‘0-1 graphs’. For either choice of graph, model selection may involve a multiple testing problem, in which edges in a graph are drawn only after rejecting hypotheses involving (saturated or lower-order) partial correlation parameters. Most multiple testing procedures applied in previously proposed graphical model selection algorithms rely on standard, marginal testing methods which do not take into account the joint distribution of the test statistics derived from (partial) correlations. We propose and implement a multiple testing framework useful when testing for edge inclusion during graphical model selection. Two features of our methodology include (i) a computationally efficient and asymptotically valid test statistics joint null distribution derived from influence curves for correlation-based parameters, and (ii) the application of empirical Bayes joint multiple testing procedures which can effectively control a variety of popular Type I error rates by incorpo- rating joint null distributions such as those described here (Dudoit and van der Laan, 2008). Using a dataset from Arabidopsis thaliana, we observe that the use of more sophisticated, modular approaches to multiple testing allows one to identify greater numbers of edges when approximating an undirected graphical model using a 0-1 graph. Our framework may also be extended to edge testing algorithms for other types of graphical models (e.g., for classical undirected, bidirected, and directed acyclic graphs).
Selecting Optimal Treatments Based On Predictive Factors, Eric C. Polley, Mark J. Van Der Laan
Selecting Optimal Treatments Based On Predictive Factors, Eric C. Polley, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
No abstract provided.
A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin
A Small Sample Correction For Estimating Attributable Risk In Case-Control Studies, Daniel B. Rubin
U.C. Berkeley Division of Biostatistics Working Paper Series
The attributable risk, often called the population attributable risk, is in many epidemiological contexts a more relevant measure of exposure-disease association than the excess risk, relative risk, or odds ratio. When estimating attributable risk with case-control data and a rare disease, we present a simple correction to the standard approach making it essentially unbiased, and also less noisy. As with analogous corrections given in Jewell (1986) for other measures of association, the adjustment often won't make a substantial difference unless the sample size is very small or point estimates are desired within fine strata, but we discuss the possible utility …
A Note On Risk Prediction For Case-Control Studies, Sherri Rose, Mark J. Van Der Laan
A Note On Risk Prediction For Case-Control Studies, Sherri Rose, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We introduce a new method for prediction in case-control study designs, which is a simple extension of the work by van der Laan (2008). Case-control samples are biased since the proportion of cases in the sample is not the same as the population of interest. The case-control weighting for prediction proposed in this paper relies on knowledge of the true incidence probability P(Y=1) to eliminate the bias of the sampling design. In many practical settings, case-control weighting will outperform an existing method for prediction, intercept adjustment.
Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans
Confidence Intervals For Negative Binomial Random Variables Of High Dispersion, David Shilane, Alan E. Hubbard, S N. Evans
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper considers the problem of constructing confidence intervals for the mean of a Negative Binomial random variable based upon sampled data. When the sample size is large, we traditionally rely upon a Normal distribution approximation to construct these intervals. However, we demonstrate that the sample mean of highly dispersed Negative Binomials exhibits a slow convergence to the Normal in distribution as a function of the sample size. As a result, standard techniques (such as the Normal approximation and bootstrap) that construct confidence intervals for the mean will typically be too narrow and significantly undercover in the case of high …
Why Match? Investigating Matched Case-Control Study Designs With Causal Effect Estimation, Sherri Rose, Mark J. Van Der Laan
Why Match? Investigating Matched Case-Control Study Designs With Causal Effect Estimation, Sherri Rose, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Matched case-control study designs are commonly implemented in the field of public health. While matching is intended to eliminate confounding, the main potential benefit of matching in case-control studies is a gain in efficiency. Methods for analyzing matched case-control studies have focused on utilizing conditional logistic regression models that provide conditional and not causal estimates of the odds ratio. This article investigates the use of case-control weighted targeted maximum likelihood estimation to obtain marginal causal effects in matched case-control study designs. We compare the use of case-control weighted targeted maximum likelihood estimation in matched and unmatched designs in an effort …
Fdr Controlling Procedure For Multi-Stage Analyses, Catherine Tuglus, Mark J. Van Der Laan
Fdr Controlling Procedure For Multi-Stage Analyses, Catherine Tuglus, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Multiple testing has become an integral component in genomic analyses involving microarray experiments where large number of hypotheses are tested simultaneously. However before applying more computationally intensive methods, it is often desirable to complete an initial truncation of the variable set using a simpler and faster supervised method such as univariate regression. Once such a truncation is completed, multiple testing methods applied to any subsequent analysis no longer control the appropriate Type I error rates. Here we propose a modified marginal Benjamini \& Hochberg step-up FDR controlling procedure for multi-stage analyses (FDR-MSA), which correctly controls Type I error in terms …
Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan
Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a new approach to studying the relationship between a very high dimensional random variable and an outcome. Our method is based on a novel concept, the supervised distance matrix, which quantifies pairwise similarity between variables based on their association with the outcome. A supervised distance matrix is derived in two stages. The first stage involves a transformation based on a particular model for association. In particular, one might regress the outcome on each variable and then use the residuals or the influence curve from each regression as a data transformation. In the second stage, a choice of distance …
Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan
Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
The validity of standard confidence intervals constructed in survey sampling is based on the central limit theorem. For small sample sizes, the central limit theorem may give a poor approximation, resulting in confidence intervals that are misleading. We discuss this issue and propose methods for constructing confidence intervals for the population mean tailored to small sample sizes.
We present a simple approach for constructing confidence intervals for the population mean based on tail bounds for the sample mean that are correct for all sample sizes. Bernstein's inequality provides one such tail bound. The resulting confidence intervals have guaranteed coverage probability …
Doubly Robust Ecological Inference, Daniel B. Rubin, Mark J. Van Der Laan
Doubly Robust Ecological Inference, Daniel B. Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
The ecological inference problem is a famous longstanding puzzle that arises in many disciplines. The usual formulation in epidemiology is that we would like to quantify an exposure-disease association by obtaining disease rates among the exposed and unexposed, but only have access to exposure rates and disease rates for several regions. The problem is generally intractable, but can be attacked under the assumptions of King's (1997) extended technique if we can correctly specify a model for a certain conditional distribution. We introduce a procedure that it is a valid approach if either this original model is correct or if we …
Estimation Based On Case-Control Designs With Known Incidence Probability, Mark J. Van Der Laan
Estimation Based On Case-Control Designs With Known Incidence Probability, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Case-control sampling is an extremely common design used to generate data to estimate effects of exposures or treatments on a binary outcome of interest when the proportion of cases (i.e., binary outcome equal to 1) in the population of interest is low. Case-control sampling represents a biased sample of a target population of interest by sampling a disproportional number of cases. Case-control studies are also commonly employed to estimate the effects of genetic markers or biomarkers on phenotypes. The typical approach used in practice is to fit (conditional) logistic regression models, ignoring the case-control sampling, in order to estimate the …
A Guide To Causal Parameters In Case-Control Designs: Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan
A Guide To Causal Parameters In Case-Control Designs: Targeted Maximum Likelihood Estimation, Sherri Rose, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Researchers of uncommon diseases are often interested in assessing potential risk factors. Given the low incidence of disease, these studies are frequently case-control in design, as this allows for a sufficient number of cases to be obtained without extensive sampling and can increase efficiency. However, these case-control samples are then biased since the proportion of cases in the sample is not the same as the population of interest. Methods for analyzing case-control studies have focused on utilizing logistic regression models that provide conditional and not causal estimates of the odds ratio. This article will demonstrate the use of the prevalence …
Targeted Methods For Biomarker Discovery, The Search For A Standard, Catherine Tuglus, Mark J. Van Der Laan
Targeted Methods For Biomarker Discovery, The Search For A Standard, Catherine Tuglus, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
More often than not biomarker studies analyze large quantities of variables with complicated and generally unknown correlation structure. There are numerous statistical methods which attempt to unravel these variables and determine the underlying mechanism through identification of causally related biomarkers. Results from these methods are generally difficult to interpret and nearly impossible to compare across studies. The FDA has currently called for a standardization of methods and protocol for biomarker detection. In response, we propose targeted variable importance (tVIM) as a standardized method for biomarker discovery. Through the use of targeted Maximum Likelihood, tVIM provides double robust estimates of variable …
The Construction And Analysis Of Adaptive Group Sequential Designs, Mark J. Van Der Laan
The Construction And Analysis Of Adaptive Group Sequential Designs, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In order to answer scientific questions of interest one often carries out an ordered sequence of experiments generating the appropriate data over time. The design of each experiment involves making various decisions such as 1) What variables to measure on the randomly sampled experimental unit?, 2) How regularly to monitor the unit, and for how long?, 3) How to randomly assign a treatment or drug-dose to the unit?, among others. That is, the design of each experiment involves selecting a so called treatment mechanism/monitoring mechanism/ missingness/censoring mechanism, where these mechanisms represent a formally defined conditional distribution of one of these …
Data-Adaptive Selection Of The Truncation Level For Inverse-Probability-Of-Treatment-Weighted Estimators, Oliver Bembom, Mark J. Van Der Laan
Data-Adaptive Selection Of The Truncation Level For Inverse-Probability-Of-Treatment-Weighted Estimators, Oliver Bembom, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Inverse-Probability-of-Treatment-Weighted (IPTW) estimators are becoming a popular analysis tool in causal inference. It is well known that these estimators suffer from high variability if some treatment probabilities are estimated to be close to zero. While it is a common recommendation for such situations to truncate the weights in order to reduce the mean squared error of the estimator, the current literature gives little guidance on how to select an appropriate truncation level. In this article, we develop a closed-form estimate for the mean squared error of a truncated IPTW estimator that can be used to select this truncation level data-adaptively. …
Data-Adaptive Selection Of The Adjustment Set In Variable Importance Estimation, Oliver Bembom, Jeffrey W. Fessel, Robert W. Shafer, Mark J. Van Der Laan
Data-Adaptive Selection Of The Adjustment Set In Variable Importance Estimation, Oliver Bembom, Jeffrey W. Fessel, Robert W. Shafer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
If estimates of the effect of a treatment variable on an outcome of interest are to be adjusted for a set of possible confounding factors, it is necessary to rely on the assumption of experimental treatment assignment (ETA) according to which each experimental unit has positive probability of being observed at any of the possible levels of the treatment variable regardless of the values the confounding factors may take on. Even if this assumption is only practically violated in the sense that certain values of the confounding factors cause some treatment levels to become not impossible, but at least highly …