When Does Combining Markers Improve Classification Performance And What Are Implications For Practice?,
2011
University of Washington
When Does Combining Markers Improve Classification Performance And What Are Implications For Practice?, Aasthaa Bansal, Margaret Sullivan Pepe
UW Biostatistics Working Paper Series
When an existing standard marker does not have sufficient classification accuracy on its own, new markers are sought with the goal of yielding a combination with better performance. The primary criterion for selecting new markers is that they have good performance on their own and preferably be uncorrelated with the standard. Most often linear combinations are considered. In this paper we investigate the increment in performance that is possible by combining a novel continuous marker with a moderately performing standard continuous marker under a variety of biologically motivated models for their joint distribution. We find that an uncorrelated continuous marker …
Targeted Maximum Likelihood Estimation Of Conditional Relative Risk In A Semi-Parametric Regression Model,
2011
University of California, Berkeley
Targeted Maximum Likelihood Estimation Of Conditional Relative Risk In A Semi-Parametric Regression Model, Cathy Tuglus, Kristin E. Porter, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
The conditional relative risk is an important measure in medical and epidemiological studies when the outcome of interest is binary (i.e. disease vs. no disease). When the outcome is common, estimation of conditional relative risk and related parameters can be problematic, especially when the exposure or covariates are continuous. We propose a new estimation procedure based on targeted maximum likelihood methodology that targets the parameters relating to the conditional relative risk for common outcomes under a log-linear, or multiplicative, semi-parametric model. In this paper, we present three possible targeted maximum likelihood estimators for relative risk parameters implied by such a …
Super Learner Based Conditional Density Estimation With Application To Marginal Structural Models,
2011
University of California, Berkeley, School of Public Health - Division of Biostatistics
Super Learner Based Conditional Density Estimation With Application To Marginal Structural Models, Ivan Diaz Munoz, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper we present a histogram-like estimator of a conditional density that uses super learner crossvalidation to estimate the histogram probabilities, as well as the optimal number and position of the bins. This estimator is an alternative to kernel density estimators when the dimension of the problem is large. We demonstrate its applicability to estimation of Marginal Structural Model (MSM) parameters in which an initial estimator of the treatment %mechanism is needed. MSM estimation based on the proposed density estimator results in less biased estimates, when compared to estimates based on a misspecified parametric model.
Comparing Roc Curves Derived From Regression Models,
2011
Memorial Sloan-Kettering Cancer Center
Comparing Roc Curves Derived From Regression Models, Venkatraman E. Seshan, Mithat Gonen, Colin B. Begg
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
In constructing predictive models, investigators frequently assess the incremental value of a predictive marker by comparing the ROC curve generated from the predictive model including the new marker with the ROC curve from the model excluding the new marker. Many commentators have noticed empirically that a test of the two ROC areas often produces a non-significant result when a corresponding Wald test from the underlying regression model is significant. A recent article showed using simulations that the widely-used ROC area test [1] produces exceptionally conservative test size and extremely low power [2]. In this article we show why the ROC …
On Causal Mediation Analysis With A Survival Outcome,
2011
Harvard University
On Causal Mediation Analysis With A Survival Outcome, Eric J. Tchetgen Tchetgen
Harvard University Biostatistics Working Paper Series
Suppose that having established a marginal total effect of a point exposure on a time-to-event outcome, an investigator wishes to decompose this effect into its direct and indirect pathways, also know as natural direct and indirect effects, mediated by a variable known to occur after the exposure and prior to the outcome. This paper proposes a theory of estimation of natural direct and indirect effects in two important semiparametric models for a failure time outcome. The underlying survival model for the marginal total effect and thus for the direct and indirect effects, can either be a marginal structural Cox proportional …
Semiparametric Estimation Of Models For Natural Direct And Indirect Effects,
2011
Harvard University
Semiparametric Estimation Of Models For Natural Direct And Indirect Effects, Eric J. Tchetgen Tchetgen, Ilya Shpitser
Harvard University Biostatistics Working Paper Series
In recent years, researchers in the health and social sciences have become increasingly interested in mediation analysis. Specifically, upon establishing a non-null total effect of an exposure, investigators routinely wish to make inferences about the direct (indirect) pathway of the effect of the exposure not through (through) a mediator variable that occurs subsequently to the exposure and prior to the outcome. Natural direct and indirect effects are of particular interest as they generally combine to produce the total effect of the exposure and therefore provide insight on the mechanism by which it operates to produce the outcome. A semiparametric theory …
Semiparametric Theory For Causal Mediation Analysis: Efficiency Bounds, Multiple Robustness, And Sensitivity Analysis,
2011
Harvard University
Semiparametric Theory For Causal Mediation Analysis: Efficiency Bounds, Multiple Robustness, And Sensitivity Analysis, Eric J. Tchetgen Tchetgen, Ilya Shpitser
Harvard University Biostatistics Working Paper Series
Whilst estimation of the marginal (total) causal effect of a point exposure on an outcome is arguably the most common objective of experimental and observational studies in the health and social sciences, in recent years, investigators have also become increasingly interested in mediation analysis. Specifically, upon establishing a non-null total effect of the exposure, investigators routinely wish to make inferences about the direct (indirect) pathway of the effect of the exposure not through (through) a mediator variable that occurs subsequently to the exposure and prior to the outcome. Although powerful semiparametric methodologies have been developed to analyze observational studies, that …
Relationship Of Vitamin D Levels To Blood Pressure In A Biethnic Cohort,
2011
Loma Linda University
Relationship Of Vitamin D Levels To Blood Pressure In A Biethnic Cohort, Rosario O. Sakamoto
Loma Linda University Electronic Theses, Dissertations & Projects
Background: Serum hydroxyvitamin D [25(OH)D] has been well-accepted as not an ordinary vitamin but a pro-hormone that has many benefits beyond its well-known effects on bone. Cardiovascular risk factors such as hypertension remain a huge health burden and Blacks have been recognized to have higher prevalence of hypertension compared to non-Hispanic Whites. Despite increasing evidence of the benefits of vitamin D on blood pressure control, there is much more to be learned about the relationship of serum 25(OH)D to blood pressure among different ethnicities.
Purpose: The goal of this study was to determine whether vitamin D serum 25(OH)D levels were …
Ranking Single Nucleotide Polymorphisms With Support Vector Regression In Continuous Phenotypes,
2011
New Jersey Institute of Technology
Ranking Single Nucleotide Polymorphisms With Support Vector Regression In Continuous Phenotypes, Seif Shahidain
Theses
Support vector machines (SVM) have been used to improve the ranking of single nucleotide polymorphisms (SNPs) over traditional chi-square tests in disease case studies [2]. In this investigation, ranking SNPs with support vector regression (SVR) was compared to the Wald test in predicting continuous phenotypes. SVR-ranked SNPs consistently outperformed the Wald test-ranked SNPs to provide a more accurate prediction of the phenotype with fewer SNPs across several methods of prediction.
A General Implementation Of Tmle For Longitudinal Data Applied To Causal Inference In Survival Analysis,
2011
University of California, Berkeley, DIvision of Biostatistics
A General Implementation Of Tmle For Longitudinal Data Applied To Causal Inference In Survival Analysis, Ori M. Stitelman, Victor De Gruttola, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In many randomized controlled trials the outcome of interest is a time to event, and one measures on each subject baseline covariates and time-dependent covariates until the subject either drops-out, the time to event is observed, or the end of study is reached. The goal of such a study is to assess the causal effect of the treatment on the survival curve. Standard methods (e.g., Kaplan-Meier estimator, Cox-proportional hazards) ignore the available baseline and time-dependent covariates, and are therefore biased if the drop-out is affected by these covariates, and are always inefficient. We present a targeted maximum likelihood estimator of …
Targeted Minimum Loss Based Estimator That Outperforms A Given Estimator,
2011
University of California, Berkeley
Targeted Minimum Loss Based Estimator That Outperforms A Given Estimator, Susan Gruber, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Targeted minimum loss based estimation (TMLE) provides a template for the construction of semiparametric locally efficient double robust substitution estimators of the target parameter of the data generating distribution in a semiparametric censored data or causal inference model (van der Laan and Rubin (2006),van der Laan (2008), van der Laan and Rose (2011)). In this article we demonstrate how to construct a TMLE that also satisfies the property that it is at least as efficient as a user supplied asymptotically linear estimator. For the sake of illustration we focus on estimation of the additive average causal effect of a point …
The Relative Performance Of Targeted Maximum Likelihood Estimators,
2011
University of California, Berkeley, DIvision of Biostatistics
The Relative Performance Of Targeted Maximum Likelihood Estimators, Kristin E. Porter, Susan Gruber, Mark J. Van Der Laan, Jasjeet S. Sekhon
U.C. Berkeley Division of Biostatistics Working Paper Series
There is an active debate in the literature on censored data about the relative performance of model based maximum likelihood estimators, IPCW-estimators, and a variety of double robust semiparametric efficient estimators. Kang and Schafer (2007) demonstrate the fragility of double robust and IPCW-estimators in a simulation study with positivity violations. They focus on a simple missing data problem with covariates where one desires to estimate the mean of an outcome that is subject to missingness. Responses by Robins et al. (2007), Tsiatis and Davidian (2007), Tan (2007a) and Ridgeway and McCaffrey (2007) further explore the challenges faced by double robust …
Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials,
2011
Laboratoire MAP5, Université Paris Descartes and CNRS
Estimation And Testing In Targeted Group Sequential Covariate-Adjusted Randomized Clinical Trials, Antoine Chambaz, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
This article is devoted to the construction and asymptotic study of adaptive group sequential covariate-adjusted randomized clinical trials analyzed through the prism of the semiparametric methodology of targeted maximum likelihood estimation (TMLE). We show how to build, as the data accrue group-sequentially, a sampling design which targets a user-supplied optimal design. We also show how to carry out a sound TMLE statistical inference based on such an adaptive sampling scheme (therefore extending some results known in the i.i.d setting only so far), and how group-sequential testing applies on top of it. The procedure is robust (i.e., consistent even if the …
Subsample Ignorable Likelihood For Accelerated Failure Time Models With Missing Predictors,
2011
University ofSouth Florida
Subsample Ignorable Likelihood For Accelerated Failure Time Models With Missing Predictors, Nanhua Zhang, Roderick J. Little
The University of Michigan Department of Biostatistics Working Paper Series
No abstract provided.
Threshold Regression Models Adapted To Case-Control Studies, And The Risk Of Lung Cancer Due To Occupational Exposure To Asbestos In France,
2011
Laboratoire MAP5, Université Paris Descartes and CNRS
Threshold Regression Models Adapted To Case-Control Studies, And The Risk Of Lung Cancer Due To Occupational Exposure To Asbestos In France, Antoine Chambaz, Dominique Choudat, Catherine Huber, Jean-Claude Pairon, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Asbestos has been known for many years as a powerful carcinogen. Our purpose is quantify the relationship between an occupational exposure to asbestos and an increase of the risk of lung cancer. Furthermore, we wish to tackle the very delicate question of the evaluation, in subjects suffering from a lung cancer, of how much the amount of exposure to asbestos explains the occurrence of the cancer. For this purpose, we rely on a recent French case-control study. We build a large collection of threshold regression models, data-adaptively select a better model in it by multi-fold likelihood-based cross-validation, then fit the …
Efektivitas Dan Efisiensi Sistem Informasi Keluarga Berencana Di Puskesmas,
2011
Direktorat Bina Kesehatan Ibu Direktorat Jenderal Bina Kesehatan Masyarakat Kementerian Kesehatan RI
Efektivitas Dan Efisiensi Sistem Informasi Keluarga Berencana Di Puskesmas, Aragar Putri, Besral Besral
Kesmas
Efisiensi dan efektivitas sistem informasi keluarga berencana (KB)-kesehatan yang telah disosialisasikan sejak tahun 2007 dibandingkan dengan sistem yang lama belum diketahui. Suatu penelitian survei dilakukan di empat provinsi, yaitu DKI Jakarta, Lampung, Kalimatan Tengah, dan Bali. Di tiap provinsi dipilih dua kabupaten/kota dan pada tiap kabupaten/kota dipilih dua puskesmas (kecamatan) yang sudah menerapkan sistem informasi KB-kesehatan tersebut. Pengumpulan data dilakukan pada bulan Juni-September 2008. Penelitian ini menemukan bahwa efektivitas dan efisiensi sistem informasi KB yang baru cukup baik, 77,8% responden menyatakan lebih efektif atau sangat lebih efektif dan 66,7% responden menyatakan lebih efisien atau sangat lebih efisien dibandingkan dengan sistem …
Bate Curve In Assessment Of Clinical Utility Of Predictive Biomarkers,
2011
University of Washington
Bate Curve In Assessment Of Clinical Utility Of Predictive Biomarkers, Xiao-Hua Zhou, Yunbei Ma
UW Biostatistics Working Paper Series
In this paper, for time-to-event data, we propose a new statistical framework for casual inference in evaluating clinical utility of predictive biomarkers and in selecting an optimal treatment for a particular patient. This new casual framework is based on a new concept, called Biomarker Adjusted Treatment Effect (BATE) curve, which can be used to represent the clinical utility of a predictive biomarker and select an optimal treatment for one particular patient. We then propose semi-parametric methods for estimating the BATE curves of biomarkers and establish asymptotic results of the proposed estimators for the BATE curves. We also conduct extensive simulation …
Tmle: An R Package For Targeted Maximum Likelihood Estimation,
2011
Harvard School of Public Health
Tmle: An R Package For Targeted Maximum Likelihood Estimation, Susan Gruber, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Targeted maximum likelihood estimation (TMLE) presents an approach for construction of an efficient double-robust semi-parametric substitution estimator of a target feature of the data generating distribution, such as a statistical association measure or a causal effect parameter. tmle is a recently developed R package that implements TMLE for estimation of the effect of a binary treatment at a single point in time on an outcome of interest, controlling for user supplied covariates: the additive treatment effect, the relative risk, the odds ratio. The package allows outcome data with missingness, and experimental units that contribute repeated records of the point-treatment data …
Causal Inference Under Multiple Versions Of Treatment,
2011
Harvard University
Causal Inference Under Multiple Versions Of Treatment, Tyler J. Vanderweele, Miguel A. Hernan
COBRA Preprint Series
In this article we discuss the no-multiple-versions-of-treatment assumption and extend the potential outcomes framework to accommodate causal inference under violations of this assumption. A variety of examples are discussed in which the assumption may be violated. Identification results are provided for the overall treatment effect and the effect of treatment on the treated when multiple versions of treatment are present and also for the causal effect comparing a version of one treatment to some other version of the same or a different treatment. Further identification and interpretative results are given for cases in which a treatment variable is dichotomized to …
Non-Homogeneous Markov Process Models With Incomplete Observations: Application To A Dementia Disease Study,
2011
University of Washington
Non-Homogeneous Markov Process Models With Incomplete Observations: Application To A Dementia Disease Study, Xiao-Hua Zhou, Baojiang Chen
UW Biostatistics Working Paper Series
Identifying risk factors for transition rates among normal cognition, mildly cognitive impairment, dementia and death in an Alzheimer's disease study is very important. It is known that transition rates among these states are strongly time dependent. While Markov process models are often used to describe these disease progressions, the literature mainly focuses on time homogeneous processes, and limited tools are available for dealing with non-homogeneity. Further, patients may choose when they want to visit the clinics, which creates informative observations. In this paper, we develop methods to deal with non-homogeneous Markov processes through time scale transformation when observation times are …
