Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 121 - 150 of 1108

Full-Text Articles in Statistics and Probability

Addressing Confounding In Predictive Models With An Application To Neuroimaging, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara Sep 2015

Addressing Confounding In Predictive Models With An Application To Neuroimaging, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara

UPenn Biostatistics Working Papers

Understanding structural changes in the brain that are caused by a particular disease is a major goal of neuroimaging research. Multivariate pattern analysis (MVPA) comprises a collection of tools that can be used to understand complex disease effects across the brain. We discuss several important issues that must be considered when analyzing data from neuroimaging studies using MVPA. In particular, we focus on the consequences of confounding by non-imaging variables such as age and sex on the results of MVPA. After reviewing current practice to address confounding in neuroimaging studies, we propose an alternative approach based on inverse probability weighting. …


Control-Group Feature Normalization For Multivariate Pattern Analysis Using The Support Vector Machine, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara Sep 2015

Control-Group Feature Normalization For Multivariate Pattern Analysis Using The Support Vector Machine, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara

UPenn Biostatistics Working Papers

Normalization of feature vector values is a common practice in machine learning. Generally, each feature value is standardized to the unit hypercube or by normalizing to zero mean and unit variance. Classification decisions based on support vector machines (SVMs) or by other methods are sensitive to the specific normalization used on the features. In the context of multivariate pattern analysis using neuroimaging data, standardization effectively up- and down-weights features based on their individual variability. Since the standard approach uses the entire data set to guide the normalization it utilizes the total variability of these features. This total variation is inevitably …


A Simple Method To Estimate The Time-Dependent Roc Curve Under Right Censoring, Liang Li, Bo Hu, Tom Greene Sep 2015

A Simple Method To Estimate The Time-Dependent Roc Curve Under Right Censoring, Liang Li, Bo Hu, Tom Greene

COBRA Preprint Series

The time-dependent Receiver Operating Characteristic (ROC) curve is often used to study the diagnostic accuracy of a single continuous biomarker, measured at baseline, on the onset of a disease condition when the disease onset may occur at different times during the follow-up and hence may be right censored. Due to censoring, the true disease onset status prior to the pre-specified time horizon may be unknown on some patients, which causes difficulty in calculating the time-dependent sensitivity and specificity. We study a simple method that adjusts for censoring by weighting the censored data by the conditional probability of disease onset prior …


On Varieties Of Doubly Robust Estimators Under Missing Not At Random With An Ancillary Variable, Wang Miao, Eric Tchetgen Tchetgen Sep 2015

On Varieties Of Doubly Robust Estimators Under Missing Not At Random With An Ancillary Variable, Wang Miao, Eric Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Pairwise Likelihood Augmented Estimator For The Cox Model Under Left-Truncation, Fan Wu, Sehee Kim, Jing Qin, Rajiv Saran, Yi Li Sep 2015

A Pairwise Likelihood Augmented Estimator For The Cox Model Under Left-Truncation, Fan Wu, Sehee Kim, Jing Qin, Rajiv Saran, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Survival data collected from prevalent cohorts are subject to left-truncation and the analysis is challenging. Conditional approaches for left-truncated data under the Cox model are inefficient as they typically ignore the information in the marginal likelihood of the truncation times. Length-biased sampling methods can improve the estimation efficiency but only when the stationarity assumption of the disease incidence holds, i.e., the truncation distribution is uniform; otherwise they may generate biased estimates. In this paper, we propose a semi-parametric method for the Cox model under general left-truncation, where the truncation distribution is unspecified. Our approach is to make inference based on …


On Partial Identification Of The Pure Direct Effect, Caleb Miles, Phyllis Kanki, Seema Meloni, Eric Tchetgen Tchetgen Sep 2015

On Partial Identification Of The Pure Direct Effect, Caleb Miles, Phyllis Kanki, Seema Meloni, Eric Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


Statistical Estimation Of T1 Relaxation Time Using Conventional Magnetic Resonance Imaging, Amanda Mejia, Elizabeth M. Sweeney, Blake Dewey, Govind Nair, Pascal Sati, Colin Shea, Daniel S. Reich, Russell T. Shinohara Aug 2015

Statistical Estimation Of T1 Relaxation Time Using Conventional Magnetic Resonance Imaging, Amanda Mejia, Elizabeth M. Sweeney, Blake Dewey, Govind Nair, Pascal Sati, Colin Shea, Daniel S. Reich, Russell T. Shinohara

UPenn Biostatistics Working Papers

Quantitative T1 maps estimate T1 relaxation times and can be used to assess diffuse tissue abnormalities within normal-appearing tissue. T1 maps are popular for studying the progression and treatment of multiple sclerosis (MS). However, their inclusion in standard imaging protocols remains limited due to the additional scanning time and expert calibration required and susceptibility to bias and noise. Here, we propose a new method of estimating T1 maps using four conventional MR images, which are intensity- normalized using cerebellar gray matter as a reference tissue and related to T1 using a smooth regression model. Using …


Historical Prediction Modeling Approach For Estimating Long-Term Concentrations Of Pm In Cohort Studies Before The 1999 Implementation Of Widespread Monitoring, Sun-Young Kim, Casey Olives, Lianne Sheppard, Paul D. Sampson, Timothy V. Larson, Joel Kaufman Aug 2015

Historical Prediction Modeling Approach For Estimating Long-Term Concentrations Of Pm In Cohort Studies Before The 1999 Implementation Of Widespread Monitoring, Sun-Young Kim, Casey Olives, Lianne Sheppard, Paul D. Sampson, Timothy V. Larson, Joel Kaufman

UW Biostatistics Working Paper Series

Introduction: Recent cohort studies use exposure prediction models to estimate the association between long-term residential concentrations of PM2.5 and health. Because these prediction models rely on PM2.5 monitoring data, predictions for times before extensive spatial monitoring present a challenge to understanding long-term exposure effects. The Environmental Protection Agency (EPA) Federal Reference Method (FRM) network for PM2.5 was established in 1999. We evaluated a novel statistical approach to produce high quality exposure predictions from 1980-2010 for epidemiological applications.

Methods: We developed spatio-temporal prediction models using geographic predictors and annual average PM2.5 data from 1999 through 2010 from …


C-Learning: A New Classification Framework To Estimate Optimal Dynamic Treatment Regimes, Baqun Zhang, Min Zhang Aug 2015

C-Learning: A New Classification Framework To Estimate Optimal Dynamic Treatment Regimes, Baqun Zhang, Min Zhang

The University of Michigan Department of Biostatistics Working Paper Series

Personalizing treatment to accommodate patient heterogeneity and the evolving nature of a disease over time has received considerable attention lately. A dynamic treatment regime is a set of decision rules, each corresponding to a decision point, that determine that next treatment based on each individual’s own available characteristics and treatment history up to that point. We show that identifying the optimal dynamic treatment regime can be recast as a sequential classification problem and is equivalent to sequentially minimizing a weighted expected misclassification error. This general classification perspective targets the exact goal of optimally individualizing treatments and is new and fundamentally …


On Simple Relations Between Difference-In-Differences And Negative Outcome Control Of Unobserved Confounding, Tamar Sofer, David B. Richardson, Elena Colincino, Joel Schwartz, Eric J. Tchetgen Tchetgen Aug 2015

On Simple Relations Between Difference-In-Differences And Negative Outcome Control Of Unobserved Confounding, Tamar Sofer, David B. Richardson, Elena Colincino, Joel Schwartz, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


Lepski's Method And Adaptive Estimation Of Nonlinear Integral Functionals Of Density, Rajarshi Mukherjee, Eric J. Tchetgen Tchetgen, James M. Robins Aug 2015

Lepski's Method And Adaptive Estimation Of Nonlinear Integral Functionals Of Density, Rajarshi Mukherjee, Eric J. Tchetgen Tchetgen, James M. Robins

Harvard University Biostatistics Working Paper Series

No abstract provided.


Computerizing Efficient Estimation Of A Pathwise Differentiable Target Parameter, Mark J. Van Der Laan, Marco Carone, Alexander R. Luedtke Jul 2015

Computerizing Efficient Estimation Of A Pathwise Differentiable Target Parameter, Mark J. Van Der Laan, Marco Carone, Alexander R. Luedtke

U.C. Berkeley Division of Biostatistics Working Paper Series

Frangakis et al. (2015) proposed a numerical method for computing the efficient influence function of a parameter in a nonparametric model at a specified distribution and observation (provided such an influence function exists). Their approach is based on the assumption that the efficient influence function is given by the directional derivative of the target parameter mapping in the direction of a perturbation of the data distribution defined as the convex line from the data distribution to a pointmass at the observation. In our discussion paper Luedtke et al. (2015) we propose a regularization of this procedure and establish the validity …


Drawing Valid Targeted Inference When Covariate-Adjusted Response-Adaptive Rct Meets Data-Adaptive Loss-Based Estimation, With An Application To The Lasso, Wenjing Zheng, Antoine Chambaz, Mark J. Van Der Laan Jul 2015

Drawing Valid Targeted Inference When Covariate-Adjusted Response-Adaptive Rct Meets Data-Adaptive Loss-Based Estimation, With An Application To The Lasso, Wenjing Zheng, Antoine Chambaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Adaptive clinical trial design methods have garnered growing attention in the recent years, in large part due to their greater flexibility over their traditional counterparts. One such design is the so-called covariate-adjusted, response-adaptive (CARA) randomized controlled trial (RCT). In a CARA RCT, the treatment randomization schemes are allowed to depend on the patient’s pre-treatment covariates, and the investigators have the opportunity to adjust these schemes during the course of the trial based on accruing information (including previous responses), in order to meet a pre-specified optimality criterion, while preserving the validity of the trial in learning its primary study parameter.

In …


Negative Outcome Control For Unobserved Confounding Under A Cox Proportional Hazards Model, Eric J. Tchetgen Tchetgen, Tamar Sofer, David Richardson Jul 2015

Negative Outcome Control For Unobserved Confounding Under A Cox Proportional Hazards Model, Eric J. Tchetgen Tchetgen, Tamar Sofer, David Richardson

Harvard University Biostatistics Working Paper Series

No abstract provided.


Survival Analysis With Functions Of Mis-Measured Covariate Histories: The Case Of Chronic Air Pollution Exposure In Relation To Mortality In The Nurses' Health Study, Xiaomei Liao, Molin Wang, Jaime E. Hart, Francine Laden, Donna Spiegelman Jul 2015

Survival Analysis With Functions Of Mis-Measured Covariate Histories: The Case Of Chronic Air Pollution Exposure In Relation To Mortality In The Nurses' Health Study, Xiaomei Liao, Molin Wang, Jaime E. Hart, Francine Laden, Donna Spiegelman

Harvard University Biostatistics Working Paper Series

Environmental epidemiologists are often interested in estimating the effect of functions of time-varying exposure histories, such as the 12-month moving average, in relation to chronic disease incidence or mortality. The individual exposure measurements that comprise such an exposure history are usually mis-measured, at least moderately, and, often, more substantially. To obtain unbiased estimates of Cox model hazard ratios for these complex mis-measured exposure functions, an extended risk set regression calibration (RRC) method for Cox models is developed and applied to a study of long-term exposure to the fine particulate matter ($PM_{2.5}$) component of air pollution in relation to all-cause mortality …


Doubly Robust Estimation Of A Marginal Average Effect Of Treatment On The Treated With An Instrumental Variable, Lan Liu, Wang Miao, Baoluo Sun, James M. Robins, Eric J. Tchetgen Tchetgen Jun 2015

Doubly Robust Estimation Of A Marginal Average Effect Of Treatment On The Treated With An Instrumental Variable, Lan Liu, Wang Miao, Baoluo Sun, James M. Robins, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


One-Step Targeted Minimum Loss-Based Estimation Based On Universal Least Favorable One-Dimensional Submodels, Mark J. Van Der Laan Jun 2015

One-Step Targeted Minimum Loss-Based Estimation Based On Universal Least Favorable One-Dimensional Submodels, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Consider a study in which one observes n independent and identically distributed random variables whose probability distribution is known to be an element of a particular statistical model, and one is concerned with estimation of a particular real valued pathwise differentiable target parameter of this data probability distribution. The canonical gradient of the pathwise derivative of the target parameter, also called the efficient influence curve, defines an asymptotically efficient estimator as an estimator that is asymptotically linear with influence curve equal to the efficient influence curve.The targeted maximum likelihood estimator is a two stage estimator obtained by constructing a so …


Identification And Doubly Robust Estimation Of Data Missing Not At Random With An Ancillary Variable, Wang Miao, Eric Tchetgen Tchetgen, Zhi Geng Jun 2015

Identification And Doubly Robust Estimation Of Data Missing Not At Random With An Ancillary Variable, Wang Miao, Eric Tchetgen Tchetgen, Zhi Geng

Harvard University Biostatistics Working Paper Series

No abstract provided.


A General Framework For Diagnosing Confounding Of Time-Varying And Other Joint Exposures, John W. Jackson May 2015

A General Framework For Diagnosing Confounding Of Time-Varying And Other Joint Exposures, John W. Jackson

Harvard University Biostatistics Working Paper Series

No abstract provided.


Second Order Inference For The Mean Of A Variable Missing At Random, Ivan Diaz, Marco Carone, Mark J. Van Der Laan May 2015

Second Order Inference For The Mean Of A Variable Missing At Random, Ivan Diaz, Marco Carone, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We present a second order estimator of the mean of a variable subject to missingness, under the missing at random assumption. The estimator improves upon existing methods by using an approximate second order expansion of the parameter functional, in addition to the first order expansion employed by standard doubly robust methods. This results in weaker assumptions about the convergence rates necessary to establish consistency, local efficiency, and asymptotic linearity. The general estimation strategy is developed under the targeted minimum loss based estimation (TMLE) framework. We present a simulation comparing the sensitivity of the first and second order estimators to the …


Adaptive Pre-Specification In Randomized Trials With And Without Pair-Matching, Laura B. Balzer, Mark J. Van Der Laan, Maya L. Petersen May 2015

Adaptive Pre-Specification In Randomized Trials With And Without Pair-Matching, Laura B. Balzer, Mark J. Van Der Laan, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

In randomized trials, adjustment for measured covariates during the analysis can reduce variance and increase power. To avoid misleading inference, the analysis plan must be pre-specified. However, it is unclear a priori which baseline covariates (if any) should be included in the analysis. Consider, for example, the Sustainable East Africa Research in Community Health (SEARCH) trial for HIV prevention and treatment. There are 16 matched pairs of communities and many potential adjustment variables, including region, HIV prevalence, male circumcision coverage and measures of community-level viral load. In this paper, we propose a rigorous procedure to data-adaptively select the adjustment set …


Double Robust Estimation Of Encouragement-Design Intervention Effects Transported Across Sites, Kara E. Rudolph, Mark J. Van Der Laan May 2015

Double Robust Estimation Of Encouragement-Design Intervention Effects Transported Across Sites, Kara E. Rudolph, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We develop double robust targeted maximum likelihood estimators (TMLE) for transporting intervention effects from one population to another. Specifically, we develop TMLE estimators for three transported estimands: intent-to-treat average treatment effect (ATE) and complier ATE, which are relevant for encouragement-design interventions and instrumental variable analyses, and the ATE of the exposure on the outcome, which is applicable to any randomized or observational study. We demonstrate finite sample performance of these TMLE estimators using simulation, including in the presence of practical violations of the positivity assumption. We then apply these methods to the Moving to Opportunity trial, a multi-site, encouragement-design intervention …


Stochastic Optimization Via Forward Slice, Bob A. Salim, Lurdes Y. T. Inoue May 2015

Stochastic Optimization Via Forward Slice, Bob A. Salim, Lurdes Y. T. Inoue

UW Biostatistics Working Paper Series

Optimization consists of maximizing or minimizing a real-valued objective function. In many problems, the objective function may not yield closed-form solutions. Over many decades, optimization methods, both deterministic and stochastic, have been developed to provide solutions to these problems. However, some common limitations of these methods are the sensitivity to the initial value and that often current methods only find a local (non-global) extremum. In this article, we propose an alternative stochastic optimization method, which we call "Forward Slice", and assess its performance relative to available optimization methods.


Distance Correlation Measures Applied To Analyze Relation Between Variables In Liver Cirrhosis Marker Data, Atanu Bhattacharjee Dr. May 2015

Distance Correlation Measures Applied To Analyze Relation Between Variables In Liver Cirrhosis Marker Data, Atanu Bhattacharjee Dr.

COBRA Preprint Series

Distance Correlation is another newer choice to compute the relation between variables. However, the Bayesian counterpart of Distance Correlation is not established. In this paper, Bayesian counterpart of Distance Correlation is pro- posed. Proposed method is illustrated with Liver Chirrhosis Marker data. The relevant studies information about relation between AST and ALT is used to formulate the prior information for Bayesian computation. The Distance Correlation between AST and ALT (both are liver performance marker) is computed with 0.44. The credible interval is observed with (0.41, 0.46).Bayesian counter- part to compute Distance correlation is simple and handy.


Adaptive Enrichment Designs For Randomized Trials With Delayed Endpoints, Using Locally Efficient Estimators To Improve Precision, Michael Rosenblum, Tianchen Qian, Yu Du, Huitong Qiu Apr 2015

Adaptive Enrichment Designs For Randomized Trials With Delayed Endpoints, Using Locally Efficient Estimators To Improve Precision, Michael Rosenblum, Tianchen Qian, Yu Du, Huitong Qiu

Johns Hopkins University, Dept. of Biostatistics Working Papers

Adaptive enrichment designs involve preplanned rules for modifying enrollment criteria based on accrued data in an ongoing trial. For example, enrollment of a subpopulation where there is sufficient evidence of treatment efficacy, futility, or harm could be stopped, while enrollment for the remaining subpopulations is continued. Most existing methods for constructing adaptive enrichment designs are limited to situations where patient outcomes are observed soon after enrollment. This is a major barrier to the use of such designs in practice, since for many diseases the outcome of most clinical importance does not occur shortly after enrollment. We propose a new class …


Simulation Of Semicompeting Risk Survival Data And Estimation Based On Multistate Frailty Model, Fei Jiang, Sebastien Haneuse Apr 2015

Simulation Of Semicompeting Risk Survival Data And Estimation Based On Multistate Frailty Model, Fei Jiang, Sebastien Haneuse

Harvard University Biostatistics Working Paper Series

We develop a simulation procedure to simulate the semicompeting risk survival data. In addition, we introduce an EM algorithm and a B–spline based estimation procedure to evaluate and implement Xu et al. (2010)’s nonparametric likelihood es- timation approach. The simulation procedure provides a route to simulate samples from the likelihood introduced in Xu et al. (2010)’s. Further, the EM algorithm and the B–spline methods stabilize the estimation and gives accurate estimation results. We illustrate the simulation and the estimation procedure with simluation examples and real data analysis.


Targeted Estimation And Inference For The Sample Average Treatment Effect, Laura B. Balzer, Maya L. Petersen, Mark J. Van Der Laan Mar 2015

Targeted Estimation And Inference For The Sample Average Treatment Effect, Laura B. Balzer, Maya L. Petersen, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

While the population average treatment effect has been the subject of extensive methods and applied research, less consideration has been given to the sample average treatment effect: the mean difference in the counterfactual outcomes for the study units. The sample parameter is easily interpretable and is arguably the most relevant when the study units are not representative of a greater population or when the exposure's impact is heterogeneous. Formally, the sample effect is not identifiable from the observed data distribution. Nonetheless, targeted maximum likelihood estimation (TMLE) can provide an asymptotically unbiased and efficient estimate of both the population and sample …


Optimal Dynamic Treatments In Resource-Limited Settings, Alexander R. Luedtke, Mark J. Van Der Laan Jan 2015

Optimal Dynamic Treatments In Resource-Limited Settings, Alexander R. Luedtke, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

A dynamic treatment rule (DTR) is a treatment rule which assigns treatments to individuals based on (a subset of) their measured covariates. An optimal DTR is the DTR which maximizes the population mean outcome. Previous works in this area have assumed that treatment is an unlimited resource so that the entire population can be treated if this strategy maximizes the population mean outcome. We consider optimal DTRs in settings where the treatment resource is limited so that there is a maximum proportion of the population which can be treated. We give a general closed-form expression for an optimal stochastic DTR …


Applying Multiple Imputation For External Calibration To Propensty Score Analysis, Yenny Webb-Vargas, Kara E. Rudolph, D. Lenis, Peter Murakami, Elizabeth A. Stuart Jan 2015

Applying Multiple Imputation For External Calibration To Propensty Score Analysis, Yenny Webb-Vargas, Kara E. Rudolph, D. Lenis, Peter Murakami, Elizabeth A. Stuart

Johns Hopkins University, Dept. of Biostatistics Working Papers

Although covariate measurement error is likely the norm rather than the exception, methods for handling covariate measurement error in propensity score methods have not been widely investigated. We consider a multiple imputation-based approach that uses an external calibration sample with information on the true and mismeasured covariates, Multiple Imputation for External Calibration (MI-EC), to correct for the measurement error, and investigate its performance using simulation studies. As expected, using the covariate measured with error leads to bias in the treatment effect estimate. In contrast, the MI-EC method can eliminate almost all the bias. We confirm that the outcome must be …


Adaptive, Group Sequential Designs That Balance The Benefits And Risks Of Wider Inclusion Criteria, Michael Rosenblum, Brandon S. Luber, Richard E. Thompson, Daniel F. Hanley Jan 2015

Adaptive, Group Sequential Designs That Balance The Benefits And Risks Of Wider Inclusion Criteria, Michael Rosenblum, Brandon S. Luber, Richard E. Thompson, Daniel F. Hanley

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose a new class of adaptive randomized trial designs aimed at gaining the advantages of wider generalizability and faster recruitment, while mitigating the risks of including a population for which there is greater a priori uncertainty. Our designs use adaptive enrichment, i.e., they have preplanned decision rules for modifying enrollment criteria based on data accrued at interim analyses. For example, enrollment can be restricted if the participants from predefined subpopulations are not benefiting from the new treatment. To the best of our knowledge, our designs are the first adaptive enrichment designs to have all of the following features: the …