Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 301 - 330 of 1108

Full-Text Articles in Statistics and Probability

Causal Inference For Networks, Mark J. Van Der Laan Oct 2012

Causal Inference For Networks, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose that we observe a population of causally connected units according to a network. On each unit we observe a set of potentially connected units that contains the true connections, and a longitudinal data structure, which includes time-dependent exposure or treatment, time-dependent covariates, a final outcome of interest. The target quantity of interest is defined as the mean outcome for this group of units if the exposures of the units would be probabilistically assigned according to a known specified mechanism, where the latter is called a stochastic intervention. Causal effects of interest are defined as contrasts of the mean of …


Targeted Learning Of The Probability Of Success Of An In Vitro Fertilization Program Controlling For Time-Dependent Confounders, Antoine Chambaz, Sherri Rose, Jean Bouyer, Mark J. Van Der Laan Oct 2012

Targeted Learning Of The Probability Of Success Of An In Vitro Fertilization Program Controlling For Time-Dependent Confounders, Antoine Chambaz, Sherri Rose, Jean Bouyer, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Infertility is a global public health issue and various treatments are available. In vitro fertilization (IVF) is an increasingly common treatment method, but accurately assessing the success of IVF programs has proven challenging since they consist of multiple cycles. We present a double robust semiparametric method that incorporates machine learning to estimate the probability of success (i.e., delivery resulting from embryo transfer) of a program of at most four IVF cycles in the French Devenir Apr`es Interruption de la FIV (DAIFI) study and several simulation studies, controlling for time-dependent confounders. We find that the probability of success in the DAIFI …


Assessing The Causal Effect Of Policies: An Approach Based On Stochastic Interventions, Iván Díaz, Mark J. Van Der Laan Oct 2012

Assessing The Causal Effect Of Policies: An Approach Based On Stochastic Interventions, Iván Díaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Stochastic interventions are a powerful tool to define parameters that measure the causal effect of a realistic intervention that intends to alter the population distribution of an exposure. In this paper we follow the approach described in D\'iaz and van der Laan (2011) to define and estimate the effect of an intervention that is expected to cause a truncation in the population distribution of the exposure. The observed data parameter that identifies the causal parameter of interest is established, as well as its efficient influence function under the non parametric model. Inverse probability of treatment weighted (IPTW), augmented IPTW and …


Quantifying Alternative Splicing From Paired-End Rna-Sequencing Data, David Rossell, Camille Stephan-Otto Attolini, Manuel Kroiss, Almond Stöcker Sep 2012

Quantifying Alternative Splicing From Paired-End Rna-Sequencing Data, David Rossell, Camille Stephan-Otto Attolini, Manuel Kroiss, Almond Stöcker

COBRA Preprint Series

RNA-sequencing has revolutionized biomedical research and, in particular, our ability to study gene alternative splicing. The problem has important implications for human health, as alternative splicing is involved in malfunctions at the cellular level and multiple diseases. However, the high-dimensional nature of the data and the existence of experimental biases pose serious data analysis challenges. We find that the standard data summaries used to study alternative splicing are severely limited, as they ignore a substantial amount of valuable information. Current data analysis methods are based on such summaries and are hence sub-optimal. Further, they have limited flexibility in accounting for …


Robust Estimation Of Pure/Natural Direct Effects With Mediator Measurement Error, Eric J. Tchetgen Tchetgen, Sheng Hsuan Lin Sep 2012

Robust Estimation Of Pure/Natural Direct Effects With Mediator Measurement Error, Eric J. Tchetgen Tchetgen, Sheng Hsuan Lin

COBRA Preprint Series

Recent developments in causal mediation analysis have offered new notions of direct and indirect effects, that formalize more traditional and informal notions of mediation analysis emanating primarily from the social sciences. The pure or natural direct effect of Robins-Greenland-Pearl quantifies the causal effect of an exposure that is not mediated by a variable on the causal pathway to the outcome, and combines with the natural indirect effect to produce the total causal effect of the exposure. Sufficient conditions for identification of natural direct effects were previously given, that assume certain independencies about potential outcomes, and a rich literature on estimation …


Robust Estimation Of Pure/Natural Direct Effects With Mediator Measurement Error, Eric J. Tchetgen Tchetgen, Sheng Hsuan Lin Sep 2012

Robust Estimation Of Pure/Natural Direct Effects With Mediator Measurement Error, Eric J. Tchetgen Tchetgen, Sheng Hsuan Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


Targeted Learning For Causality And Statistical Analysis In Medical Research, Sherri Rose, Richard J.C.M. Starmans, Mark J. Van Der Laan Aug 2012

Targeted Learning For Causality And Statistical Analysis In Medical Research, Sherri Rose, Richard J.C.M. Starmans, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The authors present the use of targeted learning methods for medical research, prepared as a chapter for the upcoming book "Statistics: Discovering Your Future Power." The targeted learning framework involves the explicit specification of the data, model, and parameter. The estimators are double robust and efficient, and can incorporate machine learning procedures such as the super learner.


Modeling Sleep Fragmentation In Populations Of Sleep Hypnograms, Bruce J. Swihart, Naresh M. Punjabi, Ciprian M. Crainiceanu Aug 2012

Modeling Sleep Fragmentation In Populations Of Sleep Hypnograms, Bruce J. Swihart, Naresh M. Punjabi, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce methods for the analysis of large populations of sleep architectures (hypnograms) that respect the 5-state 20-transition-type structure defined by the American Academy of Sleep Medicine. By applying these methods to the hypnograms of 5598 subjects from the Sleep Heart Health Study we: 1) provide the firrst analysis of sleep hypnogram data of such size and complexity in a community cohort with a 4-level comorbidity; 2) compare 5-state 20-transition-type sleep to 3-state 6-transition-type sleep for a check of feasibility and information-loss; 3) extend current approaches to multivariate survival data analysis to populations of time-to-transition processes; and 4) provide scalable …


A Phase I Bayesian Adaptive Design To Simultaneously Optimize Dose And Schedule Assignments Both Among And Within Patients, Thomas M. Braun, Jin Zhang Aug 2012

A Phase I Bayesian Adaptive Design To Simultaneously Optimize Dose And Schedule Assignments Both Among And Within Patients, Thomas M. Braun, Jin Zhang

The University of Michigan Department of Biostatistics Working Paper Series

In traditional schedule or dose-schedule finding designs, patients are assumed to receive their assigned dose-schedule combination throughout the trial even though the combination may be found to have an undesirable toxicity profile, which contradicts actual clinical practice. Since no systematic approach exists to optimize intra-patient dose-schedule as- signment, we propose a Phase I clinical trial design that extends existing approaches that optimize dose and schedule solely among patients by incorporating adaptive variations to dose-schedule assignments within patients as the study proceeds. Our design is based on a Bayesian non-mixture cure rate model that incorporates multiple administrations each patient receives with …


Fitting And Interpreting Continuous-Time Latent Markov Models For Panel Data, Jane M. Lange, Vladimir N. Minin Aug 2012

Fitting And Interpreting Continuous-Time Latent Markov Models For Panel Data, Jane M. Lange, Vladimir N. Minin

UW Biostatistics Working Paper Series

Multistate models are used to characterize disease processes within an individual. Clinical studies often observe the disease status of individuals at discrete time points, making exact times of transitions between disease states unknown. Such panel data pose considerable modeling challenges. Assuming the disease process progresses according a standard continuous-time Markov chain (CTMC) yields tractable likelihoods, but the assumption of exponential sojourn time distributions is typically unrealistic. More flexible semi-Markov models permit generic sojourn distributions yet yield intractable likelihoods for panel data in the presence of reversible transitions. One attractive alternative is to assume that the disease process is characterized by …


Transitions Among Health States Using 12 Measures Of Successful Aging: Results From The Cardiovascular Health Study, Stephen Thielke, Paula Diehr Aug 2012

Transitions Among Health States Using 12 Measures Of Successful Aging: Results From The Cardiovascular Health Study, Stephen Thielke, Paula Diehr

UW Biostatistics Working Paper Series

Introduction

Successful aging has many dimensions, which may manifest differently in men and women and at different ages. We sought to characterize one-year transitions in 12 measures of successful aging among a large cohort of older adults.

Methods

We analyzed twelve different measures of health in the Cardiovascular Health Study: self-rated health, ADLs, IADLs, depression, cognition, timed walk, number of days spent in bed, number of blocks walked, extremity strength, recent hospitalizations, feelings about life as a whole, and life satisfaction. We dichotomized responses for each variable into “healthy” or “sick”, and estimated the prevalence of the healthy state and …


Flexible Covariate-Adjusted Exact Tests For Randomized Studies, Alisa J. Stephens, Eric J. Tchetgen Tchetgen, Victor De Gruttola Aug 2012

Flexible Covariate-Adjusted Exact Tests For Randomized Studies, Alisa J. Stephens, Eric J. Tchetgen Tchetgen, Victor De Gruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Locally Efficient Estimation Of Marginal Treatment Effects When Outcomes Are Correlated: Is The Prize Worth The Chase?, Alisa J. Stephens, Eric J. Tchetgen Tchetgen, Victor De Gruttola Aug 2012

Locally Efficient Estimation Of Marginal Treatment Effects When Outcomes Are Correlated: Is The Prize Worth The Chase?, Alisa J. Stephens, Eric J. Tchetgen Tchetgen, Victor De Gruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Adaptive Matching In Randomized Trials And Observational Studies, Mark J. Van Der Laan, Laura Balzer, Maya L. Petersen Jul 2012

Adaptive Matching In Randomized Trials And Observational Studies, Mark J. Van Der Laan, Laura Balzer, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

In many randomized and observational studies the allocation of treatment among a sample of n independent and identically distributed units is a function of the covariates of all sampled units. As a result, the treatment labels among the units are possibly dependent, complicating estimation and posing challenges for statistical inference. For example, cluster randomized trials frequently sample communities from some target population, construct matched pairs of communities from those included in the sample based on some metric of similarity in baseline community characteristics, and then randomly allocate a treatment and a control intervention within each matched pair. In this case, …


Analytic Programming With Fmri Data: A Quick-Start Guide For Statisticians Using R, Ani Eloyan, Shanshan Li, John Muschelli, Jim Pekar, Stewart Mostofsky, Brian S. Caffo Jul 2012

Analytic Programming With Fmri Data: A Quick-Start Guide For Statisticians Using R, Ani Eloyan, Shanshan Li, John Muschelli, Jim Pekar, Stewart Mostofsky, Brian S. Caffo

Johns Hopkins University, Dept. of Biostatistics Working Papers

Functional magnetic resonance imaging (fMRI) is a thriving field that plays an important role in medical imaging analysis, biological and neuroscience research and practice. This manuscript gives a didactic introduction to the statistical analysis of fMRI data using the R project along with the relevant R code. The goal is to give tatisticians who would like to pursue research in this area a quick start for programming with fMRI data along with the available data visualization tools.


Causal Mediation In A Survival Setting With Time-Dependent Mediators, Wenjing Zheng, Mark J. Van Der Laan Jun 2012

Causal Mediation In A Survival Setting With Time-Dependent Mediators, Wenjing Zheng, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The effect of an expsore on an outcome of interest is often mediated by intermediate variables. The goal of causal mediation analysis is to evaluate the role of these intermediate variables (mediators) in the causal effect of the exposure on the outcome. In this paper, we consider causal mediation of a baseline exposure on a survival (or time-to-event) outcome, when the mediator is time-dependent. The challenge in this setting lies in that the event process takes places jointly with the mediator process; in particular, the length of the mediator history depends on the survival time. As a result, we argue …


A Prior-Free Framework Of Coherent Inference And Its Derivation Of Simple Shrinkage Estimators, David R. Bickel Jun 2012

A Prior-Free Framework Of Coherent Inference And Its Derivation Of Simple Shrinkage Estimators, David R. Bickel

COBRA Preprint Series

The reasoning behind uses of confidence intervals and p-values in scientific practice may be made coherent by modeling the inferring statistician or scientist as an idealized intelligent agent. With other things equal, such an agent regards a hypothesis coinciding with a confidence interval of a higher confidence level as more certain than a hypothesis coinciding with a confidence interval of a lower confidence level. The agent uses different methods of confidence intervals conditional on what information is available. The coherence requirement means all levels of certainty of hypotheses about the parameter agree with the same distribution of certainty over parameter …


On Identification Of Natural Direct Effects When A Confounder Of The Mediator Is Directly Affected By Exposure, Eric J. Tchetgen Tchetgen, Tyler J. Vanderweele Jun 2012

On Identification Of Natural Direct Effects When A Confounder Of The Mediator Is Directly Affected By Exposure, Eric J. Tchetgen Tchetgen, Tyler J. Vanderweele

COBRA Preprint Series

Natural direct and indirect effects formalize traditional notions of mediation analysis into a rigorous causal framework and have recently received considerable attention in epidemiology and in the social sciences. Sufficient conditions for identification of natural direct effects were formulated by Judea Pearl under a nonparametric structural equations model, which assumes certain independencies between potential outcomes. A common situation in epidemiology is that a confounder of the mediator is affected by the exposure, in which case, natural direct effects fail to be nonparametrically identified without additional assumptions, even under Pearl's nonparametric structural equations model. In this paper, the authors show that …


Confidence Intervals For The Selected Population In Randomized Trials That Adapt The Population Enrolled, Michael Rosenblum May 2012

Confidence Intervals For The Selected Population In Randomized Trials That Adapt The Population Enrolled, Michael Rosenblum

Johns Hopkins University, Dept. of Biostatistics Working Papers

It is a challenge to design randomized trials when it is suspected that a treatment may benefit only certain subsets of the target population. In such situations, trial designs have been proposed that modify the population enrolled based on an interim analysis, in a preplanned manner. For example, if there is early evidence that the treatment only benefits a certain subset of the population, enrollment may then be restricted to this subset. At the end of such a trial, it is desirable to draw inferences about the selected population. We focus on constructing confidence intervals for the average treatment effect …


Why Odds Ratio Estimates Of Gwas Are Almost Always Close To 1.0, Yutaka Yasui May 2012

Why Odds Ratio Estimates Of Gwas Are Almost Always Close To 1.0, Yutaka Yasui

COBRA Preprint Series

“Missing heritability” in genome-wide association studies (GWAS) refers to the seeming inability for GWAS data to capture the great majority of genetic causes of a disease in comparison to the known degree of heritability for the disease, in spite of GWAS’ genome-wide measures of genetic variations. This paper presents a simple mathematical explanation for this phenomenon, assuming that the heritability information exists in GWAS data. Specifically, it focuses on the fact that the great majority of association measures (in the form of odds ratios) from GWAS are consistently close to the value that indicates no association, explains why this occurs, …


Differential Patterns Of Interaction And Gaussian Graphical Models, Masanao Yajima, Donatello Telesca, Yuan Ji, Peter Muller Apr 2012

Differential Patterns Of Interaction And Gaussian Graphical Models, Masanao Yajima, Donatello Telesca, Yuan Ji, Peter Muller

COBRA Preprint Series

We propose a methodological framework to assess heterogeneous patterns of association amongst components of a random vector expressed as a Gaussian directed acyclic graph. The proposed framework is likely to be useful when primary interest focuses on potential contrasts characterizing the association structure between known subgroups of a given sample. We provide inferential frameworks as well as an efficient computational algorithm to fit such a model and illustrate its validity through a simulation. We apply the model to Reverse Phase Protein Array data on Acute Myeloid Leukemia patients to show the contrast of association structure between refractory patients and relapsed …


Automated Diagnoses Of Attention Deficit Hyperactive Disorder Using Magnetic Resonance Imaging, Ani Eloyan, John Muschelli, Mary Beth Nebel, Han Liu, Fang Han, Tuo Zhao, Anita Barber, Suresh Joel, James J. Pekar, Stewart Mostofsky, Brian Caffo Apr 2012

Automated Diagnoses Of Attention Deficit Hyperactive Disorder Using Magnetic Resonance Imaging, Ani Eloyan, John Muschelli, Mary Beth Nebel, Han Liu, Fang Han, Tuo Zhao, Anita Barber, Suresh Joel, James J. Pekar, Stewart Mostofsky, Brian Caffo

Johns Hopkins University, Dept. of Biostatistics Working Papers

Successful automated diagnoses of attention de.cit hyperactive disorder (ADHD) using imaging and functional biomarkers would have fundamental consequences on the public health impact of the disease. In this work, we show results on the predictability of ADHD using imaging biomarkers and discuss the scienti.c and diagnostic impacts of the research. We created a prediction model using the land­mark ADHD 200 data set focusing on resting state functional connectivity (rs-fc) and structural brain imaging. We predicted ADHD status and subtype, obtained by behavioral examination, using imaging data, intelligence quotients and other co­variates. The novel contributions of this manuscript include a thorough …


Testing For Improvement In Prediction Model Performance, Margaret S. Pepe Phd, Kathleen F. Kerr, Gary M. Longton, Zheyu Wang Mar 2012

Testing For Improvement In Prediction Model Performance, Margaret S. Pepe Phd, Kathleen F. Kerr, Gary M. Longton, Zheyu Wang

UW Biostatistics Working Paper Series

New methodology has been proposed in recent years for evaluating the improvement in prediction performance gained by adding a new predictor, Y, to a risk model containing a set of baseline predictors, X, for a binary outcome D. We prove theoretically that null hypotheses concerning no improvement in performance are equivalent to the simple null hypothesis that the coefficient for Y is zero in the risk model, P(D = 1|X, Y ). Therefore, testing for improvement in prediction performance is redundant if Y has already been shown to be a risk factor. We investigate properties of tests through simulation studies, …


Hierarchical Rank Aggregation With Applications To Nanotoxicology, Trina Patel, Donatello Telesca, Robert Rallo, Saji George, Xia Tian, Nel Andre Mar 2012

Hierarchical Rank Aggregation With Applications To Nanotoxicology, Trina Patel, Donatello Telesca, Robert Rallo, Saji George, Xia Tian, Nel Andre

COBRA Preprint Series

The development of high throughput screening (HTS) assays in the field of nanotoxicology provide new opportunities for the hazard assessment and ranking of engineered nanomaterials (ENM). It is often necessary to rank lists of materials based on multiple risk assessment parameters, often aggregated across several measures of toxicity and possibly spanning an array of experimental platforms. Bayesian models coupled with the optimization of loss functions have been shown to provide an effective framework for conducting inference on ranks. In this article we present various loss function based ranking approaches for comparing ENM within experiments and toxicity parameters. Additionally, we propose …


Robustness Of Measures Of Interaction To Unmeasured Confounding, Eric J. Tchetgen Tchetgen, Tyler J. Vanderweele Mar 2012

Robustness Of Measures Of Interaction To Unmeasured Confounding, Eric J. Tchetgen Tchetgen, Tyler J. Vanderweele

Harvard University Biostatistics Working Paper Series

No abstract provided.


Bootstrap-Based Inference On The Difference In The Means Of Two Correlated Functional Processes, Ciprian M. Crainiceanu, Ana-Maria Staicu, Shubankar Ray, Naresh Punjabi Mar 2012

Bootstrap-Based Inference On The Difference In The Means Of Two Correlated Functional Processes, Ciprian M. Crainiceanu, Ana-Maria Staicu, Shubankar Ray, Naresh Punjabi

Johns Hopkins University, Dept. of Biostatistics Working Papers

Nonparametric inference methods on the mean difference between two correlated Functional processes are proposed. We compare methods that: 1) incorporate different levels of smoothing of the mean and covariance; 2) preserve the sampling design; and 3) use parametric and nonparametric estimation of the mean functions. We apply our method to estimating the mean difference between average normalized δ-power of sleep electroencephalograms for 51 subjects with severe sleep apnea and 51 matched controls in the first 4 hours after sleep onset. Data are obtained from the Sleep Heart Health Study (SHHS), the largest community cohort study of sleep. While methods are …


Robustness Of Measures Of Interaction To Unmeasured Confounding, Eric J. Tchetgen Tchetgen, Tyler J. Vanderweele Mar 2012

Robustness Of Measures Of Interaction To Unmeasured Confounding, Eric J. Tchetgen Tchetgen, Tyler J. Vanderweele

COBRA Preprint Series

In this paper, we study the impact of unmeasured confounding on inference about a two-way interaction in a mean regression model with identity, log or logit link function. Necessary and sufficient conditions are established for a two-way interaction to be nonparametrically identified from the observed data, despite unmeasured confounding for the factors defining the interaction. A lung cancer data application illustrates the results.


On A Logistic Mixed Model Formulation Of A Quadratic Exponential Model For Correlated Binary Outcomes, Eric J. Tchetgen Tchetgen Feb 2012

On A Logistic Mixed Model Formulation Of A Quadratic Exponential Model For Correlated Binary Outcomes, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


C2bat: A Novel Method For Association Between Ge- Netic Markers And Multiple Phenotypes, Melissa Naylor, Christoph Lange Feb 2012

C2bat: A Novel Method For Association Between Ge- Netic Markers And Multiple Phenotypes, Melissa Naylor, Christoph Lange

Harvard University Biostatistics Working Paper Series

The purpose of this technical report is to describe a novel method developed to detect association between a genetic marker and multiple phenotypes. In order to obtain a one-degree of freedom test, a generalized principal component approach is suggested that aggregates the information about the genetic effect in the first prin- cipal component, while the remain principal components contain only environment noise. A limited simulation study is done validating the method. For scenarios in which the genetic effect is constant across all measurements and there is no envi- ronmental correlation between the measurements, preliminary results suggest that this method has …


Avoiding Boundary Estimates In Linear Mixed Models Through Weakly Informative Priors, Yeojin Chung, Sophia Rabe-Hesketh, Andrew Gelman, Jingchen Liu, Vincent Dorie Feb 2012

Avoiding Boundary Estimates In Linear Mixed Models Through Weakly Informative Priors, Yeojin Chung, Sophia Rabe-Hesketh, Andrew Gelman, Jingchen Liu, Vincent Dorie

U.C. Berkeley Division of Biostatistics Working Paper Series

Variance parameters in mixed or multilevel models can be difficult to estimate, especially when the number of groups is small. We propose a maximum penalized likelihood approach which is equivalent to estimating variance parameters by their marginal posterior mode, given a weakly informative prior distribution. By choosing the prior from the gamma family with at least 1 degree of freedom, we ensure that the prior density is zero at the boundary and thus the marginal posterior mode of the group-level variance will be positive. The use of a weakly informative prior allows us to stabilize our estimates while remaining faithful …