Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 91 - 120 of 1108

Full-Text Articles in Statistics and Probability

A Weighted Instrumental Variable Estimator To Control For Instrument-Outcome Confounders, Douglas Lehmann, Yun Li, Rajiv Saran, Yi Li Apr 2016

A Weighted Instrumental Variable Estimator To Control For Instrument-Outcome Confounders, Douglas Lehmann, Yun Li, Rajiv Saran, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

No abstract provided.


Recommendation To Use Exact P-Values In Biomarker Discovery Research, Margaret Sullivan Pepe, Matthew F. Buas, Christopher I. Li, Garnet L. Anderson Apr 2016

Recommendation To Use Exact P-Values In Biomarker Discovery Research, Margaret Sullivan Pepe, Matthew F. Buas, Christopher I. Li, Garnet L. Anderson

UW Biostatistics Working Paper Series

Background: In biomarker discovery studies, markers are ranked for validation using P-values. Standard P-value calculations use normal approximations that may not be valid for small P-values and small sample sizes common in discovery research.

Methods: We compared exact P-values, valid by definition, with normal and logit-normal approximations in a simulated study of 40 cases and 160 controls. The key measure of biomarker performance was sensitivity at 90% specificity. Data for 3000 uninformative markers and 30 true markers were generated randomly, with 10 replications of the simulation. We also analyzed real data on 2371 antibody array markers …


One-Step Targeted Minimum Loss-Based Estimation Based On Universal Least Favorable One-Dimensional Submodels, Mark J. Van Der Laan, Susan Gruber Mar 2016

One-Step Targeted Minimum Loss-Based Estimation Based On Universal Least Favorable One-Dimensional Submodels, Mark J. Van Der Laan, Susan Gruber

U.C. Berkeley Division of Biostatistics Working Paper Series

Consider a study in which one observes n independent and identically distributed random variables whose probability distribution is known to be an element of a particular statistical model, and one is concerned with estimation of a particular real valued pathwise differentiable target parameter of this data probability distribution. The targeted maximum likelihood estimator (TMLE) is an asymptotically efficient substitution estimator obtained by constructing a so called least favorable parametric submodel through an initial estimator with score, at zero fluctuation of the initial estimator, that spans the efficient influence curve, and iteratively maximizing the corresponding parametric likelihood till no more updates …


Maximum Likelihood Based Analysis Of Equally Spaced Longitudinal Count Data With Specified Marginal Means, First-Order Antedependence, And Linear Conditional Expectations, Victoria Gamerman, Matthew Guerra, Justine Shults Mar 2016

Maximum Likelihood Based Analysis Of Equally Spaced Longitudinal Count Data With Specified Marginal Means, First-Order Antedependence, And Linear Conditional Expectations, Victoria Gamerman, Matthew Guerra, Justine Shults

UPenn Biostatistics Working Papers

This manuscript implements a maximum likelihood based approach that is appropriate for equally spaced longitudinal count data with over-dispersion, so that the variance of the outcome variable is larger than expected for the assumed Poisson distribution. We implement the proposed method in the analysis of two data sets and make comparisons with the semi-parametric generalized estimating equations (GEE) approach that incorrectly ignores the over-dispersion. The simulations demonstrate that the proposed method has better small sample efficiency than GEE. We also provide code in R that can be used to recreate the analysis results that we provide in this manuscript.


Marginal Structural Models With Counterfactual Effect Modifiers, Wenjing Zheng, Zhehui Luo, Mark J. Van Der Laan Mar 2016

Marginal Structural Models With Counterfactual Effect Modifiers, Wenjing Zheng, Zhehui Luo, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In health and social sciences, research questions often involve systematic assessment of the modification of treatment causal effect by patient characteristics, in longitudinal settings with time-varying or post-intervention effect modifiers of interest. In this work, we investigate the robust and efficient estimation of the so-called Counterfactual-History-Adjusted Marginal Structural Model (van der Laan and Petersen (2007)), which models the conditional intervention-specific mean outcome given modifier history in an ideal experiment where, possible contrary to fact, the subject was assigned the intervention of interest, including the treatment sequence in the conditioning history. We establish the semiparametric efficiency theory for these models, and …


Conditional Screening For Ultra-High Dimensional Covariates With Survival Outcomes, Hyokyoung Grace Hong, Jian Kang, Yi Li Mar 2016

Conditional Screening For Ultra-High Dimensional Covariates With Survival Outcomes, Hyokyoung Grace Hong, Jian Kang, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Identifying important biomarkers that are predictive for cancer patients' prognosis is key in gaining better insights into the biological influences on the disease and has become a critical component of precision medicine. The emergence of large-scale biomedical survival studies, which typically involve excessive number of biomarkers, has brought high demand in designing efficient screening tools for selecting predictive biomarkers. The vast amount of biomarkers defies any existing variable selection methods via regularization. The recently developed variable screening methods, though powerful in many practical setting, fail to incorporate prior information on the importance of each biomarker and are less powerful in …


Simulating Longer Vectors Of Correlated Binary Random Variables Via Multinomial Sampling, Justine Shults Mar 2016

Simulating Longer Vectors Of Correlated Binary Random Variables Via Multinomial Sampling, Justine Shults

UPenn Biostatistics Working Papers

The ability to simulate correlated binary data is important for sample size calculation and comparison of methods for analysis of clustered and longitudinal data with dichotomous outcomes. One available approach for simulating length n vectors of dichotomous random variables is to sample from the multinomial distribution of all possible length n permutations of zeros and ones. However, the multinomial sampling method has only been implemented in general form (without first making restrictive assumptions) for vectors of length 2 and 3, because specifying the multinomial distribution is very challenging for longer vectors. I overcome this difficulty by presenting an algorithm for …


Strengthening Instrumental Variables Through Weighting, Douglas Lehmann, Yun Li, Rajiv Saran, Yi Li Mar 2016

Strengthening Instrumental Variables Through Weighting, Douglas Lehmann, Yun Li, Rajiv Saran, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Instrumental variable (IV) methods are widely used to deal with the issue of unmeasured confounding and are becoming popular in health and medical research. IV models are able to obtain consistent estimates in the presence of unmeasured confounding, but rely on assumptions that are hard to verify and often criticized. An instrument is a variable that influences or encourages individuals toward a particular treatment without directly affecting the outcome. Estimates obtained using instruments with a weak influence over the treatment are known to have larger small-sample bias and to be less robust to the critical IV assumption that the instrument …


Evaluating The Impact Of A Hiv Low-Risk Express Care Task-Shifting Program: A Case Study Of The Targeted Learning Roadmap, Linh Tran, Constantin T. Yiannoutsos, Beverly S. Musick, Kara K. Wools-Kaloustian, Abraham Siika, Sylvester Kimaiyo, Mark J. Van Der Laan, Maya L. Petersen Mar 2016

Evaluating The Impact Of A Hiv Low-Risk Express Care Task-Shifting Program: A Case Study Of The Targeted Learning Roadmap, Linh Tran, Constantin T. Yiannoutsos, Beverly S. Musick, Kara K. Wools-Kaloustian, Abraham Siika, Sylvester Kimaiyo, Mark J. Van Der Laan, Maya L. Petersen

U.C. Berkeley Division of Biostatistics Working Paper Series

In conducting studies on an exposure of interest, a systematic roadmap should be applied for translating causal questions into statistical analyses and interpreting the results. In this paper we describe an application of one such roadmap applied to estimating the joint effect of both time to availability of a nurse-based triage system (low risk express care (LREC)) and individual enrollment in the program among HIV patients in East Africa. Our study population is comprised of 16;513 subjects found eligible for this task-shifting program within 15 clinics in Kenya between 2006 and 2009, with each clinic starting the LREC program between …


Crtgeedr: An R Package For Doubly Robust Generalized Estimating Equations Estimations In Cluster Randomized Trials With Missing Data, Melanie Prague, Rui Wang, Victor De Gruttola Feb 2016

Crtgeedr: An R Package For Doubly Robust Generalized Estimating Equations Estimations In Cluster Randomized Trials With Missing Data, Melanie Prague, Rui Wang, Victor De Gruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang Feb 2016

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


Accounting For Interactions And Complex Inter-Subject Dependency In Estimating Treatment Effect In Cluster Randomized Trials With Missing Outcomes, Melanie Prague, Rui Wang, Alisa Stephens, Eric Tchetgen Tchetgen, Victor Degruttola Jan 2016

Accounting For Interactions And Complex Inter-Subject Dependency In Estimating Treatment Effect In Cluster Randomized Trials With Missing Outcomes, Melanie Prague, Rui Wang, Alisa Stephens, Eric Tchetgen Tchetgen, Victor Degruttola

Harvard University Biostatistics Working Paper Series

No abstract provided.


Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret Jan 2016

Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret

UW Biostatistics Working Paper Series

We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …


An Efficient Basket Trial Design, Kristen Cunanan, Alexia Iasonos, Ronglai Shen, Colin B. Begg, Mithat Gonen Jan 2016

An Efficient Basket Trial Design, Kristen Cunanan, Alexia Iasonos, Ronglai Shen, Colin B. Begg, Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

The landscape for early phase cancer clinical trials is changing dramatically due to the advent of targeted therapy. Increasingly, new drugs are designed to work against a target such as the presence of a specific tumor mutation. Since typically only a small proportion of cancer patients will possess the mutational target, but the mutation is present in many different cancers, a new class of basket trials is emerging, whereby the drug is tested simultaneously in different baskets, i.e., sub-groups of different tumor types. Investigators not only desire to test whether the drug works, but also to determine which types of …


Variable Selection For Case-Cohort Studies With Failure Time Outcome, Andy Ni, Jianwen Cai, Donglin Zeng Jan 2016

Variable Selection For Case-Cohort Studies With Failure Time Outcome, Andy Ni, Jianwen Cai, Donglin Zeng

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Case-cohort designs are widely used in large cohort studies to reduce the cost associated with covariate measurement. In many such studies the number of covariates is very large, so an efficient variable selection method is necessary. In this paper, we study the properties of variable selection using the smoothly clipped absolute deviation penalty in a case-cohort design with a diverging number of parameters. We establish the consistency and asymptotic normality of the maximum penalized pseudo-partial likelihood estimator, and show that the proposed variable selection procedure is consistent and has an asymptotic oracle property. Simulation studies compare the finite sample performance …


Leveraging Contact Network Structure In The Design Of Cluster Randomized Trials, Guy Harling, Rui Wang, Jukka-Pekka Onnela, Victor Degruttola Jan 2016

Leveraging Contact Network Structure In The Design Of Cluster Randomized Trials, Guy Harling, Rui Wang, Jukka-Pekka Onnela, Victor Degruttola

Harvard University Biostatistics Working Paper Series

Background: In settings like the Ebola epidemic, where proof-of-principle trials have succeeded but questions remain about the effectiveness of different possible modes of implementation, it may be useful to develop trials that not only generate information about intervention effects but also themselves provide public health benefit. Cluster randomized trials are of particular value for infectious disease prevention research by virtue of their ability to capture both direct and indirect effects of intervention; the latter of which depends heavily on the nature of contact networks within and across clusters. By leveraging information about these networks – in particular the degree …


Using Validation Data To Adjust The Inverse Probability Weighting Estimator For Misclassified Treatment, Danielle Braun, Corwin Zigler, Francesca Dominici, Malka Gorfine Jan 2016

Using Validation Data To Adjust The Inverse Probability Weighting Estimator For Misclassified Treatment, Danielle Braun, Corwin Zigler, Francesca Dominici, Malka Gorfine

Harvard University Biostatistics Working Paper Series

The inverse probability weighting (IPW) estimator is widely used to estimate the treatment effect in observational studies in which patient characteristics might not be balanced by treatment group. The estimator assumes that treatment assignment, is error-free, but in reality treatment assignment can be measured with error. This arises in the context of comparative effectiveness research, using administrative data sources in which accurate procedural or billing codes are not always available. We show the bias introduced to the estimator when using error-prone treatment assignment, and propose an adjusted estimator using a validation study to eliminate this bias. In simulations, we explore …


Estimation And Inference For The Mediation Proportion, Daniel Nevo, Xiaomei Liao, Donna Spiegelman Jan 2016

Estimation And Inference For The Mediation Proportion, Daniel Nevo, Xiaomei Liao, Donna Spiegelman

Harvard University Biostatistics Working Paper Series

In epidemiology, public health and social science, mediation analysis is often undertaken to investigate the extent to which the effect of a risk factor on an outcome of interest is mediated by other covariates. A pivotal quantity of interest in such an analysis is the mediation proportion. A common method for estimating it, termed the "difference method'', compares estimates from models with and without the hypothesized mediator. However, rigorous methodology for estimation and statistical inference for this quantity has not previously been available. We formulated the problem for the Cox model and generalized linear models, and utilize a data duplication …


A Cautionary Note On The Effect Of Treatment Misclassification On The Average Treatment Effect, Danielle Braun, Corwin Zigler, Malka Gorfine, Francesca Dominici Jan 2016

A Cautionary Note On The Effect Of Treatment Misclassification On The Average Treatment Effect, Danielle Braun, Corwin Zigler, Malka Gorfine, Francesca Dominici

Harvard University Biostatistics Working Paper Series

Comparative effectiveness research often relies on large administrative data, such as claims data. Methods to estimate treatment effects assume that treatment assignment is error-free, but in reality the inaccuracy of procedural or billing codes frequently misclassifies patients into treatment groups. Propensity score methods are widely used to analyze observational studies in which patient characteristics might not be balanced by treatment group. We evaluate the impact of treatment misclassification on 1) propensity score estimation; 2) treatment effect estimation conditional on propensity score estimation and implementation. We focus on three common propensity score implementations: subclassification, matching, and inverse probability of treatment weighting …


The Myth Of Making Inferences For An Overall Treatment Efficacy With Data From Multiple Comparative Studies Via Meta-Analysis, Takahiro Hasegawa, Brian Claggett, Lu Tian, Scott D. Solomon, Marc A. Pfeffer, Lee-Jen Wei Jan 2016

The Myth Of Making Inferences For An Overall Treatment Efficacy With Data From Multiple Comparative Studies Via Meta-Analysis, Takahiro Hasegawa, Brian Claggett, Lu Tian, Scott D. Solomon, Marc A. Pfeffer, Lee-Jen Wei

Harvard University Biostatistics Working Paper Series

Meta analysis techniques, if applied appropriately, can provide a summary of the totality of evidence regarding an overall difference between a new treatment and a control group using data from multiple comparative clinical studies. The standard meta analysis procedures, however, may not give a meaningful between-group difference summary measure or identify a meaningful patient population of interest, especially when the fixed effect model assumption is not met. Moreover, a single between-group comparison measure without a reference value obtained from patients in the control arm would likely not be informative enough for clinical decision making. In this paper, we propose a …


Moving Beyond The Conventional Stratified Analysis To Estimate An Overall Treatment Efficacy With The Data From A Comparative Randomized Clinical Study, Lu Tian, Fei Jiang, Takahiro Hasegawa, Hajime Uno, Marc Alan Pfeffer, L.J. Wei Jan 2016

Moving Beyond The Conventional Stratified Analysis To Estimate An Overall Treatment Efficacy With The Data From A Comparative Randomized Clinical Study, Lu Tian, Fei Jiang, Takahiro Hasegawa, Hajime Uno, Marc Alan Pfeffer, L.J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Semi-Parametrics Dose Finding Methods, Matthieu Clertant, John O'Quigley Jan 2016

Semi-Parametrics Dose Finding Methods, Matthieu Clertant, John O'Quigley

COBRA Preprint Series

We describe a new class of dose finding methods to be used in early phase clinical trials. Under some added parametric conditions the class reduces to the family of continual reassessment method (CRM) designs. Under some relaxation of the underlying structure the method is equivalent to the CCD, mTPI or BOIN classes of designs. These latter designs are non-parametric in nature whereas the CRM class can be viewed as being strongly parametric. The proposed class is characterized as being semi-parametric since it corresponds to CRM with a nuisance parameter. Performance is good, matching that of the CRM class and improving …


Statistical Handling Of Medical Data - An Ethical Perspective, Ajay Kumar Bansal Dr Dec 2015

Statistical Handling Of Medical Data - An Ethical Perspective, Ajay Kumar Bansal Dr

COBRA Preprint Series

Medical Science is a delicate subject and the clinical data generated from the medical trials must be reliable and of good quality. Not only the quality of generated data is important, but the management is also crucial and is to be handled very carefully. In this paper, the ethical aspect of statistical handling of such data is discussed.

Every profession has some set of norms to follow to achieve its objectives. These norms are called professional ethics which shows the essence of human behaviour. Same way, the field of medical research is expected to follow ethical norms, to obtain reliable …


Semi-Parametric Estimation And Inference For The Mean Outcome Of The Single Time-Point Intervention In A Causally Connected Population, Oleg Sofrygin, Mark J. Van Der Laan Dec 2015

Semi-Parametric Estimation And Inference For The Mean Outcome Of The Single Time-Point Intervention In A Causally Connected Population, Oleg Sofrygin, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We study the framework for semi-parametric estimation and statistical inference for the sample average treatment-specific mean effects in observational settings where data are collected on a single network of connected units (e.g., in the presence of interference or spillover). Despite recent advances, many of the current statistical methods rely on estimation techniques that assume a particular parametric model for the outcome, even though some of the most important statistical assumptions required by these models are most likely violated in the observational network settings, often resulting in invalid and anti-conservative statistical inference. In this manuscript, we rely on the recent methodological …


Statistical Estimation Of White Matter Microstructure From Conventional Mri, Leah Suttner, Amanda Mejia, Blake Dewey, Pascal Sati, Daniel S. Reich, Russell T. Shinohara Dec 2015

Statistical Estimation Of White Matter Microstructure From Conventional Mri, Leah Suttner, Amanda Mejia, Blake Dewey, Pascal Sati, Daniel S. Reich, Russell T. Shinohara

UPenn Biostatistics Working Papers

Diffusion tensor imaging (DTI) has become the predominant modality for studying white matter integrity in multiple sclerosis (MS) and other neurological disorders. Unfortunately, the use of DTI-based biomarkers in large multi-center studies is hindered by systematic biases that confound the study of disease-related changes. Furthermore, the site-to-site variability in multi-center studies is significantly higher for DTI than that for conventional MRI-based markers. In our study, we apply the Quantitative MR Estimation Employing Normalization (QuEEN) model to estimate the four DTI measures: MD, FA, RD, and AD. QuEEN uses a voxel-wise generalized additive regression model to relate the normalized intensities of …


A Generally Efficient Targeted Minimum Loss Based Estimator, Mark J. Van Der Laan Dec 2015

A Generally Efficient Targeted Minimum Loss Based Estimator, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose we observe n independent and identically distributed observations of a finite dimensional bounded random variable. This article is concerned with the construction of an efficient targeted minimum loss-based estimator (TMLE) of a pathwise differentiable target parameter based on a realistic statistical model.

The canonical gradient of the target parameter at a particular data distribution will depend on the data distribution through an infinite dimensional nuisance parameter which can be defined as the minimizer of the expectation of a loss function (e.g., log-likelihood loss). For many models and target parameters the nuisance parameter can be split up in two components, …


Inequality In Treatment Benefits: Can We Determine If A New Treatment Benefits The Many Or The Few?, Emily Huang, Ethan Fang, Daniel Hanley, Michael Rosenblum Dec 2015

Inequality In Treatment Benefits: Can We Determine If A New Treatment Benefits The Many Or The Few?, Emily Huang, Ethan Fang, Daniel Hanley, Michael Rosenblum

Johns Hopkins University, Dept. of Biostatistics Working Papers

The primary analysis in many randomized controlled trials focuses on the average treatment effect and does not address whether treatment benefits are widespread or limited to a select few. This problem affects many disease areas, since it stems from how randomized trials, often the gold standard for evaluating treatments, are designed and analyzed. Our goal is to learn about the fraction who benefit from a treatment, based on randomized trial data. We consider the case where the outcome is ordinal, with binary outcomes as a special case. In general, the fraction who benefit is a non-identifiable parameter, and the best …


Meta-Analysis Of Genome-Wide Association Studies With Correlated Individuals: Application To The Hispanic Community Health Study/Study Of Latinos (Hchs/Sol), Tamar Sofer, John R. Shaffer, Misa Graff, Qibin Qi, Adrienne M. Stilp, Stephanie M. Gogarten, Kari E. North, Carmen R. Isasi, Cathy C. Laurie, Adam A. Szpiro Nov 2015

Meta-Analysis Of Genome-Wide Association Studies With Correlated Individuals: Application To The Hispanic Community Health Study/Study Of Latinos (Hchs/Sol), Tamar Sofer, John R. Shaffer, Misa Graff, Qibin Qi, Adrienne M. Stilp, Stephanie M. Gogarten, Kari E. North, Carmen R. Isasi, Cathy C. Laurie, Adam A. Szpiro

UW Biostatistics Working Paper Series

Investigators often meta-analyze multiple genome-wide association studies (GWASs) to increase the power to detect associations of single nucleotide polymorphisms (SNPs) with a trait. Meta-analysis is also performed within a single cohort that is stratified by, e.g., sex or ancestry group. Having correlated individuals among the strata may complicate meta-analyses, limit power, and inflate Type 1 error. For example, in the Hispanic Community Health Study/Study of Latinos (HCHS/SOL), sources of correlation include genetic relatedness, shared household, and shared community. We propose a novel mixed-effect model for meta-analysis, “MetaCor", which accounts for correlation between stratum-specific effect estimates. Simulations show that MetaCor controls …


Nested Partially-Latent, Class Models For Dependent Binary Data, Estimating Disease Etiology, Zhenke Wu, Maria Deloria-Knoll, Scott L. Zeger Nov 2015

Nested Partially-Latent, Class Models For Dependent Binary Data, Estimating Disease Etiology, Zhenke Wu, Maria Deloria-Knoll, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

The Pneumonia Etiology Research for Child Health (PERCH) study seeks to use modern measurement technology to infer the causes of pneumonia for which gold-standard evidence is unavailable. The paper describes a latent variable model designed to infer from case-control data the etiology distribution for the population of cases, and for an individual case given his or her measurements. We assume each observation is drawn from a mixture model for which each component represents one cause or disease class. The model addresses a major limitation of the traditional latent class approach by taking account of residual dependence among multivariate binary outcome …


Removing Inter-Subject Technical Variability In Magnetic Resonance Imaging Studies, Jean-Philippe Fortin, Elizabeth M. Sweeney, John Muschelli, Ciprian M. Crainiceanu, Russell T. Shinohara, Alzheimer’S Disease Neuroimaging Initiative Oct 2015

Removing Inter-Subject Technical Variability In Magnetic Resonance Imaging Studies, Jean-Philippe Fortin, Elizabeth M. Sweeney, John Muschelli, Ciprian M. Crainiceanu, Russell T. Shinohara, Alzheimer’S Disease Neuroimaging Initiative

UPenn Biostatistics Working Papers

Magnetic resonance imaging (MRI) intensities are acquired in arbitrary units, making scans non-comparable across sites and between subjects. Intensity normalization is a first step for the improvement of comparability of the images across subjects. However, we show that unwanted inter-scan variability associated with imaging site, scanner effect and other technical artifacts is still present after standard intensity normalization in large multi-site neuroimaging studies. We propose RAVEL (Removal of Artificial Voxel Effect by Linear regression), a tool to remove residual technical variability after intensity normalization. As proposed by SVA and RUV [Leek and Storey, 2007, …