Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 18 of 18

Full-Text Articles in Biostatistics

Distance-Based Analysis Of Variance For Brain Connectivity, Russell T. Shinohara, Haochang Shou, Marco Carone, Robert Schultz, Birkan Tunc, Drew Parker, Ragini Verma Aug 2016

Distance-Based Analysis Of Variance For Brain Connectivity, Russell T. Shinohara, Haochang Shou, Marco Carone, Robert Schultz, Birkan Tunc, Drew Parker, Ragini Verma

UPenn Biostatistics Working Papers

The field of neuroimaging dedicated to mapping connections in the brain is increasingly being recognized as key for understanding neurodevelopment and pathology. Networks of these connections are quantitatively represented using complex structures including matrices, functions, and graphs, which require specialized statistical techniques for estimation and inference about developmental and disorder-related changes. Unfortunately, classical statistical testing procedures are not well suited to high-dimensional testing problems. In the context of global or regional tests for differences in neuroimaging data, traditional analysis of variance (ANOVA) is not directly applicable without first summarizing the data into univariate or low-dimensional features, a process that may …


Interpretable High-Dimensional Inference Via Score Maximization With An Application In Neuroimaging, Simon N. Vandekar, Philip T. Reiss, Russell T. Shinohara May 2016

Interpretable High-Dimensional Inference Via Score Maximization With An Application In Neuroimaging, Simon N. Vandekar, Philip T. Reiss, Russell T. Shinohara

UPenn Biostatistics Working Papers

In the fields of neuroimaging and genetics a key goal is testing the association of a single outcome with a very high-dimensional imaging or genetic variable. Oftentimes summary measures of the high-dimensional variable are created to sequentially test and localize the association with the outcome. In some cases, the results for summary measures are significant, but subsequent tests used to localize differences are underpowered and do not identify regions associated with the outcome. We propose a generalization of Rao's score test based on maximizing the score statistic in a linear subspace of the parameter space. If the test rejects the …


Maximum Likelihood Based Analysis Of Equally Spaced Longitudinal Count Data With Specified Marginal Means, First-Order Antedependence, And Linear Conditional Expectations, Victoria Gamerman, Matthew Guerra, Justine Shults Mar 2016

Maximum Likelihood Based Analysis Of Equally Spaced Longitudinal Count Data With Specified Marginal Means, First-Order Antedependence, And Linear Conditional Expectations, Victoria Gamerman, Matthew Guerra, Justine Shults

UPenn Biostatistics Working Papers

This manuscript implements a maximum likelihood based approach that is appropriate for equally spaced longitudinal count data with over-dispersion, so that the variance of the outcome variable is larger than expected for the assumed Poisson distribution. We implement the proposed method in the analysis of two data sets and make comparisons with the semi-parametric generalized estimating equations (GEE) approach that incorrectly ignores the over-dispersion. The simulations demonstrate that the proposed method has better small sample efficiency than GEE. We also provide code in R that can be used to recreate the analysis results that we provide in this manuscript.


Simulating Longer Vectors Of Correlated Binary Random Variables Via Multinomial Sampling, Justine Shults Mar 2016

Simulating Longer Vectors Of Correlated Binary Random Variables Via Multinomial Sampling, Justine Shults

UPenn Biostatistics Working Papers

The ability to simulate correlated binary data is important for sample size calculation and comparison of methods for analysis of clustered and longitudinal data with dichotomous outcomes. One available approach for simulating length n vectors of dichotomous random variables is to sample from the multinomial distribution of all possible length n permutations of zeros and ones. However, the multinomial sampling method has only been implemented in general form (without first making restrictive assumptions) for vectors of length 2 and 3, because specifying the multinomial distribution is very challenging for longer vectors. I overcome this difficulty by presenting an algorithm for …


Statistical Estimation Of White Matter Microstructure From Conventional Mri, Leah Suttner, Amanda Mejia, Blake Dewey, Pascal Sati, Daniel S. Reich, Russell T. Shinohara Dec 2015

Statistical Estimation Of White Matter Microstructure From Conventional Mri, Leah Suttner, Amanda Mejia, Blake Dewey, Pascal Sati, Daniel S. Reich, Russell T. Shinohara

UPenn Biostatistics Working Papers

Diffusion tensor imaging (DTI) has become the predominant modality for studying white matter integrity in multiple sclerosis (MS) and other neurological disorders. Unfortunately, the use of DTI-based biomarkers in large multi-center studies is hindered by systematic biases that confound the study of disease-related changes. Furthermore, the site-to-site variability in multi-center studies is significantly higher for DTI than that for conventional MRI-based markers. In our study, we apply the Quantitative MR Estimation Employing Normalization (QuEEN) model to estimate the four DTI measures: MD, FA, RD, and AD. QuEEN uses a voxel-wise generalized additive regression model to relate the normalized intensities of …


Removing Inter-Subject Technical Variability In Magnetic Resonance Imaging Studies, Jean-Philippe Fortin, Elizabeth M. Sweeney, John Muschelli, Ciprian M. Crainiceanu, Russell T. Shinohara, Alzheimer’S Disease Neuroimaging Initiative Oct 2015

Removing Inter-Subject Technical Variability In Magnetic Resonance Imaging Studies, Jean-Philippe Fortin, Elizabeth M. Sweeney, John Muschelli, Ciprian M. Crainiceanu, Russell T. Shinohara, Alzheimer’S Disease Neuroimaging Initiative

UPenn Biostatistics Working Papers

Magnetic resonance imaging (MRI) intensities are acquired in arbitrary units, making scans non-comparable across sites and between subjects. Intensity normalization is a first step for the improvement of comparability of the images across subjects. However, we show that unwanted inter-scan variability associated with imaging site, scanner effect and other technical artifacts is still present after standard intensity normalization in large multi-site neuroimaging studies. We propose RAVEL (Removal of Artificial Voxel Effect by Linear regression), a tool to remove residual technical variability after intensity normalization. As proposed by SVA and RUV [Leek and Storey, 2007, …


Addressing Confounding In Predictive Models With An Application To Neuroimaging, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara Sep 2015

Addressing Confounding In Predictive Models With An Application To Neuroimaging, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara

UPenn Biostatistics Working Papers

Understanding structural changes in the brain that are caused by a particular disease is a major goal of neuroimaging research. Multivariate pattern analysis (MVPA) comprises a collection of tools that can be used to understand complex disease effects across the brain. We discuss several important issues that must be considered when analyzing data from neuroimaging studies using MVPA. In particular, we focus on the consequences of confounding by non-imaging variables such as age and sex on the results of MVPA. After reviewing current practice to address confounding in neuroimaging studies, we propose an alternative approach based on inverse probability weighting. …


Control-Group Feature Normalization For Multivariate Pattern Analysis Using The Support Vector Machine, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara Sep 2015

Control-Group Feature Normalization For Multivariate Pattern Analysis Using The Support Vector Machine, Kristin A. Linn, Bilwaj Gaonkar, Jimit Doshi, Christos Davatzikos, Russell T. Shinohara

UPenn Biostatistics Working Papers

Normalization of feature vector values is a common practice in machine learning. Generally, each feature value is standardized to the unit hypercube or by normalizing to zero mean and unit variance. Classification decisions based on support vector machines (SVMs) or by other methods are sensitive to the specific normalization used on the features. In the context of multivariate pattern analysis using neuroimaging data, standardization effectively up- and down-weights features based on their individual variability. Since the standard approach uses the entire data set to guide the normalization it utilizes the total variability of these features. This total variation is inevitably …


Statistical Estimation Of T1 Relaxation Time Using Conventional Magnetic Resonance Imaging, Amanda Mejia, Elizabeth M. Sweeney, Blake Dewey, Govind Nair, Pascal Sati, Colin Shea, Daniel S. Reich, Russell T. Shinohara Aug 2015

Statistical Estimation Of T1 Relaxation Time Using Conventional Magnetic Resonance Imaging, Amanda Mejia, Elizabeth M. Sweeney, Blake Dewey, Govind Nair, Pascal Sati, Colin Shea, Daniel S. Reich, Russell T. Shinohara

UPenn Biostatistics Working Papers

Quantitative T1 maps estimate T1 relaxation times and can be used to assess diffuse tissue abnormalities within normal-appearing tissue. T1 maps are popular for studying the progression and treatment of multiple sclerosis (MS). However, their inclusion in standard imaging protocols remains limited due to the additional scanning time and expert calibration required and susceptibility to bias and noise. Here, we propose a new method of estimating T1 maps using four conventional MR images, which are intensity- normalized using cerebellar gray matter as a reference tissue and related to T1 using a smooth regression model. Using …


Regression Modeling Of Longitudinal Binary Outcomes With Outcome-Dependent Observation Times, Kay See Tan, Andrea B. Troxel, Stephen E. Kimmel, Kevin G. Volpp, Benjamin French Feb 2014

Regression Modeling Of Longitudinal Binary Outcomes With Outcome-Dependent Observation Times, Kay See Tan, Andrea B. Troxel, Stephen E. Kimmel, Kevin G. Volpp, Benjamin French

UPenn Biostatistics Working Papers

Conventional longitudinal data analysis methods assume that outcomes are independent of the data-collection schedule. However, the independence assumption may be violated, for example, when adverse events trigger additional physician visits in between prescheduled follow-ups. Observation times may therefore be associated with outcome values, which may introduce bias when estimating the eect of covariates on outcomes using standard longitudinal regression methods. Existing semi-parametric methods that accommodate outcome-dependent observation times are limited to the analysis of continuous outcomes. We develop new methods for the analysis of binary outcomes, while retaining the exibility of semi-parametric models. Our methods are based on counting process …


Normalization Techniques For Statistical Inference From Magnetic Resonance Imaging, Russell T. Shinohara, Elizabeth M. Sweeney, Jeff Goldsmith, Navid Shiee, Farrah J. Mateen, Peter A. Calabresi, Samson Jarso, Dzung L. Pham, Daniel S. Reich, Ciprian M. Crainiceanu Aug 2013

Normalization Techniques For Statistical Inference From Magnetic Resonance Imaging, Russell T. Shinohara, Elizabeth M. Sweeney, Jeff Goldsmith, Navid Shiee, Farrah J. Mateen, Peter A. Calabresi, Samson Jarso, Dzung L. Pham, Daniel S. Reich, Ciprian M. Crainiceanu

UPenn Biostatistics Working Papers

While computed tomography and other imaging techniques are measured in absolute units with physical meaning, magnetic resonance images are expressed in arbitrary units that are difficult to interpret and differ between study visits and subjects. Much work in the image processing literature on intensity normalization has focused on histogram matching and other histogram mapping techniques, with little emphasis on normalizing images to have biologically interpretable units. Furthermore, there are no formalized principles or goals for the crucial comparability of image intensities within and across subjects. To address this, we propose a set of criteria necessary for the normalization of images. …


On The Simulation Of Longitudinal Discrete Data With Specified Marginal Means And First-Order Antedependence, Matthew Guerra, Justine Shults Jan 2013

On The Simulation Of Longitudinal Discrete Data With Specified Marginal Means And First-Order Antedependence, Matthew Guerra, Justine Shults

UPenn Biostatistics Working Papers

We propose a straightforward approach for simulation of discrete random variables with overdispersion, specified marginal means, and product correlations that are plausible for longitudinal data with equal, or unequal, temporal spacings. The method stems from results we prove for variables with first-order antedependence and linearity of the conditional expectations. The proposed approach will be especially useful for assessment of methods such as generalized estimating equations, which specify separate models for the marginal means and correlation structure of measurements on a subject.


"Implementation Of Quasi-Least Squares With The R Package Qlspack", Jichun Xie, Justine Shults Jun 2009

"Implementation Of Quasi-Least Squares With The R Package Qlspack", Jichun Xie, Justine Shults

UPenn Biostatistics Working Papers

Quasi-least squares (QLS) is an alternative method for estimating the correlation parameters within the framework of generalized estimating equations (GEE) that has two main advantages over the moment estimates that are typically applied for GEE: (1) It guarantees a consistent estimate of the correlation parameter and a positive definite estimated correlation matrix, for several correlation structures; and (2) It allows for easier implementation of some correlation structures that have not yet been implemented in the framework of GEE. Furthermore, because QLS is a method in the framework of GEE, existing software can be employed within the QLS algorithm for estimation …


Analysis Of Adverse Events In Drug Safety: A Multivariate Approach Using Stratified Quasi-Least Squares, Hanjoo Kim, Justine Shults, Scott Patterson, Robert Goldberg-Alberts Dec 2008

Analysis Of Adverse Events In Drug Safety: A Multivariate Approach Using Stratified Quasi-Least Squares, Hanjoo Kim, Justine Shults, Scott Patterson, Robert Goldberg-Alberts

UPenn Biostatistics Working Papers

Safety assessment in drug development involves numerous statistical challenges, and yet statistical methodologies and their applications to safety data have not been fully developed, despite a recent increase of interest in this area. In practice, a conventional univariate approach for analysis of safety data involves application of the Fisher's exact test to compare the proportion of subjects who experience adverse events (AEs) between treatment groups; This approach ignores several common features of safety data, including the presence of multiple endpoints, longitudinal follow-up, and a possible relationship between the AEs within body systems. In this article, we propose various regression modeling …


Variable Selection For Nonparametric Varying-Coefficient Models For Analysis Of Repeated Measurements, Lifeng Wang, Hongzhe Li Jul 2007

Variable Selection For Nonparametric Varying-Coefficient Models For Analysis Of Repeated Measurements, Lifeng Wang, Hongzhe Li

UPenn Biostatistics Working Papers

Nonparametric varying-coefficient models are commonly used for analysis of data measured repeatedly over time, including longitudinal and functional responses data. While many procedures have been developed for estimating the varying-coefficients, the problem of variable selection for such models has not been addressed. In this article, we present a regularized estimation procedure for variable selection for such nonparametric varying-coefficient models using basis function approximations and a group smoothly clipped absolute deviation penalty (gSCAD). This gSCAD procedure simultaneously selects significant variables with time-varying effects and estimates unknown smooth functions using basis function approximations. With appropriate selection of the tuning parameters, we have …


Analysis Of Multi-Level Correlated Data In The Framework Of Generalized Estimating Equations Via Xtmultcorr Procedures In Stata And Qls Functions In Matlab, Justine Shults, Sarah J. Ratcliffe Jan 2007

Analysis Of Multi-Level Correlated Data In The Framework Of Generalized Estimating Equations Via Xtmultcorr Procedures In Stata And Qls Functions In Matlab, Justine Shults, Sarah J. Ratcliffe

UPenn Biostatistics Working Papers

No abstract provided.


Conditional Likelihood Methods For Haplotype-Based Association Analysis Using Matched Case-Control Data, Jinbo Chen, Carmen Rodriguez Sep 2006

Conditional Likelihood Methods For Haplotype-Based Association Analysis Using Matched Case-Control Data, Jinbo Chen, Carmen Rodriguez

UPenn Biostatistics Working Papers

Genetic epidemiologists routinely assess disease susceptibility in relation to haplotypes, i.e., combinations of alleles on a single chromosome. We study statistical methods for inferring haplotype-related disease risk using SNP genotype data from matched case-control studies, where controls are individually matched to cases on some selected factors. Assuming a logistic regression model for haplotype-disease association, we propose two conditional likelihood approaches that address the issue that haplotypes cannot be inferred with certainty from SNP genotype data (phase ambiquity). One approach is based on the likelihood of disease status conditioned on the total number of cases, genotypes, and other covariates within each …


Improved Generalized Estimating Equation Analysis Via Xtqls For Implementation Of Quasi-Least Squares In Stata, Justine Shults, Sarah J. Ratcliffe, Mary Leonard Aug 2006

Improved Generalized Estimating Equation Analysis Via Xtqls For Implementation Of Quasi-Least Squares In Stata, Justine Shults, Sarah J. Ratcliffe, Mary Leonard

UPenn Biostatistics Working Papers

No abstract provided.