Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Johns Hopkins University, Dept. of Biostatistics Working Papers

Discipline
Keyword
Publication Year

Articles 61 - 90 of 178

Full-Text Articles in Statistics and Probability

Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan Mar 2011

Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan

Johns Hopkins University, Dept. of Biostatistics Working Papers

We present a brief overview of targeted maximum likelihood for estimating the causal effect of a single time point treatment and of a two time point treatment. We focus on simple examples demonstrating how to apply the methodology developed in (van der Laan and Rubin, 2006; Moore and van der Laan, 2007; van der Laan, 2010a,b). We include R code for the single time point case.


Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett Feb 2011

Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett

Johns Hopkins University, Dept. of Biostatistics Working Papers

In this manuscript, we use a two-stage decomposition for the analysis of func- tional magnetic resonance imaging (fMRI). In the first stage, spatial independent component analysis is applied to the group fMRI data to obtain common brain networks (spatial maps) and subject-specific mixing matrices (time courses). In the second stage, functional principal component analysis is utilized to decompose the mixing matrices into population- level eigenvectors and subject-specific loadings. Inference is performed using permutation-based exact conditional logistic regression for matched pairs data. Simulation studies suggest the ability of the decomposition methods to recover population brain networks and the major direction of …


Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu Jan 2011

Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We establish a fundamental equivalence between singular value decomposition (SVD) and functional principal components analysis (FPCA) models. The constructive relationship allows to deploy the numerical efficiency of SVD to fully estimate the components of FPCA, even for extremely high-dimensional functional objects, such as brain images. As an example, a functional mixed effect model is fitted to high-resolution morphometric (RAVENS) images. The main directions of morphometric variation in brain volumes are identified and discussed.


Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert Oct 2010

Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose a novel class of models for functional data exhibiting skewness or other shape characteristics that vary with spatial or temporal location. We use copulas so that the marginal distributions and the dependence structure can be modeled independently. Dependence is modeled with a Gaussian or t-copula, so that there is an underlying latent Gaussian process. We model the marginal distributions using the skew t family. The mean, variance, and shape parameters are modeled nonparametrically as functions of location. A computationally tractable inferential framework for estimating heterogeneous asymmetric or heavy-tailed marginal distributions is introduced. This framework provides a new set …


Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov Oct 2010

Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov

Johns Hopkins University, Dept. of Biostatistics Working Papers

Images, often stored in multidimensional arrays are fast becoming ubiquitous in medical and public health research. Analyzing populations of images is a statistical problem that raises a host of daunting challenges. The most severe challenge is that data sets incorporating images recorded for hundreds or thousands of subjects at multiple visits are massive. We introduce the population value decomposition (PVD), a general method for simultaneous dimensionality reduction of large populations of massive images. We show how PVD can seamlessly be incorporated into statistical modeling and lead to a new, transparent and fast inferential framework. Our methodology was motivated by and …


Multilevel Functional Principal Component Analysis For High-Dimensional Data, Vadim Zipunnikov, Brian Caffo, Ciprian Crainiceanu, David M. Yousem, Christos Davatzikos, Brian S. Schwartz Oct 2010

Multilevel Functional Principal Component Analysis For High-Dimensional Data, Vadim Zipunnikov, Brian Caffo, Ciprian Crainiceanu, David M. Yousem, Christos Davatzikos, Brian S. Schwartz

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose fast and scalable statistical methods for the analysis of hundreds or thousands of high dimensional vectors observed at multiple visits. The proposed inferential methods avoid the difficult task of loading the entire data set at once in the computer memory and use sequential access to data. This allows deployment of our methodology on low-resource computers where computations can be done in minutes on extremely large data sets. Our methods are motivated by and applied to a study where hundreds of subjects were scanned using Magnetic Resonance Imaging (MRI) at two visits roughly five years apart. The original data …


Estimating Temporal Associations In Electrocorticographic (Ecog) Time Series With First Order Pruning, Haley Hedlin, Dana Boatman, Brian Caffo Sep 2010

Estimating Temporal Associations In Electrocorticographic (Ecog) Time Series With First Order Pruning, Haley Hedlin, Dana Boatman, Brian Caffo

Johns Hopkins University, Dept. of Biostatistics Working Papers

Granger causality (GC) is a statistical technique used to estimate temporal associations in multivariate time series. Many applications and extensions of GC have been proposed since its formulation by Granger in 1969. Here we control for potentially mediating or confounding associations between time series in the context of event-related electrocorticographic (ECoG) time series. A pruning approach to remove spurious connections and simultaneously reduce the required number of estimations to fit the effective connectivity graph is proposed. Additionally, we consider the potential of adjusted GC applied to independent components as a method to explore temporal relationships between underlying source signals. Both …


Longitudinal Penalized Functional Regression, Jeff Goldsmith, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich Sep 2010

Longitudinal Penalized Functional Regression, Jeff Goldsmith, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose a new regression model and inferential tools for the case when both the outcome and the functional exposures are observed at multiple visits. This data structure is new but increasingly present in applications where functions or images are recorded at multiple times. This raises new inferential challenges that cannot be addressed with current methods and software. Our proposed model generalizes the Generalized Linear Mixed Effects Model (GLMM) by adding functional predictors. Smoothness of the functional coefficients is ensured using roughness penalties estimated by Restricted Maximum Likelihood (REML) in a corresponding mixed effects model. This method is computationally feasible …


Mixed Effect Poisson Log-Linear Models For Clinical And Epidemiological Sleep Hypnogram Data, Bruce J. Swihart, Brian S. Caffo Phd, Ciprian Crainiceanu Phd, Naresh M. Punjabi Phd, Md Aug 2010

Mixed Effect Poisson Log-Linear Models For Clinical And Epidemiological Sleep Hypnogram Data, Bruce J. Swihart, Brian S. Caffo Phd, Ciprian Crainiceanu Phd, Naresh M. Punjabi Phd, Md

Johns Hopkins University, Dept. of Biostatistics Working Papers

Bayesian Poisson log-linear multilevel models scalable to epidemiological studies are proposed to investigate population variability in sleep state transition rates. Hierarchical random effects are used to account for pairings of individuals and repeated measures within those individuals, as comparing diseased to non-diseased subjects while minimizing bias is of importance. Essentially, non-parametric piecewise constant hazards are estimated and smoothed, allowing for time-varying covariates and segment of the night comparisons. The Bayesian Poisson regression is justified through a re-derivation of a classical algebraic likelihood equivalence of Poisson regression with a log(time) offset and survival regression assuming exponentially distributed survival times. Such re-derivation …


A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu Jul 2010

A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

Many seemingly disparate approaches for marginal modeling have been developed in recent years. We demonstrate that many current approaches for marginal modeling of correlated binary outcomes produce likelihoods that are equivalent to the proposed copula-based models herein. These general copula models of underlying latent threshold random variables yield likelihood based models for marginal fixed effects estimation and interpretation in the analysis of correlated binary data. Moreover, we propose a nomenclature and set of model relationships that substantially elucidates the complex area of marginalized models for binary data. A diverse collection of didactic mathematical and numerical examples are given to illustrate …


The Use Of Propensity Scores To Assess The Generalizability Of Results From Randomized Trials, Elizabeth A. Stuart, Stephen R. Cole, Catherine P. Bradshaw, Philip J. Leaf May 2010

The Use Of Propensity Scores To Assess The Generalizability Of Results From Randomized Trials, Elizabeth A. Stuart, Stephen R. Cole, Catherine P. Bradshaw, Philip J. Leaf

Johns Hopkins University, Dept. of Biostatistics Working Papers

Randomized trials remain the most accepted design for estimating the effects of interventions, but they do not necessarily answer a question of primary interest: Will the program be effective in a target population in which it may be implemented? In other words,are the results generalizable? There has been very little statistical research on how to assess the generalizability, or "external validity," of randomized trials. We propose the use of propensity-score-based metrics to quantify the similarity of the participants in a randomized trial and a target population. In this setting the propensity score model predicts participation in the randomized trial, given …


Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang Mar 2010

Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider likelihood ratio tests (LRT) and their modifications for homogeneity in admixture models. The admixture model is a special case of two component mixture model, where one component is indexed by an unknown parameter while the parameter value for the other component is known. It has been widely used in genetic linkage analysis under heterogeneity, in which the kernel distribution is binomial. For such models, it is long recognized that testing for homogeneity is nonstandard and the LRT statistic does not converge to a conventional 2 distribution. In this paper, we investigate the asymptotic behavior of the LRT for …


Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu Mar 2010

Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

The basic observational unit in this paper is a function. Data are assumed to have a natural hierarchy of basic units. A simple example is when functions are recorded at multiple visits for the same subject. Di et al. (2009) proposed Multilevel Functional Principal Component Analysis (MFPCA) for this type of data structure when functions are densely sampled. Here we consider the case when functions are sparsely sampled and may contain as few as 2 or 3 observations per function. As with MFPCA, we exploit the multilevel structure of covariance operators and data reduction induced by the use of principal …


Penalized Functional Regression, Jeff Goldsmith, Jennifer Feder, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich Jan 2010

Penalized Functional Regression, Jeff Goldsmith, Jennifer Feder, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich

Johns Hopkins University, Dept. of Biostatistics Working Papers

We develop fast fitting methods for generalized functional linear models. An undersmooth of the functional predictor is obtained by projecting on a large number of smooth eigenvectors and the coefficient function is estimated using penalized spline regression. Our method can be applied to many functional data designs including functions measured with and without error, sparsely or densely sampled. The methods also extend to the case of multiple functional predictors or functional predictors with a natural multilevel structure. Our approach can be implemented using standard mixed effects software and is computationally fast. Our methodology is motivated by a diffusion tensor imaging …


Regression Adjustment And Stratification By Propensty Score In Treatment Effect Estimation, Jessica A. Myers, Thomas A. Louis Jan 2010

Regression Adjustment And Stratification By Propensty Score In Treatment Effect Estimation, Jessica A. Myers, Thomas A. Louis

Johns Hopkins University, Dept. of Biostatistics Working Papers

Propensity score adjustment of effect estimates in observational studies of treatment is a common technique used to control for bias in treatment assignment. In situations where matching on propensity score is not possible or desirable, regression adjustment and stratification are two options. Regression adjustment is used most often and can be highly efficient, but it can lead to biased results when model assumptions are violated. Validity of the stratification approach depends on fewer model assumptions, but is less efficient than regression adjustment when the regression assumptions hold. To investigate these issues, by simulation we compare stratification and regression adjustments. We …


On The Behaviour Of Marginal And Conditional Akaike Information Criteria In Linear Mixed Models, Sonja Greven, Thomas Kneib Nov 2009

On The Behaviour Of Marginal And Conditional Akaike Information Criteria In Linear Mixed Models, Sonja Greven, Thomas Kneib

Johns Hopkins University, Dept. of Biostatistics Working Papers

In linear mixed models, model selection frequently includes the selection of random effects. Two versions of the Akaike information criterion (AIC) have been used, based either on the marginal or on the conditional distribution. We show that the marginal AIC is no longer an asymptotically unbiased estimator of the Akaike information, and in fact favours smaller models without random effects. For the conditional AIC, we show that ignoring estimation uncertainty in the random effects covariance matrix, as is common practice, induces a bias that leads to the selection of any random effect not predicted to be exactly zero. We derive …


Analyzing Bivariate Survival Data With Interval Sampling And Application To Cancer Epidemiology, Hong Zhu, Mei-Cheng Wang Nov 2009

Analyzing Bivariate Survival Data With Interval Sampling And Application To Cancer Epidemiology, Hong Zhu, Mei-Cheng Wang

Johns Hopkins University, Dept. of Biostatistics Working Papers

In medical follow-up studies, ordered bivariate survival data are frequently encountered when bivariate failure events are used as the outcomes to identify the progression of a disease. In cancer studies interest could be focused on bivariate failure times, for example, time from birth to cancer onset and time from cancer onset to death. This paper considers a sampling scheme where the first failure event (cancer onset) is identified within a calendar time interval, the time of the initiating event (birth) can be retrospectively confirmed, and the occurrence of the second event (death) is observed sub ject to right censoring. To …


Modeling Multilevel Sleep Transitional Data Via Poisson Log-Linear Multilevel Models, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu, Naresh M. Punjabi Nov 2009

Modeling Multilevel Sleep Transitional Data Via Poisson Log-Linear Multilevel Models, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu, Naresh M. Punjabi

Johns Hopkins University, Dept. of Biostatistics Working Papers

This paper proposes Poisson log-linear multilevel models to investigate population variability in sleep state transition rates. We specifically propose a Bayesian Poisson regression model that is more flexible, scalable to larger studies, and easily fit than other attempts in the literature. We further use hierarchical random effects to account for pairings of individuals and repeated measures within those individuals, as comparing diseased to non-diseased subjects while minimizing bias is of epidemiologic importance. We estimate essentially non-parametric piecewise constant hazards and smooth them, and allow for time varying covariates and segment of the night comparisons. The Bayesian Poisson regression is justified …


Bayesian Functional Data Analysis Using Winbugs, Ciprian M. Crainiceanu, A. Jeffrey Goldsmith Nov 2009

Bayesian Functional Data Analysis Using Winbugs, Ciprian M. Crainiceanu, A. Jeffrey Goldsmith

Johns Hopkins University, Dept. of Biostatistics Working Papers

We provide user friendly software for Bayesian analysis of Functional Data Models using WinBUGS 1.4. The excellent properties of Bayesian analysis in this context are due to: 1) dimensionality reduction, which leads to low dimensional projection bases; 2)the mixed model representation of functional models, which provides a modular approach to model extension; and 3) the orthogonality of the principal component bases, which contributes to excellent chain convergence and mixing properties. Our paper provides one more, essential, reason for using Bayesian analysis for Functional models: the existence of software.


Lasagna Plots: A Saucy Alternative To Spaghetti Plots, Bruce Swihart, Brian Caffo, Bryan D. James, Matthew Strand, Brian S. Schwartz, Naresh M. Punjabi Oct 2009

Lasagna Plots: A Saucy Alternative To Spaghetti Plots, Bruce Swihart, Brian Caffo, Bryan D. James, Matthew Strand, Brian S. Schwartz, Naresh M. Punjabi

Johns Hopkins University, Dept. of Biostatistics Working Papers

Longitudinal repeated measures data has often been visualized with spaghetti plots for continuous out- comes. For large datasets, this often leads to over-plotting and consequential obscuring of trends in the data. This is primarily due to overlapping of trajectories. Here, we suggest a framework called lasagna plot ting that constrains the subject-specific trajectories to prevent overlapping and utilizes gradients of color to depict the outcome. Dynamic sorting and visualization is demonstrated as an exploratory data analysis tool. Supplemental material in the form of sample R code additional illustrated examples are available online.


Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry Sep 2009

Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

The DNA of most vertebrates is depleted in CpG dinucleotides; C followed by a G in the 5’ to 3’ direction. CpGs are the target for DNA methylation, a chemical modification of cytosine (C) heritable during cell division and the most well characterized epigenetic mechanism. The remaining CpGs tend to cluster in regions referred to as CpG islands (CGI). Knowing CGI locations is important because they mark functionally relevant epigenetic loci in development and disease. For various mammals, including human, a readily available and widely used list of CGI is available from the UCSC Genome Browser. This list was derived …


Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani Aug 2009

Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce combinatorial mixtures - a flexible class of models for inference on mixture distributions whose component have multidimensional parameters. The key idea is to allow each element of the component-specific parameter vectors to be shared by a subset of other components. This approach allows for mixtures that range from very flexible to very parsimonious, and unifies inference on component-specific parameters with inference on the number of components. We develop Bayesian inference and computation approaches for this class of distributions, and illustrate them in an application. This work was originally motivated by the analysis of cancer subtypes: in terms of …


A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry Jun 2009

A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …


A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry Jun 2009

A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …


A Spatio-Temporal Approach For Estimating Chronic Effects Of Air Pollution, Sonja Greven, Francesca Dominici, Scott L. Zeger Jun 2009

A Spatio-Temporal Approach For Estimating Chronic Effects Of Air Pollution, Sonja Greven, Francesca Dominici, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

Estimating the health risks associated with air pollution exposure is of great importance in public health. In air pollution epidemiology, two study designs have been used mainly. Time series studies estimate acute risk associated with short-term exposure. They compare day-to-day variation of pollution concentrations and mortality rates, and have been criticized for potential confounding by time-varying covariates. Cohort studies estimate chronic effects associated with long-term exposure. They compare long-term average pollution concentrations and time-to-death across cities, and have been criticized for potential confounding by individual risk factors or city-level characteristics.

We propose a new study design and a statistical model, …


A Bayesian Shrinkage Model For Incomplete Longitudinal Binary Data With Application To The Breast Cancer Prevention Trial, C. Wang, M.J. Daniels, Daniel O. Scharfstein, S. Land May 2009

A Bayesian Shrinkage Model For Incomplete Longitudinal Binary Data With Application To The Breast Cancer Prevention Trial, C. Wang, M.J. Daniels, Daniel O. Scharfstein, S. Land

Johns Hopkins University, Dept. of Biostatistics Working Papers

We consider inference in randomized studies, in which repeatedly measured outcomes may be informatively missing due to drop out. In this setting, it is well known that full data estimands are not identified unless unverified assumptions are imposed. We assume a non-future dependence model for the drop-out mechanism and posit an exponential tilt model that links non-identifiable and identifiable distributions. This model is indexed by non-identified parameters, which are assumed to have an informative prior distribution, elicited from subject-matter experts. Under this model, full data estimands are shown to be expressed as functionals of the distribution of the observed data. …


Quantifying Uncertainty In Genotype Calls, Benilton Carvalho, Thomas A. Louis, Rafael A. Irizarry Jan 2009

Quantifying Uncertainty In Genotype Calls, Benilton Carvalho, Thomas A. Louis, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

Genome-wide association studies (GWAS) are used to discover genes underlying complex, heritable disorders for which less powerful study designs have failed in the past. The number of GWAS has skyrocketed recently with findings reported in top journals and the mainstream media. Mircorarrays are the genotype calling technology of choice in GWAS as they permit exploration of more than a million single nucleotide polymorphisms (SNPs)simultaneously. The starting point for the statistical analyses used by GWAS, to determine association between loci and disease, are genotype calls (AA, AB, or BB). However, the raw data, microarray probe intensities, are heavily processed before arriving …


Bayesian Model Averaging For Clustered Data: Imputing Missing Daily Air Pollution Concentration, Howard H. Chang, Francesca Dominici, Roger D. Peng Dec 2008

Bayesian Model Averaging For Clustered Data: Imputing Missing Daily Air Pollution Concentration, Howard H. Chang, Francesca Dominici, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

The presence of missing observations is a challenge in statistical analysis especially when data are clustered. In this paper, we develop a Bayesian model averaging (BMA) approach for imputing missing observations in clustered data. Our approach extends BMA by allowing the weights of competing regression models for missing data imputation to vary between clusters while borrowing information across clusters in estimating model parameters. Through simulation and cross-validation studies, we demonstrate that our approach outperforms the standard BMA imputation approach where model weights are assumed to be the same for all clusters. We then apply our proposed method to a national …


Spatial Misalignment In Time Series Studies Of Air Pollution And Health Data, Roger D. Peng, Michelle L. Bell Dec 2008

Spatial Misalignment In Time Series Studies Of Air Pollution And Health Data, Roger D. Peng, Michelle L. Bell

Johns Hopkins University, Dept. of Biostatistics Working Papers

Time series studies of environmental exposures often involve comparing daily changes in a toxicant measured at a point in space with daily changes in an aggregate measure of health. Spatial misalignment of the exposure and response variables can bias the estimation of health risk and the magnitude of this bias depends on the spatial variation of the exposure of interest. In air pollution epidemiology, there is an increasing focus on estimating the health effects of the chemical components of particulate matter. One issue that is raised by this new focus is the spatial misalignment error introduced by the lack of …


Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche Oct 2008

Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche

Johns Hopkins University, Dept. of Biostatistics Working Papers

Latent class analysis (LCA) and latent class regression (LCR) are widely used for modeling multivariate categorical outcomes in social sciences and biomedical studies. Standard analyses assume data of different respondents to be mutually independent, excluding application of the methods to familial and other designs in which participants are clustered. In this paper, we develop multilevel latent class model, in which subpopulation mixing probabilities are treated as random effects that vary among clusters according to a common Dirichlet distribution. We apply the Expectation-Maximization (EM) algorithm for model fitting by maximum likelihood (ML). This approach works well, but is computationally intensive when …