Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (69)
- Statistical Methodology (59)
- Statistical Models (47)
- Statistical Theory (43)
- Medicine and Health Sciences (25)
-
- Public Health (21)
- Longitudinal Data Analysis and Time Series (15)
- Epidemiology (13)
- Applied Mathematics (10)
- Microarrays (10)
- Numerical Analysis and Computation (10)
- Survival Analysis (10)
- Multivariate Analysis (9)
- Categorical Data Analysis (7)
- Life Sciences (7)
- Clinical Trials (6)
- Disease Modeling (6)
- Diseases (6)
- Genetics and Genomics (6)
- Health Services Research (6)
- Design of Experiments and Sample Surveys (4)
- Genetics (4)
- Medical Specialties (4)
- Bioinformatics (2)
- Computational Biology (2)
- Infectious Disease (2)
- International Public Health (2)
- Pediatrics (2)
- Keyword
-
- Functional data analysis (5)
- Air pollution (4)
- Longitudinal data (4)
- Penalized splines (4)
- Treatment Effect Heterogeneity (4)
-
- Bayesian hierarchical model (3)
- Causal inference (3)
- Measurement error (3)
- Multiple testing procedure (3)
- Time series (3)
- Treatment effect heterogeneity (3)
- Adaptive enrichment design (2)
- Bayesian inference (2)
- Bayesian methods (2)
- Competing risks (2)
- Conditional independence (2)
- Data augmentation (2)
- Epidemiology (2)
- Frailty (2)
- Generalizability (2)
- Generalized estimating equations (2)
- Group sequential design (2)
- Health expenditures (2)
- Heterogeneity (2)
- Hierarchical models (2)
- Interaction (2)
- Internal and external validity (2)
- Log-normal (2)
- MCMC (2)
- Mixed model (2)
Articles 61 - 90 of 178
Full-Text Articles in Statistics and Probability
Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan
Simple Examples Of Estimating Causal Effects Using Targeted Maximum Likelihood Estimation, Michael Rosenblum, Mark J. Van Der Laan
Johns Hopkins University, Dept. of Biostatistics Working Papers
We present a brief overview of targeted maximum likelihood for estimating the causal effect of a single time point treatment and of a two time point treatment. We focus on simple examples demonstrating how to apply the methodology developed in (van der Laan and Rubin, 2006; Moore and van der Laan, 2007; van der Laan, 2010a,b). We include R code for the single time point case.
Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett
Population Functional Data Analysis Of Group Ica-Based Connectivity Measures From Fmri, Shanshan Li, Brian S. Caffo, Suresh Joel, Stewart Mostofsky, James Pekar, Susan Spear Bassett
Johns Hopkins University, Dept. of Biostatistics Working Papers
In this manuscript, we use a two-stage decomposition for the analysis of func- tional magnetic resonance imaging (fMRI). In the first stage, spatial independent component analysis is applied to the group fMRI data to obtain common brain networks (spatial maps) and subject-specific mixing matrices (time courses). In the second stage, functional principal component analysis is utilized to decompose the mixing matrices into population- level eigenvectors and subject-specific loadings. Inference is performed using permutation-based exact conditional logistic regression for matched pairs data. Simulation studies suggest the ability of the decomposition methods to recover population brain networks and the major direction of …
Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu
Functional Principal Components Model For High-Dimensional Brain Imaging, Vadim Zipunnikov, Brian S. Caffo, David M. Yousem, Christos Davatzikos, Brian S. Schwartz, Ciprian Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
We establish a fundamental equivalence between singular value decomposition (SVD) and functional principal components analysis (FPCA) models. The constructive relationship allows to deploy the numerical efficiency of SVD to fully estimate the components of FPCA, even for extremely high-dimensional functional objects, such as brain images. As an example, a functional mixed effect model is fitted to high-resolution morphometric (RAVENS) images. The main directions of morphometric variation in brain volumes are identified and discussed.
Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert
Modeling Functional Data With Spatially Heterogeneous Shape Characteristics, Ana-Maria Staicu, Ciprian M. Crainiceanu, Daniel S. Reich, David Ruppert
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose a novel class of models for functional data exhibiting skewness or other shape characteristics that vary with spatial or temporal location. We use copulas so that the marginal distributions and the dependence structure can be modeled independently. Dependence is modeled with a Gaussian or t-copula, so that there is an underlying latent Gaussian process. We model the marginal distributions using the skew t family. The mean, variance, and shape parameters are modeled nonparametrically as functions of location. A computationally tractable inferential framework for estimating heterogeneous asymmetric or heavy-tailed marginal distributions is introduced. This framework provides a new set …
Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov
Population Value Decomposition, A Framework For The Analysis Of Image Populations, Ciprian M. Crainiceanu, Brian S. Caffo, Sheng Luo, Vadim Zipunnikov
Johns Hopkins University, Dept. of Biostatistics Working Papers
Images, often stored in multidimensional arrays are fast becoming ubiquitous in medical and public health research. Analyzing populations of images is a statistical problem that raises a host of daunting challenges. The most severe challenge is that data sets incorporating images recorded for hundreds or thousands of subjects at multiple visits are massive. We introduce the population value decomposition (PVD), a general method for simultaneous dimensionality reduction of large populations of massive images. We show how PVD can seamlessly be incorporated into statistical modeling and lead to a new, transparent and fast inferential framework. Our methodology was motivated by and …
Multilevel Functional Principal Component Analysis For High-Dimensional Data, Vadim Zipunnikov, Brian Caffo, Ciprian Crainiceanu, David M. Yousem, Christos Davatzikos, Brian S. Schwartz
Multilevel Functional Principal Component Analysis For High-Dimensional Data, Vadim Zipunnikov, Brian Caffo, Ciprian Crainiceanu, David M. Yousem, Christos Davatzikos, Brian S. Schwartz
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose fast and scalable statistical methods for the analysis of hundreds or thousands of high dimensional vectors observed at multiple visits. The proposed inferential methods avoid the difficult task of loading the entire data set at once in the computer memory and use sequential access to data. This allows deployment of our methodology on low-resource computers where computations can be done in minutes on extremely large data sets. Our methods are motivated by and applied to a study where hundreds of subjects were scanned using Magnetic Resonance Imaging (MRI) at two visits roughly five years apart. The original data …
Estimating Temporal Associations In Electrocorticographic (Ecog) Time Series With First Order Pruning, Haley Hedlin, Dana Boatman, Brian Caffo
Estimating Temporal Associations In Electrocorticographic (Ecog) Time Series With First Order Pruning, Haley Hedlin, Dana Boatman, Brian Caffo
Johns Hopkins University, Dept. of Biostatistics Working Papers
Granger causality (GC) is a statistical technique used to estimate temporal associations in multivariate time series. Many applications and extensions of GC have been proposed since its formulation by Granger in 1969. Here we control for potentially mediating or confounding associations between time series in the context of event-related electrocorticographic (ECoG) time series. A pruning approach to remove spurious connections and simultaneously reduce the required number of estimations to fit the effective connectivity graph is proposed. Additionally, we consider the potential of adjusted GC applied to independent components as a method to explore temporal relationships between underlying source signals. Both …
Longitudinal Penalized Functional Regression, Jeff Goldsmith, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich
Longitudinal Penalized Functional Regression, Jeff Goldsmith, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich
Johns Hopkins University, Dept. of Biostatistics Working Papers
We propose a new regression model and inferential tools for the case when both the outcome and the functional exposures are observed at multiple visits. This data structure is new but increasingly present in applications where functions or images are recorded at multiple times. This raises new inferential challenges that cannot be addressed with current methods and software. Our proposed model generalizes the Generalized Linear Mixed Effects Model (GLMM) by adding functional predictors. Smoothness of the functional coefficients is ensured using roughness penalties estimated by Restricted Maximum Likelihood (REML) in a corresponding mixed effects model. This method is computationally feasible …
Mixed Effect Poisson Log-Linear Models For Clinical And Epidemiological Sleep Hypnogram Data, Bruce J. Swihart, Brian S. Caffo Phd, Ciprian Crainiceanu Phd, Naresh M. Punjabi Phd, Md
Mixed Effect Poisson Log-Linear Models For Clinical And Epidemiological Sleep Hypnogram Data, Bruce J. Swihart, Brian S. Caffo Phd, Ciprian Crainiceanu Phd, Naresh M. Punjabi Phd, Md
Johns Hopkins University, Dept. of Biostatistics Working Papers
Bayesian Poisson log-linear multilevel models scalable to epidemiological studies are proposed to investigate population variability in sleep state transition rates. Hierarchical random effects are used to account for pairings of individuals and repeated measures within those individuals, as comparing diseased to non-diseased subjects while minimizing bias is of importance. Essentially, non-parametric piecewise constant hazards are estimated and smoothed, allowing for time-varying covariates and segment of the night comparisons. The Bayesian Poisson regression is justified through a re-derivation of a classical algebraic likelihood equivalence of Poisson regression with a log(time) offset and survival regression assuming exponentially distributed survival times. Such re-derivation …
A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu
A Unified Approach To Modeling Multivariate Binary Data Using Copulas Over Partitions, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
Many seemingly disparate approaches for marginal modeling have been developed in recent years. We demonstrate that many current approaches for marginal modeling of correlated binary outcomes produce likelihoods that are equivalent to the proposed copula-based models herein. These general copula models of underlying latent threshold random variables yield likelihood based models for marginal fixed effects estimation and interpretation in the analysis of correlated binary data. Moreover, we propose a nomenclature and set of model relationships that substantially elucidates the complex area of marginalized models for binary data. A diverse collection of didactic mathematical and numerical examples are given to illustrate …
The Use Of Propensity Scores To Assess The Generalizability Of Results From Randomized Trials, Elizabeth A. Stuart, Stephen R. Cole, Catherine P. Bradshaw, Philip J. Leaf
The Use Of Propensity Scores To Assess The Generalizability Of Results From Randomized Trials, Elizabeth A. Stuart, Stephen R. Cole, Catherine P. Bradshaw, Philip J. Leaf
Johns Hopkins University, Dept. of Biostatistics Working Papers
Randomized trials remain the most accepted design for estimating the effects of interventions, but they do not necessarily answer a question of primary interest: Will the program be effective in a target population in which it may be implemented? In other words,are the results generalizable? There has been very little statistical research on how to assess the generalizability, or "external validity," of randomized trials. We propose the use of propensity-score-based metrics to quantify the similarity of the participants in a randomized trial and a target population. In this setting the propensity score model predicts participation in the randomized trial, given …
Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang
Likelihood Ratio Testing For Admixture Models With Application To Genetic Linkage Analysis, Chong-Zhi Di, Kung-Yee Liang
Johns Hopkins University, Dept. of Biostatistics Working Papers
We consider likelihood ratio tests (LRT) and their modifications for homogeneity in admixture models. The admixture model is a special case of two component mixture model, where one component is indexed by an unknown parameter while the parameter value for the other component is known. It has been widely used in genetic linkage analysis under heterogeneity, in which the kernel distribution is binomial. For such models, it is long recognized that testing for homogeneity is nonstandard and the LRT statistic does not converge to a conventional 2 distribution. In this paper, we investigate the asymptotic behavior of the LRT for …
Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu
Multilevel Sparse Functional Principal Component Analysis, Chong-Zhi Di, Ciprian M. Crainiceanu
Johns Hopkins University, Dept. of Biostatistics Working Papers
The basic observational unit in this paper is a function. Data are assumed to have a natural hierarchy of basic units. A simple example is when functions are recorded at multiple visits for the same subject. Di et al. (2009) proposed Multilevel Functional Principal Component Analysis (MFPCA) for this type of data structure when functions are densely sampled. Here we consider the case when functions are sparsely sampled and may contain as few as 2 or 3 observations per function. As with MFPCA, we exploit the multilevel structure of covariance operators and data reduction induced by the use of principal …
Penalized Functional Regression, Jeff Goldsmith, Jennifer Feder, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich
Penalized Functional Regression, Jeff Goldsmith, Jennifer Feder, Ciprian M. Crainiceanu, Brian Caffo, Daniel Reich
Johns Hopkins University, Dept. of Biostatistics Working Papers
We develop fast fitting methods for generalized functional linear models. An undersmooth of the functional predictor is obtained by projecting on a large number of smooth eigenvectors and the coefficient function is estimated using penalized spline regression. Our method can be applied to many functional data designs including functions measured with and without error, sparsely or densely sampled. The methods also extend to the case of multiple functional predictors or functional predictors with a natural multilevel structure. Our approach can be implemented using standard mixed effects software and is computationally fast. Our methodology is motivated by a diffusion tensor imaging …
Regression Adjustment And Stratification By Propensty Score In Treatment Effect Estimation, Jessica A. Myers, Thomas A. Louis
Regression Adjustment And Stratification By Propensty Score In Treatment Effect Estimation, Jessica A. Myers, Thomas A. Louis
Johns Hopkins University, Dept. of Biostatistics Working Papers
Propensity score adjustment of effect estimates in observational studies of treatment is a common technique used to control for bias in treatment assignment. In situations where matching on propensity score is not possible or desirable, regression adjustment and stratification are two options. Regression adjustment is used most often and can be highly efficient, but it can lead to biased results when model assumptions are violated. Validity of the stratification approach depends on fewer model assumptions, but is less efficient than regression adjustment when the regression assumptions hold. To investigate these issues, by simulation we compare stratification and regression adjustments. We …
On The Behaviour Of Marginal And Conditional Akaike Information Criteria In Linear Mixed Models, Sonja Greven, Thomas Kneib
On The Behaviour Of Marginal And Conditional Akaike Information Criteria In Linear Mixed Models, Sonja Greven, Thomas Kneib
Johns Hopkins University, Dept. of Biostatistics Working Papers
In linear mixed models, model selection frequently includes the selection of random effects. Two versions of the Akaike information criterion (AIC) have been used, based either on the marginal or on the conditional distribution. We show that the marginal AIC is no longer an asymptotically unbiased estimator of the Akaike information, and in fact favours smaller models without random effects. For the conditional AIC, we show that ignoring estimation uncertainty in the random effects covariance matrix, as is common practice, induces a bias that leads to the selection of any random effect not predicted to be exactly zero. We derive …
Analyzing Bivariate Survival Data With Interval Sampling And Application To Cancer Epidemiology, Hong Zhu, Mei-Cheng Wang
Analyzing Bivariate Survival Data With Interval Sampling And Application To Cancer Epidemiology, Hong Zhu, Mei-Cheng Wang
Johns Hopkins University, Dept. of Biostatistics Working Papers
In medical follow-up studies, ordered bivariate survival data are frequently encountered when bivariate failure events are used as the outcomes to identify the progression of a disease. In cancer studies interest could be focused on bivariate failure times, for example, time from birth to cancer onset and time from cancer onset to death. This paper considers a sampling scheme where the first failure event (cancer onset) is identified within a calendar time interval, the time of the initiating event (birth) can be retrospectively confirmed, and the occurrence of the second event (death) is observed sub ject to right censoring. To …
Modeling Multilevel Sleep Transitional Data Via Poisson Log-Linear Multilevel Models, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu, Naresh M. Punjabi
Modeling Multilevel Sleep Transitional Data Via Poisson Log-Linear Multilevel Models, Bruce J. Swihart, Brian Caffo, Ciprian Crainiceanu, Naresh M. Punjabi
Johns Hopkins University, Dept. of Biostatistics Working Papers
This paper proposes Poisson log-linear multilevel models to investigate population variability in sleep state transition rates. We specifically propose a Bayesian Poisson regression model that is more flexible, scalable to larger studies, and easily fit than other attempts in the literature. We further use hierarchical random effects to account for pairings of individuals and repeated measures within those individuals, as comparing diseased to non-diseased subjects while minimizing bias is of epidemiologic importance. We estimate essentially non-parametric piecewise constant hazards and smooth them, and allow for time varying covariates and segment of the night comparisons. The Bayesian Poisson regression is justified …
Bayesian Functional Data Analysis Using Winbugs, Ciprian M. Crainiceanu, A. Jeffrey Goldsmith
Bayesian Functional Data Analysis Using Winbugs, Ciprian M. Crainiceanu, A. Jeffrey Goldsmith
Johns Hopkins University, Dept. of Biostatistics Working Papers
We provide user friendly software for Bayesian analysis of Functional Data Models using WinBUGS 1.4. The excellent properties of Bayesian analysis in this context are due to: 1) dimensionality reduction, which leads to low dimensional projection bases; 2)the mixed model representation of functional models, which provides a modular approach to model extension; and 3) the orthogonality of the principal component bases, which contributes to excellent chain convergence and mixing properties. Our paper provides one more, essential, reason for using Bayesian analysis for Functional models: the existence of software.
Lasagna Plots: A Saucy Alternative To Spaghetti Plots, Bruce Swihart, Brian Caffo, Bryan D. James, Matthew Strand, Brian S. Schwartz, Naresh M. Punjabi
Lasagna Plots: A Saucy Alternative To Spaghetti Plots, Bruce Swihart, Brian Caffo, Bryan D. James, Matthew Strand, Brian S. Schwartz, Naresh M. Punjabi
Johns Hopkins University, Dept. of Biostatistics Working Papers
Longitudinal repeated measures data has often been visualized with spaghetti plots for continuous out- comes. For large datasets, this often leads to over-plotting and consequential obscuring of trends in the data. This is primarily due to overlapping of trajectories. Here, we suggest a framework called lasagna plot ting that constrains the subject-specific trajectories to prevent overlapping and utilizes gradients of color to depict the outcome. Dynamic sorting and visualization is demonstrated as an exploratory data analysis tool. Supplemental material in the form of sample R code additional illustrated examples are available online.
Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry
Redefining Cpg Islands Using A Hideen Markov Model, Hao Wu, Brain Caffo, Harris A. Jaffee, Andrew P. Feinberg, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
The DNA of most vertebrates is depleted in CpG dinucleotides; C followed by a G in the 5’ to 3’ direction. CpGs are the target for DNA methylation, a chemical modification of cytosine (C) heritable during cell division and the most well characterized epigenetic mechanism. The remaining CpGs tend to cluster in regions referred to as CpG islands (CGI). Knowing CGI locations is important because they mark functionally relevant epigenetic loci in development and disease. For various mammals, including human, a readily available and widely used list of CGI is available from the UCSC Genome Browser. This list was derived …
Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani
Combinational Mixtures Of Multiparameter Distributions, Valeria Edefonti, Giovanni Parmigiani
Johns Hopkins University, Dept. of Biostatistics Working Papers
We introduce combinatorial mixtures - a flexible class of models for inference on mixture distributions whose component have multidimensional parameters. The key idea is to allow each element of the component-specific parameter vectors to be shared by a subset of other components. This approach allows for mixtures that range from very flexible to very parsimonious, and unifies inference on component-specific parameters with inference on the number of components. We develop Bayesian inference and computation approaches for this class of distributions, and illustrate them in an application. This work was originally motivated by the analysis of cancer subtypes: in terms of …
A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
A Multilevel Model To Address Batch Effects In Copy Number Estimation Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …
A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
A Multilevel Model To Address Batch Effects In Copy Number Using Snp Arrays, Robert B. Scharpf, Ingo Ruczinski, Benilton Carvalho, Betty Doan, Aravinda Chakravarti, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Submicroscopic changes in chromosomal DNA copy number dosage are common and have been implicated in many heritable diseases and cancers. Recent high-throughput technologies have a resolution that permits the detection of segmental changes in DNA copy number that span thousands of basepairs across the genome. Genome-wide association studies (GWAS) may simultaneously screen for copy number-phenotype and SNP-phenotype associations as part of the analytic strategy. However, genome-wide array analyses are particularly susceptible to batch effects as the logistics of preparing DNA and processing thousands of arrays often involves multiple laboratories and technicians, or changes over calendar time to the reagents and …
A Spatio-Temporal Approach For Estimating Chronic Effects Of Air Pollution, Sonja Greven, Francesca Dominici, Scott L. Zeger
A Spatio-Temporal Approach For Estimating Chronic Effects Of Air Pollution, Sonja Greven, Francesca Dominici, Scott L. Zeger
Johns Hopkins University, Dept. of Biostatistics Working Papers
Estimating the health risks associated with air pollution exposure is of great importance in public health. In air pollution epidemiology, two study designs have been used mainly. Time series studies estimate acute risk associated with short-term exposure. They compare day-to-day variation of pollution concentrations and mortality rates, and have been criticized for potential confounding by time-varying covariates. Cohort studies estimate chronic effects associated with long-term exposure. They compare long-term average pollution concentrations and time-to-death across cities, and have been criticized for potential confounding by individual risk factors or city-level characteristics.
We propose a new study design and a statistical model, …
A Bayesian Shrinkage Model For Incomplete Longitudinal Binary Data With Application To The Breast Cancer Prevention Trial, C. Wang, M.J. Daniels, Daniel O. Scharfstein, S. Land
A Bayesian Shrinkage Model For Incomplete Longitudinal Binary Data With Application To The Breast Cancer Prevention Trial, C. Wang, M.J. Daniels, Daniel O. Scharfstein, S. Land
Johns Hopkins University, Dept. of Biostatistics Working Papers
We consider inference in randomized studies, in which repeatedly measured outcomes may be informatively missing due to drop out. In this setting, it is well known that full data estimands are not identified unless unverified assumptions are imposed. We assume a non-future dependence model for the drop-out mechanism and posit an exponential tilt model that links non-identifiable and identifiable distributions. This model is indexed by non-identified parameters, which are assumed to have an informative prior distribution, elicited from subject-matter experts. Under this model, full data estimands are shown to be expressed as functionals of the distribution of the observed data. …
Quantifying Uncertainty In Genotype Calls, Benilton Carvalho, Thomas A. Louis, Rafael A. Irizarry
Quantifying Uncertainty In Genotype Calls, Benilton Carvalho, Thomas A. Louis, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
Genome-wide association studies (GWAS) are used to discover genes underlying complex, heritable disorders for which less powerful study designs have failed in the past. The number of GWAS has skyrocketed recently with findings reported in top journals and the mainstream media. Mircorarrays are the genotype calling technology of choice in GWAS as they permit exploration of more than a million single nucleotide polymorphisms (SNPs)simultaneously. The starting point for the statistical analyses used by GWAS, to determine association between loci and disease, are genotype calls (AA, AB, or BB). However, the raw data, microarray probe intensities, are heavily processed before arriving …
Bayesian Model Averaging For Clustered Data: Imputing Missing Daily Air Pollution Concentration, Howard H. Chang, Francesca Dominici, Roger D. Peng
Bayesian Model Averaging For Clustered Data: Imputing Missing Daily Air Pollution Concentration, Howard H. Chang, Francesca Dominici, Roger D. Peng
Johns Hopkins University, Dept. of Biostatistics Working Papers
The presence of missing observations is a challenge in statistical analysis especially when data are clustered. In this paper, we develop a Bayesian model averaging (BMA) approach for imputing missing observations in clustered data. Our approach extends BMA by allowing the weights of competing regression models for missing data imputation to vary between clusters while borrowing information across clusters in estimating model parameters. Through simulation and cross-validation studies, we demonstrate that our approach outperforms the standard BMA imputation approach where model weights are assumed to be the same for all clusters. We then apply our proposed method to a national …
Spatial Misalignment In Time Series Studies Of Air Pollution And Health Data, Roger D. Peng, Michelle L. Bell
Spatial Misalignment In Time Series Studies Of Air Pollution And Health Data, Roger D. Peng, Michelle L. Bell
Johns Hopkins University, Dept. of Biostatistics Working Papers
Time series studies of environmental exposures often involve comparing daily changes in a toxicant measured at a point in space with daily changes in an aggregate measure of health. Spatial misalignment of the exposure and response variables can bias the estimation of health risk and the magnitude of this bias depends on the spatial variation of the exposure of interest. In air pollution epidemiology, there is an increasing focus on estimating the health effects of the chemical components of particulate matter. One issue that is raised by this new focus is the spatial misalignment error introduced by the lack of …
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Multilevel Latent Class Models With Dirichlet Mixing Distribution, Chongzhi Di, Karen Bandeen-Roche
Johns Hopkins University, Dept. of Biostatistics Working Papers
Latent class analysis (LCA) and latent class regression (LCR) are widely used for modeling multivariate categorical outcomes in social sciences and biomedical studies. Standard analyses assume data of different respondents to be mutually independent, excluding application of the methods to familial and other designs in which participants are clustered. In this paper, we develop multilevel latent class model, in which subpopulation mixing probabilities are treated as random effects that vary among clusters according to a common Dirichlet distribution. We apply the Expectation-Maximization (EM) algorithm for model fitting by maximum likelihood (ML). This approach works well, but is computationally intensive when …