Open Access. Powered by Scholars. Published by Universities.®

Statistical Models Commons

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 61 - 90 of 242

Full-Text Articles in Statistical Models

Joint Spatial Modeling Of Recurrent Infection And Growth With Processes Under Intermittent Observation, Farouk S. Nathoo Aug 2008

Joint Spatial Modeling Of Recurrent Infection And Growth With Processes Under Intermittent Observation, Farouk S. Nathoo

COBRA Preprint Series

In this article we present new statistical methodology for longitudinal studies in forestry where trees are subject to recurrent infection and the hazard of infection depends on tree growth over time. Understanding the nature of this dependence has important implications for reforestation and breeding programs. Challenges arise for statistical analysis in this setting with sampling schemes leading to panel data, exhibiting dynamic spatial variability, and incomplete covariate histories for hazard regression. In addition, data are collected at a large number of locations which poses computational difficulties for spatiotemporal modeling. A joint model for infection and growth is developed; wherein, a …


Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan Jun 2008

Supervised Distance Matrices: Theory And Applications To Genomics, Katherine S. Pollard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a new approach to studying the relationship between a very high dimensional random variable and an outcome. Our method is based on a novel concept, the supervised distance matrix, which quantifies pairwise similarity between variables based on their association with the outcome. A supervised distance matrix is derived in two stages. The first stage involves a transformation based on a particular model for association. In particular, one might regress the outcome on each variable and then use the residuals or the influence curve from each regression as a data transformation. In the second stage, a choice of distance …


Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan Jun 2008

Confidence Intervals For The Population Mean Tailored To Small Sample Sizes, With Applications To Survey Sampling, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

The validity of standard confidence intervals constructed in survey sampling is based on the central limit theorem. For small sample sizes, the central limit theorem may give a poor approximation, resulting in confidence intervals that are misleading. We discuss this issue and propose methods for constructing confidence intervals for the population mean tailored to small sample sizes.

We present a simple approach for constructing confidence intervals for the population mean based on tail bounds for the sample mean that are correct for all sample sizes. Bernstein's inequality provides one such tail bound. The resulting confidence intervals have guaranteed coverage probability …


A Method For Visualizing Multivariate Time Series Data, Roger D. Peng Feb 2008

A Method For Visualizing Multivariate Time Series Data, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

Visualization and exploratory analysis is an important part of any data analysis and is made more challenging when the data are voluminous and high-dimensional. One such example is environmental monitoring data, which are often collected over time and at multiple locations, resulting in a geographically indexed multivariate time series. Financial data, although not necessarily containing a geographic component, present another source of high-volume multivariate time series data. We present the mvtsplot function which provides a method for visualizing multivariate time series data. We outline the basic design concepts and provide some examples of its usage by applying it to a …


Jointly Modeling Continuous And Binary Outcomes For Boolean Outcomes: An Application To Modeling Hypertension, Xianbin Li, Brian S. Caffo, Elizabeth Stuart Feb 2008

Jointly Modeling Continuous And Binary Outcomes For Boolean Outcomes: An Application To Modeling Hypertension, Xianbin Li, Brian S. Caffo, Elizabeth Stuart

Johns Hopkins University, Dept. of Biostatistics Working Papers

Binary outcomes defined by logical (Boolean) "and" or "or" operations on original continuous and discrete outcomes arise commonly in medical diagnoses and epidemiological research. In this manuscript,we consider applying the “or” operator to two continuous variables above a threshold and a binary variable, a setting that occurs frequently in the modeling of hypertension. Rather than modeling the resulting composite outcome defined by the logical operator, we present a method that models the original outcomes thus utilizing all information in the data, yet continues to yield conclusions on the composite scale. A stratified propensity score adjustment is proposed to account for …


Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan Jan 2008

Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Regression models are often used to test for cause-effect relationships from data collected in randomized trials or experiments. This practice has deservedly come under heavy scrutiny, since commonly used models such as linear and logistic regression will often not capture the actual relationships between variables, and incorrectly specified models potentially lead to incorrect conclusions. In this paper, we focus on hypothesis test of whether the treatment given in a randomized trial has any effect on the mean of the primary outcome, within strata of baseline variables such as age, sex, and health status. Our primary concern is ensuring that such …


Bayesian Analysis For Penalized Spline Regression Using Win Bugs, Ciprian M. Crainiceanu, David Ruppert, M.P. Wand Dec 2007

Bayesian Analysis For Penalized Spline Regression Using Win Bugs, Ciprian M. Crainiceanu, David Ruppert, M.P. Wand

Johns Hopkins University, Dept. of Biostatistics Working Papers

Penalized splines can be viewed as BLUPs in a mixed model framework, which allows the use of mixed model software for smoothing. Thus, software originally developed for Bayesian analysis of mixed models can be used for penalized spline regression. Bayesian inference for nonparametric models enjoys the flexibility of nonparametric models and the exact inference provided by the Bayesian inferential machinery. This paper provides a simple, yet comprehensive, set of programs for the implementation of nonparametric Bayesian analysis in WinBUGS. MCMC mixing is substantially improved from the previous versions by using low{rank thin{plate splines instead of truncated polynomial basis. Simulation time …


Assessing Population Level Genetic Instability Via Moving Average, Samuel Mcdaniel, Rebecca Betensky, Tianxi Cai Nov 2007

Assessing Population Level Genetic Instability Via Moving Average, Samuel Mcdaniel, Rebecca Betensky, Tianxi Cai

Harvard University Biostatistics Working Paper Series

No abstract provided.


Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard Jul 2007

Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

Previous articles (van der Laan and Dudoit (2003); van der Laan et al. (2006); Sinisi et al. (2007)) advertised and theoretically validated the use of cross-validation to select among many candidate estimators to compute a so called super learner which outperforms any of the given candidate estimators. The theoretical basis was provided for this super learner based on oracle results for the cross-validation selector (e.g., van der Laan and Dudoit (2003); van der Laan et al. (2006)) and in Sinisi et al. (2007). In addition, these papers contained a practical demonstration of the adaptivity of this so called super learner …


Random Effects Models In A Meta-Analysis Of The Accuracy Of Diagnostic Tests Within A Gold Standard In The Presence Of Missing Data, Haitao Chu, Sining Chen, Thomas A. Louis Jun 2007

Random Effects Models In A Meta-Analysis Of The Accuracy Of Diagnostic Tests Within A Gold Standard In The Presence Of Missing Data, Haitao Chu, Sining Chen, Thomas A. Louis

Johns Hopkins University, Dept. of Biostatistics Working Papers

In evaluating the accuracy of diagnosis tests, it is common to apply two imperfect tests jointly or sequentially to a study population. In a recent meta-analysis of the accuracy of microsatellite instability testing (MSI) and traditional mutation analysis (MUT) in predicting germline mutations of the mismatch repair (MMR) genes, a Bayesian approach (Chen, Watson, and Parmigiani 2005) was proposed to handle missing data resulting from partial testing and the lack of a gold standard. In this paper, we demonstrate an improved estimation of the sensitivities and specificities of MSI and MUT by using a nonlinear mixed model and a Bayesian …


Bayesian Bivariate Image Analysis With Application To Dual Autoradiography, Timothy D. Johnson, Morand Piert May 2007

Bayesian Bivariate Image Analysis With Application To Dual Autoradiography, Timothy D. Johnson, Morand Piert

The University of Michigan Department of Biostatistics Working Paper Series

We present a Bayesian bivariate image model and apply it to a study that was designed to investigate the relationship between hypoxia and angiogenesis in an animal tumor model. Two radiolabeled tracers (one measuring angio- genesis, the other measuring hypoxia) were simultaneously injected into the animals, the tumors removed and autoradiographic images of the tracer concentrations were obtained. We model correlation between tracers with a mixture of bivariate normal distributions and the spatial correlation inherent in the images by means of the celebrated Potts model. Although the Potts model is typically used for image segmentation, we use it solely as …


Quantitative Magnetic Resonance Image Analysis Via The Em Algorithm With Stochastic Variation, Xiaoxi Zhang, Timothy D. Johnson, Roderick J.A. Little May 2007

Quantitative Magnetic Resonance Image Analysis Via The Em Algorithm With Stochastic Variation, Xiaoxi Zhang, Timothy D. Johnson, Roderick J.A. Little

The University of Michigan Department of Biostatistics Working Paper Series

Quantitative Magnetic Resonance Imaging (qMRI) provides researchers insight into pathological and physiological alterations of living tissue, with the help of which, researchers hope to predict (local) therapeutic efficacy early and determine optimal treatment schedule. However, the analysis of qMRI has been limited to ad-hoc heuristic methods. Our research provides a powerful statistical framework for image analysis and sheds light on future localized adaptive treatment regimes tailored to the individual’s response. We assume in an imperfect world we only observe a blurred and noisy version of the underlying “true” scene via qMRI, due to measurement errors or unpredictable influences. We use …


Bayesian Spatial Modeling Of Fmri Data: A Multiple-Subject Analysis, Lei Xu, Timothy Johnson, Thomas Nichols Apr 2007

Bayesian Spatial Modeling Of Fmri Data: A Multiple-Subject Analysis, Lei Xu, Timothy Johnson, Thomas Nichols

The University of Michigan Department of Biostatistics Working Paper Series

The aim of this work is to develop a spatial model for multi-subject fMRI data. While there has been much work on univariate modeling of each voxel for single- and multi-subject data, and some work on spatial modeling for single-subject data, there has been no work on spatial models that explicitly account for intersubject variability in activation location. We use a Bayesian hierarchical spatial model to fit the data. At the first level we model "population centers" that mark the centers of regions of activation. For a given population center each subject may have zero or more associated "individual components". …


A Survey Of The Likelihood Approach To Bioequivalence Trials, Leena Choi, Brian S. Caffo, Charles Rohde Feb 2007

A Survey Of The Likelihood Approach To Bioequivalence Trials, Leena Choi, Brian S. Caffo, Charles Rohde

Johns Hopkins University, Dept. of Biostatistics Working Papers

Bioequivalence trials are abbreviated clinical trials whereby a generic drug or new formulation is evaluated to determine if it is "equivalent" to a corresponding previously approved brand-name drug or formulation. In this manuscript, we survey the process of testing bioequivalence and advocate the likelihood paradigm for representing the resulting data as evidence. We emphasize the unique conflicts between hypothesis testing and confidence intervals in this area - which we believe are indicative of the existence of the systemic defects in the frequentist approach - that the likelihood paradigm avoids. We suggest the direct use of profile likelihoods for evaluating bioequivalence …


Mortality In The Medicare Population And Chronic Exposure To Fine Particulate Air Pollution , Scott L. Zeger, Francesca Dominici, Aidan Mcdermott, Jonathan M. Samet Jan 2007

Mortality In The Medicare Population And Chronic Exposure To Fine Particulate Air Pollution , Scott L. Zeger, Francesca Dominici, Aidan Mcdermott, Jonathan M. Samet

Johns Hopkins University, Dept. of Biostatistics Working Papers

Prospective cohort studies have provided evidence on longer-term mortality risks of fine particulate matter (PM2.5), but due to their complexity and costs, only a few have been conducted.

By linking monitoring data to the U.S. Medicare system by county of residence, we developed a retrospective cohort study, the Medicare Air Pollution Cohort Study (MCAPS), comprising over 20 million enrollees in the 250 largest counties during 2000-2002. We estimated log-linear regression models having as outcome the age-specific mortality rate for each county and as the main predictor, the average level for the study period 2000. Area-level covariates were used to adjust …


A Bayesian Hierarchical Model For Constrained Distributed Lag Functions: Estimating The Time Course Of Hospitalization Associated With Air Pollution Exposure, Roger Peng, Francesca Dominici, Leah J. Welty Jan 2007

A Bayesian Hierarchical Model For Constrained Distributed Lag Functions: Estimating The Time Course Of Hospitalization Associated With Air Pollution Exposure, Roger Peng, Francesca Dominici, Leah J. Welty

Johns Hopkins University, Dept. of Biostatistics Working Papers

Numerous time series studies have provided strong evidence of an association between increased levels of ambient air pollution and increased levels of hospital admissions, typically at 0, 1, or 2 days after an air pollution episode. An important research aim is to extend existing statistical models so that a more detailed understanding of the time course of hospitalization after exposure to air pollution can be obtained. Information about this time course, combined with prior knowledge about biological mechanisms, could provide the basis for hypotheses concerning the mechanism by which air pollution causes disease. Previous studies have identified two important methodological …


Gamma Shape Mixtures For Heavy-Tailed Distributions, Sergio Venturini, Francesca Dominici, Giovanni Parmigiani Dec 2006

Gamma Shape Mixtures For Heavy-Tailed Distributions, Sergio Venturini, Francesca Dominici, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

An important question in health services research is the estimation of the proportion of medical expenditures that exceed a given threshold. Typically, medical expenditures present highly skewed, heavy tailed distributions, for which a) simple variable transformations are insufficient to achieve a tractable low- dimensional parametric form and b) nonparametric methods are not efficient in estimating exceedance probabilities for large thresholds. Motivated by this context, in this paper we propose a general Bayesian approach for the estimation of tail probabilities of heavy-tailed distributions,based on a mixture of gamma distributions in which the mixing occurs over the shape parameter. This family provides …


Spatio-Temporal Analysis Of Areal Data And Discovery Of Neighborhood Relationships In Conditionally Autoregressive Models, Subharup Guha, Louise Ryan Nov 2006

Spatio-Temporal Analysis Of Areal Data And Discovery Of Neighborhood Relationships In Conditionally Autoregressive Models, Subharup Guha, Louise Ryan

Harvard University Biostatistics Working Paper Series

No abstract provided.


Semiparametric Regression Of Multi-Dimensional Genetic Pathway Data: Least Squares Kernel Machines And Linear Mixed Models, Dawei Liu, Xihong Lin, Debashis Ghosh Nov 2006

Semiparametric Regression Of Multi-Dimensional Genetic Pathway Data: Least Squares Kernel Machines And Linear Mixed Models, Dawei Liu, Xihong Lin, Debashis Ghosh

Harvard University Biostatistics Working Paper Series

No abstract provided.


Statistical Analysis Of Air Pollution Panel Studies: An Illustration, Holly Janes, Lianne Sheppard, Kristen Shepherd Oct 2006

Statistical Analysis Of Air Pollution Panel Studies: An Illustration, Holly Janes, Lianne Sheppard, Kristen Shepherd

UW Biostatistics Working Paper Series

The panel study design is commonly used to evaluate the short-term health effects of air pollution. Standard statistical methods for analyzing longitudinal data are available, but the literature reveals that the techniques are not well understood by practitioners. We illustrate these methods using data from the 1999 to 2002 Seattle panel study. Marginal, conditional, and transitional approaches for modeling longitudinal data are reviewed and contrasted with respect to their parameter interpretation and methods for accounting for correlation and dealing with missing data. We also discuss and illustrate techniques for controlling for time-dependent and time-independent confounding, and for exploring and summarizing …


Cox Models With Nonlinear Effect Of Covariates Measured With Error: A Case Study Of Chronic Kidney Disease Incidence, Ciprian M. Crainiceanu, David Ruppert, Josef Coresh Sep 2006

Cox Models With Nonlinear Effect Of Covariates Measured With Error: A Case Study Of Chronic Kidney Disease Incidence, Ciprian M. Crainiceanu, David Ruppert, Josef Coresh

Johns Hopkins University, Dept. of Biostatistics Working Papers

We propose, develop and implement the simulation extrapolation (SIMEX) methodology for Cox regression models when the log hazard function is linear in the model parameters but nonlinear in the variables measured with error (LPNE). The class of LPNE functions contains but is not limited to strata indicators, splines, quadratic and interaction terms. The first order bias correction method proposed here has the advantage that it remains computationally feasible even when the number of observations is very large and multiple models need to be explored. Theoretical and simulation results show that the SIMEX method outperforms the naive method even with small …


Spatial Cluster Detection For Censored Outcome Data, Andrea J. Cook, Diane Gold, Yi Li Sep 2006

Spatial Cluster Detection For Censored Outcome Data, Andrea J. Cook, Diane Gold, Yi Li

Harvard University Biostatistics Working Paper Series

No abstract provided.


Adjustment Uncertainty In Effect Estimation, Ciprian M. Crainiceanu, Francesca Dominici, Giovanni Parmigiani Aug 2006

Adjustment Uncertainty In Effect Estimation, Ciprian M. Crainiceanu, Francesca Dominici, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

The selection of confounders and their functional relationship with the out- come affects exposure effect estimates. In practice, there is often substantial uncertainty about this selection, which we define here as “adjustment uncertainty.” We address the problem of estimating the effect of exposure on an outcome with focus on quantifying the effect of unknown confounders from a large set of potential confounders. We propose a general statistical framework for handling adjustment uncertainty in exposure effect estimation, a specific implementation called "Structured Estimation under Adjustment Uncertainty (STEADy)", and associated visualization tools. Theoretical results and simulation studies show that STEADy consistently estimates …


Bayesian Smoothing Of Irregularly-Spaced Data Using Fourier Basis Functions, Christopher J. Paciorek Aug 2006

Bayesian Smoothing Of Irregularly-Spaced Data Using Fourier Basis Functions, Christopher J. Paciorek

Harvard University Biostatistics Working Paper Series

No abstract provided.


Predicting Future Responses Based On Possibly Misspecified Working Models, Tianxi Cai, Lu Tian, Scott D. Solomon, L.J. Wei Aug 2006

Predicting Future Responses Based On Possibly Misspecified Working Models, Tianxi Cai, Lu Tian, Scott D. Solomon, L.J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


An Informative Bayesian Structural Equation Model To Assess Source-Specific Health Effects Of Air Pollution, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski Jul 2006

An Informative Bayesian Structural Equation Model To Assess Source-Specific Health Effects Of Air Pollution, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski

Harvard University Biostatistics Working Paper Series

No abstract provided.


Mixed Multiplicative Factor Analysis Model For Air Pollution Exposure Assessment, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski Jul 2006

Mixed Multiplicative Factor Analysis Model For Air Pollution Exposure Assessment, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski

Harvard University Biostatistics Working Paper Series

No abstract provided.


Relative Risk Regression In Medical Research: Models, Contrasts, Estimators, And Algorithms, Thomas Lumley, Richard Kronmal, Shuangge Ma Jul 2006

Relative Risk Regression In Medical Research: Models, Contrasts, Estimators, And Algorithms, Thomas Lumley, Richard Kronmal, Shuangge Ma

UW Biostatistics Working Paper Series

The relative risk or prevalence ratio is a natural and familiar summary of association between a binary outcome and an exposure or intervention. For rare events, the relative risk can be approximately estimated by logistic regression. For common events estimation is more difficult. We review proposed estimation algorithms for relative risk regression. Some of these give inconsistent estimates or invalid standard errors. We show that the methods that give correct inference can be viewed as arising from a family of quasilikelihood estimating functions for the same generalized linear model, differing in their efficiency and in their robustness to outlying values …


Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan Jul 2006

Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan

COBRA Preprint Series

In behavioral medicine trials, such as smoking cessation trials, two or more active treatments are often compared. Noncompliance by some subjects with their assigned treatment poses a challenge to the data analyst. Causal parameters of interest might include those defined by subpopulations based on their potential compliance status under each assignment, using the principal stratification framework (e.g., causal effect of new therapy compared to standard therapy among subjects that would comply with either intervention). Even if subjects in one arm do not have access to the other treatment(s), the causal effect of each treatment typically can only be identified from …


Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen May 2006

Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Bioequivalence trials are usually conducted to compare two or more formulations of a drug. Simultaneous assessment of bioequivalence on multiple endpoints is called multivariate bioequivalence. Despite the fact that some tests for multivariate bioequivalence are suggested, current practice usually involves univariate bioequivalence assessments ignoring the correlations between the endpoints such as AUC and Cmax. In this paper we develop a semiparametric Bayesian test for bioequivalence under multiple endpoints. Specifically, we show how the correlation between the endpoints can be incorporated in the analysis and how this correlation affects the inference. Resulting estimates and posterior probabilities ``borrow strength'' from one another …