Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 241 - 270 of 567

Full-Text Articles in Biostatistics

Vertically Shifted Mixture Models For Clustering Longitudinal Data By Shape, Brianna C. Heggeseth, Nicholas P. Jewell Mar 2013

Vertically Shifted Mixture Models For Clustering Longitudinal Data By Shape, Brianna C. Heggeseth, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

Longitudinal studies play a prominent role in health, social and behavioral sciences as well as in the biological sciences, economics, and marketing. By following subjects over time, temporal changes in an outcome of interest can be directly observed and studied. An important question concerns the existence of distinct trajectory patterns. One way to determine these distinct patterns is through cluster analysis, which seeks to separate objects (subjects, patients, observational units) into homogeneous groups. Many methods have been adapted for longitudinal data, but almost all of them fail to explicitly group trajectories according to distinct pattern shapes. To fulfill the need …


Efficient Estimation Of Risk Ratios From Clustered Binary Data, Matthew Cefalu, Eric Tchetgen Tchetgen Mar 2013

Efficient Estimation Of Risk Ratios From Clustered Binary Data, Matthew Cefalu, Eric Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


Predicting Human Movement Type Based On Multiple Accelerometers Using Movelets, Bing He, Jiawei Bai, Annemarie Koster, Casserotti Paolo, Nancy Glynn, Tamara B. Harris, Ciprian Crainiceanu Mar 2013

Predicting Human Movement Type Based On Multiple Accelerometers Using Movelets, Bing He, Jiawei Bai, Annemarie Koster, Casserotti Paolo, Nancy Glynn, Tamara B. Harris, Ciprian Crainiceanu

Johns Hopkins University, Dept. of Biostatistics Working Papers

We introduce statistical methods for prediction of types of human movement based on three tri-axial accelerometers worn simultaneously at the hip, left, and right wrist. We compare the individual performance of the three accelerometers using movelets and propose a new prediction algorithm that integrates the information from all three accelerometers. The development is motivated by a study of 20 older subjects who were instructed to perform 15 different types of activities during in-laboratory sessions. The differences in the prediction performance for different activity types among the three accelerometers reveal subtle yet important insights into how the intrinsic physical features of …


A Bayesian Regression Tree Approach To Identify The Effect Of Nanoparticles Properties On Toxicity Profiles, Cecile Low-Kam, Haiyuan Zhang, Zhaoxia Ji, Tian Xia, Jeffrey I. Zinc, Andre Nel, Donatello Telesca Mar 2013

A Bayesian Regression Tree Approach To Identify The Effect Of Nanoparticles Properties On Toxicity Profiles, Cecile Low-Kam, Haiyuan Zhang, Zhaoxia Ji, Tian Xia, Jeffrey I. Zinc, Andre Nel, Donatello Telesca

COBRA Preprint Series

We introduce a Bayesian multiple regression tree model to characterize relationships between physico-chemical properties of nanoparticles and their in-vitro toxicity over multiple doses and times of exposure. Unlike conventional models that rely on data summaries, our model solves the low sample size issue and avoids arbitrary loss of information by combining all measurements from a general exposure experiment across doses, times of exposure, and replicates. The proposed technique integrates Bayesian trees for modeling threshold effects and interactions, and penalized B-splines for dose and time-response surfaces smoothing. The resulting posterior distribution is sampled via a Markov Chain Monte Carlo algorithm. This …


Asymptotic And Finite Sample Behavior Of Net Reclassification Indices, Zheyu Wang Feb 2013

Asymptotic And Finite Sample Behavior Of Net Reclassification Indices, Zheyu Wang

UW Biostatistics Working Paper Series

The Net Reclassification Index (NRI) introduced by Pencina and colleagues [1, 2] is designed to quantify the prediction increment provided by a new biomarker. It has become popular for evaluating and selecting novel markers. The published variance formulae for NRI statistics do not account for the fact that risks are estimated based on risk models fit to data, and thus are not valid in practice when estimated risks are used [3]. Kerr and colleagues [4] showed that the confidence intervals constructed based on a bootstrap estimate of the variance and Normal approximation had the best performance among various methods they …


On The Restricted Mean Event Time In Survival Analysis, Lu Tian, Lihui Zhao, L. J. Wei Feb 2013

On The Restricted Mean Event Time In Survival Analysis, Lu Tian, Lihui Zhao, L. J. Wei

Harvard University Biostatistics Working Paper Series

No abstract provided.


Surrogacy Assessment Using Principal Stratification When Surrogate And Outcome Measures Are Multivariate Normal, Anna Conlon, Jeremy M.G. Taylor, Michael R. Elliott Feb 2013

Surrogacy Assessment Using Principal Stratification When Surrogate And Outcome Measures Are Multivariate Normal, Anna Conlon, Jeremy M.G. Taylor, Michael R. Elliott

The University of Michigan Department of Biostatistics Working Paper Series

No abstract provided.


Targeted Estimation Of Variable Importance Measures With Interval-Censored Outcomes, Stephanie Sapp, Mark J. Van Der Laan, Kimberly Page Feb 2013

Targeted Estimation Of Variable Importance Measures With Interval-Censored Outcomes, Stephanie Sapp, Mark J. Van Der Laan, Kimberly Page

U.C. Berkeley Division of Biostatistics Working Paper Series

In most experimental and observational studies, participants are not followed in continuous time. Instead, data is collected about participants only at certain monitoring times. These monitoring times are random, and often participant specific. As a result, outcomes are only known up to random time intervals, resulting in interval-censored data. In contrast, when estimating variable importance measures on interval-censored outcomes, practitioners often ignore the presence of interval-censoring, and instead treat the data as continuous or right-censored, applying ad-hoc approaches to mask the true interval-censoring. In this paper, we describe Targeted Minimum Loss-based Estimation methods tailored for estimation of variable importance measures …


Missing At Random And Ignorability For Inferences About Subsets Of Parameters With Missing Data, Roderick J. Little, Sahar Zanganeh Feb 2013

Missing At Random And Ignorability For Inferences About Subsets Of Parameters With Missing Data, Roderick J. Little, Sahar Zanganeh

The University of Michigan Department of Biostatistics Working Paper Series

For likelihood-based inferences from data with missing values, Rubin (1976) showed that the missing data mechanism can be ignored when (a) the missing data are missing at random (MAR), in the sense that missingness does not depend on the missing values after conditioning on the observed data, and (b) the parameters of the data model and the missing-data mechanism are distinct; that is, there are no a priori ties, via parameter space restrictions or prior distributions, between the parameters of the data model and the parameters of the model for the mechanism. Rubin described (a) and (b) as the "weakest …


Targeted Data Adaptive Estimation Of The Causal Dose Response Curve, Iván Díaz, Mark J. Van Der Laan Jan 2013

Targeted Data Adaptive Estimation Of The Causal Dose Response Curve, Iván Díaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Estimation of the causal dose-response curve is an old problem in statistics. In a non parametric model, if the treatment is continuous, the dose-response curve is not a pathwise differentiable parameter, and no root-n-consistent estimator is available. However, the risk of a candidate algorithm for estimation of the dose response curve is a pathwise differentiable parameter, whose consistent and efficient estimation is possible. In this work, we review the cross validated augmented inverse probability of treatment weighted estimator (CV A-IPTW) of the risk, and present a cross validated targeted minimum loss based estimator (CV-TMLE) counterpart. These estimators are proven consistent …


Statistical Methods For Evaluating And Comparing Biomarkers For Patient Treatment Selection, Holly Janes, Marshall D. Brown, Margaret Pepe, Ying Huang Jan 2013

Statistical Methods For Evaluating And Comparing Biomarkers For Patient Treatment Selection, Holly Janes, Marshall D. Brown, Margaret Pepe, Ying Huang

UW Biostatistics Working Paper Series

Despite the heightened interest in developing biomarkers predicting treatment response that are used to optimize patient treatment decisions, there has been relatively little development of statistical methodology to evaluate these markers. There is currently no unified statistical framework for marker evaluation. This paper proposes a suite of descriptive and inferential methods designed to evaluate individual markers and to compare candidate markers. An R software package has been developed which implements these methods. Their utility is illustrated in the breast cancer treatment context, where candidate markers are evaluated for their ability to identify a subset of women who do not benefit …


A General Regression Framework For A Secondary Outcome In Case-Control Studies, Eric J. Tchetgen Tchetgen Jan 2013

A General Regression Framework For A Secondary Outcome In Case-Control Studies, Eric J. Tchetgen Tchetgen

Harvard University Biostatistics Working Paper Series

No abstract provided.


In Praise Of Simplicity Not Mathematistry! Ten Simple Powerful Ideas For The Statistical Scientist, Roderick J. Little Jan 2013

In Praise Of Simplicity Not Mathematistry! Ten Simple Powerful Ideas For The Statistical Scientist, Roderick J. Little

The University of Michigan Department of Biostatistics Working Paper Series

Ronald Fisher was by all accounts a first-rate mathematician, but he saw himself as a scientist, not a mathematician, and he railed against what George Box called (in his Fisher lecture) "mathematistry". Mathematics is the indispensable foundation for statistics, but our subject is constantly under assault by people who want to turn statistics into a branch of mathematics, making the subject as impenetrable to non-mathematicians as possible. Valuing simplicity, I describe ten simple and powerful ideas that have influenced my thinking about statistics, in my areas of research interest: missing data, causal inference, survey sampling, and statistical modeling in general. …


Visualizing Longitudinal Data With Dropouts, Mithat Gonen Jan 2013

Visualizing Longitudinal Data With Dropouts, Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

A triangle plot is proposed to display longitudinal data with dropouts. The triangle plot is a tool of data visualization that can also serve as a graphical check for informativeness of the dropout process. There are similarities between the lasagna plot and the triangle plot but the explicit use of dropout time as an axis is an advantage of the triangle plot over the more commonly used graphical strategies for longitudinal data. It is possible to interpret the triangle plot as a trellis plot 1 which gives rise to several extensions such as the triangle histogram and the triangle boxplot. …


On The Simulation Of Longitudinal Discrete Data With Specified Marginal Means And First-Order Antedependence, Matthew Guerra, Justine Shults Jan 2013

On The Simulation Of Longitudinal Discrete Data With Specified Marginal Means And First-Order Antedependence, Matthew Guerra, Justine Shults

UPenn Biostatistics Working Papers

We propose a straightforward approach for simulation of discrete random variables with overdispersion, specified marginal means, and product correlations that are plausible for longitudinal data with equal, or unequal, temporal spacings. The method stems from results we prove for variables with first-order antedependence and linearity of the conditional expectations. The proposed approach will be especially useful for assessment of methods such as generalized estimating equations, which specify separate models for the marginal means and correlation structure of measurements on a subject.


Mixtures Of Receiver Operating Characteristic Curves, Mithat Gonen Jan 2013

Mixtures Of Receiver Operating Characteristic Curves, Mithat Gonen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Rationale and Objectives: ROC curves are ubiquitous in the analysis of imaging metrics as markers of both diagnosis and prognosis. While empirical estimation of ROC curves remains the most popular method, there are several reasons to consider smooth estimates based on a parametric model.

Materials and Methods: A mixture model is considered for modeling the distribution of the marker in the diseased population motivated by the biological observation that there is more heterogeneity in the diseased population than there is in the normal one. It is shown that this model results in an analytically tractable ROC curve which is itself …


A Frailty Approach For Survival Analysis With Error-Prone Covariate, Sehee Kim, Yi Li, Donna Spiegelman Jan 2013

A Frailty Approach For Survival Analysis With Error-Prone Covariate, Sehee Kim, Yi Li, Donna Spiegelman

The University of Michigan Department of Biostatistics Working Paper Series

This paper discovers an inherent relationship between the survival model with covariate measurement error and the frailty model. The discovery motivates our using a frailty-based estimating equation to draw inference for the proportional hazards model with error-prone covariates. Our established framework accommodates general distributional structures for the error-prone covariates, not restricted to a linear additive measurement error model or Gaussian measurement error. When the conditional distribution of the frailty given the surrogate is unknown, it is estimated through a semiparametric copula function. The proposed copula-based approach enables us to fit flexible measurement error models without the curse of dimensionality as …


Ultrahigh Dimensional Time Course Feature Selection, Peirong Xu, Lixing Zhu, Yi Li Jan 2013

Ultrahigh Dimensional Time Course Feature Selection, Peirong Xu, Lixing Zhu, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Statistical challenges arise from modern biomedical studies that produce time course genomic data with ultrahigh dimensions. In a renal cancer study that motivated this paper, the pharmacokinetic measures of a tumor suppressor (CCI-779) and expression levels of 12625 genes were measured for each of 33 patients at 8 and 16 weeks after the start of treatments, with the goal of identifying predictive gene transcripts and the interactions with time in peripheral blood mononuclear cells for pharmacokinetics over the time course. The resulting dataset defies analysis even with regularized regression. Although some remedies have been proposed for both linear and generalized …


Selection Of Latent Variables For Multiple Mixed-Outcome Models, Ling Zhou, Huazhen Lin, Xin-Yuan Song, Yi Li Jan 2013

Selection Of Latent Variables For Multiple Mixed-Outcome Models, Ling Zhou, Huazhen Lin, Xin-Yuan Song, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Latent variable models have been widely used for modeling the dependence structure of multiple outcomes data. As the formulation of a latent variable model is often unknown a priori, misspecification could distort the dependence structure and lead to unreliable model inference. More- over, the multiple outcomes are often of varying types (e.g., continuous and ordinal), which presents analytical challenges. In this article, we present a class of general latent variable models that can accommodate mixed types of outcomes, and further propose a novel selection approach that simultaneously selects latent variables and estimates model parameters. We show that the proposed estimators …


Semiparametric Latent Variable Transformation Models For Multiple Mixed Outcomes, Huazhen Lin, Ling Zhou, Robert Elashoff, Yi Li Jan 2013

Semiparametric Latent Variable Transformation Models For Multiple Mixed Outcomes, Huazhen Lin, Ling Zhou, Robert Elashoff, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

No abstract provided.


Semiparametric Transformation Models For Semicompeting Survival Data, Huazhen Lin, Ling Zhou, Chunhong Li, Yi Li Jan 2013

Semiparametric Transformation Models For Semicompeting Survival Data, Huazhen Lin, Ling Zhou, Chunhong Li, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Semicompeting risk outcome data, e.g. time to disease progression and time to death, are commonly collected in clinical trials, but complicated analytical tools hamper the analysis and the interpretation of the results. We propose a novel semiparametric transformation model for such data. Compared with the existing models, our model is advantageous in the following distinctive ways. First, it allows us to provide direct estimators of the regression analysis and the association parameter. Second, the measure of surrogacy, for example, the proportion of treatment effect and relative effect, can also be directly obtained. We propose a two-stage estimation procedure for inference …


Score Test Variable Screening, Sihai Dave Zhao, Yi Li Jan 2013

Score Test Variable Screening, Sihai Dave Zhao, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Variable screening has emerged as a crucial first step in the analysis of high-throughput data, but existing procedures can be computationally cumbersome, difficult to justify theoretically, or inapplicable to certain types of analyses. Motivated by a high-dimensional censored quantile regression problem in multiple myeloma genomics, this paper makes three contributions. First, we establish a score test-based screening framework, which is widely applicable, extremely computationally efficient, and relatively simple to justify. Secondly, we propose a resampling-based procedure for selecting the number of variables to retain after screening according to the principle of reproducibility. Finally, we propose a new iterative score test …


A Latent Variable Transformation Model Approach For Exploring Dysphagia, Anna Snavely, David P. Harrington, Yi Li Jan 2013

A Latent Variable Transformation Model Approach For Exploring Dysphagia, Anna Snavely, David P. Harrington, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

No abstract provided.


Covariance-Enhanced Discriminant Analysis, Peirong Xu, Ji Zhu, Lixing Zhu, Yi Li Jan 2013

Covariance-Enhanced Discriminant Analysis, Peirong Xu, Ji Zhu, Lixing Zhu, Yi Li

The University of Michigan Department of Biostatistics Working Paper Series

Linear discriminant analysis (LDA), a classical method in pattern recognition and machine learning, has been widely used to characterize or separate multiple classes via linear combinations of features. However, the high-dimensionality of the high-throughput features obtained from modern biological experiments, for example, microarray or proteomics, defies traditional discriminant analysis techniques. The possible interfeature correlations present additional challenges and are often under-utilized in modeling. In this paper, by incorporating the possible inter-feature correlations, we propose a Covariance-Enhanced Discriminant Analysis (CEDA) method that simultaneously and consistently selects informative features and identifies the corresponding discriminable classes. We show that, under mild regularity conditions, …


Optimal Spatial Prediction Using Ensemble Machine Learning, Molly M. Davies, Mark J. Van Der Laan Dec 2012

Optimal Spatial Prediction Using Ensemble Machine Learning, Molly M. Davies, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Spatial prediction is an important problem in many scientific disciplines. Super Learner is an ensemble prediction approach related to stacked generalization that uses cross-validation to search for the optimal predictor amongst all convex combinations of a heterogeneous candidate set. It has been applied to non-spatial data, where theoretical results demonstrate it will perform asymptotically at least as well as the best candidate under consideration. We review these optimality properties and discuss the assumptions required in order for them to hold for spatial prediction problems. We present results of a simulation study confirming Super Learner works well in practice under a …


Relating Nanoparticle Properties To Biological Outcomes In Exposure Escalation Experiments, Trina Patel, Cecile Low-Kam, Zhaoxia Ji, Haiyuan Zhang, Tian Xia, Andre E. Nel, Jeffrey I. Zinc, Donatello Telesca Dec 2012

Relating Nanoparticle Properties To Biological Outcomes In Exposure Escalation Experiments, Trina Patel, Cecile Low-Kam, Zhaoxia Ji, Haiyuan Zhang, Tian Xia, Andre E. Nel, Jeffrey I. Zinc, Donatello Telesca

COBRA Preprint Series

A fundamental goal in nano-toxicology is that of identifying particle physical and chemical properties, which are likely to explain biological hazard. The first line of screening for potentially adverse outcomes often consists of exposure escalation experiments, involving the exposure of micro-organisms or cell lines to a battery of nanomaterials. We discuss a modeling strategy, that relates the outcome of an exposure escalation experiment to nanoparticle properties. Our approach makes use of a hierarchical decision process, where we jointly identify particles that initiate adverse biological outcomes and explain the probability of this event in terms of the particle physico-chemical descriptors. The …


Sensitivity Analysis For Causal Inference Under Unmeasured Confounding And Measurement Error Problems, Iván Díaz, Mark J. Van Der Laan Dec 2012

Sensitivity Analysis For Causal Inference Under Unmeasured Confounding And Measurement Error Problems, Iván Díaz, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In this paper we present a sensitivity analysis for drawing inferences about parameters that are not estimable from observed data without additional assumptions. We present the methodology using two different examples: a causal parameter that is not identifiable due to violations of the randomization assumption, and a parameter that is not estimable in the nonparametric model due to measurement error. Existing methods for tackling these problems assume a parametric model for the type of violation to the identifiability assumption, and require the development of new estimators and inference for every new model. The method we present can be used in …


Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan Dec 2012

Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In binary classification problems, the area under the ROC curve (AUC), is an effective means of measuring the performance of your model. Most often, cross-validation is also used, in order to assess how the results will generalize to an independent data set. In order to evaluate the quality of an estimate for cross-validated AUC, we must obtain an estimate for its variance. For massive data sets, the process of generating a single performance estimate can be computationally expensive. Additionally, when using a complex prediction method, calculating the cross-validated AUC on even a relatively small data set can still require a …


A National Model Built With Partial Least Squares And Universal Kriging And Bootstrap-Based Measurement Error Correction Techniques: An Application To The Multi-Ethnic Study Of Atherosclerosis, Silas Bergen, Lianne Sheppard, Paul D. Sampson, Sun-Young Kim, Mark Richards, Sverre Vedal, Joel Kaufman, Adam A. Szpiro Dec 2012

A National Model Built With Partial Least Squares And Universal Kriging And Bootstrap-Based Measurement Error Correction Techniques: An Application To The Multi-Ethnic Study Of Atherosclerosis, Silas Bergen, Lianne Sheppard, Paul D. Sampson, Sun-Young Kim, Mark Richards, Sverre Vedal, Joel Kaufman, Adam A. Szpiro

UW Biostatistics Working Paper Series

Studies estimating health effects of long-term air pollution exposure often use a two-stage approach, building exposure models to assign individual-level exposures which are then used in regression analyses. This requires accurate exposure modeling and careful treatment of exposure measurement error. To illustrate the importance of carefully accounting for exposure model characteristics in two-stage air pollution studies, we consider a case study based on data from the Multi-Ethnic Study of Atherosclerosis (MESA). We present national spatial exposure models that use partial least squares and universal kriging to estimate annual average concentrations of four PM2.5 components: elemental carbon (EC), organic carbon (OC), …


Nonparametric Inference For Meta Analysis With Fixed Unknown, Study-Specific Parameters, Brian Claggett, Minge Xie, Lu Tian Nov 2012

Nonparametric Inference For Meta Analysis With Fixed Unknown, Study-Specific Parameters, Brian Claggett, Minge Xie, Lu Tian

Harvard University Biostatistics Working Paper Series

No abstract provided.