Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

UW Biostatistics Working Paper Series

Discipline
Keyword
Publication Year

Articles 91 - 120 of 215

Full-Text Articles in Statistics and Probability

Nonparametric Heteroscedastic Transformation Regression Models For Skewed Data With An Application To Health Care Costs, Xiao-Hua Zhou, Huazhen Lin, Eric Johnson Apr 2008

Nonparametric Heteroscedastic Transformation Regression Models For Skewed Data With An Application To Health Care Costs, Xiao-Hua Zhou, Huazhen Lin, Eric Johnson

UW Biostatistics Working Paper Series

No abstract provided.


Semiparametric Inferential Procedures For Comparing Multivariate Roc Curves With Interaction Terms, Liansheng Tang, Xiao-Hua Zhou Apr 2008

Semiparametric Inferential Procedures For Comparing Multivariate Roc Curves With Interaction Terms, Liansheng Tang, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Multivariate ROC curve models that include an interaction term be- tween biomarker type and false positive rate is important in comparative biomarker studies, because such interaction allows ROC curves of different biomarkers to cross each other. However, there has been limited work in drawing inference for comparing multivariate ROC curves, especially when the interaction terms are present. In this article we derive the asymptotic covariance of three estimators for multivariate ROC models. These covariance estimates have not been readily available in the literature, and bootstrap methods have to be used to obtain co- variance estimates. With the readily available variance …


Semi-Parametric Maximum Likelihood Estimates For Roc Curves Of Continuous-Scale Tests, Xiao-Hua Zhou, Huazhen Lin Apr 2008

Semi-Parametric Maximum Likelihood Estimates For Roc Curves Of Continuous-Scale Tests, Xiao-Hua Zhou, Huazhen Lin

UW Biostatistics Working Paper Series

No abstract provided.


Multiple Imputation Of Timing Of Mother-To-Child Transmission Of Hiv, Elizabeth Brown, Ying Qing Chen Feb 2008

Multiple Imputation Of Timing Of Mother-To-Child Transmission Of Hiv, Elizabeth Brown, Ying Qing Chen

UW Biostatistics Working Paper Series

In this paper, we present a model for imputing timing of mother-to- child transmission (MTCT) of HIV. The method re ects the three modes of MTCT of HIV: in utero, during delivery and via breastfeeding and can accomodate shapes for the baseline hazard that vary between infants. Ad- ditionally, it allows that the majority of infants do not experience MTCT of HIV. Final analyses from the imputed data sets are combined in a mul- tiple imputation framework. The methods is illustrated on a large trial designed to assess the use of antibiotics in preventing MTCT of HIV and is validated …


Accommodating Covariates In Roc Analysis, Holly Janes, Gary M. Longton, Margaret Pepe Jan 2008

Accommodating Covariates In Roc Analysis, Holly Janes, Gary M. Longton, Margaret Pepe

UW Biostatistics Working Paper Series

Classification accuracy is the ability of a marker or diagnostic test to discriminate between two groups of individuals, cases and controls, and is commonly summarized using the receiver operating characteristic (ROC) curve. In studies of classification accuracy, there are often covariates that should be incorporated into the ROC analysis. We describe three different ways of using covariate informa- tion. For factors that affect marker observations among controls, we present a method for covariate adjustment. For factors that affect discrimination (ie the ROC curve), we describe methods for mod- elling the ROC curve as a function of covariates. Finally, for factors …


Estimation And Comparison Of Receiver Operating Characteristic Curves, Margaret Pepe, Gary M. Longton, Holly Janes Jan 2008

Estimation And Comparison Of Receiver Operating Characteristic Curves, Margaret Pepe, Gary M. Longton, Holly Janes

UW Biostatistics Working Paper Series

The receiver operating characteristic (ROC) curve displays the capacity of a marker or diagnostic test to discriminate between two groups of subjects, cases versus controls. We present a comprehensive suite of Stata commands for performing ROC analysis. Non-parametric, semiparametric and parametric estimators are calculated. Comparisons between curves are based on the area or partial area under the ROC curve. Alternatively pointwise comparisons between ROC curves or inverse ROC curves can be made. Options to adjust these analyses for covariates, and to perform ROC regression are described in a companion article. We use a unified framework by representing the ROC curve …


Model-Robust Bayesian Regression And The Sandwich Estimator, Adam A. Szpiro, Kenneth M. Rice, Thomas Lumley Dec 2007

Model-Robust Bayesian Regression And The Sandwich Estimator, Adam A. Szpiro, Kenneth M. Rice, Thomas Lumley

UW Biostatistics Working Paper Series

PLEASE NOTE THAT AN UPDATED VERSION OF THIS RESEARCH IS AVAILABLE AS WORKING PAPER 338 IN THE UNIVERSITY OF WASHINGTON BIOSTATISTICS WORKING PAPER SERIES (http://www.bepress.com/uwbiostat/paper338).

In applied regression problems there is often sufficient data for accurate estimation, but standard parametric models do not accurately describe the source of the data, so associated uncertainty estimates are not reliable. We describe a simple Bayesian approach to inference in linear regression that recovers least-squares point estimates while providing correct uncertainty bounds by explicitly recognizing that standard modeling assumptions need not be valid. Our model-robust development parallels frequentist estimating equations and leads to intervals …


Estimating Sensitivity And Specificity From A Phase 2 Biomarker Study That Allows For Early Termination, Margaret S. Pepe Phd Dec 2007

Estimating Sensitivity And Specificity From A Phase 2 Biomarker Study That Allows For Early Termination, Margaret S. Pepe Phd

UW Biostatistics Working Paper Series

Development of a disease screening biomarker involves several phases. In phase 2 its sensitivity and specificity is compared with established thresholds for minimally acceptable performance. Since we anticipate that most candidate markers will not prove to be useful and availability of specimens and funding is limited, early termination of a study is appropriate if accumulating data indicate that the marker is inadequate. Yet, for markers that complete phase 2, we seek estimates of sensitivity and specificity to proceed with the design of subsequent phase 3 studies.

We suggest early stopping criteria and estimation procedures that adjust for bias caused by …


Longitudinal Data With Follow-Up Truncated By Death: Finding A Match Between Analysis Method And Research Aims, Brenda Kurland, Laura Lee Johnson, Paula Diehr Nov 2007

Longitudinal Data With Follow-Up Truncated By Death: Finding A Match Between Analysis Method And Research Aims, Brenda Kurland, Laura Lee Johnson, Paula Diehr

UW Biostatistics Working Paper Series

Diverse analysis approaches have been proposed to distinguish data missing due to death from nonresponse, and to summarize trajectories of longitudinal data truncated by death. We demonstrate how these analysis approaches arise from factorizations of the distribution of longitudinal data and survival information. Models are illustrated using hypothetical data examples (cognitive functioning in older adults, and quality of life under hospice care) and up to 10 annual assessments of longitudinal cognitive functioning data for 3814 participants in an observational study. For unconditional models, deaths do not occur, deaths are independent of the longitudinal response, or the unconditional longitudinal response averages …


A Parametric Roc Model Based Approach For Evaluating The Predictiveness Of Continuous Markers In Case-Control Studies, Ying Huang, Margaret Pepe Nov 2007

A Parametric Roc Model Based Approach For Evaluating The Predictiveness Of Continuous Markers In Case-Control Studies, Ying Huang, Margaret Pepe

UW Biostatistics Working Paper Series

The predictiveness curve shows the population distribution of risk endowed by a marker or risk prediction model. It provides a means for assessing the model's capacity for risk stratification. Methods for making inference about the predictiveness curve have been developed using cross-sectional or cohort data. Here we consider inference based on case-control studies and prior knowledge about prevalence or incidence of the outcome. We exploit the relationship between the ROC curve and the predictiveness curve given disease prevalence. Methods are developed for deriving the predictiveness curve from a parametric ROC model. Estimation of the whole range and of a portion …


Identifiability And Estimation Of Causal Effects In Randomized Trials With Noncompliance And Completely Non-Ignorable Missing-Data, Hua Chen, Zhi Geng, Xiao-Hua Zhou Nov 2007

Identifiability And Estimation Of Causal Effects In Randomized Trials With Noncompliance And Completely Non-Ignorable Missing-Data, Hua Chen, Zhi Geng, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In this paper we first studied parameter identifiability in randomized clinical trials with noncompliance and missing outcomes. We showed that under certain conditions the parameters of interest were identifiable even under different types of completely non-ignorable missing data, that is, the missing mechanism depends on the outcome.We then derived their maximum likelihood (ML) and moment estimators and evaluated their finite-sample properties in simulation studies in terms of bias, efficiency and robustness. Our sensitive analysis showed the assumed non-ignorable missing- data model had an important impact on the estimated complier average causal effect (CACE) parameter. Our new method provides some new …


Nonparametric And Semiparametric Group Sequential Methods For Comparing Accuracy Of Diagnostic Tests, Liansheng Tang, Scott S. Emerson, Xiao-Hua Zhou Oct 2007

Nonparametric And Semiparametric Group Sequential Methods For Comparing Accuracy Of Diagnostic Tests, Liansheng Tang, Scott S. Emerson, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Comparison of the accuracy of two diagnostic tests using the receiver operating characteristic (ROC) curves from two diagnostic tests has been typically conducted using fixed sample designs. On the other hand, the human experimentation inherent in a comparison of diagnostic modalities argues for periodic monitoring of the accruing data to address many issues related to the ethics and efficiency of the medical study. To date, very little research has been done in the use of sequential sampling plans for comparative ROC studies, even when these studies may use expensive and unsafe diagnostic procedures. In this paper, we propose a nonparametric …


Roc Surfaces In The Presence Of Verification Bias, Yueh-Yun Chi, Xiao-Hua (Andrew) Zhou Sep 2007

Roc Surfaces In The Presence Of Verification Bias, Yueh-Yun Chi, Xiao-Hua (Andrew) Zhou

UW Biostatistics Working Paper Series

In diagnostic medicine, the Receiver Operating Characteristic (ROC) surface is one of the established tools for assessing the accuracy of a diagnostic test in discriminating three disease states, and the volume under the ROC surface has served as a summary index for diagnostic accuracy. In practice, the selection for definitive disease examination may be based on initial test measurements, and induces verification bias in the assessment. We propose here a nonparametric likelihood-based approach to construct the empirical ROC surface in the presence of differential verification, and to estimate the volume under the ROC surface. Estimators of the standard deviation are …


A Censored Multinomial Regression Model For Perinatal Mother To Child Transmission Of Hiv, Charlotte C. Gard, Elizabeth R. Brown Jul 2007

A Censored Multinomial Regression Model For Perinatal Mother To Child Transmission Of Hiv, Charlotte C. Gard, Elizabeth R. Brown

UW Biostatistics Working Paper Series

In studies designed to estimate rates of perinatal mother to child transmission of HIV, HIV assays are scheduled at multiple points in time. Still infection status for some infants at some time points is often unknown, particularly when interim analyses are conducted. Logistic regression and Cox proportional hazards regression are commonly used to estimate covariate-adjusted transmission rates, but their methods for handling missing data may be inadequate. Here, we propose using censored multinomial regression models to estimate cumulative and conditional rates of HIV transmission. Through simulation, we show that the proposed methods perform better than standard logistic models in terms …


Reporting And Interpretation In Genome-Wide Association Studies, Jon Wakefield Jul 2007

Reporting And Interpretation In Genome-Wide Association Studies, Jon Wakefield

UW Biostatistics Working Paper Series

In the context of genome-wide association studies we critique a number of methods that have been suggested for flagging associations for further investigation. The p-value is by far the most commonly used measure, but requires careful calibration when the a priori probability of an association is small, and discards information by not considering the power associated with each test. The q-value is a frequentist method by which the false discovery rate (FDR) may be controlled. We advocate the use of the Bayes factor as a summary of the information in the data with respect to the comparison of the null …


Evaluating The Roc Performance Of Markers For Future Events, Margaret Pepe, Yingye Zheng, Yuying Jin May 2007

Evaluating The Roc Performance Of Markers For Future Events, Margaret Pepe, Yingye Zheng, Yuying Jin

UW Biostatistics Working Paper Series

Receiver operating characteristic (ROC) curves play a central role in the evaluation of biomarkers and tests for disease diagnosis. Predictors for event time outcomes can also be evaluated with ROC curves, but the time lag between marker measurement and event time must be acknowledged. We discuss different definitions of time-dependent ROC curves in the context of real applications. Several approaches have been proposed for estimation. We contrast retrospective versus prospective methods in regards to assumptions and flexibility, including their capacities to incorporate censored data, competing risks and different sampling schemes. Applications to two datasets are presented.


Adjusting For Covariates In Studies Of Diagnostic, Screening, Or Prognostic Markers: An Old Concept In A New Setting, Holly Janes, Margaret Pepe May 2007

Adjusting For Covariates In Studies Of Diagnostic, Screening, Or Prognostic Markers: An Old Concept In A New Setting, Holly Janes, Margaret Pepe

UW Biostatistics Working Paper Series

The concept of covariate adjustment is well established in therapeutic and etiologic studies. However, it has received little attention in the growing area of medical research devoted to the development of markers for disease diagnosis, screening, or prognosis, where classification accuracy, rather than association, is of primary interest. In this paper, we demonstrate the need for covariate adjustment in studies of classification accuracy, discuss methods for adjusting for covariates, and distinguish covariate adjustment from several other related but fundamentally different uses for covariates. We draw analogies and contrasts throughout with studies of association.


Ecologic Studies Revisited, Jon Wakefield May 2007

Ecologic Studies Revisited, Jon Wakefield

UW Biostatistics Working Paper Series

Ecologic studies use data aggregated over groups, rather than data on individuals. Such studies are popular since they may make use of existing data bases, and can offer large exposure variation if based on broad geographical areas. Unfortunately the aggregation of data that defines ecologic studies results in a loss of information that can lead to ecologic bias. Specifically, ecologic bias arises from the inability of ecologic data to characterize within-area variability in exposures and confounders. We describe in detail particular forms of ecologic bias so that their potential impact on any particular study may be assessed. The only way …


Gamma Generalized Linear Models For Pharmacokinetic Data, Ruth Salway, Jon Wakefield May 2007

Gamma Generalized Linear Models For Pharmacokinetic Data, Ruth Salway, Jon Wakefield

UW Biostatistics Working Paper Series

This paper considers the modeling of single dose pharmacoki- netic data. Traditionally, so-called compartmental models have been used to analyze such data. Unfortunately the mean function of such models are sums of exponentials for which inference and computation may not be straightfor- ward. We present an alternative to these models based on generalized linear models, for which desirable statistical properties exist, with a logarithmic link and gamma distribution. The latter has a constant coefficient of variation which is often appropriate for pharmacokinetic data. Inference is convenient from either a likelihood or a Bayesian perspective. We consider models for both single …


Evaluating A Group Sequential Design In The Setting Of Nonproportional Hazards, Daniel L. Gillen, Scott S. Emerson May 2007

Evaluating A Group Sequential Design In The Setting Of Nonproportional Hazards, Daniel L. Gillen, Scott S. Emerson

UW Biostatistics Working Paper Series

Group sequential methods have been widely described and implemented in a clinical trial setting where parametric and semiparametric models are deemed suitable. In these situations, the evaluation of the operating characteristics of a group sequential stopping rule remains relatively straightforward. However, in the presence of nonproportional hazards survival data nonparametric methods are often used, and the evaluation of stopping rules is no longer a trivial task. Specifically, nonparametric test statistics do not necessarily correspond to a parameter of clinical interest, thus making it difficult to characterize alternatives at which operating characteristics are to be computed. We describe an approach for …


Biomarker Evaluation Using The Controls As A Reference Population, Ying Huang, Margaret Pepe Apr 2007

Biomarker Evaluation Using The Controls As A Reference Population, Ying Huang, Margaret Pepe

UW Biostatistics Working Paper Series

The classification accuracy of a continuous marker is typically evaluated with the Receiver Operating Characteristic Curve. In this paper, we study an alternative conceptual framework, the "percentile value". In particular the controls only provide a reference distribution to standardize the marker. The analysis proceeds by analyzing the standardized marker only in cases. The approach is shown to be equivalent to ROC analysis. Advantages are that it provides a framework more familiar to biostatisticians and it opens up avenues for new statistical techniques in biomarker evaluation. We develop several new procedures based on this framework for comparing biomarkers and for comparing …


What Is The Best Reference Rna? And Other Questions Regarding The Design And Analysis Of Two-Color Microarray Experiments, Kathleen F. Kerr, Kyle A. Serikawa, Caimiao Wei, Mette A. Peters, Roger E. Bumgarner Apr 2007

What Is The Best Reference Rna? And Other Questions Regarding The Design And Analysis Of Two-Color Microarray Experiments, Kathleen F. Kerr, Kyle A. Serikawa, Caimiao Wei, Mette A. Peters, Roger E. Bumgarner

UW Biostatistics Working Paper Series

The reference design is a practical and popular choice for microarray studies using two-color platforms. In the reference design, the reference RNA uses half of all array resources, leading investigators to ask: What is the best reference RNA? We propose a novel method for evaluating reference RNAs and present the results of an experiment that was specially designed to evaluate three common choices of reference RNA. We found no compelling evidence in favor of any particular reference. In particular, a commercial reference showed no advantage in our data. Our experimental design also enabled a new way to test the effectiveness …


Power Boosting In Genome-Wide Studies Via Methods For Multivariate Outcomes, Mary J. Emond Feb 2007

Power Boosting In Genome-Wide Studies Via Methods For Multivariate Outcomes, Mary J. Emond

UW Biostatistics Working Paper Series

Whole-genome studies are becoming a mainstay of biomedical research. Examples include expression array experiments, comparative genomic hybridization analyses and large case-control studies for detecting polymorphism/disease associations. The tactic of applying a regression model to every locus to obtain test statistics is useful in such studies. However, this approach ignores potential correlation structure in the data that could be used to gain power, particularly when a Bonferroni correction is applied to adjust for multiple testing. In this article, we propose using regression techniques for misspecified multivariate outcomes to increase statistical power over independence-based modeling at each locus. Even when the outcome …


A Semiparametric Approach For The Nonparametric Transformation Survival Model With Multiple Covariates, Xiao Song, Shuangge Ma, Jian Huang, Xiao-Hua Zhou Dec 2006

A Semiparametric Approach For The Nonparametric Transformation Survival Model With Multiple Covariates, Xiao Song, Shuangge Ma, Jian Huang, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

The nonparametric transformation model for survival time that makes no parametric assumptions on both the transformation function and the error is appealing in its flexibility. The nonparametric transformation model makes no assumption on the forms of the transformation function and the error distribution. This model is appealing in its flexibility for modeling censored survival data. Current approaches for estimation of the regression parameters involve maximizing discontinuous objective functions, which are numerically infeasible to implement in the case of multiple covariates. Based on the partial rank estimator (Khan & Tamer, 2004), we propose a smoothed partial rank estimator which maximizes a …


Large Cluster Asymptotics For Gee: Working Correlation Models, Hyoju Chung, Thomas Lumley Oct 2006

Large Cluster Asymptotics For Gee: Working Correlation Models, Hyoju Chung, Thomas Lumley

UW Biostatistics Working Paper Series

This paper presents large cluster asymptotic results for generalized estimating equations. The complexity of working correlation model is characterized in terms of the number of working correlation components to be estimated. When the cluster size is relatively large, we may encounter a situation where a high-dimensional working correlation matrix is modeled and estimated from the data. In the present asymptotic setting, the cluster size and the complexity of working correlation model grow with the number of independent clusters. We show the existence, weak consistency and asymptotic normality of marginal regression parameter estimators using the results of empirical process theory and …


Statistical Analysis Of Air Pollution Panel Studies: An Illustration, Holly Janes, Lianne Sheppard, Kristen Shepherd Oct 2006

Statistical Analysis Of Air Pollution Panel Studies: An Illustration, Holly Janes, Lianne Sheppard, Kristen Shepherd

UW Biostatistics Working Paper Series

The panel study design is commonly used to evaluate the short-term health effects of air pollution. Standard statistical methods for analyzing longitudinal data are available, but the literature reveals that the techniques are not well understood by practitioners. We illustrate these methods using data from the 1999 to 2002 Seattle panel study. Marginal, conditional, and transitional approaches for modeling longitudinal data are reviewed and contrasted with respect to their parameter interpretation and methods for accounting for correlation and dealing with missing data. We also discuss and illustrate techniques for controlling for time-dependent and time-independent confounding, and for exploring and summarizing …


Covariate Specific Roc Curve With Survival Outcome, Xiao Song, Xiao-Hua Zhou Sep 2006

Covariate Specific Roc Curve With Survival Outcome, Xiao Song, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

The receiver operating characteristic (ROC) curve has been extended to survival data recently, including the nonparametric approach by Heagerty, Lumley and Pepe (2000) and the semiparametric approach by Heagerty and Zheng (2005) using standard survival analysis techniques based on two different time-dependent ROC curve definitions. However, both approaches cannot adjust for the effect of covariates on the accuracy of the biomarker. To account for the covariate effect, we propose semiparametric models for covariate specific ROC curves corresponding to the two time-dependent ROC curve definitions, respectively. We show that the estimators are consistent and converge to Gaussian processes. In the case …


Generalized Confidence Intervals For The Ratio Or Difference Of Two Means For Lognormal Populations With Zeros, Yea-Hung Chen, Xiao-Hua Zhou Sep 2006

Generalized Confidence Intervals For The Ratio Or Difference Of Two Means For Lognormal Populations With Zeros, Yea-Hung Chen, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

We discuss in this article methods for analyzing lognormal data that may include zeros. Specifically, we are interested in interval estimation for the ratio or difference of the population means. We propose here two generalized pivotal (GP) approaches: a ``true'' GP method and an ``approximate'' GP method. Additionally, we propose two likelihood-based approaches: a signed log-likelihood ratio (SLLR) method and a modified SLLR method. Our simulation studies suggest that the approximate generalized pivotal approach outperforms all other known methods; it results in highly accurate coverage frequencies and fairly low bias, even in small sample settings.


Multiple Imputation - Review Of Theory, Implementation And Software, Ofer Harel, Xiao-Hua Zhou Sep 2006

Multiple Imputation - Review Of Theory, Implementation And Software, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Missing data is a common complication in data analysis. In many medical settings missing data can cause difficulties in estimation, precision and inference. Multiple imputation (MI) \cite{Rubin87} is a simulation based approach to deal with incomplete data. Although there are many different methods to deal with incomplete data, MI has become one of the leading methods. Since the late 80's we observed a constant increase in the use and publication of MI related research. This tutorial does not attempt to cover all the material concerning MI, but rather provides an overview and combines together the theory behind MI, the implementation …


Multiple Imputation For The Comparison Of Two Screening Tests In Two-Phase Alzheimer Studies, Ofer Harel, Xiao-Hua Zhou Sep 2006

Multiple Imputation For The Comparison Of Two Screening Tests In Two-Phase Alzheimer Studies, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Two-phase designs are common in epidemiological studies of dementia, and especially in Alzheimer research. In the first phase, all subjects are screened using a common screening test(s), while in the second phase, only a subset of these subjects is tested using a more definitive verification assessment, i.e. golden standard test. When comparing the accuracy of two screening tests in a two-phase study of dementia, inferences are commonly made using only the verified sample. It is well documented that in that case, there is a risk for bias, called verification bias. When the two screening tests have only two values (e.g. …