Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

UW Biostatistics Working Paper Series

Articles 61 - 90 of 102

Full-Text Articles in Biostatistics

Semiparametric Inferential Procedures For Comparing Multivariate Roc Curves With Interaction Terms, Liansheng Tang, Xiao-Hua Zhou Apr 2008

Semiparametric Inferential Procedures For Comparing Multivariate Roc Curves With Interaction Terms, Liansheng Tang, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Multivariate ROC curve models that include an interaction term be- tween biomarker type and false positive rate is important in comparative biomarker studies, because such interaction allows ROC curves of different biomarkers to cross each other. However, there has been limited work in drawing inference for comparing multivariate ROC curves, especially when the interaction terms are present. In this article we derive the asymptotic covariance of three estimators for multivariate ROC models. These covariance estimates have not been readily available in the literature, and bootstrap methods have to be used to obtain co- variance estimates. With the readily available variance …


Semi-Parametric Maximum Likelihood Estimates For Roc Curves Of Continuous-Scale Tests, Xiao-Hua Zhou, Huazhen Lin Apr 2008

Semi-Parametric Maximum Likelihood Estimates For Roc Curves Of Continuous-Scale Tests, Xiao-Hua Zhou, Huazhen Lin

UW Biostatistics Working Paper Series

No abstract provided.


Accommodating Covariates In Roc Analysis, Holly Janes, Gary M. Longton, Margaret Pepe Jan 2008

Accommodating Covariates In Roc Analysis, Holly Janes, Gary M. Longton, Margaret Pepe

UW Biostatistics Working Paper Series

Classification accuracy is the ability of a marker or diagnostic test to discriminate between two groups of individuals, cases and controls, and is commonly summarized using the receiver operating characteristic (ROC) curve. In studies of classification accuracy, there are often covariates that should be incorporated into the ROC analysis. We describe three different ways of using covariate informa- tion. For factors that affect marker observations among controls, we present a method for covariate adjustment. For factors that affect discrimination (ie the ROC curve), we describe methods for mod- elling the ROC curve as a function of covariates. Finally, for factors …


Estimation And Comparison Of Receiver Operating Characteristic Curves, Margaret Pepe, Gary M. Longton, Holly Janes Jan 2008

Estimation And Comparison Of Receiver Operating Characteristic Curves, Margaret Pepe, Gary M. Longton, Holly Janes

UW Biostatistics Working Paper Series

The receiver operating characteristic (ROC) curve displays the capacity of a marker or diagnostic test to discriminate between two groups of subjects, cases versus controls. We present a comprehensive suite of Stata commands for performing ROC analysis. Non-parametric, semiparametric and parametric estimators are calculated. Comparisons between curves are based on the area or partial area under the ROC curve. Alternatively pointwise comparisons between ROC curves or inverse ROC curves can be made. Options to adjust these analyses for covariates, and to perform ROC regression are described in a companion article. We use a unified framework by representing the ROC curve …


Longitudinal Data With Follow-Up Truncated By Death: Finding A Match Between Analysis Method And Research Aims, Brenda Kurland, Laura Lee Johnson, Paula Diehr Nov 2007

Longitudinal Data With Follow-Up Truncated By Death: Finding A Match Between Analysis Method And Research Aims, Brenda Kurland, Laura Lee Johnson, Paula Diehr

UW Biostatistics Working Paper Series

Diverse analysis approaches have been proposed to distinguish data missing due to death from nonresponse, and to summarize trajectories of longitudinal data truncated by death. We demonstrate how these analysis approaches arise from factorizations of the distribution of longitudinal data and survival information. Models are illustrated using hypothetical data examples (cognitive functioning in older adults, and quality of life under hospice care) and up to 10 annual assessments of longitudinal cognitive functioning data for 3814 participants in an observational study. For unconditional models, deaths do not occur, deaths are independent of the longitudinal response, or the unconditional longitudinal response averages …


A Parametric Roc Model Based Approach For Evaluating The Predictiveness Of Continuous Markers In Case-Control Studies, Ying Huang, Margaret Pepe Nov 2007

A Parametric Roc Model Based Approach For Evaluating The Predictiveness Of Continuous Markers In Case-Control Studies, Ying Huang, Margaret Pepe

UW Biostatistics Working Paper Series

The predictiveness curve shows the population distribution of risk endowed by a marker or risk prediction model. It provides a means for assessing the model's capacity for risk stratification. Methods for making inference about the predictiveness curve have been developed using cross-sectional or cohort data. Here we consider inference based on case-control studies and prior knowledge about prevalence or incidence of the outcome. We exploit the relationship between the ROC curve and the predictiveness curve given disease prevalence. Methods are developed for deriving the predictiveness curve from a parametric ROC model. Estimation of the whole range and of a portion …


Identifiability And Estimation Of Causal Effects In Randomized Trials With Noncompliance And Completely Non-Ignorable Missing-Data, Hua Chen, Zhi Geng, Xiao-Hua Zhou Nov 2007

Identifiability And Estimation Of Causal Effects In Randomized Trials With Noncompliance And Completely Non-Ignorable Missing-Data, Hua Chen, Zhi Geng, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

In this paper we first studied parameter identifiability in randomized clinical trials with noncompliance and missing outcomes. We showed that under certain conditions the parameters of interest were identifiable even under different types of completely non-ignorable missing data, that is, the missing mechanism depends on the outcome.We then derived their maximum likelihood (ML) and moment estimators and evaluated their finite-sample properties in simulation studies in terms of bias, efficiency and robustness. Our sensitive analysis showed the assumed non-ignorable missing- data model had an important impact on the estimated complier average causal effect (CACE) parameter. Our new method provides some new …


Nonparametric And Semiparametric Group Sequential Methods For Comparing Accuracy Of Diagnostic Tests, Liansheng Tang, Scott S. Emerson, Xiao-Hua Zhou Oct 2007

Nonparametric And Semiparametric Group Sequential Methods For Comparing Accuracy Of Diagnostic Tests, Liansheng Tang, Scott S. Emerson, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Comparison of the accuracy of two diagnostic tests using the receiver operating characteristic (ROC) curves from two diagnostic tests has been typically conducted using fixed sample designs. On the other hand, the human experimentation inherent in a comparison of diagnostic modalities argues for periodic monitoring of the accruing data to address many issues related to the ethics and efficiency of the medical study. To date, very little research has been done in the use of sequential sampling plans for comparative ROC studies, even when these studies may use expensive and unsafe diagnostic procedures. In this paper, we propose a nonparametric …


Roc Surfaces In The Presence Of Verification Bias, Yueh-Yun Chi, Xiao-Hua (Andrew) Zhou Sep 2007

Roc Surfaces In The Presence Of Verification Bias, Yueh-Yun Chi, Xiao-Hua (Andrew) Zhou

UW Biostatistics Working Paper Series

In diagnostic medicine, the Receiver Operating Characteristic (ROC) surface is one of the established tools for assessing the accuracy of a diagnostic test in discriminating three disease states, and the volume under the ROC surface has served as a summary index for diagnostic accuracy. In practice, the selection for definitive disease examination may be based on initial test measurements, and induces verification bias in the assessment. We propose here a nonparametric likelihood-based approach to construct the empirical ROC surface in the presence of differential verification, and to estimate the volume under the ROC surface. Estimators of the standard deviation are …


Reporting And Interpretation In Genome-Wide Association Studies, Jon Wakefield Jul 2007

Reporting And Interpretation In Genome-Wide Association Studies, Jon Wakefield

UW Biostatistics Working Paper Series

In the context of genome-wide association studies we critique a number of methods that have been suggested for flagging associations for further investigation. The p-value is by far the most commonly used measure, but requires careful calibration when the a priori probability of an association is small, and discards information by not considering the power associated with each test. The q-value is a frequentist method by which the false discovery rate (FDR) may be controlled. We advocate the use of the Bayes factor as a summary of the information in the data with respect to the comparison of the null …


Adjusting For Covariates In Studies Of Diagnostic, Screening, Or Prognostic Markers: An Old Concept In A New Setting, Holly Janes, Margaret Pepe May 2007

Adjusting For Covariates In Studies Of Diagnostic, Screening, Or Prognostic Markers: An Old Concept In A New Setting, Holly Janes, Margaret Pepe

UW Biostatistics Working Paper Series

The concept of covariate adjustment is well established in therapeutic and etiologic studies. However, it has received little attention in the growing area of medical research devoted to the development of markers for disease diagnosis, screening, or prognosis, where classification accuracy, rather than association, is of primary interest. In this paper, we demonstrate the need for covariate adjustment in studies of classification accuracy, discuss methods for adjusting for covariates, and distinguish covariate adjustment from several other related but fundamentally different uses for covariates. We draw analogies and contrasts throughout with studies of association.


Ecologic Studies Revisited, Jon Wakefield May 2007

Ecologic Studies Revisited, Jon Wakefield

UW Biostatistics Working Paper Series

Ecologic studies use data aggregated over groups, rather than data on individuals. Such studies are popular since they may make use of existing data bases, and can offer large exposure variation if based on broad geographical areas. Unfortunately the aggregation of data that defines ecologic studies results in a loss of information that can lead to ecologic bias. Specifically, ecologic bias arises from the inability of ecologic data to characterize within-area variability in exposures and confounders. We describe in detail particular forms of ecologic bias so that their potential impact on any particular study may be assessed. The only way …


Gamma Generalized Linear Models For Pharmacokinetic Data, Ruth Salway, Jon Wakefield May 2007

Gamma Generalized Linear Models For Pharmacokinetic Data, Ruth Salway, Jon Wakefield

UW Biostatistics Working Paper Series

This paper considers the modeling of single dose pharmacoki- netic data. Traditionally, so-called compartmental models have been used to analyze such data. Unfortunately the mean function of such models are sums of exponentials for which inference and computation may not be straightfor- ward. We present an alternative to these models based on generalized linear models, for which desirable statistical properties exist, with a logarithmic link and gamma distribution. The latter has a constant coefficient of variation which is often appropriate for pharmacokinetic data. Inference is convenient from either a likelihood or a Bayesian perspective. We consider models for both single …


Biomarker Evaluation Using The Controls As A Reference Population, Ying Huang, Margaret Pepe Apr 2007

Biomarker Evaluation Using The Controls As A Reference Population, Ying Huang, Margaret Pepe

UW Biostatistics Working Paper Series

The classification accuracy of a continuous marker is typically evaluated with the Receiver Operating Characteristic Curve. In this paper, we study an alternative conceptual framework, the "percentile value". In particular the controls only provide a reference distribution to standardize the marker. The analysis proceeds by analyzing the standardized marker only in cases. The approach is shown to be equivalent to ROC analysis. Advantages are that it provides a framework more familiar to biostatisticians and it opens up avenues for new statistical techniques in biomarker evaluation. We develop several new procedures based on this framework for comparing biomarkers and for comparing …


A Semiparametric Approach For The Nonparametric Transformation Survival Model With Multiple Covariates, Xiao Song, Shuangge Ma, Jian Huang, Xiao-Hua Zhou Dec 2006

A Semiparametric Approach For The Nonparametric Transformation Survival Model With Multiple Covariates, Xiao Song, Shuangge Ma, Jian Huang, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

The nonparametric transformation model for survival time that makes no parametric assumptions on both the transformation function and the error is appealing in its flexibility. The nonparametric transformation model makes no assumption on the forms of the transformation function and the error distribution. This model is appealing in its flexibility for modeling censored survival data. Current approaches for estimation of the regression parameters involve maximizing discontinuous objective functions, which are numerically infeasible to implement in the case of multiple covariates. Based on the partial rank estimator (Khan & Tamer, 2004), we propose a smoothed partial rank estimator which maximizes a …


Covariate Specific Roc Curve With Survival Outcome, Xiao Song, Xiao-Hua Zhou Sep 2006

Covariate Specific Roc Curve With Survival Outcome, Xiao Song, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

The receiver operating characteristic (ROC) curve has been extended to survival data recently, including the nonparametric approach by Heagerty, Lumley and Pepe (2000) and the semiparametric approach by Heagerty and Zheng (2005) using standard survival analysis techniques based on two different time-dependent ROC curve definitions. However, both approaches cannot adjust for the effect of covariates on the accuracy of the biomarker. To account for the covariate effect, we propose semiparametric models for covariate specific ROC curves corresponding to the two time-dependent ROC curve definitions, respectively. We show that the estimators are consistent and converge to Gaussian processes. In the case …


Generalized Confidence Intervals For The Ratio Or Difference Of Two Means For Lognormal Populations With Zeros, Yea-Hung Chen, Xiao-Hua Zhou Sep 2006

Generalized Confidence Intervals For The Ratio Or Difference Of Two Means For Lognormal Populations With Zeros, Yea-Hung Chen, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

We discuss in this article methods for analyzing lognormal data that may include zeros. Specifically, we are interested in interval estimation for the ratio or difference of the population means. We propose here two generalized pivotal (GP) approaches: a ``true'' GP method and an ``approximate'' GP method. Additionally, we propose two likelihood-based approaches: a signed log-likelihood ratio (SLLR) method and a modified SLLR method. Our simulation studies suggest that the approximate generalized pivotal approach outperforms all other known methods; it results in highly accurate coverage frequencies and fairly low bias, even in small sample settings.


Multiple Imputation - Review Of Theory, Implementation And Software, Ofer Harel, Xiao-Hua Zhou Sep 2006

Multiple Imputation - Review Of Theory, Implementation And Software, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Missing data is a common complication in data analysis. In many medical settings missing data can cause difficulties in estimation, precision and inference. Multiple imputation (MI) \cite{Rubin87} is a simulation based approach to deal with incomplete data. Although there are many different methods to deal with incomplete data, MI has become one of the leading methods. Since the late 80's we observed a constant increase in the use and publication of MI related research. This tutorial does not attempt to cover all the material concerning MI, but rather provides an overview and combines together the theory behind MI, the implementation …


Multiple Imputation For The Comparison Of Two Screening Tests In Two-Phase Alzheimer Studies, Ofer Harel, Xiao-Hua Zhou Sep 2006

Multiple Imputation For The Comparison Of Two Screening Tests In Two-Phase Alzheimer Studies, Ofer Harel, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Two-phase designs are common in epidemiological studies of dementia, and especially in Alzheimer research. In the first phase, all subjects are screened using a common screening test(s), while in the second phase, only a subset of these subjects is tested using a more definitive verification assessment, i.e. golden standard test. When comparing the accuracy of two screening tests in a two-phase study of dementia, inferences are commonly made using only the verified sample. It is well documented that in that case, there is a risk for bias, called verification bias. When the two screening tests have only two values (e.g. …


Evaluating Causal Effect Predictiveness Of Candidate Surrogate Endpoints, Peter B. Gilbert, Michael Hudgens Jul 2006

Evaluating Causal Effect Predictiveness Of Candidate Surrogate Endpoints, Peter B. Gilbert, Michael Hudgens

UW Biostatistics Working Paper Series

Most methods for evaluating surrogate endpoints measure validity in terms of net effects (i.e., treatment effects adjusted for the biomarker measured after randomization). Frangakis and Rubin (2002, Biometrics) criticized these approaches because net effects may reflect selection bias, and suggested an alternative definition of a surrogate endpoint (a "principal" surrogate) based on causal effects. For evaluating principal surrogates we introduce a causal effect predictiveness (CEP) surface, which quantifies how well causal treatment effects on the biomarker predict causal treatment effects on the clinical endpoint. The CEP surface is not identifiable in general due to missing potential outcomes. However, by incorporating …


Age- And Sex-Specific Transformations Of Health Status Measures To Incorporate Death, Ann M. Derleth, Paula Diehr Jun 2006

Age- And Sex-Specific Transformations Of Health Status Measures To Incorporate Death, Ann M. Derleth, Paula Diehr

UW Biostatistics Working Paper Series

Introduction: Measures of health status and physical function do not usually include a specific code for death. This can cause problems in longitudinal studies because analyses limited to survivors may bias the results. One approach is to recode the status variables to include a reasonable value for death. One method that has been used is to replace each scale value with the estimated probability that a person with this value will be “healthy”. “Healthy” has been defined as being above a particular threshold on the variable of interest one year later, or alternatively as being in excellent, very good, or …


Integrating The Predictiveness Of A Marker With Its Performance As A Classifier, Margaret S. Pepe, Ziding Feng, Ying Huang, Gary M. Longton, Ross Prentice, Ian M. Thompson, Yingye Zheng Jun 2006

Integrating The Predictiveness Of A Marker With Its Performance As A Classifier, Margaret S. Pepe, Ziding Feng, Ying Huang, Gary M. Longton, Ross Prentice, Ian M. Thompson, Yingye Zheng

UW Biostatistics Working Paper Series

There are two popular statistical approaches to biomarker evaluation. One models the risk of disease (or disease outcome) using, for example, logistic regression. A marker is useful if it has a strong effect on risk. The second evaluates classification performance using measures such as sensitivity, specificity, predictive values and ROC curves. There is controversy about which approach is most appropriate. Moreover, the two approaches often give contradictory results on the same data. We present a new graphic, the predictiveness curve, that complements the risk modeling approach. It assesses the usefulness of a risk model when applied to the population. In …


Reliability, Effect Size, And Responsiveness And Intraclass Correlation Of Health Status Measures Used In Randomized And Cluster-Randomized Trials, Paula Diehr, Lu Chen, Donald L. Patrick, Ziding Feng, Yutaka Yasui Mar 2006

Reliability, Effect Size, And Responsiveness And Intraclass Correlation Of Health Status Measures Used In Randomized And Cluster-Randomized Trials, Paula Diehr, Lu Chen, Donald L. Patrick, Ziding Feng, Yutaka Yasui

UW Biostatistics Working Paper Series

Background: New health status instruments are described by psychometric properties, such as Reliability, Effect Size, and Responsiveness. For cluster-randomized trials, another important statistic is the Intraclass Correlation for the instrument within clusters. Studies using better instruments can be performed with smaller sample sizes, but better instruments may be more expensive in terms of dollars, lost opportunities, or poorer data quality due to the response burden of longer instruments. Investigators often need to estimate the psychometric properties of a new instrument, or of an established instrument in a new setting. Optimal sample sizes for estimating these properties have not been studied …


Adjusting For Covariate Effects On Classification Accuracy Using The Covariate-Adjusted Roc Curve, Holly Janes, Margaret S. Pepe Mar 2006

Adjusting For Covariate Effects On Classification Accuracy Using The Covariate-Adjusted Roc Curve, Holly Janes, Margaret S. Pepe

UW Biostatistics Working Paper Series

Recent scientific and technological innovations have produced an abundance of potential markers which are being investigated for their use in disease screen- ing and diagnosis. In evaluating these markers, it is often necessary to account for covariates which are associated with the marker of interest. These covariates may include subject characteristics, expertise of the test operator, test proce- dures, or aspects of specimen handling. In this paper, we propose the AROC, a covariate-adjusted measure of the classification accuracy. The AROC is the common covariate-specific ROC curve, when the covariate does not affect dis- crimination, and a weighted average of covariate-specific …


Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes Mar 2006

Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes

UW Biostatistics Working Paper Series

Consider a placebo-controlled preventive HIV vaccine efficacy trial. An HIV amino acid sequence is measured from each volunteer who acquires HIV, and these sequences are aligned together with the reference HIV sequence represented in the vaccine. We develop genome scanning methods to identify HIV positions at which the amino acids in sequences from infected vaccine recipients tend to be more divergent from the corresponding reference amino acid than the amino acids in sequences from infected placebo recipients. We consider five two-sample test statistics, based on Euclidean, Mahalanobis, and Kullback-Leibler divergence measures. Weights are incorporated to reflect biological information contained in …


Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng Mar 2006

Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng

UW Biostatistics Working Paper Series

Consider a continuous marker for predicting a binary outcome. For example, serum concentration of prostate specific antigen (PSA) may be used to calculate the risk of finding prostate cancer in a biopsy. In this paper we argue that the predictive capacity of a marker has to do with the population distribution of risk given the marker and suggest a graphical tool, the predictiveness curve, that displays this distribution. The display provides a common meaningful scale for comparing markers that may not be comparable on their original scales. Some existing measures of predictiveness are shown to be summary indices derived from …


Comparison Of Haplotype-Based And Tree-Based Snp Imputation In Association Studies, James Y. Dai, Ingo Ruczinski, Michael Leblanc, Charles Kooperberg Jan 2006

Comparison Of Haplotype-Based And Tree-Based Snp Imputation In Association Studies, James Y. Dai, Ingo Ruczinski, Michael Leblanc, Charles Kooperberg

UW Biostatistics Working Paper Series

Missing single nucleotide polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we compare two haplotype-based imputation approaches and one regression tree-based imputation approach for association studies. The goal is to assess the imputation accuracy, and to evaluate the impact of imputation on parameter estimation. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to …


Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng Nov 2005

Estimating A Treatment Effect With Repeated Measurements Accounting For Varying Effectiveness Duration, Ying Qing Chen, Jingrong Yang, Su-Chun Cheng

UW Biostatistics Working Paper Series

To assess treatment efficacy in clinical trials, certain clinical outcomes are repeatedly measured for same subject over time. They can be regarded as function of time. The difference in their mean functions between the treatment arms usually characterises a treatment effect. Due to the potential existence of subject-specific treatment effectiveness lag and saturation times, erosion of treatment effect in the difference may occur during the observation period of time. Instead of using ad hoc parametric or purely nonparametric time-varying coefficients in statistical modeling, we first propose to model the treatment effectiveness durations, which are the varying time intervals between the …


Is The Number Of Sick Persons In A Cohort Constant Over Time?, Paula Diehr, Ann Derleth, Anne Newman, Liming Cai Oct 2005

Is The Number Of Sick Persons In A Cohort Constant Over Time?, Paula Diehr, Ann Derleth, Anne Newman, Liming Cai

UW Biostatistics Working Paper Series

Objectives: To estimate the number of persons in a cohort who are sick, over time.

Methods: We calculated the number of sick persons in the Cardiovascular Health Study (CHS), a cohort study of older adults followed up to 14 years, using eight definitions of “healthy” and “sick”. We projected the number in each health state over time for a birth cohort.

Results: The number of sick persons in CHS was approximately constant for 14 years, for all definitions of “sick”. The estimated number of sick persons in the birth cohort was approximately constant from ages 55-75, after which it decreased. …


Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang Jul 2005

Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang

UW Biostatistics Working Paper Series

Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.