Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (567)
- Statistical Methodology (362)
- Statistical Theory (336)
- Statistical Models (242)
- Medicine and Health Sciences (176)
-
- Survival Analysis (147)
- Public Health (142)
- Epidemiology (99)
- Life Sciences (92)
- Longitudinal Data Analysis and Time Series (89)
- Genetics and Genomics (88)
- Clinical Trials (82)
- Microarrays (78)
- Multivariate Analysis (78)
- Applied Mathematics (57)
- Genetics (57)
- Numerical Analysis and Computation (57)
- Bioinformatics (50)
- Computational Biology (50)
- Categorical Data Analysis (49)
- Design of Experiments and Sample Surveys (39)
- Clinical Epidemiology (36)
- Diseases (30)
- Disease Modeling (28)
- Medical Specialties (23)
- Health Services Research (17)
- Applied Statistics (13)
- Vital and Health Statistics (11)
- Keyword
-
- Causal inference (30)
- Cross-validation (25)
- Prediction (23)
- Genetics (21)
- Longitudinal data (19)
-
- Survival analysis (16)
- Classification (14)
- Influence curve (14)
- Model selection (14)
- Sensitivity (14)
- Bootstrap (13)
- Gene expression (13)
- Clinical trials (12)
- Targeted maximum likelihood estimation (12)
- Counterfactual (11)
- Efficient influence curve (11)
- Multiple testing (11)
- Confounding (10)
- Loss function (10)
- Missing data (10)
- Variable selection (10)
- Causal effect (9)
- Estimating equation (9)
- Measurement error (9)
- Regression (9)
- Specificity (9)
- Adjusted p-value (8)
- Air pollution (8)
- Asymptotic linearity (8)
- Censoring (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (242)
- UW Biostatistics Working Paper Series (215)
- Harvard University Biostatistics Working Paper Series (212)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (178)
- The University of Michigan Department of Biostatistics Working Paper Series (111)
Articles 271 - 300 of 1108
Full-Text Articles in Statistics and Probability
A General Regression Framework For A Secondary Outcome In Case-Control Studies, Eric J. Tchetgen Tchetgen
A General Regression Framework For A Secondary Outcome In Case-Control Studies, Eric J. Tchetgen Tchetgen
Harvard University Biostatistics Working Paper Series
No abstract provided.
In Praise Of Simplicity Not Mathematistry! Ten Simple Powerful Ideas For The Statistical Scientist, Roderick J. Little
In Praise Of Simplicity Not Mathematistry! Ten Simple Powerful Ideas For The Statistical Scientist, Roderick J. Little
The University of Michigan Department of Biostatistics Working Paper Series
Ronald Fisher was by all accounts a first-rate mathematician, but he saw himself as a scientist, not a mathematician, and he railed against what George Box called (in his Fisher lecture) "mathematistry". Mathematics is the indispensable foundation for statistics, but our subject is constantly under assault by people who want to turn statistics into a branch of mathematics, making the subject as impenetrable to non-mathematicians as possible. Valuing simplicity, I describe ten simple and powerful ideas that have influenced my thinking about statistics, in my areas of research interest: missing data, causal inference, survey sampling, and statistical modeling in general. …
Visualizing Longitudinal Data With Dropouts, Mithat Gonen
Visualizing Longitudinal Data With Dropouts, Mithat Gonen
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
A triangle plot is proposed to display longitudinal data with dropouts. The triangle plot is a tool of data visualization that can also serve as a graphical check for informativeness of the dropout process. There are similarities between the lasagna plot and the triangle plot but the explicit use of dropout time as an axis is an advantage of the triangle plot over the more commonly used graphical strategies for longitudinal data. It is possible to interpret the triangle plot as a trellis plot 1 which gives rise to several extensions such as the triangle histogram and the triangle boxplot. …
On The Simulation Of Longitudinal Discrete Data With Specified Marginal Means And First-Order Antedependence, Matthew Guerra, Justine Shults
On The Simulation Of Longitudinal Discrete Data With Specified Marginal Means And First-Order Antedependence, Matthew Guerra, Justine Shults
UPenn Biostatistics Working Papers
We propose a straightforward approach for simulation of discrete random variables with overdispersion, specified marginal means, and product correlations that are plausible for longitudinal data with equal, or unequal, temporal spacings. The method stems from results we prove for variables with first-order antedependence and linearity of the conditional expectations. The proposed approach will be especially useful for assessment of methods such as generalized estimating equations, which specify separate models for the marginal means and correlation structure of measurements on a subject.
Mixtures Of Receiver Operating Characteristic Curves, Mithat Gonen
Mixtures Of Receiver Operating Characteristic Curves, Mithat Gonen
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
Rationale and Objectives: ROC curves are ubiquitous in the analysis of imaging metrics as markers of both diagnosis and prognosis. While empirical estimation of ROC curves remains the most popular method, there are several reasons to consider smooth estimates based on a parametric model.
Materials and Methods: A mixture model is considered for modeling the distribution of the marker in the diseased population motivated by the biological observation that there is more heterogeneity in the diseased population than there is in the normal one. It is shown that this model results in an analytically tractable ROC curve which is itself …
A Frailty Approach For Survival Analysis With Error-Prone Covariate, Sehee Kim, Yi Li, Donna Spiegelman
A Frailty Approach For Survival Analysis With Error-Prone Covariate, Sehee Kim, Yi Li, Donna Spiegelman
The University of Michigan Department of Biostatistics Working Paper Series
This paper discovers an inherent relationship between the survival model with covariate measurement error and the frailty model. The discovery motivates our using a frailty-based estimating equation to draw inference for the proportional hazards model with error-prone covariates. Our established framework accommodates general distributional structures for the error-prone covariates, not restricted to a linear additive measurement error model or Gaussian measurement error. When the conditional distribution of the frailty given the surrogate is unknown, it is estimated through a semiparametric copula function. The proposed copula-based approach enables us to fit flexible measurement error models without the curse of dimensionality as …
Ultrahigh Dimensional Time Course Feature Selection, Peirong Xu, Lixing Zhu, Yi Li
Ultrahigh Dimensional Time Course Feature Selection, Peirong Xu, Lixing Zhu, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Statistical challenges arise from modern biomedical studies that produce time course genomic data with ultrahigh dimensions. In a renal cancer study that motivated this paper, the pharmacokinetic measures of a tumor suppressor (CCI-779) and expression levels of 12625 genes were measured for each of 33 patients at 8 and 16 weeks after the start of treatments, with the goal of identifying predictive gene transcripts and the interactions with time in peripheral blood mononuclear cells for pharmacokinetics over the time course. The resulting dataset defies analysis even with regularized regression. Although some remedies have been proposed for both linear and generalized …
Selection Of Latent Variables For Multiple Mixed-Outcome Models, Ling Zhou, Huazhen Lin, Xin-Yuan Song, Yi Li
Selection Of Latent Variables For Multiple Mixed-Outcome Models, Ling Zhou, Huazhen Lin, Xin-Yuan Song, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Latent variable models have been widely used for modeling the dependence structure of multiple outcomes data. As the formulation of a latent variable model is often unknown a priori, misspecification could distort the dependence structure and lead to unreliable model inference. More- over, the multiple outcomes are often of varying types (e.g., continuous and ordinal), which presents analytical challenges. In this article, we present a class of general latent variable models that can accommodate mixed types of outcomes, and further propose a novel selection approach that simultaneously selects latent variables and estimates model parameters. We show that the proposed estimators …
Semiparametric Latent Variable Transformation Models For Multiple Mixed Outcomes, Huazhen Lin, Ling Zhou, Robert Elashoff, Yi Li
Semiparametric Latent Variable Transformation Models For Multiple Mixed Outcomes, Huazhen Lin, Ling Zhou, Robert Elashoff, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
No abstract provided.
Semiparametric Transformation Models For Semicompeting Survival Data, Huazhen Lin, Ling Zhou, Chunhong Li, Yi Li
Semiparametric Transformation Models For Semicompeting Survival Data, Huazhen Lin, Ling Zhou, Chunhong Li, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Semicompeting risk outcome data, e.g. time to disease progression and time to death, are commonly collected in clinical trials, but complicated analytical tools hamper the analysis and the interpretation of the results. We propose a novel semiparametric transformation model for such data. Compared with the existing models, our model is advantageous in the following distinctive ways. First, it allows us to provide direct estimators of the regression analysis and the association parameter. Second, the measure of surrogacy, for example, the proportion of treatment effect and relative effect, can also be directly obtained. We propose a two-stage estimation procedure for inference …
Score Test Variable Screening, Sihai Dave Zhao, Yi Li
Score Test Variable Screening, Sihai Dave Zhao, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Variable screening has emerged as a crucial first step in the analysis of high-throughput data, but existing procedures can be computationally cumbersome, difficult to justify theoretically, or inapplicable to certain types of analyses. Motivated by a high-dimensional censored quantile regression problem in multiple myeloma genomics, this paper makes three contributions. First, we establish a score test-based screening framework, which is widely applicable, extremely computationally efficient, and relatively simple to justify. Secondly, we propose a resampling-based procedure for selecting the number of variables to retain after screening according to the principle of reproducibility. Finally, we propose a new iterative score test …
A Latent Variable Transformation Model Approach For Exploring Dysphagia, Anna Snavely, David P. Harrington, Yi Li
A Latent Variable Transformation Model Approach For Exploring Dysphagia, Anna Snavely, David P. Harrington, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
No abstract provided.
Covariance-Enhanced Discriminant Analysis, Peirong Xu, Ji Zhu, Lixing Zhu, Yi Li
Covariance-Enhanced Discriminant Analysis, Peirong Xu, Ji Zhu, Lixing Zhu, Yi Li
The University of Michigan Department of Biostatistics Working Paper Series
Linear discriminant analysis (LDA), a classical method in pattern recognition and machine learning, has been widely used to characterize or separate multiple classes via linear combinations of features. However, the high-dimensionality of the high-throughput features obtained from modern biological experiments, for example, microarray or proteomics, defies traditional discriminant analysis techniques. The possible interfeature correlations present additional challenges and are often under-utilized in modeling. In this paper, by incorporating the possible inter-feature correlations, we propose a Covariance-Enhanced Discriminant Analysis (CEDA) method that simultaneously and consistently selects informative features and identifies the corresponding discriminable classes. We show that, under mild regularity conditions, …
Optimal Spatial Prediction Using Ensemble Machine Learning, Molly M. Davies, Mark J. Van Der Laan
Optimal Spatial Prediction Using Ensemble Machine Learning, Molly M. Davies, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Spatial prediction is an important problem in many scientific disciplines. Super Learner is an ensemble prediction approach related to stacked generalization that uses cross-validation to search for the optimal predictor amongst all convex combinations of a heterogeneous candidate set. It has been applied to non-spatial data, where theoretical results demonstrate it will perform asymptotically at least as well as the best candidate under consideration. We review these optimality properties and discuss the assumptions required in order for them to hold for spatial prediction problems. We present results of a simulation study confirming Super Learner works well in practice under a …
Relating Nanoparticle Properties To Biological Outcomes In Exposure Escalation Experiments, Trina Patel, Cecile Low-Kam, Zhaoxia Ji, Haiyuan Zhang, Tian Xia, Andre E. Nel, Jeffrey I. Zinc, Donatello Telesca
Relating Nanoparticle Properties To Biological Outcomes In Exposure Escalation Experiments, Trina Patel, Cecile Low-Kam, Zhaoxia Ji, Haiyuan Zhang, Tian Xia, Andre E. Nel, Jeffrey I. Zinc, Donatello Telesca
COBRA Preprint Series
A fundamental goal in nano-toxicology is that of identifying particle physical and chemical properties, which are likely to explain biological hazard. The first line of screening for potentially adverse outcomes often consists of exposure escalation experiments, involving the exposure of micro-organisms or cell lines to a battery of nanomaterials. We discuss a modeling strategy, that relates the outcome of an exposure escalation experiment to nanoparticle properties. Our approach makes use of a hierarchical decision process, where we jointly identify particles that initiate adverse biological outcomes and explain the probability of this event in terms of the particle physico-chemical descriptors. The …
A Regionalized National Universal Kriging Model Using Partial Least Squares Regression For Estimating Annual Pm2.5 Concentrations In Epidemiology, Paul D. Sampson, Mark Richards, Adam A. Szpiro, Silas Bergen, Lianne Sheppard, Timothy V. Larson, Joel Kaufman
A Regionalized National Universal Kriging Model Using Partial Least Squares Regression For Estimating Annual Pm2.5 Concentrations In Epidemiology, Paul D. Sampson, Mark Richards, Adam A. Szpiro, Silas Bergen, Lianne Sheppard, Timothy V. Larson, Joel Kaufman
UW Biostatistics Working Paper Series
Many cohort studies in environmental epidemiology require accurate modeling and prediction of fine scale spatial variation in ambient air quality across the U.S. This modeling requires the use of small spatial scale geographic or “land use” regression covariates and some degree of spatial smoothing. Furthermore, the details of the prediction of air quality by land use regression and the spatial variation in ambient air quality not explained by this regression should be allowed to vary across the continent due to the large scale heterogeneity in topography, climate, and sources of air pollution. This paper introduces a regionalized national universal kriging …
Sensitivity Analysis For Causal Inference Under Unmeasured Confounding And Measurement Error Problems, Iván Díaz, Mark J. Van Der Laan
Sensitivity Analysis For Causal Inference Under Unmeasured Confounding And Measurement Error Problems, Iván Díaz, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper we present a sensitivity analysis for drawing inferences about parameters that are not estimable from observed data without additional assumptions. We present the methodology using two different examples: a causal parameter that is not identifiable due to violations of the randomization assumption, and a parameter that is not estimable in the nonparametric model due to measurement error. Existing methods for tackling these problems assume a parametric model for the type of violation to the identifiability assumption, and require the development of new estimators and inference for every new model. The method we present can be used in …
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In binary classification problems, the area under the ROC curve (AUC), is an effective means of measuring the performance of your model. Most often, cross-validation is also used, in order to assess how the results will generalize to an independent data set. In order to evaluate the quality of an estimate for cross-validated AUC, we must obtain an estimate for its variance. For massive data sets, the process of generating a single performance estimate can be computationally expensive. Additionally, when using a complex prediction method, calculating the cross-validated AUC on even a relatively small data set can still require a …
A National Model Built With Partial Least Squares And Universal Kriging And Bootstrap-Based Measurement Error Correction Techniques: An Application To The Multi-Ethnic Study Of Atherosclerosis, Silas Bergen, Lianne Sheppard, Paul D. Sampson, Sun-Young Kim, Mark Richards, Sverre Vedal, Joel Kaufman, Adam A. Szpiro
A National Model Built With Partial Least Squares And Universal Kriging And Bootstrap-Based Measurement Error Correction Techniques: An Application To The Multi-Ethnic Study Of Atherosclerosis, Silas Bergen, Lianne Sheppard, Paul D. Sampson, Sun-Young Kim, Mark Richards, Sverre Vedal, Joel Kaufman, Adam A. Szpiro
UW Biostatistics Working Paper Series
Studies estimating health effects of long-term air pollution exposure often use a two-stage approach, building exposure models to assign individual-level exposures which are then used in regression analyses. This requires accurate exposure modeling and careful treatment of exposure measurement error. To illustrate the importance of carefully accounting for exposure model characteristics in two-stage air pollution studies, we consider a case study based on data from the Multi-Ethnic Study of Atherosclerosis (MESA). We present national spatial exposure models that use partial least squares and universal kriging to estimate annual average concentrations of four PM2.5 components: elemental carbon (EC), organic carbon (OC), …
Nonparametric Inference For Meta Analysis With Fixed Unknown, Study-Specific Parameters, Brian Claggett, Minge Xie, Lu Tian
Nonparametric Inference For Meta Analysis With Fixed Unknown, Study-Specific Parameters, Brian Claggett, Minge Xie, Lu Tian
Harvard University Biostatistics Working Paper Series
No abstract provided.
Statistical Inference When Using Data Adaptive Estimators Of Nuisance Parameters, Mark J. Van Der Laan
Statistical Inference When Using Data Adaptive Estimators Of Nuisance Parameters, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In order to be concrete we focus on estimation of the treatment specific mean, controlling for all measured baseline covariates, based on observing n independent and identically distributed copies of a random variable consisting of baseline covariates, a subsequently assigned binary treatment, and a final outcome. The statistical model only assumes possible restrictions on the conditional distribution of treatment, given the covariates, the so called propensity score. Estimators of the treatment specific mean involve estimation of the propensity score and/or estimation of the conditional mean of the outcome, given the treatment and covariates. In order to make these estimators asymptotically …
Treatment Selections Using Risk-Benefit Profiles Based On Data From Comparative Randomized Clinical Trials With Multiple Endpoints, Brian Claggett, Lu Tian, Davide Castagno, L. J. Wei
Treatment Selections Using Risk-Benefit Profiles Based On Data From Comparative Randomized Clinical Trials With Multiple Endpoints, Brian Claggett, Lu Tian, Davide Castagno, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
Likelihood Ratio Tests For The Mean Structure Of Correlated Functional Processes, Ana-Maria Staicu, Yingxing Li, Ciprian Crainiceanu, David M. Ruppert
Likelihood Ratio Tests For The Mean Structure Of Correlated Functional Processes, Ana-Maria Staicu, Yingxing Li, Ciprian Crainiceanu, David M. Ruppert
Johns Hopkins University, Dept. of Biostatistics Working Papers
The paper introduces a general framework for testing hypotheses about the structure of the mean function of complex functional processes. Important particular cases of the proposed framework are: 1) testing the null hypotheses that the mean of a functional process is parametric against a nonparametric alternative; and 2) testing the null hypothesis that the means of two possibly correlated functional processes are equal or differ by only a simple parametric function. A global pseudo likelihood ratio test is proposed and its asymptotic distribution is derived. The size and power properties of the test are confirmed in realistic simulation scenarios. Finite …
Longitudinal Functional Models With Structured Penalties, Madan G. Kundu, Jaroslaw Harezlak, Timothy W. Randolph
Longitudinal Functional Models With Structured Penalties, Madan G. Kundu, Jaroslaw Harezlak, Timothy W. Randolph
Johns Hopkins University, Dept. of Biostatistics Working Papers
Collection of functional data is becoming increasingly common including longitudinal observations in many studies. For example, we use magnetic resonance (MR) spectra collected over a period of time from late stage HIV patients. MR spectroscopy (MRS) produces a spectrum which is a mixture of metabolite spectra, instrument noise and baseline profile. Analysis of such data typically proceeds in two separate steps: feature extraction and regression modeling. In contrast, a recently-proposed approach, called partially empirical eigenvectors for regression (PEER) (Randolph, Harezlak and Feng, 2012), for functional linear models incorporates a priori knowledge via a scientifically-informed penalty operator in the regression function …
Pls-Rog: Partial Least Squares With Rank Order Of Groups, Hiroyuki Yamamoto
Pls-Rog: Partial Least Squares With Rank Order Of Groups, Hiroyuki Yamamoto
COBRA Preprint Series
Partial least squares (PLS), which is an unsupervised dimensionality reduction method, has been widely used in metabolomics. PLS can separate score depend on groups in a low dimensional subspace. However, this cannot use the information about rank order of groups. This information is often provided in which concentration of administered drugs to animals is gradually varies. In this study, we proposed partial least squares for rank order of groups (PLS-ROG). PLS-ROG can consider both separation and rank order of groups.
Statistical Hypothesis Test Of Factor Loading In Principal Component Analysis And Its Application To Metabolite Set Enrichment Analysis, Hiroyuki Yamamoto, Tamaki Fujimori, Hajime Sato, Gen Ishikawa, Kenjiro Kami, Yoshiaki Ohashi
Statistical Hypothesis Test Of Factor Loading In Principal Component Analysis And Its Application To Metabolite Set Enrichment Analysis, Hiroyuki Yamamoto, Tamaki Fujimori, Hajime Sato, Gen Ishikawa, Kenjiro Kami, Yoshiaki Ohashi
COBRA Preprint Series
Principal component analysis (PCA) has been widely used to visualize high-dimensional metabolomic data in a two- or three-dimensional subspace. In metabolomics, some metabolites (e.g. top 10 metabolites) have been subjectively selected when using factor loading in PCA, and biological inferences for these metabolites are made. However, this approach is possible to lead biased biological inferences because these metabolites are not objectively selected by statistical criterion. We proposed a statistical procedure to pick up metabolites by statistical hypothesis test of factor loading in PCA and make biological inferences by metabolite set enrichment analysis (MSEA) for these significant metabolites. This procedure depends …
Decline In Health For Older Adults: 5-Year Change In 13 Key Measures Of Standardized Health, Paula H. Diehr, Stephen M. Thielke, Anne B. Newman, Calvin H. Hirsch, Russell Tracy
Decline In Health For Older Adults: 5-Year Change In 13 Key Measures Of Standardized Health, Paula H. Diehr, Stephen M. Thielke, Anne B. Newman, Calvin H. Hirsch, Russell Tracy
UW Biostatistics Working Paper Series
Introduction
The health of older adults declines over time, but there are many ways of measuring health. We examined whether all measures declined at the same rate, or whether some aspects of health were less sensitive to aging than others.
Methods
We compared the decline in 13 measures of physical, mental, and functional health from the Cardiovascular Health Study: hospitalization, bed days, cognition, extremity strength, feelings about life as a whole, satisfaction with the purpose of life, self-rated health, depression, digit symbol substitution test, grip strength, ADLs, IADLs, and gait speed. Each measure was standardized against self-rated health. We compared …
Methods For Evaluating Prediction Performance Of Biomarkers And Tests, Margaret Pepe, Holly Janes
Methods For Evaluating Prediction Performance Of Biomarkers And Tests, Margaret Pepe, Holly Janes
UW Biostatistics Working Paper Series
This chapter describes and critiques methods for evaluating the performance of markers to predict risk of a current or future clinical outcome. We consider three criteria that are important for evaluating a risk model: calibration, benefit for decision making and accurate classification. We also describe and discuss a variety of summary measures in common use for quantifying predictive information such as the area under the ROC curve and R-squared. The roles and problems with recently proposed risk reclassification approaches are discussed in detail.
The Impact Of Covariance Misspecification In Multivariate Gaussian Mixtures On Estimation And Inference: An Application To Longitudinal Modeling, Brianna C. Heggeseth, Nicholas P. Jewell
The Impact Of Covariance Misspecification In Multivariate Gaussian Mixtures On Estimation And Inference: An Application To Longitudinal Modeling, Brianna C. Heggeseth, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
Multivariate Gaussian mixtures are a class of models that provide a flexible parametric approach for the representation of heterogeneous multivariate outcomes. When the outcome is a vector of repeated measurements taken on the same subject, there is often inherent dependence between observations. However, a common covariance assumption is conditional independence---that is, given the mixture component label, the outcomes for subjects are independent. In this paper, we study, through asymptotic bias calculations and simulation, the impact of covariance misspecification in multivariate Gaussian mixtures. Although maximum likelihood estimators of regression and mixing probability parameters are not consistent under misspecification, they have little …
Borrowing Information Across Populations In Estimating Positive And Negative Predictive Values, Ying Huang, Youyi Fong, John Wei, Ziding Feng
Borrowing Information Across Populations In Estimating Positive And Negative Predictive Values, Ying Huang, Youyi Fong, John Wei, Ziding Feng
UW Biostatistics Working Paper Series
A marker's capacity to predict risk of a disease depends on disease prevalence in the target population and its classification accuracy, i.e. its ability to discriminate diseased subjects from non-diseased subjects. The latter is often considered an intrinsic property of the marker; it is independent of disease prevalence and hence more likely to be similar across populations than risk prediction measures. In this paper, we are interested in evaluating the population-specific performance of a risk prediction marker in terms of positive predictive value (PPV) and negative predictive value (NPV) at given thresholds, when samples are available from the target population …