Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (102)
- Statistical Methodology (56)
- Statistical Theory (55)
- Medicine and Health Sciences (44)
- Public Health (44)
-
- Statistical Models (31)
- Clinical Epidemiology (23)
- Epidemiology (21)
- Clinical Trials (20)
- Multivariate Analysis (18)
- Survival Analysis (17)
- Genetics and Genomics (16)
- Life Sciences (16)
- Longitudinal Data Analysis and Time Series (16)
- Microarrays (15)
- Categorical Data Analysis (11)
- Bioinformatics (10)
- Computational Biology (10)
- Design of Experiments and Sample Surveys (10)
- Health Services Research (10)
- Genetics (8)
- Vital and Health Statistics (6)
- Applied Mathematics (4)
- Medical Specialties (4)
- Numerical Analysis and Computation (4)
- Disease Modeling (3)
- Diseases (3)
- Probability (2)
- Keyword
-
- Sensitivity (12)
- Classification (11)
- Prediction (9)
- Specificity (9)
- Biomarker (7)
-
- Diagnostic test (6)
- Diagnostic tests (6)
- ROC curve (5)
- Biomarkers (4)
- Clinical trials (4)
- Estimating equations (4)
- Interim analyses (4)
- Longitudinal data (4)
- Sample size (4)
- Self-rated health (4)
- Skewed data (4)
- Verification bias (4)
- Aging (3)
- Air pollution (3)
- Biased sampling (3)
- Bootstrap (3)
- Confidence intervals (3)
- EM algorithm (3)
- Informative follow-up (3)
- Likelihood (3)
- Measurement error (3)
- Microarray (3)
- Missing data (3)
- Operating characteristics (3)
- ROC curves (3)
Articles 151 - 180 of 215
Full-Text Articles in Statistics and Probability
Semiparametric Loglinear Regression For Longitudinal Measurements Subject To Irregular, Biased Follow-Up, Petra Buzkova, Thomas Lumley
Semiparametric Loglinear Regression For Longitudinal Measurements Subject To Irregular, Biased Follow-Up, Petra Buzkova, Thomas Lumley
UW Biostatistics Working Paper Series
We propose a method for analysis of loglinear regression models for longitudinal data that are subject to continuous and irregular follow-up. Frequently, if the follow-up is irregular, the availability of outcome data may be related to the outcome measure or other covariates that are related to the outcome measure. Under such biased sampling designs unadjusted regression analysis yield biased estimates. We examine the marginal association of the covariates X at time t and the logarithm of the mean of response Y at time t. We focus on semiparametric regression with unspecified baseline function of time. To predict the follow-up times …
The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey
The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey
UW Biostatistics Working Paper Series
Significance testing is one of the main objectives of statistics. The Neyman-Pearson lemma provides a simple rule for optimally testing a single hypothesis when the null and alternative distributions are known. This result has played a major role in the development of significance testing strategies that are used in practice. Most of the work extending single testing strategies to multiple tests has focused on formulating and estimating new types of significance measures, such as the false discovery rate. These methods tend to be based on p-values that are calculated from each test individually, ignoring information from the other tests. As …
The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek
The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek
UW Biostatistics Working Paper Series
As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the Optimal Discovery Procedure (ODP), which has recently been introduced and theoretically shown …
Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang
Linear Regression Of Censored Length-Biased Lifetimes, Ying Qing Chen, Yan Wang
UW Biostatistics Working Paper Series
Length-biased lifetimes may be collected in observational studies or sample surveys due to biased sampling scheme. In this article, we use a linear regression model, namely, the accelerated failure time model, for the population lifetime distributions in regression analysis of the length-biased lifetimes. It is discovered that the associated regression parameters are invariant under the length-biased sampling scheme. According to this discovery, we propose the quasi partial score estimating equations to estimate the population regression parameters. The proposed methodologies are evaluated and demonstrated by simulation studies and an application to actual data set.
On Additive Regression Of Expectancy, Ying Qing Chen
On Additive Regression Of Expectancy, Ying Qing Chen
UW Biostatistics Working Paper Series
Regression models have been important tools to study the association between outcome variables and their covariates. The traditional linear regression models usually specify such an association by the expectations of the outcome variables as function of the covariates and some parameters. In reality, however, interests often focus on their expectancies characterized by the conditional means. In this article, a new class of additive regression models is proposed to model the expectancies. The model parameters carry practical implication, which may allow the models to be useful in applications such as treatment assessment, resource planning or short-term forecasting. Moreover, the new model …
An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley
An Empirical Process Limit Theorem For Sparsely Correlated Data, Thomas Lumley
UW Biostatistics Working Paper Series
We consider data that are dependent, but where most small sets of observations are independent. By extending Bernstein's inequality we prove a strong law of law numbers and an empirical process central limit theorem under bracketing entropy conditions.
A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe
A Linear Regression Framework For Receiver Operating Characteristic(Roc) Curve Analysis, Zheng Zhang, Margaret S. Pepe
UW Biostatistics Working Paper Series
In the field of medical diagnostic testing, the receiver operating characteristics(ROC) curve has long been used as a standard statistical tool to assess the accuracy of tests that yield continuous results. Although previous research in this area focused mostly on estimating the ROC curve, recently it has been recognized that the accuracy of a given test may fluctuate depending on certain factors, which motivates modelling covariate effects on the ROC curve. Comparing the corresponding ROC curves between two or more tests is a special case of covariate effect modelling. In this manuscript, we introduce a linear regression framework to model …
Attributable Risk Function In The Proportional Hazards Model, Ying Qing Chen, Chengcheng Hu, Yan Wang
Attributable Risk Function In The Proportional Hazards Model, Ying Qing Chen, Chengcheng Hu, Yan Wang
UW Biostatistics Working Paper Series
As an epidemiological parameter, the population attributable fraction is an important measure to quantify the public health attributable risk of an exposure to morbidity and mortality. In this article, we extend this parameter to the attributable fraction function in survival analysis of time-to-event outcomes, and further establish its estimation and inference procedures based on the widely used proportional hazards models. Numerical examples and simulations studies are presented to validate and demonstrate the proposed methods.
Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou
Multiple Imputation For Correcting Verification Bias, Ofer Harel, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
In the case in which all subjects are screened using a common test, and only a subset of these subjects are tested using a golden standard test, it is well documented that there is a risk for bias, called verification bias. When the test has only two levels (e.g. positive and negative) and we are trying to estimate the sensitivity and specificity of the test, one is actually constructing a confidence interval for a binomial proportion. Since it is well documented that this estimation is not trivial even with complete data, we adopt Multiple imputation (MI) framework for verification bias …
A Comparison Of Parametric And Coarsened Bayesian Interval Estimation In The Presence Of A Known Mean-Variance Relationship, Kent Koprowicz, Scott S. Emerson, Peter Hoff
A Comparison Of Parametric And Coarsened Bayesian Interval Estimation In The Presence Of A Known Mean-Variance Relationship, Kent Koprowicz, Scott S. Emerson, Peter Hoff
UW Biostatistics Working Paper Series
While the use of Bayesian methods of analysis have become increasingly common, classical frequentist hypothesis testing still holds sway in medical research - especially clinical trials. One major difference between a standard frequentist approach and the most common Bayesian approaches is that even when a frequentist hypothesis test is derived from parametric models, the interpretation and operating characteristics of the test may be considered in a distribution-free manner. Bayesian inference, on the other hand, is often conducted in a parametric setting where the interpretation of the results is dependent on the parametric model. Here we consider a Bayesian counterpart to …
Application Of The Time-Dependent Roc Curves For Prognostic Accuracy With Multiple Biomarkers, Yingye Zheng, Tianxi Cai, Ziding Feng
Application Of The Time-Dependent Roc Curves For Prognostic Accuracy With Multiple Biomarkers, Yingye Zheng, Tianxi Cai, Ziding Feng
UW Biostatistics Working Paper Series
The rapid advancement in molecule technology has lead to the discovery of many markers that have potential applications in disease diagnosis and prognosis. In a prospective cohort study, information on a panel of biomarkers as well as the disease status for a patient are routinely collected over time. Such information is useful to predict patients' prognosis and select patients for targeted therapy. In this paper, we develop procedures for constructing a composite test with optimal discrimination power when there are multiple markers available to assist in prediction and characterize the accuracy of the resulting test by extending the time-dependent receiver …
New Confidence Intervals For The Difference Between Two Sensitivities At A Fixed Level Of Specificity, Gengsheng Qin, Yu-Sheng Hsu, Xiao-Hua Zhou
New Confidence Intervals For The Difference Between Two Sensitivities At A Fixed Level Of Specificity, Gengsheng Qin, Yu-Sheng Hsu, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
For two continuous-scale diagnostic tests, it is of interest to compare their sensitivities at a predetermined level of specificity. In this paper we propose three new intervals for the difference between two sensitivities at a fixed level of specificity. These intervals are easy to compute. We also conduct simulation studies to compare the relative performance of the new intervals with the existing normal approximation based interval proposed by Wieand et al (1989). Our simulation results show that the newly proposed intervals perform better than the existing normal approximation based interval in terms of coverage accuracy and interval length.
Frequentist Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
Frequentist Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
UW Biostatistics Working Paper Series
Group sequential stopping rules are often used as guidelines in the monitoring of clinical trials in order to address the ethical and efficiency issues inherent in human testing of a new treatment or preventive agent for disease. Such stopping rules have been proposed based on a variety of different criteria, both scientific (e.g., estimates of treatment effect) and statistical (e.g., frequentist type I error, Bayesian posterior probabilities, stochastic curtailment). It is easily shown, however, that a stopping rule based on one of those criteria induces a stopping rule on all other criteria. Thus the basis used to initially define a …
Bayesian Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
Bayesian Evaluation Of Group Sequential Clinical Trial Designs, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
UW Biostatistics Working Paper Series
Clincal trial designs often incorporate a sequential stopping rule to serve as a guide in the early termination of a study. When choosing a particular stopping rule, it is most common to examine frequentist operating characteristics such as type I error, statistical power, and precision of confi- dence intervals (Emerson, et al. [1]). Increasingly, however, clinical trials are designed and analyzed in the Bayesian paradigm. In this paper we describe how the Bayesian operating characteristics of a particular stopping rule might be evaluated and communicated to the scientific community. In particular, we consider a choice of probability models and a …
On The Use Of Stochastic Curtailment In Group Sequential Clinical Trials, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
On The Use Of Stochastic Curtailment In Group Sequential Clinical Trials, Scott S. Emerson, John M. Kittelson, Daniel L. Gillen
UW Biostatistics Working Paper Series
Many different criteria have been proposed for the selection of a stopping rule for group sequen- tial trials. These include both scientific (e.g., estimates of treatment effect) and statistical (e.g., frequentist type I error, Bayesian posterior probabilities, stochastic curtailment) measures of the evidence for or against beneficial treatment effects. Because a stopping rule based on one of those criteria induces a stopping rule on all other criteria, the utility of any particular scale relates to the ease with which it allows a clinical trialist to search for sequential sampling plans having de- sirable operating characteristics. In this paper we examine …
The Clustering Of Regression Models Method With Applications In Gene Expression Data, Li-Xuan Qin, Steven G. Self
The Clustering Of Regression Models Method With Applications In Gene Expression Data, Li-Xuan Qin, Steven G. Self
UW Biostatistics Working Paper Series
Identification of differentially expressed genes and clustering of genes are two important and complementary objectives addressed with gene expression data. For the differential expression question, many "per-gene" analytic methods have been proposed. These methods can generally be characterized as using a regression function to independently model the observations for each gene; various adjustments for multiplicity are then used to interpret the statistical significance of these per-gene regression models over the collection of genes analyzed. Motivated by this common structure of per-gene models, we propose a new model-based clustering method -- the clustering of regression models method, which groups genes that …
Insights Into Latent Class Analysis, Margaret S. Pepe, Holly Janes
Insights Into Latent Class Analysis, Margaret S. Pepe, Holly Janes
UW Biostatistics Working Paper Series
Latent class analysis is a popular statistical technique for estimating disease prevalence and test sensitivity and specificity. It is used when a gold standard assessment of disease is not available but results of multiple imperfect tests are. We derive analytic expressions for the parameter estimates in terms of the raw data, under the conditional independence assumption. These expressions indicate explicitly how observed two- and three-way associations between test results are used to infer disease prevalence and test operating characteristics. Although reasonable if the conditional independence model holds, the estimators have no basis when it fails. We therefore caution against using …
Standardizing Markers To Evaluate And Compare Their Performances, Margaret S. Pepe, Gary M. Longton
Standardizing Markers To Evaluate And Compare Their Performances, Margaret S. Pepe, Gary M. Longton
UW Biostatistics Working Paper Series
Introduction: Markers that purport to distinguish subjects with a condition from those without a condition must be evaluated rigorously for their classification accuracy. A single approach to statistically evaluating and comparing markers is not yet established.
Methods: We suggest a standardization that uses the marker distribution in unaffected subjects as a reference. For an affected subject with marker value Y, the standardized placement value is the proportion of unaffected subjects with marker values that exceed Y.
Results: We apply the standardization to two illustrative datasets. In patients with pancreatic cancer placement values calculated for the CA 19-9 marker are smaller …
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang, Gary M. Longton
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang, Gary M. Longton
UW Biostatistics Working Paper Series
No single biomarker for cancer is considered adequately sensitive and specific for cancer screening. It is expected that the results of multiple markers will need to be combined in order to yield adequately accurate classification. Typically the objective function that is optimized for combining markers is the likelihood function. In this paper we consider an alternative objective function -- the area under the empirical receiver operating characteristic curve (AUC). We note that it yields consistent estimates of parameters in a generalized linear model for the risk score but does not require specifying the link function. Like logistic regression it yields …
Referent Selection Strategies In Case-Crossover Analyses Of Air Pollution Exposure Data: Implications For Bias, Holly Janes, Lianne Sheppard, Thomas Lumley
Referent Selection Strategies In Case-Crossover Analyses Of Air Pollution Exposure Data: Implications For Bias, Holly Janes, Lianne Sheppard, Thomas Lumley
UW Biostatistics Working Paper Series
The case-crossover design has been widely used to study the association between short term air pollution exposure and the risk of an acute adverse health event. The design uses cases only, and, for each individual, compares exposure just prior to the event with exposure at other control, or “referent” times. By making within-subject comparisons, time invariant confounders are controlled by design. Even more important in the air pollution setting is that, by matching referents to the index time, time varying confounders can also be controlled by design. Yet, the referent selection strategy is important for reasons other than control of …
Semi-Parametric Single-Index Two-Part Regression Models, Xiao-Hua Zhou, Hua Liang
Semi-Parametric Single-Index Two-Part Regression Models, Xiao-Hua Zhou, Hua Liang
UW Biostatistics Working Paper Series
In this paper, we proposed a semi-parametric single-index two-part regression model to weaken assumptions in parametric regression methods that were frequently used in the analysis of skewed data with additional zero values. The estimation procedure for the parameters of interest in the model was easily implemented. The proposed estimators were shown to be consistent and asymptotically normal. Through a simulation study, we showed that the proposed estimators have reasonable finite-sample performance. We illustrated the application of the proposed method in one real study on the analysis of health care costs.
Estimating The Retransformed Mean In A Heteroscedastic Two-Part Model, Alan H. Welsh, Xiao-Hua Zhou
Estimating The Retransformed Mean In A Heteroscedastic Two-Part Model, Alan H. Welsh, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
Two distribution free estimators are proposed to estimate the mean of a dependent variable after fitting a semiparametric two-part heteroscedastic regression model to a transformation of the dependent variable. We show that the proposed estimators are consistent and have asymptotic normal distributions. We also compare their finite-sample performance in a simulation study. Finally, we illustrate the proposed methods in a real-world example of predicting in-patient health care costs.
A Marginal Model Approach For Analysis Of Multi-Reader Multi-Test Receiver Operating Characteristic (Roc) Data, Xiao Song, Xiao-Hua Zhou
A Marginal Model Approach For Analysis Of Multi-Reader Multi-Test Receiver Operating Characteristic (Roc) Data, Xiao Song, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
The receiver operating characteristic (ROC) curve is a popular tool to characterize the capabilities of diagnostic tests with continuous or ordinal responses. One common design for assessing the accuracy of diagnostic tests is to have each patient examined by multiple readers with multiple tests; this design is most commonly used in a radiology setting, where the results of diagnostic tests depend on a radiologist's subjective interpretation. The most widely used approach for analyzing data from such a study is the Dorfman-Berbaum-Metz (DBM) method (Dorfman, Berbaum and Metz, 1992) which utilizes a standard analysis of variance (ANOVA) model for the jackknife …
Nonparametric Confidence Intervals For The One- And Two-Sample Problems, Xiao-Hua Zhou, Phillip Dinh
Nonparametric Confidence Intervals For The One- And Two-Sample Problems, Xiao-Hua Zhou, Phillip Dinh
UW Biostatistics Working Paper Series
Confidence intervals for the mean of one sample and the difference in means of two independent samples based on the ordinary-t statistic suffer deficiencies when samples come from skewed distributions. In this article, we evaluate several existing techniques and propose new methods to improve coverage accuracy. The methods examined include the ordinary-t, the bootstrap-t, the biased-corrected acceleration (BCa) bootstrap, and three new intervals based on transformation of the t-statistic. Our study shows that our new transformation intervals and the bootstrap-t intervals give best coverage accuracy for a variety of skewed distributions; and that our new transformation intervals have shorter interval …
Significance Analysis Of Time Course Microarray Experiments, John D. Storey, Wenzhong Xiao, Jeffrey T. Leek, Ronald G. Tompkins, Ron W. Davis
Significance Analysis Of Time Course Microarray Experiments, John D. Storey, Wenzhong Xiao, Jeffrey T. Leek, Ronald G. Tompkins, Ron W. Davis
UW Biostatistics Working Paper Series
Characterizing the genome-wide dynamic regulation of gene expression is important and will be of much interest in the future. However, there is currently no established method for identifying differentially expressed genes in a time course study. Here we propose a significance method for analyzing time course microarray studies that can be applied to the typical types of comparisons and sampling schemes. This method is applied to two studies on humans. In one study, genes are identified that show differential expression over time in response to in vivo endotoxin administration. Using our method 7409 genes are called significant at a 1% …
Non-Parametric Estimation Of Roc Curves In The Absence Of A Gold Standard, Xiao-Hua Zhou, Pete Castelluccio, Chuan Zhou
Non-Parametric Estimation Of Roc Curves In The Absence Of A Gold Standard, Xiao-Hua Zhou, Pete Castelluccio, Chuan Zhou
UW Biostatistics Working Paper Series
In evaluation of diagnostic accuracy of tests, a gold standard on the disease status is required. However, in many complex diseases, it is impossible or unethical to obtain such the gold standard. If an imperfect standard is used as if it were a gold standard, the estimated accuracy of the tests would be biased. This type of bias is called imperfect gold standard bias. In this paper we develop a maximum likelihood (ML) method for estimating ROC curves and their areas of ordinal-scale tests in the absence of a gold standard. Our simulation study shows the proposed estimates for the …
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang
Combining Predictors For Classification Using The Area Under The Roc Curve, Margaret S. Pepe, Tianxi Cai, Zheng Zhang
UW Biostatistics Working Paper Series
We compare simple logistic regression with an alternative robust procedure for constructing linear predictors to be used for the two state classification task. Theoritical advantages of the robust procedure over logistic regression are: (i) although it assumes a generalized linear model for the dichotomous outcome variable, it does not require specification of the link function; (ii) it accommodates case-control designs even when the model is not logistic; and (iii) it yields sensible results even when the generalized linear model assumption fails to hold. Surprisingly, we find that the linear predictor derived from the logistic regression likelihood is very robust in …
On Corrected Score Approach For Proportional Hazards Model With Covariate Measurement Error, Xiao Song, Yijian Huang
On Corrected Score Approach For Proportional Hazards Model With Covariate Measurement Error, Xiao Song, Yijian Huang
UW Biostatistics Working Paper Series
In the presence of covariate measurement error with the proportional hazards model, several functional modeling methods have been proposed. These include the conditional score estimator (Tsiatis and Davidian, 2001), the parametric correction estimator (Nakamura, 1992) and the nonparametric correction estimator (Huang and Wang, 2000, 2003) in the order of weaker assumptions on the error. Although they are all consistent, each suffers from potential difficulties with small samples and substantial measurement error. In this article, upon noting that the conditional score and parametric correction estimators are asymptotically equivalent in the case of normal error, we investigate their relative finite sample performance …
Evaluating Markers For Selecting A Patient's Treatment, Xiao Song, Margaret S. Pepe
Evaluating Markers For Selecting A Patient's Treatment, Xiao Song, Margaret S. Pepe
UW Biostatistics Working Paper Series
Selecting the best treatment for a patient's disease may be facilitated by evaluating clinical characteristics or biomarker measurements at diagnosis. We consider how to evaluate the potential of such measurements to impact on treatment selection algorithms. For example, magnetic resonance neurographic imaging is potentially useful for deciding whether a patient should be treated surgically for carpal tunnel syndrome or if he/she should receive less invasive conservative therapy. We propose a graphical display, the selection impact (SI) curve, that shows the population response rate as a function of treatment selection criteria based on the marker. The curve can be useful for …
Calibrating Observed Differential Gene Expression For The Multiplicity Of Genes On The Array, Yingye Zheng, Margaret S. Pepe
Calibrating Observed Differential Gene Expression For The Multiplicity Of Genes On The Array, Yingye Zheng, Margaret S. Pepe
UW Biostatistics Working Paper Series
In a gene expression array study, the expression levels of thousands of genes are monitored simultaneously across various biological conditions on a small set of subjects. One goal of such studies is to explore a large pool of genes in order to select a subset of genes that appear to be differently expressed for further investigation. Of particular interest here is how to select the top k genes once genes are ranked based on their evidence for differential expression in two tissue types. We consider statistical methods that provide a more rigorous and intuitively appealing selection process for k. We …