Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

COBRA

Discipline
Keyword
Publication Year
Publication

Articles 781 - 810 of 1108

Full-Text Articles in Statistics and Probability

The Two-Sample Problem For Failure Rates Depending On A Continuous Mark: An Application To Vaccine Efficacy, Peter B. Gilbert, Ian W. Mckeague, Yanqing Sun Mar 2006

The Two-Sample Problem For Failure Rates Depending On A Continuous Mark: An Application To Vaccine Efficacy, Peter B. Gilbert, Ian W. Mckeague, Yanqing Sun

UW Biostatistics Working Paper Series

The efficacy of an HIV vaccine to prevent infection is likely to depend on the genetic variation of the exposing virus. This paper addresses the problem of using data on the HIV sequences that infect vaccine efficacy trial participants to 1) test for vaccine efficacy more powerfully than procedures that ignore the sequence data; and 2) evaluate the dependence of vaccine efficacy on the divergence of infecting HIV strains from the HIV strain that is contained in the vaccine. Because hundreds of amino acid sites in each HIV genome are sequenced, it is natural to treat the divergence (defined in …


Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes Mar 2006

Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes

UW Biostatistics Working Paper Series

Consider a placebo-controlled preventive HIV vaccine efficacy trial. An HIV amino acid sequence is measured from each volunteer who acquires HIV, and these sequences are aligned together with the reference HIV sequence represented in the vaccine. We develop genome scanning methods to identify HIV positions at which the amino acids in sequences from infected vaccine recipients tend to be more divergent from the corresponding reference amino acid than the amino acids in sequences from infected placebo recipients. We consider five two-sample test statistics, based on Euclidean, Mahalanobis, and Kullback-Leibler divergence measures. Weights are incorporated to reflect biological information contained in …


On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger Mar 2006

On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger

Johns Hopkins University, Dept. of Biostatistics Working Papers

Time series and case-crossover methods are often viewed as competing alternatives in environmental epidemiologic studies. Several recent studies have compared the time series and case-crossover methods. In this paper, we show that case-crossover using conditional logistic regression is a special case of time series analysis when there is a common exposure such as in air pollution studies. This equivalence provides computational convenience for case-crossover analyses and a better understanding of time series models. Time series log-linear regression accounts for over-dispersion of the Poisson variance, while case-crossover analyses typically do not. This equivalence also permits model checking for case-crossover data using …


Evaluating Prediction Rules For T-Year Survivors With Censored Regression Models, Hajime Uno, Tianxi Cai, Lu Tian, L.J. Wei Mar 2006

Evaluating Prediction Rules For T-Year Survivors With Censored Regression Models, Hajime Uno, Tianxi Cai, Lu Tian, L.J. Wei

Harvard University Biostatistics Working Paper Series

Suppose that we are interested in establishing simple, but reliable rules for predicting future t-year survivors via censored regression models. In this article, we present inference procedures for evaluating such binary classification rules based on various prediction precision measures quantified by the overall misclassification rate, sensitivity and specificity, and positive and negative predictive values. Specifically, under various working models we derive consistent estimators for the above measures via substitution and cross validation estimation procedures. Furthermore, we provide large sample approximations to the distributions of these nonsmooth estimators without assuming that the working model is correctly specified. Confidence intervals, for example, …


2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr Mar 2006

2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr

UW Biostatistics Working Paper Series

When a two-level design must be run in blocks of size two, there is a unique blocking scheme that enables estimation of all the main effects. Unfortunately this design does not enable estimation of any two-factor interactions. When the experimental goal is to estimate all main effects and two-factor interactions, it is necessary to combine replicates of the experiment that use different blocking schemes. In this paper we identify such designs for up to eight factors that enable estimation of all main effects and two-factor interactions with the fewest number of replications. In addition, we give a construction for general …


Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan Mar 2006

Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

We propose a general and formal statistical framework for the multiple tests of associations between known fixed features of a genome and unknown parameters of the distribution of variable features of this genome in a population of interest. The known fixed gene-annotation profiles, corresponding to the fixed features of the genome, may concern Gene Ontology (GO) annotation, pathway membership, regulation by particular transcription factors, nucleotide sequences, or protein sequences. The unknown gene-parameter profiles, corresponding to the variable features of the genome, may be, for example, regression coefficients relating genome-wide transcript levels or DNA copy numbers to possibly censored biological and …


A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull Mar 2006

A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull

Harvard University Biostatistics Working Paper Series

We introduce a diagnostic test for the mixing distribution in a generalised linear mixed model. The test is based on the difference between the marginal maximum likelihood and conditional maximum likelihood estimates of a subset of the fixed effects in the model. We derive the asymptotic variance of this difference, and propose a test statistic that has a limiting chi-square distribution under the null hypothesis that the mixing distribution is correctly specified. For the important special case of the logistic regression model with random intercepts, we evaluate via simulation the power of the test in finite samples under several alternative …


Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng Mar 2006

Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng

UW Biostatistics Working Paper Series

Consider a continuous marker for predicting a binary outcome. For example, serum concentration of prostate specific antigen (PSA) may be used to calculate the risk of finding prostate cancer in a biopsy. In this paper we argue that the predictive capacity of a marker has to do with the population distribution of risk given the marker and suggest a graphical tool, the predictiveness curve, that displays this distribution. The display provides a common meaningful scale for comparing markers that may not be comparable on their original scales. Some existing measures of predictiveness are shown to be summary indices derived from …


Different Public Health Interventions Have Varying Effects, Paula Diehr, Anne B. Newman, Liming Cai, Ann Derleth Feb 2006

Different Public Health Interventions Have Varying Effects, Paula Diehr, Anne B. Newman, Liming Cai, Ann Derleth

UW Biostatistics Working Paper Series

Objective: To compare performance of one-time health interventions to those that change the probability of transitioning from one health state to another. Study Design and Setting: We used multi-state life table methods to estimate the impact of eight types of interventions on several outcomes. Results: In a cohort beginning at age 65, curing all the sick persons at baseline would increase life expectancy by 0.23 years and increase years of healthy life by .54 years. An equal amount of improvement could be obtained with a 12% decrease in the probability of getting sick, a 16% increase in the probability of …


Survival Analysis Methods In Genetic Epidemiology, Hongzhe Li Feb 2006

Survival Analysis Methods In Genetic Epidemiology, Hongzhe Li

UPenn Biostatistics Working Papers

Mapping genes for complex human diseases is a challenging problem due to the fact that many such diseases are due to both genetic and enviromental risk factors and many also exhibit phenotypic heterogeneity, such as variable age of onset. Information on variable age of disease onset is often a good indicator for disease heterogeneity and incorporation of such information together with enviromental risk factors into genetic analysis should lead to more powerful tests for genetic analysis. Due to the problem of censoring, survival analysis methods have proved to be very useful for genetic analysis. In this paper, I review some …


On The Violation Of Bounds For The Correlation In Generalized Estimating Equation Analyses Of Binary Data From Longitudinal Trials, Justine Shults, Wenguang Sun, Xin Tu, Jay Amsterdam Feb 2006

On The Violation Of Bounds For The Correlation In Generalized Estimating Equation Analyses Of Binary Data From Longitudinal Trials, Justine Shults, Wenguang Sun, Xin Tu, Jay Amsterdam

UPenn Biostatistics Working Papers

It is well-known that the correlation among binary outcomes is constrained by the marginal means, yet approaches such as generalized estimating equations (GEE) do not check that the constraints for the correlations are satisfied. We explore this issue for Markovian dependence in the context of a GEE analysis of a clinical trial that compares Venlafaxine with Lithium in the prevention of major depressive episode. We obtain simplified expressions for the constraints for the logistic model and the equicorrelated and first-order autoregressive correlation structures. We then obtain the limiting values of the GEE and quasi-least squares (QLS) estimates of the correlation …


Comparing The Predictive Values Of Diagnostic Tests: Sample Size And Analysis For Paired Study Designs, Chaya S. Moskowitz, Margaret S. Pepe Feb 2006

Comparing The Predictive Values Of Diagnostic Tests: Sample Size And Analysis For Paired Study Designs, Chaya S. Moskowitz, Margaret S. Pepe

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

In this paper we consider the design and analysis of studies comparing the positive and negative predictive values of two diagnostic tests that are measured on all subjects. Although statistical methodology is well developed for comparing diagnostic tests in terms of their sensitivities and specificities, comparative inference about predictive values is not. We derive analytic variance expressions for the relative predictive values. Sample size formulas for study design ensue. In addition, two new methods for analyzing the resulting data are presented and compared with an existing marginal regression methodology.


Regression Analysis For The Partial Area Under The Roc Curve, Tianxi Cai, Lori E. Dodd Feb 2006

Regression Analysis For The Partial Area Under The Roc Curve, Tianxi Cai, Lori E. Dodd

Harvard University Biostatistics Working Paper Series

No abstract provided.


Use Of Unbiased Estimating Equations To Estimate Correlation In Generalized Estimating Equation Analysis Of Longitudinal Trials, Wenguang Sun, Justine Shults, Mary Leonard Jan 2006

Use Of Unbiased Estimating Equations To Estimate Correlation In Generalized Estimating Equation Analysis Of Longitudinal Trials, Wenguang Sun, Justine Shults, Mary Leonard

UPenn Biostatistics Working Papers

In a recent publication, Wang and Carey (Journal of the American Statistical Association, 99, pp. 845-853, 2004) presented a new approach for estimation of the correlation parameters in the framework of generalized estimating equations (GEE). They considered correlated continuous, binary and count data with a generalized Markov correlation structure that includes the first-order autoregressive AR(1) and Markov structures as special cases. They made detailed comparisons with pseudo-likelihood (PL) and the first stage of quasi-least squares (QLS), a two-stage approach in the framework of generalized estimating equations (GEE). In this note we extend their comparisons for the second (bias corrected) stage …


Case-Cohort Methods For Survival Data On Families From Routine Registers, Tron Anders Moger, Yudi Pawitan, Ørnulf Borgan Jan 2006

Case-Cohort Methods For Survival Data On Families From Routine Registers, Tron Anders Moger, Yudi Pawitan, Ørnulf Borgan

UW Biostatistics Working Paper Series

In the Nordic countries, there exist several registers containing information on diseases and risk factors for millions of individuals. This information can be linked into families by use of personal identification numbers, and represent a great opportunity for studying diseases that show familial aggregation. Due to the size of the registers, it is difficult to analyze the data by using traditional methods for multivariate survival analysis, such as frailty or copula models. Since the size of the cohort is known, case-cohort methods based on pseudo-likelihoods are suitable for analyzing the data. We present methods for sampling control families both with …


Comparison Of Haplotype-Based And Tree-Based Snp Imputation In Association Studies, James Y. Dai, Ingo Ruczinski, Michael Leblanc, Charles Kooperberg Jan 2006

Comparison Of Haplotype-Based And Tree-Based Snp Imputation In Association Studies, James Y. Dai, Ingo Ruczinski, Michael Leblanc, Charles Kooperberg

UW Biostatistics Working Paper Series

Missing single nucleotide polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we compare two haplotype-based imputation approaches and one regression tree-based imputation approach for association studies. The goal is to assess the imputation accuracy, and to evaluate the impact of imputation on parameter estimation. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to …


Semiparametric Approaches For Joint Modeling Of Longitudinal And Survival Data With Time Varying Coefficients, Xiao Song, C.Y. Wang Dec 2005

Semiparametric Approaches For Joint Modeling Of Longitudinal And Survival Data With Time Varying Coefficients, Xiao Song, C.Y. Wang

UW Biostatistics Working Paper Series

We study joint modeling of survival and longitudinal data. There are two regression models of interest. The primary model is for survival outcomes, which are assumed to follow a time varying coefficient proportional hazards model. The second model is for longitudinal data, which are assumed to follow a random effects model. Based on the trajectory of a subject's longitudinal data, some covariates in the survival model are functions of the unobserved random effects. Estimated random effects are generally different from the unobserved random effects and hence this leads to covariate measurement error. To deal with covariate measurement error, we propose …


Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson Dec 2005

Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson

UW Biostatistics Working Paper Series

In this paper, we illustrate that combining ecological data with subsample data in situations in which a linear model is appropriate provides three main benefits. First, by including the individual level subsample data, the biases associated with linear ecological inference can be eliminated. Second, by supplementing the subsample data with ecological data, the information about parameters will be increased. Third, we can use readily available ecological data to design optimal subsampling schemes, so as to further increase the information about parameters. We present an application of this methodology to the classic problem of estimating the effect of a college degree …


Bayesian Analysis Of Cell-Cycle Gene Expression Data, Chuan Zhou, Jon Wakefield, Linda Breeden Dec 2005

Bayesian Analysis Of Cell-Cycle Gene Expression Data, Chuan Zhou, Jon Wakefield, Linda Breeden

UW Biostatistics Working Paper Series

The study of the cell-cycle is important in order to aid in our understanding of the basic mechanisms of life, yet progress has been slow due to the complexity of the process and our lack of ability to study it at high resolution. Recent advances in microarray technology have enabled scientists to study the gene expression at the genome-scale with a manageable cost, and there has been an increasing effort to identify cell-cycle regulated genes. In this chapter, we discuss the analysis of cell-cycle gene expression data, focusing on a model-based Bayesian approaches. The majority of the models we describe …


Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou Dec 2005

Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

For a continuous-scale diagnostic test, the most commonly used summary index of the receiver operating characteristic (ROC) curve is the area under the curve (AUC) that measures the accuracy of the diagnostic test. In this paper we propose an empirical likelihood approach for the inference of AUC. We first define an empirical likelihood ratio for AUC and show that its limiting distribution is a scaled chi-square distribution. We then obtain an empirical likelihood based confidence interval for AUC using the scaled chi-square distribution. This empirical likelihood inference for AUC can be extended to stratified samples and the resulting limiting distribution …


Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou Dec 2005

Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou

UW Biostatistics Working Paper Series

Health research often gives rise to data that follow lognormal distributions. In two sample situations, researchers are likely to be interested in estimating the difference or ratio of the population means. Several methods have been proposed for providing confidence intervals for these parameters. However, it is not clear which techniques are most appropriate, or how their performance might vary. Additionally, methods for the difference of means have not been adequately explored. We discuss in the present article five methods of analysis. These include two methods based on the log-likelihood ratio statistic and a generalized pivotal approach. Additionally, we provide and …


Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li Dec 2005

Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li

UW Biostatistics Working Paper Series

In many studies of health economics, we are interested in the expected total cost over a certain period for a patient with given characteristics. Problems can arise if cost estimation models do not account for distributional aspects of costs. Two such problems are 1) the skewed nature of the data and 2) censored observations. In this paper we propose an empirical likelihood (EL) method for constructing a confidence region for the vector of regression parameters and a confidence interval for the expected total cost of a patient with the given covariates. We show that this new method has good theoretical …


Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau Dec 2005

Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau

UW Biostatistics Working Paper Series

The accuracy of a binary-scale diagnostic test can be represented by sensitivity (Se), specificity (Sp) and positive and negative predictive values (PPV and NPV). Although Se and Sp measure the intrinsic accuracy of a diagnostic test that does not depend on the prevalence rate, they do not provide information on the diagnostic accuracy of a particular patient. To obtain this information we need to use PPV and NPV. Since PPV and NPV are functions of both the intrinsic accuracy and the prevalence of the disease, constructing confidence intervals for PPV and NPV for a particular patient in a population with …


Model Checking For Roc Regression Analysis, Tianxi Cai, Yingye Zheng Dec 2005

Model Checking For Roc Regression Analysis, Tianxi Cai, Yingye Zheng

Harvard University Biostatistics Working Paper Series

The Receiver Operating Characteristic (ROC) curve is a prominent tool for characterizing the accuracy of continuous diagnostic test. To account for factors that might invluence the test accuracy, various ROC regression methods have been proposed. However, as in any regression analysis, when the assumed models do not fit the data well, these methods may render invalid and misleading results. To date practical model checking techniques suitable for validating existing ROC regression models are not yet available. In this paper, we develop cumulative residual based procedures to graphically and numerically assess the goodness-of-fit for some commonly used ROC regression models, and …


On The Use Of Non-Euclidean Isotropy In Geostatistics, Frank C. Curriero Dec 2005

On The Use Of Non-Euclidean Isotropy In Geostatistics, Frank C. Curriero

Johns Hopkins University, Dept. of Biostatistics Working Papers

This paper investigates the use of non-Euclidean distances to characterize isotropic spatial dependence for geostatistical related applications. A simple example is provided to demonstrate there are no guarantees that existing covariogram and variogram functions remain valid (i.e.\ positive definite or conditionally negative definite) when used with a non-Euclidean distance measure. Furthermore, satisfying the conditions of a metric is not sufficient to ensure the distance measure can be used with existing functions. Current literature is not clear on these topics. There are certain distance measures that when used with existing covariogram and variogram functions remain valid, an issue that is explored. …


Gradient Directed Regularization For Sparse Gaussian Concentration Graphs, With Applications To Inference Of Genetic Networks, Hongzhe Li, Jiang Gui Dec 2005

Gradient Directed Regularization For Sparse Gaussian Concentration Graphs, With Applications To Inference Of Genetic Networks, Hongzhe Li, Jiang Gui

UPenn Biostatistics Working Papers

Large-scale microarray gene expression data provide the possibility of constructing genetic networks or biological pathways. Gaussian graphical models have been suggested to provide an effective method for constructing such genetic networks. However, most of the available methods for constructing Gaussian graphs do not account for the sparsity of the networks and are computationally more demanding or infeasible, especially in the settings of high-dimension and low sample size. We introduce a threshold gradient descent regularization procedure for estimating the sparse precision matrix in the setting of Gaussian graphical models and demonstrate its application to identifying genetic networks. Such a procedure is …


Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith Dec 2005

Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith

U.C. Berkeley Division of Biostatistics Working Paper Series

A new data filtering method for SELDI-TOF MS proteomic spectra data is described. We examined technical repeats (2 per subject) of intensity versus m/z (mass/charge) of bone marrow cell lysate for two groups of childhood leukemia patients: acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). As others have noted, the type of data processing as well as experimental variability can have a disproportionate impact on the list of "interesting" proteins (see Baggerly et al. (2004)). We propose a list of processing and multiple testing techniques to correct for 1) background drift; 2) filtering using smooth regression and cross-validated bandwidth …


Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard Nov 2005

Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard

U.C. Berkeley Division of Biostatistics Working Paper Series

Simultaneously testing a collection of null hypotheses about a data generating distribution based on a sample of independent and identically distributed observations is a fundamental and important statistical problem involving many applications. Methods based on marginal null distributions (i.e., marginal p-values) are attractive since the marginal p-values can be based on a user supplied choice of marginal null distributions and they are computationally trivial, but they, by necessity, are known to either be conservative or to rely on assumptions about the dependence structure between the test-statistics. Resampling based multiple testing (Westfall and Young, 1993) involves sampling from a joint null …


Data Adaptive Pathway Testing, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan Nov 2005

Data Adaptive Pathway Testing, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

A majority of diseases are caused by a combination of factors, for example, composite genetic mutation profiles have been found in many cases to predict a deleterious outcome. There are several statistical techniques that have been used to analyze these types of biological data. This article implements a general strategy which uses data adaptive regression methods to build a specific pathway model, thus predicting a disease outcome by a combination of biological factors and assesses the significance of this model, or pathway, by using a permutation based null distribution. We also provide several simulation comparisons with other techniques. In addition, …


Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey Nov 2005

Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey

UW Biostatistics Working Paper Series

Nearest centroid classifiers have recently been successfully employed in high-dimensional applications. A necessary step when building a classifier for high-dimensional data is feature selection. Feature selection is typically carried out by computing univariate statistics for each feature individually, without consideration for how a subset of features performs as a whole. For subsets of a given size, we characterize the optimal choice of features, corresponding to those yielding the smallest misclassification rate. Furthermore, we propose an algorithm for estimating this optimal subset in practice. Finally, we investigate the applicability of shrinkage ideas to nearest centroid classifiers. We use gene-expression microarrays for …