Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (567)
- Statistical Methodology (362)
- Statistical Theory (336)
- Statistical Models (242)
- Medicine and Health Sciences (176)
-
- Survival Analysis (147)
- Public Health (142)
- Epidemiology (99)
- Life Sciences (92)
- Longitudinal Data Analysis and Time Series (89)
- Genetics and Genomics (88)
- Clinical Trials (82)
- Microarrays (78)
- Multivariate Analysis (78)
- Applied Mathematics (57)
- Genetics (57)
- Numerical Analysis and Computation (57)
- Bioinformatics (50)
- Computational Biology (50)
- Categorical Data Analysis (49)
- Design of Experiments and Sample Surveys (39)
- Clinical Epidemiology (36)
- Diseases (30)
- Disease Modeling (28)
- Medical Specialties (23)
- Health Services Research (17)
- Applied Statistics (13)
- Vital and Health Statistics (11)
- Keyword
-
- Causal inference (30)
- Cross-validation (25)
- Prediction (23)
- Genetics (21)
- Longitudinal data (19)
-
- Survival analysis (16)
- Classification (14)
- Influence curve (14)
- Model selection (14)
- Sensitivity (14)
- Bootstrap (13)
- Gene expression (13)
- Clinical trials (12)
- Targeted maximum likelihood estimation (12)
- Counterfactual (11)
- Efficient influence curve (11)
- Multiple testing (11)
- Confounding (10)
- Loss function (10)
- Missing data (10)
- Variable selection (10)
- Causal effect (9)
- Estimating equation (9)
- Measurement error (9)
- Regression (9)
- Specificity (9)
- Adjusted p-value (8)
- Air pollution (8)
- Asymptotic linearity (8)
- Censoring (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (242)
- UW Biostatistics Working Paper Series (215)
- Harvard University Biostatistics Working Paper Series (212)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (178)
- The University of Michigan Department of Biostatistics Working Paper Series (111)
Articles 781 - 810 of 1108
Full-Text Articles in Statistics and Probability
The Two-Sample Problem For Failure Rates Depending On A Continuous Mark: An Application To Vaccine Efficacy, Peter B. Gilbert, Ian W. Mckeague, Yanqing Sun
The Two-Sample Problem For Failure Rates Depending On A Continuous Mark: An Application To Vaccine Efficacy, Peter B. Gilbert, Ian W. Mckeague, Yanqing Sun
UW Biostatistics Working Paper Series
The efficacy of an HIV vaccine to prevent infection is likely to depend on the genetic variation of the exposing virus. This paper addresses the problem of using data on the HIV sequences that infect vaccine efficacy trial participants to 1) test for vaccine efficacy more powerfully than procedures that ignore the sequence data; and 2) evaluate the dependence of vaccine efficacy on the divergence of infecting HIV strains from the HIV strain that is contained in the vaccine. Because hundreds of amino acid sites in each HIV genome are sequenced, it is natural to treat the divergence (defined in …
Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes
Genome Scanning Methods For Comparing Sequences Between Groups, With Application To Hiv Vaccine Trials, Peter B. Gilbert, Chunyuan Wu, David V. Jobes
UW Biostatistics Working Paper Series
Consider a placebo-controlled preventive HIV vaccine efficacy trial. An HIV amino acid sequence is measured from each volunteer who acquires HIV, and these sequences are aligned together with the reference HIV sequence represented in the vaccine. We develop genome scanning methods to identify HIV positions at which the amino acids in sequences from infected vaccine recipients tend to be more divergent from the corresponding reference amino acid than the amino acids in sequences from infected placebo recipients. We consider five two-sample test statistics, based on Euclidean, Mahalanobis, and Kullback-Leibler divergence measures. Weights are incorporated to reflect biological information contained in …
On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger
On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger
Johns Hopkins University, Dept. of Biostatistics Working Papers
Time series and case-crossover methods are often viewed as competing alternatives in environmental epidemiologic studies. Several recent studies have compared the time series and case-crossover methods. In this paper, we show that case-crossover using conditional logistic regression is a special case of time series analysis when there is a common exposure such as in air pollution studies. This equivalence provides computational convenience for case-crossover analyses and a better understanding of time series models. Time series log-linear regression accounts for over-dispersion of the Poisson variance, while case-crossover analyses typically do not. This equivalence also permits model checking for case-crossover data using …
Evaluating Prediction Rules For T-Year Survivors With Censored Regression Models, Hajime Uno, Tianxi Cai, Lu Tian, L.J. Wei
Evaluating Prediction Rules For T-Year Survivors With Censored Regression Models, Hajime Uno, Tianxi Cai, Lu Tian, L.J. Wei
Harvard University Biostatistics Working Paper Series
Suppose that we are interested in establishing simple, but reliable rules for predicting future t-year survivors via censored regression models. In this article, we present inference procedures for evaluating such binary classification rules based on various prediction precision measures quantified by the overall misclassification rate, sensitivity and specificity, and positive and negative predictive values. Specifically, under various working models we derive consistent estimators for the above measures via substitution and cross validation estimation procedures. Furthermore, we provide large sample approximations to the distributions of these nonsmooth estimators without assuming that the working model is correctly specified. Confidence intervals, for example, …
2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr
2^K Factorials In Blocks Of Size 2, With Application To Two-Color Microarray Experiments, Kathleen F. Kerr
UW Biostatistics Working Paper Series
When a two-level design must be run in blocks of size two, there is a unique blocking scheme that enables estimation of all the main effects. Unfortunately this design does not enable estimation of any two-factor interactions. When the experimental goal is to estimate all main effects and two-factor interactions, it is necessary to combine replicates of the experiment that use different blocking schemes. In this paper we identify such designs for up to eight factors that enable estimation of all main effects and two-factor interactions with the fewest number of replications. In addition, we give a construction for general …
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a general and formal statistical framework for the multiple tests of associations between known fixed features of a genome and unknown parameters of the distribution of variable features of this genome in a population of interest. The known fixed gene-annotation profiles, corresponding to the fixed features of the genome, may concern Gene Ontology (GO) annotation, pathway membership, regulation by particular transcription factors, nucleotide sequences, or protein sequences. The unknown gene-parameter profiles, corresponding to the variable features of the genome, may be, for example, regression coefficients relating genome-wide transcript levels or DNA copy numbers to possibly censored biological and …
A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull
A Diagnostic Test For The Mixing Distribution In A Generalised Linear Mixed Model, Eric J. Tchetgen, Brent A. Coull
Harvard University Biostatistics Working Paper Series
We introduce a diagnostic test for the mixing distribution in a generalised linear mixed model. The test is based on the difference between the marginal maximum likelihood and conditional maximum likelihood estimates of a subset of the fixed effects in the model. We derive the asymptotic variance of this difference, and propose a test statistic that has a limiting chi-square distribution under the null hypothesis that the mixing distribution is correctly specified. For the important special case of the logistic regression model with random intercepts, we evaluate via simulation the power of the test in finite samples under several alternative …
Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng
Evaluating The Predictiveness Of A Continuous Marker, Ying Huang, Margaret S. Pepe, Ziding Feng
UW Biostatistics Working Paper Series
Consider a continuous marker for predicting a binary outcome. For example, serum concentration of prostate specific antigen (PSA) may be used to calculate the risk of finding prostate cancer in a biopsy. In this paper we argue that the predictive capacity of a marker has to do with the population distribution of risk given the marker and suggest a graphical tool, the predictiveness curve, that displays this distribution. The display provides a common meaningful scale for comparing markers that may not be comparable on their original scales. Some existing measures of predictiveness are shown to be summary indices derived from …
Different Public Health Interventions Have Varying Effects, Paula Diehr, Anne B. Newman, Liming Cai, Ann Derleth
Different Public Health Interventions Have Varying Effects, Paula Diehr, Anne B. Newman, Liming Cai, Ann Derleth
UW Biostatistics Working Paper Series
Objective: To compare performance of one-time health interventions to those that change the probability of transitioning from one health state to another. Study Design and Setting: We used multi-state life table methods to estimate the impact of eight types of interventions on several outcomes. Results: In a cohort beginning at age 65, curing all the sick persons at baseline would increase life expectancy by 0.23 years and increase years of healthy life by .54 years. An equal amount of improvement could be obtained with a 12% decrease in the probability of getting sick, a 16% increase in the probability of …
Survival Analysis Methods In Genetic Epidemiology, Hongzhe Li
Survival Analysis Methods In Genetic Epidemiology, Hongzhe Li
UPenn Biostatistics Working Papers
Mapping genes for complex human diseases is a challenging problem due to the fact that many such diseases are due to both genetic and enviromental risk factors and many also exhibit phenotypic heterogeneity, such as variable age of onset. Information on variable age of disease onset is often a good indicator for disease heterogeneity and incorporation of such information together with enviromental risk factors into genetic analysis should lead to more powerful tests for genetic analysis. Due to the problem of censoring, survival analysis methods have proved to be very useful for genetic analysis. In this paper, I review some …
On The Violation Of Bounds For The Correlation In Generalized Estimating Equation Analyses Of Binary Data From Longitudinal Trials, Justine Shults, Wenguang Sun, Xin Tu, Jay Amsterdam
On The Violation Of Bounds For The Correlation In Generalized Estimating Equation Analyses Of Binary Data From Longitudinal Trials, Justine Shults, Wenguang Sun, Xin Tu, Jay Amsterdam
UPenn Biostatistics Working Papers
It is well-known that the correlation among binary outcomes is constrained by the marginal means, yet approaches such as generalized estimating equations (GEE) do not check that the constraints for the correlations are satisfied. We explore this issue for Markovian dependence in the context of a GEE analysis of a clinical trial that compares Venlafaxine with Lithium in the prevention of major depressive episode. We obtain simplified expressions for the constraints for the logistic model and the equicorrelated and first-order autoregressive correlation structures. We then obtain the limiting values of the GEE and quasi-least squares (QLS) estimates of the correlation …
Comparing The Predictive Values Of Diagnostic Tests: Sample Size And Analysis For Paired Study Designs, Chaya S. Moskowitz, Margaret S. Pepe
Comparing The Predictive Values Of Diagnostic Tests: Sample Size And Analysis For Paired Study Designs, Chaya S. Moskowitz, Margaret S. Pepe
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
In this paper we consider the design and analysis of studies comparing the positive and negative predictive values of two diagnostic tests that are measured on all subjects. Although statistical methodology is well developed for comparing diagnostic tests in terms of their sensitivities and specificities, comparative inference about predictive values is not. We derive analytic variance expressions for the relative predictive values. Sample size formulas for study design ensue. In addition, two new methods for analyzing the resulting data are presented and compared with an existing marginal regression methodology.
Regression Analysis For The Partial Area Under The Roc Curve, Tianxi Cai, Lori E. Dodd
Regression Analysis For The Partial Area Under The Roc Curve, Tianxi Cai, Lori E. Dodd
Harvard University Biostatistics Working Paper Series
No abstract provided.
Use Of Unbiased Estimating Equations To Estimate Correlation In Generalized Estimating Equation Analysis Of Longitudinal Trials, Wenguang Sun, Justine Shults, Mary Leonard
Use Of Unbiased Estimating Equations To Estimate Correlation In Generalized Estimating Equation Analysis Of Longitudinal Trials, Wenguang Sun, Justine Shults, Mary Leonard
UPenn Biostatistics Working Papers
In a recent publication, Wang and Carey (Journal of the American Statistical Association, 99, pp. 845-853, 2004) presented a new approach for estimation of the correlation parameters in the framework of generalized estimating equations (GEE). They considered correlated continuous, binary and count data with a generalized Markov correlation structure that includes the first-order autoregressive AR(1) and Markov structures as special cases. They made detailed comparisons with pseudo-likelihood (PL) and the first stage of quasi-least squares (QLS), a two-stage approach in the framework of generalized estimating equations (GEE). In this note we extend their comparisons for the second (bias corrected) stage …
Case-Cohort Methods For Survival Data On Families From Routine Registers, Tron Anders Moger, Yudi Pawitan, Ørnulf Borgan
Case-Cohort Methods For Survival Data On Families From Routine Registers, Tron Anders Moger, Yudi Pawitan, Ørnulf Borgan
UW Biostatistics Working Paper Series
In the Nordic countries, there exist several registers containing information on diseases and risk factors for millions of individuals. This information can be linked into families by use of personal identification numbers, and represent a great opportunity for studying diseases that show familial aggregation. Due to the size of the registers, it is difficult to analyze the data by using traditional methods for multivariate survival analysis, such as frailty or copula models. Since the size of the cohort is known, case-cohort methods based on pseudo-likelihoods are suitable for analyzing the data. We present methods for sampling control families both with …
Comparison Of Haplotype-Based And Tree-Based Snp Imputation In Association Studies, James Y. Dai, Ingo Ruczinski, Michael Leblanc, Charles Kooperberg
Comparison Of Haplotype-Based And Tree-Based Snp Imputation In Association Studies, James Y. Dai, Ingo Ruczinski, Michael Leblanc, Charles Kooperberg
UW Biostatistics Working Paper Series
Missing single nucleotide polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we compare two haplotype-based imputation approaches and one regression tree-based imputation approach for association studies. The goal is to assess the imputation accuracy, and to evaluate the impact of imputation on parameter estimation. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to …
Semiparametric Approaches For Joint Modeling Of Longitudinal And Survival Data With Time Varying Coefficients, Xiao Song, C.Y. Wang
Semiparametric Approaches For Joint Modeling Of Longitudinal And Survival Data With Time Varying Coefficients, Xiao Song, C.Y. Wang
UW Biostatistics Working Paper Series
We study joint modeling of survival and longitudinal data. There are two regression models of interest. The primary model is for survival outcomes, which are assumed to follow a time varying coefficient proportional hazards model. The second model is for longitudinal data, which are assumed to follow a random effects model. Based on the trajectory of a subject's longitudinal data, some covariates in the survival model are functions of the unobserved random effects. Estimated random effects are generally different from the unobserved random effects and hence this leads to covariate measurement error. To deal with covariate measurement error, we propose …
Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson
Alleviating Linear Ecological Bias And Optimal Design With Subsample Data, Adam Glynn, Jon Wakefield, Mark Handcock, Thomas Richardson
UW Biostatistics Working Paper Series
In this paper, we illustrate that combining ecological data with subsample data in situations in which a linear model is appropriate provides three main benefits. First, by including the individual level subsample data, the biases associated with linear ecological inference can be eliminated. Second, by supplementing the subsample data with ecological data, the information about parameters will be increased. Third, we can use readily available ecological data to design optimal subsampling schemes, so as to further increase the information about parameters. We present an application of this methodology to the classic problem of estimating the effect of a college degree …
Bayesian Analysis Of Cell-Cycle Gene Expression Data, Chuan Zhou, Jon Wakefield, Linda Breeden
Bayesian Analysis Of Cell-Cycle Gene Expression Data, Chuan Zhou, Jon Wakefield, Linda Breeden
UW Biostatistics Working Paper Series
The study of the cell-cycle is important in order to aid in our understanding of the basic mechanisms of life, yet progress has been slow due to the complexity of the process and our lack of ability to study it at high resolution. Recent advances in microarray technology have enabled scientists to study the gene expression at the genome-scale with a manageable cost, and there has been an increasing effort to identify cell-cycle regulated genes. In this chapter, we discuss the analysis of cell-cycle gene expression data, focusing on a model-based Bayesian approaches. The majority of the models we describe …
Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou
Empirical Likelihood Inference For The Area Under The Roc Curve, Gengsheng Qin, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
For a continuous-scale diagnostic test, the most commonly used summary index of the receiver operating characteristic (ROC) curve is the area under the curve (AUC) that measures the accuracy of the diagnostic test. In this paper we propose an empirical likelihood approach for the inference of AUC. We first define an empirical likelihood ratio for AUC and show that its limiting distribution is a scaled chi-square distribution. We then obtain an empirical likelihood based confidence interval for AUC using the scaled chi-square distribution. This empirical likelihood inference for AUC can be extended to stratified samples and the resulting limiting distribution …
Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou
Interval Estimation For The Ratio And Difference Of Two Lognormal Means, Yea-Hung Chen, Xiao-Hua Zhou
UW Biostatistics Working Paper Series
Health research often gives rise to data that follow lognormal distributions. In two sample situations, researchers are likely to be interested in estimating the difference or ratio of the population means. Several methods have been proposed for providing confidence intervals for these parameters. However, it is not clear which techniques are most appropriate, or how their performance might vary. Additionally, methods for the difference of means have not been adequately explored. We discuss in the present article five methods of analysis. These include two methods based on the log-likelihood ratio statistic and a generalized pivotal approach. Additionally, we provide and …
Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li
Inferences In Censored Cost Regression Models With Empirical Likelihood, Xiao-Hua Zhou, Gengsheng Qin, Huazhen Lin, Gang Li
UW Biostatistics Working Paper Series
In many studies of health economics, we are interested in the expected total cost over a certain period for a patient with given characteristics. Problems can arise if cost estimation models do not account for distributional aspects of costs. Two such problems are 1) the skewed nature of the data and 2) censored observations. In this paper we propose an empirical likelihood (EL) method for constructing a confidence region for the vector of regression parameters and a confidence interval for the expected total cost of a patient with the given covariates. We show that this new method has good theoretical …
Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau
Confidence Intervals For Predictive Values Using Data From A Case Control Study, Nathaniel David Mercaldo, Xiao-Hua Zhou, Kit F. Lau
UW Biostatistics Working Paper Series
The accuracy of a binary-scale diagnostic test can be represented by sensitivity (Se), specificity (Sp) and positive and negative predictive values (PPV and NPV). Although Se and Sp measure the intrinsic accuracy of a diagnostic test that does not depend on the prevalence rate, they do not provide information on the diagnostic accuracy of a particular patient. To obtain this information we need to use PPV and NPV. Since PPV and NPV are functions of both the intrinsic accuracy and the prevalence of the disease, constructing confidence intervals for PPV and NPV for a particular patient in a population with …
Model Checking For Roc Regression Analysis, Tianxi Cai, Yingye Zheng
Model Checking For Roc Regression Analysis, Tianxi Cai, Yingye Zheng
Harvard University Biostatistics Working Paper Series
The Receiver Operating Characteristic (ROC) curve is a prominent tool for characterizing the accuracy of continuous diagnostic test. To account for factors that might invluence the test accuracy, various ROC regression methods have been proposed. However, as in any regression analysis, when the assumed models do not fit the data well, these methods may render invalid and misleading results. To date practical model checking techniques suitable for validating existing ROC regression models are not yet available. In this paper, we develop cumulative residual based procedures to graphically and numerically assess the goodness-of-fit for some commonly used ROC regression models, and …
On The Use Of Non-Euclidean Isotropy In Geostatistics, Frank C. Curriero
On The Use Of Non-Euclidean Isotropy In Geostatistics, Frank C. Curriero
Johns Hopkins University, Dept. of Biostatistics Working Papers
This paper investigates the use of non-Euclidean distances to characterize isotropic spatial dependence for geostatistical related applications. A simple example is provided to demonstrate there are no guarantees that existing covariogram and variogram functions remain valid (i.e.\ positive definite or conditionally negative definite) when used with a non-Euclidean distance measure. Furthermore, satisfying the conditions of a metric is not sufficient to ensure the distance measure can be used with existing functions. Current literature is not clear on these topics. There are certain distance measures that when used with existing covariogram and variogram functions remain valid, an issue that is explored. …
Gradient Directed Regularization For Sparse Gaussian Concentration Graphs, With Applications To Inference Of Genetic Networks, Hongzhe Li, Jiang Gui
Gradient Directed Regularization For Sparse Gaussian Concentration Graphs, With Applications To Inference Of Genetic Networks, Hongzhe Li, Jiang Gui
UPenn Biostatistics Working Papers
Large-scale microarray gene expression data provide the possibility of constructing genetic networks or biological pathways. Gaussian graphical models have been suggested to provide an effective method for constructing such genetic networks. However, most of the available methods for constructing Gaussian graphs do not account for the sparsity of the networks and are computationally more demanding or infeasible, especially in the settings of high-dimension and low sample size. We introduce a threshold gradient descent regularization procedure for estimating the sparse precision matrix in the setting of Gaussian graphical models and demonstrate its application to identifying genetic networks. Such a procedure is …
Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith
Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith
U.C. Berkeley Division of Biostatistics Working Paper Series
A new data filtering method for SELDI-TOF MS proteomic spectra data is described. We examined technical repeats (2 per subject) of intensity versus m/z (mass/charge) of bone marrow cell lysate for two groups of childhood leukemia patients: acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). As others have noted, the type of data processing as well as experimental variability can have a disproportionate impact on the list of "interesting" proteins (see Baggerly et al. (2004)). We propose a list of processing and multiple testing techniques to correct for 1) background drift; 2) filtering using smooth regression and cross-validated bandwidth …
Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard
Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing a collection of null hypotheses about a data generating distribution based on a sample of independent and identically distributed observations is a fundamental and important statistical problem involving many applications. Methods based on marginal null distributions (i.e., marginal p-values) are attractive since the marginal p-values can be based on a user supplied choice of marginal null distributions and they are computationally trivial, but they, by necessity, are known to either be conservative or to rely on assumptions about the dependence structure between the test-statistics. Resampling based multiple testing (Westfall and Young, 1993) involves sampling from a joint null …
Data Adaptive Pathway Testing, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
Data Adaptive Pathway Testing, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
A majority of diseases are caused by a combination of factors, for example, composite genetic mutation profiles have been found in many cases to predict a deleterious outcome. There are several statistical techniques that have been used to analyze these types of biological data. This article implements a general strategy which uses data adaptive regression methods to build a specific pathway model, thus predicting a disease outcome by a combination of biological factors and assesses the significance of this model, or pathway, by using a permutation based null distribution. We also provide several simulation comparisons with other techniques. In addition, …
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
Nearest centroid classifiers have recently been successfully employed in high-dimensional applications. A necessary step when building a classifier for high-dimensional data is feature selection. Feature selection is typically carried out by computing univariate statistics for each feature individually, without consideration for how a subset of features performs as a whole. For subsets of a given size, we characterize the optimal choice of features, corresponding to those yielding the smallest misclassification rate. Furthermore, we propose an algorithm for estimating this optimal subset in practice. Finally, we investigate the applicability of shrinkage ideas to nearest centroid classifiers. We use gene-expression microarrays for …