Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (567)
- Statistical Methodology (362)
- Statistical Theory (336)
- Statistical Models (242)
- Medicine and Health Sciences (176)
-
- Survival Analysis (147)
- Public Health (142)
- Epidemiology (99)
- Life Sciences (92)
- Longitudinal Data Analysis and Time Series (89)
- Genetics and Genomics (88)
- Clinical Trials (82)
- Microarrays (78)
- Multivariate Analysis (78)
- Applied Mathematics (57)
- Genetics (57)
- Numerical Analysis and Computation (57)
- Bioinformatics (50)
- Computational Biology (50)
- Categorical Data Analysis (49)
- Design of Experiments and Sample Surveys (39)
- Clinical Epidemiology (36)
- Diseases (30)
- Disease Modeling (28)
- Medical Specialties (23)
- Health Services Research (17)
- Applied Statistics (13)
- Vital and Health Statistics (11)
- Keyword
-
- Causal inference (30)
- Cross-validation (25)
- Prediction (23)
- Genetics (21)
- Longitudinal data (19)
-
- Survival analysis (16)
- Classification (14)
- Influence curve (14)
- Model selection (14)
- Sensitivity (14)
- Bootstrap (13)
- Gene expression (13)
- Clinical trials (12)
- Targeted maximum likelihood estimation (12)
- Counterfactual (11)
- Efficient influence curve (11)
- Multiple testing (11)
- Confounding (10)
- Loss function (10)
- Missing data (10)
- Variable selection (10)
- Causal effect (9)
- Estimating equation (9)
- Measurement error (9)
- Regression (9)
- Specificity (9)
- Adjusted p-value (8)
- Air pollution (8)
- Asymptotic linearity (8)
- Censoring (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (242)
- UW Biostatistics Working Paper Series (215)
- Harvard University Biostatistics Working Paper Series (212)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (178)
- The University of Michigan Department of Biostatistics Working Paper Series (111)
Articles 751 - 780 of 1108
Full-Text Articles in Statistics and Probability
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
COBRA Preprint Series
In behavioral medicine trials, such as smoking cessation trials, two or more active treatments are often compared. Noncompliance by some subjects with their assigned treatment poses a challenge to the data analyst. Causal parameters of interest might include those defined by subpopulations based on their potential compliance status under each assignment, using the principal stratification framework (e.g., causal effect of new therapy compared to standard therapy among subjects that would comply with either intervention). Even if subjects in one arm do not have access to the other treatment(s), the causal effect of each treatment typically can only be identified from …
A Computationally Tractable Multivariate Random Effects Model For Clustered Binary Data, Brent A. Coull, E. Andres Houseman, Rebecca A. Betensky
A Computationally Tractable Multivariate Random Effects Model For Clustered Binary Data, Brent A. Coull, E. Andres Houseman, Rebecca A. Betensky
Harvard University Biostatistics Working Paper Series
No abstract provided.
Longitudinal Nested Compliance Class Model In The Presence Of Time-Varying Noncompliance, Julia Y. Lin, Thomas R. Tenhave, Michael R. Elliott
Longitudinal Nested Compliance Class Model In The Presence Of Time-Varying Noncompliance, Julia Y. Lin, Thomas R. Tenhave, Michael R. Elliott
UPenn Biostatistics Working Papers
This article discusses a nested latent class model for analyzing longitudinal randomized trials when subjects do not always adhere to the treatment to which they are randomized. In the "Prevention of Suicide in Primary Care Elderly: Collaborative Trial" (PROSPECT) study, subjects were randomized to either the control treatment, where they received standard care, or to the intervention, where they received standard care in addition to meeting with depression health specialists. The health specialists educate patients, their families, and physicians about depression and monitor their treatment. Those randomized to the control treatment have no access to the health specialists; however, those …
Hierarchical Lévy Frailty Models And A Frailty Analysis Of Data On Infant Mortality In Norwegian Siblings, Tron Anders Moger, Odd O. Aalen
Hierarchical Lévy Frailty Models And A Frailty Analysis Of Data On Infant Mortality In Norwegian Siblings, Tron Anders Moger, Odd O. Aalen
UW Biostatistics Working Paper Series
Distributions determined by non-negative Lévy processes, which include the power variance function (PVF) distributions among others, are commonly used as frailty distributions to model dependent survival times in family data. We present a hierarchical frailty model constructed by randomizing scale parameters, corresponding to time parameters of Lévy processes, in the Lévy frailty distributions. In its simplest form, this yields a two-model with heterogeneity the individual and family level. The family level frailty is shared within families, creating dependence. In the more complex models, it is extended to allow for several levels of dependence. This yields models with nested dependence structures …
Age- And Sex-Specific Transformations Of Health Status Measures To Incorporate Death, Ann M. Derleth, Paula Diehr
Age- And Sex-Specific Transformations Of Health Status Measures To Incorporate Death, Ann M. Derleth, Paula Diehr
UW Biostatistics Working Paper Series
Introduction: Measures of health status and physical function do not usually include a specific code for death. This can cause problems in longitudinal studies because analyses limited to survivors may bias the results. One approach is to recode the status variables to include a reasonable value for death. One method that has been used is to replace each scale value with the estimated probability that a person with this value will be “healthy”. “Healthy” has been defined as being above a particular threshold on the variable of interest one year later, or alternatively as being in excellent, very good, or …
Doubly Robust Censoring Unbiased Transformations, Daniel Rubin, Mark J. Van Der Laan
Doubly Robust Censoring Unbiased Transformations, Daniel Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We consider random design nonparametric regression when the response variable is subject to right censoring. Following the work of Fan and Gijbels (1994), a common approach to this problem is to apply what has been termed a censoring unbiased transformation to the data to obtain surrogate responses, and then enter these surrogate responses with covariate data into standard smoothing algorithms. Existing censoring unbiased transformations generally depend on either the conditional survival function of the response of interest, or that of the censoring variable. We show that a mapping introduced in another statistical context is in fact a censoring unbiased transformation …
New Spiked-In Probe Sets For The Affymetrix Hgu-133a Latin Square Experiment, Monnie Mcgee, Zhongxue Chen
New Spiked-In Probe Sets For The Affymetrix Hgu-133a Latin Square Experiment, Monnie Mcgee, Zhongxue Chen
COBRA Preprint Series
The Affymetrix HGU-133A spike in data set has been used for determining the sensitivity and specificity of various methods for the analysis of microarray data. We show that there are 22 additional probe sets that detect spike in RNAs that should be considered as spike in probe sets. We assign each proposed spiked-in probe set to a concentration group within the Latin Square design, and examine the effects of the additional spiked-in probe sets on assessing the accuracy of analysis methods currently in use. We show that several popular preprocessing methods are more sensitive and specific when the new spike-ins …
A Method To Increase The Power Of Multiple Testing Procedures Through Sample Splitting, Daniel Rubin, Sandrine Dudoit, Mark J. Van Der Laan
A Method To Increase The Power Of Multiple Testing Procedures Through Sample Splitting, Daniel Rubin, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Consider the standard multiple testing problem where many hypotheses are to be tested, each hypothesis is associated with a test statistic, and large test statistics provide evidence against the null hypotheses. One proposal to provide probabilistic control of Type-I errors is the use of procedures ensuring that the expected number of false positives does not exceed a user-supplied threshold. Among such multiple testing procedures, we derive the ``most powerful'' method, meaning the test statistic cutoffs that maximize the expected number of true positives. Unfortunately, these optimal cutoffs depend on the true unknown data generating distribution, so could never be used …
Integrating The Predictiveness Of A Marker With Its Performance As A Classifier, Margaret S. Pepe, Ziding Feng, Ying Huang, Gary M. Longton, Ross Prentice, Ian M. Thompson, Yingye Zheng
Integrating The Predictiveness Of A Marker With Its Performance As A Classifier, Margaret S. Pepe, Ziding Feng, Ying Huang, Gary M. Longton, Ross Prentice, Ian M. Thompson, Yingye Zheng
UW Biostatistics Working Paper Series
There are two popular statistical approaches to biomarker evaluation. One models the risk of disease (or disease outcome) using, for example, logistic regression. A marker is useful if it has a strong effect on risk. The second evaluates classification performance using measures such as sensitivity, specificity, predictive values and ROC curves. There is controversy about which approach is most appropriate. Moreover, the two approaches often give contradictory results on the same data. We present a new graphic, the predictiveness curve, that complements the risk modeling approach. It assesses the usefulness of a risk model when applied to the population. In …
Using Profile Likelihood For Semiparametric Model Selection With Application To Proportional Hazards Mixed Models, Ronghui Xu, Anthony Gamst, Michael Donohue, Florin Vaida, David P. Harrington
Using Profile Likelihood For Semiparametric Model Selection With Application To Proportional Hazards Mixed Models, Ronghui Xu, Anthony Gamst, Michael Donohue, Florin Vaida, David P. Harrington
Harvard University Biostatistics Working Paper Series
No abstract provided.
Individualized Treatment Rules: Generating Candidate Clinical Trials, Maya L. Petersen, Steven G. Deeks, Mark J. Van Der Laan
Individualized Treatment Rules: Generating Candidate Clinical Trials, Maya L. Petersen, Steven G. Deeks, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Statistical methods have rarely been applied to learn individualized treatment rules, or rules for altering treatments over time in response to changes in individual covariates. Termed dynamic treatment regimes in the statistical literature, such individualized treatment rules are of primary importance in the practice of clinical medicine. History-Adjusted Marginal Structural Models (HA-MSM) estimate individualized treatment rules that assign, at each time point, the first action of the future static treatment plan that optimizes expected outcome given a patient's covariates. However, as we discuss here, the optimality of these rules can depend on the way in which treatment was assigned in …
A Marginalized Diffusion Model For Estimating Age At First Endoscopy Examination From Current Status Data, Diana Miglioretti, Elizabeth Brown
A Marginalized Diffusion Model For Estimating Age At First Endoscopy Examination From Current Status Data, Diana Miglioretti, Elizabeth Brown
UW Biostatistics Working Paper Series
We propose an approach for estimating the age at first endoscopy examination from current status data collected via two series of cross-sectional surveys. To model the national probability of ever having an endoscopy, we incorporate birth cohort effects into a mixed-influence diffusion model. We link a state-specific model to the national-level diffusion model using a marginalized modeling approach. In future research, results from our model will be used as microsimulation model inputs to estimate the contribution of endoscopy examinations to observed changes in colorectal cancer incidence and mortality.
Posterior Simulation In The Generalized Linear Model With Semiparmetric Random Effects, Subharup Guha
Posterior Simulation In The Generalized Linear Model With Semiparmetric Random Effects, Subharup Guha
Harvard University Biostatistics Working Paper Series
Generalized linear mixed models with semiparametric random effects are useful in a wide variety of Bayesian applications. When the random effects arise from a mixture of Dirichlet process (MDP) model, normal base measures and Gibbs sampling procedures based on the Pólya urn scheme are often used to simulate posterior draws. These algorithms are applicable in the conjugate case when (for a normal base measure) the likelihood is normal. In the non-conjugate case, the algorithms proposed by MacEachern and Müller (1998) and Neal (2000) are often applied to generate posterior samples. Some common problems associated with simulation algorithms for non-conjugate MDP …
Principal Stratification Designs To Estimate Input Data Missing Due To Death, Constantine E. Frangakis, Donald B. Rubin, Ming-Wen An, Ellen Mackenzie
Principal Stratification Designs To Estimate Input Data Missing Due To Death, Constantine E. Frangakis, Donald B. Rubin, Ming-Wen An, Ellen Mackenzie
Johns Hopkins University, Dept. of Biostatistics Working Papers
We consider studies of cohorts of individuals after a critical event, such as an injury, with the following characteristics. First, the studies are designed to measure “input” variables, which describe the period before the critical event, and to characterize the distribution of the input variables in the cohort. Second, the studies are designed to measure “output” variables, primarily mortality after the critical event, and to characterize the predictive (conditional) distribution of mortality given the input variables in the cohort. Such studies often possess the complication that the input data are missing for those who die shortly after the critical event …
Multiple Imputation In The Presence Of Outliers, Michael Elliott
Multiple Imputation In The Presence Of Outliers, Michael Elliott
The University of Michigan Department of Biostatistics Working Paper Series
We consider the problem of obtaining population-based inference in the presence of missing data and outliers in the context of estimating obesity prevalence and body-mass index (BMI) measures from the Healthy For Life Study. Identifying multiple outliers in a multivariate setting is problematic because of problems such as masking, in which groups of outliers inflate the covariance matrix in a fashion that prevents their identification when included, and swamping, in which outliers skew covariances in a fashion that make non-outling observations appear to be outliers. We develop a latent class model that assumes each observation belongs to one of $K$ …
Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen
Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
Bioequivalence trials are usually conducted to compare two or more formulations of a drug. Simultaneous assessment of bioequivalence on multiple endpoints is called multivariate bioequivalence. Despite the fact that some tests for multivariate bioequivalence are suggested, current practice usually involves univariate bioequivalence assessments ignoring the correlations between the endpoints such as AUC and Cmax. In this paper we develop a semiparametric Bayesian test for bioequivalence under multiple endpoints. Specifically, we show how the correlation between the endpoints can be incorporated in the analysis and how this correlation affects the inference. Resulting estimates and posterior probabilities ``borrow strength'' from one another …
Combining Information From Two Surveys To Estimate County-Level Prevalence Rates Of Cancer Risk Factors And Screening, Trivellore E. Raghuanthan, Dawei Xie, Nathaniel Schenker, Van Parsons, William W. Davis, Kevin W. Dodd, Eric J. Feuer
Combining Information From Two Surveys To Estimate County-Level Prevalence Rates Of Cancer Risk Factors And Screening, Trivellore E. Raghuanthan, Dawei Xie, Nathaniel Schenker, Van Parsons, William W. Davis, Kevin W. Dodd, Eric J. Feuer
The University of Michigan Department of Biostatistics Working Paper Series
Cancer surveillance requires estimates of the prevalence of cancer risk factors and screening for small areas such as counties. Two popular data sources are the Behavioral Risk Factor Surveillance System (BRFSS), a telephone survey conducted by state agencies, and the National Health Interview Survey (NHIS), an area probability sample survey conducted through face-to-face interviews. Both data sources have advantages and disadvantages. The BRFSS is a larger survey, and almost every county is included in the survey; but it has lower response rates as is typical with telephone surveys, and it does not include subjects who live in households with no …
Semiparametric Latent Variable Regression Models For Spatio-Temporal Modeling Of Mobile Source Particles In The Greater Boston Area, Alexandros Gryparis, Brent A. Coull, Joel Schwartz, Helen H. Suh
Semiparametric Latent Variable Regression Models For Spatio-Temporal Modeling Of Mobile Source Particles In The Greater Boston Area, Alexandros Gryparis, Brent A. Coull, Joel Schwartz, Helen H. Suh
Harvard University Biostatistics Working Paper Series
Traffic particle concentrations show considerable spatial variability within a metropolitan area. We consider latent variable semiparametric regression models for modeling the spatial and temporal variability of black carbon and elemental carbon concentrations in the greater Boston area. Measurements of these pollutants, which are markers of traffic particles, were obtained from several individual exposure studies conducted at specific household locations as well as 15 ambient monitoring sites in the city. The models allow for both flexible, nonlinear effects of covariates and for unexplained spatial and temporal variability in exposure. In addition, the different individual exposure studies recorded different surrogates of traffic …
Estimating The Integrated Likelihood Via Posterior Simulation Using The Harmonic Mean Identity, Adrian E. Raftery, Michael A. Newton, Jaya M. Satagopan, Pavel N. Krivitsky
Estimating The Integrated Likelihood Via Posterior Simulation Using The Harmonic Mean Identity, Adrian E. Raftery, Michael A. Newton, Jaya M. Satagopan, Pavel N. Krivitsky
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
The integrated likelihood (also called the marginal likelihood or the normalizing constant) is a central quantity in Bayesian model selection and model averaging. It is defined as the integral over the parameter space of the likelihood times the prior density. The Bayes factor for model comparison and Bayesian testing is a ratio of integrated likelihoods, and the model weights in Bayesian model averaging are proportional to the integrated likelihoods. We consider the estimation of the integrated likelihood from posterior simulation output, aiming at a generic method that uses only the likelihoods from the posterior simulation iterations. The key is the …
Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan
Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many statistical methods exist that can be used to learn a predictor based on observed data. Examples include decision trees, neural networks, support vector regression, least angle regression, Logic Regression, and the Deletion/Substitution/Addition algorithm. The optimal algorithm for prediction will vary depending on the underlying data-generating distribution. In this article, we introduce a "super learner," a prediction algorithm that applies any set of candidate learners and uses cross-validation to select among them. Theory shows that asymptotically the super learner performs essentially as well or better than any of the candidate learners. We briefly present the theory behind the super learner, …
Recurrent Event Models In The Presence Of A Terminal Event: Comparison, Inference And Data Analysis, Xianghua Luo, Mei-Cheng Wang
Recurrent Event Models In The Presence Of A Terminal Event: Comparison, Inference And Data Analysis, Xianghua Luo, Mei-Cheng Wang
Johns Hopkins University, Dept. of Biostatistics Working Papers
This article focuses on statistical implications of proportional rate models for recurrent event data in the presence of a terminal event. In such circumstances, various definitions of the recurrent rate function have been adopted in the proportional rate models. Although these rate functions have quite different interpretations, recognition of the differences has been lacking theoretically and practically. We compare three types of rate functions from both conceptual and quantitative perspectives; conclude that the inappropriate choice of a rate function may lead to misleading scientific conclusions. Simulations are conducted for comparisons of the focused models. Analysis of data from an AIDS …
Empirical Bayes Approach To Controlling Familywise Error: An Application To Hiv Resistance Data, Rhoderick N. Machekano, Alan E. Hubbard
Empirical Bayes Approach To Controlling Familywise Error: An Application To Hiv Resistance Data, Rhoderick N. Machekano, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Statistical challenges arise in identifying meaningful patterns and structures from high dimensional genomic data sets. Relating HIV genotype (sequence of amino acids) to phenotypic resistance presents a typical problem. When the HIV virus is under antiretroviral drug pressure, unfavorable mutations of the target genes often lead to greatly increased resistance of the virus to drugs, including drugs the virus has not been exposed to. Identification of mutation combinations and their correlation to drug resistance is critical in guiding efficient prescription of HIV drugs. The identification of a subset of codons associated with drug resistance from a set of several hundreds …
Survival Analysis With Change Point Hazard Functions, Melody S. Goodman, Yi Li, Ram C. Tiwari
Survival Analysis With Change Point Hazard Functions, Melody S. Goodman, Yi Li, Ram C. Tiwari
Harvard University Biostatistics Working Paper Series
No abstract provided.
Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan
Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
An important class of models in causal inference are the so-called marginal structural models which model the comparison between counterfactual outcome distributions corresponding with a static treatment intervention, conditional on user supplied baseline covariates, based on observing a longitudinal data structure on a sample of n independent and identically distributed experimental units. Identification of a static treatment regimen specific outcome distribution based on observational data requires beyond the so-called sequential randomization assumption that each experimental unit has positive probability of following the static treatment regimen. The latter assumption is called the experimental treatment assignment assumption (ETA) (which is parameter specific). …
Reliability, Effect Size, And Responsiveness And Intraclass Correlation Of Health Status Measures Used In Randomized And Cluster-Randomized Trials, Paula Diehr, Lu Chen, Donald L. Patrick, Ziding Feng, Yutaka Yasui
Reliability, Effect Size, And Responsiveness And Intraclass Correlation Of Health Status Measures Used In Randomized And Cluster-Randomized Trials, Paula Diehr, Lu Chen, Donald L. Patrick, Ziding Feng, Yutaka Yasui
UW Biostatistics Working Paper Series
Background: New health status instruments are described by psychometric properties, such as Reliability, Effect Size, and Responsiveness. For cluster-randomized trials, another important statistic is the Intraclass Correlation for the instrument within clusters. Studies using better instruments can be performed with smaller sample sizes, but better instruments may be more expensive in terms of dollars, lost opportunities, or poorer data quality due to the response burden of longer instruments. Investigators often need to estimate the psychometric properties of a new instrument, or of an established instrument in a new setting. Optimal sample sizes for estimating these properties have not been studied …
Detecting Pulsatile Hormone Secretion Events: A Bayesian Approach, Tim Johnson
Detecting Pulsatile Hormone Secretion Events: A Bayesian Approach, Tim Johnson
The University of Michigan Department of Biostatistics Working Paper Series
Many challenges arise in the analysis of pulsatile, or episodic, hormone concentration time series data. Among these challenges is the determination of the number and location of pulsatile events and the discrimination of events from noise. Analyses of these data are typically performed in two stages. In the first stage, the number and approximate location of the pulses are determined. In the second stage, a model (typically a deconvolution model) is fit to the data conditional on the number of pulses. Any error made in the first stage is carried over to the second stage. Furthermore, current methods, except two, …
Semiparametric Analysis For Correlated Recurrent And Terminal Events, Yining Ye, Jack Kalbfleisch, Doug E. Schaubel
Semiparametric Analysis For Correlated Recurrent And Terminal Events, Yining Ye, Jack Kalbfleisch, Doug E. Schaubel
The University of Michigan Department of Biostatistics Working Paper Series
In clinical and observational studies, recurrent event data (e.g. hospitalization) with a terminal event (e.g. death) are often encountered. In many instances, the terminal event is strongly correlated with the recurrent event process. In this article, we propose a semiparametric method to jointly model the recurrent and terminal event processes. The dependence is modeled by a shared gamma frailty that is included in both the recurrent event rate and terminal event hazard function. Marginal models are used to estimate the regression effects on the terminal and recurrent event processes and a Poisson model is used to estimate the dispersion of …
Adjusting For Covariate Effects On Classification Accuracy Using The Covariate-Adjusted Roc Curve, Holly Janes, Margaret S. Pepe
Adjusting For Covariate Effects On Classification Accuracy Using The Covariate-Adjusted Roc Curve, Holly Janes, Margaret S. Pepe
UW Biostatistics Working Paper Series
Recent scientific and technological innovations have produced an abundance of potential markers which are being investigated for their use in disease screen- ing and diagnosis. In evaluating these markers, it is often necessary to account for covariates which are associated with the marker of interest. These covariates may include subject characteristics, expertise of the test operator, test proce- dures, or aspects of specimen handling. In this paper, we propose the AROC, a covariate-adjusted measure of the classification accuracy. The AROC is the common covariate-specific ROC curve, when the covariate does not affect dis- crimination, and a weighted average of covariate-specific …
Censored Data Regression In High-Dimension And Low-Sample Size Settings For Genomic Applications, Hongzhe Li
Censored Data Regression In High-Dimension And Low-Sample Size Settings For Genomic Applications, Hongzhe Li
UPenn Biostatistics Working Papers
New high-throughput technologies are generating various types of high-dimensional genomic and proteomic data and meta-data (e.g., networks and pathways) in order to obtain a systems-level understanding of various complex diseases such as human cancers and cardiovascular diseases. As the amount and complexity of the data increase and as the questions being addressed become more sophisticated, we face the great challenge of how to model such data in order to draw valid statistical and biological conclusions. One important problem in genomic research is to relate these high-throughput genomic data to various clinical outcomes, including possibly censored survival outcomes such as age …
A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska
A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper proposes a statistical methodology for comparing the performance of evolutionary computation algorithms. A two-fold sampling scheme for collecting performance data is introduced, and these data are analyzed using bootstrap-based multiple hypothesis testing procedures. The proposed method is sufficiently flexible to allow the researcher to choose how performance is measured, does not rely upon distributional assumptions, and can be extended to analyze many other randomized numeric optimization routines. As a result, this approach offers a convenient, flexible, and reliable technique for comparing algorithms in a wide variety of applications.