Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (118)
- Statistical Theory (116)
- Statistical Methodology (114)
- Statistical Models (60)
- Survival Analysis (48)
-
- Medicine and Health Sciences (31)
- Epidemiology (26)
- Public Health (26)
- Life Sciences (20)
- Multivariate Analysis (20)
- Longitudinal Data Analysis and Time Series (19)
- Genetics and Genomics (18)
- Applied Mathematics (17)
- Numerical Analysis and Computation (17)
- Genetics (15)
- Clinical Trials (13)
- Microarrays (12)
- Design of Experiments and Sample Surveys (7)
- Categorical Data Analysis (5)
- Bioinformatics (4)
- Computational Biology (4)
- Disease Modeling (4)
- Diseases (4)
- Laboratory and Basic Science Research (4)
- Applied Statistics (3)
- Medical Specialties (3)
- Other Statistics and Probability (1)
- Vital and Health Statistics (1)
- Keyword
-
- Cross-validation (23)
- Causal inference (20)
- Influence curve (14)
- Prediction (12)
- Efficient influence curve (11)
-
- Model selection (11)
- Targeted maximum likelihood estimation (11)
- Loss function (10)
- Bootstrap (9)
- Causal effect (9)
- Counterfactual (9)
- Adjusted p-value (8)
- Asymptotic linearity (8)
- Multiple testing (8)
- Super-learning (8)
- Type I error rate (8)
- Confounding (7)
- Empirical process (7)
- One-step estimator (7)
- Survival analysis (7)
- Counting process (6)
- Estimating equation (6)
- False discovery rate (6)
- Gene expression (6)
- Null distribution (6)
- Pathwise differentiable parameter (6)
- Semiparametric statistical model (6)
- Asymptotic control (5)
- Censored data (5)
- Censoring (5)
Articles 121 - 150 of 242
Full-Text Articles in Statistics and Probability
Covariate Adjustment For The Intention-To-Treat Parameter With Empirical Efficiency Maximization, Daniel B. Rubin, Mark J. Van Der Laan
Covariate Adjustment For The Intention-To-Treat Parameter With Empirical Efficiency Maximization, Daniel B. Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In randomized experiments, the intention-to-treat parameter is defined as the difference in expected outcomes between groups assigned to treatment and control arms. There is a large literature focusing on how (possibly misspecified) working models can sometimes exploit baseline covariate measurements to gain precision, although covariate adjustment is not strictly necessary. In Rubin and van der Laan (2008), we proposed the technique of empirical efficiency maximization for improving estimation by forming nonstandard fits of such working models. Considering a more realistic randomization scheme than in our original article, we suggest a new class of working models for utilizing covariate information, show …
Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan
Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Regression models are often used to test for cause-effect relationships from data collected in randomized trials or experiments. This practice has deservedly come under heavy scrutiny, since commonly used models such as linear and logistic regression will often not capture the actual relationships between variables, and incorrectly specified models potentially lead to incorrect conclusions. In this paper, we focus on hypothesis test of whether the treatment given in a randomized trial has any effect on the mean of the primary outcome, within strata of baseline variables such as age, sex, and health status. Our primary concern is ensuring that such …
Loss-Based Estimation With Evolutionary Algorithms And Cross-Validation, David Shilane, Richard H. Liang, Sandrine Dudoit
Loss-Based Estimation With Evolutionary Algorithms And Cross-Validation, David Shilane, Richard H. Liang, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Many statistical inference methods rely upon selection procedures to estimate a parameter of the joint distribution of explanatory and outcome data, such as the regression function. Within the general framework for loss-based estimation of Dudoit and van der Laan, this project proposes an evolutionary algorithm (EA) as a procedure for risk optimization. We also analyze the size of the parameter space for polynomial regression under an interaction constraints along with constraints on either the polynomial or variable degree.
Resampling-Based Empirical Bayes Multiple Testing Procedures For Controlling Generalized Tail Probability And Expected Value Error Rates: , Sandrine Dudoit, Houston N. Gilbert, Mark J. Van Der Laan
Resampling-Based Empirical Bayes Multiple Testing Procedures For Controlling Generalized Tail Probability And Expected Value Error Rates: , Sandrine Dudoit, Houston N. Gilbert, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
This article proposes resampling-based empirical Bayes multiple testing procedures for controlling a broad class of Type I error rates, defined as generalized tail probability (gTP) error rates, gTP(q,g) = Pr(g(Vn,Sn) > q), and generalized expected value (gEV) error rates, gEV(g) = [g(Vn,Sn)], for arbitrary functions g(Vn,Sn) of the numbers of false positives Vn and true positives Sn. Of particular interest are error rates based on the …
A Note On Targeted Maximum Likelihood And Right Censored Data, Mark J. Van Der Laan, Daniel Rubin
A Note On Targeted Maximum Likelihood And Right Censored Data, Mark J. Van Der Laan, Daniel Rubin
U.C. Berkeley Division of Biostatistics Working Paper Series
A popular way to estimate an unknown parameter is with substitution, or evaluating the parameter at a likelihood based fit of the data generating density. In many cases, such estimators have substantial bias and can fail to converge at the parametric rate. van der Laan and Rubin (2006) introduced targeted maximum likelihood learning, removing these shackles from substitution estimators, which were made in full agreement with the locally efficient estimating equation procedures as presented in Robins and Rotnitzsky (1992) and van der Laan and Robins (2003). This note illustrates how targeted maximum likelihood can be applied in right censored data …
Detailed Version: Analyzing Direct Effects In Randomized Trials With Secondary Interventions: An Application To Hiv Prevention Trials, Michael A. Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
Detailed Version: Analyzing Direct Effects In Randomized Trials With Secondary Interventions: An Application To Hiv Prevention Trials, Michael A. Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
U.C. Berkeley Division of Biostatistics Working Paper Series
This is the detailed technical report that accompanies the paper “Analyzing Direct Effects in Randomized Trials with Secondary Interventions: An Application to HIV Prevention Trials” (an unpublished, technical report version of which is available online at http://www.bepress.com/ucbbiostat/paper223).
The version here gives full details of the models for the time-dependent analysis, and presents further results in the data analysis section. The Methods for Improving Reproductive Health in Africa (MIRA) trial is a recently completed randomized trial that investigated the effect of diaphragm and lubricant gel use in reducing HIV infection among susceptible women. 5,045 women were randomly assigned to either the …
Analyzing Direct Effects In Randomized Trials With Secondary Interventions , Michael Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
Analyzing Direct Effects In Randomized Trials With Secondary Interventions , Michael Rosenblum, Nicholas P. Jewell, Mark J. Van Der Laan, Stephen Shiboski, Ariane Van Der Straten, Nancy Padian
U.C. Berkeley Division of Biostatistics Working Paper Series
The Methods for Improving Reproductive Health in Africa (MIRA) trial is a recently completed randomized trial that investigated the effect of diaphragm and lubricant gel use in reducing HIV infection among susceptible women. 5,045 women were randomly assigned to either the active treatment arm or not. Additionally, all subjects in both arms received intensive condom counselling and provision, the "gold standard" HIV prevention barrier method. There was much lower reported condom use in the intervention arm than in the control arm, making it difficult to answer important public health questions based solely on the intention-to-treat analysis. We adapt an analysis …
Biomarker Discovery Using Targeted Maximum Likelihood Estimation: Application To The Treatment Of Antiretroviral Resistant Hiv Infection, Oliver Bembom, Maya L. Petersen , Soo-Yon Rhee , W. Jeffrey Fessel , Sandra E. Sinisi, Robert W. Shafer, Mark J. Van Der Laan
Biomarker Discovery Using Targeted Maximum Likelihood Estimation: Application To The Treatment Of Antiretroviral Resistant Hiv Infection, Oliver Bembom, Maya L. Petersen , Soo-Yon Rhee , W. Jeffrey Fessel , Sandra E. Sinisi, Robert W. Shafer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Researchers in clinical science and bioinformatics frequently aim to learn which of a set of candidate biomarkers is important in determining a given outcome, and to rank the contributions of the candidates accordingly. This article introduces a new approach to research questions of this type, based on targeted maximum likelihood estimation of variable importance measures.
The methodology is illustrated using an example drawn from the treatment of HIV infection. Specifically, given a list of candidate mutations in the protease enzyme of HIV, we aim to discover mutations that reduce clinical virologic response to antiretroviral regimens containing the protease inhibitor lopinavir. …
Empirical Efficiency Maximization, Daniel B. Rubin, Mark J. Van Der Laan
Empirical Efficiency Maximization, Daniel B. Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
It has long been recognized that covariate adjustment can increase precision, even when it is not strictly necessary. The phenomenon is particularly emphasized in clinical trials, whether using continuous, categorical, or censored time-to-event outcomes. Adjustment is often straightforward when a discrete covariate partitions the sample into a handful of strata, but becomes more involved when modern studies collect copious amounts of baseline information on each subject.
The dilemma helped motivate locally efficient estimation for coarsened data structures, as surveyed in the books of van der Laan and Robins (2003) and Tsiatis (2006). Here one fits a relatively small working model …
Regression Analysis Of A Disease Onset Distribution Using Diagnosis Data, Jessica G. Young, Nicholas P. Jewell, Steven J. Samuels
Regression Analysis Of A Disease Onset Distribution Using Diagnosis Data, Jessica G. Young, Nicholas P. Jewell, Steven J. Samuels
U.C. Berkeley Division of Biostatistics Working Paper Series
We consider methods for estimating the effect of a covariate on a disease onset distribution when the observed data structure consists of right-censored data on diagnosis times and current status data on onset times amongst individuals who have not yet been diagnosed. Dunson and Baird (2001) approached this problem using maximum likelihood, under the assumption that the ratio of the diagnosis and onset distributions is monotonic non-decreasing. As an alternative, we propose a two-step estimator, an extension of the approach of van der Laan, Jewell and Petersen (1997) in the single sample setting, that is computationally much simpler and requires …
Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard
Super Learner, Mark J. Van Der Laan, Eric C. Polley, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Previous articles (van der Laan and Dudoit (2003); van der Laan et al. (2006); Sinisi et al. (2007)) advertised and theoretically validated the use of cross-validation to select among many candidate estimators to compute a so called super learner which outperforms any of the given candidate estimators. The theoretical basis was provided for this super learner based on oracle results for the cross-validation selector (e.g., van der Laan and Dudoit (2003); van der Laan et al. (2006)) and in Sinisi et al. (2007). In addition, these papers contained a practical demonstration of the adaptivity of this so called super learner …
Estimating The Effect Of Vigorous Physical Activity On Mortality In The Elderly Based On Realistic Individualized Treatment And Intention-To-Treat Rules, Oliver Bembom, Mark J. Van Der Laan
Estimating The Effect Of Vigorous Physical Activity On Mortality In The Elderly Based On Realistic Individualized Treatment And Intention-To-Treat Rules, Oliver Bembom, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
The effect of vigorous physical activity on mortality in the elderly is difficult to estimate using conventional approaches to causal inference that define this effect by comparing the mortality risks corresponding to hypothetical scenarios in which all subjects in the target population engage in a given level of vigorous physical activity. A causal effect defined on the basis of such a static treatment intervention can only be identified from observed data if all subjects in the target population have a positive probability of selecting each of the candidate treatment options, an assumption that is highly unrealistic in this case since …
Analyzing Sequentially Randomized Trials Based On Causal Effect Models For Realistic Individualized Treatment Rules, Oliver Bembom, Mark J. Van Der Laan
Analyzing Sequentially Randomized Trials Based On Causal Effect Models For Realistic Individualized Treatment Rules, Oliver Bembom, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper, we argue that causal effect models for realistic individualized treatment rules represent an attractive tool for analyzing sequentially randomized trials. Unlike a number of methods proposed previously, this approach does not rely on the assumption that intermediate outcomes are discrete or that models for the distributions of these intermediate outcomes given the observed past are correctly specified. In addition, it generalizes the methodology for performing pairwise comparisons between individualized treatment rules by allowing the user to posit a marginal structural model for all candidate treatment rules simultaneously. If only a small number of candidate treatment rules are …
Covariate Adjustment In Randomized Trials With Binary Outcomes: Targeted Maximum Likelihood Estimation, Kelly L. Moore, Mark J. Van Der Laan
Covariate Adjustment In Randomized Trials With Binary Outcomes: Targeted Maximum Likelihood Estimation, Kelly L. Moore, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Covariate adjustment using linear models for continuous outcomes in randomized trials has been shown to increase efficiency and power over the unadjusted method in estimating the marginal effect of treatment. However, for binary outcomes, investigators generally rely on the unadjusted estimate as the literature indicates that covariate-adjusted estimates based on logistic regression models are less efficient. The crucial step that has been missing when adjusting for covariates is that one must integrate/average the adjusted estimate over those covariates in order to obtain the marginal effect. We apply the method of targeted maximum likelihood estimation (MLE), as presented in van der …
Targeted Maximum Likelihood Learning, Mark J. Van Der Laan, Daniel Rubin
Targeted Maximum Likelihood Learning, Mark J. Van Der Laan, Daniel Rubin
U.C. Berkeley Division of Biostatistics Working Paper Series
Suppose one observes a sample of independent and identically distributed observations from a particular data generating distribution. Suppose that one has available an estimate of the density of the data generating distribution such as a maximum likelihood estimator according to a given or data adaptively selected model. Suppose that one is concerned with estimation of a particular pathwise differentiable Euclidean parameter. A substitution estimator evaluating the parameter of the density estimator is typically too biased and might not even converge at the parametric rate: that is, the density estimator was targeted to be a good estimator of the density and …
Diagnosing Bias In The Inverse Probability Of Treatment Weighted Estimator Resulting From Violation Of Experimental Treatment Assignment, Yue Wang, Maya L. Petersen, David Bangsberg, Mark J. Van Der Laan
Diagnosing Bias In The Inverse Probability Of Treatment Weighted Estimator Resulting From Violation Of Experimental Treatment Assignment, Yue Wang, Maya L. Petersen, David Bangsberg, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Inverse probability of treatment weighting (IPTW) is frequently used to estimate the causal effects of treatments and interventions. The consistency of the IPTW estimator relies not only on the well-recognized assumption of no unmeasured confounders (Sequential Randomization Assumption or SRA), but also on the assumption of experimentation in the assignment of treatment (Experimental Treatment Assignment or ETA). In finite samples, violations in the ETA assumption can occur due simply to chance; certain treatments become rare or non-existent for certain strata of the population. Such practical violations of the ETA assumption occur frequently in real data, and can result in significant …
Extending Marginal Structural Models Through Local, Penalized, And Additive Learning, Daniel Rubin, Mark J. Van Der Laan
Extending Marginal Structural Models Through Local, Penalized, And Additive Learning, Daniel Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Marginal structural models (MSMs) allow one to form causal inferences from data, by specifying a relationship between a treatment and the marginal distribution of a corresponding counterfactual outcome. Following their introduction in Robins (1997), MSMs have typically been fit after assuming a semiparametric model, and then estimating a finite dimensional parameter. van der Laan and Dudoit (2003) proposed to instead view MSM fitting not as a task of semiparametric parameter estimation, but of nonparametric function approximation. They introduced a class of causal effect estimators based on mapping loss functions suitable for the unavailable counterfactual data to those suitable for the …
Statistical Learning Of Origin-Specific Statically Optimal Individualized Treatment Rules, Mark J. Van Der Laan, Maya L. Petersen
Statistical Learning Of Origin-Specific Statically Optimal Individualized Treatment Rules, Mark J. Van Der Laan, Maya L. Petersen
U.C. Berkeley Division of Biostatistics Working Paper Series
Consider a longitudinal observational or controlled study in which one collects chronological data over time on n randomly sampled subjects. The time-dependent process one observes on each randomly sampled subject contains time-dependent covariates, time-dependent treatment actions, and an outcome process or single final outcome of interest. A statically optimal individualized treatment rule (as introduced in van der Laan, Petersen & Joffe (2005), Petersen & van der Laan (2006)) is a (unknown) treatment rule which at any point in time conditions on a user-supplied subset of the past, computes the future static treatment regimen that maximizes a (conditional) mean future outcome …
Doubly Robust Censoring Unbiased Transformations, Daniel Rubin, Mark J. Van Der Laan
Doubly Robust Censoring Unbiased Transformations, Daniel Rubin, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We consider random design nonparametric regression when the response variable is subject to right censoring. Following the work of Fan and Gijbels (1994), a common approach to this problem is to apply what has been termed a censoring unbiased transformation to the data to obtain surrogate responses, and then enter these surrogate responses with covariate data into standard smoothing algorithms. Existing censoring unbiased transformations generally depend on either the conditional survival function of the response of interest, or that of the censoring variable. We show that a mapping introduced in another statistical context is in fact a censoring unbiased transformation …
A Method To Increase The Power Of Multiple Testing Procedures Through Sample Splitting, Daniel Rubin, Sandrine Dudoit, Mark J. Van Der Laan
A Method To Increase The Power Of Multiple Testing Procedures Through Sample Splitting, Daniel Rubin, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Consider the standard multiple testing problem where many hypotheses are to be tested, each hypothesis is associated with a test statistic, and large test statistics provide evidence against the null hypotheses. One proposal to provide probabilistic control of Type-I errors is the use of procedures ensuring that the expected number of false positives does not exceed a user-supplied threshold. Among such multiple testing procedures, we derive the ``most powerful'' method, meaning the test statistic cutoffs that maximize the expected number of true positives. Unfortunately, these optimal cutoffs depend on the true unknown data generating distribution, so could never be used …
Individualized Treatment Rules: Generating Candidate Clinical Trials, Maya L. Petersen, Steven G. Deeks, Mark J. Van Der Laan
Individualized Treatment Rules: Generating Candidate Clinical Trials, Maya L. Petersen, Steven G. Deeks, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Statistical methods have rarely been applied to learn individualized treatment rules, or rules for altering treatments over time in response to changes in individual covariates. Termed dynamic treatment regimes in the statistical literature, such individualized treatment rules are of primary importance in the practice of clinical medicine. History-Adjusted Marginal Structural Models (HA-MSM) estimate individualized treatment rules that assign, at each time point, the first action of the future static treatment plan that optimizes expected outcome given a patient's covariates. However, as we discuss here, the optimality of these rules can depend on the way in which treatment was assigned in …
Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan
Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many statistical methods exist that can be used to learn a predictor based on observed data. Examples include decision trees, neural networks, support vector regression, least angle regression, Logic Regression, and the Deletion/Substitution/Addition algorithm. The optimal algorithm for prediction will vary depending on the underlying data-generating distribution. In this article, we introduce a "super learner," a prediction algorithm that applies any set of candidate learners and uses cross-validation to select among them. Theory shows that asymptotically the super learner performs essentially as well or better than any of the candidate learners. We briefly present the theory behind the super learner, …
Empirical Bayes Approach To Controlling Familywise Error: An Application To Hiv Resistance Data, Rhoderick N. Machekano, Alan E. Hubbard
Empirical Bayes Approach To Controlling Familywise Error: An Application To Hiv Resistance Data, Rhoderick N. Machekano, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Statistical challenges arise in identifying meaningful patterns and structures from high dimensional genomic data sets. Relating HIV genotype (sequence of amino acids) to phenotypic resistance presents a typical problem. When the HIV virus is under antiretroviral drug pressure, unfavorable mutations of the target genes often lead to greatly increased resistance of the virus to drugs, including drugs the virus has not been exposed to. Identification of mutation combinations and their correlation to drug resistance is critical in guiding efficient prescription of HIV drugs. The identification of a subset of codons associated with drug resistance from a set of several hundreds …
Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan
Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
An important class of models in causal inference are the so-called marginal structural models which model the comparison between counterfactual outcome distributions corresponding with a static treatment intervention, conditional on user supplied baseline covariates, based on observing a longitudinal data structure on a sample of n independent and identically distributed experimental units. Identification of a static treatment regimen specific outcome distribution based on observational data requires beyond the so-called sequential randomization assumption that each experimental unit has positive probability of following the static treatment regimen. The latter assumption is called the experimental treatment assignment assumption (ETA) (which is parameter specific). …
A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska
A General Framework For Statistical Performance Comparison Of Evolutionary Computation Algorithms, David Shilane, Jarno Martikainen, Sandrine Dudoit, Seppo Ovaska
U.C. Berkeley Division of Biostatistics Working Paper Series
This paper proposes a statistical methodology for comparing the performance of evolutionary computation algorithms. A two-fold sampling scheme for collecting performance data is introduced, and these data are analyzed using bootstrap-based multiple hypothesis testing procedures. The proposed method is sufficiently flexible to allow the researcher to choose how performance is measured, does not rely upon distributional assumptions, and can be extended to analyze many other randomized numeric optimization routines. As a result, this approach offers a convenient, flexible, and reliable technique for comparing algorithms in a wide variety of applications.
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a general and formal statistical framework for the multiple tests of associations between known fixed features of a genome and unknown parameters of the distribution of variable features of this genome in a population of interest. The known fixed gene-annotation profiles, corresponding to the fixed features of the genome, may concern Gene Ontology (GO) annotation, pathway membership, regulation by particular transcription factors, nucleotide sequences, or protein sequences. The unknown gene-parameter profiles, corresponding to the variable features of the genome, may be, for example, regression coefficients relating genome-wide transcript levels or DNA copy numbers to possibly censored biological and …
Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith
Issues Of Processing And Multiple Testing Of Seldi-Tof Ms Proteomic Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan, Christine F. Skibola, Christine M. Hegedus, Martyn T. Smith
U.C. Berkeley Division of Biostatistics Working Paper Series
A new data filtering method for SELDI-TOF MS proteomic spectra data is described. We examined technical repeats (2 per subject) of intensity versus m/z (mass/charge) of bone marrow cell lysate for two groups of childhood leukemia patients: acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). As others have noted, the type of data processing as well as experimental variability can have a disproportionate impact on the list of "interesting" proteins (see Baggerly et al. (2004)). We propose a list of processing and multiple testing techniques to correct for 1) background drift; 2) filtering using smooth regression and cross-validated bandwidth …
Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard
Quantile-Function Based Null Distribution In Resampling Based Multiple Testing, Mark J. Van Der Laan, Alan E. Hubbard
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing a collection of null hypotheses about a data generating distribution based on a sample of independent and identically distributed observations is a fundamental and important statistical problem involving many applications. Methods based on marginal null distributions (i.e., marginal p-values) are attractive since the marginal p-values can be based on a user supplied choice of marginal null distributions and they are computationally trivial, but they, by necessity, are known to either be conservative or to rely on assumptions about the dependence structure between the test-statistics. Resampling based multiple testing (Westfall and Young, 1993) involves sampling from a joint null …
Data Adaptive Pathway Testing, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
Data Adaptive Pathway Testing, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
A majority of diseases are caused by a combination of factors, for example, composite genetic mutation profiles have been found in many cases to predict a deleterious outcome. There are several statistical techniques that have been used to analyze these types of biological data. This article implements a general strategy which uses data adaptive regression methods to build a specific pathway model, thus predicting a disease outcome by a combination of biological factors and assesses the significance of this model, or pathway, by using a permutation based null distribution. We also provide several simulation comparisons with other techniques. In addition, …
Correspondences Between Regression Models For Complex Binary Outcomes And Those For Structured Multivariate Survival Analyses, Nicholas P. Jewell
Correspondences Between Regression Models For Complex Binary Outcomes And Those For Structured Multivariate Survival Analyses, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
Doksum and Gasko [5] described a one-to-one correspondence between regression models for binary outcomes and those for continuous time survival analyses. This correspondence has been exploited heavily in the analysis of current status data (Jewell and van der Laan [11], Shiboski [18]). Here, we explore similar correspondences for complex survival models and categorical regression models for polytomous data. We include discussion of competing risks and progressive multi-state survival random variables.