Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (374)
- Statistical Methodology (350)
- Social and Behavioral Sciences (254)
- Statistical Theory (216)
- Medicine and Health Sciences (197)
-
- Biostatistics (192)
- Data Science (181)
- Multivariate Analysis (162)
- Life Sciences (145)
- Computer Sciences (144)
- Longitudinal Data Analysis and Time Series (141)
- Applied Mathematics (136)
- Probability (129)
- Survival Analysis (128)
- Engineering (123)
- Categorical Data Analysis (112)
- Public Health (111)
- Mathematics (110)
- Business (108)
- Economics (93)
- Other Statistics and Probability (83)
- Artificial Intelligence and Robotics (78)
- Environmental Sciences (72)
- Design of Experiments and Sample Surveys (66)
- Epidemiology (66)
- Numerical Analysis and Computation (66)
- Genetics and Genomics (54)
- Institution
-
- COBRA (242)
- University of Kentucky (58)
- Southern Methodist University (52)
- Central Bank of Nigeria (29)
- City University of New York (CUNY) (29)
-
- Virginia Commonwealth University (29)
- University of Nebraska - Lincoln (28)
- Old Dominion University (27)
- World Maritime University (26)
- Illinois State University (25)
- University of Arkansas, Fayetteville (24)
- Georgia Southern University (22)
- Claremont Colleges (20)
- East Tennessee State University (20)
- Air Force Institute of Technology (19)
- Purdue University (19)
- University of Nevada, Las Vegas (17)
- Kennesaw State University (16)
- Technological University Dublin (16)
- The Texas Medical Center Library (16)
- The University of Akron (16)
- Embry-Riddle Aeronautical University (15)
- California Polytechnic State University, San Luis Obispo (14)
- University of Louisville (14)
- Utah State University (14)
- Clemson University (13)
- Michigan Technological University (12)
- Western Michigan University (12)
- Wright State University (12)
- Loma Linda University (11)
- Keyword
-
- Statistics (73)
- Machine Learning (32)
- Machine learning (28)
- Regression (28)
- Simulation (18)
-
- Modeling (17)
- Prediction (15)
- Bayesian (14)
- Classification (13)
- Logistic regression (13)
- Mathematics (13)
- Deep Learning (12)
- Forecasting (12)
- Survival analysis (12)
- COVID-19 (11)
- Data Science (11)
- Gene expression (11)
- Models (11)
- Causal inference (10)
- Mathematical models (10)
- R (10)
- Epidemiology (9)
- Genetics (9)
- Model selection (9)
- Psychology (9)
- Time series (9)
- Artificial Intelligence (8)
- Cross-validation (8)
- Deep learning (8)
- Linear regression (8)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (60)
- Theses and Dissertations (59)
- Electronic Theses and Dissertations (48)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (47)
- Harvard University Biostatistics Working Paper Series (43)
-
- Theses and Dissertations--Statistics (42)
- The University of Michigan Department of Biostatistics Working Paper Series (40)
- SMU Data Science Review (37)
- UW Biostatistics Working Paper Series (31)
- CBN Journal of Applied Statistics (JAS) (29)
- World Maritime University Dissertations (26)
- Annual Symposium on Biomathematics and Ecology Education and Research (21)
- College of Graduate Studies: Theses & Dissertations (21)
- Articles (17)
- Dissertations (17)
- Graduate Theses and Dissertations (17)
- CMC Senior Theses (16)
- Williams Honors College, Honors Research Projects (16)
- COBRA Preprint Series (15)
- Dissertations and Theses (Open Access) (15)
- Psychology Faculty Publications (14)
- Statistical Science Theses and Dissertations (14)
- Dissertations, Master's Theses and Master's Reports (12)
- All Dissertations (11)
- Loma Linda University Electronic Theses, Dissertations & Projects (11)
- Master's Theses (11)
- Dissertations, Theses, and Capstone Projects (10)
- LSU New Orleans Theses and Dissertations (10)
- Publications (10)
- Publications and Research (9)
- Publication Type
- File Type
Articles 1111 - 1140 of 1308
Full-Text Articles in Statistical Models
An Informative Bayesian Structural Equation Model To Assess Source-Specific Health Effects Of Air Pollution, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski
An Informative Bayesian Structural Equation Model To Assess Source-Specific Health Effects Of Air Pollution, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski
Harvard University Biostatistics Working Paper Series
No abstract provided.
Mixed Multiplicative Factor Analysis Model For Air Pollution Exposure Assessment, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski
Mixed Multiplicative Factor Analysis Model For Air Pollution Exposure Assessment, Margaret C. Nikolov, Brent A. Coull, Paul J. Catalano, John J. Godleski
Harvard University Biostatistics Working Paper Series
No abstract provided.
Relative Risk Regression In Medical Research: Models, Contrasts, Estimators, And Algorithms, Thomas Lumley, Richard Kronmal, Shuangge Ma
Relative Risk Regression In Medical Research: Models, Contrasts, Estimators, And Algorithms, Thomas Lumley, Richard Kronmal, Shuangge Ma
UW Biostatistics Working Paper Series
The relative risk or prevalence ratio is a natural and familiar summary of association between a binary outcome and an exposure or intervention. For rare events, the relative risk can be approximately estimated by logistic regression. For common events estimation is more difficult. We review proposed estimation algorithms for relative risk regression. Some of these give inconsistent estimates or invalid standard errors. We show that the methods that give correct inference can be viewed as arising from a family of quasilikelihood estimating functions for the same generalized linear model, differing in their efficiency and in their robustness to outlying values …
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
Causal Comparisons In Randomized Trials Of Two Active Treatments: The Effect Of Supervised Exercise To Promote Smoking Cessation, Jason Roy, Joseph W. Hogan
COBRA Preprint Series
In behavioral medicine trials, such as smoking cessation trials, two or more active treatments are often compared. Noncompliance by some subjects with their assigned treatment poses a challenge to the data analyst. Causal parameters of interest might include those defined by subpopulations based on their potential compliance status under each assignment, using the principal stratification framework (e.g., causal effect of new therapy compared to standard therapy among subjects that would comply with either intervention). Even if subjects in one arm do not have access to the other treatment(s), the causal effect of each treatment typically can only be identified from …
Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen
Semiparametric Bayesian Modeling Of Multivariate Average Bioequivalence, Pulak Ghosh Dr., Mithat Gonen
Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series
Bioequivalence trials are usually conducted to compare two or more formulations of a drug. Simultaneous assessment of bioequivalence on multiple endpoints is called multivariate bioequivalence. Despite the fact that some tests for multivariate bioequivalence are suggested, current practice usually involves univariate bioequivalence assessments ignoring the correlations between the endpoints such as AUC and Cmax. In this paper we develop a semiparametric Bayesian test for bioequivalence under multiple endpoints. Specifically, we show how the correlation between the endpoints can be incorporated in the analysis and how this correlation affects the inference. Resulting estimates and posterior probabilities ``borrow strength'' from one another …
Profile Likelihood Estimation Of Partially Linear Panel Data Models With Fixed Effects, Liangjun Su, Aman Ullah
Profile Likelihood Estimation Of Partially Linear Panel Data Models With Fixed Effects, Liangjun Su, Aman Ullah
Research Collection School Of Economics
We consider consistent estimation of partially linear panel data models with fixed effects. We propose profile-likelihood-based estimators for both the parametric and nonparametric components in the models and establish convergence rates and asymptotic normality for both estimators.
Semiparametric Latent Variable Regression Models For Spatio-Temporal Modeling Of Mobile Source Particles In The Greater Boston Area, Alexandros Gryparis, Brent A. Coull, Joel Schwartz, Helen H. Suh
Semiparametric Latent Variable Regression Models For Spatio-Temporal Modeling Of Mobile Source Particles In The Greater Boston Area, Alexandros Gryparis, Brent A. Coull, Joel Schwartz, Helen H. Suh
Harvard University Biostatistics Working Paper Series
Traffic particle concentrations show considerable spatial variability within a metropolitan area. We consider latent variable semiparametric regression models for modeling the spatial and temporal variability of black carbon and elemental carbon concentrations in the greater Boston area. Measurements of these pollutants, which are markers of traffic particles, were obtained from several individual exposure studies conducted at specific household locations as well as 15 ambient monitoring sites in the city. The models allow for both flexible, nonlinear effects of covariates and for unexplained spatial and temporal variability in exposure. In addition, the different individual exposure studies recorded different surrogates of traffic …
Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan
Super Learning: An Application To Prediction Of Hiv-1 Drug Susceptibility, Sandra E. Sinisi, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many statistical methods exist that can be used to learn a predictor based on observed data. Examples include decision trees, neural networks, support vector regression, least angle regression, Logic Regression, and the Deletion/Substitution/Addition algorithm. The optimal algorithm for prediction will vary depending on the underlying data-generating distribution. In this article, we introduce a "super learner," a prediction algorithm that applies any set of candidate learners and uses cross-validation to select among them. Theory shows that asymptotically the super learner performs essentially as well or better than any of the candidate learners. We briefly present the theory behind the super learner, …
Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan
Causal Effect Models For Intention To Treat And Realistic Individualized Treatment Rules, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
An important class of models in causal inference are the so-called marginal structural models which model the comparison between counterfactual outcome distributions corresponding with a static treatment intervention, conditional on user supplied baseline covariates, based on observing a longitudinal data structure on a sample of n independent and identically distributed experimental units. Identification of a static treatment regimen specific outcome distribution based on observational data requires beyond the so-called sequential randomization assumption that each experimental unit has positive probability of following the static treatment regimen. The latter assumption is called the experimental treatment assignment assumption (ETA) (which is parameter specific). …
On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger
On The Equivalence Of Case-Crossover And Time Series Methods In Environmental Epidemiology, Yun Lu, Scott L. Zeger
Johns Hopkins University, Dept. of Biostatistics Working Papers
Time series and case-crossover methods are often viewed as competing alternatives in environmental epidemiologic studies. Several recent studies have compared the time series and case-crossover methods. In this paper, we show that case-crossover using conditional logistic regression is a special case of time series analysis when there is a common exposure such as in air pollution studies. This equivalence provides computational convenience for case-crossover analyses and a better understanding of time series models. Time series log-linear regression accounts for over-dispersion of the Poisson variance, while case-crossover analyses typically do not. This equivalence also permits model checking for case-crossover data using …
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
Multiple Tests Of Association With Biological Annotation Metadata, Sandrine Dudoit, Sunduz Keles, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We propose a general and formal statistical framework for the multiple tests of associations between known fixed features of a genome and unknown parameters of the distribution of variable features of this genome in a population of interest. The known fixed gene-annotation profiles, corresponding to the fixed features of the genome, may concern Gene Ontology (GO) annotation, pathway membership, regulation by particular transcription factors, nucleotide sequences, or protein sequences. The unknown gene-parameter profiles, corresponding to the variable features of the genome, may be, for example, regression coefficients relating genome-wide transcript levels or DNA copy numbers to possibly censored biological and …
Optimization Of A Multi-Echelon Repair System Via Generalized Pattern Search With Ranking And Selection: A Computational Study, Derek D. Tharaldson
Optimization Of A Multi-Echelon Repair System Via Generalized Pattern Search With Ranking And Selection: A Computational Study, Derek D. Tharaldson
Theses and Dissertations
With increasing developments in computer technology and available software, simulation is becoming a widely used tool to model, analyze, and improve a real world system or process. However, simulation in itself is not an optimization approach. Common optimization procedures require either an explicit mathematical formulation or numerous function evaluations at improving iterative points. Mathematical formulation is generally impossible for problems where simulation is relevant, which are characteristically the types of problems that arise in practical applications. Further complicating matters is the variability in the simulation response which can cause problems in iterative techniques using the simulation model as a function …
Model Checking For Roc Regression Analysis, Tianxi Cai, Yingye Zheng
Model Checking For Roc Regression Analysis, Tianxi Cai, Yingye Zheng
Harvard University Biostatistics Working Paper Series
The Receiver Operating Characteristic (ROC) curve is a prominent tool for characterizing the accuracy of continuous diagnostic test. To account for factors that might invluence the test accuracy, various ROC regression methods have been proposed. However, as in any regression analysis, when the assumed models do not fit the data well, these methods may render invalid and misleading results. To date practical model checking techniques suitable for validating existing ROC regression models are not yet available. In this paper, we develop cumulative residual based procedures to graphically and numerically assess the goodness-of-fit for some commonly used ROC regression models, and …
On The Use Of Non-Euclidean Isotropy In Geostatistics, Frank C. Curriero
On The Use Of Non-Euclidean Isotropy In Geostatistics, Frank C. Curriero
Johns Hopkins University, Dept. of Biostatistics Working Papers
This paper investigates the use of non-Euclidean distances to characterize isotropic spatial dependence for geostatistical related applications. A simple example is provided to demonstrate there are no guarantees that existing covariogram and variogram functions remain valid (i.e.\ positive definite or conditionally negative definite) when used with a non-Euclidean distance measure. Furthermore, satisfying the conditions of a metric is not sufficient to ensure the distance measure can be used with existing functions. Current literature is not clear on these topics. There are certain distance measures that when used with existing covariogram and variogram functions remain valid, an issue that is explored. …
Casual Mediation Analyses With Structural Mean Models, Thomas R. Tenhave, Marshall Joffe, Kevin Lynch, Greg Brown, Stephen Maisto
Casual Mediation Analyses With Structural Mean Models, Thomas R. Tenhave, Marshall Joffe, Kevin Lynch, Greg Brown, Stephen Maisto
UPenn Biostatistics Working Papers
We represent a linear structural mean model (SMM)approach for analyzing mediation of a randomized baseline intervention's effect on a univariate follow-up outcome. Unlike standard mediation analyses, our approach does not assume that the mediating factor is randomly assigned to individuals (i.e., sequential ignorability). Hence, a comparison of the results of the proposed and standard approaches in with respect to mediation offers a sensitivity analyses of the sequential ignorability assumption. The G-estimation procedure for the proposed SMM represents an extension of the work on direct effects of randomized treatment effects for survival outcomes by Robins and Greenland (1994) (Section 5.0 and …
Model Evaluation Based On The Distribution Of Estimated Absolute Prediction Error, Lu Tian, Tianxi Cai, Els Goetghebeur, L. J. Wei
Model Evaluation Based On The Distribution Of Estimated Absolute Prediction Error, Lu Tian, Tianxi Cai, Els Goetghebeur, L. J. Wei
Harvard University Biostatistics Working Paper Series
The construction of a reliable, practically useful prediction rule for future response is heavily dependent on the "adequacy" of the fitted regression model. In this article, we consider the absolute prediction error, the expected value of the absolute difference between the future and predicted responses, as the model evaluation criterion. This prediction error is easier to interpret than the average squared error and is equivalent to the mis-classification error for the binary outcome. We show that the distributions of the apparent error and its cross-validation counterparts are approximately normal even under a misspecified fitted model. When the prediction rule is …
A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit
A Fine-Scale Linkage Disequilibrium Measure Based On Length Of Haplotype Sharing, Yan Wang, Lue Ping Zhao, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
High-throughput genotyping technologies for single nucleotide polymorphisms (SNP) have enabled the recent completion of the International HapMap Project (Phase I), which has stimulated much interest in studying genome-wide linkage disequilibrium (LD) patterns. Conventional LD measures, such as D' and r-square, are two-point measurements, and their relationship with physical distance is highly noisy. We propose a new LD measure, defined in terms of the correlation coefficient for shared haplotype lengths around two loci, thereby borrowing information from multiple loci. A U-statistic-based estimator of the new LD measure, which takes into consideration the dependence structure of the observed data, is developed and …
Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan
Population Intervention Models In Causal Inference, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Marginal structural models (MSM) provide a powerful tool for estimating the causal effect of a] treatment variable or risk variable on the distribution of a disease in a population. These models, as originally introduced by Robins (e.g., Robins (2000a), Robins (2000b), van der Laan and Robins (2002)), model the marginal distributions of treatment-specific counterfactual outcomes, possibly conditional on a subset of the baseline covariates, and its dependence on treatment. Marginal structural models are particularly useful in the context of longitudinal data structures, in which each subject's treatment and covariate history are measured over time, and an outcome is recorded at …
Gauss-Seidel Estimation Of Generalized Linear Mixed Models With Application To Poisson Modeling Of Spatially Varying Disease Rates, Subharup Guha, Louise Ryan
Gauss-Seidel Estimation Of Generalized Linear Mixed Models With Application To Poisson Modeling Of Spatially Varying Disease Rates, Subharup Guha, Louise Ryan
Harvard University Biostatistics Working Paper Series
Generalized linear mixed models (GLMMs) provide an elegant framework for the analysis of correlated data. Due to the non-closed form of the likelihood, GLMMs are often fit by computational procedures like penalized quasi-likelihood (PQL). Special cases of these models are generalized linear models (GLMs), which are often fit using algorithms like iterative weighted least squares (IWLS). High computational costs and memory space constraints often make it difficult to apply these iterative procedures to data sets with very large number of cases.
This paper proposes a computationally efficient strategy based on the Gauss-Seidel algorithm that iteratively fits sub-models of the GLMM …
Computational Techniques For Spatial Logistic Regression With Large Datasets, Christopher J. Paciorek, Louise Ryan
Computational Techniques For Spatial Logistic Regression With Large Datasets, Christopher J. Paciorek, Louise Ryan
Harvard University Biostatistics Working Paper Series
In epidemiological work, outcomes are frequently non-normal, sample sizes may be large, and effects are often small. To relate health outcomes to geographic risk factors, fast and powerful methods for fitting spatial models, particularly for non-normal data, are required. We focus on binary outcomes, with the risk surface a smooth function of space. We compare penalized likelihood models, including the penalized quasi-likelihood (PQL) approach, and Bayesian models based on fit, speed, and ease of implementation.
A Bayesian model using a spectral basis representation of the spatial surface provides the best tradeoff of sensitivity and specificity in simulations, detecting real spatial …
Is The Number Of Sick Persons In A Cohort Constant Over Time?, Paula Diehr, Ann Derleth, Anne Newman, Liming Cai
Is The Number Of Sick Persons In A Cohort Constant Over Time?, Paula Diehr, Ann Derleth, Anne Newman, Liming Cai
UW Biostatistics Working Paper Series
Objectives: To estimate the number of persons in a cohort who are sick, over time.
Methods: We calculated the number of sick persons in the Cardiovascular Health Study (CHS), a cohort study of older adults followed up to 14 years, using eight definitions of “healthy” and “sick”. We projected the number in each health state over time for a birth cohort.
Results: The number of sick persons in CHS was approximately constant for 14 years, for all definitions of “sick”. The estimated number of sick persons in the birth cohort was approximately constant from ages 55-75, after which it decreased. …
A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky
A Pseudolikelihood Approach For Simultaneous Analysis Of Array Comparative Genomic Hybridizations (Acgh), David A. Engler, Gayatry Mohapatra, David N. Louis, Rebecca Betensky
Harvard University Biostatistics Working Paper Series
DNA sequence copy number has been shown to be associated with cancer development and progression. Array-based Comparative Genomic Hybridization (aCGH) is a recent development that seeks to identify the copy number ratio at large numbers of markers across the genome. Due to experimental and biological variations across chromosomes and across hybridizations, current methods are limited to analyses of single chromosomes. We propose a more powerful approach that borrows strength across chromosomes and across hybridizations. We assume a Gaussian mixture model, with a hidden Markov dependence structure, and with random effects to allow for intertumoral variation, as well as intratumoral clonal …
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
A Nonstationary Negative Binomial Time Series With Time-Dependent Covariates: Enterococcus Counts In Boston Harbor, E. Andres Houseman, Brent Coull, James P. Shine
Harvard University Biostatistics Working Paper Series
Boston Harbor has had a history of poor water quality, including contamination by enteric pathogens. We conduct a statistical analysis of data collected by the Massachusetts Water Resources Authority (MWRA) between 1996 and 2002 to evaluate the effects of court-mandated improvements in sewage treatment. Motivated by the ineffectiveness of standard Poisson mixture models and their zero-inflated counterparts, we propose a new negative binomial model for time series of Enterococcus counts in Boston Harbor, where nonstationarity and autocorrelation are modeled using a nonparametric smooth function of time in the predictor. Without further restrictions, this function is not identifiable in the presence …
Cross-Validated Bagged Prediction Of Survival, Sandra E. Sinisi, Romain Neugebauer, Mark J. Van Der Laan
Cross-Validated Bagged Prediction Of Survival, Sandra E. Sinisi, Romain Neugebauer, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In this article, we show how to apply our previously proposed Deletion/Substitution/Addition algorithm in the context of right-censoring for the prediction of survival. Furthermore, we introduce how to incorporate bagging into the algorithm to obtain a cross-validated bagged estimator. The method is used for predicting the survival time of patients with diffuse large B-cell lymphoma based on gene expression variables.
Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin
Semiparametric Normal Transformation Models For Spatially Correlated Survival Data, Yi Li, Xihong Lin
Harvard University Biostatistics Working Paper Series
There is an emerging interest in modeling spatially correlated survival data in biomedical and epidemiological studies. In this paper, we propose a new class of semiparametric normal transformation models for right censored spatially correlated survival data. This class of models assumes that survival outcomes marginally follow a Cox proportional hazard model with unspecified baseline hazard, and their joint distribution is obtained by transforming survival outcomes to normal random variables, whose joint distribution is assumed to be multivariate normal with a spatial correlation structure. A key feature of the class of semiparametric normal transformation models is that it provides a rich …
Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen
Direct Effect Models, Mark J. Van Der Laan, Maya L. Petersen
U.C. Berkeley Division of Biostatistics Working Paper Series
The causal effect of a treatment on an outcome is generally mediated by several intermediate variables. Estimation of the component of the causal effect of a treatment that is mediated by a given intermediate variable (the indirect effect of the treatment), and the component that is not mediated by that intermediate variable (the direct effect of the treatment) is often relevant to mechanistic understanding and to the design of clinical and public health interventions. Under the assumption of no-unmeasured confounders for treatment and the intermediate variable, Robins & Greenland (1992) define an individual direct effect as the counterfactual effect of …
Survival Point Estimate Prediction In Matched And Non-Matched Case-Control Subsample Designed Studies, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore, Karla Kerlikowske
Survival Point Estimate Prediction In Matched And Non-Matched Case-Control Subsample Designed Studies, Annette M. Molinaro, Mark J. Van Der Laan, Dan H. Moore, Karla Kerlikowske
U.C. Berkeley Division of Biostatistics Working Paper Series
Providing information about the risk of disease and clinical factors that may increase or decrease a patient's risk of disease is standard medical practice. Although case-control studies can provide evidence of strong associations between diseases and risk factors, clinicians need to be able to communicate to patients the age-specific risks of disease over a defined time interval for a set of risk factors.
An estimate of absolute risk cannot be determined from case-control studies because cases are generally chosen from a population whose size is not known (necessary for calculation of absolute risk) and where duration of follow-up is not …
Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
Application Of A Multiple Testing Procedure Controlling The Proportion Of False Positives To Protein And Bacterial Data, Merrill D. Birkner, Alan E. Hubbard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Simultaneously testing multiple hypotheses is important in high-dimensional biological studies. In these situations, one is often interested in controlling the Type-I error rate, such as the proportion of false positives to total rejections (TPPFP) at a specific level, alpha. This article will present an application of the E-Bayes/Bootstrap TPPFP procedure, presented in van der Laan et al. (2005), which controls the tail probability of the proportion of false positives (TPPFP), on two biological datasets. The two data applications include firstly, the application to a mass-spectrometry dataset of two leukemia subtypes, AML and ALL. The protein data measurements include intensity and …
Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan
Cross-Validating And Bagging Partitioning Algorithms With Variable Importance, Annette M. Molinaro, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
We present a cross-validated bagging scheme in the context of partitioning algorithms. To explore the benefits of the various bagging scheme, we compare via simulations the predictive ability of single Classification and Regression (CART) Tree with several previously suggested bagging schemes and with our proposed approach. Additionally, a variable importance measure is explained and illustrated.
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
Test Statistics Null Distributions In Multiple Testing: Simulation Studies And Applications To Genomics, Katherine S. Pollard, Merrill D. Birkner, Mark J. Van Der Laan, Sandrine Dudoit
U.C. Berkeley Division of Biostatistics Working Paper Series
Multiple hypothesis testing problems arise frequently in biomedical and genomic research, for instance, when identifying differentially expressed or co-expressed genes in microarray experiments. We have developed generally applicable resampling-based single-step and stepwise multiple testing procedures (MTP) for control of a broad class of Type I error rates, defined as tail probabilities and expected values for arbitrary functions of the numbers of false positives and rejected hypotheses (Dudoit and van der Laan, 2005; Dudoit et al., 2004a,b; Pollard and van der Laan, 2004; van der Laan et al., 2005, 2004a,b). As argued in the early article of Pollard and van der …