Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (9)
- Public Health (7)
- Applied Statistics (5)
- Data Science (4)
- Statistical Methodology (4)
-
- Statistical Models (4)
- Categorical Data Analysis (3)
- Epidemiology (3)
- Life Sciences (3)
- Longitudinal Data Analysis and Time Series (3)
- Survival Analysis (3)
- Clinical Trials (2)
- Immunology and Infectious Disease (2)
- Multivariate Analysis (2)
- Artificial Intelligence and Robotics (1)
- Bioinformatics (1)
- Computer Sciences (1)
- Disease Modeling (1)
- Diseases (1)
- Health Services Research (1)
- Immunology of Infectious Disease (1)
- Medical Specialties (1)
- Microarrays (1)
- Oncology (1)
- Organisms (1)
- Other Public Health (1)
- Parasitology (1)
- Institution
- Publication Year
- Publication
-
- Statistical Science Theses and Dissertations (11)
- Theses and Dissertations (8)
- UPenn Biostatistics Working Papers (2)
- COBRA Preprint Series (1)
- Capstone Experience: Master of Public Health (1)
-
- Department of Medicine Faculty Papers (1)
- Dissertations and Theses (Open Access) (1)
- Electronic Theses and Dissertations (1)
- Legacy Theses & Dissertations (2009 - 2024) (1)
- Senior Honors Theses (1)
- Statistical Sciences and Operations Research Publications (1)
- Statistics (1)
- U.C. Berkeley Division of Biostatistics Working Paper Series (1)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (1)
- USF Tampa Graduate Theses and Dissertations (1)
- Publication Type
Articles 31 - 33 of 33
Full-Text Articles in Biostatistics
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In binary classification problems, the area under the ROC curve (AUC), is an effective means of measuring the performance of your model. Most often, cross-validation is also used, in order to assess how the results will generalize to an independent data set. In order to evaluate the quality of an estimate for cross-validated AUC, we must obtain an estimate for its variance. For massive data sets, the process of generating a single performance estimate can be computationally expensive. Additionally, when using a complex prediction method, calculating the cross-validated AUC on even a relatively small data set can still require a …
Hypothesis Testing And Power Calculations For Taxonomic-Based Human Microbiome Data, P. S. Larossa, J. Paul Brooks, Elena Deych, Edward L. Boone, David J. Edwards, Qin Wang, Erica Sodergren, George Weinstock, William D. Shannon
Hypothesis Testing And Power Calculations For Taxonomic-Based Human Microbiome Data, P. S. Larossa, J. Paul Brooks, Elena Deych, Edward L. Boone, David J. Edwards, Qin Wang, Erica Sodergren, George Weinstock, William D. Shannon
Statistical Sciences and Operations Research Publications
This paper presents new biostatistical methods for the analysis of microbiome data based on a fully parametric approach using all the data. The Dirichlet-multinomial distribution allows the analyst to calculate power and sample sizes for experimental design, perform tests of hypotheses (e.g., compare microbiomes across groups), and to estimate parameters describing microbiome properties. The use of a fully parametric model for these data has the benefit over alternative non-parametric approaches such as bootstrapping and permutation testing, in that this model is able to retain more information contained in the data. This paper details the statistical approaches for several tests of …
Normal Mixture Models For Gene Cluster Identification In Two Dimensional Microarray Data, Eric Scott Harvey
Normal Mixture Models For Gene Cluster Identification In Two Dimensional Microarray Data, Eric Scott Harvey
Theses and Dissertations
This dissertation focuses on methodology specific to microarray data analyses that organize the data in preliminary steps and proposes a cluster analysis method which improves the interpretability of the cluster results. Cluster analysis of microarray data allows samples with similar gene expression values to be discovered and may serve as a useful diagnostic tool. Since microarray data is inherently noisy, data preprocessing steps including smoothing and filtering are discussed. Comparing the results of different clustering methods is complicated by the arbitrariness of the cluster labels. Methods for re-labeling clusters to assess the agreement between the results of different clustering techniques …