Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Life Sciences (15)
- Bioinformatics (8)
- Medicine and Health Sciences (8)
- Applied Statistics (7)
- Data Science (7)
-
- Computer Sciences (6)
- Statistical Methodology (5)
- Statistical Models (5)
- Clinical Trials (4)
- Engineering (4)
- Genetics and Genomics (4)
- Mathematics (3)
- Public Health (3)
- Analysis (2)
- Animal Sciences (2)
- Applied Mathematics (2)
- Artificial Intelligence and Robotics (2)
- Epidemiology (2)
- Medical Specialties (2)
- Multivariate Analysis (2)
- Oncology (2)
- Other Computer Sciences (2)
- Plant Sciences (2)
- Agriculture (1)
- Agronomy and Crop Sciences (1)
- Aquaculture and Fisheries (1)
- Atomic, Molecular and Optical Physics (1)
- Institution
-
- COBRA (4)
- University of Louisville (2)
- University of Nebraska - Lincoln (2)
- University of South Carolina (2)
- California Polytechnic State University, San Luis Obispo (1)
-
- Dartmouth College (1)
- Kennesaw State University (1)
- LSU Health New Orleans (1)
- Michigan Technological University (1)
- New Jersey Institute of Technology (1)
- Old Dominion University (1)
- Purdue University (1)
- The Texas Medical Center Library (1)
- University of Kentucky (1)
- University of Montana (1)
- University of Nevada, Las Vegas (1)
- University of New Hampshire (1)
- University of New Mexico (1)
- University of South Dakota (1)
- University of South Florida (1)
- University of Texas at El Paso (1)
- Virginia Commonwealth University (1)
- Publication Year
- Publication
-
- U.C. Berkeley Division of Biostatistics Working Paper Series (3)
- Electronic Theses and Dissertations (2)
- Theses and Dissertations (2)
- Biostatistics Faculty Publications (1)
- Dartmouth College Ph.D Dissertations (1)
-
- Department of Agricultural and Biological Systems Engineering: Faculty Publications (1)
- Department of Statistics: Dissertations, Theses, and Student Research (1)
- Dissertations (1)
- Dissertations and Theses (1)
- Dissertations and Theses (Open Access) (1)
- Dissertations, Master's Theses and Master's Reports (1)
- Electrical & Computer Engineering Theses & Dissertations (1)
- Faculty Articles (1)
- Faculty Publications (1)
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Honors Theses and Capstones (1)
- Master's Theses (1)
- Mathematics & Statistics ETDs (1)
- Open Access Dissertations (1)
- Open Access Theses & Dissertations (1)
- School of Public Health Faculty Publications (1)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (1)
- USF Tampa Graduate Theses and Dissertations (1)
- UW Biostatistics Working Paper Series (1)
- Publication Type
Articles 1 - 28 of 28
Full-Text Articles in Biostatistics
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Master's Theses
Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.
Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.
We use time splitting and Mel-frequency cepstrum …
Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal
Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal
Faculty Articles
Efficacy testing is a cornerstone of clinical trials, ensuring that medical interventions achieve their intended therapeutic effects. Over the decades, a wide range of statistical methodologies have been developed to address the complexities of clinical trial data, including parametric, nonparametric, Bayesian, and machine learning approaches. Parametric methods, such as t-tests, ANOVA, and LMMs, have traditionally been the foundation of efficacy testing due to their efficiency under well-defined assumptions. Nonparametric techniques, including the Friedman test, Brunner-Munzel test, and modern extensions like nparLD, have emerged as robust alternatives, particularly for skewed, ordinal, or non-normal data. Bayesian methodologies have enabled the incorporation of …
Predicting Sleep And Sleep Stage In Children Using Actigraphy And Heartrate Via A Long Short-Term Memory Deep Learning Algorithm: A Performance Evaluation, Robert Weaver Med, Phd, James White, Olivia Finnegan, Hongpeng Yang, Zifei Zhong, Keagan Kiely, Catherine Jones, Yan Tong, Srihari Nelakuditi, Rahul Ghosal, David E. Brown, Russell R. Pate Ph.D., Gregory J. Welk, Massimiliano De Zambotti, Yuan Wang, Sarah Burkart, Elizabeth L. Adams Phd, Bridget Armstrong, Michael Beets Med, Mph, Phd
Predicting Sleep And Sleep Stage In Children Using Actigraphy And Heartrate Via A Long Short-Term Memory Deep Learning Algorithm: A Performance Evaluation, Robert Weaver Med, Phd, James White, Olivia Finnegan, Hongpeng Yang, Zifei Zhong, Keagan Kiely, Catherine Jones, Yan Tong, Srihari Nelakuditi, Rahul Ghosal, David E. Brown, Russell R. Pate Ph.D., Gregory J. Welk, Massimiliano De Zambotti, Yuan Wang, Sarah Burkart, Elizabeth L. Adams Phd, Bridget Armstrong, Michael Beets Med, Mph, Phd
Faculty Publications
Children's ambulatory sleep is commonly measured via actigraphy. However, traditional actigraphy measured sleep (e.g., Sadeh algorithm) struggles to predict wake (i.e., specificity, values typically < 70) and cannot predict sleep stages. Long short-term memory (LSTM) is a machine learning algorithm that may address these deficiencies. This study evaluated the agreement of LSTM sleep estimates from actigraphy and heartrate (HR) data with polysomnography (PSG). Children (N = 238, 5–12 years,52.8% male, 50% Black 31.9% White) participated in an overnight laboratory polysomnography. Participants were referred be-cause of suspected sleep disruptions. Children wore an ActiGraph GT9X accelerometer and two of three consumer wearables(i.e., Apple Watch Series 7, Fitbit Sense, Garmin Vivoactive 4) on their non-dominant wrist during the polysomnogram. LSTM estimated sleep versus wake and sleep stage (wake, not-REM, REM) using raw actigraphy and HR data for each 30-s epoch. Logistic regression and random forest were also estimated as a benchmark for performance with which to compare the LSTM results. A 10-fold cross-validation technique was employed, and confusion matrices were constructed. Sensitivity and specificity were calculated to assess the agreement between research-grade and consumer wearables with the criterion polysomnography. For sleep versus wake classification, LSTM outperformed logistic regression and random forest with accuracy ranging from 94.1to 95.1, sensitivity ranging from 94.9 to 95.9 across different devices, and specificity ranging from 84.5 to 89.6. The addition of HR improved the prediction of sleep stages but not binary sleep versus wake. LSTM is promising for predicting sleep and sleep staging from actigraphy data, and HR may improve sleep stage prediction.
Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng
Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng
Theses and Dissertations
Lesion-symptom mapping (LSM) studies offer insight into the brain areas involved in various aspects of cognition. This is commonly done via behavioral testing in patients with a naturally occurring brain injury or lesions (e.g., strokes or brain tumors). This results in high-dimensional observational data where lesion status (present/absent) is non-uniformly distributed, with some voxels having lesions in very few (or no) subjects. In this situation, mass univariate hypothesis tests have severe power heterogeneity where many tests are known a priori to have little to no power. Additionally, high-dimensional observational data can be grouped according to brain anatomical structure.
In this …
Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul
Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul
School of Public Health Faculty Publications
Diabetes is a growing global health concern, affecting millions and leading to severe complications if not properly managed. The primary challenge in diabetes management is maintaining blood glucose levels (BGLs) within a safe range to prevent complications such as renal failure, cardiovascular disease, and neuropathy. Traditional methods, such as finger-prick testing, often result in low patient adherence due to discomfort, invasiveness, and inconvenience. Consequently, there is an increasing need for non-invasive techniques that provide accurate BGL measurements. Photoplethysmography (PPG), a photosensitive method that detects blood volume variations, has shown promise for non-invasive glucose monitoring. Deep neural networks (DNNs) applied to …
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Dissertations, Master's Theses and Master's Reports
Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.
In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …
Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du
Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du
Dissertations
While immune therapies achieve remarkable success in treating various cancers, only a subset of patients achieves a durable clinical response, and many exhibit innate or acquired resistance. Precision medicine aims to tailor treatments to individual patients based on specific biological markers, ensuring that each patient receives the therapy most likely to be effective. Predictive biomarkers and gene signatures offer potential for more personalized treatment strategies by identifying patients likely to benefit. Recent studies suggest that gene signatures, comprising sets of genes, hold predictive value for certain clinical variables. Typically derived from biological expert knowledge, these signatures demonstrate substantial predictive potential, …
Time Series Models For Predicting Application Gpu Utilization And Power Draw Based On Trace Data, Dorothy Xiaoshuang Parry
Time Series Models For Predicting Application Gpu Utilization And Power Draw Based On Trace Data, Dorothy Xiaoshuang Parry
Electrical & Computer Engineering Theses & Dissertations
This work explores collecting performance metrics and leveraging various statistical and machine learning time series predictive models on a memory-intensive application, Inception v3. Trace data collected using nvidia-smi measured GPU utilization and power draw for two runs of Inception3. Experimental results from the statistical and machine learning-based time series predictive algorithms showed that the predictions from statistical-based models were unable to capture the complex changes in the trace data. The Probabilistic TNN model provided the best results for the power draw trace, according to the test evaluation metrics. For the GPU utilization trace, the RNN models produced the most accurate …
Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang
Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang
Dissertations and Theses (Open Access)
In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.
My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the …
Statistical And Machine Learning Approaches To Describe Factors Affecting Preweaning Mortality Of Piglets, Md Towfiqur Rahman, Tami M. Brown-Brandl, Gary A. Rohrer, Sudhendu R. Sharma, Vamsi Manthena, Yeyin Shi
Statistical And Machine Learning Approaches To Describe Factors Affecting Preweaning Mortality Of Piglets, Md Towfiqur Rahman, Tami M. Brown-Brandl, Gary A. Rohrer, Sudhendu R. Sharma, Vamsi Manthena, Yeyin Shi
Department of Agricultural and Biological Systems Engineering: Faculty Publications
High preweaning mortality (PWM) rates for piglets are a significant concern for the worldwide pork industries, causing economic loss and well-being issues. This study focused on identifying the factors affecting PWM, overlays, and predicting PWM using historical production data with statistical and machine learning models. Data were collected from 1,982 litters from the United States Meat Animal Research Center, Nebraska, over the years 2016 to 2021. Sows were housed in a farrowing building with three rooms, each with 20 farrowing crates, and taken care of by well-trained animal caretakers. A generalized linear model was used to analyze the various sow, …
Polygenic Risk Score Development And Validation For Early Detection And Risk Stratification Of Rheumatoid Arthritis And Osteoarthritis In Postmenopausal Women, Yingke Xu
UNLV Theses, Dissertations, Professional Papers, and Capstones
Introduction: Around one in four adults worldwide suffer from arthritis. There are more than one hundred different forms of arthritis; the two most common forms of arthritis are rheumatoid arthritis (RA) and osteoarthritis (OA). RA is an autoimmune disease that can cause joint inflammation. Around 1.3 million adults in the US suffer from RA, representing 0.6%–1% of the population. The RA diagnosis in its early stages is difficult since its signs and symptoms are similar to other arthritis. OA is the most common form of arthritis. In the US, around 30.8 million people are affected by this disease. However, OA …
Predicting Hiv Vaccine-Mediated Protection Level To Identify Immune Correlates Using Positive Unlabeled Learning, Shiwei Xu
Dartmouth College Ph.D Dissertations
The development of a vaccine for Human Immunodeficiency Virus type 1 (HIV-1) is a crucial step in preventing the global spread of AIDS. To ensure the progress and effectiveness of this vaccine, it is essential to establish efficient biotechnology platforms and data mining methods. These methods would help identify immune characteristics that distinguish individuals with varying levels of vaccine-induced protection and determine the underlying factors that contribute to protection against HIV acquisition. While previous studies have focused on identifying immune markers associated with infection outcomes among vaccinated patients, it is important to acknowledge the limitations of traditional case-control analytical protocols …
Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi
Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi
Mathematics & Statistics ETDs
The piezoelectric response has been a measure of interest in density functional theory (DFT) for micro-electromechanical systems (MEMS) since the inception of MEMS technology. Piezoelectric-based MEMS devices find wide applications in automobiles, mobile phones, healthcare devices, and silicon chips for computers, to name a few. Piezoelectric properties of doped aluminum nitride (AlN) have been under investigation in materials science for piezoelectric thin films because of its wide range of device applicability. In this research using rigorous DFT calculations, high throughput ab-initio simulations for 23 AlN alloys are generated.
This research is the first to report strong enhancements of piezoelectric properties …
Using Fine-Scale Aquatic Habitat Data To Construct Dreissenid Sdms In The Laurentian Great Lakes, Grace C. Henderson
Using Fine-Scale Aquatic Habitat Data To Construct Dreissenid Sdms In The Laurentian Great Lakes, Grace C. Henderson
USF Tampa Graduate Theses and Dissertations
The invasion of the Laurentian Great Lakes by aquatic invasive species (AIS) has been the subject of investigation for decades, due to their dramatic alterations to the ecosystem and high economic costs. Two AIS with the largest impacts are dreissenid zebra and quagga mussels, and though these species have been studied extensively, questions remain about what factors control their distributions, and whether lake warming will alter these distributions. Species distribution models (SDMs) offer a powerful tool to examine the relationship between species presences and environmental variables, which are typically bioclimactic data. The creation of the Aquatic Habitat (AqHab) dataset containing …
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Honors Theses and Capstones
COVID-19 caused state and nation-wide lockdowns, which altered human foot traffic, especially in restaurants. The seafood sector in particular suffered greatly as there was an increase in illegal fishing, it is made up of perishable goods, it is seasonal in some places, and imports and exports were slowed. Foot traffic data is useful for business owners to have to know how much to order, how many employees to schedule, etc. One issue is that the data is very expensive, hard to get, and not available until months after it is recorded. Our goal is to not only find covariates that …
Framework For The Evaluation Of Perturbations In The Systems Biology Landscape And Inter-Sample Similarity From Transcriptomic Datasets — A Digital Twin Perspective, Mariah Marie Hoffman
Framework For The Evaluation Of Perturbations In The Systems Biology Landscape And Inter-Sample Similarity From Transcriptomic Datasets — A Digital Twin Perspective, Mariah Marie Hoffman
Dissertations and Theses
One approach to interrogating the complexities of human systems in their well-regulated and dysregulated states is through the use of digital twins. Digital twins are virtual representations of physical systems that are descriptive of an individual's state of health, an object fundamentally related to precision medicine. A key element for building a functional digital twin type for a disease or predicting the therapeutic efficacy of a potential treatment is harmonized, machine-parsable domain knowledge. Hypothesis-driven investigations are the gold standard for representing subsystems, but their results encompass a limited knowledge of the full biosystem. Multi-omics data is one rich source of …
Comparing Machine Learning Techniques With State-Of-The-Art Parametric Prediction Models For Predicting Soybean Traits, Susweta Ray
Department of Statistics: Dissertations, Theses, and Student Research
Soybean is a significant source of protein and oil, and also widely used as animal feed. Thus, developing lines that are superior in terms of yield, protein and oil content is important to feed the ever-growing population. As opposed to the high-cost phenotyping, genotyping is both cost and time efficient for breeders while evaluating new lines in different environments (location-year combinations) can be costly. Several Genomic prediction (GP) methods have been developed to use the marker and environment data effectively to predict the yield or other relevant phenotypic traits of crops. Our study compares a conventional GP method (GBLUP), a …
Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil
Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil
Open Access Theses & Dissertations
With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …
Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis
Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis
Theses and Dissertations
High-throughput chromosome conformation capture technology (Hi-C) has revealed extensive DNA looping and folding into discrete 3D domains. These include Topologically Associating Domains (TADs) and chromatin loops, the 3D domains critical for cellular processes like gene regulation and cell differentiation. The relatively low resolution of Hi-C data (regions of several kilobases in size) prevents precise mapping of domain boundaries by conventional TAD/loop-callers. However, high resolution genomic annotations associated with boundaries, such as CTCF and members of cohesin complex, suggest a computational approach for precise location of domain boundaries.
We developed preciseTAD, an optimized machine learning framework that leverages a random …
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Graduate Student Theses, Dissertations, & Professional Papers
The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …
Modified-Half-Normal Distribution And Different Methods To Estimate Average Treatment Effect., Jingchao Sun
Modified-Half-Normal Distribution And Different Methods To Estimate Average Treatment Effect., Jingchao Sun
Electronic Theses and Dissertations
This dissertation consists of three projects related to Modified-Half-Normal distribution and causal inference. In my first project, a new distribution called Modified-Half-Normal distribution was introduced. I explored a few of its distributional properties, the procedures for generating random samples based on Bayesian approaches, and the parameter estimation based on the method of moments. The second project deals with the problem of selection bias of average treatment effect (ATE) if we use the observational data. I combined the propensity score based inverse probability of treatment weighting (IPTW) method and the directed acyclic graph (DAG) to solve this problem. The third project …
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya
Electronic Theses and Dissertations
Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …
Nonparametric Variable Importance Assessment Using Machine Learning Techniques, Brian D. Williamson, Peter B. Gilbert, Noah Simon, Marco Carone
Nonparametric Variable Importance Assessment Using Machine Learning Techniques, Brian D. Williamson, Peter B. Gilbert, Noah Simon, Marco Carone
UW Biostatistics Working Paper Series
In a regression setting, it is often of interest to quantify the importance of various features in predicting the response. Commonly, the variable importance measure used is determined by the regression technique employed. For this reason, practitioners often only resort to one of a few regression techniques for which a variable importance measure is naturally defined. Unfortunately, these regression techniques are often sub-optimal for predicting response. Additionally, because the variable importance measures native to different regression techniques generally have a different interpretation, comparisons across techniques can be difficult. In this work, we study a novel variable importance measure that can …
Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun
Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun
Biostatistics Faculty Publications
In contrast to feature selection and gene set analysis, bi-level selection is a process of selecting not only important gene sets but also important genes within those gene sets. Depending on the order of selections, a bi-level selection method can be classified into three categories – forward selection, which first selects relevant gene sets followed by the selection of relevant individual genes; backward selection which takes the reversed order; and simultaneous selection, which performs the two tasks simultaneously usually with the aids of a penalized regression model. To test the existence of subtype-specific prognostic genes for non-small cell lung cancer …
Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan
Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Online estimators update a current estimate with a new incoming batch of data without having to revisit past data thereby providing streaming estimates that are scalable to big data. We develop flexible, ensemble-based online estimators of an infinite-dimensional target parameter, such as a regression function, in the setting where data are generated sequentially by a common conditional data distribution given summary measures of the past. This setting encompasses a wide range of time-series models and as special case, models for independent and identically distributed data. Our estimator considers a large library of candidate online estimators and uses online cross-validation to …
Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier
Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier
Open Access Dissertations
Increasingly, new sources of data are being incorporated into plant breeding pipelines. Enormous amounts of data from field phenomics and genotyping technologies places data mining and analysis into a completely different level that is challenging from practical and theoretical standpoints. Intelligent decision-making relies on our capability of extracting from data useful information that may help us to achieve our goals more efficiently. Many plant breeders, agronomists and geneticists perform analyses without knowing relevant underlying assumptions, strengths or pitfalls of the employed methods. The study endeavors to assess statistical learning properties and plant breeding applications of supervised and unsupervised machine learning …
Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen
Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper we present prediction and variable importance (VIM) methods for longitudinal data sets containing both continuous and binary exposures subject to missingness. We demonstrate the use of these methods for prognosis of medical outcomes of severe trauma patients, a field in which current medical practice involves rules of thumb and scoring methods that only use a few variables and ignore the dynamic and high-dimensional nature of trauma recovery. Well-principled prediction and VIM methods can thus provide a tool to make care decisions informed by the high-dimensional patient’s physiological and clinical history. Our VIM parameters can be causally interpreted …
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In binary classification problems, the area under the ROC curve (AUC), is an effective means of measuring the performance of your model. Most often, cross-validation is also used, in order to assess how the results will generalize to an independent data set. In order to evaluate the quality of an estimate for cross-validated AUC, we must obtain an estimate for its variance. For massive data sets, the process of generating a single performance estimate can be computationally expensive. Additionally, when using a complex prediction method, calculating the cross-validated AUC on even a relatively small data set can still require a …