Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Machine learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 28 of 28

Full-Text Articles in Biostatistics

Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski Jun 2026

Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski

Master's Theses

Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.

Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.

We use time splitting and Mel-frequency cepstrum …


Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal Apr 2026

Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal

Faculty Articles

Efficacy testing is a cornerstone of clinical trials, ensuring that medical interventions achieve their intended therapeutic effects. Over the decades, a wide range of statistical methodologies have been developed to address the complexities of clinical trial data, including parametric, nonparametric, Bayesian, and machine learning approaches. Parametric methods, such as t-tests, ANOVA, and LMMs, have traditionally been the foundation of efficacy testing due to their efficiency under well-defined assumptions. Nonparametric techniques, including the Friedman test, Brunner-Munzel test, and modern extensions like nparLD, have emerged as robust alternatives, particularly for skewed, ordinal, or non-normal data. Bayesian methodologies have enabled the incorporation of …


Predicting Sleep And Sleep Stage In Children Using Actigraphy And Heartrate Via A Long Short-Term Memory Deep Learning Algorithm: A Performance Evaluation, Robert Weaver Med, Phd, James White, Olivia Finnegan, Hongpeng Yang, Zifei Zhong, Keagan Kiely, Catherine Jones, Yan Tong, Srihari Nelakuditi, Rahul Ghosal, David E. Brown, Russell R. Pate Ph.D., Gregory J. Welk, Massimiliano De Zambotti, Yuan Wang, Sarah Burkart, Elizabeth L. Adams Phd, Bridget Armstrong, Michael Beets Med, Mph, Phd Jul 2025

Predicting Sleep And Sleep Stage In Children Using Actigraphy And Heartrate Via A Long Short-Term Memory Deep Learning Algorithm: A Performance Evaluation, Robert Weaver Med, Phd, James White, Olivia Finnegan, Hongpeng Yang, Zifei Zhong, Keagan Kiely, Catherine Jones, Yan Tong, Srihari Nelakuditi, Rahul Ghosal, David E. Brown, Russell R. Pate Ph.D., Gregory J. Welk, Massimiliano De Zambotti, Yuan Wang, Sarah Burkart, Elizabeth L. Adams Phd, Bridget Armstrong, Michael Beets Med, Mph, Phd

Faculty Publications

Children's ambulatory sleep is commonly measured via actigraphy. However, traditional actigraphy measured sleep (e.g., Sadeh algorithm) struggles to predict wake (i.e., specificity, values typically < 70) and cannot predict sleep stages. Long short-term memory (LSTM) is a machine learning algorithm that may address these deficiencies. This study evaluated the agreement of LSTM sleep estimates from actigraphy and heartrate (HR) data with polysomnography (PSG). Children (N = 238, 5–12 years,52.8% male, 50% Black 31.9% White) participated in an overnight laboratory polysomnography. Participants were referred be-cause of suspected sleep disruptions. Children wore an ActiGraph GT9X accelerometer and two of three consumer wearables(i.e., Apple Watch Series 7, Fitbit Sense, Garmin Vivoactive 4) on their non-dominant wrist during the polysomnogram. LSTM estimated sleep versus wake and sleep stage (wake, not-REM, REM) using raw actigraphy and HR data for each 30-s epoch. Logistic regression and random forest were also estimated as a benchmark for performance with which to compare the LSTM results. A 10-fold cross-validation technique was employed, and confusion matrices were constructed. Sensitivity and specificity were calculated to assess the agreement between research-grade and consumer wearables with the criterion polysomnography. For sleep versus wake classification, LSTM outperformed logistic regression and random forest with accuracy ranging from 94.1to 95.1, sensitivity ranging from 94.9 to 95.9 across different devices, and specificity ranging from 84.5 to 89.6. The addition of HR improved the prediction of sleep stages but not binary sleep versus wake. LSTM is promising for predicting sleep and sleep staging from actigraphy data, and HR may improve sleep stage prediction.


Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng Jul 2025

Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng

Theses and Dissertations

Lesion-symptom mapping (LSM) studies offer insight into the brain areas involved in various aspects of cognition. This is commonly done via behavioral testing in patients with a naturally occurring brain injury or lesions (e.g., strokes or brain tumors). This results in high-dimensional observational data where lesion status (present/absent) is non-uniformly distributed, with some voxels having lesions in very few (or no) subjects. In this situation, mass univariate hypothesis tests have severe power heterogeneity where many tests are known a priori to have little to no power. Additionally, high-dimensional observational data can be grouped according to brain anatomical structure.

In this …


Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul Apr 2025

Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul

School of Public Health Faculty Publications

Diabetes is a growing global health concern, affecting millions and leading to severe complications if not properly managed. The primary challenge in diabetes management is maintaining blood glucose levels (BGLs) within a safe range to prevent complications such as renal failure, cardiovascular disease, and neuropathy. Traditional methods, such as finger-prick testing, often result in low patient adherence due to discomfort, invasiveness, and inconvenience. Consequently, there is an increasing need for non-invasive techniques that provide accurate BGL measurements. Photoplethysmography (PPG), a photosensitive method that detects blood volume variations, has shown promise for non-invasive glucose monitoring. Deep neural networks (DNNs) applied to …


Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah Jan 2025

Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah

Dissertations, Master's Theses and Master's Reports

Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.

In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …


Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du Dec 2024

Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du

Dissertations

While immune therapies achieve remarkable success in treating various cancers, only a subset of patients achieves a durable clinical response, and many exhibit innate or acquired resistance. Precision medicine aims to tailor treatments to individual patients based on specific biological markers, ensuring that each patient receives the therapy most likely to be effective. Predictive biomarkers and gene signatures offer potential for more personalized treatment strategies by identifying patients likely to benefit. Recent studies suggest that gene signatures, comprising sets of genes, hold predictive value for certain clinical variables. Typically derived from biological expert knowledge, these signatures demonstrate substantial predictive potential, …


Time Series Models For Predicting Application Gpu Utilization And Power Draw Based On Trace Data, Dorothy Xiaoshuang Parry Apr 2024

Time Series Models For Predicting Application Gpu Utilization And Power Draw Based On Trace Data, Dorothy Xiaoshuang Parry

Electrical & Computer Engineering Theses & Dissertations

This work explores collecting performance metrics and leveraging various statistical and machine learning time series predictive models on a memory-intensive application, Inception v3. Trace data collected using nvidia-smi measured GPU utilization and power draw for two runs of Inception3. Experimental results from the statistical and machine learning-based time series predictive algorithms showed that the predictions from statistical-based models were unable to capture the complex changes in the trace data. The Probabilistic TNN model provided the best results for the power draw trace, according to the test evaluation metrics. For the GPU utilization trace, the RNN models produced the most accurate …


Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang Dec 2023

Addressing The Analytical And Computational Challenges Using Machine Learning In Biomedical Research, Yizhuo Wang

Dissertations and Theses (Open Access)

In the contemporary healthcare field, professionals are confronted with an ever-growing volume of clinical data stored in electronic health records, alongside the genomic data stemming from laboratory experiments. As a response to this deluge of data, the application of machine learning (ML) techniques is gaining popularity since ML techniques have demonstrated an exceptional proficiency in processing big data and deciphering complex nonlinear patterns that are intrinsic to biomedical research.

My research leverages ML's capabilities to address the computational challenges spanning diverse areas, including adaptive clinical trial designs, survival analysis, and high-dimensional genetic data analysis. Specifically, Chapter 2 focused on the …


Statistical And Machine Learning Approaches To Describe Factors Affecting Preweaning Mortality Of Piglets, Md Towfiqur Rahman, Tami M. Brown-Brandl, Gary A. Rohrer, Sudhendu R. Sharma, Vamsi Manthena, Yeyin Shi Oct 2023

Statistical And Machine Learning Approaches To Describe Factors Affecting Preweaning Mortality Of Piglets, Md Towfiqur Rahman, Tami M. Brown-Brandl, Gary A. Rohrer, Sudhendu R. Sharma, Vamsi Manthena, Yeyin Shi

Department of Agricultural and Biological Systems Engineering: Faculty Publications

High preweaning mortality (PWM) rates for piglets are a significant concern for the worldwide pork industries, causing economic loss and well-being issues. This study focused on identifying the factors affecting PWM, overlays, and predicting PWM using historical production data with statistical and machine learning models. Data were collected from 1,982 litters from the United States Meat Animal Research Center, Nebraska, over the years 2016 to 2021. Sows were housed in a farrowing building with three rooms, each with 20 farrowing crates, and taken care of by well-trained animal caretakers. A generalized linear model was used to analyze the various sow, …


Polygenic Risk Score Development And Validation For Early Detection And Risk Stratification Of Rheumatoid Arthritis And Osteoarthritis In Postmenopausal Women, Yingke Xu Aug 2023

Polygenic Risk Score Development And Validation For Early Detection And Risk Stratification Of Rheumatoid Arthritis And Osteoarthritis In Postmenopausal Women, Yingke Xu

UNLV Theses, Dissertations, Professional Papers, and Capstones

Introduction: Around one in four adults worldwide suffer from arthritis. There are more than one hundred different forms of arthritis; the two most common forms of arthritis are rheumatoid arthritis (RA) and osteoarthritis (OA). RA is an autoimmune disease that can cause joint inflammation. Around 1.3 million adults in the US suffer from RA, representing 0.6%–1% of the population. The RA diagnosis in its early stages is difficult since its signs and symptoms are similar to other arthritis. OA is the most common form of arthritis. In the US, around 30.8 million people are affected by this disease. However, OA …


Predicting Hiv Vaccine-Mediated Protection Level To Identify Immune Correlates Using Positive Unlabeled Learning, Shiwei Xu Jun 2023

Predicting Hiv Vaccine-Mediated Protection Level To Identify Immune Correlates Using Positive Unlabeled Learning, Shiwei Xu

Dartmouth College Ph.D Dissertations

The development of a vaccine for Human Immunodeficiency Virus type 1 (HIV-1) is a crucial step in preventing the global spread of AIDS. To ensure the progress and effectiveness of this vaccine, it is essential to establish efficient biotechnology platforms and data mining methods. These methods would help identify immune characteristics that distinguish individuals with varying levels of vaccine-induced protection and determine the underlying factors that contribute to protection against HIV acquisition. While previous studies have focused on identifying immune markers associated with infection outcomes among vaccinated patients, it is important to acknowledge the limitations of traditional case-control analytical protocols …


Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi Jun 2022

Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi

Mathematics & Statistics ETDs

The piezoelectric response has been a measure of interest in density functional theory (DFT) for micro-electromechanical systems (MEMS) since the inception of MEMS technology. Piezoelectric-based MEMS devices find wide applications in automobiles, mobile phones, healthcare devices, and silicon chips for computers, to name a few. Piezoelectric properties of doped aluminum nitride (AlN) have been under investigation in materials science for piezoelectric thin films because of its wide range of device applicability. In this research using rigorous DFT calculations, high throughput ab-initio simulations for 23 AlN alloys are generated.

This research is the first to report strong enhancements of piezoelectric properties …


Using Fine-Scale Aquatic Habitat Data To Construct Dreissenid Sdms In The Laurentian Great Lakes, Grace C. Henderson Mar 2022

Using Fine-Scale Aquatic Habitat Data To Construct Dreissenid Sdms In The Laurentian Great Lakes, Grace C. Henderson

USF Tampa Graduate Theses and Dissertations

The invasion of the Laurentian Great Lakes by aquatic invasive species (AIS) has been the subject of investigation for decades, due to their dramatic alterations to the ecosystem and high economic costs. Two AIS with the largest impacts are dreissenid zebra and quagga mussels, and though these species have been studied extensively, questions remain about what factors control their distributions, and whether lake warming will alter these distributions. Species distribution models (SDMs) offer a powerful tool to examine the relationship between species presences and environmental variables, which are typically bioclimactic data. The creation of the Aquatic Habitat (AqHab) dataset containing …


Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu Jan 2022

Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu

Honors Theses and Capstones

COVID-19 caused state and nation-wide lockdowns, which altered human foot traffic, especially in restaurants. The seafood sector in particular suffered greatly as there was an increase in illegal fishing, it is made up of perishable goods, it is seasonal in some places, and imports and exports were slowed. Foot traffic data is useful for business owners to have to know how much to order, how many employees to schedule, etc. One issue is that the data is very expensive, hard to get, and not available until months after it is recorded. Our goal is to not only find covariates that …


Framework For The Evaluation Of Perturbations In The Systems Biology Landscape And Inter-Sample Similarity From Transcriptomic Datasets — A Digital Twin Perspective, Mariah Marie Hoffman Jan 2022

Framework For The Evaluation Of Perturbations In The Systems Biology Landscape And Inter-Sample Similarity From Transcriptomic Datasets — A Digital Twin Perspective, Mariah Marie Hoffman

Dissertations and Theses

One approach to interrogating the complexities of human systems in their well-regulated and dysregulated states is through the use of digital twins. Digital twins are virtual representations of physical systems that are descriptive of an individual's state of health, an object fundamentally related to precision medicine. A key element for building a functional digital twin type for a disease or predicting the therapeutic efficacy of a potential treatment is harmonized, machine-parsable domain knowledge. Hypothesis-driven investigations are the gold standard for representing subsystems, but their results encompass a limited knowledge of the full biosystem. Multi-omics data is one rich source of …


Comparing Machine Learning Techniques With State-Of-The-Art Parametric Prediction Models For Predicting Soybean Traits, Susweta Ray Dec 2021

Comparing Machine Learning Techniques With State-Of-The-Art Parametric Prediction Models For Predicting Soybean Traits, Susweta Ray

Department of Statistics: Dissertations, Theses, and Student Research

Soybean is a significant source of protein and oil, and also widely used as animal feed. Thus, developing lines that are superior in terms of yield, protein and oil content is important to feed the ever-growing population. As opposed to the high-cost phenotyping, genotyping is both cost and time efficient for breeders while evaluating new lines in different environments (location-year combinations) can be costly. Several Genomic prediction (GP) methods have been developed to use the marker and environment data effectively to predict the yield or other relevant phenotypic traits of crops. Our study compares a conventional GP method (GBLUP), a …


Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil May 2021

Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil

Open Access Theses & Dissertations

With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …


Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis Jan 2021

Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis

Theses and Dissertations

High-throughput chromosome conformation capture technology (Hi-C) has revealed extensive DNA looping and folding into discrete 3D domains. These include Topologically Associating Domains (TADs) and chromatin loops, the 3D domains critical for cellular processes like gene regulation and cell differentiation. The relatively low resolution of Hi-C data (regions of several kilobases in size) prevents precise mapping of domain boundaries by conventional TAD/loop-callers. However, high resolution genomic annotations associated with boundaries, such as CTCF and members of cohesin complex, suggest a computational approach for precise location of domain boundaries.

We developed preciseTAD, an optimized machine learning framework that leverages a random …


Ensemble Protein Inference Evaluation, Kyle Lee Lucke Jan 2021

Ensemble Protein Inference Evaluation, Kyle Lee Lucke

Graduate Student Theses, Dissertations, & Professional Papers

The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …


Modified-Half-Normal Distribution And Different Methods To Estimate Average Treatment Effect., Jingchao Sun Dec 2020

Modified-Half-Normal Distribution And Different Methods To Estimate Average Treatment Effect., Jingchao Sun

Electronic Theses and Dissertations

This dissertation consists of three projects related to Modified-Half-Normal distribution and causal inference. In my first project, a new distribution called Modified-Half-Normal distribution was introduced. I explored a few of its distributional properties, the procedures for generating random samples based on Bayesian approaches, and the parameter estimation based on the method of moments. The second project deals with the problem of selection bias of average treatment effect (ATE) if we use the observational data. I combined the propensity score based inverse probability of treatment weighting (IPTW) method and the directed acyclic graph (DAG) to solve this problem. The third project …


Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya May 2020

Novel Inference Methods For Generalized Linear Models Using Shrinkage Priors And Data Augmentation., Arinjita Bhattacharyya

Electronic Theses and Dissertations

Generalized linear models have broad applications in biostatistics and sociology. In a regression setup, the main target is to find a relevant set of predictors out of a large collection of covariates. Sparsity is the assumption that only a few of these covariates in a regression setup have a meaningful correlation with an outcome variate of interest. Sparsity is incorporated by regularizing the irrelevant slopes towards zero without changing the relevant predictors and keeping the resulting inferences intact. Frequentist variable selection and sparsity are addressed by popular techniques like Lasso, Elastic Net. Bayesian penalized regression can tackle the curse of …


Nonparametric Variable Importance Assessment Using Machine Learning Techniques, Brian D. Williamson, Peter B. Gilbert, Noah Simon, Marco Carone Aug 2017

Nonparametric Variable Importance Assessment Using Machine Learning Techniques, Brian D. Williamson, Peter B. Gilbert, Noah Simon, Marco Carone

UW Biostatistics Working Paper Series

In a regression setting, it is often of interest to quantify the importance of various features in predicting the response. Commonly, the variable importance measure used is determined by the regression technique employed. For this reason, practitioners often only resort to one of a few regression techniques for which a variable importance measure is naturally defined. Unfortunately, these regression techniques are often sub-optimal for predicting response. Additionally, because the variable importance measures native to different regression techniques generally have a different interpretation, comparisons across techniques can be difficult. In this work, we study a novel variable importance measure that can …


Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun Apr 2017

Identification Of Prognostic Genes And Gene Sets For Early-Stage Non-Small Cell Lung Cancer Using Bi-Level Selection Methods, Suyan Tian, Chi Wang, Howard H. Chang, Jianguo Sun

Biostatistics Faculty Publications

In contrast to feature selection and gene set analysis, bi-level selection is a process of selecting not only important gene sets but also important genes within those gene sets. Depending on the order of selections, a bi-level selection method can be classified into three categories – forward selection, which first selects relevant gene sets followed by the selection of relevant individual genes; backward selection which takes the reversed order; and simultaneous selection, which performs the two tasks simultaneously usually with the aids of a penalized regression model. To test the existence of subtype-specific prognostic genes for non-small cell lung cancer …


Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan Oct 2016

Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Online estimators update a current estimate with a new incoming batch of data without having to revisit past data thereby providing streaming estimates that are scalable to big data. We develop flexible, ensemble-based online estimators of an infinite-dimensional target parameter, such as a regression function, in the setting where data are generated sequentially by a common conditional data distribution given summary measures of the past. This setting encompasses a wide range of time-series models and as special case, models for independent and identically distributed data. Our estimator considers a large library of candidate online estimators and uses online cross-validation to …


Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier Aug 2016

Learning From Data: Plant Breeding Applications Of Machine Learning, Alencar Xavier

Open Access Dissertations

Increasingly, new sources of data are being incorporated into plant breeding pipelines. Enormous amounts of data from field phenomics and genotyping technologies places data mining and analysis into a completely different level that is challenging from practical and theoretical standpoints. Intelligent decision-making relies on our capability of extracting from data useful information that may help us to achieve our goals more efficiently. Many plant breeders, agronomists and geneticists perform analyses without knowing relevant underlying assumptions, strengths or pitfalls of the employed methods. The study endeavors to assess statistical learning properties and plant breeding applications of supervised and unsupervised machine learning …


Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen Oct 2013

Variable Importance And Prediction Methods For Longitudinal Problems With Missing Variables, Ivan Diaz, Alan E. Hubbard, Anna Decker, Mitchell Cohen

U.C. Berkeley Division of Biostatistics Working Paper Series

In this paper we present prediction and variable importance (VIM) methods for longitudinal data sets containing both continuous and binary exposures subject to missingness. We demonstrate the use of these methods for prognosis of medical outcomes of severe trauma patients, a field in which current medical practice involves rules of thumb and scoring methods that only use a few variables and ignore the dynamic and high-dimensional nature of trauma recovery. Well-principled prediction and VIM methods can thus provide a tool to make care decisions informed by the high-dimensional patient’s physiological and clinical history. Our VIM parameters can be causally interpreted …


Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan Dec 2012

Computationally Efficient Confidence Intervals For Cross-Validated Area Under The Roc Curve Estimates, Erin Ledell, Maya L. Petersen, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

In binary classification problems, the area under the ROC curve (AUC), is an effective means of measuring the performance of your model. Most often, cross-validation is also used, in order to assess how the results will generalize to an independent data set. In order to evaluate the quality of an estimate for cross-validated AUC, we must obtain an estimate for its variance. For massive data sets, the process of generating a single performance estimate can be computationally expensive. Additionally, when using a complex prediction method, calculating the cross-validated AUC on even a relatively small data set can still require a …