Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Machine learning

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 153

Full-Text Articles in Statistics and Probability

Research On Mechanical Properties Of Concrete At High Temperatures Based On Machine Learning, Liu Junhua, Liu Bin, Cao Haifeng, Liu Zhiguang, Li Zhiyong Jun 2026

Research On Mechanical Properties Of Concrete At High Temperatures Based On Machine Learning, Liu Junhua, Liu Bin, Cao Haifeng, Liu Zhiguang, Li Zhiyong

Journal of China & Foreign Highway

Structural safety is directly affected by the mechanical properties of concrete at high temperatures. Firstly, based on the existing compression and tension test data of concrete at high temperatures, the Abaqus finite element software was adopted for numerical simulation reproduction, and the reliability of the simulation method was verified. Secondly, by simulating the uniaxial tension-compression and confining pressure tests of normal concrete with different strength grades under high temperatures of 20‒800 ℃, the influence rules of temperature on the compressive strength, splitting tensile strength, elastic modulus, and stress ‒ strain relationship of concrete were elucidated. Finally, based on three commonly …


Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski Jun 2026

Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski

Master's Theses

Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.

Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.

We use time splitting and Mel-frequency cepstrum …


Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi Jun 2026

Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi

Master's Theses

Unsupervised clustering algorithms today are used across a wide variety of fields such as biology, engineering, and industry in order to classify observations into groups where labels are not provided. This can provide important latent information regarding the observations within groups, as well as insight regarding the groups themselves. In order to judge the optimal number of clusters for an unsupervised clustering algorithm, many methods exist such as the Elbow Method and Silhouette Score; however, these methods come with drawbacks and are not necessarily flexible across many unsupervised methods. We present a novel clustering score framework relying on a resampling-based …


Multimodal Machine Learning For Soil Burn Severity Mapping Across California Wildfires, Sanjana Checker Jun 2026

Multimodal Machine Learning For Soil Burn Severity Mapping Across California Wildfires, Sanjana Checker

Master's Theses

Accurate mapping of soil burn severity (SBS) is critical for post-fire watershed management, erosion risk assessment, and ecological recovery planning, yet traditional field-based approaches remain costly, time-intensive, and spatially limited. This thesis presents a machine learning pipeline for wall-to-wall SBS classification across California wildfires using multi-sensor satellite imagery, terrain derivatives, and bioclimatic covariates. Field-collected SBS observations (n = 2,180) from 52 wildfires occur- ring between 2013 and 2025, sourced from the U.S. Forest Service and CAL FIRE, were used to train and evaluate multiple classification architectures within a Google Earth Engine and Google Cloud-based prediction framework. After upsampling the unburned …


A Computer Vision Approach To Analyzing Taxane Effects On Prostate Cancer Cells, Diana Elizabeth Dancea Apr 2026

A Computer Vision Approach To Analyzing Taxane Effects On Prostate Cancer Cells, Diana Elizabeth Dancea

Electronic Theses and Dissertations

Actin is a family of proteins that help create the structure of the cytoskeleton, which gives shape to the cell. In many chemotherapy treatments, researchers target actin because it controls the cell division process. Therefore, if they are able to understand the actin fibers, that may help in formulating methods to stop or slow down cancer cells from reproducing. Another important protein is PAK6, which regulates actin. In our research, a collaborative effort with Prof. Michael Lu’s lab at Florida Atlantic University, we use machine learning techniques to analyze cells which had their PAK6 protein knocked out, and compare them …


Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal Apr 2026

Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal

Faculty Articles

Efficacy testing is a cornerstone of clinical trials, ensuring that medical interventions achieve their intended therapeutic effects. Over the decades, a wide range of statistical methodologies have been developed to address the complexities of clinical trial data, including parametric, nonparametric, Bayesian, and machine learning approaches. Parametric methods, such as t-tests, ANOVA, and LMMs, have traditionally been the foundation of efficacy testing due to their efficiency under well-defined assumptions. Nonparametric techniques, including the Friedman test, Brunner-Munzel test, and modern extensions like nparLD, have emerged as robust alternatives, particularly for skewed, ordinal, or non-normal data. Bayesian methodologies have enabled the incorporation of …


Comparative Machine Learning Models For Disease Risk Prediction, Mercy Mawusi Agbley Jan 2026

Comparative Machine Learning Models For Disease Risk Prediction, Mercy Mawusi Agbley

Theses, Dissertations and Capstones

Accurate prediction of disease outcomes is crucial for improving clinical decision-making and enabling early intervention. This study compares the performance of various statistical and machine learning models for clinical risk prediction using two healthcare datasets: diabetic retinopathy and heart disease. The models assessed include Logistic Regression, LASSO, k-Nearest Neighbors (KNN), Support Vector Machines (SVM), Neural Networks, Random Forests, Gradient Boosting Machines (GBM), and a stacked ensemble model. Prior to modeling, datasets were split into train and test sets. Standardization was applied to numeric features whilst categorical features were one-hot encoded. These transformations were later applied to the test set. Principal …


Serum Biomarker Trajectory Clusters Predict Functional Outcome And Quality Of Life For Traumatic Brain Injury, Thanh Son Do, Chantal Carnes, Zhihui Yang, Firas Kobeissy, Hamad Yadikar, Gayla R. Olbricht, Olli Tenovuo, Jussi P. Posti, Ewout W. Steyerberg, Lindsay Wilson, Nicole Von Steinbüchel, Endre Czeiter, Andras Buki, David K. Menon Jan 2026

Serum Biomarker Trajectory Clusters Predict Functional Outcome And Quality Of Life For Traumatic Brain Injury, Thanh Son Do, Chantal Carnes, Zhihui Yang, Firas Kobeissy, Hamad Yadikar, Gayla R. Olbricht, Olli Tenovuo, Jussi P. Posti, Ewout W. Steyerberg, Lindsay Wilson, Nicole Von Steinbüchel, Endre Czeiter, Andras Buki, David K. Menon

Mathematics and Statistics Faculty Research & Creative Works

Serum brain-enriched biomarkers are increasingly employed in the clinical evaluation of traumatic brain injury (TBI) to assist with triage, neuroimaging decisions, and prognostication. However, the potential of temporal biomarker trajectories to inform disease monitoring and long-term outcomes remains underexplored. We aim to identify distinct biomarker trajectory (TRAJ) profiles in traumatic brain injury patients and to examine their associations with long-term clinical outcomes. The study included 373, CT-positive Intensive Care Unit (ICU) traumatic brain injury patients (256 with initial Glasgow Coma Scale 3–12) from the Collaborative European Neurotrauma Effectiveness Research in TBI (CENTER-TBI) core study who had at least two serum …


An Integrated Data-Driven Framework For Arctic Shipping: Analyzing Vessel Speed, Environmental And Ecological Factors Through Innovative Statistical Spatio-Temporal Methods, Inverse Optimization And Machine Learning, Mauli Pant Jan 2026

An Integrated Data-Driven Framework For Arctic Shipping: Analyzing Vessel Speed, Environmental And Ecological Factors Through Innovative Statistical Spatio-Temporal Methods, Inverse Optimization And Machine Learning, Mauli Pant

Theses and Dissertations

This dissertation develops an integrated data-driven framework to analyze vessel navigation and ecological risk in the United States Arctic from 2010 to 2019. As environmental change and maritime activity increase in the region, understanding how vessels respond to dynamic conditions and how those responses interact with marine ecosystems has become increasingly important. A central theme of this dissertation is the treatment of vessel speed as both an observed outcome and a decision variable reflecting trade- offs among operational, environmental, and ecological factors. The first chapter develops a predictive framework for vessel speed over ground (SOG) using Gaussian Process Boosting (GPBoost), …


Machine Learning-Based Intrusion Detection System For Iot Networks Using The Rt-Iot 2022 Dataset, Bukunmi Ebenezer Afolabi Jan 2026

Machine Learning-Based Intrusion Detection System For Iot Networks Using The Rt-Iot 2022 Dataset, Bukunmi Ebenezer Afolabi

Theses, Dissertations and Capstones

The rapid expansion of the Internet of Things (IoT) has transformed modern computing by enabling seamless connectivity among heterogeneous devices across diverse application domains. However, this increased interconnectivity has significantly enlarged the attack surface of IoT networks, exposing them to a wide range of sophisticated cyber threats. Conventional security mechanisms often lack the capability to detect emerging attacks in real time, thereby necessitating the development of intelligent Intrusion Detection Systems (IDS) capable of accurately identifying malicious network activities. This study developed and evaluated a machine learning-based intrusion detection framework for multiclass IoT attack detection using the RT-IoT2022 dataset. The dataset …


Quantitative Evaluation Of Tunnel Rock Mass Integrity Based On Mwd Technology, Zhang Kunmu, Peng Hao, Liang Ming, Han Yu, Song Guanxian Dec 2025

Quantitative Evaluation Of Tunnel Rock Mass Integrity Based On Mwd Technology, Zhang Kunmu, Peng Hao, Liang Ming, Han Yu, Song Guanxian

Journal of China & Foreign Highway

In tunnel construction,the quantitative evaluation of rock mass integrity heavily relies on information from the exposed face,and there are challenges when drilling data is used for integrity evaluation.To this end,this study introduced a novel method for quantitative evaluation of rock mass integrity during drilling,integrating numerical statistics with machine learning.A substantial dataset of digital drilling data was collected,covering three common types of rock mass integrity:relatively intact,relatively fractured,and fractured.Subsequently,a high-performance random forest model for the classification of rock mass integrity was developed through data preprocessing and hyperparameter optimization.The interpretability of the model ’s predictive results was enhanced using Shapley additive explanations (SHAP …


Is Complexity Virtuous?, Ryan Elmore, Jack Strauss Nov 2025

Is Complexity Virtuous?, Ryan Elmore, Jack Strauss

Business Information and Analytics: Faculty Scholarship

(Kelly et al., 2024) show that increasing complexity in linear models, with potentially thousands of predictors, is "virtuous". Their work contradicts the dogma of model selection, including the Principles of Parsimony and Occam's Razor. They find that when the number of predictors far exceeds the number of observations, the bias-variance trade-off breaks down, the variance declines, and the Sharpe ratio increases. In the context of ridge regression, we find that very high complexity coupled with large penalty terms (excessive shrinkage) generate forecasts that converge to a rolling window of past returns. For example, we show the past twelve-month moving average …


Trends And Predictive Modeling Of Real Estate Prices In Major Saudi Arabia Cities, Meshal S. Aldahas Sep 2025

Trends And Predictive Modeling Of Real Estate Prices In Major Saudi Arabia Cities, Meshal S. Aldahas

Theses and Dissertations

his research examines historical trends and explanatory modeling of real estate prices in major Saudi cities, with a focus on Riyadh, Jeddah, and Dammam. Using a mixed-methods approach, the study integrates quantitative data from 2010–2023, including housing and macroeconomic indicators, with qualitative insights drawn from over 320 survey responses that captured consumer sentiment on affordability, job security, and housing policies. A combination of descriptive statistics, ARIMA and Exponential Smoothing techniques was applied to detect long-term patterns, seasonal variations, and market shocks. Predictive modeling was conducted using Linear Regression, Decision Trees, and Neural Networks, with results showing that job security consistently …


Research On Autonomous Deviation Correction Of Tunnel Boring Machines And Parameters Based On Machine Learning, Zhang Jun, Li Maopeng Aug 2025

Research On Autonomous Deviation Correction Of Tunnel Boring Machines And Parameters Based On Machine Learning, Zhang Jun, Li Maopeng

Journal of China & Foreign Highway

In order to solve the problem of realizing the autonomous deviation correction of tunnel boring machines (TBMs ), a TBM deviation correction control method that integrated the random forest (RF) algorithm with the genetic algorithm (GA) was proposed based on actual engineering data.The method combined a prediction model with an optimization model,using target deviation values as input to invert and output the required TBM deviation correction parameter values,thereby further improving the automation level of TBM deviation correction.By comparing it with the actual data,the feasibility of the model was verified.The results show that the RF algorithm-based prediction model achieves an R2 …


Comparative Analysis Of Sequential And Non-Sequential Modeling Techniques For Ddos Attack Detection With Explainable Ai, Vincent Agbenyeavu Aug 2025

Comparative Analysis Of Sequential And Non-Sequential Modeling Techniques For Ddos Attack Detection With Explainable Ai, Vincent Agbenyeavu

Theses and Dissertations

Cybersecurity is known today as one of the greatest challenges of the modern era. Among the various types of cyber-attacks that threaten our security, the Distributed Denial of Service (DDoS) attack is among some of the most common, effective, and well-recognized attack strategies. Since this form of attack is meant to disrupt the availability factor covertly, it can be detrimental to the targeted machines and difficult to discover. Because of that, there have been several approaches, as well as solutions that have been devised to detect it as accurately and efficiently as possible. In this study, four sequential data modeling …


Unveiling Insights From Complexity: Advanced Computational Techniques For High-Dimensional Medical Data, Devin P. Eddington Aug 2025

Unveiling Insights From Complexity: Advanced Computational Techniques For High-Dimensional Medical Data, Devin P. Eddington

All Graduate Theses and Dissertations, Fall 2023 to Present

Healthcare generates vast amounts of data daily, from genetic profiles to hospital records, but much of it remains untapped due to its complexity. This dissertation develops new computational tools to unlock this data’s potential, aiming to improve patient care and medical research. Five projects tackle different challenges: Project 1 creates Deep MAGIC, a method to fill in missing genetic and image data accurately, vital for understanding diseases like cancer. Project 2 analyzes how the COVID-19 pandemic disrupted surgeries, finding a 27% drop and temporary complication rises in 2020, guiding future crisis planning. Projects 3 and 4 study kidney disease trials, …


Predicting Sleep And Sleep Stage In Children Using Actigraphy And Heartrate Via A Long Short-Term Memory Deep Learning Algorithm: A Performance Evaluation, Robert Weaver Med, Phd, James White, Olivia Finnegan, Hongpeng Yang, Zifei Zhong, Keagan Kiely, Catherine Jones, Yan Tong, Srihari Nelakuditi, Rahul Ghosal, David E. Brown, Russell R. Pate Ph.D., Gregory J. Welk, Massimiliano De Zambotti, Yuan Wang, Sarah Burkart, Elizabeth L. Adams Phd, Bridget Armstrong, Michael Beets Med, Mph, Phd Jul 2025

Predicting Sleep And Sleep Stage In Children Using Actigraphy And Heartrate Via A Long Short-Term Memory Deep Learning Algorithm: A Performance Evaluation, Robert Weaver Med, Phd, James White, Olivia Finnegan, Hongpeng Yang, Zifei Zhong, Keagan Kiely, Catherine Jones, Yan Tong, Srihari Nelakuditi, Rahul Ghosal, David E. Brown, Russell R. Pate Ph.D., Gregory J. Welk, Massimiliano De Zambotti, Yuan Wang, Sarah Burkart, Elizabeth L. Adams Phd, Bridget Armstrong, Michael Beets Med, Mph, Phd

Faculty Publications

Children's ambulatory sleep is commonly measured via actigraphy. However, traditional actigraphy measured sleep (e.g., Sadeh algorithm) struggles to predict wake (i.e., specificity, values typically < 70) and cannot predict sleep stages. Long short-term memory (LSTM) is a machine learning algorithm that may address these deficiencies. This study evaluated the agreement of LSTM sleep estimates from actigraphy and heartrate (HR) data with polysomnography (PSG). Children (N = 238, 5–12 years,52.8% male, 50% Black 31.9% White) participated in an overnight laboratory polysomnography. Participants were referred be-cause of suspected sleep disruptions. Children wore an ActiGraph GT9X accelerometer and two of three consumer wearables(i.e., Apple Watch Series 7, Fitbit Sense, Garmin Vivoactive 4) on their non-dominant wrist during the polysomnogram. LSTM estimated sleep versus wake and sleep stage (wake, not-REM, REM) using raw actigraphy and HR data for each 30-s epoch. Logistic regression and random forest were also estimated as a benchmark for performance with which to compare the LSTM results. A 10-fold cross-validation technique was employed, and confusion matrices were constructed. Sensitivity and specificity were calculated to assess the agreement between research-grade and consumer wearables with the criterion polysomnography. For sleep versus wake classification, LSTM outperformed logistic regression and random forest with accuracy ranging from 94.1to 95.1, sensitivity ranging from 94.9 to 95.9 across different devices, and specificity ranging from 84.5 to 89.6. The addition of HR improved the prediction of sleep stages but not binary sleep versus wake. LSTM is promising for predicting sleep and sleep staging from actigraphy data, and HR may improve sleep stage prediction.


Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng Jul 2025

Approaches To Enhancing Multiple Hypothesis Testing Methods With Side-Information, Siyu Zheng

Theses and Dissertations

Lesion-symptom mapping (LSM) studies offer insight into the brain areas involved in various aspects of cognition. This is commonly done via behavioral testing in patients with a naturally occurring brain injury or lesions (e.g., strokes or brain tumors). This results in high-dimensional observational data where lesion status (present/absent) is non-uniformly distributed, with some voxels having lesions in very few (or no) subjects. In this situation, mass univariate hypothesis tests have severe power heterogeneity where many tests are known a priori to have little to no power. Additionally, high-dimensional observational data can be grouped according to brain anatomical structure.

In this …


Shape-Based Nanoparticle Classification Using Machine Learning, Caitlin Caitlin Robertson, Hender Lopez Jun 2025

Shape-Based Nanoparticle Classification Using Machine Learning, Caitlin Caitlin Robertson, Hender Lopez

SAML-25 Workshop on Statistical and Machine Learning

The accurate classification of nanoparticles (NPs) based on their shapes is crucial for understanding their physical-chemical properties and predict their bioactivity. Nowadays, synthesis method are able to produce a broad range of shapes, such as spheres, cubes and branched NPs and commonly these NP shapes are only described qualitative. This study presents NP descriptors obtained from NPs contours extracted from electron microscopy images. Descriptors such as Fourier descriptors, aspect ratio, and compactness are then used as input for machine learning classifiers. In particular, XGBoost, Random Forest, and neural networks are explored and the their performances are compared and discussed.


Analyzing Option Chain Bid–Ask Spreads With Machine Learning, Brian Byrne, Qianru Shang Jun 2025

Analyzing Option Chain Bid–Ask Spreads With Machine Learning, Brian Byrne, Qianru Shang

SAML-25 Workshop on Statistical and Machine Learning

This paper investigates the determinants of option bid–ask spreads using machine learning techniques. We analyze a cross-sectional dataset of Apple Inc. (AAPL) call options, focusing on the relative bid–ask spread as the target variable. By comparing linear models with ensemble methods such as Random Forests and XGBoost, we find that nonlinear machine learning methods significantly outperform traditional OLS regression. The most influential factors are moneyness, implied volatility, and time to expiration, while volume and open interest have limited predictive power. Results suggest that spreads are driven by a mix of market microstructure dynamics, capital constraints, and regulatory requirements such as …


Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun Apr 2025

Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun

SMU Data Science Review

Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …


Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul Apr 2025

Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul

School of Public Health Faculty Publications

Diabetes is a growing global health concern, affecting millions and leading to severe complications if not properly managed. The primary challenge in diabetes management is maintaining blood glucose levels (BGLs) within a safe range to prevent complications such as renal failure, cardiovascular disease, and neuropathy. Traditional methods, such as finger-prick testing, often result in low patient adherence due to discomfort, invasiveness, and inconvenience. Consequently, there is an increasing need for non-invasive techniques that provide accurate BGL measurements. Photoplethysmography (PPG), a photosensitive method that detects blood volume variations, has shown promise for non-invasive glucose monitoring. Deep neural networks (DNNs) applied to …


On Large Language Models In National Security Applications, William N. Caballero, Phillip R. Jenkins Mar 2025

On Large Language Models In National Security Applications, William N. Caballero, Phillip R. Jenkins

Faculty Publications

The overwhelming success of GPT-4 in early 2023 highlighted the transformative potential of large language models (LLMs) across various sectors, including national security. This article explores the implications of LLM integration within national security contexts, analyzing their potential to revolutionize information processing, decision-making, and operational efficiency. Whereas LLMs offer substantial benefits, such as automating tasks and enhancing data analysis, they also pose significant risks, including hallucinations, data privacy concerns, and vulnerability to adversarial attacks. Through their coupling with decision-theoretic principles and Bayesian reasoning, LLMs can significantly improve decision-making processes within national security organizations. Namely, LLMs can facilitate the transition from …


Supplementary Files For: "Structure Identification For High-Dimensional Data In The Vicinity Of Bear Lake", Ben Shaw, Haley Burger, Brennan Bean, Kevin Moon Jan 2025

Supplementary Files For: "Structure Identification For High-Dimensional Data In The Vicinity Of Bear Lake", Ben Shaw, Haley Burger, Brennan Bean, Kevin Moon

Browse all Datasets

This report focuses on seven water quality measurements taken at 43 different depths on the Bear Lake for the months of June - November in the years 2018 - 2023. These measurements create a high-dimensional dataset on which we apply state-of-the-art machine learning (ML) techniques to look for low-dimensional structure in the data. A similar effort was made for weather measurements taken near the lake. Our analysis revealed that water quality measurements tend to cluster (i.e., group together) by year, while weather measurements tend to cluster by time of the year. This suggests that the structure observed in the water …


Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah Jan 2025

Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah

Dissertations, Master's Theses and Master's Reports

Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.

In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …


Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi Jan 2025

Machine Learning Models Leveraging Patient-Similarity And Clinical Temporality For Disease Prognoses, Ahmad F. Al Musawi

Theses and Dissertations

Electronic Health Records (EHRs) constitute a comprehensive and high-dimensional repository of clinical data, encompassing a wide array of patient-level information such as diagnoses, procedures, medications, laboratory results, and unstructured clinical narratives. These data hold immense potential for advancing predictive modeling in healthcare, including tasks such as disease progression modeling, hospital readmission prediction, and length of stay (LoS) estimation. However, the intrinsic complexity of EHR data—manifested in its heterogeneity, sparsity, and temporal dynamics—poses significant analytical challenges that limit the generalizability and interpretability of conventional machine learning models. Recent methodological advancements in deep learning and graph-based learning, particularly Graph Neural Networks (GNNs), …


Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du Dec 2024

Machine Learning Methods For Pattern Recognition Analysis Of Genomic And Molecular Data, Kuang Du

Dissertations

While immune therapies achieve remarkable success in treating various cancers, only a subset of patients achieves a durable clinical response, and many exhibit innate or acquired resistance. Precision medicine aims to tailor treatments to individual patients based on specific biological markers, ensuring that each patient receives the therapy most likely to be effective. Predictive biomarkers and gene signatures offer potential for more personalized treatment strategies by identifying patients likely to benefit. Recent studies suggest that gene signatures, comprising sets of genes, hold predictive value for certain clinical variables. Typically derived from biological expert knowledge, these signatures demonstrate substantial predictive potential, …


Optimizing Transport Predictive Modeling With Simulation-Based Statistical Inference Authors, Mamunur Rashid, Quyen Tran Dec 2024

Optimizing Transport Predictive Modeling With Simulation-Based Statistical Inference Authors, Mamunur Rashid, Quyen Tran

Mathematics Faculty Publications

Simulation-based statistical inference (SBI) leverages computer simulations to help scientists understand and analyze complex data. This paper explores how SBI techniques can be used to analyze transportation data. We use modern computational methods, including machine learning models, to improve the accuracy of predictions and decision-making in transportation planning. Our study focuses on two SBI methods, Approximate Bayesian Computation - Markov Chain Monte Carlo and Synthetic Likelihood, to create synthetic data for training machine learning models. These models show the potential of SBI to handle uncertain transportation data. It also highlights the practical benefits of SBI in making better decisions for …


Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni Dec 2024

Decoding Neural Networks: An Information-Theoretic Guide To Interpretability, Error Analysis And Efficiency, Mackenzie J. Meni

Theses and Dissertations

This dissertation addresses critical challenges in neural network design by leveraging entropy-based techniques to improve model efficiency, interpretability, and bias reduction. Focusing on the unique demands of computer vision applications, particularly object detection and classification for real-time systems, this work introduces a series of innovative methods centered on information theory. At the core of these methods is the Probabilistic Explanations of Entropic Knowledge (PEEK) framework, a tool developed to analyze and visualize entropy distributions across feature maps. PEEK offers insights into information flow within neural networks, making it possible to pinpoint layers that contribute meaningfully to decision-making or identify those …


Pooling And Winsorizing Machine Learning Forecasts To Predict Stock Returns With High-Dimensional Data, Erik Mekelburg, Jack Strauss Sep 2024

Pooling And Winsorizing Machine Learning Forecasts To Predict Stock Returns With High-Dimensional Data, Erik Mekelburg, Jack Strauss

Finance: Faculty Scholarship

We evaluate US market return predictability using a novel data set of several hundred ag- gregated firm-level characteristics. We apply LASSO, Elastic Net, Random Forest, Neural Net, Extreme Gradient Boosting, and Light Gradient Boosting Machine methods and find these models experience large prediction errors that lead to forecast failures. However, winsorizing and pooling machine learning model forecasts provides consistent out-of-sample predictability. To assess robustness, we apply machine learning methods to high-dimensional data for Canada, China, Germany and the UK as well as the Goyal-Welch data. All machine learning models we consider, except for the ensemble pooled methods, fail to significantly …