Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches,
2026
Kennesaw State University
Efficacy Analysis In Clinical Trials: A Comprehensive Review Of Statistical And Machine Learning Approaches, Dhrubajyoti Ghosh, Samhita Pal
Faculty Articles
Efficacy testing is a cornerstone of clinical trials, ensuring that medical interventions achieve their intended therapeutic effects. Over the decades, a wide range of statistical methodologies have been developed to address the complexities of clinical trial data, including parametric, nonparametric, Bayesian, and machine learning approaches. Parametric methods, such as t-tests, ANOVA, and LMMs, have traditionally been the foundation of efficacy testing due to their efficiency under well-defined assumptions. Nonparametric techniques, including the Friedman test, Brunner-Munzel test, and modern extensions like nparLD, have emerged as robust alternatives, particularly for skewed, ordinal, or non-normal data. Bayesian methodologies have enabled the incorporation of …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury,
2026
Belmont University
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation,
2026
Rochester Institute of Technology
Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue
Articles
We introduce Saturated Hierarchical Atomic Incremental Learning (sHAIL), a learning paradigm in which complex tasks are approached through a sequence of simpler atomic subtasks, each mastered to saturation before progression. The central mechanism is a saturation criterion that detects when learning dynamics enter a plateau region, triggering consolidation and subsequent ascent to a higher level of task complexity. We develop a theoretical framework for sHAIL and show that it naturally gives rise to \emph{staircased convergence}: alternating phases of rapid improvement and genuine plateau. Within each level, classical convergence guarantees apply under standard smoothness conditions, while the hierarchical transitions are driven …
Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning,
2026
Rochester Institute of Technology
Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental
Articles
Social insect colonies and ensemble machine learning methods represent two of the most successful examples of decentralized information processing in nature and computation respectively. Here we develop a rigorous mathematical framework demonstrating that ant colony decision-making and random forest learning are isomorphic under a common formalism of stochastic ensemble intelligence. We show that the mechanisms by which genetically identical ants achieve functional differentiation— through stochastic response to local cues and positive feedback—map precisely onto the bootstrap aggregation and random feature subsampling that decorrelate decision trees. Using tools from Bayesian inference, multi-armed bandit theory, and statistical learning theory, we prove that …
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure,
2026
Rochester Institute of Technology
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue
Articles
Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. …
Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data,
2026
Rochester Institute of Technology
Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue
Articles
Ordinal data arise ubiquitously in survey research, psychology, medicine, economics, and recommender systems, yet kernel methods for such data typically rely on either nominal encodings or arbitrary numeric codings. The former discards order information; the lat- ter imposes a fictitious metric structure. This paper develops a principled framework for kernel design on ordinal scales and introduces a new class of Semantic–Aware Ordinal Ker- nels (SAOK) that simultaneously capture ordinal order and semantic proximity between categories. We begin by formalizing order–preserving embeddings of finite chains and characterizing a broad family of chain distances that are conditionally negative definite. Through Schoen- berg …
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems,
2026
Southern Methodist University
Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh
SMU Data Science Review
Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …
Selecting Without Replacement From A Population Of Bands Of Serially Connected Objects,
2026
Rochester Institute of Technology
Selecting Without Replacement From A Population Of Bands Of Serially Connected Objects, James E. Marengo, Dominick Banasik, Joseph Voelkel, David L. Farnsworth
Articles
The sampling procedure from a finite population of objects that are serially attached into bands is described and analyzed. One object is randomly selected and removed at a time, which results in that object’s band being broken into two bands or shortened by one object. The main result gives the probability of choosing an object that is part of a band of serially connected objects of any specified size at each stage of the selection process.
Using Ensemble Disagreement To Stabilize Conformal Prediction Under Distribution Shift,
2026
California Polytechnic State University, San Luis Obispo
Using Ensemble Disagreement To Stabilize Conformal Prediction Under Distribution Shift, Patrick D. Murphy
Master's Theses
Semantic segmentation of eelgrass from drone imagery is crucial for coastal habitat monitoring, restoration, and management, as these habitats continue to see rapid changes due to climate change and human influence. However, the reliability of generalizing a deployed classification model relies on both high-accuracy segmentation as well as robust uncertainty quantification that holds up when conditions change over years or locations. Conformal prediction (CP) is a method that converts a classifier's output into prediction sets with a guaranteed average coverage level for in-distribution data. However, the “vanilla” conformal score can often under-cover in hard or out-of-distribution (OOD) regions under drift. …
Estimasi Proporsi Pekerja Anak Pulau Maluku & Papua: Pendekatan Small Area Estimation Hierarchical Bayes Distribusi Beta,
2026
Politeknik Statistika STIS
Estimasi Proporsi Pekerja Anak Pulau Maluku & Papua: Pendekatan Small Area Estimation Hierarchical Bayes Distribusi Beta, Apriani Sofiana, Fauzana Afininnas, Fachrol Mochti, Angga Prayoga, Shafira Husna, Nofita Istiana
Jurnal Ekonomi Kependudukan dan Keluarga
Pekerja anak merupakan isu krusial yang memerlukan penanganan segera untuk mendukung pencapaian target pembangunan global. Pengentasan isu ini menuntut ketersediaan data yang akurat hingga wilayah kecil guna mendukung perumusan kebijakan yang tepat sasaran. Penelitian ini bertujuan menduga proporsi pekerja anak usia 5–17 tahun di kabupaten/kota Pulau Maluku dan Papua tahun 2024 menggunakan metode Small Area Estimation (SAE) Hierarchical Bayes (HB) distribusi Beta. Lima variabel penyerta dari PODES dan regsosek dipilih melalui stepwise regression dan dieksplorasi secara spasial. Hasil pemodelan HB Beta Pulau Maluku dan Papua menunjukkan sebagian besar wilayah masih memiliki RSE tinggi. Untuk meningkatkan presisi, dilakukan klasterisasi wilayah sebelum …
Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa,
2026
University of the Free State
Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng
Journal of Aviation Technology and Engineering
This study evaluates the effectiveness of log transformation in enhancing multiple regression models used to forecast air traffic movements (ATMs) in South Africa during the COVID-19 pandemic. Using 60 monthly observations from October 2016 to September 2021, the analysis incorporates variables such as revenue, lockdown levels, COVID-19 metrics, exchange rates, gross domestic product, and population. Two models are compared: one using raw ATMs and another with log-transformed ATMs as the dependent variable.
While the untransformed model shows stronger explanatory power (R² = 0.904, adjusted R² = 0.891) compared to the log-transformed model (R² = 0.772, adjusted R² = 0.741), the …
Distribution Of New Statistics Of Parking Functions And Their Generalizations,
2026
Uppsala Universitet
Distribution Of New Statistics Of Parking Functions And Their Generalizations, Stephan Wagner, Catherine H. Yan, Mei Yin
Mathematics: Faculty Scholarship
In this paper we present new results on the enumeration of parking functions and labeled forests. We introduce new statistics on parking functions, which are then extended to labeled forests via bijective correspondences. We determine the joint distribution of two statistics on parking functions and their counterparts on labeled forests. Our results on labeled forests also serve to explain the mysterious equidistribution between two seemingly unrelated statistics in parking functions recently identified by Stanley and Yin and give an explicit bijection between the two statistics. Extensions of our techniques are discussed, including joint distribution on further refinement of these new …
Comparative Machine Learning Models For Disease Risk Prediction,
2026
Marshall University
Comparative Machine Learning Models For Disease Risk Prediction, Mercy Mawusi Agbley
Theses, Dissertations and Capstones
Accurate prediction of disease outcomes is crucial for improving clinical decision-making and enabling early intervention. This study compares the performance of various statistical and machine learning models for clinical risk prediction using two healthcare datasets: diabetic retinopathy and heart disease. The models assessed include Logistic Regression, LASSO, k-Nearest Neighbors (KNN), Support Vector Machines (SVM), Neural Networks, Random Forests, Gradient Boosting Machines (GBM), and a stacked ensemble model. Prior to modeling, datasets were split into train and test sets. Standardization was applied to numeric features whilst categorical features were one-hot encoded. These transformations were later applied to the test set. Principal …
Swimming In Uncertainty: Filling Data Gaps And Providing An Educational Platform For Beach Water Quality At Tybee Island, Georgia,
2026
Georgia Southern University
Swimming In Uncertainty: Filling Data Gaps And Providing An Educational Platform For Beach Water Quality At Tybee Island, Georgia, Lukas Roberson
College of Graduate Studies: Theses & Dissertations
@font-face {font-family:"Cambria Math"; panose-1:2 4 5 3 5 4 6 3 2 4; mso-font-charset:0; mso-generic-font-family:roman; mso-font-pitch:variable; mso-font-signature:-536870145 1107305727 0 0 415 0;}p.MsoNormal, li.MsoNormal, div.MsoNormal {mso-style-unhide:no; mso-style-qformat:yes; mso-style-parent:""; margin:0in; mso-pagination:widow-orphan; font-size:12.0pt; font-family:"Times New Roman",serif; mso-fareast-font-family:"Times New Roman";}.MsoChpDefault {mso-style-type:export-only; mso-default-props:yes; mso-font-kerning:0pt; mso-ligatures:none;}div.WordSection1 {page:WordSection1;}
Swimming in beaches water contaminated with high levels of bacteria can make you sick. Current monitoring at the public beaches on Tybee Island consists of weekly monitoring and enumeration of fecal indicator bacteria that takes 24 hours for results. If the number of bacteria exceed regulatory limits, a public health advisory is issued, and affected waters are retested until …
Serum Biomarker Trajectory Clusters Predict Functional Outcome And Quality Of Life For Traumatic Brain Injury,
2026
Missouri University of Science and Technology
Serum Biomarker Trajectory Clusters Predict Functional Outcome And Quality Of Life For Traumatic Brain Injury, Thanh Son Do, Chantal Carnes, Zhihui Yang, Firas Kobeissy, Hamad Yadikar, Gayla R. Olbricht, Olli Tenovuo, Jussi P. Posti, Ewout W. Steyerberg, Lindsay Wilson, Nicole Von Steinbüchel, Endre Czeiter, Andras Buki, David K. Menon
Mathematics and Statistics Faculty Research & Creative Works
Serum brain-enriched biomarkers are increasingly employed in the clinical evaluation of traumatic brain injury (TBI) to assist with triage, neuroimaging decisions, and prognostication. However, the potential of temporal biomarker trajectories to inform disease monitoring and long-term outcomes remains underexplored. We aim to identify distinct biomarker trajectory (TRAJ) profiles in traumatic brain injury patients and to examine their associations with long-term clinical outcomes. The study included 373, CT-positive Intensive Care Unit (ICU) traumatic brain injury patients (256 with initial Glasgow Coma Scale 3–12) from the Collaborative European Neurotrauma Effectiveness Research in TBI (CENTER-TBI) core study who had at least two serum …
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts,
2026
Claremont McKenna College
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha
CMC Senior Theses
This thesis documents the design, deployment, and forward-test evaluation of an evolutionary multi-agent algorithmic trading system on Polymarket, the largest decentralized prediction market. The system pairs a locally-hosted 72-billion-parameter language model with a gradient-boosted statistical filter and an evolutionary selection mechanism that maintains a population of approximately 500 autonomous trading agents. Each agent generates a probability estimate for an event, compares it to the prevailing market price, and trades the resulting disagreement.
The central empirical exercise estimates a panel regression of trade-level profit on the absolute disagreement between the agent's probability estimate and the market price, controlling for agent identity, …
Data-Driven Partitioning In Distributed Optimization For Networked Systems,
2026
Illinois State University
Data-Driven Partitioning In Distributed Optimization For Networked Systems, Prosper Azameti
Theses and Dissertations
The convergence behavior of distributed optimal power flow (OPF) depends strongly on how the power network is partitioned into regions. Classical graph-based methods such as METIS are widely used, but they rely mainly on static topological criteria and do not explicitly incorporate operating-point-dependent information that may affect distributed optimization performance. This thesis develops a data-driven partitioning framework for distributed OPF using graph neural networks (GNNs). Each OPF scenario is represented as a graph in which buses are nodes and transmission lines are edges. Node and edge features capture both structural and operational characteristics of the network. Partition prediction is formulated …
Teaching Effectiveness On Secondary Mathematics: Evidence From Pisa—Shanghai-China,
2026
Missouri University of Science and Technology
Teaching Effectiveness On Secondary Mathematics: Evidence From Pisa—Shanghai-China, Ting Shen
Psychological Science Faculty Research & Creative Works
Educational researchers and policymakers around the world have a strong interest in understanding the underlying reasons for the remarkable academic achievement of Chinese students in the Programme for International Student Assessment (PISA). Although teachers have a significant impact on student achievement, empirical evidence on teaching effectiveness in the Chinese education system has been scarce. This study uses the PISA 2012 Shanghai-China data and employs both multilevel models and quantile regression models to investigate effective teaching factors and their differential effects for students at different mathematics achievement levels. The results reveal the importance of cognitive activation and disciplinary climate as consistent, …
Multivariate Quantile Autoregression-Mixed Data Sampling (Mvqar-Midas) Modeling Of Cost Of Living And Supply Chain Dynamics In Canada.,
2026
Wilfrid Laurier University
Multivariate Quantile Autoregression-Mixed Data Sampling (Mvqar-Midas) Modeling Of Cost Of Living And Supply Chain Dynamics In Canada., Patrick Gbolonyo
Theses and Dissertations (Comprehensive)
In recent years, the rising cost of living as a result of persistent inflationary pressures, disruptions in the global supply chains, and changes in the macroeconomic landscape has become a critical topic of discussion. To address this, we move beyond a mean-based framework and employ a quantile regression approach. This allows the persistence of each series and the transmis- sion of shocks between the Consumer Price Index (CPI) (the total CPI which is a percentage change over the past 12 months), the Interest Rate (IR)(the target for the overnight rate), the New Housing Price Index (NHPI), and high-frequency supply chain …
Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls,
2026
Claremont McKenna College
Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls, Jason Liang
CMC Senior Theses
Polling results from traditional single-choice plurality elections are readily interpretable. Simple frequentist population parameters are estimated, including each candidate’s total support and the size of the front runner's lead. If the point estimate for the size of the front runner's lead exceeds the margin of error of the lead, we can conclude that the poll shows a statistically significant front runner. However, the interpretability of these population statistics disappears when applied to ranked-choice voting elections. Because ballots rank multiple candidates and candidates are eliminated in rounds, simple population-wide parameters are not well-defined. In RCV elections, a candidate’s ability to win …
