Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Smith College (49)
- Southern Methodist University (37)
- Kennesaw State University (27)
- Central Bank of Nigeria (23)
- Old Dominion University (23)
-
- City University of New York (CUNY) (18)
- University of Central Florida (16)
- Chapman University (13)
- West Virginia University (13)
- Department of Primary Industries and Regional Development, Western Australia (12)
- Illinois State University (12)
- California Polytechnic State University, San Luis Obispo (11)
- East Tennessee State University (11)
- LSU Health New Orleans (11)
- Claremont Colleges (10)
- Georgia Southern University (10)
- University of Arkansas, Fayetteville (10)
- Embry-Riddle Aeronautical University (9)
- University of Kentucky (9)
- Virginia Commonwealth University (9)
- Rochester Institute of Technology (7)
- Clemson University (6)
- Dartmouth College (6)
- Purdue University (6)
- Binghamton University (5)
- Murray State University (5)
- The University of Akron (5)
- University of Louisville (5)
- University of New Mexico (5)
- University of South Florida (5)
- Keyword
-
- Machine learning (36)
- Machine Learning (35)
- Statistics (27)
- Deep learning (15)
- Data Science (14)
-
- Data science (13)
- Classification (12)
- Time series (11)
- COVID-19 (10)
- Deep Learning (9)
- Artificial Intelligence (8)
- Forecasting (8)
- Regression (8)
- Prediction (7)
- Western Australia (7)
- Logistic regression (6)
- Neural Network (6)
- Clustering (5)
- Data analysis (5)
- Natural language processing (5)
- Sentiment analysis (5)
- Simulation (5)
- Time Series (5)
- Analysis (4)
- Analytics (4)
- Baseball (4)
- Bayesian (4)
- Bioinformatics (4)
- Biostatistics (4)
- CNN (4)
- Publication Year
- Publication
-
- Statistical and Data Sciences: Faculty Publications (45)
- SMU Data Science Review (27)
- CBN Journal of Applied Statistics (JAS) (20)
- Symposium of Student Scholars (20)
- Electronic Theses and Dissertations (18)
-
- Theses and Dissertations (17)
- Master's Theses (11)
- Annual Symposium on Biomathematics and Ecology Education and Research (10)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (10)
- Mathematics & Statistics Faculty Publications (10)
- School of Public Health Faculty Publications (10)
- College of Graduate Studies: Theses & Dissertations (9)
- Data Science and Data Mining (9)
- Dissertations, Theses, and Capstone Projects (8)
- Articles (7)
- Computational and Data Sciences (PhD) Dissertations (7)
- Dissertations (7)
- CMC Senior Theses (6)
- All Dissertations (5)
- Honors College Theses (5)
- Northeast Journal of Complex Systems (NEJCS) (5)
- Publications and Research (5)
- Statistical Science Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- Dartmouth College Ph.D Dissertations (4)
- Dissertations, Master's Theses and Master's Reports (4)
- Doctor of Data Science and Analytics Dissertations (4)
- Electronic Theses & Dissertations (2024 - present) (4)
- Fisheries Research Articles (4)
- Honors Projects (4)
- Publication Type
- File Type
Articles 1 - 30 of 549
Full-Text Articles in Data Science
Individualized Bayesian Inference Identifies Novel Genetic Variants For Parkinson's Disease, Jin Ren, Yasaman J. Soofi, Md Asad Rahman, Qing Lu, Jinling Liu
Individualized Bayesian Inference Identifies Novel Genetic Variants For Parkinson's Disease, Jin Ren, Yasaman J. Soofi, Md Asad Rahman, Qing Lu, Jinling Liu
Engineering Management and Systems Engineering Faculty Research & Creative Works
Parkinson's disease (PD) is a complex neurodegenerative disorder with a significant genetic component. While genome-wide association studies (GWAS) have been instrumental in identifying genetic variants associated with PD, the reliance on large sample sizes and population-level analyses may overlook variants with lower minor allele frequencies or individual-specific relevance. Individualized Bayesian Inference (IBI) offers a promising method to complement GWAS by identifying and prioritizing candidate genetic markers at both the individual and patients-like-me subgroup levels. This study evaluates the application of IBI to PD genetics, using GWAS as a baseline for comparison. We analyzed genetic data from the Fox Insight online …
Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor
Predicting Remaining Useful Life Using Multivariate Time-Series Data, Anayah Smith, Victoria Gaibor
Discovery Day - Daytona Beach
Accurate prediction of Remaining Useful Life (RUL) is critical for enabling predictive maintenance, improving system reliability, and reducing operational costs in degrading systems. This project addresses the problem of modeling and predicting RUL using multivariate time-series sensor data from the NASA CMAPSS turbofan engine dataset, with a focus on understanding how predictive performance changes across datasets of varying complexity. The objective is to develop a reproducible machine learning pipeline that captures degradation patterns and produces reliable time-to-failure predictions. The approach includes data preprocessing, exploratory data analysis, feature engineering, dimensionality reduction, and model evaluation. RUL values are computed and capped to …
Unifying And Expanding Global And Local Variable Importance Methods For Explainable Machine Learning, Kelvyn K. Bladen
Unifying And Expanding Global And Local Variable Importance Methods For Explainable Machine Learning, Kelvyn K. Bladen
All Graduate Theses and Dissertations, Fall 2023 to Present
Machine learning methods are powerful analytical tools used across all scientific disciplines and many other fields of investigation for prediction and inference from diverse data sources. Despite their broad applicability, machine learning methods are often highly complex and difficult to interpret. Developing a greater understanding of which variables most influence a response is essential for increasing the interpretability of these models and supporting informed decision-making. This research focuses on improving how we evaluate the importance of these variables.
One common approach is to shuffle the values of a variable and see how much the model accuracy gets worse. Another approach …
Machine Learning For Predictive Energy And Emissions Modeling Of Vehicles And Power Grids In The United States, S M Tanvir Faysal Alam Chowdhoury
Machine Learning For Predictive Energy And Emissions Modeling Of Vehicles And Power Grids In The United States, S M Tanvir Faysal Alam Chowdhoury
Dissertations
The environmental benefits of electric vehicle (EV) adoption depend on more than replacing internal combustion engine vehicles with electric powertrains. EV adoption reshapes electricity demand, interacts with regional generation mixes, and influences travel behavior and congestion, creating a coupled transportation-energy system in which vehicle and power-plant emissions must be evaluated together. This dissertation develops machine-learning frameworks for predicting energy consumption and emissions from vehicles and power grids under rising EV adoption. The first component forecasts grid emissions from EV charging. Using simulation data from NREL's Cambium database, a Prophet-based time-series framework predicts carbon dioxide, nitrous oxide, and methane emission rates …
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Evaluating Machine Learning Models On Classification Of Novel Cyber Attacks In The Healthcare Domain, Promise Ehimen
Dissertations, Theses, and Projects
The increasing adoption of the Internet of Medical Things (IoMT) has improved healthcare delivery through connected medical devices while simultaneously expanding the cybersecurity risks facing healthcare organizations. Although machine learning based intrusion detection systems have demonstrated high detection accuracy, their ability to respond reliably to previously unseen cyberattacks remains uncertain. This study investigated how a Neural Network model and a Logistic Regression model classified novel cyberattacks within the IoMT environment. The Neural Network and Logistic Regression models were both trained and tested using a subset of the CICIoMT2024 benchmark dataset. The Neural Network achieved 99.82% test accuracy and a 0.94 …
A Mathematical Decision-Making Framework For Athlete Development In A Collegiate Taekwondo Community: Prioritizing Coaching Interventions Using Statistical Analysis And The Analytic Hierarchy Process, King Harold A. Recto, Hazel Jade L. Antonio, Jhyrald Anthony P. Dalida
A Mathematical Decision-Making Framework For Athlete Development In A Collegiate Taekwondo Community: Prioritizing Coaching Interventions Using Statistical Analysis And The Analytic Hierarchy Process, King Harold A. Recto, Hazel Jade L. Antonio, Jhyrald Anthony P. Dalida
Electronics, Computer, and Communications Engineering Faculty Publications
Athlete development within collegiate sports communities requires informed decisions regarding the prioritization of coaching interventions and allocation of developmental resources. However, such decisions are frequently guided by experience and intuition, limiting opportunities for systematic and evidence-based decision-making. This study develops a mathematical decision-making framework for athlete development by integrating statistical analysis and the Analytic Hierarchy Process (AHP) within a collegiate taekwondo community. Data were collected from 25 collegiate taekwondo athletes who satisfied established eligibility criteria, including participation in University Athletic Association of the Philippines (UAAP) competitions during the previous three seasons. Athletes evaluated coaching practices across five dimensions: Training and …
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
University Honors Theses
This paper summarizes and describes the development of a set of learning materials that were created to educate students and professionals from other fields in statistical modeling techniques. These materials are primarily aimed at biology students, but are still intended to be useful for anyone who is interested in incorporating decision trees and random forest models into their personal research in the future. By directing the reader towards the JMP software, these materials navigate around the statistical knowledge base and coding implementation practices that otherwise would serve as a barrier to learning statistical modeling techniques, and instead focus on the …
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Master's Theses
Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.
Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.
We use time splitting and Mel-frequency cepstrum …
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Master's Theses
Unsupervised clustering algorithms today are used across a wide variety of fields such as biology, engineering, and industry in order to classify observations into groups where labels are not provided. This can provide important latent information regarding the observations within groups, as well as insight regarding the groups themselves. In order to judge the optimal number of clusters for an unsupervised clustering algorithm, many methods exist such as the Elbow Method and Silhouette Score; however, these methods come with drawbacks and are not necessarily flexible across many unsupervised methods. We present a novel clustering score framework relying on a resampling-based …
Transportation Deserts And Structural Mobility Access In New York City, John Cruz
Transportation Deserts And Structural Mobility Access In New York City, John Cruz
Student Theses
This study examines structural mobility access across New York City census tracts by constructing a tract-level Mobility Access Index (MAI) that integrates employment accessibility, hospital accessibility, and first-mile subway walking burden using MTA GTFS transit data, NYC Taxi and Limousine Commission trip records, and US Census American Community Survey demographic estimates. Three OLS regression models, supplemented by Lasso, Elastic Net, and Random Forest specifications, test whether structural access gaps are associated with short-distance connector trip intensity and per-worker connector cost burden. Results show that MAI varies substantially across tracts, with high access concentrated in Manhattan and along major subway corridors. …
Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama
Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama
Northeast Journal of Complex Systems (NEJCS)
Financial markets play a critical role in resource allocation. Their performance depends on the decisions of millions of independent investors constantly reacting to one another. Their informational efficiency remains a subject of debate across economic systems. When informational efficiency is present at the weak form, historical price information should not consistently predict future returns. Several empirical tests of this hypothesis often focus on the behavior of aggregate market indices, and use individual efficiency proxies such as autocorrelation, GARCH-type volatility, or entropy-based measures to measure efficiency. This has often yielded mixed results, particularly in emerging markets. Here we show that testing …
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
Civil and Environmental Engineering Theses and Dissertations
Urban areas are increasingly exposed to natural hazards while accommodating a growing share of the global population, yet a consistent science-based framework for quantifying urban and community resilience remains lacking. This dissertation develops a physics-based analytical framework grounded in statistical mechanics and the quantitative theory of Brownian motion. A city is conceptualized as a complex medium in which citizens move analogously to Brownian particles within a viscoelastic environment, influenced by socioeconomic interactions and infrastructure functionality.
A central premise is that urban resilience, interpreted as engineering resilience (an outcome), can be quantified through a single metric: the mean-square displacement MSD=⟨r²(t)⟩, of …
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
Assessing Trends In Medical Students’ Perceptions Regarding Statistical Analysis, Ethan Noble, Valeriy Kozmenko, Paul Thompson
Assessing Trends In Medical Students’ Perceptions Regarding Statistical Analysis, Ethan Noble, Valeriy Kozmenko, Paul Thompson
Scholarship Pathways Program
Assessing Trends in Medical Students’ Perceptions Regarding Data Analysis and Statistics Knowledge and Skills
Ethan Noble, MD | Mentors: Valeriy Kozmenko, MD, Paul Thompson, PhD
Introduction: The use of evidence-based medicine requires that physicians are able to properly analyze and interpret the results of new research. The development of new research and medical knowledge is swift, and a strong foundation in statistics and research is needed for physicians and medical students to keep up with new research. Curriculum in medical education often lacks in-depth coverage of the subject, and additional curriculum has been shown to enhance student confidence and ability …
Freshman 15? Freshman 50?? The Reality Of Daily Life Habits Of A First Year College Student, Gregorio R. Salgado
Freshman 15? Freshman 50?? The Reality Of Daily Life Habits Of A First Year College Student, Gregorio R. Salgado
Student Scholar Symposium Abstracts and Posters
This project presents a personal data tracking study in which I collected daily self-reported metrics over the course of the Spring semester using Excel. The variables tracked include sleep duration, caloric intake, screen time, social media usage, phone checks per day, family communication, and personal spending. The goal of this project is to identify meaningful patterns and correlations between daily habits and personal well-being.
Data was collected through a combination of manual logging and smartphone-generated daily reports. This study explores potential relationships between variables such as sleep duration and social media usage, as well as the association between family communication …
The Impatience Of Winning: An Analysis Of Time Discounting, Predictive Modeling, And The Nba Draft, Alec R. Plante
The Impatience Of Winning: An Analysis Of Time Discounting, Predictive Modeling, And The Nba Draft, Alec R. Plante
Business and Economics Honors Papers
This paper examines whether NBA draft decisions can be better explained by incorporating non-geometric time discounting into a model of general manager decision making. Using a dataset of 285 NBA draft prospects over a 12-year period, the impact of college statistics on Value Over Replacement Player (VORP) is determined, and these impact values are then used to create a “predicted” VORP for the first 4 seasons of each player’s career: a projection of what a general manager might think of a prospect’s future value given their college statistics. Following this, geometric and hyperbolic time discounting models are applied to estimate …
Fossil-Fuels In A Decarbonized Country? Modeling The Drivers Of Icelandic Oil Sales, Inbal Armony
Fossil-Fuels In A Decarbonized Country? Modeling The Drivers Of Icelandic Oil Sales, Inbal Armony
Environmental Studies Honors Projects
Although 100% of Iceland’s electricity comes from renewable energy sources, it still relies on fossil fuels for land transportation, marine transportation, aviation, and some industry. Understanding geographic nuances in oil use is critical to achieving Iceland’s goals of carbon neutrality by 2040. As the island has one primary urban center with two thirds of the population, information is lacking about oil use in non-Capital areas and a gap between state and municipal climate plans. Using newly available data of oil sales at the municipality-level in a Small Area Estimation model, we analyze drivers of oil sales across Icelandic municipalities. We …
Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder
Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder
Electrical Engineering and Computer Science Undergraduate Honors Theses
Splines are used for representing complex functions. In statistics, splines can be used for distributional shapes that are difficult to model by traditional parametric approaches. Ramsay (1) uses M-Spline bases to estimate continuous distributions. Semi-Parametric Quantile Regression (SPQR), developed by Xu and Reich (2), models conditional distributions where a neural network is used to estimate the basis function weights that depend on covariates. (3) implements a package for SPQR in R. We build on this by implementing a version of SPQR in Python with PyTorch. By using PyTorch, we can use more sophisticated deep learning architectures than those available in …
High Throughput Phenomics Pipeline For Pulse Crop Nutritional Breeding, Amod Udayanga Madurapperumage
High Throughput Phenomics Pipeline For Pulse Crop Nutritional Breeding, Amod Udayanga Madurapperumage
All Dissertations
Dry pea (Pisum sativum L.), lentil (Lens culinaris Medik.), and chickpea (Cicer arietinum L.) are major pulse crops valued for their high nutritional composition and importance to global food systems. Pulses are rich in carbohydrates, protein, and essential minerals, making them ideal whole foods and critical contributors to food and nutrition security. Due to these advantages, pulse breeding programs are increasingly focusing on enhancing nutritional traits, such as protein quality, amino acid balance, and micronutrient density, through the process of biofortification. However, improvement of agronomic traits remains equally essential. Characteristics such as plant height, standability, stress tolerance, …
A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari
A New Approach To Generate Combinatorial Patterns In Logical Analysis Of Data And Its Application To Predict College Retention, Salihah Ahmed E. Jaafari
Theses and Dissertations
Student retention and degree completion remain central challenges for higher-education institutions, with significant implications for student success, institutional effectiveness, and public accountability. While advances in predictive analytics have enabled earlier identification of students at risk of withdrawal, many commonly used machine learning approaches suffer from limited interpretability, constraining their practical usefulness for advising, intervention, and policy decision making. This dissertation addresses the problem of predicting student persistence by developing and evaluating optimization based, interpretable classification models within the Logical Analysis of Data (LAD) framework. Building on existing LAD formulations, this research introduces two novel pattern generation models, the Best Term …
Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg
Understanding Delays In Emergency Department Care: A National Analysis Of Wait Times, Gregory Forsberg
Mathematics, Statistics, and Computer Science Honors Projects
Emergency department (ED) wait times remain a persistent bottleneck in the United States healthcare system, impacting patient outcomes, hospital efficiency, and equitable access to care. This study analyzes nationally representative data from the National Hospital Ambulatory Medical Care Survey (NHAMCS), a complex, multi-stage probability sample. Using survey-weighted analyses and predictive modeling, we examine the effects of patient characteristics, triage acuity, and visit timing. Results indicate that operational and system-level factors, including hospital capacity, geographic region, and temporal variation, are among the most influential predictors of ED wait times
The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue
The Item Response Warehouse: What It Is, How To Use It, And Targets For Potential Improvements, Savira D. Nadela, Hansol Lee, Nishka Jain, Ayaan Gupta, Xingyi Zhang, Benjamin W. Domingue
Chinese/English Journal of Educational Measurement and Evaluation | 教育测量与评估双语期刊
The Item Response Warehouse (IRW) is a repository of harmonized item response datasets designed to support secondary analysis and methodological research in psychological and educational measurement. This paper serves as a practical guide for researchers interested in using the IRW. We describe the structure of IRW datasets and the quantitative and qualitative metadata available for dataset selection, and we demonstrate how researchers can navigate the IRW website to explore and compare available tables. We further show how the IRW R and Python packages can be used to filter datasets programmatically, download response-level data, and generate standardized citations for reproducible research …
Complex Systems Mapping Of Fiscal Growth Dynamics At Strategic Maritime Chokepoints Using Time-Series Slopes, Rahul Balamurugan, Preethi Nanjundan, Avichal Sharma
Complex Systems Mapping Of Fiscal Growth Dynamics At Strategic Maritime Chokepoints Using Time-Series Slopes, Rahul Balamurugan, Preethi Nanjundan, Avichal Sharma
Northeast Journal of Complex Systems (NEJCS)
This study examines how maritime and trading states allocate public resources between defence, health, and economic growth around three strategic chokepoints the Strait of Malacca, the Strait of Hormuz, and the Suez Canal. The analysis extends the classic “guns versus butter” framing by treating defence and health spending as co-evolving components of an interconnected fiscal-growth system. Using World Development Indicators data (1999-2024), trend slopes are estimated for military spending (% of GDP), healthcare spending (% of GDP), and GDP growth (annual %). Two derived indicators are computed, a defence-to-health slope ratio (military slope/health slope) and a fiscal-balance proxy (health slope …
Information Theory Analysis Of The Solar Wind Magnetic Structures For Space Weather Prediction, Katherine Holland
Information Theory Analysis Of The Solar Wind Magnetic Structures For Space Weather Prediction, Katherine Holland
Doctoral Dissertations and Master's Theses
Forecasting space weather at Earth is highly complicated, because of the limited measurements of the dynamic processes in the Sun that span multiple temporal, spatial, and energy-scales. The solar wind is a highly structured, multi-scale, evolving plasma and consists of coronal mass ejections (CMEs), stream interaction regions (SIRs), expanding flux tubes (Borovsky, 2008), and interplanetary magnetic field (IMF) discontinuities and fluctuations. The aim of this research is to improve our understanding of the evolution and dissipation of different scale-size solar wind magnetic structures as they move from the Sun-Earth Lagrange point 1 (L1) to Earth's bow shock and, ultimately, to …
Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue
Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue
Articles
We introduce Saturated Hierarchical Atomic Incremental Learning (sHAIL), a learning paradigm in which complex tasks are approached through a sequence of simpler atomic subtasks, each mastered to saturation before progression. The central mechanism is a saturation criterion that detects when learning dynamics enter a plateau region, triggering consolidation and subsequent ascent to a higher level of task complexity. We develop a theoretical framework for sHAIL and show that it naturally gives rise to \emph{staircased convergence}: alternating phases of rapid improvement and genuine plateau. Within each level, classical convergence guarantees apply under standard smoothness conditions, while the hierarchical transitions are driven …
No Intelligence Without Statistics: The Invisible Backbone Of Artificial Intelligence, Ernest Fokoue
No Intelligence Without Statistics: The Invisible Backbone Of Artificial Intelligence, Ernest Fokoue
Articles
The rapid ascent of artificial intelligence (AI) is often portrayed as a revolution born from computer science and engineering. This narrative, however, obscures a fundamental truth: the theoretical and methodological core of AI is, and has always been, statistical. This paper systematically argues that the field of statistics provides the indispensable foundation for machine learning and modern AI. We deconstruct AI into nine foundational pillars—Inference, Density Estimation, Sequential Learning, Generalization, Representation Learning, Interpretability, Causality, Optimization, and Unification—demonstrating that each is built upon century-old statistical principles. From the inferential frameworks of hypothesis testing and estimation that underpin model evaluation, to the …
Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental
Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental
Articles
Social insect colonies and ensemble machine learning methods represent two of the most successful examples of decentralized information processing in nature and computation respectively. Here we develop a rigorous mathematical framework demonstrating that ant colony decision-making and random forest learning are isomorphic under a common formalism of stochastic ensemble intelligence. We show that the mechanisms by which genetically identical ants achieve functional differentiation— through stochastic response to local cues and positive feedback—map precisely onto the bootstrap aggregation and random feature subsampling that decorrelate decision trees. Using tools from Bayesian inference, multi-armed bandit theory, and statistical learning theory, we prove that …
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue
A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue
Articles
Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. …
On The Scientific Stature Of Data Science: The Epistemological Unicorn, Ernest Fokoue
On The Scientific Stature Of Data Science: The Epistemological Unicorn, Ernest Fokoue
Articles
Data Science has ignited unprecedented academic, industrial, and pedagogical fervor, yet its status as a \textit{science} in the classical sense---comparable to physics or biology---remains profoundly unsettled. This article interrogates the epistemological foundations of Data Science by examining its hybrid theoretical lineage, from the Universal Approximation Theorem to the No-Free-Lunch Theorems, with special emphasis on the fundamental Bayesian optimality results for both regression and classification. We argue that Data Science is in a vigorous \textit{gestational period}, characterized not by an absence of principles but by a creative tension between empirical pragmatism and deep mathematical theory. The Cross-Validation score emerges as the …
On Fibonacci Ensembles: An Alternative Approach To Ensemble Learning Inspired By The Timeless Architecture Of The Golden Ratio, Ernest Fokoue
On Fibonacci Ensembles: An Alternative Approach To Ensemble Learning Inspired By The Timeless Architecture Of The Golden Ratio, Ernest Fokoue
Articles
Nature rarely reveals her secrets bluntly, yet in the Fibonacci sequence she grants us a glimpse of her quiet architecture of growth, harmony, and recursive stability \citep{Koshy2001Fibonacci, Livio2002GoldenRatio}. From spiral galaxies to the unfolding of leaves, this humble sequence reflects a universal grammar of balance. In this work, we introduce \emph{Fibonacci Ensembles}, a mathematically principled yet philosophically inspired framework for ensemble learning that complements and extends classical aggregation schemes such as bagging, boosting, and random forests \citep{Breiman1996Bagging, Breiman2001RandomForests, Friedman2001GBM, Zhou2012Ensemble, HastieTibshiraniFriedman2009ESL}. Two intertwined formulations unfold: (1) the use of normalized Fibonacci weights -- tempered through orthogonalization and Rao--Blackwell optimization -- …