Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles

Discipline
Institution
Keyword
Publication Year

Articles 1 - 23 of 23

Full-Text Articles in Data Science

Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue Mar 2026

Saturated Hierarchical Atomic Incremental Learning (Shail): A Behavioral Learning Perspective On Staged Mastery And Saturation, Ernest Fokoue

Articles

We introduce Saturated Hierarchical Atomic Incremental Learning (sHAIL), a learning paradigm in which complex tasks are approached through a sequence of simpler atomic subtasks, each mastered to saturation before progression. The central mechanism is a saturation criterion that detects when learning dynamics enter a plateau region, triggering consolidation and subsequent ascent to a higher level of task complexity. We develop a theoretical framework for sHAIL and show that it naturally gives rise to \emph{staircased convergence}: alternating phases of rapid improvement and genuine plateau. Within each level, classical convergence guarantees apply under standard smoothness conditions, while the hierarchical transitions are driven …


No Intelligence Without Statistics: The Invisible Backbone Of Artificial Intelligence, Ernest Fokoue Mar 2026

No Intelligence Without Statistics: The Invisible Backbone Of Artificial Intelligence, Ernest Fokoue

Articles

The rapid ascent of artificial intelligence (AI) is often portrayed as a revolution born from computer science and engineering. This narrative, however, obscures a fundamental truth: the theoretical and methodological core of AI is, and has always been, statistical. This paper systematically argues that the field of statistics provides the indispensable foundation for machine learning and modern AI. We deconstruct AI into nine foundational pillars—Inference, Density Estimation, Sequential Learning, Generalization, Representation Learning, Interpretability, Causality, Optimization, and Unification—demonstrating that each is built upon century-old statistical principles. From the inferential frameworks of hypothesis testing and estimation that underpin model evaluation, to the …


Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental Mar 2026

Decorrelation, Diversity, And Emergent Intelligence: The Isomorphism Between Social Insect Colonies And Ensemble Machine Learning, Ernest Fokoue, Gregory Babbitt, Yuval Levental

Articles

Social insect colonies and ensemble machine learning methods represent two of the most successful examples of decentralized information processing in nature and computation respectively. Here we develop a rigorous mathematical framework demonstrating that ant colony decision-making and random forest learning are isomorphic under a common formalism of stochastic ensemble intelligence. We show that the mechanisms by which genetically identical ants achieve functional differentiation— through stochastic response to local cues and positive feedback—map precisely onto the bootstrap aggregation and random feature subsampling that decorrelate decision trees. Using tools from Bayesian inference, multi-armed bandit theory, and statistical learning theory, we prove that …


A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue Mar 2026

A General Weighting Theory For Ensemble Learning: Beyond Variance Reduction Via Spectral And Geometric Structure, Ernest Fokoue

Articles

Ensemble learning is traditionally justified as a variance-reduction strategy, explaining its strong performance for unstable predictors such as decision trees. This explanation, however, does not account for ensembles constructed from intrinsically stable estimators-including smoothing splines, kernel ridge regression, Gaussian process regression, and other regularized reproducing kernel Hilbert space (RKHS) methods whose variance is already tightly controlled by regularization and spectral shrinkage. This paper develops a general weighting theory for ensemble learning that moves beyond classical variance-reduction arguments. We formalize ensembles as linear operators acting on a hypothesis space and endow the space of weighting sequences with geometric and spectral constraints. …


On The Scientific Stature Of Data Science: The Epistemological Unicorn, Ernest Fokoue Mar 2026

On The Scientific Stature Of Data Science: The Epistemological Unicorn, Ernest Fokoue

Articles

Data Science has ignited unprecedented academic, industrial, and pedagogical fervor, yet its status as a \textit{science} in the classical sense---comparable to physics or biology---remains profoundly unsettled. This article interrogates the epistemological foundations of Data Science by examining its hybrid theoretical lineage, from the Universal Approximation Theorem to the No-Free-Lunch Theorems, with special emphasis on the fundamental Bayesian optimality results for both regression and classification. We argue that Data Science is in a vigorous \textit{gestational period}, characterized not by an absence of principles but by a creative tension between empirical pragmatism and deep mathematical theory. The Cross-Validation score emerges as the …


On Fibonacci Ensembles: An Alternative Approach To Ensemble Learning Inspired By The Timeless Architecture Of The Golden Ratio, Ernest Fokoue Mar 2026

On Fibonacci Ensembles: An Alternative Approach To Ensemble Learning Inspired By The Timeless Architecture Of The Golden Ratio, Ernest Fokoue

Articles

Nature rarely reveals her secrets bluntly, yet in the Fibonacci sequence she grants us a glimpse of her quiet architecture of growth, harmony, and recursive stability \citep{Koshy2001Fibonacci, Livio2002GoldenRatio}. From spiral galaxies to the unfolding of leaves, this humble sequence reflects a universal grammar of balance. In this work, we introduce \emph{Fibonacci Ensembles}, a mathematically principled yet philosophically inspired framework for ensemble learning that complements and extends classical aggregation schemes such as bagging, boosting, and random forests \citep{Breiman1996Bagging, Breiman2001RandomForests, Friedman2001GBM, Zhou2012Ensemble, HastieTibshiraniFriedman2009ESL}. Two intertwined formulations unfold: (1) the use of normalized Fibonacci weights -- tempered through orthogonalization and Rao--Blackwell optimization -- …


Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue Mar 2026

Learning Ordinal Geometry: Semantic–Aware Kernels For Ordered Categorical Data, Ernest Fokoue

Articles

Ordinal data arise ubiquitously in survey research, psychology, medicine, economics, and recommender systems, yet kernel methods for such data typically rely on either nominal encodings or arbitrary numeric codings. The former discards order information; the lat- ter imposes a fictitious metric structure. This paper develops a principled framework for kernel design on ordinal scales and introduces a new class of Semantic–Aware Ordinal Ker- nels (SAOK) that simultaneously capture ordinal order and semantic proximity between categories. We begin by formalizing order–preserving embeddings of finite chains and characterizing a broad family of chain distances that are conditionally negative definite. Through Schoen- berg …


A Method For Generating A Non-Manual Feature Model For Sign Language Processing, Robert G. Smith Dr, Markus Hofmann Dr Aug 2023

A Method For Generating A Non-Manual Feature Model For Sign Language Processing, Robert G. Smith Dr, Markus Hofmann Dr

Articles

While recent approaches to sign language processing have shifted to the domain of Machine Learning (ML), the treatment of Non-Manual Features (NMFs) remains an open question. The principal challenge facing this method is the comparatively small sign language corpora available for training machine learning models. This study produces a statistical model which may be used in future ML, rules-based, and hybrid-learning approaches for sign language processing tasks. In doing so, this research explores the emerging patterns of non-manual articulation concerning grammatical classes in Irish Sign Language (ISL). The experimental method applied here is a novel implementation of an association rules …


Determining The Proportionality Of Ischemic Stroke Risk Factors To Age, Elizabeth Hunter, John D. Kelleher Jan 2023

Determining The Proportionality Of Ischemic Stroke Risk Factors To Age, Elizabeth Hunter, John D. Kelleher

Articles

While age is an important risk factor, there are some disadvantages to including it in a stroke risk model: age can dominate the risk score and lead to over-or under-predictions in some age groups. There is evidence to suggest that some of these disadvantages are due to the non-proportionality of other risk factors with age, eg, risk factors contribute differently to stroke risk based on an individual’s age. In this paper, we present a framework to test if risk factors are proportional with age. We then apply the framework to a set of risk factors using Framingham heart study data …


The Interaction Of Normalisation And Clustering In Sub-Domain Definition For Multi-Source Transfer Learning Based Time Series Anomaly Detection, Matthew Nicholson, Rahul Agrahari, Clare Conran, Haythem Assem, John D. Kelleher Dec 2022

The Interaction Of Normalisation And Clustering In Sub-Domain Definition For Multi-Source Transfer Learning Based Time Series Anomaly Detection, Matthew Nicholson, Rahul Agrahari, Clare Conran, Haythem Assem, John D. Kelleher

Articles

This paper examines how data normalisation and clustering interact in the definition of sub-domains within multi-source transfer learning systems for time series anomaly detection. The paper introduces a distinction between (i) clustering as a primary/direct method for anomaly detection, and (ii) clustering as a method for identifying sub-domains within the source or target datasets. Reporting the results of three sets of experiments, we find that normalisation after feature extraction and before clustering results in the best performance for anomaly detection. Interestingly, we find that in the multi-source transfer learning scenario clustering on the target dataset and identifying subdomains in the …


“Be A Pattern For The World”: The Development Of A Dark Patterns Detection Tool To Prevent Online User Loss, Jordan Donnelly, Alan Dowley, Yunpeng Liu, Yufei Su, Quanwei Sun, Lan Zeng, Andrea Curley, Damian Gordon, Paul Kelly, Dympna O'Sullivan, Anna Becevel Sep 2022

“Be A Pattern For The World”: The Development Of A Dark Patterns Detection Tool To Prevent Online User Loss, Jordan Donnelly, Alan Dowley, Yunpeng Liu, Yufei Su, Quanwei Sun, Lan Zeng, Andrea Curley, Damian Gordon, Paul Kelly, Dympna O'Sullivan, Anna Becevel

Articles

Dark Patterns are designed to trick users into sharing more information or spending more money than they had intended to do, by configuring online interactions to confuse or add pressure to the users. They are highly varied in their form, and are therefore difficult to classify and detect. Therefore, this research is designed to develop a framework for the automated detection of potential instances of web-based dark patterns, and from there to develop a software tool that will provide a highly useful defensive tool that helps detect and highlight these patterns.


Self-Supervised Learning For Invariant Representations From Multi-Spectral And Sar Images, Pallavi Jain, Bianca Schoen Phelan, Robert J. Ross Sep 2022

Self-Supervised Learning For Invariant Representations From Multi-Spectral And Sar Images, Pallavi Jain, Bianca Schoen Phelan, Robert J. Ross

Articles

Self-Supervised learning (SSL) has become the new state of the art in several domain classification and segmentation tasks. One popular category of SSL are distillation networks such as Bootstrap Your Own Latent (BYOL). This work proposes RS-BYOL, which builds on BYOL in the remote sensing (RS) domain where data are non-trivially different from natural RGB images. Since multi-spectral (MS) and synthetic aperture radar (SAR) sensors provide varied spectral and spatial resolution information, we utilise them as an implicit augmentation to learn invariant feature embeddings. In order to learn RS based invariant features with SSL, we trained RS-BYOL in two ways, …


Assessing Feature Representations For Instance-Based Cross-Domain Anomaly Detection In Cloud Services Univariate Time Series Data, Rahul Agrahari, Matthew Nicholson, Clare Conran, Haythem Assem, John D. Kelleher Jan 2022

Assessing Feature Representations For Instance-Based Cross-Domain Anomaly Detection In Cloud Services Univariate Time Series Data, Rahul Agrahari, Matthew Nicholson, Clare Conran, Haythem Assem, John D. Kelleher

Articles

In this paper, we compare and assess the efficacy of a number of time-series instance feature representations for anomaly detection. To assess whether there are statistically significant differences between different feature representations for anomaly detection in a time series, we calculate and compare confidence intervals on the average performance of different feature sets across a number of different model types and cross-domain time-series datasets. Our results indicate that the catch22 time-series feature set augmented with features based on rolling mean and variance performs best on average, and that the difference in performance between this feature set and the next best …


Towards Exchanging Wearable-Pghd With Ehrs: Developing A Standardized Information Model For Wearable-Based Patient Generated Health Data, Abdullahi Abubakar Kawu, Dympna O'Sullivan, Lucy Hederman Jan 2022

Towards Exchanging Wearable-Pghd With Ehrs: Developing A Standardized Information Model For Wearable-Based Patient Generated Health Data, Abdullahi Abubakar Kawu, Dympna O'Sullivan, Lucy Hederman

Articles

Wearables have become commonplace for tracking and making sense of patient lifestyle, wellbeing and health data. Most of this tracking is done by individuals outside of clinical settings, however some data from wearables may be useful in a clinical context. As such, wearables may be considered a prominent source of Patient Generated Health Data (PGHD). Studies have attempted to maximize the use of the data from wearables including integrating with Electronic Health Records (EHRs). However, usually a limited number of wearables are considered for integration and, in many cases, only one brand is investigated. In addition, we find limited studies …


Provenance: An Intermediary-Free Solution For Digital Content Verification, Bilal Yousuf, M. Atif Qureshi, Brendan Spillane, Gary Munnelly, Oisin Carroll, Matthew Runswick, Kirsty Park, Eileen Culloty, Owen Conlan, Jane Suiter Nov 2021

Provenance: An Intermediary-Free Solution For Digital Content Verification, Bilal Yousuf, M. Atif Qureshi, Brendan Spillane, Gary Munnelly, Oisin Carroll, Matthew Runswick, Kirsty Park, Eileen Culloty, Owen Conlan, Jane Suiter

Articles

The threat posed by misinformation and disinformation is one of the defining challenges of the 21st century. Provenance is designed to help combat this threat by warning users when the content they are looking at may be misinformation or disinformation. It is also designed to improve media literacy among its users and ultimately reduce susceptibility to the threat among vulnerable groups within society. The Provenance browser plugin checks the content that users see on the Internet and social media and provides warnings in their browser or social media feed. Unlike similar plugins, which require human experts to provide evaluations and …


Virtual Network Function Embedding Under Nodal Outage Using Deep Q-Learning, Swarna Bindu Chetty, Hamed Ahmadi, Sachin Sharma, Avishek Nag Mar 2021

Virtual Network Function Embedding Under Nodal Outage Using Deep Q-Learning, Swarna Bindu Chetty, Hamed Ahmadi, Sachin Sharma, Avishek Nag

Articles

With the emergence of various types of applications such as delay-sensitive applications, future communication networks are expected to be increasingly complex and dynamic. Network Function Virtualization (NFV) provides the necessary support towards efficient management of such complex networks, by virtualizing network functions and placing them on shared commodity servers. However, one of the critical issues in NFV is the resource allocation for the highly complex services; moreover, this problem is classified as an NP-Hard problem. To solve this problem, our work investigates the potential of Deep Reinforcement Learning (DRL) as a swift yet accurate approach (as compared to integer linear …


An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill Jan 2021

An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill

Articles

This paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple part-of-speech taggers to annotate natural language text at a higher quality than the part-of-speech taggers are able to obtain independently. Our ensemble uses three state-of-the-art part-of-speech taggers: SWUM, POSSE, and Stanford. We study the quality of the ensemble's annotations on five different types of identifier names: function, class, attribute, parameter, and declaration statement at the level of both individual words and full identifier names. We also study and discuss the weaknesses of our tagger to …


Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao Dec 2020

Data: The Good, The Bad And The Ethical, John D. Kelleher, Filipe Cabral Pinto, Luis M. Cortesao

Articles

It is often the case with new technologies that it is very hard to predict their long-term impacts and as a result, although new technology may be beneficial in the short term, it can still cause problems in the longer term. This is what happened with oil by-products in different areas: the use of plastic as a disposable material did not take into account the hundreds of years necessary for its decomposition and its related long-term environmental damage. Data is said to be the new oil. The message to be conveyed is associated with its intrinsic value. But as in …


Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr. Oct 2020

Comparing Variable Importance In Prediction Of Silence Behaviours Between Random Forest And Conditional Inference Forest Models., Stephen Barrett Dr, Geraldine Gray Dr, Colm Mcguinness Dr, Michael Knoll Dr.

Articles

This paper explores variable importance metrics of Conditional Inference Trees (CIT) and classical Classification And Regression Trees (CART) based Random Forests. The paper compares both algorithms variable importance rankings and highlights why CIT should be used when dealing with data with different levels of aggregation. The models analysed explored the role of cultural factors at individual and societal level when predicting Organisational Silence behaviours.


Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha Aug 2020

Multi‑View Clustering For Multi‑Omics Data Using Unifed Embedding, Mohammed Hasanuzzaman, Sayantan Mitra, Sriparna Saha

Articles

In real world applications, data sets are often comprised of multiple views, which provide consensus and complementary information to each other. Embedding learning is an effective strategy for nearest neighbour search and dimensionality reduction in large data sets. This paper attempts to learn a unified probability distribution of the points across different views and generates a unified embedding in a low-dimensional space to optimally preserve neighbourhood identity. Probability distributions generated for each point for each view are combined by conflation method to create a single unified distribution. The goal is to approximate this unified distribution as much as possible when …


An Application Of Machine Learning To Explore Relationships Between Factors Of Organisational Silence And Culture, With Specific Focus On Predicting Silence Behaviours, Stephen Barrett Dr May 2020

An Application Of Machine Learning To Explore Relationships Between Factors Of Organisational Silence And Culture, With Specific Focus On Predicting Silence Behaviours, Stephen Barrett Dr

Articles

Research indicates that there are many individual reasons why people do not speak up when confronted with situations that may concern them within their working environment. One of the areas that requires more focused research is the role culture plays in why a person may remain silent when such situations arise. The purpose of this study is to use data science techniques to explore the patterns in a data set that would lead a person to engage in organisational silence. The main research question the thesis asks is: Is Machine Learning a tool that Social Scientists can use with respect …


Finding Common Ground For Citizen Empowerment In The Smart City, John D. Kelleher, Aphra Kerr Jan 2020

Finding Common Ground For Citizen Empowerment In The Smart City, John D. Kelleher, Aphra Kerr

Articles

Corporate smart city initiatives are just one example of the contemporary culture of surveillance. They rely on extensive information gathering systems and Big Data analysis to predict citizen behaviour and optimise city services. In this paper we argue that many smart city and social media technologies result in a paradox whereby digital inclusion for the purposes of service provision also results in marginalisation and disempowerment of citizens. Drawing upon insights garnered from a digital inclusion workshop conducted in the Galapagos islands, we propose that critically and creatively unpacking the computational techniques embedded in data services is needed as a first …


On The Exactitude Of Big Data: La Bêtise And Artificial Intelligence, Noel Fitzpatrick, John D. Kelleher Dec 2018

On The Exactitude Of Big Data: La Bêtise And Artificial Intelligence, Noel Fitzpatrick, John D. Kelleher

Articles

This article revisits the question of ‘la bêtise’ or stupidity in the era of Artificial Intelligence driven by Big Data, it extends on the questions posed by Gille Deleuze and more recently by Bernard Stiegler. However, the framework for revisiting the question of la bêtise will be through the lens of contemporary computer science, in particular the development of data science as a mode of analysis, sometimes, misinterpreted as a mode of intelligence. In particular, this article will argue that with the advent of forms of hype (sometimes referred to as the hype cycle) in relation to big data and …