Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons™

Open Access. Powered by Scholars. Published by Universities.®

3,268 Full-Text Articles 9,368 Authors 1,358,113 Downloads 222 Institutions

All Articles in Data Science

Faceted Search

3,268 full-text articles. Page 9 of 157.

Evaluating The Effects Of Anti-Forensic Activities In Additive Manufacturing Devices, Daniel B. Miller 2026 University of South Alabama

Evaluating The Effects Of Anti-Forensic Activities In Additive Manufacturing Devices, Daniel B. Miller

Shelby Hall Graduate Research Forum Posters

Additive Manufacturing (AM) is a newer famlily of production tecchologies that constructs objects by fusing layers of material into the desired shape. Methods for achieving this as described in Gibson et al. [1] are varied and include Fused Filament Deposition, Selective Laser Sintering, Stereolithography (SLA), and Powder Bed Fusion. Computers are integral to the processes being responsible for creating and decoding design instructions, collecting and processing sensor data, and, ultimately, directing the activity of the machines which implement the process. Additionally, the AM industry is rapidly expanding, worth an estimated $23 billion in 2023 and projected to reach $88 billion …


Cellebrite Reliability In Digital Forensics, Christina Huynh 2026 University of South Alabama

Cellebrite Reliability In Digital Forensics, Christina Huynh

Shelby Hall Graduate Research Forum Posters

Forensic tools like Cellebrite are commonly used in court to gather and interpret raw data for evidence. Cellebrite does not only collect data but creates and interprets the artifacts of data to create a scene of the process it has been through. With this, evidence can be influenced by software designs and not just the data on the mobile device. Courts and police use Cellebrite to gather evidence and reconstruct it to create an easily readable dataset. These tools lack reproducibility, transparency, integrity, and chain of evidence command. Cellebrite is often used in court and by police without further vetting …


Evaluating Regularized Logistic Regression And K-Nn On Mnist Under Increasing Random Missingness, Daniel Markwei 2026 University of Central Florida

Evaluating Regularized Logistic Regression And K-Nn On Mnist Under Increasing Random Missingness, Daniel Markwei

Data Science and Data Mining

This paper investigates the effect of random missingness on the performance of regularized multinomial logistic regression and the k-nearest neighbors (k-NN) classifier for handwritten digit recognition on the MNIST dataset. In particular, we study L1-regularized (LASSO) logistic regression and L2-regularized (Ridge) logistic regression alongside k-NN. Varying percentages of random missingness were introduced into the original dataset, and each model was evaluated in terms of its classification performance. The results show that random missingness degrades the performance of all three classifiers. Overall, k-NN consistently achieves higher accuracy than both L1- and L2-regularized logistic regression across all missingness levels; however, its performance …


Optimization Of Image Quality Of Simulated Multiple Detectors Computed Tomography Acquisition Parameters Using Machine Learning And Catphan Phantom, Ali O. Masoud, Najat K. Mohammed, Khamis O. Amour, Ahmed M. Jusabani, Denise Mwalongo, Mwingereza John Kumwenda 2026 Directorate of Radiation Control Unit, Tanzania Atomic Energy Commission (TAEC) P, O Box 1585, Dodoma, Tanzania

Optimization Of Image Quality Of Simulated Multiple Detectors Computed Tomography Acquisition Parameters Using Machine Learning And Catphan Phantom, Ali O. Masoud, Najat K. Mohammed, Khamis O. Amour, Ahmed M. Jusabani, Denise Mwalongo, Mwingereza John Kumwenda

Tanzania Journal of Science

The study successfully employed Monte Carlo (MC) simulation and a Machine Learning (ML) approach using a Random Forest Regression (RFR) model to develop optimized Multi-Detector CT (MDCT) protocols that significantly reduce radiation dose while maintaining diagnostic image quality. The MC engine accurately modeled X-ray spectra, and the RFR model demonstrated high predictive power for key metrics, achieving R2 scores of 0.97 for CTDIvol and over 0.92 for image quality metrics (Noise, CNR). Through multi-objective optimization guided by the RFR, the final protocol (Optimization-3) was found on the Pareto front, achieving a notable 35% dose reduction (from 15.5 mGy to 9.9 …


Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell 2026 The University of Alabama, Tuscaloosa

Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell

Proceedings from the Document Academy

Generative Artificial Intelligences (AIs) and current advanced large language models (LLMs) are algorithmically designed to generate text-based conversations as conversational agents (CAs), by replicating human language and conversational communication. Pairing human cognition with generative computationally coded cognition. We have never been here before: cerebral and artificial information collaborations and processing producing expressions that may or may not become visible as second-hand/secondary source documents.

Sensemaking or sense(un)making is a unique autonomous human drive cognitively, our information processing is sensemaking in action and expressions and articulations are evidence of the sensemaking cycle. Documentation [expressed or articulated through various mediums] are a product …


Effective Wordle Heuristics, Ronald I. Greenberg 2026 Loyola University Chicago

Effective Wordle Heuristics, Ronald I. Greenberg

Computer Science: Faculty Publications and Other Works

While previous researchers have performed an exhaustive search to determine an optimal Wordle strategy, that computation is very time consuming and produced a strategy using words that are unfamiliar to most people. With Wordle solutions being gradually eliminated (with a new puzzle each day and no reuse), an improved strategy could be generated each day, but the computation time makes a daily exhaustive search impractical. This paper shows that simple heuristics allow for fast generation of effective strategies and that little is lost by guessing only words that are possible solution words rather than more obscure words.


Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra 2026 Smith College

Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra

Statistical and Data Sciences: Faculty Publications

We present tidychangepoint, a new R package for changepoint detection analysis. Most R packages for segmenting univariate time series focus on providing one or two algorithms for changepoint detection that work with a small set of models and penalized objective functions, and all of them return a custom, nonstandard object type. This makes comparing results across various algorithms, models, and penalized objective functions unnecessarily difficult. tidychangepoint solves this problem by wrapping functions from a variety of existing packages and storing the results in a common S3 class called tidycpt. The package then provides functionality for easily extracting comparable numeric or …


Bibliography For Love Data Week 2026, Annikah Carpio, Sally Park 2026 Chapman University

Bibliography For Love Data Week 2026, Annikah Carpio, Sally Park

Library Displays and Bibliographies

A bibliography created to support a display about research data and Love Data Week during February 2026 at the Leatherby Libraries at Chapman University.


Courts Of New York: A Visual Atlas Of The City’S Public Basketball Spaces, Nathaniel Rattner 2026 CUNY Graduate Center

Courts Of New York: A Visual Atlas Of The City’S Public Basketball Spaces, Nathaniel Rattner

Dissertations, Theses, and Capstone Projects

Basketball courts in New York City are recreation facilities, community anchors and part of the city’s cultural image. In the basketball capital of the world, New Yorkers are rarely more than a few blocks away from a court. The visual diversity of these courts, however, is not widely documented in systematic ways.

This project makes that diversity visible to the public, combining open data, aerial imagery and computational analysis to document this important public space across the five boroughs. It is a narrative story and digital atlas of New York City’s public basketball courts, using surface color as a way …


Typeface: Machine-Viewing Gentrification On Storefront Imagery In Bedford-Stuyvesant, Brooklyn, Alexander McQuilkin 2026 CUNY Graduate Center

Typeface: Machine-Viewing Gentrification On Storefront Imagery In Bedford-Stuyvesant, Brooklyn, Alexander Mcquilkin

Dissertations, Theses, and Capstone Projects

Gentrification—broadly, the replacement of a less powerful group by a more powerful one in an urban context—is oft-discussed in the popular press, but its definition is much-debated in the urban planning literature. Furthermore, academic treatments of displacement understandably focus on measurable yet fairly abstract indicators like changes in rent or income, whereas neighborhood change is often registered by residents on the ground using visual, but difficult-to-quantify markers like retail turnover. This project uses image recognition technology on a set of storefront photos to index the visual streetscape of a neighborhood, as well as to track changes to that portrait over …


Removed: When Taxi Drivers Meet Dynamic Pricing: A Lesson From Singapore's Justgrab Program, Shih-Fen CHENG, Wen-Tai HSU, Jing LI 2026 Singapore Management University

Removed: When Taxi Drivers Meet Dynamic Pricing: A Lesson From Singapore's Justgrab Program, Shih-Fen Cheng, Wen-Tai Hsu, Jing Li

Research Collection School Of Economics

This paper studies how dynamic pricing influences taxi drivers’ behaviors using a unique event, the inception of the JustGrab program in Singapore in 2017, which introduces dynamic pricing to some, but not all, taxi drivers. This is the first time in history that traditional taxi drivers have access to dynamic pricing. Using data covering the universe of taxi trips before and after the inception of JustGrab, we find that there is spatial reallocation that directs more taxi drivers to the previously less-served areas, that there is also a temporal reallocation that directs more taxi drivers to rush hours, as well …


Interpretable Linear Models For Heart Disease Prediction: A Comparative Study, Dipok Deb, Emran Hossain 2026 PhD Student, Big Data Analytics, UCF

Interpretable Linear Models For Heart Disease Prediction: A Comparative Study, Dipok Deb, Emran Hossain

Data Science and Data Mining

Heart disease remains a leading cause of mortality worldwide, underscoring the importance of accurate and transparent methods for early diagnosis. While many machine learning and artificial intelligence models have demonstrated strong predictive performance, their limited interpretability poses challenges for clinical adoption. In this study, we evaluate three interpretable linear classification models—Generalized Linear Model (GLM) logistic regression, L1-regularized (Lasso) logistic regression, and Linear Discriminant Analysis (LDA)—for heart disease prediction using the Cleveland Heart Disease dataset. Following comprehensive data preprocessing, the models are assessed on a held-out test set using standard evaluation metrics, including accuracy, precision, recall, F1-score, and the area under …


Correlation Analysis Of Factors Associated With Students’ Numeracy Skills Using Decision Tree Algorithms, Yenita Roza, Arisman Adnan, Zul Indra, Tuti Alawiyah 2026 Universitas Riau

Correlation Analysis Of Factors Associated With Students’ Numeracy Skills Using Decision Tree Algorithms, Yenita Roza, Arisman Adnan, Zul Indra, Tuti Alawiyah

Numeracy

Numeracy is a critical competency for academic and everyday functioning. This study investigates the key factors associated with students’ numeracy skills by employing decision tree algorithms as a data mining technique. The dataset used in this study is educational assessment data from Indonesia. Utilizing a dataset comprising 6,953 entries and 60 variables from Education Report, the research adopts an exploratory approach involving data preprocessing, exploratory data analysis, and decision tree model construction. The findings reveal that students’ literacy skills serve as the most dominant predictor of numeracy proficiency, emerging as the root node in the decision tree structure. Additional associated …


Predicting Male Flowering Time In Maize Using Machine Learning Technique, Dipok Deb 2026 PhD Student, Big Data Analytics, UCF

Predicting Male Flowering Time In Maize Using Machine Learning Technique, Dipok Deb

Data Science and Data Mining

This study compares three machine learning approaches—Elastic Net, Principal Component Regression (PCR), and Partial Least Squares (PLS)—for variable selection and prediction within a high-dimensional Maize-GWAS framework. The goal was to accurately predict the complex polygenic trait of time to male flowering while managing the challenges of numerous, highly correlated genetic markers. The ENET model, which combines l1 and l2 penalties, delivered the highest predictive accuracy and successfully identified a select subset of the most influential genetic variants. In contrast, PCR and PLS, both utilizing dimension reduction, offered a significant advantage in computational speed and model stability. The findings confirm that …


Fairmaterials: Ontology Tools With Data Fairification In Development, Alexander Harding Bradley, Jonathan E. Gordon, Balashanmuga Priyan Rajamohan, Van D. Tran, Nathaniel Hahn, Kiefer Lin, Hayden W. Caldwell, Arafath Nihar, Quynh D. Tran, Yinghui Wu, Laura S. Bruckman, Erika I. Barcelos, Roger H. French 2026 Case Western Reserve University

Fairmaterials: Ontology Tools With Data Fairification In Development, Alexander Harding Bradley, Jonathan E. Gordon, Balashanmuga Priyan Rajamohan, Van D. Tran, Nathaniel Hahn, Kiefer Lin, Hayden W. Caldwell, Arafath Nihar, Quynh D. Tran, Yinghui Wu, Laura S. Bruckman, Erika I. Barcelos, Roger H. French

Student Scholarship

The bilingual FAIRmaterials package simplifies the creation and visualization of materials and data science ontologies. FAIRmaterials, available in the Python and R languages, addresses the complexities associated with traditional ontology editors based on manual user input such as Protege (Musen, 2015) with an intuitive workflow and easy-to-use templates, making it accessible to users both experienced and inexperienced with ontologies. The FAIRmaterials package is its ability to programatically convert simple and structured CSV inputs into rich, well-defined ontologies. This capability is designed to support the findability, accessibility, interoperability, and reusability (FAIR) (Wilkinson et al., 2016) of research data and serve as …


Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng 2026 University of the Free State

Statistical Analysis Of Log Transformation Effectiveness In Air Traffic Movement Forecasting During Covid-19 In South Africa, John Lehlaka Masekoameng

Journal of Aviation Technology and Engineering

This study evaluates the effectiveness of log transformation in enhancing multiple regression models used to forecast air traffic movements (ATMs) in South Africa during the COVID-19 pandemic. Using 60 monthly observations from October 2016 to September 2021, the analysis incorporates variables such as revenue, lockdown levels, COVID-19 metrics, exchange rates, gross domestic product, and population. Two models are compared: one using raw ATMs and another with log-transformed ATMs as the dependent variable.

While the untransformed model shows stronger explanatory power (R² = 0.904, adjusted R² = 0.891) compared to the log-transformed model (R² = 0.772, adjusted R² = 0.741), the …


Topological Data Analysis Of New York City Taxi Trip Data Using Advanced Dimensionality Reduction And Clustering Techniques – Uncovering Structure And Timeliness, Mahalakshmi Sakthivel 2026 University of Arkansas Little Rock

Topological Data Analysis Of New York City Taxi Trip Data Using Advanced Dimensionality Reduction And Clustering Techniques – Uncovering Structure And Timeliness, Mahalakshmi Sakthivel

Theses and Dissertations

In this project, we used Topological Data Analysis (TDA) to explore the shape and structure of high-dimensional data and the timeliness dimension of information quality through topological data analysis, with the long-term goal of automatically computing timeliness values that reflect how useful data items are for decision making. The project followed a two-phase approach: in the first half, we employed the Kepler Mapper library along with techniques like Principal Component Analysis (PCA), Uniform Manifold Approximation and Projection (UMAP), T-SNE (t-Distributed Stochastic Neighbor Embedding) and Customized Embedding to analyze and visualize complex datasets. In the second phase, we specifically applied our …


Reproducible Semantic Data Management Workflow For Materials Data Science: Generating Knowledge Graphs With Robust Fairifcation Pipelines, Kyle R. Henrikson, Van D. Tran, Meredith Francis, Isabella Giammattei, Quynh D. Tran, Laura S. Bruckman, Erika I. Barcelos, Roger H. French 2026 Case Western Reserve University

Reproducible Semantic Data Management Workflow For Materials Data Science: Generating Knowledge Graphs With Robust Fairifcation Pipelines, Kyle R. Henrikson, Van D. Tran, Meredith Francis, Isabella Giammattei, Quynh D. Tran, Laura S. Bruckman, Erika I. Barcelos, Roger H. French

Student Scholarship

Combining data from multiple sources is crucial for efficient knowledge aggregation in materials data science. FAIR data from ontology and Linked Data principles enable this. Semantic data management streamlines data exchange and aggregation, ensuring information is available and extractable. FAIRLinked and GraphDB provide solutions for consolidating, hosting, and extracting meaningful insight from multimodal data.


Data-Driven Prediction Of Superconducting Critical Temperature: A Linear And Regularized Linear Modeling Approach, Dipok Deb 2026 PhD Student, Big Data Analytics, UCF

Data-Driven Prediction Of Superconducting Critical Temperature: A Linear And Regularized Linear Modeling Approach, Dipok Deb

Data Science and Data Mining

This study adopts a data-driven approach to estimate the critical temperature of superconducting materials using linear machine learning models. A comprehensive dataset derived from material physico-chemical properties was analyzed after systematic preprocessing and standardization. Three linear modeling strategies—Linear Regression, Ridge Regression, and Linear Regression with Subset Selection—were developed and evaluated using standard regression performance metrics. The findings demonstrate that both basic and regularized linear models can effectively capture the relationship between material features and superconducting behavior, offering robust and interpretable predictions. While feature selection enhances model transparency, it comes with a modest reduction in predictive capability. Overall, this work emphasizes …


Identifying Relevant Covariates In Rna-Seq Analysis By Pseudo-Variable Augmentation, Yet Nguyen, Dan Nettleton 2026 Old Dominion University

Identifying Relevant Covariates In Rna-Seq Analysis By Pseudo-Variable Augmentation, Yet Nguyen, Dan Nettleton

Mathematics & Statistics Faculty Publications

RNA-sequencing (RNA-seq) technology allows for the identification of differentially expressed genes, which are genes whose mean transcript abundance levels vary across conditions. In practice, RNA-seq datasets often include covariates that are of primary interest in addition to a set of covariates that are subject to selection. Some of these covariates may be relevant to gene expression levels, while others may be irrelevant. Ignoring relevant covariates or attempting to adjust for the effect of irrelevant covariates can compromise the identification of differentially expressed genes. To address this issue, we propose a variable selection method that uses pseudo-variables to control the expected …


Digital Commons powered by bepress