Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,231 Full-Text Articles 9,305 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,231 full-text articles. Page 3 of 155.

A Novel, Embedding-Based Approach To Longitudinal Survey Data Imputation, Julia Rezvani 2026 Portland State University

A Novel, Embedding-Based Approach To Longitudinal Survey Data Imputation, Julia Rezvani

University Honors Theses

Longitudinal surveys are ubiquitous in the social sciences as a means of tracking changes in behavior and opinions with time and identifying potential causal mechanisms. These surveys are frequently plagued by missing data and semantic drift, both of which limit their effectiveness and scientific utility. Imputation algorithms allow researchers to fill gaps in collected survey datasets, imperfectly reconstructing lost data. Although deep learning algorithms have been used in imputation to great success, approaches which simultaneously leverage the semantic and temporal structure of longitudinal surveys have not yet been developed. We propose a novel imputation architecture which is capable of leveraging …


A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson 2026 Department of Mathematical Sciences, Purdue University Fort Wayne

A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson

CODEE Journal

In the age of data-driven decision making, ordinary differential equations (ODEs) remain a powerful and interpretable framework for modeling dynamic processes, especially when integrated with modern tools from statistical learning and data-driven dynamical systems. Yet, general undergraduate and graduate curricula do not typically address key opportunities in data-driven dynamical systems.

This first paper in a series focuses on the mathematical and methodological core of a professional development course first developed in the academic year 2025-2026 at a Primarily Undergraduate Institution, Purdue University Fort Wayne. The curriculum developed in this course emphasized how regression, regularization, and sparse identification can be used …


Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski 2026 California Polytechnic State University, San Luis Obispo

Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski

Master's Theses

Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.

Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.

We use time splitting and Mel-frequency cepstrum …


Fairlinked: Data Fairification Tools For Materials Data Science, Van D. Tran, Brandon Lee, Ritika Lamba, Henry Dirks, Quynh D. Tran, Balashanmuga Priyan Rajamohan, Ozan Dernek, Laura S. Bruckman, Yinghui Wu, Erika I. Barcelos, Roger H. French 2026 Case Western Reserve University

Fairlinked: Data Fairification Tools For Materials Data Science, Van D. Tran, Brandon Lee, Ritika Lamba, Henry Dirks, Quynh D. Tran, Balashanmuga Priyan Rajamohan, Ozan Dernek, Laura S. Bruckman, Yinghui Wu, Erika I. Barcelos, Roger H. French

Student Scholarship

FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of …


Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera 2026 California Polytechnic State University, San Luis Obispo

Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera

Master's Theses

Unsupervised clustering often faces the challenge of determining the correct number of clusters in the absence of a true target variable. Traditional methods such as the Elbow Method and the Silhouette Score can produce ambiguous results and rely on assumptions about cluster shape or separation. To address this, we created the Clustering Rivals and Buddies (CRAB) algorithm which evaluates clusters based on stability across multiple subsamples. CRAB uses pairwise classifications to identify points that consistently group together called “Buddies” and points that remain separated called “Rivals.” Applied with K-means, CRAB accurately recovers underlying cluster structures in both spherical and non-spherical …


The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala 2026 CUNY John Jay College

The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala

Student Theses

The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …


Can Generative Ai Make Farming Decisions? Current Status And Future Pathways: A Case Study In Row Crop Production With Chatgpt, Nipuna Chamara, Yufeng Ge, Joe Luck, Yu Pan, Saleh Taghvaeian, Cory Walters, Christopher Proctor, Daran Rudnick, Daren Redfearn 2026 University of Nebraska-Lincoln

Can Generative Ai Make Farming Decisions? Current Status And Future Pathways: A Case Study In Row Crop Production With Chatgpt, Nipuna Chamara, Yufeng Ge, Joe Luck, Yu Pan, Saleh Taghvaeian, Cory Walters, Christopher Proctor, Daran Rudnick, Daren Redfearn

Department of Agricultural and Biological Systems Engineering: Faculty Publications

The agricultural decision-making process is experience-based, knowledge-dependent, time-sensitive, complex, and driven by historical data. Planting, fertilization, irrigation, and chemigation are key categories in farm decision-making, and currently there is no one-shot decision-support tool that covers all these activities. Generative Artificial Intelligence (AI) models are more advanced than traditional machine learning and deep learning models. These models have been trained on vast amounts of data from the internet, allowing them to accept unstructured data in various forms and generate human-like text, solutions to problems, and scenario predictions. Given this capability, we became interested in exploring the potential of generative AI in …


Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain 2026 CUNY Graduate Center

Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain

Dissertations, Theses, and Capstone Projects

About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …


Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin 2026 CUNY Graduate Center

Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin

Dissertations, Theses, and Capstone Projects

We perform field-level likelihood-free inference of the matter density parameter Ωm from simulated galaxy catalogs using machine learning models with differing inductive biases. Using features extracted from hydrodynamic simulations in the CAMELS suite, we investigate how both observable choice and model architecture govern the extraction of cosmological information. We consider galaxy positions and line-of-sight peculiar velocities, both separately and in combination, and compare permutation-invariant Deep Sets, implemented with either standard multilayer perceptrons (MLPs) or Kolmogorov–Arnold Networks (KANs), to graph neural networks (GNNs) implemented with MLPs, which explicitly encode spatial relations. We evaluate inference performance under both in-distribution and out-of-distribution (OOD) …


Bridging Data Gaps In Retinal Imaging: From Structural Domain Adaptation To Topology-Aware Synthesis, Gözde Merve Demirci 2026 CUNY Graduate Center

Bridging Data Gaps In Retinal Imaging: From Structural Domain Adaptation To Topology-Aware Synthesis, Gözde Merve Demirci

Dissertations, Theses, and Capstone Projects

Comprehensive visualization of the retina is essential for diagnosing and monitoring blinding diseases such as Diabetic Retinopathy and Retinopathy of Prematurity (ROP), where pathological changes often extend beyond a single field of view. Despite significant advances in automated retinal image analysis, clinical deployment remains limited by two fundamental data gaps: a structural learning gap, arising from scarce expert annotations and poor generalization across imaging domains, and a spatial coverage gap, caused by the difficulty of acquiring multi-view retinal images in fragile populations. Although these challenges are often addressed independently, this dissertation argues that they are tightly coupled: accurate, …


Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi 2026 California Polytechnic State University, San Luis Obispo

Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi

Master's Theses

Unsupervised clustering algorithms today are used across a wide variety of fields such as biology, engineering, and industry in order to classify observations into groups where labels are not provided. This can provide important latent information regarding the observations within groups, as well as insight regarding the groups themselves. In order to judge the optimal number of clusters for an unsupervised clustering algorithm, many methods exist such as the Elbow Method and Silhouette Score; however, these methods come with drawbacks and are not necessarily flexible across many unsupervised methods. We present a novel clustering score framework relying on a resampling-based …


Edge Co-Occurrence Regularization For Node Classification, Kadir Altunel 2026 New Jersey Institute of Technology

Edge Co-Occurrence Regularization For Node Classification, Kadir Altunel

Theses

We propose a simple yet effective regularization technique for node classification on graphs that leverages edge-based label co-occurrence patterns. We first train an MLP on node features to produce class probability distributions, then compute a fixed penalty matrix from edge-based co-occurrence statistics of these predictions. This penalty matrix, which captures unlikely class combinations on connected nodes, is then used to regularize GNN training without further updates. We evaluate this approach across multiple homophilic datasets (Cora, CiteSeer, PubMed, ogbn-arxiv) and heterophilic benchmarks (Chameleon, Squirrel, Actor, Roman-Empire) using three GNN architectures: GCN, GraphSAGE, and H2GCN. Results show consistent improvements on homophilic graphs, …


A Generative Ai-Driven Computational Framework For Industry-Scale Discovery Of Novel Battery Materials, Joy Datta 2026 New Jersey Institute of Technology

A Generative Ai-Driven Computational Framework For Industry-Scale Discovery Of Novel Battery Materials, Joy Datta

Dissertations

The growing demand for sustainable, high-energy-density electrochemical storage has motivated the exploration of multivalent-ion batteries based on earth-abundant elements such as aluminum, calcium, magnesium, and zinc. While multivalent charge carriers offer higher theoretical energy density than lithium, their practical deployment is hindered by sluggish ion transport, strong ion-host interactions, and structural degradation of electrode materials. Identifying host materials that can reversibly accommodate multivalent ions while maintaining structural integrity remains a fundamental challenge. The dissertation develops a scalable, end-to-end computational framework that integrates density functional theory (DFT), machine learning (ML), and generative artificial intelligence (GenAI) to accelerate the discovery of next-generation …


Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan 2026 New Jersey Institute of Technology

Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan

Dissertations

Digital three-dimensional (3D) models are central to engineering design, analysis, and manufacturing, but learning pipelines for man-made geometry often operate on sampled carriers that do not preserve all of the structure present in exact CAD representations. This dissertation studies learning-based reconstruction and part decomposition for structured man-made 3D geometry, from general object benchmarks to CAD-derived datasets, with a focus on neural implicit representations trained from signed-distance samples, point clouds, and tessellated meshes. The goal is to make these models more accurate, more part-aware, and more consistently supervised.

First, signed distance function (SDF) reconstruction with implicit neural representations is improved through …


Transportation Deserts And Structural Mobility Access In New York City, John Cruz 2026 CUNY School of Professional Studies

Transportation Deserts And Structural Mobility Access In New York City, John Cruz

Student Theses

This study examines structural mobility access across New York City census tracts by constructing a tract-level Mobility Access Index (MAI) that integrates employment accessibility, hospital accessibility, and first-mile subway walking burden using MTA GTFS transit data, NYC Taxi and Limousine Commission trip records, and US Census American Community Survey demographic estimates. Three OLS regression models, supplemented by Lasso, Elastic Net, and Random Forest specifications, test whether structural access gaps are associated with short-distance connector trip intensity and per-worker connector cost burden. Results show that MAI varies substantially across tracts, with high access concentrated in Manhattan and along major subway corridors. …


Forecasting Interborough Express Ridership Using Network-Based Station Typologies And Direct Demand Models, Fomba Kassoh 2026 CUNY School of Professional Studies

Forecasting Interborough Express Ridership Using Network-Based Station Typologies And Direct Demand Models, Fomba Kassoh

Student Theses

Forecasting ridership for new transit infrastructure is difficult in the absence of observed outcomes, particularly under domain shift between an existing system and a proposed corridor. This study develops a station-level direct demand modeling (DDM) framework to forecast average weekday ridership for the proposed Interborough Express (IBX) in New York City — a 14-mile circumferential rapid transit corridor connecting Brooklyn and Queens. The approach pairs unsupervised learning with supervised estimation in a common feature space defined by transit service, accessibility, and built-environment characteristics. K-means clustering identifies latent station typologies (node–place regimes), and IBX stations are projected into this topology to …


Predicting The Outcome Of Ischemic Hepatitis With Real-Patient Data Using Machine Learning Tools, Christiana Beard, Madison Utterback, Olcay Akman, Priya Kohli, William M. Lee, Aditi Ghosh 2026 Illinois State University

Predicting The Outcome Of Ischemic Hepatitis With Real-Patient Data Using Machine Learning Tools, Christiana Beard, Madison Utterback, Olcay Akman, Priya Kohli, William M. Lee, Aditi Ghosh

Spora: A Journal of Biomathematics

Ischemic hepatitis (IH) results from shock-related conditions that impair oxygenated blood flow to the liver, causing hepatocyte death. Diagnosis relies largely on clinical history due to the absence of specific diagnostic tests and limited ability to predict outcomes. This study applies machine learning methods to real-world IH patient data to improve outcome prediction. Biomedical indicators analyzed include creatinine, international normalized ratio (INR), aspartate aminotransferase (AST), alanine transaminase (ALT), and bilirubin. Data were collected from multiple U.S. centers through the Acute Liver Failure Study Group (ALFSG), a multicenter network focused on this rare condition. We implemented logistic regression, regression tree methods …


Evaluating Soil Health And Crop Yield In Louisiana Agricultural Systems: Impacts Of Best Management Practices And Prediction Models, Hector J. Mendoza Lagos 2026 Louisiana State University and Agricultural and Mechanical College

Evaluating Soil Health And Crop Yield In Louisiana Agricultural Systems: Impacts Of Best Management Practices And Prediction Models, Hector J. Mendoza Lagos

LSU Doctoral Dissertations

The adoption of conservation management practices is critical for improving soil health, enhancing nutrient use efficiency, and sustaining crop productivity in row crop systems in Louisiana. This study evaluated the role of conservation agronomic practices, soil biochemical indicators, and machine learning predictive models to improve soil nutrient dynamics, soil health indicators, microbial communities (MC), and crop productivity on a corn (Zea mays L.) research plot scale and in a commercial forty-hectare cotton (Gassypium hirsutum L.)-corn-soybean (Glycine max L.) rotation system in northeast Louisiana. The objectives of the study were to evaluate soil nutrient dynamics and MCs under …


Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama 2026 Binghamton University--SUNY

Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama

Northeast Journal of Complex Systems (NEJCS)

Financial markets play a critical role in resource allocation. Their performance depends on the decisions of millions of independent investors constantly reacting to one another. Their informational efficiency remains a subject of debate across economic systems. When informational efficiency is present at the weak form, historical price information should not consistently predict future returns. Several empirical tests of this hypothesis often focus on the behavior of aggregate market indices, and use individual efficiency proxies such as autocorrelation, GARCH-type volatility, or entropy-based measures to measure efficiency. This has often yielded mixed results, particularly in emerging markets. Here we show that testing …


Material Costs, Karima Weinman 2026 Rhode Island School of Design

Material Costs, Karima Weinman

Masters Theses

This thesis investigates how migration fatality and disappearance data can be reinterpreted through material craft to create a more reflective encounter with information. Working with the Missing Migrants Project's dataset, this project asks how design can communicate dimensions of human loss that conventional data visualization cannot reach.

The work situates contemporary border violence within a longer colonial history, arguing that the logics of surveillance and quantification that structured European imperial expansion persist in the databases that govern mobility in the Mediterranean today.

Terrazzo is a 15th-century Venetian flooring technique built from discarded fragments bound together into a unified surface. This …


Digital Commons powered by bepress