Data-Driven Characterization Of Counties In The Prison Industrial Complex Using Clustering Analysis,
2026
Mississippi State University
Data-Driven Characterization Of Counties In The Prison Industrial Complex Using Clustering Analysis, Riley N. Tuccio
Capstone Projects
This project investigates the complex relationship between counties that house prisons in the United States and the rurality associated with them. The central research question explores how both county characteristics, such as variables corresponding to cost of living and demographics of a county, and prison characteristics, such as programming available to inmates and staffing levels, differ across the census-designated rural-urban distinctions. Furthermore, the study examines whether modern data science methods can more accurately define and distinguish these characteristics, providing a nuanced understanding of the Prison Industrial Complex (PIC) and its manifestation across various American communities. The motivation for this research …
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing,
2026
Southern Methodist University
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Mathematics Theses and Dissertations
This dissertation presents a computational framework for high-frequency options trading that combines Cross-Data-Type 1-D Convolutional Neural Networks (CDT-1D CNN) with Simpson-Sobolev regularization for directional prediction, and finite element methods (FEM) for realistic option pricing during backtesting. The core innovation lies in developing a mathematically rigorous regularization approach that maintains the adaptability of modern deep learning while enabling accurate evaluation through stochastic volatility models. The primary contribution is the Simpson-Sobolev regularization scheme, which extends traditional Sobolev regularization by incorporating Simpson’s rule for numerical integration. This approach achieves higher-order accuracy in approximating the Sobolev norms that control function smoothness. Simpson’s rule attains …
Unmasking Twitter Bots: An Applied Machine Learning Approach,
2026
Department of Mathematics and Computer Science, Faculty of Science, Beirut Arab University, Lebanon
Unmasking Twitter Bots: An Applied Machine Learning Approach, Rayane El Raba’A, Layal Abu Daher
BAU Journal - Science and Technology
The rapid growth of social networks has led to increased challenges, such as fraud, cyberbullying, and the spread of automated accounts (bots). Detecting anomalies within these networks is essential to maintaining security and trust. This study explored machine learning algorithms: Random Forest, XGBoost, Support Vector Machine (SVM), and Logistic Regression for anomaly detection in social networks, specifically focusing on Twitter bot identification, By applying AI-driven data mining techniques to a dataset of 37,438 Twitter bot accounts dataset, the research evaluates the effectiveness of these models in detecting unusual patterns. XGBoost achieved the highest accuracy (84.9%), with an ROA_AUC of 0.87, …
Ai For Regression Analysis And More,
2026
Washington University in St. Louis
Ai For Regression Analysis And More, Eli Snir
Generative AI Teaching Activities
Students use Copilot and NotebookLM to create a dataset and develop statistical analyses including regression.
Integration Of Intraoperative Data In Interpretable Machine Learning Models To Predict Postoperative Aki In Noncardiac Surgery Patients,
2026
Thomas Jefferson University
Integration Of Intraoperative Data In Interpretable Machine Learning Models To Predict Postoperative Aki In Noncardiac Surgery Patients, Justin Do, Karan H. Shah, Melissa Xu, Andrew Hyunwoo Kim, Vivaswat Suresh, Nidhir Guggilla, Michael Li, Rishi Kothari
Department of Anesthesiology Faculty Papers
OBJECTIVES: We aimed to (1) quantify changes in discrimination when adding intraoperative data to preoperative data and (2) compare tabular machine learning with feature engineering against a time-aware LSTM-based model.
MATERIALS AND METHODS: Retrospective cohort of 46 204 adults undergoing 57 055 eligible noncardiac surgery in the INSPIRE database. We extracted 38 preoperative and 49 intraoperative variables; acute kidney injury (AKI) was defined by KDIGO serum creatinine criteria and modeled as stage 2/3 postoperative AKI. Models were trained on preoperative-only and combined pre- and intraoperative data. Intraoperative series were summarized using eight statistical features for tabular models or integrated directly …
Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks,
2026
Thomas Jefferson High School for Science and Technology
Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer
CODEE Journal
Undergraduate instruction in ordinary differential equations (ODEs) is typically organized around the forward problem: finding solution trajectories when the governing equation and its parameters are known. In scientific practice, however, inverse problems are often more relevant, requiring unknown parameters to be inferred from noisy observations while assessing whether a proposed model is consistent with the data. We introduce PINNLab, an open-source MATLAB dashboard designed to help undergraduate students explore inverse modeling through physics-informed neural networks (PINNs). PINNLab presents PINNs as a complementary data-driven framework that connects differential equations, optimization, empirical data, and scientific machine learning. The instructional sequence is organized …
Pitching Fwar Vs Bwar As Predictors Of Team Success In The Mlb Regular Season,
2026
Portland State University
Pitching Fwar Vs Bwar As Predictors Of Team Success In The Mlb Regular Season, Jason Lee
University Honors Theses
Pitching Wins Above Replacement (WAR) is an area of sabermetrics capable of being used to predict regular season success in Major League Baseball. Fangraphs WAR (fWAR) and Baseball Reference WAR (bWAR) were used to construct regression models to predict regular season winning percentage, to build logistic models to establish a relationship between pitching WAR and the probability to win an individual regular season game, and to overlay density plots to consider WAR accumulation by pitching role and observe the difference of impact between starter and relief pitchers. This research finds that while pitching fWAR and pitching bWAR are both statistically …
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms,
2026
Portland State University
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
University Honors Theses
This paper summarizes and describes the development of a set of learning materials that were created to educate students and professionals from other fields in statistical modeling techniques. These materials are primarily aimed at biology students, but are still intended to be useful for anyone who is interested in incorporating decision trees and random forest models into their personal research in the future. By directing the reader towards the JMP software, these materials navigate around the statistical knowledge base and coding implementation practices that otherwise would serve as a barrier to learning statistical modeling techniques, and instead focus on the …
A Novel, Embedding-Based Approach To Longitudinal Survey Data Imputation,
2026
Portland State University
A Novel, Embedding-Based Approach To Longitudinal Survey Data Imputation, Julia Rezvani
University Honors Theses
Longitudinal surveys are ubiquitous in the social sciences as a means of tracking changes in behavior and opinions with time and identifying potential causal mechanisms. These surveys are frequently plagued by missing data and semantic drift, both of which limit their effectiveness and scientific utility. Imputation algorithms allow researchers to fill gaps in collected survey datasets, imperfectly reconstructing lost data. Although deep learning algorithms have been used in imputation to great success, approaches which simultaneously leverage the semantic and temporal structure of longitudinal surveys have not yet been developed. We propose a novel imputation architecture which is capable of leveraging …
A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content,
2026
Department of Mathematical Sciences, Purdue University Fort Wayne
A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson
CODEE Journal
In the age of data-driven decision making, ordinary differential equations (ODEs) remain a powerful and interpretable framework for modeling dynamic processes, especially when integrated with modern tools from statistical learning and data-driven dynamical systems. Yet, general undergraduate and graduate curricula do not typically address key opportunities in data-driven dynamical systems.
This first paper in a series focuses on the mathematical and methodological core of a professional development course first developed in the academic year 2025-2026 at a Primarily Undergraduate Institution, Purdue University Fort Wayne. The curriculum developed in this course emphasized how regression, regularization, and sparse identification can be used …
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models,
2026
California Polytechnic State University, San Luis Obispo
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Master's Theses
Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.
Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.
We use time splitting and Mel-frequency cepstrum …
Fairlinked: Data Fairification Tools For Materials Data Science,
2026
Case Western Reserve University
Fairlinked: Data Fairification Tools For Materials Data Science, Van D. Tran, Brandon Lee, Ritika Lamba, Henry Dirks, Quynh D. Tran, Balashanmuga Priyan Rajamohan, Ozan Dernek, Laura S. Bruckman, Yinghui Wu, Erika I. Barcelos, Roger H. French
Student Scholarship
FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of …
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs,
2026
CUNY Graduate Center
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin
Dissertations, Theses, and Capstone Projects
We perform field-level likelihood-free inference of the matter density parameter Ωm from simulated galaxy catalogs using machine learning models with differing inductive biases. Using features extracted from hydrodynamic simulations in the CAMELS suite, we investigate how both observable choice and model architecture govern the extraction of cosmological information. We consider galaxy positions and line-of-sight peculiar velocities, both separately and in combination, and compare permutation-invariant Deep Sets, implemented with either standard multilayer perceptrons (MLPs) or Kolmogorov–Arnold Networks (KANs), to graph neural networks (GNNs) implemented with MLPs, which explicitly encode spatial relations. We evaluate inference performance under both in-distribution and out-of-distribution (OOD) …
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study,
2026
CUNY Graduate Center
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Dissertations, Theses, and Capstone Projects
About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …
Determining K Clusters In K-Means Clustering With The Crab Algorithm,
2026
California Polytechnic State University, San Luis Obispo
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Master's Theses
Unsupervised clustering often faces the challenge of determining the correct number of clusters in the absence of a true target variable. Traditional methods such as the Elbow Method and the Silhouette Score can produce ambiguous results and rely on assumptions about cluster shape or separation. To address this, we created the Clustering Rivals and Buddies (CRAB) algorithm which evaluates clusters based on stability across multiple subsamples. CRAB uses pairwise classifications to identify points that consistently group together called “Buddies” and points that remain separated called “Rivals.” Applied with K-means, CRAB accurately recovers underlying cluster structures in both spherical and non-spherical …
The Security Of Llm-Generated Code,
2026
CUNY John Jay College
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
Student Theses
The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …
Can Generative Ai Make Farming Decisions? Current Status And Future Pathways: A Case Study In Row Crop Production With Chatgpt,
2026
University of Nebraska-Lincoln
Can Generative Ai Make Farming Decisions? Current Status And Future Pathways: A Case Study In Row Crop Production With Chatgpt, Nipuna Chamara, Yufeng Ge, Joe Luck, Yu Pan, Saleh Taghvaeian, Cory Walters, Christopher Proctor, Daran Rudnick, Daren Redfearn
Department of Agricultural and Biological Systems Engineering: Faculty Publications
The agricultural decision-making process is experience-based, knowledge-dependent, time-sensitive, complex, and driven by historical data. Planting, fertilization, irrigation, and chemigation are key categories in farm decision-making, and currently there is no one-shot decision-support tool that covers all these activities. Generative Artificial Intelligence (AI) models are more advanced than traditional machine learning and deep learning models. These models have been trained on vast amounts of data from the internet, allowing them to accept unstructured data in various forms and generate human-like text, solutions to problems, and scenario predictions. Given this capability, we became interested in exploring the potential of generative AI in …
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning,
2026
California Polytechnic State University, San Luis Obispo
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Master's Theses
Unsupervised clustering algorithms today are used across a wide variety of fields such as biology, engineering, and industry in order to classify observations into groups where labels are not provided. This can provide important latent information regarding the observations within groups, as well as insight regarding the groups themselves. In order to judge the optimal number of clusters for an unsupervised clustering algorithm, many methods exist such as the Elbow Method and Silhouette Score; however, these methods come with drawbacks and are not necessarily flexible across many unsupervised methods. We present a novel clustering score framework relying on a resampling-based …
Edge Co-Occurrence Regularization For Node Classification,
2026
New Jersey Institute of Technology
Edge Co-Occurrence Regularization For Node Classification, Kadir Altunel
Theses
We propose a simple yet effective regularization technique for node classification on graphs that leverages edge-based label co-occurrence patterns. We first train an MLP on node features to produce class probability distributions, then compute a fixed penalty matrix from edge-based co-occurrence statistics of these predictions. This penalty matrix, which captures unlikely class combinations on connected nodes, is then used to regularize GNN training without further updates. We evaluate this approach across multiple homophilic datasets (Cora, CiteSeer, PubMed, ogbn-arxiv) and heterophilic benchmarks (Chameleon, Squirrel, Actor, Roman-Empire) using three GNN architectures: GCN, GraphSAGE, and H2GCN. Results show consistent improvements on homophilic graphs, …
A Generative Ai-Driven Computational Framework For Industry-Scale Discovery Of Novel Battery Materials,
2026
New Jersey Institute of Technology
A Generative Ai-Driven Computational Framework For Industry-Scale Discovery Of Novel Battery Materials, Joy Datta
Dissertations
The growing demand for sustainable, high-energy-density electrochemical storage has motivated the exploration of multivalent-ion batteries based on earth-abundant elements such as aluminum, calcium, magnesium, and zinc. While multivalent charge carriers offer higher theoretical energy density than lithium, their practical deployment is hindered by sluggish ion transport, strong ion-host interactions, and structural degradation of electrode materials. Identifying host materials that can reversibly accommodate multivalent ions while maintaining structural integrity remains a fundamental challenge. The dissertation develops a scalable, end-to-end computational framework that integrates density functional theory (DFT), machine learning (ML), and generative artificial intelligence (GenAI) to accelerate the discovery of next-generation …
