Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (1156)
- Medicine and Health Sciences (780)
- Life Sciences (765)
- Bioinformatics (568)
- Statistics and Probability (549)
-
- Biomedical Informatics (530)
- Engineering (527)
- Artificial Intelligence and Robotics (525)
- Social and Behavioral Sciences (519)
- Databases and Information Systems (212)
- Computer Engineering (208)
- Electrical and Computer Engineering (204)
- Applied Statistics (193)
- Medical Sciences (190)
- Business (189)
- Statistical Models (181)
- Applied Mathematics (175)
- Medical Specialties (173)
- Theory and Algorithms (149)
- Environmental Sciences (147)
- Mathematics (144)
- Other Computer Sciences (127)
- Data Storage Systems (123)
- Systems and Communications (120)
- Numerical Analysis and Scientific Computing (116)
- Public Health (116)
- Public Affairs, Public Policy and Public Administration (109)
- Statistical Methodology (109)
- Institution
-
- The Texas Medical Center Library (523)
- Old Dominion University (173)
- Southern Methodist University (144)
- Universitas Negeri Malang (113)
- City University of New York (CUNY) (100)
-
- CCT College Dublin (82)
- Chapman University (66)
- Kennesaw State University (63)
- University of Central Florida (62)
- Smith College (60)
- Air Force Institute of Technology (57)
- Embry-Riddle Aeronautical University (52)
- Singapore Management University (45)
- University of Arkansas, Fayetteville (45)
- Chinese Academy of Sciences (44)
- Purdue University (44)
- California Polytechnic State University, San Luis Obispo (39)
- Technological University Dublin (39)
- Illinois State University (38)
- University of Kentucky (38)
- University of Nebraska - Lincoln (38)
- New Jersey Institute of Technology (37)
- West Virginia University (37)
- Claremont Colleges (36)
- Virginia Commonwealth University (35)
- Clemson University (32)
- Dartmouth College (31)
- University of Texas at Arlington (27)
- East Tennessee State University (26)
- Minnesota State University, Mankato (26)
- Keyword
-
- Humans (278)
- Machine learning (241)
- Machine Learning (215)
- Deep learning (115)
- Computer Science (98)
-
- Deep Learning (93)
- Artificial Intelligence (65)
- Data science (58)
- Data Science (57)
- Natural Language Processing (56)
- COVID-19 (55)
- Artificial intelligence (53)
- Female (52)
- Male (50)
- Classification (49)
- Natural language processing (46)
- Animals (41)
- Data (41)
- Electronic Health Records (41)
- Neural Networks (40)
- Algorithms (38)
- Big data (37)
- Data mining (37)
- Statistics (36)
- Clustering (32)
- Computer science (31)
- Adult (30)
- NLP (30)
- Neural networks (30)
- AI (29)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (508)
- SMU Data Science Review (124)
- Knowledge Engineering and Data Science (113)
- Theses and Dissertations (111)
- ICT (82)
-
- Data Science and Data Mining (53)
- Dissertations (53)
- Statistical and Data Sciences: Faculty Publications (53)
- Electronic Theses and Dissertations (49)
- Bulletin of Chinese Academy of Sciences (Chinese Version) (44)
- Dissertations, Theses, and Capstone Projects (44)
- Research Collection School Of Computing and Information Systems (37)
- Master's Theses (35)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (34)
- Data Science Undergraduate Honors Theses (31)
- Annual Symposium on Biomathematics and Ecology Education and Research (30)
- Computer Science Faculty Publications (30)
- Publications and Research (30)
- All Graduate Theses, Dissertations, and Other Capstone Projects (24)
- Computational and Data Sciences (PhD) Dissertations (24)
- Symposium of Student Scholars (24)
- All Dissertations (23)
- Articles (23)
- Electrical & Computer Engineering Faculty Publications (22)
- CBN Journal of Applied Statistics (JAS) (21)
- College of Graduate Studies: Theses & Dissertations (20)
- CMC Senior Theses (19)
- Theses (19)
- Electronic Theses, Projects, and Dissertations (18)
- Faculty Publications (18)
- Publication Type
- File Type
Articles 31 - 60 of 3231
Full-Text Articles in Data Science
Human-Driven, Autonomous, Or Hybrid? The Optimal Fleet Configurations For Ride-Hailing Platforms, Wenjing Li, Yali Zhang, Jun Sun, Zhaojun Yang
Human-Driven, Autonomous, Or Hybrid? The Optimal Fleet Configurations For Ride-Hailing Platforms, Wenjing Li, Yali Zhang, Jun Sun, Zhaojun Yang
Information Systems Faculty Publications
The growing commercialization of autonomous vehicles (AVs) is reshaping consumer service preferences and prompting ride-hailing platforms to redesign fleet structures that accommodate the coexistence of human-driven vehicles (HVs) and AVs. This article develops a queueing game framework that incorporates vehicle heterogeneity and consumer preference differences to systematically compare three fleet configuration strategies: the pure HV (PHV) strategy (HVs only), the pure AV (PAV) strategy (AVs only), and the hybrid strategy (both HVs and AVs). The analysis highlights how consumer mismatch losses, AV operating costs, and service rates jointly shape equilibrium outcomes. Results show that when consumer mismatch losses are moderate, …
Data-Driven Characterization Of Counties In The Prison Industrial Complex Using Clustering Analysis, Riley N. Tuccio
Data-Driven Characterization Of Counties In The Prison Industrial Complex Using Clustering Analysis, Riley N. Tuccio
Capstone Projects
This project investigates the complex relationship between counties that house prisons in the United States and the rurality associated with them. The central research question explores how both county characteristics, such as variables corresponding to cost of living and demographics of a county, and prison characteristics, such as programming available to inmates and staffing levels, differ across the census-designated rural-urban distinctions. Furthermore, the study examines whether modern data science methods can more accurately define and distinguish these characteristics, providing a nuanced understanding of the Prison Industrial Complex (PIC) and its manifestation across various American communities. The motivation for this research …
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Cdt-1d Cnn Integration With Simpson-Sobolev Regularization For High-Frequency Options Trading: With Fem-Based Heston Option Pricing, Daniel M. Margolis, Johannes Tausch, Arthur K. Selender
Mathematics Theses and Dissertations
This dissertation presents a computational framework for high-frequency options trading that combines Cross-Data-Type 1-D Convolutional Neural Networks (CDT-1D CNN) with Simpson-Sobolev regularization for directional prediction, and finite element methods (FEM) for realistic option pricing during backtesting. The core innovation lies in developing a mathematically rigorous regularization approach that maintains the adaptability of modern deep learning while enabling accurate evaluation through stochastic volatility models. The primary contribution is the Simpson-Sobolev regularization scheme, which extends traditional Sobolev regularization by incorporating Simpson’s rule for numerical integration. This approach achieves higher-order accuracy in approximating the Sobolev norms that control function smoothness. Simpson’s rule attains …
Chi Meta-Project Ecosystem Overview - Spring 2026, David B. Smith
Chi Meta-Project Ecosystem Overview - Spring 2026, David B. Smith
Publications and Research
This paper offers a high-level account of the Center for Holistic Integration’s (CHI) meta-project ecosystem as visualized in the included system map. CHI provides an organizational structure framed around persistent meta-projects that support and extend individual initiatives across curriculum, scholarly and applied research, infrastructure, artistic production, AI development, cultural inquiry, and external partnerships. Rather than presenting the map as a static inventory of projects, the paper examines how its core domains function as living systems through which knowledge, tools, documentation, participants, and collaborations can accumulate over time. It also considers how CHI-mediated connectivity, institutional integration, and external funding allow the …
Unmasking Twitter Bots: An Applied Machine Learning Approach, Rayane El Raba’A, Layal Abu Daher
Unmasking Twitter Bots: An Applied Machine Learning Approach, Rayane El Raba’A, Layal Abu Daher
BAU Journal - Science and Technology
The rapid growth of social networks has led to increased challenges, such as fraud, cyberbullying, and the spread of automated accounts (bots). Detecting anomalies within these networks is essential to maintaining security and trust. This study explored machine learning algorithms: Random Forest, XGBoost, Support Vector Machine (SVM), and Logistic Regression for anomaly detection in social networks, specifically focusing on Twitter bot identification, By applying AI-driven data mining techniques to a dataset of 37,438 Twitter bot accounts dataset, the research evaluates the effectiveness of these models in detecting unusual patterns. XGBoost achieved the highest accuracy (84.9%), with an ROA_AUC of 0.87, …
Ai For Regression Analysis And More, Eli Snir
Ai For Regression Analysis And More, Eli Snir
Generative AI Teaching Activities
Students use Copilot and NotebookLM to create a dataset and develop statistical analyses including regression.
Integration Of Intraoperative Data In Interpretable Machine Learning Models To Predict Postoperative Aki In Noncardiac Surgery Patients, Justin Do, Karan H. Shah, Melissa Xu, Andrew Hyunwoo Kim, Vivaswat Suresh, Nidhir Guggilla, Michael Li, Rishi Kothari
Integration Of Intraoperative Data In Interpretable Machine Learning Models To Predict Postoperative Aki In Noncardiac Surgery Patients, Justin Do, Karan H. Shah, Melissa Xu, Andrew Hyunwoo Kim, Vivaswat Suresh, Nidhir Guggilla, Michael Li, Rishi Kothari
Department of Anesthesiology Faculty Papers
OBJECTIVES: We aimed to (1) quantify changes in discrimination when adding intraoperative data to preoperative data and (2) compare tabular machine learning with feature engineering against a time-aware LSTM-based model.
MATERIALS AND METHODS: Retrospective cohort of 46 204 adults undergoing 57 055 eligible noncardiac surgery in the INSPIRE database. We extracted 38 preoperative and 49 intraoperative variables; acute kidney injury (AKI) was defined by KDIGO serum creatinine criteria and modeled as stage 2/3 postoperative AKI. Models were trained on preoperative-only and combined pre- and intraoperative data. Intraoperative series were summarized using eight statistical features for tabular models or integrated directly …
Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer
Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer
CODEE Journal
Undergraduate instruction in ordinary differential equations (ODEs) is typically organized around the forward problem: finding solution trajectories when the governing equation and its parameters are known. In scientific practice, however, inverse problems are often more relevant, requiring unknown parameters to be inferred from noisy observations while assessing whether a proposed model is consistent with the data. We introduce PINNLab, an open-source MATLAB dashboard designed to help undergraduate students explore inverse modeling through physics-informed neural networks (PINNs). PINNLab presents PINNs as a complementary data-driven framework that connects differential equations, optimization, empirical data, and scientific machine learning. The instructional sequence is organized …
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
Mycelial Modeling: Teaching Biology Students Statistical Modeling With Mushrooms, Colette Wolf
University Honors Theses
This paper summarizes and describes the development of a set of learning materials that were created to educate students and professionals from other fields in statistical modeling techniques. These materials are primarily aimed at biology students, but are still intended to be useful for anyone who is interested in incorporating decision trees and random forest models into their personal research in the future. By directing the reader towards the JMP software, these materials navigate around the statistical knowledge base and coding implementation practices that otherwise would serve as a barrier to learning statistical modeling techniques, and instead focus on the …
Pitching Fwar Vs Bwar As Predictors Of Team Success In The Mlb Regular Season, Jason Lee
Pitching Fwar Vs Bwar As Predictors Of Team Success In The Mlb Regular Season, Jason Lee
University Honors Theses
Pitching Wins Above Replacement (WAR) is an area of sabermetrics capable of being used to predict regular season success in Major League Baseball. Fangraphs WAR (fWAR) and Baseball Reference WAR (bWAR) were used to construct regression models to predict regular season winning percentage, to build logistic models to establish a relationship between pitching WAR and the probability to win an individual regular season game, and to overlay density plots to consider WAR accumulation by pitching role and observe the difference of impact between starter and relief pitchers. This research finds that while pitching fWAR and pitching bWAR are both statistically …
A Novel, Embedding-Based Approach To Longitudinal Survey Data Imputation, Julia Rezvani
A Novel, Embedding-Based Approach To Longitudinal Survey Data Imputation, Julia Rezvani
University Honors Theses
Longitudinal surveys are ubiquitous in the social sciences as a means of tracking changes in behavior and opinions with time and identifying potential causal mechanisms. These surveys are frequently plagued by missing data and semantic drift, both of which limit their effectiveness and scientific utility. Imputation algorithms allow researchers to fill gaps in collected survey datasets, imperfectly reconstructing lost data. Although deep learning algorithms have been used in imputation to great success, approaches which simultaneously leverage the semantic and temporal structure of longitudinal surveys have not yet been developed. We propose a novel imputation architecture which is capable of leveraging …
A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson
A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson
CODEE Journal
In the age of data-driven decision making, ordinary differential equations (ODEs) remain a powerful and interpretable framework for modeling dynamic processes, especially when integrated with modern tools from statistical learning and data-driven dynamical systems. Yet, general undergraduate and graduate curricula do not typically address key opportunities in data-driven dynamical systems.
This first paper in a series focuses on the mathematical and methodological core of a professional development course first developed in the academic year 2025-2026 at a Primarily Undergraduate Institution, Purdue University Fort Wayne. The curriculum developed in this course emphasized how regression, regularization, and sparse identification can be used …
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Developing A Humpback Whale Vocalization Detector Using Machine Learning Models, Lucas Kantorowski
Master's Theses
Humpback whale songs are notoriously complex. Identification of humpback whale song units requires bioacousticians to tediously listen, analyze, and annotate collected sound data. Even sparse data requires listening to the entirety of the collected acoustic data. In this study, three hours of audio containing over one-thousand humpback whale song units was collected in Monterey Bay, California.
Prior studies have seen success using convolutional neural networks by performing image classification on hundreds of hours worth of spectrograms. Our study uses traditional machine learning models, as they are less computationally demanding, and require less data.
We use time splitting and Mel-frequency cepstrum …
Fairlinked: Data Fairification Tools For Materials Data Science, Van D. Tran, Brandon Lee, Ritika Lamba, Henry Dirks, Quynh D. Tran, Balashanmuga Priyan Rajamohan, Ozan Dernek, Laura S. Bruckman, Yinghui Wu, Erika I. Barcelos, Roger H. French
Fairlinked: Data Fairification Tools For Materials Data Science, Van D. Tran, Brandon Lee, Ritika Lamba, Henry Dirks, Quynh D. Tran, Balashanmuga Priyan Rajamohan, Ozan Dernek, Laura S. Bruckman, Yinghui Wu, Erika I. Barcelos, Roger H. French
Student Scholarship
FAIRLinked is a software package created to support the FAIRification of materials science data, ensuring proper alignment with FAIR principles: Findable, Accessible, Interoperable, and Reusable. It is built to be compatible with MDS-Onto, an ontology designed to capture the semantics of various types of materials data, enabling integration and sharing across different research workflows. The package is subdivided into three subpackages: InterfaceMDS, RDFTableConversion, and QBWorkflow. The first subpackage, InterfaceMDS allows users to search for terms using either string search or various filters, explore different domains and subdomains, and add terms to MDS-Onto. RDFTableConversion is used for serialization and deserialization of …
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Determining K Clusters In K-Means Clustering With The Crab Algorithm, Jasmine Kristine S. Cabrera
Master's Theses
Unsupervised clustering often faces the challenge of determining the correct number of clusters in the absence of a true target variable. Traditional methods such as the Elbow Method and the Silhouette Score can produce ambiguous results and rely on assumptions about cluster shape or separation. To address this, we created the Clustering Rivals and Buddies (CRAB) algorithm which evaluates clusters based on stability across multiple subsamples. CRAB uses pairwise classifications to identify points that consistently group together called “Buddies” and points that remain separated called “Rivals.” Applied with K-means, CRAB accurately recovers underlying cluster structures in both spherical and non-spherical …
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala
Student Theses
The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …
Can Generative Ai Make Farming Decisions? Current Status And Future Pathways: A Case Study In Row Crop Production With Chatgpt, Nipuna Chamara, Yufeng Ge, Joe Luck, Yu Pan, Saleh Taghvaeian, Cory Walters, Christopher Proctor, Daran Rudnick, Daren Redfearn
Can Generative Ai Make Farming Decisions? Current Status And Future Pathways: A Case Study In Row Crop Production With Chatgpt, Nipuna Chamara, Yufeng Ge, Joe Luck, Yu Pan, Saleh Taghvaeian, Cory Walters, Christopher Proctor, Daran Rudnick, Daren Redfearn
Department of Agricultural and Biological Systems Engineering: Faculty Publications
The agricultural decision-making process is experience-based, knowledge-dependent, time-sensitive, complex, and driven by historical data. Planting, fertilization, irrigation, and chemigation are key categories in farm decision-making, and currently there is no one-shot decision-support tool that covers all these activities. Generative Artificial Intelligence (AI) models are more advanced than traditional machine learning and deep learning models. These models have been trained on vast amounts of data from the internet, allowing them to accept unstructured data in various forms and generate human-like text, solutions to problems, and scenario predictions. Given this capability, we became interested in exploring the potential of generative AI in …
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Identifying Textual Predictors Of Early Termination In Clinical Trials In Medicine: An Explainable Machine-Learning Study, Rohan Ramnarain
Dissertations, Theses, and Capstone Projects
About one in five clinical trials in medicine ends early, wasting valuable resources and reducing the evidence available for developing life-saving medical treatments. This project uses a method called Trial2Vec, which is a self-supervised machine-learning method that converts clinical trial documents into dense numerical representations that capture their key design and clinical characteristics, to turn each proposed clinical trial’s written protocol into a compact numerical profile (a process referred to as embedding). These profiles are then paired with a predictive machine learning models to identify the words and phrases in the trial documents that can signal a higher risk of …
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin
Inductive Biases In Field-Level Cosmological Inference From Galaxy Catalogs, James O'Connor Baldwin
Dissertations, Theses, and Capstone Projects
We perform field-level likelihood-free inference of the matter density parameter Ωm from simulated galaxy catalogs using machine learning models with differing inductive biases. Using features extracted from hydrodynamic simulations in the CAMELS suite, we investigate how both observable choice and model architecture govern the extraction of cosmological information. We consider galaxy positions and line-of-sight peculiar velocities, both separately and in combination, and compare permutation-invariant Deep Sets, implemented with either standard multilayer perceptrons (MLPs) or Kolmogorov–Arnold Networks (KANs), to graph neural networks (GNNs) implemented with MLPs, which explicitly encode spatial relations. We evaluate inference performance under both in-distribution and out-of-distribution (OOD) …
Bridging Data Gaps In Retinal Imaging: From Structural Domain Adaptation To Topology-Aware Synthesis, Gözde Merve Demirci
Bridging Data Gaps In Retinal Imaging: From Structural Domain Adaptation To Topology-Aware Synthesis, Gözde Merve Demirci
Dissertations, Theses, and Capstone Projects
Comprehensive visualization of the retina is essential for diagnosing and monitoring blinding diseases such as Diabetic Retinopathy and Retinopathy of Prematurity (ROP), where pathological changes often extend beyond a single field of view. Despite significant advances in automated retinal image analysis, clinical deployment remains limited by two fundamental data gaps: a structural learning gap, arising from scarce expert annotations and poor generalization across imaging domains, and a spatial coverage gap, caused by the difficulty of acquiring multi-view retinal images in fragile populations. Although these challenges are often addressed independently, this dissertation argues that they are tightly coupled: accurate, …
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Crab: A Novel Clustering Score Using Clustering With Rivals And Buddies For Unsupervised Learning, Allen Choi
Master's Theses
Unsupervised clustering algorithms today are used across a wide variety of fields such as biology, engineering, and industry in order to classify observations into groups where labels are not provided. This can provide important latent information regarding the observations within groups, as well as insight regarding the groups themselves. In order to judge the optimal number of clusters for an unsupervised clustering algorithm, many methods exist such as the Elbow Method and Silhouette Score; however, these methods come with drawbacks and are not necessarily flexible across many unsupervised methods. We present a novel clustering score framework relying on a resampling-based …
Edge Co-Occurrence Regularization For Node Classification, Kadir Altunel
Edge Co-Occurrence Regularization For Node Classification, Kadir Altunel
Theses
We propose a simple yet effective regularization technique for node classification on graphs that leverages edge-based label co-occurrence patterns. We first train an MLP on node features to produce class probability distributions, then compute a fixed penalty matrix from edge-based co-occurrence statistics of these predictions. This penalty matrix, which captures unlikely class combinations on connected nodes, is then used to regularize GNN training without further updates. We evaluate this approach across multiple homophilic datasets (Cora, CiteSeer, PubMed, ogbn-arxiv) and heterophilic benchmarks (Chameleon, Squirrel, Actor, Roman-Empire) using three GNN architectures: GCN, GraphSAGE, and H2GCN. Results show consistent improvements on homophilic graphs, …
A Generative Ai-Driven Computational Framework For Industry-Scale Discovery Of Novel Battery Materials, Joy Datta
Dissertations
The growing demand for sustainable, high-energy-density electrochemical storage has motivated the exploration of multivalent-ion batteries based on earth-abundant elements such as aluminum, calcium, magnesium, and zinc. While multivalent charge carriers offer higher theoretical energy density than lithium, their practical deployment is hindered by sluggish ion transport, strong ion-host interactions, and structural degradation of electrode materials. Identifying host materials that can reversibly accommodate multivalent ions while maintaining structural integrity remains a fundamental challenge. The dissertation develops a scalable, end-to-end computational framework that integrates density functional theory (DFT), machine learning (ML), and generative artificial intelligence (GenAI) to accelerate the discovery of next-generation …
Toward Learning-Based Reconstruction And Part Decomposition Of Man-Made 3d Geometry: Neural Implicit Representations And Scalable Supervision, Shen Fan
Dissertations
Digital three-dimensional (3D) models are central to engineering design, analysis, and manufacturing, but learning pipelines for man-made geometry often operate on sampled carriers that do not preserve all of the structure present in exact CAD representations. This dissertation studies learning-based reconstruction and part decomposition for structured man-made 3D geometry, from general object benchmarks to CAD-derived datasets, with a focus on neural implicit representations trained from signed-distance samples, point clouds, and tessellated meshes. The goal is to make these models more accurate, more part-aware, and more consistently supervised.
First, signed distance function (SDF) reconstruction with implicit neural representations is improved through …
Transportation Deserts And Structural Mobility Access In New York City, John Cruz
Transportation Deserts And Structural Mobility Access In New York City, John Cruz
Student Theses
This study examines structural mobility access across New York City census tracts by constructing a tract-level Mobility Access Index (MAI) that integrates employment accessibility, hospital accessibility, and first-mile subway walking burden using MTA GTFS transit data, NYC Taxi and Limousine Commission trip records, and US Census American Community Survey demographic estimates. Three OLS regression models, supplemented by Lasso, Elastic Net, and Random Forest specifications, test whether structural access gaps are associated with short-distance connector trip intensity and per-worker connector cost burden. Results show that MAI varies substantially across tracts, with high access concentrated in Manhattan and along major subway corridors. …
Forecasting Interborough Express Ridership Using Network-Based Station Typologies And Direct Demand Models, Fomba Kassoh
Forecasting Interborough Express Ridership Using Network-Based Station Typologies And Direct Demand Models, Fomba Kassoh
Student Theses
Forecasting ridership for new transit infrastructure is difficult in the absence of observed outcomes, particularly under domain shift between an existing system and a proposed corridor. This study develops a station-level direct demand modeling (DDM) framework to forecast average weekday ridership for the proposed Interborough Express (IBX) in New York City — a 14-mile circumferential rapid transit corridor connecting Brooklyn and Queens. The approach pairs unsupervised learning with supervised estimation in a common feature space defined by transit service, accessibility, and built-environment characteristics. K-means clustering identifies latent station typologies (node–place regimes), and IBX stations are projected into this topology to …
Predicting The Outcome Of Ischemic Hepatitis With Real-Patient Data Using Machine Learning Tools, Christiana Beard, Madison Utterback, Olcay Akman, Priya Kohli, William M. Lee, Aditi Ghosh
Predicting The Outcome Of Ischemic Hepatitis With Real-Patient Data Using Machine Learning Tools, Christiana Beard, Madison Utterback, Olcay Akman, Priya Kohli, William M. Lee, Aditi Ghosh
Spora: A Journal of Biomathematics
Ischemic hepatitis (IH) results from shock-related conditions that impair oxygenated blood flow to the liver, causing hepatocyte death. Diagnosis relies largely on clinical history due to the absence of specific diagnostic tests and limited ability to predict outcomes. This study applies machine learning methods to real-world IH patient data to improve outcome prediction. Biomedical indicators analyzed include creatinine, international normalized ratio (INR), aspartate aminotransferase (AST), alanine transaminase (ALT), and bilirubin. Data were collected from multiple U.S. centers through the Acute Liver Failure Study Group (ALFSG), a multicenter network focused on this rare condition. We implemented logistic regression, regression tree methods …
Evaluating Soil Health And Crop Yield In Louisiana Agricultural Systems: Impacts Of Best Management Practices And Prediction Models, Hector J. Mendoza Lagos
Evaluating Soil Health And Crop Yield In Louisiana Agricultural Systems: Impacts Of Best Management Practices And Prediction Models, Hector J. Mendoza Lagos
LSU Doctoral Dissertations
The adoption of conservation management practices is critical for improving soil health, enhancing nutrient use efficiency, and sustaining crop productivity in row crop systems in Louisiana. This study evaluated the role of conservation agronomic practices, soil biochemical indicators, and machine learning predictive models to improve soil nutrient dynamics, soil health indicators, microbial communities (MC), and crop productivity on a corn (Zea mays L.) research plot scale and in a commercial forty-hectare cotton (Gassypium hirsutum L.)-corn-soybean (Glycine max L.) rotation system in northeast Louisiana. The objectives of the study were to evaluate soil nutrient dynamics and MCs under …
Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama
Measuring Stock Market Inefficiency Using A Multilayer Composite Efficiency Index: A Case Of The Egyptian Exchange, Patrick K. Owido, Hiroki Sayama
Northeast Journal of Complex Systems (NEJCS)
Financial markets play a critical role in resource allocation. Their performance depends on the decisions of millions of independent investors constantly reacting to one another. Their informational efficiency remains a subject of debate across economic systems. When informational efficiency is present at the weak form, historical price information should not consistently predict future returns. Several empirical tests of this hypothesis often focus on the behavior of aggregate market indices, and use individual efficiency proxies such as autocorrelation, GARCH-type volatility, or entropy-based measures to measure efficiency. This has often yielded mixed results, particularly in emerging markets. Here we show that testing …
Material Costs, Karima Weinman
Material Costs, Karima Weinman
Masters Theses
This thesis investigates how migration fatality and disappearance data can be reinterpreted through material craft to create a more reflective encounter with information. Working with the Missing Migrants Project's dataset, this project asks how design can communicate dimensions of human loss that conventional data visualization cannot reach.
The work situates contemporary border violence within a longer colonial history, arguing that the logics of surveillance and quantification that structured European imperial expansion persist in the databases that govern mobility in the Mediterranean today.
Terrazzo is a 15th-century Venetian flooring technique built from discarded fragments bound together into a unified surface. This …