Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 391 - 420 of 3232

Full-Text Articles in Data Science

Historical Perspectives In Volatility Forecasting Methods With Machine Learning, Zhiang Qiu, Clemens Kownatzki, Fabien Scalzo, Eun Sang Cha May 2025

Historical Perspectives In Volatility Forecasting Methods With Machine Learning, Zhiang Qiu, Clemens Kownatzki, Fabien Scalzo, Eun Sang Cha

All Faculty Open Access Publications

Volatility forecasting for financial institutions plays a pivotal role across a wide range of domains, such as risk management, option pricing, and market making. For instance, banks can incorporate volatility forecasts into stress testing frameworks to ensure they are holding sufficient capital during extreme market conditions. However, volatility forecasting is challenging because volatility can only be estimated, and different factors influence volatility, ranging from macroeconomic indicators to investor sentiments. While recent works show promising advances in machine learning and artificial intelligence for volatility forecasting, a comprehensive assessment of current statistical and learning-based methods is lacking. Thus, this paper aims to …


Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao May 2025

Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao

Faculty, Staff and Student Publications

Bone metastasis is a major cause of cancer death; however, the epigenetic determinants driving this process remain elusive. Here, we report that histone methyltransferase ASH1L is genetically amplified and is required for bone metastasis in men with prostate cancer. ASH1L rewires histone methylations and cooperates with HIF-1α to induce pro-metastatic transcriptome in invading cancer cells, resulting in monocyte differentiation into lipid-associated macrophage (LA-TAM) and enhancing their pro-tumoral phenotype in the metastatic bone niche. We identified IGF-2 as a direct target of ASH1L/HIF-1α and mediates LA-TAMs' differentiation and phenotypic changes by reprogramming oxidative phosphorylation. Pharmacologic inhibition of the ASH1L-HIF-1α-macrophages axis elicits …


Challenging $\Lambda$Cdm: Unraveling Cosmic Distances, Dark Sector Phenomenology, And Alternative Primordial B-Mode Sources, Kylar L. Greene May 2025

Challenging $\Lambda$Cdm: Unraveling Cosmic Distances, Dark Sector Phenomenology, And Alternative Primordial B-Mode Sources, Kylar L. Greene

Physics & Astronomy ETDs

The dominant Lambda Cold Dark Matter (LCDM) cosmological model, while remarkably successful, increasingly shows signs that it may not fully describe our Universe, as persistent tensions in expansion rates and structure formation remain unresolved. In this thesis, I challenge the LCDM paradigm using novel theoretical frameworks combined with rigorous numerical analyses. I demonstrate that the expansion-rate tension fundamentally reflects underlying distance disagreements, and that the Thomson scattering rate strongly restricts higher pre-recombination expansion rates without additional physics. Further, I present a novel cosmological model using a mirror dark sector and varying fundamental constants, revealing an observational degeneracy allowing significantly higher …


Data Encoding, Compilation, And Algorithms For Quantum Machine Learning, Aviraj Sinha May 2025

Data Encoding, Compilation, And Algorithms For Quantum Machine Learning, Aviraj Sinha

Computer Science and Engineering Theses and Dissertations

Quantum computing enables new approaches to data processing, especially in quantum machine learning. Unlike classical systems, quantum data must be synthesized through operations and can exist in superposition. Encoding choices affect efficiency, noise resilience, and trainability—key factors in quantum machine learning models. This dissertation enhances quantum data encodings by extending quantum read-only memory (QROM) beyond binary representations, improving efficiency and parallelism. It introduces new compilation methods for quantum random number generators (QRNGs), supporting non-parametric distributions for post-quantum cryptography. Additionally, it explores Cayley graph-based encodings to extract spectral features for quantum machine learning.


Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani May 2025

Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani

Computer Science and Engineering Theses and Dissertations

The rapid expansion of scientific literature has intensified the challenge of identifying relevant citations, particularly for newly published or under-cited papers. Traditional citation recommendation systems typically model static relationships or respond to past citation activity, offering limited predictive power for emerging works. In response, this thesis presents a temporal modeling framework for citation recommendation that anticipates future scholarly relevance by forecasting the latent representations of academic papers.

Building on prior work that utilized Temporal Graph Networks (TGNs) to model dynamic citation flows, we propose Graph-Time, a hybrid architecture that integrates a Graph Transformer with a GRU-based time series predictor. The …


A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer May 2025

A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer

Computer Science ETDs

Modern drug discovery and chemical biology research relies heavily on analyzing bioassay data. One of the many challenges in bioassay data analysis is identifying false trails, i.e., chemical compounds which initially appear to have desirable activity but are found to be problematic upon further investigation. Badapple (the BioAssay-Data Associative Promiscuity Pattern Learning Engine) was created over ten years ago to help researchers identify promiscuous compounds and thus avoid a common source of these false trails. Through an effort involving software engineering, cheminformatics, and biomedical data science we have developed Badapple 2.0, which incorporates updated assay records and expanded data semantics. …


Incorporating Latent Survival Trajectories And Covariate Heterogeneity In Time-To-Event Data Analysis: A Joint Mixture Model Approach, Fu-Wen Liang, Wenyaw Chan, Michael D Swartz, Bouthaina S Dabaja May 2025

Incorporating Latent Survival Trajectories And Covariate Heterogeneity In Time-To-Event Data Analysis: A Joint Mixture Model Approach, Fu-Wen Liang, Wenyaw Chan, Michael D Swartz, Bouthaina S Dabaja

Faculty, Staff and Student Publications

Background: Finite mixture models have been recently applied in time-to-event data to identify subgroups with distinct hazard functions, yet they often assume differing covariate effects on failure times across latent classes but homogeneous covariate distributions. This study aimed to develop a method for analyzing time-to-event data while accounting for unobserved heterogeneity within a mixture modeling framework.

Methods: A joint model was developed to incorporate latent survival trajectories and observed information for the joint analysis of time-to-event outcomes, correlated discrete and continuous covariates, and a latent class variable. It assumed covariate effects on survival times and covariate distributions vary across latent …


Gnns For Network Classification In Single Cell Rna Sequencing Data, Reid C. Sewell May 2025

Gnns For Network Classification In Single Cell Rna Sequencing Data, Reid C. Sewell

Capstone Projects

A common technique when investigating a disease is to profile gene expression, as this gives unique insights into the functions of a cell. Gene expression data gathered from single cell RNA sequencing can be encoded into a gene co-expression network, which is a graph of potential relationships between different genes. One method for interpreting data encoded as a graph is to use a graph neural network, or GNN. This project designs and implements a GNN architecture to accomplish classification tasks on graph data. Then, given a dataset of gene co-expression networks made from multiple single cell RNA sequencing studies, the …


Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow May 2025

Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow

Capstone Projects

This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …


Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi May 2025

Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi

Open Educational Resources

This open-access machine learning course is a comprehensive 15-week curriculum developed and published on GitHub with full Google Colab compatibility. It combines theoretical concepts with hands-on Python coding, real-world datasets, and structured projects covering regression, classification, clustering, deep learning, transformers, and multimodal AI. The course is designed for students, educators, and researchers interested in applied machine learning, including biomedical applications. It includes explainable AI components and ethical discussions to align with modern AI standards. The course is maintained by BioMind AI Lab at CUNY.


Optimized Student Grouping For Enhanced Classroom Performance, Kathryn E. Reardon May 2025

Optimized Student Grouping For Enhanced Classroom Performance, Kathryn E. Reardon

Honors Theses

Effective grouping methods enhance classroom collaboration and allow for a student-centered teaching approach; however, traditional grouping methods are time-consuming, subjective, and can create inconsistent group dynamics. This project addresses these challenges by employing a data-driven approach to optimize student groups based on academic performance, behavior, attendance, language barriers, and teacher preferences. The minimum viable product is a web application with an algorithm-driven system to group students and a database storage for group results. During the initiation phase, a problem was defined with a proposed solution. During the planning phase, potential design choices and grouping methods were researched and assessed. During …


Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson May 2025

Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson

Electronic Theses and Dissertations

The beef industry plays a vital role in global agriculture, with carcass quality and consumer preference being key determinants of market success. This thesis examines predictive modeling techniques for estimating the Total Score of beef carcasses, a composite measure representing yield and quality, primarily used by the Nebraska Cattlemen Association. Using data from the Nebraska Cattlemen’s Foundation Retail Value Steer Challenge (2000–2023), the study compares the performance of First Order Multiple Linear Regression (MLR) with three machine learning techniques: K-Nearest Neighbors (KNN), Random Forest, and Gradient Boosting Machine (GBM).

The analysis focuses on six key predictors: Hot Carcass Weight, Back …


Performance Of Lasso And Ridge Regression For Variable Selection In Genome-Wide Association Studies Of Maize Flowering Time, Dipok Deb May 2025

Performance Of Lasso And Ridge Regression For Variable Selection In Genome-Wide Association Studies Of Maize Flowering Time, Dipok Deb

Data Science and Data Mining

Genome-Wide Association Studies (GWAS) are instrumental in identifying genetic variants linked to complex traits, providing valuable insights into trait heritability and biological mechanisms. This study applies GWAS to investigate flowering time in maize, a critical adaptive trait, using a diverse dataset of 5,000 recombinant inbred lines across eight environments. Traditional GWAS methods often encounter challenges in high-dimensional datasets due to the presence of multiple small-effect genetic loci. To address this, we compared two penalized regression methods—LASSO and Ridge regression—to perform variable selection and regression analysis within a GWAS framework. LASSO effectively reduced the number of predictors by selecting the most …


Data Science Job Salary Prediction Using Linear Regression, Dipok Deb May 2025

Data Science Job Salary Prediction Using Linear Regression, Dipok Deb

Data Science and Data Mining

In the evolving landscape of data science, accurate salary prediction plays a crucial role in shaping career expectations, informing educational strategies, and guiding organizational hiring decisions. This study investigates the key factors influencing entry-level data science salaries in the United States by applying a multiple linear regression model to a recent dataset spanning from 2020 to 2024. Through data preprocessing, transformation, and diagnostic evaluation, we identify how job roles, experience levels, employment types, work arrangements, residency status, and company size impact compensation. Despite challenges such as outliers, heteroscedasticity, and non-normal residuals, model refinements like the Box-Cox transformation and variable selection …


Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville May 2025

Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville

Student Scholar Symposium Abstracts and Posters

For my Introduction to Statistics Class, I have been tasked with collecting unique, personal data to give insight into my daily routine. I decided to record nine different outcomes (two qualitative and seven quantitative). On February 6, 2025, I began with a blank Excel sheet, and so far, I have 57 full days of data collected. I will continue monitoring my findings for the remainder of the Spring 2025 Semester. Per my project instructions, I must include tables and graphs for my qualitative and quantitative outcomes. So far, I have collected daily quantitative data on my screen time (Instagram and …


Two-Sample Bi-Directional Causality Between Two Traits With Some Invalid Ivs In Both Directions Using Gwas Summary Statistics, Siyi Chen May 2025

Two-Sample Bi-Directional Causality Between Two Traits With Some Invalid Ivs In Both Directions Using Gwas Summary Statistics, Siyi Chen

School of Public Health Faculty Publications

Mendelian randomization (MR) is a widely used method for assessing causal relationships between risk factors and outcomes using genetic variants as instrumental variables (IVs). While traditional MR assumes uni-directional causality, bi-directional MR aims to identify the true causal direction. In uni-directional MR, invalid IVs due to pleiotropy can violate assumptions and introduce biases. In bi-directional MR, traditional MR can be performed separately for each direction, but the presence of invalid IVs poses even greater challenges. We introduce a new bi-directional MR method incorporating stepwise selection (Bidir-SW) designed to address these challenges. Our approach leverages public genome-wide association study (GWAS) datasets …


Handwritten Digit Recognition Using Machine Learning, Dipok Deb May 2025

Handwritten Digit Recognition Using Machine Learning, Dipok Deb

Data Science and Data Mining

Handwritten Digit Recognition (HDR) remains a fundamental benchmark in pattern recognition and machine learning due to its practical applications and inherent classification challenges posed by diverse handwriting styles. This study investigates and compares two classical statistical classifiers—Gaussian Naive Bayes (GNB) and Linear Discriminant Analysis (LDA)—to recognize the digits from the MNIST dataset. Both models assume underlying normality in feature distributions and offer computational efficiency, making them suitable for high-dimensional input such as image pixels. Using 60,000 training and 10,000 test samples, we evaluate model performance through accuracy, precision, recall, F1 score, and confusion matrices. The results reveal that while GNB …


Clustering Dataset Using K-Mean Clustering, Dipok Deb May 2025

Clustering Dataset Using K-Mean Clustering, Dipok Deb

Data Science and Data Mining

Clustering is a fundamental technique in unsupervised machine learning, widely applied in various domains such as pattern recognition, data segmentation, and anomaly detection. This study evaluates the performance of the K-Means clustering algorithm on multiple benchmark datasets, including low-dimensional, high-dimensional, and imbalanced datasets. The clustering results are assessed using four key evaluation metrics: Mean Squared Error (MSE), Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), and Silhouette Score. Experimental results demonstrate that K-Means performs effectively on datasets with well-separated clusters, particularly in high-dimensional spaces, where it achieves near-perfect clustering accuracy. However, its performance deteriorates in datasets with overlapping clusters and …


Dynamic Approaches To Missing Data In Healthcare: Evaluating Ensemble Models, Feature Selection, And Meta-Features, Dylan Dominguez Sulca May 2025

Dynamic Approaches To Missing Data In Healthcare: Evaluating Ensemble Models, Feature Selection, And Meta-Features, Dylan Dominguez Sulca

Theses and Dissertations

Missing data is pervasive in healthcare, where incomplete observations commonly arise from patient dropout, sensor failures, or privacy constraints. This research presents an investigation into handling such data, focusing on (1) Missingness-Aware Dynamic Ensemble Weighting (MDEW), (2) feature selection under varying missing rates, (3) autoencoder-based imputation (ODAE), and (4) a meta-feature analysis guiding pipeline selection. We evaluate our experiments on four diverse datasets, Cleveland Heart Disease, Diabetic Retinopathy, Breast Cancer Wisconsin, EEG Eye State. Our research shows that MDEW adaptively selects imputer classifier pipelines, outperforming single model and uniform averaging baselines at moderate to high missingness 10% to 50%. Filter …


Achieving Professional, Effective Graphics Efficiently With Available Clemson Resources, Stacie P. Powell, Brad Landon Walters, Yaswanth Mulakala May 2025

Achieving Professional, Effective Graphics Efficiently With Available Clemson Resources, Stacie P. Powell, Brad Landon Walters, Yaswanth Mulakala

Presentations

No abstract provided.


A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster May 2025

A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster

Undergraduate Honors Theses

Machine learning is a method that employs statistical algorithms to identify patterns and make predictions from data. This study applies machine learning techniques to analyze data from Major League Baseball (MLB) teams between 1998 and 2024, with the goal of determining which factors strongly influence a team's likelihood of reaching the postseason and in accurately predicting the teams that do and do not qualify for the postseason. Data exploration and unsupervised machine learning methods such as clustering were used to identify underlying patterns in team performance metrics and determine potential significant contributors to team success. Many different supervised learning methods …


The Little Diagram That Could: Geometric Properties And Statistical Applications Of Persistence Diagrams In Topological Data Analysis, Eugene Kler May 2025

The Little Diagram That Could: Geometric Properties And Statistical Applications Of Persistence Diagrams In Topological Data Analysis, Eugene Kler

McKelvey School of Engineering Graduate Student Theses & Dissertations

Topological Data Analysis (TDA) is a collection of techniques for data analysis that leverages topological invariants of spaces formed from data points. These methods excel at extracting useful information from noisy or sparse data, making them attractive to many mathematicians, statisticians, and scientists. In this thesis, we explore TDA on three fronts: algebraic foundations, statistical applications, and metric properties. Throughout, the central object of study is the Persistence Diagram (PD), a summary of the changes in homology that occur as one builds simplicial complexes from the data by increasing a parameter.


Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran May 2025

Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran

Data Science Undergraduate Honors Theses

Early and accurate wildfire detection is critical for minimizing environmental damage and ensuring a timely response. However, existing satellite-based wildfire datasets suffer from limitations such as coarse ground truth, poor spectral coverage, and class imbalance, which hinder progress in developing robust segmentation models. In this paper, we introduce Land8Fire, a new large-scale wildfire segmentation dataset composed of over 20,000 multispectral image patches derived from Landsat 8 and manually annotated for high-quality fire masks. Building on the ActiveFire dataset, Land8Fire improves ground truth reliability and offers predefined splits for consistent benchmarking. We evaluate a range of state-of-the-art convolutional and transformer-based models, …


Advancing Drug-Drug Interaction Prediction Using Multi-Modal Feature Integration With Graph Neural Networks, Ernest C. Chianumba May 2025

Advancing Drug-Drug Interaction Prediction Using Multi-Modal Feature Integration With Graph Neural Networks, Ernest C. Chianumba

Theses, Dissertations and Culminating Projects

Pharmaceutical treatments are essential for managing medical conditions, but drug-drug interactions (DDIs) pose significant risks to patient safety and healthcare outcomes. This research integrates Knowledge Graphs and Graph Neural Networks to predict DDIs by exploring complex drug relationships. We construct a comprehensive knowledge graph using DrugBank data (1,000 drugs, 155,774 interactions) enriched with molecular features from PubChem. Our methodology introduces a novel multi-modal approach by integrating transformer-based embeddings (ChemBERTa, SPECTER, and SBERT) to create 1152-dimensional feature vectors that capture structural, biomedical literature, and semantic properties of drugs. Formulating DDI prediction as a link prediction task, we compare three Graph Neural …


Don't Let Lead Lead On Environmental Justice: A Simulative Approach To Lead Remediation In The Big Data Era, Charles C. Knoble Ii May 2025

Don't Let Lead Lead On Environmental Justice: A Simulative Approach To Lead Remediation In The Big Data Era, Charles C. Knoble Ii

Theses, Dissertations and Culminating Projects

Environmental justice, as both a movement and a theoretical construct, continues to evolve in response to shifting societal, environmental, and technological conditions. This dissertation investigates the integration of big data, such as social media, remote sensing imagery, and internet search frequencies, into the identification, analysis, and remediation of environmental injustices. Framing environmental justice through the lenses of distributive and data justice, the project explores both the promises and pitfalls of using emergent data sources to enhance the spatial and temporal precision of environmental equity investigations. Through a combination of systematic literature review, spatial analysis, system dynamics simulation, and policy evaluation, …


From Seasonality To Causality: Understanding Urban Water Usage Using Statistical And Machine Learning Models, Kelsey Hawkins May 2025

From Seasonality To Causality: Understanding Urban Water Usage Using Statistical And Machine Learning Models, Kelsey Hawkins

Electrical Engineering and Computer Science (MS) Theses

This study examines the relationship between climate conditions and residential water usage, focusing on how seasonal and environmental changes influence water consumption. Utilizing data from over 100,000 households across three micro-climate zones for over a five-year period, we apply statistical analysis and machine learning techniques to assess the impact of temperature, precipitation, evapotranspiration, and location on water usage. By integrating climate and billing data, this research provides a data-driven approach on water usage behaviors in Irvine, CA, in collaboration with Irvine Ranch Water District (IRWD).

Our analysis utilizes time series modeling, including a Seasonal Autoregressive Integrated Moving Average (SARIMA) and …


An Analysis Of Bias Towards Women In Large Language Models Using Likert Scale Evaluations, Sarah T. Fieck May 2025

An Analysis Of Bias Towards Women In Large Language Models Using Likert Scale Evaluations, Sarah T. Fieck

Electrical Engineering and Computer Science (MS) Theses

Closed-source large language models (LLMs) developed by large technology companies continue to grow in popularity. However, ethical conversations surrounding the safety of model outputs have been a prominent topic of discussion. This project aims to assess three leading closed-source LLMs: OpenAI’s ChatGPT, Google’s Gemini, and Anthropic’s Claude, to analyze how their outputs perform when treated as a subject of several psychological evaluation scales measuring biased behaviors against women. The Ambivalent Sexism Index, Modern Sexism Scale, and Belief in Sexism Shift evaluations were used to get descriptions of how the LLMs respond to traditional and modern prompts involving sexism and gender …


A Narrative-Focused Machine Learning Approach To Predicting Feature Film Success, Arisa T. Trombley May 2025

A Narrative-Focused Machine Learning Approach To Predicting Feature Film Success, Arisa T. Trombley

Electrical Engineering and Computer Science (MS) Theses

For decades, the field of film production has been driven by marketability, and it has relied on gut feelings and subjectivity to produce feature films. The analysis of the relationship between a screenplay’s narrative and a film’s success has been widely overlooked due to the challenges involved in data acquisition and complexity. This study investigates the predictive power of narrative structure on film success and aims to build evidence for hypothesized narrative principles. The results suggest that narrative structural elements exhibit moderate predictive power, with strong support for the alignment of the 2nd act crucial moments and the 2nd act …


Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar May 2025

Enhancing The University Of Arkansas' Operations Through Data Science, Aura L. Pinto-Avelar

Data Science Undergraduate Honors Theses

As universities navigate financial constraints and resource allocation challenges, data driven financial analysis has become increasingly important. Universities employ various methods to assess financial efficiency, predict future expenditures, and optimize student credit hour distribution. However, the approaches to financial analysis vary widely, with some institutions leveraging advanced predictive modeling and business intelligence tools, while others rely on traditional budgeting techniques and manual forecasting.

This thesis examines how the University of Arkansas' (“Uark”) financial analysis methods compare to those of other institutions and alternative data-driven approaches. Using four years of financial and student credit hour data, this study evaluates cost trends …


Optimizing Fire Station Placement In Sugar Land, Tx: A Socioeconomic Risk-Based Approach, Alicia Gallemore May 2025

Optimizing Fire Station Placement In Sugar Land, Tx: A Socioeconomic Risk-Based Approach, Alicia Gallemore

Data Science Undergraduate Honors Theses

Fire station placement has a critical role in emergency response efficiency and community safety. Traditional optimization models focus on mainly the minimization of response times and the maximization of coverage. However, this approach may overlook potential socioeconomic disparities that can influence emergency demand. This study seeks to expand upon the existing project of zoning a fire station in Sugar Land, TX, by integrating spatial road network analysis and publicly available census data—including population density, median household income, and age-based vulnerability—into a Maximal Coverage Location Problem (MCLP) framework. Using a road network-based travel time with realistic constraints, the goal is to …