Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 151 - 180 of 504

Full-Text Articles in Data Science

Surface Characterization Of Asian Lacquers Using Surface Metrology And Data Science: Introducing The Roughness Spectrum, Ravines Patrick, H. David Sheets, Marianne Webb, Joy Mazurek, Michael R. Schilling, Herant Khanjian May 2025

Surface Characterization Of Asian Lacquers Using Surface Metrology And Data Science: Introducing The Roughness Spectrum, Ravines Patrick, H. David Sheets, Marianne Webb, Joy Mazurek, Michael R. Schilling, Herant Khanjian

Computer and Data Science Faculty Publications

No abstract provided.


Protein Word Detection Using Text Segmentation Techniques, Ashish V. Tendulkar, Sutanu Chakraborti, Ganesh Devi May 2025

Protein Word Detection Using Text Segmentation Techniques, Ashish V. Tendulkar, Sutanu Chakraborti, Ganesh Devi

Journal of Global Awareness

Literature in Molecular Biology is abundant with linguistic metaphors. There has been works in the past that attempt to draw parallels between linguistics and biology, driven by the fundamental premise that proteins have a language of their own. Since word detection is crucial to the decipherment of any unknown language, we attempt to establish a problem mapping from natural language text to protein sequences at the level of words. Towards this end, we explore the use of an unsupervised text segmentation algorithm for the task of extracting "biological words” from protein sequences. We demonstrate the effectiveness of using domain knowledge …


Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi May 2025

Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi

Open Educational Resources

Data analysis using standard statistical methods and relevant computer software. Emphasis on real-world data, interpretation, and misinterpretation of computer output.

This syllabus contains open source notebook about data analysis content.


Dynamate: Leveraging Ai-Agents For Customized Research Workflows, Orlando A. Mendible-Barreto, Misael Díaz-Maldonado, Fernando J. Carmona Esteva, J. Emmanuel Torres, Ubaldo M. Córdova-Figueroa, Yamil J. Colón May 2025

Dynamate: Leveraging Ai-Agents For Customized Research Workflows, Orlando A. Mendible-Barreto, Misael Díaz-Maldonado, Fernando J. Carmona Esteva, J. Emmanuel Torres, Ubaldo M. Córdova-Figueroa, Yamil J. Colón

Computer and Data Science Faculty Publications

No abstract provided.


A Complete Transfer Learning-Based Pipeline For Discriminating Between Select Pathogenic Yeasts From Microscopy Photographs, Ryan A. Parker, Danielle S. Hannagan, Jan H. Strydom, Christopher J. Boon, Jessica Fussell, Chelbie A. Mitchell, Katie L. Moerschel, Aura G. Valter-Franco, Christopher Cornelison May 2025

A Complete Transfer Learning-Based Pipeline For Discriminating Between Select Pathogenic Yeasts From Microscopy Photographs, Ryan A. Parker, Danielle S. Hannagan, Jan H. Strydom, Christopher J. Boon, Jessica Fussell, Chelbie A. Mitchell, Katie L. Moerschel, Aura G. Valter-Franco, Christopher Cornelison

Faculty Articles

Pathogenic yeasts are an increasing concern in healthcare, with species like Candida auris often displaying drug resistance and causing high mortality in immunocompromised patients. The need for rapid and accessible diagnostic methods for accurate yeast identification is critical, especially in resource-limited settings. This study presents a convolutional neural network (CNN)-based approach for classifying pathogenic yeast species from microscopy images. Using transfer learning, we trained the model to identify six yeast species from simple micrographs, achieving high classification accuracy (93.91% at the patch level, 99.09% at the whole image level) and low misclassification rates across species, with the best performing model. …


Historical Perspectives In Volatility Forecasting Methods With Machine Learning, Zhiang Qiu, Clemens Kownatzki, Fabien Scalzo, Eun Sang Cha May 2025

Historical Perspectives In Volatility Forecasting Methods With Machine Learning, Zhiang Qiu, Clemens Kownatzki, Fabien Scalzo, Eun Sang Cha

All Faculty Open Access Publications

Volatility forecasting for financial institutions plays a pivotal role across a wide range of domains, such as risk management, option pricing, and market making. For instance, banks can incorporate volatility forecasts into stress testing frameworks to ensure they are holding sufficient capital during extreme market conditions. However, volatility forecasting is challenging because volatility can only be estimated, and different factors influence volatility, ranging from macroeconomic indicators to investor sentiments. While recent works show promising advances in machine learning and artificial intelligence for volatility forecasting, a comprehensive assessment of current statistical and learning-based methods is lacking. Thus, this paper aims to …


Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao May 2025

Histone Methyltransferase Ash1l Primes Metastases And Metabolic Reprogramming Of Macrophages In The Bone Niche, Chenling Meng, Kevin Lin, Wei Shi, Hongqi Teng, Xinhai Wan, Anna Debruine, Yin Wang, Xin Liang, Javier Leo, Feiyu Chen, Qianlin Gu, Jie Zhang, Vivien Van, Kiersten L Maldonado, Boyi Gan, Li Ma, Yue Lu, Di Zhao

Faculty, Staff and Student Publications

Bone metastasis is a major cause of cancer death; however, the epigenetic determinants driving this process remain elusive. Here, we report that histone methyltransferase ASH1L is genetically amplified and is required for bone metastasis in men with prostate cancer. ASH1L rewires histone methylations and cooperates with HIF-1α to induce pro-metastatic transcriptome in invading cancer cells, resulting in monocyte differentiation into lipid-associated macrophage (LA-TAM) and enhancing their pro-tumoral phenotype in the metastatic bone niche. We identified IGF-2 as a direct target of ASH1L/HIF-1α and mediates LA-TAMs' differentiation and phenotypic changes by reprogramming oxidative phosphorylation. Pharmacologic inhibition of the ASH1L-HIF-1α-macrophages axis elicits …


Challenging $\Lambda$Cdm: Unraveling Cosmic Distances, Dark Sector Phenomenology, And Alternative Primordial B-Mode Sources, Kylar L. Greene May 2025

Challenging $\Lambda$Cdm: Unraveling Cosmic Distances, Dark Sector Phenomenology, And Alternative Primordial B-Mode Sources, Kylar L. Greene

Physics & Astronomy ETDs

The dominant Lambda Cold Dark Matter (LCDM) cosmological model, while remarkably successful, increasingly shows signs that it may not fully describe our Universe, as persistent tensions in expansion rates and structure formation remain unresolved. In this thesis, I challenge the LCDM paradigm using novel theoretical frameworks combined with rigorous numerical analyses. I demonstrate that the expansion-rate tension fundamentally reflects underlying distance disagreements, and that the Thomson scattering rate strongly restricts higher pre-recombination expansion rates without additional physics. Further, I present a novel cosmological model using a mirror dark sector and varying fundamental constants, revealing an observational degeneracy allowing significantly higher …


Data Encoding, Compilation, And Algorithms For Quantum Machine Learning, Aviraj Sinha May 2025

Data Encoding, Compilation, And Algorithms For Quantum Machine Learning, Aviraj Sinha

Computer Science and Engineering Theses and Dissertations

Quantum computing enables new approaches to data processing, especially in quantum machine learning. Unlike classical systems, quantum data must be synthesized through operations and can exist in superposition. Encoding choices affect efficiency, noise resilience, and trainability—key factors in quantum machine learning models. This dissertation enhances quantum data encodings by extending quantum read-only memory (QROM) beyond binary representations, improving efficiency and parallelism. It introduces new compilation methods for quantum random number generators (QRNGs), supporting non-parametric distributions for post-quantum cryptography. Additionally, it explores Cayley graph-based encodings to extract spectral features for quantum machine learning.


Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani May 2025

Hybrid Graph-Recurrent Architecture For Citation Recommendation Via Future Embedding Forecasting, Mohammad Ausaf Ali Haqqani

Computer Science and Engineering Theses and Dissertations

The rapid expansion of scientific literature has intensified the challenge of identifying relevant citations, particularly for newly published or under-cited papers. Traditional citation recommendation systems typically model static relationships or respond to past citation activity, offering limited predictive power for emerging works. In response, this thesis presents a temporal modeling framework for citation recommendation that anticipates future scholarly relevance by forecasting the latent representations of academic papers.

Building on prior work that utilized Temporal Graph Networks (TGNs) to model dynamic citation flows, we propose Graph-Time, a hybrid architecture that integrates a Graph Transformer with a GRU-based time series predictor. The …


A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer May 2025

A Computational Method For Detecting Compound Promiscuity In Early-Stage Pharmaceutical Discovery, John Allen Ringer

Computer Science ETDs

Modern drug discovery and chemical biology research relies heavily on analyzing bioassay data. One of the many challenges in bioassay data analysis is identifying false trails, i.e., chemical compounds which initially appear to have desirable activity but are found to be problematic upon further investigation. Badapple (the BioAssay-Data Associative Promiscuity Pattern Learning Engine) was created over ten years ago to help researchers identify promiscuous compounds and thus avoid a common source of these false trails. Through an effort involving software engineering, cheminformatics, and biomedical data science we have developed Badapple 2.0, which incorporates updated assay records and expanded data semantics. …


Incorporating Latent Survival Trajectories And Covariate Heterogeneity In Time-To-Event Data Analysis: A Joint Mixture Model Approach, Fu-Wen Liang, Wenyaw Chan, Michael D Swartz, Bouthaina S Dabaja May 2025

Incorporating Latent Survival Trajectories And Covariate Heterogeneity In Time-To-Event Data Analysis: A Joint Mixture Model Approach, Fu-Wen Liang, Wenyaw Chan, Michael D Swartz, Bouthaina S Dabaja

Faculty, Staff and Student Publications

Background: Finite mixture models have been recently applied in time-to-event data to identify subgroups with distinct hazard functions, yet they often assume differing covariate effects on failure times across latent classes but homogeneous covariate distributions. This study aimed to develop a method for analyzing time-to-event data while accounting for unobserved heterogeneity within a mixture modeling framework.

Methods: A joint model was developed to incorporate latent survival trajectories and observed information for the joint analysis of time-to-event outcomes, correlated discrete and continuous covariates, and a latent class variable. It assumed covariate effects on survival times and covariate distributions vary across latent …


Gnns For Network Classification In Single Cell Rna Sequencing Data, Reid C. Sewell May 2025

Gnns For Network Classification In Single Cell Rna Sequencing Data, Reid C. Sewell

Capstone Projects

A common technique when investigating a disease is to profile gene expression, as this gives unique insights into the functions of a cell. Gene expression data gathered from single cell RNA sequencing can be encoded into a gene co-expression network, which is a graph of potential relationships between different genes. One method for interpreting data encoded as a graph is to use a graph neural network, or GNN. This project designs and implements a GNN architecture to accomplish classification tasks on graph data. Then, given a dataset of gene co-expression networks made from multiple single cell RNA sequencing studies, the …


Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow May 2025

Resale Revolution: Trend Implications From Media Presence Transcended To Luxury Retail Markets, Penelope Prochnow

Capstone Projects

This study aims to deepen understanding of fashion trend decline from peak popularity to obsolescence, with implications for sustainability and producer profit margins. It investigates how the attributes and media presence of fashion items influence their journey from high-end editorial coverage to resale platforms. Using survival analysis to model trend lifetimes and cosine similarity metrics to compare resale and magazine keyword frequencies, alongside machine learning for price prediction, the study uncovers critical temporal patterns. Results show that resale trends reflect magazine content with a lag of approximately 18 to 30 months and draw from long-wave revivals spanning 6 to 14 …


Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi May 2025

Machine Learning Course: A 15-Week Interactive Curriculum With Code And Case Studies, Pegah Khosravi

Open Educational Resources

This open-access machine learning course is a comprehensive 15-week curriculum developed and published on GitHub with full Google Colab compatibility. It combines theoretical concepts with hands-on Python coding, real-world datasets, and structured projects covering regression, classification, clustering, deep learning, transformers, and multimodal AI. The course is designed for students, educators, and researchers interested in applied machine learning, including biomedical applications. It includes explainable AI components and ethical discussions to align with modern AI standards. The course is maintained by BioMind AI Lab at CUNY.


Optimized Student Grouping For Enhanced Classroom Performance, Kathryn E. Reardon May 2025

Optimized Student Grouping For Enhanced Classroom Performance, Kathryn E. Reardon

Honors Theses

Effective grouping methods enhance classroom collaboration and allow for a student-centered teaching approach; however, traditional grouping methods are time-consuming, subjective, and can create inconsistent group dynamics. This project addresses these challenges by employing a data-driven approach to optimize student groups based on academic performance, behavior, attendance, language barriers, and teacher preferences. The minimum viable product is a web application with an algorithm-driven system to group students and a database storage for group results. During the initiation phase, a problem was defined with a proposed solution. During the planning phase, potential design choices and grouping methods were researched and assessed. During …


Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson May 2025

Evaluating Predictive Models For Predicting Total Score Of Beef Carcasses, Emmanuel Forson

Electronic Theses and Dissertations

The beef industry plays a vital role in global agriculture, with carcass quality and consumer preference being key determinants of market success. This thesis examines predictive modeling techniques for estimating the Total Score of beef carcasses, a composite measure representing yield and quality, primarily used by the Nebraska Cattlemen Association. Using data from the Nebraska Cattlemen’s Foundation Retail Value Steer Challenge (2000–2023), the study compares the performance of First Order Multiple Linear Regression (MLR) with three machine learning techniques: K-Nearest Neighbors (KNN), Random Forest, and Gradient Boosting Machine (GBM).

The analysis focuses on six key predictors: Hot Carcass Weight, Back …


Performance Of Lasso And Ridge Regression For Variable Selection In Genome-Wide Association Studies Of Maize Flowering Time, Dipok Deb May 2025

Performance Of Lasso And Ridge Regression For Variable Selection In Genome-Wide Association Studies Of Maize Flowering Time, Dipok Deb

Data Science and Data Mining

Genome-Wide Association Studies (GWAS) are instrumental in identifying genetic variants linked to complex traits, providing valuable insights into trait heritability and biological mechanisms. This study applies GWAS to investigate flowering time in maize, a critical adaptive trait, using a diverse dataset of 5,000 recombinant inbred lines across eight environments. Traditional GWAS methods often encounter challenges in high-dimensional datasets due to the presence of multiple small-effect genetic loci. To address this, we compared two penalized regression methods—LASSO and Ridge regression—to perform variable selection and regression analysis within a GWAS framework. LASSO effectively reduced the number of predictors by selecting the most …


Data Science Job Salary Prediction Using Linear Regression, Dipok Deb May 2025

Data Science Job Salary Prediction Using Linear Regression, Dipok Deb

Data Science and Data Mining

In the evolving landscape of data science, accurate salary prediction plays a crucial role in shaping career expectations, informing educational strategies, and guiding organizational hiring decisions. This study investigates the key factors influencing entry-level data science salaries in the United States by applying a multiple linear regression model to a recent dataset spanning from 2020 to 2024. Through data preprocessing, transformation, and diagnostic evaluation, we identify how job roles, experience levels, employment types, work arrangements, residency status, and company size impact compensation. Despite challenges such as outliers, heteroscedasticity, and non-normal residuals, model refinements like the Box-Cox transformation and variable selection …


Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville May 2025

Statistics - What Does My Data Say About Me?, Taylor Gadsden-Deterville

Student Scholar Symposium Abstracts and Posters

For my Introduction to Statistics Class, I have been tasked with collecting unique, personal data to give insight into my daily routine. I decided to record nine different outcomes (two qualitative and seven quantitative). On February 6, 2025, I began with a blank Excel sheet, and so far, I have 57 full days of data collected. I will continue monitoring my findings for the remainder of the Spring 2025 Semester. Per my project instructions, I must include tables and graphs for my qualitative and quantitative outcomes. So far, I have collected daily quantitative data on my screen time (Instagram and …


Two-Sample Bi-Directional Causality Between Two Traits With Some Invalid Ivs In Both Directions Using Gwas Summary Statistics, Siyi Chen May 2025

Two-Sample Bi-Directional Causality Between Two Traits With Some Invalid Ivs In Both Directions Using Gwas Summary Statistics, Siyi Chen

School of Public Health Faculty Publications

Mendelian randomization (MR) is a widely used method for assessing causal relationships between risk factors and outcomes using genetic variants as instrumental variables (IVs). While traditional MR assumes uni-directional causality, bi-directional MR aims to identify the true causal direction. In uni-directional MR, invalid IVs due to pleiotropy can violate assumptions and introduce biases. In bi-directional MR, traditional MR can be performed separately for each direction, but the presence of invalid IVs poses even greater challenges. We introduce a new bi-directional MR method incorporating stepwise selection (Bidir-SW) designed to address these challenges. Our approach leverages public genome-wide association study (GWAS) datasets …


Handwritten Digit Recognition Using Machine Learning, Dipok Deb May 2025

Handwritten Digit Recognition Using Machine Learning, Dipok Deb

Data Science and Data Mining

Handwritten Digit Recognition (HDR) remains a fundamental benchmark in pattern recognition and machine learning due to its practical applications and inherent classification challenges posed by diverse handwriting styles. This study investigates and compares two classical statistical classifiers—Gaussian Naive Bayes (GNB) and Linear Discriminant Analysis (LDA)—to recognize the digits from the MNIST dataset. Both models assume underlying normality in feature distributions and offer computational efficiency, making them suitable for high-dimensional input such as image pixels. Using 60,000 training and 10,000 test samples, we evaluate model performance through accuracy, precision, recall, F1 score, and confusion matrices. The results reveal that while GNB …


Clustering Dataset Using K-Mean Clustering, Dipok Deb May 2025

Clustering Dataset Using K-Mean Clustering, Dipok Deb

Data Science and Data Mining

Clustering is a fundamental technique in unsupervised machine learning, widely applied in various domains such as pattern recognition, data segmentation, and anomaly detection. This study evaluates the performance of the K-Means clustering algorithm on multiple benchmark datasets, including low-dimensional, high-dimensional, and imbalanced datasets. The clustering results are assessed using four key evaluation metrics: Mean Squared Error (MSE), Adjusted Rand Index (ARI), Normalized Mutual Information (NMI), and Silhouette Score. Experimental results demonstrate that K-Means performs effectively on datasets with well-separated clusters, particularly in high-dimensional spaces, where it achieves near-perfect clustering accuracy. However, its performance deteriorates in datasets with overlapping clusters and …


Dynamic Approaches To Missing Data In Healthcare: Evaluating Ensemble Models, Feature Selection, And Meta-Features, Dylan Dominguez Sulca May 2025

Dynamic Approaches To Missing Data In Healthcare: Evaluating Ensemble Models, Feature Selection, And Meta-Features, Dylan Dominguez Sulca

Theses and Dissertations

Missing data is pervasive in healthcare, where incomplete observations commonly arise from patient dropout, sensor failures, or privacy constraints. This research presents an investigation into handling such data, focusing on (1) Missingness-Aware Dynamic Ensemble Weighting (MDEW), (2) feature selection under varying missing rates, (3) autoencoder-based imputation (ODAE), and (4) a meta-feature analysis guiding pipeline selection. We evaluate our experiments on four diverse datasets, Cleveland Heart Disease, Diabetic Retinopathy, Breast Cancer Wisconsin, EEG Eye State. Our research shows that MDEW adaptively selects imputer classifier pipelines, outperforming single model and uniform averaging baselines at moderate to high missingness 10% to 50%. Filter …


Achieving Professional, Effective Graphics Efficiently With Available Clemson Resources, Stacie P. Powell, Brad Landon Walters, Yaswanth Mulakala May 2025

Achieving Professional, Effective Graphics Efficiently With Available Clemson Resources, Stacie P. Powell, Brad Landon Walters, Yaswanth Mulakala

Presentations

No abstract provided.


A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster May 2025

A Machine Learning Analysis Of Factors Leading To Major League Baseball Postseason Berths, Chase S. Foster

Undergraduate Honors Theses

Machine learning is a method that employs statistical algorithms to identify patterns and make predictions from data. This study applies machine learning techniques to analyze data from Major League Baseball (MLB) teams between 1998 and 2024, with the goal of determining which factors strongly influence a team's likelihood of reaching the postseason and in accurately predicting the teams that do and do not qualify for the postseason. Data exploration and unsupervised machine learning methods such as clustering were used to identify underlying patterns in team performance metrics and determine potential significant contributors to team success. Many different supervised learning methods …


The Little Diagram That Could: Geometric Properties And Statistical Applications Of Persistence Diagrams In Topological Data Analysis, Eugene Kler May 2025

The Little Diagram That Could: Geometric Properties And Statistical Applications Of Persistence Diagrams In Topological Data Analysis, Eugene Kler

McKelvey School of Engineering Graduate Student Theses & Dissertations

Topological Data Analysis (TDA) is a collection of techniques for data analysis that leverages topological invariants of spaces formed from data points. These methods excel at extracting useful information from noisy or sparse data, making them attractive to many mathematicians, statisticians, and scientists. In this thesis, we explore TDA on three fronts: algebraic foundations, statistical applications, and metric properties. Throughout, the central object of study is the Persistence Diagram (PD), a summary of the changes in homology that occur as one builds simplicial complexes from the data by increasing a parameter.


Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran May 2025

Land8fire: A Complete Study On Wildfire Segmentation Through Comprehensive Review, Human-Annotated Multispectral Dataset, And Extensive Benchmarking, Anh Tran

Data Science Undergraduate Honors Theses

Early and accurate wildfire detection is critical for minimizing environmental damage and ensuring a timely response. However, existing satellite-based wildfire datasets suffer from limitations such as coarse ground truth, poor spectral coverage, and class imbalance, which hinder progress in developing robust segmentation models. In this paper, we introduce Land8Fire, a new large-scale wildfire segmentation dataset composed of over 20,000 multispectral image patches derived from Landsat 8 and manually annotated for high-quality fire masks. Building on the ActiveFire dataset, Land8Fire improves ground truth reliability and offers predefined splits for consistent benchmarking. We evaluate a range of state-of-the-art convolutional and transformer-based models, …


Advancing Drug-Drug Interaction Prediction Using Multi-Modal Feature Integration With Graph Neural Networks, Ernest C. Chianumba May 2025

Advancing Drug-Drug Interaction Prediction Using Multi-Modal Feature Integration With Graph Neural Networks, Ernest C. Chianumba

Theses, Dissertations and Culminating Projects

Pharmaceutical treatments are essential for managing medical conditions, but drug-drug interactions (DDIs) pose significant risks to patient safety and healthcare outcomes. This research integrates Knowledge Graphs and Graph Neural Networks to predict DDIs by exploring complex drug relationships. We construct a comprehensive knowledge graph using DrugBank data (1,000 drugs, 155,774 interactions) enriched with molecular features from PubChem. Our methodology introduces a novel multi-modal approach by integrating transformer-based embeddings (ChemBERTa, SPECTER, and SBERT) to create 1152-dimensional feature vectors that capture structural, biomedical literature, and semantic properties of drugs. Formulating DDI prediction as a link prediction task, we compare three Graph Neural …


Don't Let Lead Lead On Environmental Justice: A Simulative Approach To Lead Remediation In The Big Data Era, Charles C. Knoble Ii May 2025

Don't Let Lead Lead On Environmental Justice: A Simulative Approach To Lead Remediation In The Big Data Era, Charles C. Knoble Ii

Theses, Dissertations and Culminating Projects

Environmental justice, as both a movement and a theoretical construct, continues to evolve in response to shifting societal, environmental, and technological conditions. This dissertation investigates the integration of big data, such as social media, remote sensing imagery, and internet search frequencies, into the identification, analysis, and remediation of environmental injustices. Framing environmental justice through the lenses of distributive and data justice, the project explores both the promises and pitfalls of using emergent data sources to enhance the spatial and temporal precision of environmental equity investigations. Through a combination of systematic literature review, spatial analysis, system dynamics simulation, and policy evaluation, …