Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Claremont Colleges

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 36

Full-Text Articles in Data Science

Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer Jun 2026

Pinnlab: An Interactive Dashboard For Teaching Data-Driven Parameter Estimation In Differential Equations Using Physics-Informed Neural Networks, Mohan J. Parthasarathy, Padmanabhan Seshaiyer

CODEE Journal

Undergraduate instruction in ordinary differential equations (ODEs) is typically organized around the forward problem: finding solution trajectories when the governing equation and its parameters are known. In scientific practice, however, inverse problems are often more relevant, requiring unknown parameters to be inferred from noisy observations while assessing whether a proposed model is consistent with the data. We introduce PINNLab, an open-source MATLAB dashboard designed to help undergraduate students explore inverse modeling through physics-informed neural networks (PINNs). PINNLab presents PINNs as a complementary data-driven framework that connects differential equations, optimization, empirical data, and scientific machine learning. The instructional sequence is organized …


A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson Jun 2026

A Professional Development Course On Data-Driven Dynamical Systems At A Primarily Undergraduate Institution: Part A - Scientific Content, Alessandro M. Selvitella, Jeffrey R. Anderson

CODEE Journal

In the age of data-driven decision making, ordinary differential equations (ODEs) remain a powerful and interpretable framework for modeling dynamic processes, especially when integrated with modern tools from statistical learning and data-driven dynamical systems. Yet, general undergraduate and graduate curricula do not typically address key opportunities in data-driven dynamical systems.

This first paper in a series focuses on the mathematical and methodological core of a professional development course first developed in the academic year 2025-2026 at a Primarily Undergraduate Institution, Purdue University Fort Wayne. The curriculum developed in this course emphasized how regression, regularization, and sparse identification can be used …


From Vibration To Visualization: Building An Real-Time Audio Visualization System For Learning And Exploration Using Pyqt5, Aidan Roach Apr 2026

From Vibration To Visualization: Building An Real-Time Audio Visualization System For Learning And Exploration Using Pyqt5, Aidan Roach

The Transdisciplinary STEAM+ Journal

This paper presents the design, development, and analysis of my real-time audio visualization system created entirely in Python using PyQt5, called WaveCatcher. The system captures live audio input from a microphone and provides simultaneous visual feedback through multiple signal representations: including a time-domain waveform, a frequency-domain FFT spectrum, a scrolling spectrogram, harmonic peak visualization, dynamic range, amplitude envelope, spectral centroid, and spectral bandwidth. These features offer insight not just into the raw structure of sound, but into how humans perceive its qualities–like timbre! This terminology may seem intimidating—it certainly was when I first began learning it—but I’ll explain all of …


Analysis Of Δ¹¹B As A Seawater Ph Proxy: Comparing Ocean Circulation Inverse Model Output With Marine Calcifier Geochemistry, Jesse I. Dong Jan 2026

Analysis Of Δ¹¹B As A Seawater Ph Proxy: Comparing Ocean Circulation Inverse Model Output With Marine Calcifier Geochemistry, Jesse I. Dong

CMC Senior Theses

Increasing anthropogenic carbon flux into the oceans decreases seawater pH, alters dissolved inorganic carbon speciation, and reduces biogenic calcification. The marine calcifiers— specifically corals and coralline algae—incorporate elements from surrounding seawater into their carbonate structures, which preserve past records of ocean carbon chemistry. In particular, boron in biogenic carbonate is a potentially valuable proxy for historical ocean pH across human timescales. Within seawater, boron primarily exists as boric acid B(OH)3 and borate ions B(OH)4 - , where higher pH favors the formation of borate ions. Borate ions preferentially incorporate the heavier ¹¹B isotope over 10B. On the other hand, if …


Microbial Community Structure In Global Soils, Matthew Jabro Jan 2026

Microbial Community Structure In Global Soils, Matthew Jabro

CMC Senior Theses

Soil harbors the most diverse microbial communities on Earth, yet whether predictable community types exist across biomes and whether taxonomic composition encodes habitat of origin remain open questions at global scale. This thesis addresses both questions by applying unsupervised clustering and supervised classification to transformed 16S ribosomal RNA (rRNA) amplicon profiles from two independent datasets: the global topsoil survey of Bahram et al. (193 samples) and the Earth Microbiome Project (EMP) soil subset of Thompson et al. (2,209 samples). Application of a sample clustering method based on a mixture of Gaussian Graphical Models (MixGGM) identified 19 clusters in the topsoil …


Volatility Modeling With An Application To Risk Parity Portfolios, Kenneth Hou Jan 2026

Volatility Modeling With An Application To Risk Parity Portfolios, Kenneth Hou

CMC Senior Theses

This thesis studies volatility modeling in the context of risk parity portfolio construction. I compare three risk parity portfolios that differ only in their underlying volatility model: a historical covariance baseline, a Bayesian stochastic volatility model, and a GRU–GARCH hybrid neural network. Using daily returns on Kenneth French’s five industry portfolios from January 2016 through December 2025, I construct monthly rebalanced portfolios under each model, with the SV and GRU forecasts embedded in hybrid covariance matrices that combine forecasted volatilities with rolling historical correlations. The results document a divergence between forecast accuracy and portfolio performance: the SV model is the …


Random Graph Models For Dual Graphs, Anne Friedman Jan 2025

Random Graph Models For Dual Graphs, Anne Friedman

Scripps Senior Theses

This paper aims to better characterize dual graphs derived from state districting maps by developing random graph models that replicate their structural properties. Dual graphs provide a simplified way to represent districting maps, making it computationally feasible to analyze their structure. These representations enable researchers, legislators, and courts to assess district compactness, detect signs of gerrymandering, and generate alternative districting plans. A deeper understanding of the structural patterns of these dual graphs can help researchers choose or design more effective algorithms for redistricting analysis. The random graph models developed in this study serve as testbeds for evaluating algorithmic approaches to …


A Proof Of Np-Completeness For The K-Means Clustering Algorithm, Brooke C. Feinberg Jan 2025

A Proof Of Np-Completeness For The K-Means Clustering Algorithm, Brooke C. Feinberg

Scripps Senior Theses

The k-means clustering algorithm is one of the most widely used clustering techniques in data analysis and machine learning, yet its exact computational complexity remains subject to ongoing theoretical investiga- tion. This work establishes the NP-completeness of k-means by proving (1) it is NP-hard and (2) it lies in NP. To demonstrate NP-hardness, we construct a series of polynomial-time reductions from well-known NP-complete problems. Specifically, we reduce 3sat to Vertex Cover, and then reduce Vertex Cover to k-means, thereby establishing the computational hardness of the k-means clustering problem. We then prove k-means is in NP, and thus conclude it is …


Optimizing Decision-Making In A Cerebral Palsy Model Using Reinforcement Learning, Richard Ampah Jan 2025

Optimizing Decision-Making In A Cerebral Palsy Model Using Reinforcement Learning, Richard Ampah

Pitzer Senior Theses

This study presents an original interdisciplinary investigation into how reinforcement learning (RL) can model motor and cognitive defects and potentially improve motor and cognitive functions in individuals with cerebral palsy (CP), a non-progressive neurological disorder that impairs movement and adaptability. Integrating computational neuroscience and machine learning, the research applies policy gradient methods and Markov Decision Processes (MDPs) to simulate adaptive learning in agents with and without CP-related constraints.

The central aim is to compare the cumulative rewards of optimal policies, derived from value iteration, and human-like learning policies using the REINFORCE algorithm, both with and without the Bellman baseline. The …


Empirical Analysis Of Political Districting Splitability Via Uniform Spanning Trees In Polynomial Time, Brooke C. Feinberg Jan 2025

Empirical Analysis Of Political Districting Splitability Via Uniform Spanning Trees In Polynomial Time, Brooke C. Feinberg

Scripps Senior Theses

This work expands a recently proven conjecture that a polynomial fraction of all uniform spanning trees (USTs) are splittable into k balanced partitions on grid graphs to real-world political districting plans. We investigate whether similar structural properties hold for the planar dual graphs of U.S. counties (cnty) and tracts (t), using Wilson’s algorithm to generate uniform random spanning trees and Breadth- First Search (BFS) to check for splitability into balanced partitions. Our empirical findings suggest that real-world districting plans can be split into 2-balanced, connected partitions in a fraction of polynomial time. This result highlights the potential for scalable redistricting …


Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey Jan 2025

Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey

CMC Senior Theses

This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …


Neural Correlates Of Attentional Biases In Dietary Choice: Role Of Childhood Socioeconomic Status, Justine Jamie N. Gotico Jan 2025

Neural Correlates Of Attentional Biases In Dietary Choice: Role Of Childhood Socioeconomic Status, Justine Jamie N. Gotico

CMC Senior Theses

Childhood poverty has been shown to increase adult risk for obesity above and beyond its direct effects on adult socioeconomic status (SES). One proposed mechanism of these effects is by shifting behavioral patterns of dietary consumption and choice, for example by increasing rapid attention to high-calorie unhealthy foods. Yet, whether such neural mechanisms can explain observed differences in dietary behavior based on childhood SES remains an open question. Here we used event-related potentials (ERPs) to examine early attentional correlates of low childhood SES during a dietary choice task, based on research suggesting that early attentional biases toward high-calorie foods emerge …


Cmc Thesis Chatbot, Luis Gomez Jan 2025

Cmc Thesis Chatbot, Luis Gomez

CMC Senior Theses

This GitHub repo is a senior thesis for Claremont McKenna College; it is a thesis about theses. The project is an interactive RAG-based chatbot that helps students, researchers, and faculty explore Claremont McKenna College senior theses. The goal was to create a domain-specific chatbot to show that it is possible to combat the limitations of AI, including hallucinations, outdated data, and lack of domain expertise. The website link is:

CMCThesisChatbot


Mathematical Modeling, Analysis, And Simulation Of Patient Addiction Journey, Adan Baca, Diego Gonzalez, Alonso G. Ogueda, Holly C. Matto, Padmanabhan Seshaiyer Aug 2024

Mathematical Modeling, Analysis, And Simulation Of Patient Addiction Journey, Adan Baca, Diego Gonzalez, Alonso G. Ogueda, Holly C. Matto, Padmanabhan Seshaiyer

CODEE Journal

This paper aims to develop a mathematical model to study the dynamics of addiction as individuals go through their detox journey. The motivation for this work is three fold. First, there has been a significant increase in drug overdose and drug addiction following the COVID-19 pandemic, and addiction may be interpreted as a infectious disease. Secondly, the dynamics of infectious disease could be modeled via compartmental models described by differential equations and one can therefore leverage the existing analytical and numerical methods to model addiction as a disease. Finally, the work helps to inform how mathematical models governed by differential …


Book Review: How To Expect The Unexpected: The Science Of Making Predictions -- And The Art Of Knowing When Not To By Kit Yates, Mark Huber Jul 2024

Book Review: How To Expect The Unexpected: The Science Of Making Predictions -- And The Art Of Knowing When Not To By Kit Yates, Mark Huber

Journal of Humanistic Mathematics

Humans think about the future all the time. Prediction is a part of how we prepare for the coming of both good and bad events in our lives. Kit Yates' book, How to expect the unexpected, concentrates primarily on the question of why prediction is difficult, and what mental shortcuts people take in prediction that can lead to incorrect results. Unfortunately, a lack of concern for details and several omissions undermine the quality of the book.


Towards Algorithmic Justice: Human Centered Approaches To Artificial Intelligence Design To Support Fairness And Mitigate Bias In The Financial Services Sector, Jihyun Kim Jan 2024

Towards Algorithmic Justice: Human Centered Approaches To Artificial Intelligence Design To Support Fairness And Mitigate Bias In The Financial Services Sector, Jihyun Kim

CMC Senior Theses

Artificial Intelligence (AI) has positively transformed the Financial services sector but also introduced AI biases against protected groups, amplifying existing prejudices against marginalized communities. The financial decisions made by biased algorithms could cause life-changing ramifications in applications such as lending and credit scoring. Human Centered AI (HCAI) is an emerging concept where AI systems seek to augment, not replace human abilities while preserving human control to ensure transparency, equity and privacy. The evolving field of HCAI shares a common ground with and can be enhanced by the Human Centered Design principles in that they both put humans, the user, at …


Exploring U.S. Natural Disasters And Psychological Distress: From Time Series Trends To Machine Learning Insights On Hurricane Helene, Sarah Jane Fullerton Jan 2024

Exploring U.S. Natural Disasters And Psychological Distress: From Time Series Trends To Machine Learning Insights On Hurricane Helene, Sarah Jane Fullerton

CMC Senior Theses

This research investigates the historical trends of psychological distress in the U.S. in relation to natural disaster occurrences. By analyzing long-term data, we examine how significant natural disasters relate to levels of psychological distress over time. The research employs Exploratory Data Analysis (EDA) and Time Series Analysis to identify patterns and trends between the frequency and intensity of natural disasters and the rise of psychological distress across various periods in U.S. history. Additionally, real-time data from Reddit was collected through a custom-built Reddit web scraper specialized for Hurricane Helene. This dataset was labeled for sentiment and used to train machine …


Reducing Generalization Error In Multiclass Classification Through Factorized Cross Entropy Loss, Oleksandr Horban Jan 2024

Reducing Generalization Error In Multiclass Classification Through Factorized Cross Entropy Loss, Oleksandr Horban

CMC Senior Theses

This paper introduces Factorized Cross Entropy Loss, a novel approach to multiclass classification which modifies the standard cross entropy loss by decomposing its weight matrix W into two smaller matrices, U and V, where UV is a low rank approximation of W. Factorized Cross Entropy Loss reduces generalization error from the conventional O( sqrt(k / n) ) to O( sqrt(r / n) ), where k is the number of classes, n is the sample size, and r is the reduced inner dimension of U and V.


Math And Democracy, Kimberly A. Roth, Erika L. Ward Aug 2023

Math And Democracy, Kimberly A. Roth, Erika L. Ward

Journal of Humanistic Mathematics

Math and Democracy is a math class containing topics such as voting theory, weighted voting, apportionment, and gerrymandering. It was first designed by Erika Ward for math master’s students, mostly educators, but then adapted separately by both Erika Ward and Kim Roth for a general audience of undergraduates. The course contains materials that can be explored in mathematics classes from those for non-majors through graduate students. As such, it serves students from all majors and allows for discussion of fairness, racial justice, and politics while exploring mathematics that non-major students might not otherwise encounter. This article serves as a guide …


Responsible Data Science For Genocide Prevention, Victor Piercey Aug 2023

Responsible Data Science For Genocide Prevention, Victor Piercey

Journal of Humanistic Mathematics

The term "genocide" emerged out of an effort to describe mass atrocities committed in the first half of the 20th century. Despite a convention of the United Nations outlawing genocide as a matter of international law, the problem persists. Some organizations (including the United Nations) are developing indicator frameworks and “early-warning” systems that leverage data science to produce risk assessments of countries where conflict is present. These tools raise questions about responsible data use, specifically regarding the data sources and social biases built into algorithms through their training data. This essay seeks to engage mathematicians in discussing these concerns.


Application Of Sentiment Analysis And Machine Learning Techniques To Predict Daily Cryptocurrency Price Returns, Edward Wu Jan 2023

Application Of Sentiment Analysis And Machine Learning Techniques To Predict Daily Cryptocurrency Price Returns, Edward Wu

CMC Senior Theses

This paper examines the effects of social media sentiment relating to Bitcoin on the daily price returns of Bitcoin and other popular cryptocurrencies by utilizing sentiment analysis and machine learning techniques to predict daily price returns. Many investors think that social media sentiment affects cryptocurrency prices. However, the results of this paper find that social media sentiment relating to Bitcoin does not add significant predictive value to forecasting daily price returns for each of the six cryptocurrencies used for analysis and that machine learning models that do not assume linearity between the current day price return and previous daily price …


Utilizing Machine Learning In Healthcare In An Ethical Fashion, Nishka Ayyar Jan 2023

Utilizing Machine Learning In Healthcare In An Ethical Fashion, Nishka Ayyar

CMC Senior Theses

This thesis paper explores the ethical considerations surrounding the use of machine learning (ML) solutions in healthcare. The background section discusses the basics of machine learning techniques and algorithms, and the increasing interest in their utilization in the healthcare sector. The paper then reviews and critically analyzes four studies that highlight concerns related to using ML in healthcare, including issues of bias, privacy, accountability, and transparency. Based on the analysis of these studies, the paper presents several recommendations for addressing these concerns. The paper concludes with a discussion on the potential benefits of using machine learning technology in healthcare. Ultimately, …


Defining The "Quadruple-A" Player: What Makes A Baseball Player Succeed In The Minor Leagues And Fail In The Major Leagues?, Sam Bogen Jan 2023

Defining The "Quadruple-A" Player: What Makes A Baseball Player Succeed In The Minor Leagues And Fail In The Major Leagues?, Sam Bogen

CMC Senior Theses

The "Quadruple-A" player is defined as one who is too good to play in Triple-A (the league one step down from Major League Baseball) but not good enough to play consistently in Major League Baseball. This thesis paper attempts to explain the phenomenon of the "Quadruple-A" player. Using Triple-A data from 2013-2022 and Major League data from the "Statcast Era" (2015-2022), I build logistic and linear regression models to predict Major League success based on Triple-A performance data as well as Major League Statcast data, discovering that statistics related to how a player hits the ball such as the speed …


Maximizing Productivity And Quality In Senior Thesis Writing With Artificial Intelligence And Natural Language Processing Driven Tools, Lauren Leadbetter Jan 2023

Maximizing Productivity And Quality In Senior Thesis Writing With Artificial Intelligence And Natural Language Processing Driven Tools, Lauren Leadbetter

CMC Senior Theses

This project is a Python program designed to generate a senior thesis on a user-
inputted topic using natural language processing techniques. The program takes in a
topic from the user and then uses OpenAI API to deploy text models for text genera-
tion and evaluation, such as GPT-3 and Davinci-003. The resulting output is in .tex
format and includes a first-draft outline and paper, followed by self-generated assessment, with scoring, revisions, and feedback comments instructing manual revisions.

This submission is a sample using one available model of the project, meant to
demonstrate it’s functionality and limitations. Further model versions …


A Study On Global Reef Deterioration: Exploring Coral Bleaching, Emily Fernandez Jan 2023

A Study On Global Reef Deterioration: Exploring Coral Bleaching, Emily Fernandez

CMC Senior Theses

This thesis is a study on coral bleaching and coral mortality, studying the relationship between variables such as depth, exposure, distance to shore, and temperature for percent bleaching. All of the analyses were made using two different data sets, that contain information about bleaching events in specific regions, and dates, and provide information factors such as depth, temperature, and exposure. Models were created for different relationships of variables for eco-regions, recent data, and countries. I attempted to find relationships between variables such as depth, temperature, exposure, and distance to shore, and how they affect coral bleaching. Unfortunately, I did not …


Warehouses In The Inland Empire: Displacing Land And Life, Katherine Gelsey Jan 2023

Warehouses In The Inland Empire: Displacing Land And Life, Katherine Gelsey

Pomona Senior Theses

The Inland Empire in Southern California embodies unique spatial and social configurations as a consequence of how settler colonialism has manifested locally in the region since the Spanish Mission Period. This work uses GIS software to estimate patterns of land conversion for residential, agricultural, and warehouse land from 2012 to 2022. Preliminary analysis suggests that thousands of people have been displaced by warehouse expansion over the ten-year period. In the twenty-first century, the Southern California logistics industry continues processes of land dispossession and racialized labor exploitation through displacing agricultural and residential land, exposing disproportionately low-income Black and Latine communities living …


Quantifying The Carbon Stored And Sequestered By The Trees On Pomona College’S Campus, Paola A. Giron-Carson Jan 2023

Quantifying The Carbon Stored And Sequestered By The Trees On Pomona College’S Campus, Paola A. Giron-Carson

Scripps Senior Theses

We are experiencing a climate crisis that must be confronted with strategic mitigation. Pomona College contributes to the climate crisis through its emissions for which there is a baseline record. However there is no baseline record of the climate mitigation currently performed by the trees on Pomona’s campus through carbon storage. This study seeks to determine a current baseline quantity of carbon stored and sequestrated by Pomona’s trees as well as possible courses of climate mitigation for Pomona College to take. Initial information gathering was conducted through interviews with several stakeholders. This study was conducted using data collected prior to …


Check Yourself Before You Wrek Yourself: Unpacking And Generalizing Randomized Extended Kaczmarz, William Gilroy Jan 2022

Check Yourself Before You Wrek Yourself: Unpacking And Generalizing Randomized Extended Kaczmarz, William Gilroy

HMC Senior Theses

Linear systems are fundamental in many areas of science and engineering. With the advent of computers there now exist extremely large linear systems that we are interested in. Such linear systems lend themselves to iterative methods. One such method is the family of algorithms called Randomized Kaczmarz methods.
Among this family, there exists a Randomized Kaczmarz variant called Randomized
Extended Kaczmarz which solves for least squares solutions in inconsistent linear systems.
Among Kaczmarz variants, Randomized Extended Kaczmarz is unique in that it modifies input system in a special way to solve for the least squares solution. In this work we …


Predicting Outcomes Of El Clásico Using Random Forests And Extreme Gradient Boosting, Emanuel Jarquin Jan 2022

Predicting Outcomes Of El Clásico Using Random Forests And Extreme Gradient Boosting, Emanuel Jarquin

CMC Senior Theses

In the modern era, sports betting is becoming increasingly popular. This is especially true in the realm of soccer (or ‘football’ as it is known outside the United States). As a result, the concept of attempting to predict the outcomes of soccer matches using machine learning has garnered much attention in recent years. In this thesis, I utilize well-known machine learning techniques to predict the outcomes of El Clásico matchups and compare the predictive performance of these techniques. The predictive methods employed for this thesis are random forests using the party package in R and extreme gradient boosting using the …


Advanced Full-Text Search Based On Synonyms In Postgres, Joey Bodoia Jan 2022

Advanced Full-Text Search Based On Synonyms In Postgres, Joey Bodoia

CMC Senior Theses

This paper discusses the advanced full-text search queries based on synonyms that are supported in Chajda, which is a postgres extension and corresponding python library for highly multi-lingual full-text search in postgres. This discussion will include the motivations for using advanced queries based on synonyms, examples of how to use these advanced queries in Chajda, current limitiations of the advanced queries, and performance testing of the advanced queries.