Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Claremont Colleges

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 1 - 30 of 36

Full-Text Articles in Applied Statistics

When Three Points Aren't Enough: Parameter Estimation In Newton's Law Of Cooling, Alberto A. Condori, Cara D. Brooks, Madeline R. Goldberg Sep 2026

When Three Points Aren't Enough: Parameter Estimation In Newton's Law Of Cooling, Alberto A. Condori, Cara D. Brooks, Madeline R. Goldberg

CODEE Journal

We begin with a paradoxical three-point problem where the standard parameter estimation formula fails because temperature data must satisfy a concavity condition reflecting Newton's Law of Cooling's exponential structure. Although a closed-form solution for A exists for equally-spaced measurements, high sensitivity to error motivates the use of overdetermined systems with many measurements. The "profiling over A" technique transforms this three-parameter nonlinear problem into a sequence of simple linear regressions, providing computational efficiency and conceptual transparency. This approach can be generalized to many parameter estimation problems in science and engineering, making it a valuable tool for undergraduates interested in applied mathematics. …


Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha Jan 2026

Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha

CMC Senior Theses

This thesis documents the design, deployment, and forward-test evaluation of an evolutionary multi-agent algorithmic trading system on Polymarket, the largest decentralized prediction market. The system pairs a locally-hosted 72-billion-parameter language model with a gradient-boosted statistical filter and an evolutionary selection mechanism that maintains a population of approximately 500 autonomous trading agents. Each agent generates a probability estimate for an event, compares it to the prevailing market price, and trades the resulting disagreement.

The central empirical exercise estimates a panel regression of trade-level profit on the absolute disagreement between the agent's probability estimate and the market price, controlling for agent identity, …


Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls, Jason Liang Jan 2026

Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls, Jason Liang

CMC Senior Theses

Polling results from traditional single-choice plurality elections are readily interpretable. Simple frequentist population parameters are estimated, including each candidate’s total support and the size of the front runner's lead. If the point estimate for the size of the front runner's lead exceeds the margin of error of the lead, we can conclude that the poll shows a statistically significant front runner. However, the interpretability of these population statistics disappears when applied to ranked-choice voting elections. Because ballots rank multiple candidates and candidates are eliminated in rounds, simple population-wide parameters are not well-defined. In RCV elections, a candidate’s ability to win …


Rank Rebalancing In Commodity And Foreign Exchange Markets, Prateek D. Vyas Jan 2026

Rank Rebalancing In Commodity And Foreign Exchange Markets, Prateek D. Vyas

CMC Senior Theses

This thesis empirically tests the rank-rebalancing mechanism of Stochastic Portfolio Theory (SPT) across commodity futures, foreign exchange futures, and equity ETFs. The Reverse Price-Weighted strategy (RPW) assigns, to each asset, the market weight of the asset at the opposite price rank, and generates an annualized excess return of 2.90% over the price-weighted (MKT) commodity benchmark, during the period of November 1977 to October 2025 (HAC t = 2.058, p = 0.040). The differential Sharpe ratio (dSharpe) of 0.245 is confirmed by a stationary block bootstrap, with a 𝑝-value of 0.015, and factor regressions controlling for carry, momentum, and value yield …


“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King Jan 2025

“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King

Pomona Senior Theses

The work of this thesis is twofold — first, qualitatively characterizing the confluence between the British eugenics and statistics movements in the late 19th and early 20th centuries, and second, quantitatively analyzing the effect of this foundation on pedagogical materials in the growing field of statistics between 1880 and 1970. Towards the first goal, the history of the method of least squares, state statistics, and positive and negative eugenics are outlined, followed by a close reading of the foundational texts authored by Francis Galton and Karl Pearson that introduced linear regression. Towards the latter goal, English-language statistics textbooks published between …


Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri Jan 2025

Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri

CMC Senior Theses

Over the past decades, the gaming industry has managed to evolve into a multi-billion-dollar enterprise. Gaming platforms such as Steam foster unprecedented amounts of engagement among players worldwide daily. In this thesis, we investigate the effect of incorporating sentiment-driven metrics, specifically YouTube view counts and positive reviews, into predictive models for game popularity. In addition, by comparing our linear regression sentiment-based approach to the Bayesian hierarchical folded normal model used by De Luisa et al. (2021), we can understand the many differences, strengths, and limitations of each methodology. In our thesis, we focus on three games. Each is of varying …


Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov Jan 2025

Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov

CMC Senior Theses

Traditional beta estimates are constructed from historical stock‑and‑market returns and therefore adjust only as fast as realized data accrue. This thesis investigates whether the forward‑looking information embedded in equity‑option prices can enhance beta forecasts. Using near‑end‑of‑day quotes for 236 S&P 500 firms between 2007 and 2024, I extract risk‑neutral variance and skewness, construct five alternative beta estimators (historical, option‑implied, and three hybrids), and evaluate them against realized betas over six‑, twelve‑, and twenty‑four‑month windows. Rolling‑OLS beta remains the most accurate benchmark at short horizons, yet option‑implied moments add economically and statistically significant value when systematic exposure is expected to change …


Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey Jan 2025

Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey

CMC Senior Theses

This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …


The "Benfordness" Of Bach Music, Chadrack Bantange, Darby Burgett, Luke Haws, Sybil Prince Nelson Aug 2023

The "Benfordness" Of Bach Music, Chadrack Bantange, Darby Burgett, Luke Haws, Sybil Prince Nelson

Journal of Humanistic Mathematics

In this paper we analyze the distribution of musical note frequencies in Hertz to see whether they follow the logarithmic Benford distribution. Our results show that the music of Johann Sebastian Bach and Johann Christian Bach is Benford distributed while the computer-generated music is not. We also find that computer-generated music is statistically less Benford distributed than human- composed music.


Math And Democracy, Kimberly A. Roth, Erika L. Ward Aug 2023

Math And Democracy, Kimberly A. Roth, Erika L. Ward

Journal of Humanistic Mathematics

Math and Democracy is a math class containing topics such as voting theory, weighted voting, apportionment, and gerrymandering. It was first designed by Erika Ward for math master’s students, mostly educators, but then adapted separately by both Erika Ward and Kim Roth for a general audience of undergraduates. The course contains materials that can be explored in mathematics classes from those for non-majors through graduate students. As such, it serves students from all majors and allows for discussion of fairness, racial justice, and politics while exploring mathematics that non-major students might not otherwise encounter. This article serves as a guide …


Using Short Bursts To Optimize Redistricting In Georgia, Vedika Vishweshwar Jan 2022

Using Short Bursts To Optimize Redistricting In Georgia, Vedika Vishweshwar

CMC Senior Theses

Identifying extreme outliers in large state spaces is a difficult prob-
lem. I consider this problem in the context of finding political district-
ing plans that maximize the number of districts in which the majority
of the population is from a minority group, such as African Americans.
Since the set of all possible districting plans is enormous and unfeasi-
ble to examine in practice, this paper proposes a sampling method to
find these outlying plans. Specifically, this paper experiments with short
bursts in the context of minority voting rights in Georgia. Short bursts
are a type of Markov Chain in …


Neither “Post-War” Nor Post-Pregnancy Paranoia: How America’S War On Drugs Continues To Perpetuate Disparate Incarceration Outcomes For Pregnant, Substance-Involved Offenders, Becca S. Zimmerman Jan 2021

Neither “Post-War” Nor Post-Pregnancy Paranoia: How America’S War On Drugs Continues To Perpetuate Disparate Incarceration Outcomes For Pregnant, Substance-Involved Offenders, Becca S. Zimmerman

Pitzer Senior Theses

This thesis investigates the unique interactions between pregnancy, substance involvement, and race as they relate to the War on Drugs and the hyper-incarceration of women. Using ordinary least square regression analyses and data from the Bureau of Justice Statistics’ 2016 Survey of Prison Inmates, I examine if (and how) pregnancy status, drug use, race, and their interactions influence two length of incarceration outcomes: sentence length and amount of time spent in jail between arrest and imprisonment. The results collectively indicate that pregnancy decreases length of incarceration outcomes for those offenders who are not substance-involved but not evenhandedly -- benefitting white …


Information Prioritization: A Comparison Between Utility Maximizers And Probability Matchers, Yusuf Ismaeel Jan 2021

Information Prioritization: A Comparison Between Utility Maximizers And Probability Matchers, Yusuf Ismaeel

CMC Senior Theses

This thesis examines the differences between probability matchers and utility maximizers in their preferences for information sources in a lab environment. In this paper, we consider the best source of information to be the most connected one. We conducted several linear probability model type regressions along with logit regressions. Furthermore, we also attempted to control and fix any potential misclassifications in classifying the cognitive strategy by using instrumental variables. The results show that utility maximizers will almost always choose the most informed node. Probability matchers, on the other hand, do not exhibit such a behavior as the probability matching strategy …


How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller Jan 2020

How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller

CMC Senior Theses

In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …


Choose Your Own Adventure: An Analysis Of Interactive Gamebooks Using Graph Theory, D'Andre Adams, Daniela Beckelhymer, Alison Marr Jul 2019

Choose Your Own Adventure: An Analysis Of Interactive Gamebooks Using Graph Theory, D'Andre Adams, Daniela Beckelhymer, Alison Marr

Journal of Humanistic Mathematics

"BEWARE and WARNING! This book is different from other books. You and YOU ALONE are in charge of what happens in this story." This is the captivating introduction to every book in the interactive novel series, Choose Your Own Adventure (CYOA). Our project uses the mathematical field of graph theory to analyze forty books from the CYOA book series for ages 9-12. We first began by drawing the digraphs of each book. Then we analyzed these digraphs by collecting structural data such as longest path length (i.e. longest story length) and number of vertices with outdegree zero (i.e. number …


Step Away From Stepwise, Gary N. Smith Jan 2019

Step Away From Stepwise, Gary N. Smith

Pomona Economics

Stepwise regression is a popular data-mining tool that uses statistical significance to select the explanatory variables to be used in a multiple-regression model. A fundamental problem with stepwise regression is that some real explanatory variables that have causal effects on the dependent variable may happen to not be statistically significant, while nuisance variables may be coincidentally significant. As a result, the model may fit the data well in-sample, but do poorly out-of-sample. Many Big-Data researchers believe that, the larger the number of possible explanatory variables, the more useful is stepwise regression for selecting explanatory variables. The reality is that stepwise …


Be Wary Of Black-Box Trading Algorithms, Gary N. Smith Jan 2019

Be Wary Of Black-Box Trading Algorithms, Gary N. Smith

Pomona Economics

Black-box algorithms now account for nearly a third of all U. S. stock trades. It is a mistake to think that these algorithms possess superhuman intelligence. In reality, computers do not have the common sense and wisdom that humans have accumulated by living. Trading algorithms are particularly dangerous because they are so efficient at discovering statistical patterns—but so utterly useless in judging whether the discovered patterns are meaningful.


On Cluster Robust Models, José Bayoán Santiago Calderón Jan 2019

On Cluster Robust Models, José Bayoán Santiago Calderón

CGU Theses & Dissertations

Cluster robust models are a kind of statistical models that attempt to estimate parameters considering potential heterogeneity in treatment effects. Absent heterogeneity in treatment effects, the partial and average treatment effect are the same. When heterogeneity in treatment effects occurs, the average treatment effect is a function of the various partial treatment effects and the composition of the population of interest. The first chapter explores the performance of common estimators as a function of the presence of heterogeneity in treatment effects and other characteristics that may influence their performance for estimating average treatment effects. The second chapter examines various approaches …


Bayesian Hierarchical Meta-Analysis Of Asymptomatic Ebola Seroprevalence, Peter Brody-Moore Jan 2019

Bayesian Hierarchical Meta-Analysis Of Asymptomatic Ebola Seroprevalence, Peter Brody-Moore

CMC Senior Theses

The continued study of asymptomatic Ebolavirus infection is necessary to develop a more complete understanding of Ebola transmission dynamics. This paper conducts a meta-analysis of eight studies that measure seroprevalence (the number of subjects that test positive for anti-Ebolavirus antibodies in their blood) in subjects with household exposure or known case-contact with Ebola, but that have shown no symptoms. In our two random effects Bayesian hierarchical models, we find estimated seroprevalences of 8.76% and 9.72%, significantly higher than the 3.3% found by a previous meta-analysis of these eight studies. We also produce a variation of this meta-analysis where we exclude …


Snap Scholar: The User Experience Of Engaging With Academic Research Through A Tappable Stories Medium, Ieva Burk Jan 2019

Snap Scholar: The User Experience Of Engaging With Academic Research Through A Tappable Stories Medium, Ieva Burk

CMC Senior Theses

With the shift to learn and consume information through our mobile devices, most academic research is still only presented in long-form text. The Stanford Scholar Initiative has explored the segment of content creation and consumption of academic research through video. However, there has been another popular shift in presenting information from various social media platforms and media outlets in the past few years. Snapchat and Instagram have introduced the concept of tappable “Stories” that have gained popularity in the realm of content consumption.

To accelerate the growth of the creation of these research talks, I propose an alternative to video: …


Predicting The Next Us President By Simulating The Electoral College, Boyan Kostadinov Jan 2018

Predicting The Next Us President By Simulating The Electoral College, Boyan Kostadinov

Journal of Humanistic Mathematics

We develop a simulation model for predicting the outcome of the US Presidential election based on simulating the distribution of the Electoral College. The simulation model has two parts: (a) estimating the probabilities for a given candidate to win each state and DC, based on state polls, and (b) estimating the probability that a given candidate will win at least 270 electoral votes, and thus win the White House. All simulations are coded using the high-level, open-source programming language R. One of the goals of this paper is to promote computational thinking in any STEM field by illustrating how probabilistic …


Gene × Environment Interaction: What Exactly Are We Talking About?, David S. Moore Jan 2018

Gene × Environment Interaction: What Exactly Are We Talking About?, David S. Moore

Pitzer Faculty Publications and Research

An ambiguity exists in how psychological scientists use the word “interaction.” This word can refer to physical interactions between components that constitute the mechanisms in complex systems, but it can also refer to statistical interactions revealed by General Linear Statistical Models (e.g., Analyses of Variance). Statistical interactions indicate that the nature of the relationship between two variables depends on a third variable, but the discovery of such interactions does not constitute evidence of physical interactions between components in a system. Studies conducted using traditional behavioral genetics methods sometimes reveal statistical interactions between genes and environments, but the presence or absence …


Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar Jan 2018

Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar

CMC Senior Theses

Understanding what factors influence wildlife movement allows landscape planners to make informed decisions that benefit both animals and humans. New quantitative methods, such as step-selection functions, provide valuable objective analyses of wildlife connectivity. This paper provides a framework for creating a step-selection function and demonstrates its use in a case study. The first section provides a general introduction about wildlife connectivity research. The second section explains the math behind the step-selection function using a simple example. The last section gives the results of a step-selection model for African buffalo in the Kavango Zambezi Transfrontier Conservation Area. Buffalo were found to …


Moneyball For Creative Writers: A Statistical Strategy For Publishing Your Work, Jon Wesick Feb 2017

Moneyball For Creative Writers: A Statistical Strategy For Publishing Your Work, Jon Wesick

Journal of Humanistic Mathematics

Writers face a challenge getting their poems and stories published. Rather than following the traditional strategy I model creative writing submission as a statistical process and explore the use of numerical metrics to maximize publications.


Quantum Foundations With Astronomical Photons, Calvin Leung Jan 2017

Quantum Foundations With Astronomical Photons, Calvin Leung

HMC Senior Theses

Bell's inequalities impose an upper limit on correlations between measurements of two-photon states under the assumption that the photons play by a set of local rules rather than by quantum mechanics. Quantum theory and decades of experiments both violate this limit.

Recent theoretical work in quantum foundations has demonstrated that a local realist model can explain the non-local correlations observed in experimental tests of Bell's inequality if the underlying probability distribution of the local hidden variable depends on the choice of measurement basis, or ``setting choice''. By using setting choices determined by astrophysical events in the distant past, it is …


The Document Similarity Network: A Novel Technique For Visualizing Relationships In Text Corpora, Dylan Baker Jan 2017

The Document Similarity Network: A Novel Technique For Visualizing Relationships In Text Corpora, Dylan Baker

HMC Senior Theses

With the abundance of written information available online, it is useful to be able to automatically synthesize and extract meaningful information from text corpora. We present a unique method for visualizing relationships between documents in a text corpus. By using Latent Dirichlet Allocation to extract topics from the corpus, we create a graph whose nodes represent individual documents and whose edge weights indicate the distance between topic distributions in documents. These edge lengths are then scaled using multidimensional scaling techniques, such that more similar documents are clustered together. Applying this method to several datasets, we demonstrate that these graphs are …


Machine Learning On Statistical Manifold, Bo Zhang Jan 2017

Machine Learning On Statistical Manifold, Bo Zhang

HMC Senior Theses

This senior thesis project explores and generalizes some fundamental machine learning algorithms from the Euclidean space to the statistical manifold, an abstract space in which each point is a probability distribution. In this thesis, we adapt the optimal separating hyperplane, the k-means clustering method, and the hierarchical clustering method for classifying and clustering probability distributions. In these modifications, we use the statistical distances as a measure of the dissimilarity between objects. We describe a situation where the clustering of probability distributions is needed and useful. We present many interesting and promising empirical clustering results, which demonstrate the statistical-distance-based clustering algorithms …


Gathering Steam In Health Care: A Student History, Michael J. Leach Nov 2016

Gathering Steam In Health Care: A Student History, Michael J. Leach

The Transdisciplinary STEAM+ Journal

In this reflection, I demonstrate STEAM in health care by outlining my 15 years as a university student engaged in formal education, extracurricular learning, research, and employment.


Teaching The Quandary Of Statistical Jurisprudence: A Review-Essay On Math On Trial By Schneps And Colmez, Noah Giansiracusa Jul 2016

Teaching The Quandary Of Statistical Jurisprudence: A Review-Essay On Math On Trial By Schneps And Colmez, Noah Giansiracusa

Journal of Humanistic Mathematics

This review-essay on the mother-and-daughter collaboration Math on Trial stems from my recent experience using this book as the basis for a college freshman seminar on the interactions between math and law. I discuss the strengths and weaknesses of this book as an accessible introduction to this enigmatic yet deeply important topic. For those considering teaching from this text (a highly recommended endeavor) I offer some curricular suggestions.


Applications Of Monte Carlo Methods In Statistical Inference Using Regression Analysis, Ji Young Huh Jan 2015

Applications Of Monte Carlo Methods In Statistical Inference Using Regression Analysis, Ji Young Huh

CMC Senior Theses

This paper studies the use of Monte Carlo simulation techniques in the field of econometrics, specifically statistical inference. First, I examine several estimators by deriving properties explicitly and generate their distributions through simulations. Here, simulations are used to illustrate and support the analytical results. Then, I look at test statistics where derivations are costly because of the sensitivity of their critical values to the data generating processes. Simulations here establish significance and necessity for drawing statistical inference. Overall, the paper examines when and how simulations are needed in studying econometric theories.