Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (10)
- Social and Behavioral Sciences (8)
- Economics (6)
- Applied Mathematics (4)
- Computer Sciences (3)
-
- Econometrics (3)
- Statistical Methodology (3)
- American Politics (2)
- Artificial Intelligence and Robotics (2)
- Behavioral Economics (2)
- Biostatistics (2)
- Business (2)
- Categorical Data Analysis (2)
- Data Science (2)
- Finance (2)
- Longitudinal Data Analysis and Time Series (2)
- Mathematics (2)
- Other Applied Mathematics (2)
- Political Science (2)
- Probability (2)
- Software Engineering (2)
- Theory and Algorithms (2)
- African American Studies (1)
- Arts and Humanities (1)
- Business Law, Public Responsibility, and Ethics (1)
- Community-Based Research (1)
- Corporate Finance (1)
- Keyword
-
- Statistics (2)
- 2024-presidential election (1)
- Academic papers (1)
- Adjusted statistics (1)
- African Buffalo (1)
-
- Algorithmic trading (1)
- Bayesian (1)
- Bayesian Linear Model (1)
- Bayesian Methods (1)
- Beta (1)
- Betting (1)
- Commodity futures (1)
- Conservation (1)
- Decentralized finance (1)
- Dickey-Fuller (1)
- Durbin-Watson (1)
- Ebola (1)
- Econometrics (1)
- Economics (1)
- Effects of EITC and Minimum Wage on Poverty levels by Race and Age (1)
- Efficient Market Hypothesis (1)
- Estimator (1)
- Evolutionary algorithms (1)
- Finance (1)
- Foreign exchange futures (1)
- Forward-Looking Beta (1)
- Game Theory (1)
- Gaming (1)
- Gibbs Sampler (1)
- HCI (1)
Articles 1 - 16 of 16
Full-Text Articles in Applied Statistics
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha
CMC Senior Theses
This thesis documents the design, deployment, and forward-test evaluation of an evolutionary multi-agent algorithmic trading system on Polymarket, the largest decentralized prediction market. The system pairs a locally-hosted 72-billion-parameter language model with a gradient-boosted statistical filter and an evolutionary selection mechanism that maintains a population of approximately 500 autonomous trading agents. Each agent generates a probability estimate for an event, compares it to the prevailing market price, and trades the resulting disagreement.
The central empirical exercise estimates a panel regression of trade-level profit on the absolute disagreement between the agent's probability estimate and the market price, controlling for agent identity, …
Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls, Jason Liang
Interpretable Sample Uncertainty Measures For Ranked-Choice Election Polls, Jason Liang
CMC Senior Theses
Polling results from traditional single-choice plurality elections are readily interpretable. Simple frequentist population parameters are estimated, including each candidate’s total support and the size of the front runner's lead. If the point estimate for the size of the front runner's lead exceeds the margin of error of the lead, we can conclude that the poll shows a statistically significant front runner. However, the interpretability of these population statistics disappears when applied to ranked-choice voting elections. Because ballots rank multiple candidates and candidates are eliminated in rounds, simple population-wide parameters are not well-defined. In RCV elections, a candidate’s ability to win …
Rank Rebalancing In Commodity And Foreign Exchange Markets, Prateek D. Vyas
Rank Rebalancing In Commodity And Foreign Exchange Markets, Prateek D. Vyas
CMC Senior Theses
This thesis empirically tests the rank-rebalancing mechanism of Stochastic Portfolio Theory (SPT) across commodity futures, foreign exchange futures, and equity ETFs. The Reverse Price-Weighted strategy (RPW) assigns, to each asset, the market weight of the asset at the opposite price rank, and generates an annualized excess return of 2.90% over the price-weighted (MKT) commodity benchmark, during the period of November 1977 to October 2025 (HAC t = 2.058, p = 0.040). The differential Sharpe ratio (dSharpe) of 0.245 is confirmed by a stationary block bootstrap, with a 𝑝-value of 0.015, and factor regressions controlling for carry, momentum, and value yield …
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
CMC Senior Theses
Over the past decades, the gaming industry has managed to evolve into a multi-billion-dollar enterprise. Gaming platforms such as Steam foster unprecedented amounts of engagement among players worldwide daily. In this thesis, we investigate the effect of incorporating sentiment-driven metrics, specifically YouTube view counts and positive reviews, into predictive models for game popularity. In addition, by comparing our linear regression sentiment-based approach to the Bayesian hierarchical folded normal model used by De Luisa et al. (2021), we can understand the many differences, strengths, and limitations of each methodology. In our thesis, we focus on three games. Each is of varying …
Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov
Forecasting Equity Betas Using Option-Implied Moments, Ivan Kolesnikov
CMC Senior Theses
Traditional beta estimates are constructed from historical stock‑and‑market returns and therefore adjust only as fast as realized data accrue. This thesis investigates whether the forward‑looking information embedded in equity‑option prices can enhance beta forecasts. Using near‑end‑of‑day quotes for 236 S&P 500 firms between 2007 and 2024, I extract risk‑neutral variance and skewness, construct five alternative beta estimators (historical, option‑implied, and three hybrids), and evaluate them against realized betas over six‑, twelve‑, and twenty‑four‑month windows. Rolling‑OLS beta remains the most accurate benchmark at short horizons, yet option‑implied moments add economically and statistically significant value when systematic exposure is expected to change …
Analyzing Political Sentiment On Micro-Blogging Data: A Lexicon And Machine Learning Approach To The 2024 U.S. Presidential Election, Ava Grey
CMC Senior Theses
This paper explores the trends in sentiment towards U.S. presidential candidates Kamala Harris and Donald Trump through micro-blogging social media text during the five months leading up to the election. Two datasets of varying sizes and origins were used to contextualize and validate analysis findings. The analyses include both a lexicon-based approach and a machine learning predictive method. Common sentiment analysis techniques like term frequency, term frequency inverse, various lexicons, and n-grams were utilized during the lexicon approach. During the modeling, a random forest was utilized in addition to the methods used during the lexicon approach. Results showed that overall …
Using Short Bursts To Optimize Redistricting In Georgia, Vedika Vishweshwar
Using Short Bursts To Optimize Redistricting In Georgia, Vedika Vishweshwar
CMC Senior Theses
Identifying extreme outliers in large state spaces is a difficult prob-
lem. I consider this problem in the context of finding political district-
ing plans that maximize the number of districts in which the majority
of the population is from a minority group, such as African Americans.
Since the set of all possible districting plans is enormous and unfeasi-
ble to examine in practice, this paper proposes a sampling method to
find these outlying plans. Specifically, this paper experiments with short
bursts in the context of minority voting rights in Georgia. Short bursts
are a type of Markov Chain in …
Information Prioritization: A Comparison Between Utility Maximizers And Probability Matchers, Yusuf Ismaeel
Information Prioritization: A Comparison Between Utility Maximizers And Probability Matchers, Yusuf Ismaeel
CMC Senior Theses
This thesis examines the differences between probability matchers and utility maximizers in their preferences for information sources in a lab environment. In this paper, we consider the best source of information to be the most connected one. We conducted several linear probability model type regressions along with logit regressions. Furthermore, we also attempted to control and fix any potential misclassifications in classifying the cognitive strategy by using instrumental variables. The results show that utility maximizers will almost always choose the most informed node. Probability matchers, on the other hand, do not exhibit such a behavior as the probability matching strategy …
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
CMC Senior Theses
In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …
Bayesian Hierarchical Meta-Analysis Of Asymptomatic Ebola Seroprevalence, Peter Brody-Moore
Bayesian Hierarchical Meta-Analysis Of Asymptomatic Ebola Seroprevalence, Peter Brody-Moore
CMC Senior Theses
The continued study of asymptomatic Ebolavirus infection is necessary to develop a more complete understanding of Ebola transmission dynamics. This paper conducts a meta-analysis of eight studies that measure seroprevalence (the number of subjects that test positive for anti-Ebolavirus antibodies in their blood) in subjects with household exposure or known case-contact with Ebola, but that have shown no symptoms. In our two random effects Bayesian hierarchical models, we find estimated seroprevalences of 8.76% and 9.72%, significantly higher than the 3.3% found by a previous meta-analysis of these eight studies. We also produce a variation of this meta-analysis where we exclude …
Snap Scholar: The User Experience Of Engaging With Academic Research Through A Tappable Stories Medium, Ieva Burk
CMC Senior Theses
With the shift to learn and consume information through our mobile devices, most academic research is still only presented in long-form text. The Stanford Scholar Initiative has explored the segment of content creation and consumption of academic research through video. However, there has been another popular shift in presenting information from various social media platforms and media outlets in the past few years. Snapchat and Instagram have introduced the concept of tappable “Stories” that have gained popularity in the realm of content consumption.
To accelerate the growth of the creation of these research talks, I propose an alternative to video: …
Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar
Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar
CMC Senior Theses
Understanding what factors influence wildlife movement allows landscape planners to make informed decisions that benefit both animals and humans. New quantitative methods, such as step-selection functions, provide valuable objective analyses of wildlife connectivity. This paper provides a framework for creating a step-selection function and demonstrates its use in a case study. The first section provides a general introduction about wildlife connectivity research. The second section explains the math behind the step-selection function using a simple example. The last section gives the results of a step-selection model for African buffalo in the Kavango Zambezi Transfrontier Conservation Area. Buffalo were found to …
Applications Of Monte Carlo Methods In Statistical Inference Using Regression Analysis, Ji Young Huh
Applications Of Monte Carlo Methods In Statistical Inference Using Regression Analysis, Ji Young Huh
CMC Senior Theses
This paper studies the use of Monte Carlo simulation techniques in the field of econometrics, specifically statistical inference. First, I examine several estimators by deriving properties explicitly and generate their distributions through simulations. Here, simulations are used to illustrate and support the analytical results. Then, I look at test statistics where derivations are costly because of the sensitivity of their critical values to the data generating processes. Simulations here establish significance and necessity for drawing statistical inference. Overall, the paper examines when and how simulations are needed in studying econometric theories.
Nfl Betting Market: Using Adjusted Statistics To Test Market Efficiency And Build A Betting Model, James P. Donnelly
Nfl Betting Market: Using Adjusted Statistics To Test Market Efficiency And Build A Betting Model, James P. Donnelly
CMC Senior Theses
The use of statistical analysis has been prevalent in the sports gambling industry for years. More recently, we have seen the emergence of "adjusted statistics", a more sophisticated way to examine each play and each result (further explanation below). And while adjusted statistics have become commonplace for professional and recreational bettors alike, little research has been done to justify their use. In this paper the effectiveness of this data is tested on the most heavily wagered sport in the world – the National Football League (NFL). The results are studied with two central questions in mind: Does the market account …
State Level Earned Income Tax Credit’S Effects On Race And Age: An Effective Poverty Reduction Policy, Anthony J. Barone
State Level Earned Income Tax Credit’S Effects On Race And Age: An Effective Poverty Reduction Policy, Anthony J. Barone
CMC Senior Theses
In this paper, I analyze the effectiveness of state level Earned Income Tax Credit programs on improving of poverty levels. I conducted this analysis for the years 1991 through 2011 using a panel data model with fixed effects. The main independent variables of interest were the state and federal EITC rates, minimum wage, gross state product, population, and unemployment all by state. I determined increases to the state EITC rates provided only a slight decrease to both the overall white below-poverty population and the corresponding white childhood population under 18, while both the overall and the under-18 black population for …
Applying Localized Realized Volatility Modeling To Futures Indices, Luella Fu
Applying Localized Realized Volatility Modeling To Futures Indices, Luella Fu
CMC Senior Theses
This thesis extends the application of the localized realized volatility model created by Ying Chen, Wolfgang Karl Härdle, and Uta Pigorsch to other futures markets, particularly the CAC 40 and the NI 225. The research attempted to replicate results though ultimately, those results were invalidated by procedural difficulties.