Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Mathematics (45)
- Social and Behavioral Sciences (37)
- Applied Statistics (36)
- Probability (24)
- Statistical Models (20)
-
- Economics (19)
- Applied Mathematics (17)
- Arts and Humanities (16)
- Computer Sciences (13)
- Econometrics (12)
- Medicine and Health Sciences (11)
- Data Science (10)
- Education (10)
- Statistical Methodology (9)
- Other Mathematics (8)
- Other Statistics and Probability (8)
- Psychology (8)
- Biostatistics (7)
- Finance (7)
- Life Sciences (7)
- Other Applied Mathematics (7)
- Artificial Intelligence and Robotics (6)
- Categorical Data Analysis (6)
- Longitudinal Data Analysis and Time Series (6)
- Multivariate Analysis (6)
- Political Science (6)
- Sociology (6)
- Design of Experiments and Sample Surveys (5)
- Keyword
-
- Statistics (16)
- Probability (9)
- Discrepancy (5)
- Machine Learning (5)
- Random walk (4)
-
- Mathematics (3)
- Monte Carlo (3)
- Sentiment Analysis (3)
- 62-07 (2)
- Algorithmic trading (2)
- Bayesian (2)
- Bayesian analysis (2)
- Big Data (2)
- Big data (2)
- Bounds (2)
- DNA (2)
- Data analysis (2)
- Data mining (2)
- Data reduction (2)
- Econometrics (2)
- Education (2)
- Factor analysis (2)
- Internet in higher education (2)
- Machine learning (2)
- Markov chains (2)
- Microarray (2)
- Microarrays (2)
- PCA (2)
- Pedagogy (2)
- Poetry (2)
- Publication Year
- Publication
-
- CMC Senior Theses (28)
- Journal of Humanistic Mathematics (22)
- All HMC Faculty Publications and Research (13)
- HMC Senior Theses (11)
- Pomona Faculty Publications and Research (11)
-
- CGU Theses & Dissertations (9)
- CGU Faculty Publications and Research (5)
- Pomona Economics (5)
- Humanistic Mathematics Network Journal (3)
- Pitzer Senior Theses (3)
- Scripps Senior Theses (3)
- CODEE Journal (2)
- Pomona Senior Theses (2)
- The Transdisciplinary STEAM+ Journal (2)
- Pitzer Faculty Publications and Research (1)
- Publication Type
Articles 31 - 60 of 120
Full-Text Articles in Statistics and Probability
A Gender And Race Theoretical And Probabilistic Analysis Of The Recent Title Ix Policy Changes, Jordan Wellington
A Gender And Race Theoretical And Probabilistic Analysis Of The Recent Title Ix Policy Changes, Jordan Wellington
Scripps Senior Theses
On May 6th, 2020, after extensive public comment and review, the Department of Education published the final rule for the new Title IX regulations, which took effect in schools on August 14th. Title IX is the nearly fifty year old piece of the Education Amendments that prohibits sexual discrimination in federally funded schools. Several of these changes, such as the inclusion of live hearings and cross examination of witnesses, have been widely criticized by victims’ rights advocates for potentially retraumatizing victims of sexual assault and discouraging students from pursuing a Title IX claim. While the impact of the new regulations …
Information Prioritization: A Comparison Between Utility Maximizers And Probability Matchers, Yusuf Ismaeel
Information Prioritization: A Comparison Between Utility Maximizers And Probability Matchers, Yusuf Ismaeel
CMC Senior Theses
This thesis examines the differences between probability matchers and utility maximizers in their preferences for information sources in a lab environment. In this paper, we consider the best source of information to be the most connected one. We conducted several linear probability model type regressions along with logit regressions. Furthermore, we also attempted to control and fix any potential misclassifications in classifying the cognitive strategy by using instrumental variables. The results show that utility maximizers will almost always choose the most informed node. Probability matchers, on the other hand, do not exhibit such a behavior as the probability matching strategy …
Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee
Feature Investigation For Stock Returns Prediction Using Xgboost And Deep Learning Sentiment Classification, Seungho (Samuel) Lee
CMC Senior Theses
This paper attempts to quantify predictive power of social media sentiment and financial data in stock prediction by utilizing a comprehensive set of stock-related fundamental and technical variables and social media sentiments. For conducting sentiment analysis, this study employs a pretrained finBERT model that provides three different sentiment classifications and respective softmax scores. Hence, the significance of these variables is evaluated with XGBoost regression and Shapley Additive exPlanations (SHAP) frameworks. Through investigating feature importance, this study finds that statistical properties of sentiment variables provide a stronger predictive power than a weighted sentiment score and that it is possible to quantify …
Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard
Using Twitter Api To Solve The Goat Debate: Michael Jordan Vs. Lebron James, Jordan Trey Leonard
CMC Senior Theses
Using a Twitter API, I gather and analyze tweets by performing sentiment analysis to solve the GOAT debate among professional athletes with the primary focus on comparing Michael Jordan and LeBron James. Athletes from the National Football League (NFL), the National Basketball Association (NBA), Major League Baseball (MLB), and the National Collegiate Athletic Association (NCAA) Division 1 Men's and Women's Basketball were selected to compare how sentiment polarity varies across sports. Sentiment polarity is measured by labeling text as "positive", "neutral", or "negative" which allows us to determine which athlete/sport is highly favored among the Twitter community when it comes …
An Evaluation Of Knot Placement Strategies For Spline Regression, William Klein
An Evaluation Of Knot Placement Strategies For Spline Regression, William Klein
CMC Senior Theses
Regression splines have an established value for producing quality fit at a relatively low-degree polynomial. This paper explores the implications of adopting new methods for knot selection in tandem with established methodology from the current literature. Structural features of generated datasets, as well as residuals collected from sequential iterative models are used to augment the equidistant knot selection process. From analyzing a simulated dataset and an application onto the Racial Animus dataset, I find that a B-spline basis paired with equally-spaced knots remains the best choice when data are evenly distributed, even when structural features of a dataset are known …
Quantifying Controllability In Temporal Networks With Uncertainty, James C. Boerkoel Jr., Lindsay Popowski, Michael Gao, Hemeng Li, Savana Ammons, Shyan Akmal
Quantifying Controllability In Temporal Networks With Uncertainty, James C. Boerkoel Jr., Lindsay Popowski, Michael Gao, Hemeng Li, Savana Ammons, Shyan Akmal
All HMC Faculty Publications and Research
Controllability for Simple Temporal Networks with Uncertainty (STNUs) has thus far been limited to three levels: strong, dynamic, and weak. Because of this, there is currently no systematic way for an agent to assess just how far from being controllable an uncontrollable STNU is. We provide new insights inspired by a geometric interpretation of STNUs to introduce the degrees of strong and dynamic controllability - continuous metrics that measure how far a network is from being controllable. We utilize these metrics to approximate the probabilities that an STNU can be dispatched successfully offline and online respectively. We introduce new methods …
Dynamic Control Of Probabilistic Simple Temporal Networks, James C. Boerkoel Jr., Michael Gao, Lindsay Popowski
Dynamic Control Of Probabilistic Simple Temporal Networks, James C. Boerkoel Jr., Michael Gao, Lindsay Popowski
All HMC Faculty Publications and Research
The controllability of a temporal network is defined as an agent’s ability to navigate around the uncertainty in its schedule and is well-studied for certain networks of temporal constraints. However, many interesting real-world problems can be better represented as Probabilistic Simple Temporal Networks (PSTNs) in which the uncertain durations are represented using potentially-unbounded probability density functions. This can make it inherently impossible to control for all eventualities. In this paper, we propose two new dynamic controllability algorithms that attempt to maximize the likelihood of successfully executing a schedule within a PSTN. The first approach, which we call MIN-LOSS DC, finds …
Causal Effect Random Forest Of Interaction Trees For Learning Individualized Treatment Regimes In Observational Studies: With Applications To Education Study Data, Luo Li
CGU Theses & Dissertations
Learning individualized treatment regimes (ITR) using observational data holds great interest in various fields, as treatment recommendations based on individual characteristics may improve individual treatment benefits with a reduced cost. It has long been observed that different individuals may respond to a certain treatment with significant heterogeneity. ITR can be defined as a mapping between individual characteristics to a treatment assignment. The optimal ITR is the treatment assignment that maximizes expected individual treatment effects. Rooted from personalized medicine, many studies and applications of ITR are in medical fields and clinical practice. Heterogeneous responses are also well documented in educational interventions. …
A Multinational Study Of The Etiology And Clinical Teleology Of Moral Evaluations Of Patient Behaviors, Anna Yu Lee
A Multinational Study Of The Etiology And Clinical Teleology Of Moral Evaluations Of Patient Behaviors, Anna Yu Lee
CGU Theses & Dissertations
This dissertation is a collection of four studies which collectively explore a hypothesized construct of ‘moral evaluation of patient behaviors’ (MEPB) as a driver of health professionals’ readiness to interact humanistically with their patients. In these studies, ‘humanistic interactions’ refer to the non-technical, intangible skills and factors of clinical competence; the factors specifically explored in these studies were compassion toward patients, self-efficacy for treating patients, and optimism toward patient treatment. For the purpose of specificity, all factors were examined as they pertained to patients with substance use disorders. Survey data from a convenience sample of 524 health professionals (i.e. physicians, …
Novel Random Forest Methods And Algorithms For Autism Spectrum Disorders Research, Afrooz Jahedi
Novel Random Forest Methods And Algorithms For Autism Spectrum Disorders Research, Afrooz Jahedi
CGU Theses & Dissertations
Random Forest (RF) is a flexible, easy to use machine learning algorithm that was proposed by Leo Breiman in 2001 for building a predictor ensemble with a set of decision trees that grow in randomly selected subspaces of data. Its superior prediction accuracy has made it the most used algorithms in the machine learning field. In this dissertation, we use the random forest as the main building block for creating a proximity matrix for multivariate matching and diagnostic classification problems that are used for autism research (as an exemplary application). In observational studies, matching is used to optimize the balance …
K-Means Stock Clustering Analysis Based On Historical Price Movements And Financial Ratios, Shu Bin
K-Means Stock Clustering Analysis Based On Historical Price Movements And Financial Ratios, Shu Bin
CMC Senior Theses
The 2015 article Creating Diversified Portfolios Using Cluster Analysis proposes an algorithm that uses the Sharpe ratio and results from K-means clustering conducted on companies' historical financial ratios to generate stock market portfolios. This project seeks to evaluate the performance of the portfolio-building algorithm during the beginning period of the COVID-19 recession. S&P 500 companies' historical stock price movement and their historical return on assets and asset turnover ratios are used as dissimilarity metrics for K-means clustering. After clustering, stock with the highest Sharpe ratio from each cluster is picked to become a part of the portfolio. The economic and …
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
How Machine Learning And Probability Concepts Can Improve Nba Player Evaluation, Harrison Miller
CMC Senior Theses
In this paper I will be breaking down a scholarly article, written by Sameer K. Deshpande and Shane T. Jensen, that proposed a new method to evaluate NBA players. The NBA is the highest level professional basketball league in America and stands for the National Basketball Association. They proposed to build a model that would result in how NBA players impact their teams chances of winning a game, using machine learning and probability concepts. I preface that by diving into these concepts and their mathematical backgrounds. These concepts include building a linear model using ordinary least squares method, the bias …
Mathematics Versus Statistics, Mindy B. Capaldi
Mathematics Versus Statistics, Mindy B. Capaldi
Journal of Humanistic Mathematics
Mathematics and statistics are both important and useful subjects, but the former has maintained prominence in the American education system. On the other hand, statistics is more prevalent in daily life and is an increasingly marketable subject to know. This article gives a personal history of one mathematician’s bumpy road to learning and teaching statistics. Additionally, arguments for how and why to include statistics in the K-12 and college curricula are provided.
Choose Your Own Adventure: An Analysis Of Interactive Gamebooks Using Graph Theory, D'Andre Adams, Daniela Beckelhymer, Alison Marr
Choose Your Own Adventure: An Analysis Of Interactive Gamebooks Using Graph Theory, D'Andre Adams, Daniela Beckelhymer, Alison Marr
Journal of Humanistic Mathematics
"BEWARE and WARNING! This book is different from other books. You and YOU ALONE are in charge of what happens in this story." This is the captivating introduction to every book in the interactive novel series, Choose Your Own Adventure (CYOA). Our project uses the mathematical field of graph theory to analyze forty books from the CYOA book series for ages 9-12. We first began by drawing the digraphs of each book. Then we analyzed these digraphs by collecting structural data such as longest path length (i.e. longest story length) and number of vertices with outdegree zero (i.e. number …
The Principal Problem With Principal Components Regression, Gary N. Smith, Heidi Margaret Artigue
The Principal Problem With Principal Components Regression, Gary N. Smith, Heidi Margaret Artigue
Pomona Economics
No abstract provided.
Step Away From Stepwise, Gary N. Smith
Step Away From Stepwise, Gary N. Smith
Pomona Economics
Stepwise regression is a popular data-mining tool that uses statistical significance to select the explanatory variables to be used in a multiple-regression model. A fundamental problem with stepwise regression is that some real explanatory variables that have causal effects on the dependent variable may happen to not be statistically significant, while nuisance variables may be coincidentally significant. As a result, the model may fit the data well in-sample, but do poorly out-of-sample. Many Big-Data researchers believe that, the larger the number of possible explanatory variables, the more useful is stepwise regression for selecting explanatory variables. The reality is that stepwise …
Using Neural Networks To Classify Discrete Circular Probability Distributions, Madelyn Gaumer
Using Neural Networks To Classify Discrete Circular Probability Distributions, Madelyn Gaumer
HMC Senior Theses
Given the rise in the application of neural networks to all sorts of interesting problems, it seems natural to apply them to statistical tests. This senior thesis studies whether neural networks built to classify discrete circular probability distributions can outperform a class of well-known statistical tests for uniformity for discrete circular data that includes the Rayleigh Test1, the Watson Test2, and the Ajne Test3. Each neural network used is relatively small with no more than 3 layers: an input layer taking in discrete data sets on a circle, a hidden layer, and an output …
Be Wary Of Black-Box Trading Algorithms, Gary N. Smith
Be Wary Of Black-Box Trading Algorithms, Gary N. Smith
Pomona Economics
Black-box algorithms now account for nearly a third of all U. S. stock trades. It is a mistake to think that these algorithms possess superhuman intelligence. In reality, computers do not have the common sense and wisdom that humans have accumulated by living. Trading algorithms are particularly dangerous because they are so efficient at discovering statistical patterns—but so utterly useless in judging whether the discovered patterns are meaningful.
The Paradox Of Big Data, Gary N. Smith
The Paradox Of Big Data, Gary N. Smith
Pomona Economics
Data-mining is often used to discover patterns in Big Data. It is tempting believe that because an unearthed pattern is unusual it must be meaningful, but patterns are inevitable in Big Data and usually meaningless. The paradox of Big Data is that data mining is most seductive when there are a large number of variables, but a large number of variables exacerbates the perils of data mining.
On Cluster Robust Models, José Bayoán Santiago Calderón
On Cluster Robust Models, José Bayoán Santiago Calderón
CGU Theses & Dissertations
Cluster robust models are a kind of statistical models that attempt to estimate parameters considering potential heterogeneity in treatment effects. Absent heterogeneity in treatment effects, the partial and average treatment effect are the same. When heterogeneity in treatment effects occurs, the average treatment effect is a function of the various partial treatment effects and the composition of the population of interest. The first chapter explores the performance of common estimators as a function of the presence of heterogeneity in treatment effects and other characteristics that may influence their performance for estimating average treatment effects. The second chapter examines various approaches …
Bayesian Hierarchical Meta-Analysis Of Asymptomatic Ebola Seroprevalence, Peter Brody-Moore
Bayesian Hierarchical Meta-Analysis Of Asymptomatic Ebola Seroprevalence, Peter Brody-Moore
CMC Senior Theses
The continued study of asymptomatic Ebolavirus infection is necessary to develop a more complete understanding of Ebola transmission dynamics. This paper conducts a meta-analysis of eight studies that measure seroprevalence (the number of subjects that test positive for anti-Ebolavirus antibodies in their blood) in subjects with household exposure or known case-contact with Ebola, but that have shown no symptoms. In our two random effects Bayesian hierarchical models, we find estimated seroprevalences of 8.76% and 9.72%, significantly higher than the 3.3% found by a previous meta-analysis of these eight studies. We also produce a variation of this meta-analysis where we exclude …
A Tacticians Guide To Conflict, Vol. 1: Advancing Explanations & Predictions Of Intrastate Conflict, Khaled Eid
A Tacticians Guide To Conflict, Vol. 1: Advancing Explanations & Predictions Of Intrastate Conflict, Khaled Eid
CGU Theses & Dissertations
Intrastate conflict is an ever-evolving problem – causes, explanation, and predictions are increasingly murky as traditional methods of analysis focus on structural issues as precursors of conflict. Often times these theories do not consider the underlying meso and micro dynamics that can provide vital insights into the phenomena. Tactical decision-makers are left using models that rely on highly aggregated, country level data to create proper courses of actions (COAs) to address or predict conflict. The shortcoming is that conflicts morph quite rapidly and structural variables can struggle capture such dynamic changes. To address this some tacticians are using big data …
Snap Scholar: The User Experience Of Engaging With Academic Research Through A Tappable Stories Medium, Ieva Burk
CMC Senior Theses
With the shift to learn and consume information through our mobile devices, most academic research is still only presented in long-form text. The Stanford Scholar Initiative has explored the segment of content creation and consumption of academic research through video. However, there has been another popular shift in presenting information from various social media platforms and media outlets in the past few years. Snapchat and Instagram have introduced the concept of tappable “Stories” that have gained popularity in the realm of content consumption.
To accelerate the growth of the creation of these research talks, I propose an alternative to video: …
The Principal Problem With Principal Components Regression, Heidi Margaret Artigue, Gary Smith
The Principal Problem With Principal Components Regression, Heidi Margaret Artigue, Gary Smith
Pomona Faculty Publications and Research
Principal components regression (PCR) reduces a large number of explanatory variables down to a small number of principal components. PCR is thought to be more useful, the more numerous the potential explanatory variables. The reality is that a large number of candidate explanatory variables does not make PCR more valuable; instead, it magnifies the failings of PCR.
A Math Research Project Inspired By Twin Motherhood, Tiffany N. Kolba
A Math Research Project Inspired By Twin Motherhood, Tiffany N. Kolba
Journal of Humanistic Mathematics
The phenomenon of twins, triplets, quadruplets, and other higher order multiples has fascinated humans for centuries and has even captured the attention of mathematicians who have sought to model the probabilities of multiple births. However, there has not been extensive research into the phenomenon of polyovulation, which is one of the biological mechanisms that produces multiple births. In this paper, I describe how my own experience becoming a mother to twins led me on a quest to better understand the scientific processes going on inside my own body and motivated me to conduct research on polyovulation frequencies. An overview of …
Predicting The Next Us President By Simulating The Electoral College, Boyan Kostadinov
Predicting The Next Us President By Simulating The Electoral College, Boyan Kostadinov
Journal of Humanistic Mathematics
We develop a simulation model for predicting the outcome of the US Presidential election based on simulating the distribution of the Electoral College. The simulation model has two parts: (a) estimating the probabilities for a given candidate to win each state and DC, based on state polls, and (b) estimating the probability that a given candidate will win at least 270 electoral votes, and thus win the White House. All simulations are coded using the high-level, open-source programming language R. One of the goals of this paper is to promote computational thinking in any STEM field by illustrating how probabilistic …
Gene × Environment Interaction: What Exactly Are We Talking About?, David S. Moore
Gene × Environment Interaction: What Exactly Are We Talking About?, David S. Moore
Pitzer Faculty Publications and Research
An ambiguity exists in how psychological scientists use the word “interaction.” This word can refer to physical interactions between components that constitute the mechanisms in complex systems, but it can also refer to statistical interactions revealed by General Linear Statistical Models (e.g., Analyses of Variance). Statistical interactions indicate that the nature of the relationship between two variables depends on a third variable, but the discovery of such interactions does not constitute evidence of physical interactions between components in a system. Studies conducted using traditional behavioral genetics methods sometimes reveal statistical interactions between genes and environments, but the presence or absence …
Sequential Probing With A Random Start, Joshua Miller
Sequential Probing With A Random Start, Joshua Miller
HMC Senior Theses
Processing user requests quickly requires not only fast servers, but also demands methods to quickly locate idle servers to process those requests. Methods of finding idle servers are analogous to open addressing in hash tables, but with the key difference that servers may return to an idle state after having been busy rather than staying busy. Probing sequences for open addressing are well-studied, but algorithms for locating idle servers are less understood. We investigate sequential probing with a random start as a method for finding idle servers, especially in cases of heavy traffic. We present a procedure for finding the …
Iterative Matrix Factorization Method For Social Media Data Location Prediction, Natchanon Suaysom
Iterative Matrix Factorization Method For Social Media Data Location Prediction, Natchanon Suaysom
HMC Senior Theses
Since some of the location of where the users posted their tweets collected by social media company have varied accuracy, and some are missing. We want to use those tweets with highest accuracy to help fill in the data of those tweets with incomplete information. To test our algorithm, we used the sets of social media data from a city, we separated them into training sets, where we know all the information, and the testing sets, where we intentionally pretend to not know the location. One prediction method that was used in (Dukler, Han and Wang, 2016) requires appending one-hot …
Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar
Step-Selection Functions For Modeling Animal Movement -- Case Study: African Buffalo, Maia Adar
CMC Senior Theses
Understanding what factors influence wildlife movement allows landscape planners to make informed decisions that benefit both animals and humans. New quantitative methods, such as step-selection functions, provide valuable objective analyses of wildlife connectivity. This paper provides a framework for creating a step-selection function and demonstrates its use in a case study. The first section provides a general introduction about wildlife connectivity research. The second section explains the math behind the step-selection function using a simple example. The last section gives the results of a step-selection model for African buffalo in the Kavango Zambezi Transfrontier Conservation Area. Buffalo were found to …