Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Statistics

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 31 - 60 of 412

Full-Text Articles in Statistics and Probability

“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King Jan 2025

“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King

Pomona Senior Theses

The work of this thesis is twofold — first, qualitatively characterizing the confluence between the British eugenics and statistics movements in the late 19th and early 20th centuries, and second, quantitatively analyzing the effect of this foundation on pedagogical materials in the growing field of statistics between 1880 and 1970. Towards the first goal, the history of the method of least squares, state statistics, and positive and negative eugenics are outlined, followed by a close reading of the foundational texts authored by Francis Galton and Karl Pearson that introduced linear regression. Towards the latter goal, English-language statistics textbooks published between …


Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri Jan 2025

Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri

CMC Senior Theses

Over the past decades, the gaming industry has managed to evolve into a multi-billion-dollar enterprise. Gaming platforms such as Steam foster unprecedented amounts of engagement among players worldwide daily. In this thesis, we investigate the effect of incorporating sentiment-driven metrics, specifically YouTube view counts and positive reviews, into predictive models for game popularity. In addition, by comparing our linear regression sentiment-based approach to the Bayesian hierarchical folded normal model used by De Luisa et al. (2021), we can understand the many differences, strengths, and limitations of each methodology. In our thesis, we focus on three games. Each is of varying …


Mathematics In Contemporary Society, Patrick J. Wallach Oct 2024

Mathematics In Contemporary Society, Patrick J. Wallach

Open Educational Resources

Mathematics in Contemporary Society is the textbook that corresponds to MA-321, the course of the same name. The course is designed to provide students with mathematical ideas and methods found in the social sciences, the arts, and in business. Topics will include fundamentals of statistics, scatterplots, graphics in the media, problem solving strategies, dimensional analysis, mathematics in music and art, and mathematical modeling. EXCEL is used to explore real world applications.


The Impact Of “Multiple Looks” When Performing Survival Analysis, Quentin Eloise Aug 2024

The Impact Of “Multiple Looks” When Performing Survival Analysis, Quentin Eloise

Electronic Theses and Dissertations

Survival analysis is a critical statistical method in healthcare to assess patient treatment effects and disease progression. Another critical area of statistical methodology in health care is the practice of adaptive designs. Adaptive designs allow for interim analyses to take place during a study and various decisions and actions can take place more ethically. This is beneficial for studies that take multiple years to complete and allows administrators and healthcare providers to make sound decisions as early as possible. A challenging aspect of adaptive designs is that the number of interim analyses is known in advance which is applicable in …


Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu Aug 2024

Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu

Statistical Science Theses and Dissertations

This dissertation investigates (1) Variational Bayesian Semi-supervised Keyword Extraction and (2) Variational Bayesian Multimodal Multiple Instance Classification.

The expansion of textual data, stemming from various sources such as online product reviews and scholarly publications on scientific discoveries, has created a demand for the extraction of succinct yet comprehensive information. As a result, in recent years, efforts have been spent in developing novel methodologies for keyword extraction. Although many methods have been proposed to automatically extract keywords in the contexts of both unsupervised and fully supervised learning, how to effectively use partially observed keywords, such as author-specified keywords, remains an under-explored …


Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny Aug 2024

Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny

All Theses

High blood pressure, also known as hypertension, significantly increases the risk of heart disease and stroke, which are leading causes of death in the United States. While contributing to over 691,000 deaths in 2021 alone in the United States (U.S.), it also imposes immense economic burden on the healthcare system, costing approximately $131 billion annually. One way to address this issue is for increased self-care behaviors and medication adherence, both of which require sufficient health literacy. Despite the importance of health literacy, 90% of U.S. adults struggle with health-related subjects. Overcoming the issues associated with health literacy requires addressing the …


Oh Statistics!, Heather L. Cook Jul 2024

Oh Statistics!, Heather L. Cook

Journal of Humanistic Mathematics

This poem was written about statistics and the usefulness thereof.


Learning Statistics With R: A Tutorial For Psychology Students And Other Beginners, Leslie Bain Jun 2024

Learning Statistics With R: A Tutorial For Psychology Students And Other Beginners, Leslie Bain

ATU Faculty OER Book Reviews

Review of OER Statistics textbook by Danielle Navarro, available at https://open.umn.edu/opentextbooks/textbooks/learning-statistics-with-r-a-tutorial-for-psychology-students-and-other-beginners


Introduction To Statistical Thinking, Leslie Bain Jun 2024

Introduction To Statistical Thinking, Leslie Bain

ATU Faculty OER Book Reviews

Review of OER Statistics textbook by Benjamin Yakir, available at https://open.umn.edu/opentextbooks/textbooks/introduction-to-statistical-thinking


The Impact Of Video Assistant Referee (Var) On The English Premier League, Jack Kenyon Brown Jun 2024

The Impact Of Video Assistant Referee (Var) On The English Premier League, Jack Kenyon Brown

Master's Theses

The aim of this study is to examine how the introduction of the Video Assisted Referee (VAR) system influenced the English Premier League (EPL). Since its implementation in the English Premier League in 2019, VAR has been a constant source of debate and controversy. Many studies have been done on the immediate impact of VAR on other elite professional soccer leagues, but the scope of results is very limited and due to be updated. The data for the ensuing analysis consists of 3800 matches played in the English Premier League during the five seasons before (14/15, 15/16, 16/17, 17/18, and …


Recursive Marix Game Analysis: Optimal, Simplified, And Human Strategies In Brave Rats, William A. Medwid Jun 2024

Recursive Marix Game Analysis: Optimal, Simplified, And Human Strategies In Brave Rats, William A. Medwid

Master's Theses

Brave Rats is a short game with simple rules, yet establishing a comprehensive strategy is very challenging without extensive computation. After explaining the rules, this paper begins by calculating the optimal strategy by recursively solving each turn’s Minimax strategy. It then provides summary statistics about the complex, branching Minimax solution. Next, we examine six other strategy models and evaluate their performance against each other. These models’ flaws highlight the key elements that contribute to the effectiveness of the Minimax strategy and offer insight into simpler strategies that human players could mimic. Finally, we analyze 123 games of human data collected …


Unraveling The History Of Deforestation In The Amazon Rainforest With Statistical Modeling, Ryan Destefano Jun 2024

Unraveling The History Of Deforestation In The Amazon Rainforest With Statistical Modeling, Ryan Destefano

Master's Theses

The Amazon rainforest, a vital ecosystem of immense biodiversity and global climate significance, faces the ongoing threat of deforestation driven by agricultural expansion. This thesis employs remote sensing techniques, focusing on the Enhanced Vegetation Index (EVI) derived from Landsat satellite imagery, to track land cover dynamics within the Amazon. The study examines historical land cover changes in current plantations in Peru and Brazil, regions where the exact timing of deforestation is uncertain. By analyzing EVI measurements dating back to 1984, inflection points indicative of deforestation events preceding plantation establishment are identified. Statistical modeling techniques, including spline fitting to analyze time …


Descriptions Of Interglacial Mastodons From Snowmass, Colorado, Connor White May 2024

Descriptions Of Interglacial Mastodons From Snowmass, Colorado, Connor White

Electronic Theses and Dissertations

The Ziegler Reservoir fossil site (ZRFS) in Colorado contains over 4000 mastodon bones that date from 140,000 to 100,000 years ago. At an elevation of ~2705 meters above sea level, ZRFS represents an alpine ecosystem dated to Marine Isotope Stage (MIS) 5. Formal descriptions of cheek teeth, mandibles, crania, and femora were completed. Statistical analyses of the upper and lower third molars, including a novel measurement of interloph(id) distances, indicate significant differences between ZRFS mastodons and Mammut pacificus, while falling within the ranges for Mammut americanum. This study agrees with the taxonomic assignment of ZRFS mastodons to Mammut …


Assessing Extant Methods For Generating G-Optimal Designs And A Novel Methodology To Compute The G-Score Of A Candidate Design, Hyrum John Hansen May 2024

Assessing Extant Methods For Generating G-Optimal Designs And A Novel Methodology To Compute The G-Score Of A Candidate Design, Hyrum John Hansen

All Graduate Theses and Dissertations, Fall 2023 to Present

Experimental designs are used by scientists to allocate treatments such that statistical inference is appropriate. Most traditional experimental designs have mathematical properties that make them desirable under certain conditions. Optimal experimental designs are those where the researcher can exercise total control over the treatment levels to maximize a chosen mathematical property. As is common in literature, the experimental design is represented as a matrix where each column represents a variable, and each row represents a trial. We define a function that takes as input the design matrix and outputs its score. We then algorithmically adjust each entry until a design …


Using Probability Theory To Calculate The Odds That Either Candidate Wins The 2024 Presidential Election, Andrew Ruggero Apr 2024

Using Probability Theory To Calculate The Odds That Either Candidate Wins The 2024 Presidential Election, Andrew Ruggero

Department of Applied Mathematics & Statistics Faculty Publications

In the U.S., where the electoral college is used to determine the votes of an election, calculating the odds of a president winning is not as simple as looking at the total vote percentages. With each state not exactly having a proportionally linear amount of votes per its population, the total percentage doesn’t mean much. As such, in order to calculate these odds, we must look at a variety of winning combinations per each candidate. First, we must take into account that swing states are the only ones that matter. Defined as a 5% difference between the two main candidates, …


A Survey Of The Murray State University Csis Department Of Student And Instructor Attitudes In Relation To Earlier Introduction Of Version Control Systems, Gavin Johnson Apr 2024

A Survey Of The Murray State University Csis Department Of Student And Instructor Attitudes In Relation To Earlier Introduction Of Version Control Systems, Gavin Johnson

Honors College Theses

Over the previous 20 years, the software development industry has overseen an evolution in application of Version Control Systems (VCS) from a Centralized Version Control System (CVCS) format to a Decentralized Version Control Format (DVCS). Examples of the former include Perforce and Subversion whilst the latter of the two include Github and BitBucket. As DVCS models allow software contributors to maintain their respective local repositories of relevant code bases, developers are able to work offline and maintain their work with relative fault tolerance. This contrasts to CVCS models, which require software contributors to be connected online to a main server. …


"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson Apr 2024

"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson

Senior Honors Theses

The authorship of Hebrews has been a point of contention for scholars for the past two millennia. While the epistle is traditionally attributed to Paul, many scholars assert that it carries thematic, structural, and stylistic differences from the remainder of his extant epistles; therefore, many other possible authors have been proposed. Of these, only Luke has other New Testament writings. Therefore, this project conducts a statistical comparison of Hebrews to the Pauline and Lukan corpora using stylometric authorial analysis methods. This analysis demonstrates that Hebrews is stylistically closer to Lukan literature than Pauline (but not to a significant degree), and …


Identifying Rural Health Clinics Within The Transformed Medicaid Statistical Information System (T-Msis) Analytic Files, Katherine Ahrens Mph, Phd, Zachariah Croll, Yvonne Jonk Phd, John Gale Ms, Heidi O'Connor Ms Mar 2024

Identifying Rural Health Clinics Within The Transformed Medicaid Statistical Information System (T-Msis) Analytic Files, Katherine Ahrens Mph, Phd, Zachariah Croll, Yvonne Jonk Phd, John Gale Ms, Heidi O'Connor Ms

Rural Health Clinics

Researchers at the Maine Rural Health Research Center describe a methodology for identifying Rural Health Clinic encounters within the Medicaid claims data using Transformed Medicaid Statistical Information System (T-MSIS) Analytic Files.

Background: There is limited information on the extent to which Rural Health Clinics (RHC) provide pediatric and pregnancy-related services to individuals enrolled in state Medicaid/CHIP programs. In part this is because methods to identify RHC encounters within Medicaid claims data are outdated.

Methods: We used a 100% sample of the 2018 Medicaid Demographic and Eligibility and Other Services Transformed Medicaid Statistical Information System (T-MSIS) Analytic Files for 20 states …


Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles Jan 2024

Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles

Honors Theses and Capstones

With the analytics revolution in sports in the past 20 years, it seems that everything that can be quantified is. In basketball though, trying to break the game down into a set of numbers comes with a unique problem. While we've come up with a good set of advanced numbers to measure offensive efficiency, defense is fundamentally harder to quantify. The game is played five on five, but it has often been popular or convenient to model defense as a set of five one on one games. As defenses became more complex into the 2010s, this methodology became more insignificant. …


Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky Jan 2024

Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky

Williams Honors College, Honors Research Projects

This project will examine the impact of using second-order terms in regression. For illustration, we use an example of regression where a baseball player's three by three heat map, including the height and distance from inside to outside of the pitch, are variables used to predict batting average. We find that second-order terms are crucial in discovering nonlinear relationships and interaction effects in regression models, and maintain that the common practice of using first-order additive models is insufficient.


To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis Jan 2024

To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis

Williams Honors College, Honors Research Projects

Regression to the mean is a statistical phenomenon that can hide important characteristics of what is truly happening in a research study. Caused by statistical randomness, regression to the mean occurs when extreme values, high or low, are followed by less extreme values. To correctly deal with it, one must understand what it is and how to distinguish its effect on conclusions made from the data. This paper provides examples of regression to the mean in both a medical and academic performance study and explains simple identifiers one can observe. Those are then followed up by the introduction of the …


Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn Jan 2024

Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn

Williams Honors College, Honors Research Projects

The random forest model proposed by Dr. Leo Breiman in 2001 is an ensemble machine learning method for classification prediction and regression. In the following paper, we will conduct an analysis on the random forest model with a focus on how the model works, how it is applied in software, and how it performs on a set of data. To fully understand the model, we will introduce the concept of decision trees, give a summary of the CART model, explain in detail how the random forest model operates, discuss how the model is implemented in software, demonstrate the model by …


Bayesian Analyses For Time-Varying Data And Opinion Dynamics, Torin Quinlivan Jan 2024

Bayesian Analyses For Time-Varying Data And Opinion Dynamics, Torin Quinlivan

Graduate Research Theses & Dissertations

We consider Bayesian analyses for time-varying data from the opinion dynamics in social networks and the processes of sensorimotor learning. Firstly, understanding the underlying opinions of social media users is a difficult process, as they must be understood filtered through the messages sent. To understand how to best estimate the latent opinions, we use a Bayesian approach to estimate various characteristics of the users and the network. We present a model for using Bayesian methods for this problem, along with a simulation study to demonstrate effectiveness. Secondly, incentivization with punishments or rewards may affect human skill learning. To investigate the …


Interpretable Word-Level Sentiment Analysis With Attention-Based Multiple Instance Classification Models, Chenyu Yang Dec 2023

Interpretable Word-Level Sentiment Analysis With Attention-Based Multiple Instance Classification Models, Chenyu Yang

Statistical Science Theses and Dissertations

In this study, our main objective is to tackle the black-box nature of popular machine learning models in sentiment analysis and enhance model interpretability. We aim to gain more insight into the decision-making process of sentiment analysis models, which is often obscure in those complex models. To achieve this goal, we introduce two word-level sentiment analysis models.

The first model is called the attention-based multiple instance classification (AMIC) model. It combines the transparent model structure of multiple instance classification and the self-attention mechanism in deep learning to incorporate the contextual information from documents. As demonstrated by a wine review dataset …


Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang Dec 2023

Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang

Statistical Science Theses and Dissertations

The human microbiome, comprising trillions of microorganisms, plays a pivotal role in modulating host physiology via molecular and metabolite exchanges. One of the major challenges in this field lies in the effective integration of microbiome and metabolomics data, an achievement that holds the promise of substantially enhancing the precision of disease prediction. However, many datasets prioritize microbiome data while neglecting paired metabolome information. Additionally, the prevalent analytical tools face challenges in effectively merging these intricate datasets, leading to possible misinterpretations and reduced prediction accuracies.

To address these challenges, the first part of this research introduces the Microbiome-based Supervised Contrastive Learning …


Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre Dec 2023

Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre

SMU Data Science Review

Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …


Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang Oct 2023

Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang

Statistical Science Theses and Dissertations

Spatially resolved transcriptomics (SRT) quantifies expression levels at different spatial locations, providing a new and powerful tool to investigate novel biological insights. As experimental technologies enhance both in capacity and efficiency, there arises a growing demand for the development of analytical methodologies.

One question in SRT data analysis is to identify genes whose expressions exhibit spatially correlated patterns, called spatially variable (SV) genes. Most current methods to identify SV genes are built upon the geostatistical model with Gaussian process, which could limit the models' ability to identify complex spatial patterns. In order to overcome this challenge and capture more types …


Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz Sep 2023

Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz

Beyond: Undergraduate Research Journal

When it comes to registering to vote, Hispanic voters can only register as “Hispanic” in the “Race/Ethnicity” category, causing difficulties when analyzing voting trends amongst the Hispanic community. Upon the recent idea that not all Hispanic Groups vote the same, the goal is to create a model that can possibly identify a voter’s Hispanic Group with the information provided on the public Florida voter file. This is accomplished using name and zip code data for all voters in Palm Beach, Florida. This paper will explore the model implemented, its findings and limitations. Palm Beach, Florida, is met with low confidence …


Traditional Vs Machine Learning Approaches: A Comparison Of Time Series Modeling Methods, Miguel E. Bonilla Jr., Jason Mcdonald, Tamas Toth, Bivin Sadler Aug 2023

Traditional Vs Machine Learning Approaches: A Comparison Of Time Series Modeling Methods, Miguel E. Bonilla Jr., Jason Mcdonald, Tamas Toth, Bivin Sadler

SMU Data Science Review

In recent years, various new Machine Learning and Deep Learning algorithms have been introduced, claiming to offer better performance than traditional statistical approaches when forecasting time series. Studies seeking evidence to support the usage of ML/DL over statistical approaches have been limited to comparing the forecasting performance of univariate, linear time series data. This research compares the performance of traditional statistical-based and ML/DL methods for forecasting multivariate and nonlinear time series.


Sentiment Analysis Before And During The Covid-19 Pandemic, Emily Musgrove Jul 2023

Sentiment Analysis Before And During The Covid-19 Pandemic, Emily Musgrove

Mathematics Summer Fellows

This study examines the change in connotative language use before and during the Covid-19 pandemic. By analyzing news articles from several major US newspapers, we found that there is a statistically significant correlation between the sentiment of the text and the publication period. Specifically, we document a large, systematic, and statistically significant decline in the overall sentiment of articles published in major news outlets. While our results do not directly gauge the sentiment of the population, our findings have important implications regarding the social responsibility of journalists and media outlets especially in times of crisis.