Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (41)
- Social and Behavioral Sciences (27)
- Statistical Methodology (25)
- Biostatistics (19)
- Mathematics (19)
-
- Categorical Data Analysis (17)
- Computer Sciences (15)
- Data Science (15)
- Statistical Theory (15)
- Other Statistics and Probability (13)
- Probability (13)
- Applied Mathematics (12)
- Design of Experiments and Sample Surveys (10)
- Engineering (9)
- Medicine and Health Sciences (8)
- Multivariate Analysis (8)
- Arts and Humanities (6)
- Other Applied Mathematics (6)
- Other Mathematics (6)
- Psychology (6)
- Survival Analysis (6)
- Business (5)
- Education (5)
- Life Sciences (5)
- Analysis (4)
- Artificial Intelligence and Robotics (4)
- Communication (4)
- Institution
-
- Southern Methodist University (16)
- California Polytechnic State University, San Luis Obispo (14)
- The University of Akron (8)
- Wayne State University (6)
- Claremont Colleges (5)
-
- University of Nebraska - Lincoln (5)
- Utah State University (5)
- Western Kentucky University (5)
- University of Arkansas, Fayetteville (4)
- University of South Carolina (3)
- Ursinus College (3)
- City University of New York (CUNY) (2)
- Embry-Riddle Aeronautical University (2)
- Georgia College (2)
- Georgia Southern University (2)
- Misericordia University (2)
- University of Kentucky (2)
- University of Nevada, Las Vegas (2)
- University of New Hampshire (2)
- University of South Florida (2)
- University of Southern Maine (2)
- Air Force Institute of Technology (1)
- Belmont University (1)
- Bowling Green State University (1)
- Bridgewater State University (1)
- Chapman University (1)
- Clemson University (1)
- East Tennessee State University (1)
- Eastern Washington University (1)
- Edith Cowan University (1)
- Publication Year
- Publication
-
- Statistical Science Theses and Dissertations (14)
- Statistics (8)
- Williams Honors College, Honors Research Projects (8)
- Journal of Modern Applied Statistical Methods (6)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (3)
-
- Electronic Theses and Dissertations (3)
- Graduate Theses and Dissertations (3)
- Master's Theses (3)
- Senior Theses (3)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (2)
- CMC Senior Theses (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Department of Industrial and Management Systems Engineering: Instructional Materials (2)
- Honors Theses and Capstones (2)
- Masters Theses & Specialist Projects (2)
- Numeracy (2)
- Physics (2)
- Student Research Poster Presentations 2020 (2)
- Theses and Dissertations (2)
- 2026 Symposium (1)
- All Theses (1)
- Beyond: Undergraduate Research Journal (1)
- Civil and Environmental Engineering Theses and Dissertations (1)
- DU Undergraduate Research Journal Archive (1)
- Department of Applied Mathematics & Statistics Faculty Publications (1)
- Department of Mathematics Publications (1)
- Department of Statistics: Dissertations, Theses, and Student Research (1)
- Departmental Technical Reports (CS) (1)
- Dissertations, Theses, and Capstone Projects (1)
- Doctor of Business Administration Dissertations (1)
- Publication Type
- File Type
Articles 1 - 30 of 122
Full-Text Articles in Applied Statistics
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Methodologies For Count Time Series Analysis And Topological Data Analysis Of Medical Images, Yuhyeong Jang
Statistical Science Theses and Dissertations
This dissertation addresses two distinct topics related to count time series analysis and topological medical image analysis, respectively. The first part of the dissertation comprises an application of a count time series model to analysis of US monthly sex trafficking data and development of a new model for multivariate count data that exhibits serial dependence and overdispersion. By imposing a family of multivariate mixed Poisson distributions on the count random vector, the proposed model can accommodate a broad range of overdispersion as well as positive contemporaneous correlations. For maximum likelihood estimation, a computationally feasible EM-type algorithm is derived based on …
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
Civil and Environmental Engineering Theses and Dissertations
Urban areas are increasingly exposed to natural hazards while accommodating a growing share of the global population, yet a consistent science-based framework for quantifying urban and community resilience remains lacking. This dissertation develops a physics-based analytical framework grounded in statistical mechanics and the quantitative theory of Brownian motion. A city is conceptualized as a complex medium in which citizens move analogously to Brownian particles within a viscoelastic environment, influenced by socioeconomic interactions and infrastructure functionality.
A central premise is that urban resilience, interpreted as engineering resilience (an outcome), can be quantified through a single metric: the mean-square displacement MSD=⟨r²(t)⟩, of …
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
A Statistical Analysis Of Current And Future Hurricane Activity In The North Indian Ocean, Basil Lund
A Statistical Analysis Of Current And Future Hurricane Activity In The North Indian Ocean, Basil Lund
2026 Symposium
A hurricane is defined as a tropical storm with winds sustained at 74 mph or greater. I examined major (category 3 and above) hurricane activity over the North Indian Ocean from the years 1972-2019 as reported by Colorado State University Hurricane Forecast Archive. Using RStudio, I conducted a binomial analysis of the CSU dataset to calculate probabilities of zero to ten years with one or more major North Indian Ocean hurricanes in the next decade. I conducted a geometric analysis to determine probabilities associated with waiting periods for the next year with a major hurricane, as well as a Poisson …
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
Bayesball : A Comprehensive Framework For Predicting Ucl Injury, Brady M. Pinter, Will Best Ph.D.
SPARK Symposium Presentations
Ulnar Collateral Ligament (UCL) reconstruction, commonly referred to as Tommy John Surgery, has seen a significant rise among Major League Baseball (MLB) pitchers, prompting growing interest in identifying the mechanical and performance-based factors that contribute to injury risk. While previous studies have examined these relationships using traditional frequentist approaches separately, this study combines multiple different model techniques to present a broad framework for finding significant predictors of UCL Surgery. These models include Lasso and Ridge Regression, Principal Component Regression (PCR) , Partial Least Squares Regression (PLS) , Random Forest, Multiple Linear Regression, and a Bayesian Statistical Model. Using these models, …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
Honors Undergraduate Theses
With their deserts, castles, and ghost houses, the environments of Super Mario games are colorful, whimsical, and charming, but why are they so compelling, and what happens when our analysis of these environments extends beyond individual levels to expansive game worlds? Drawing on Cresswell’s theory of place (2014) and recent work on musical place-building in Mario Kart 8 (Heazlewood-Dale, 2024), I propose a spectrum between localized and globalized scale in games. As game environments become increasingly globalized, the music may be similarly altered to account for this shift in scale. Consequently, players may then encounter a broader, less musically congruent …
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Methods For Joint Outcome Modeling And Dynamic Assessment Of Recurrent Events, Zifang Kong
Statistical Science Theses and Dissertations
Recurrent event data frequently arise in clinical studies where individuals experience repeated, possibly related, events over time. These data are often accompanied by sparse and irregular longitudinal measurements, creating challenges for traditional joint modeling approaches that struggle to account for time-dependent associations and within-subject correlations. We propose FRAILTY (Functional Regression with AutoRegressIve fraiLTY), a novel two-step framework that integrates functional principal component analysis (PACE) with a dynamic frailty model featuring autoregressive structure. FRAILTY accommodates both scalar and functional predictors and captures within-subject dependence across recurrent events. To further extend its utility, we develop a multivariate joint modeling framework that simultaneously …
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Open Educational Resources
Data analysis using standard statistical methods and relevant computer software. Emphasis on real-world data, interpretation, and misinterpretation of computer output.
This syllabus contains open source notebook about data analysis content.
Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo
Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo
Graduate Theses and Dissertations
This thesis explores the use of latent factor models to uncover hidden structures in pair wise outcomes derived from Over/Under betting markets in sports betting. Specifically, we implement and evaluate the Eigen model, a latent space model that represents dyadic data using node-specific vectors whose inner product govern edge probabilities. By modeling relationships between teams as adjacency matrices of binary outcomes, we investigate the extent to which the Eigen model captures both homophily, the tendency of similar teams to yield consistent betting results, and stochastic equivalence, where different teams exhibit indistinguishable patterns of Over/Under outcomes. A Bayesian formulation of the …
A Text Mining And Sentiment Analysis Of Valuable Cie Texts Using R, Eric Sugarman, Ethan Turber-Ortiz, Hannah Quinn
A Text Mining And Sentiment Analysis Of Valuable Cie Texts Using R, Eric Sugarman, Ethan Turber-Ortiz, Hannah Quinn
Mathematics, Computer Science & Statistics Presentations
The purpose of this project was to perform a sentiment analysis of three texts used in Ursinus College's Common Intellectual Experience (CIE) course: Between the World and Me by Ta-Nehisi Coates, The New Jim Crow by Michelle Alexander and Discourse on Method by Rene Descartes. Word count and word cloud analysis were also performed on the texts as well as term frequency and bigram analysis.
Repositioning The Game: Traditional Positions Vs Tracking-Based Archetypes In Nba Performance Models, Jacob Floyd
Repositioning The Game: Traditional Positions Vs Tracking-Based Archetypes In Nba Performance Models, Jacob Floyd
Senior Theses
Driven by the rise of advanced analytics and player tracking technologies, the NBA has transitioned away from traditional positional roles and toward more fluid player archetypes. This investigation uses principal component analysis and k-means clustering to group players based on season-long tracking data, creating new pseudo-positions that more accurately reflect modern playing styles. Predictive models were then built using both the classic position system and the newly generated clusters to forecast player scoring performance. Across every model comparison, both in terms of fit and predictive accuracy, the cluster-based system significantly outperformed the traditional position-based model. These results reinforce the idea …
Robust Spacecraft Autonomy For Deep Space Exploration In Special Euclidean Group Se(3), Matthew Wittal
Robust Spacecraft Autonomy For Deep Space Exploration In Special Euclidean Group Se(3), Matthew Wittal
Doctoral Dissertations and Master's Theses
Over the past half-century, humanity has gained extensive experience conducting manned spaceflight near Earth. Arguably, "near Earth" could even include the Moon — the most distant destination humans have reached. However, "near" in this work primarily refers low Earth orbit (LEO). One could argue that we have not truly left Earth since the Apollo, as spacecraft in some LEOs remain subject to atmospheric drag thus emphasizing their continued connection to Earth's immediate environment. Reflecting on this, it becomes clear that humanity has largely remained bound to Earth’s immediate vicinity since the Apollo missions reached the Moon. However, that is set …
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
“Regression To The Mean”: The Confluence Of Eugenics And Statistics In The 19th And 20th Centuries, Emrys G. King
Pomona Senior Theses
The work of this thesis is twofold — first, qualitatively characterizing the confluence between the British eugenics and statistics movements in the late 19th and early 20th centuries, and second, quantitatively analyzing the effect of this foundation on pedagogical materials in the growing field of statistics between 1880 and 1970. Towards the first goal, the history of the method of least squares, state statistics, and positive and negative eugenics are outlined, followed by a close reading of the foundational texts authored by Francis Galton and Karl Pearson that introduced linear regression. Towards the latter goal, English-language statistics textbooks published between …
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
Integrating Sentiment Analysis In Predictive Models: A Comparative Study On Game Popularity On Steam, Khaleefa Alhemeiri
CMC Senior Theses
Over the past decades, the gaming industry has managed to evolve into a multi-billion-dollar enterprise. Gaming platforms such as Steam foster unprecedented amounts of engagement among players worldwide daily. In this thesis, we investigate the effect of incorporating sentiment-driven metrics, specifically YouTube view counts and positive reviews, into predictive models for game popularity. In addition, by comparing our linear regression sentiment-based approach to the Bayesian hierarchical folded normal model used by De Luisa et al. (2021), we can understand the many differences, strengths, and limitations of each methodology. In our thesis, we focus on three games. Each is of varying …
The Impact Of “Multiple Looks” When Performing Survival Analysis, Quentin Eloise
The Impact Of “Multiple Looks” When Performing Survival Analysis, Quentin Eloise
Electronic Theses and Dissertations
Survival analysis is a critical statistical method in healthcare to assess patient treatment effects and disease progression. Another critical area of statistical methodology in health care is the practice of adaptive designs. Adaptive designs allow for interim analyses to take place during a study and various decisions and actions can take place more ethically. This is beneficial for studies that take multiple years to complete and allows administrators and healthcare providers to make sound decisions as early as possible. A challenging aspect of adaptive designs is that the number of interim analyses is known in advance which is applicable in …
Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu
Bayesian Variational Inference In Keyword Identification And Multiple Instance Classification, Yaofang Hu
Statistical Science Theses and Dissertations
This dissertation investigates (1) Variational Bayesian Semi-supervised Keyword Extraction and (2) Variational Bayesian Multimodal Multiple Instance Classification.
The expansion of textual data, stemming from various sources such as online product reviews and scholarly publications on scientific discoveries, has created a demand for the extraction of succinct yet comprehensive information. As a result, in recent years, efforts have been spent in developing novel methodologies for keyword extraction. Although many methods have been proposed to automatically extract keywords in the contexts of both unsupervised and fully supervised learning, how to effectively use partially observed keywords, such as author-specified keywords, remains an under-explored …
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
All Theses
High blood pressure, also known as hypertension, significantly increases the risk of heart disease and stroke, which are leading causes of death in the United States. While contributing to over 691,000 deaths in 2021 alone in the United States (U.S.), it also imposes immense economic burden on the healthcare system, costing approximately $131 billion annually. One way to address this issue is for increased self-care behaviors and medication adherence, both of which require sufficient health literacy. Despite the importance of health literacy, 90% of U.S. adults struggle with health-related subjects. Overcoming the issues associated with health literacy requires addressing the …
Recursive Marix Game Analysis: Optimal, Simplified, And Human Strategies In Brave Rats, William A. Medwid
Recursive Marix Game Analysis: Optimal, Simplified, And Human Strategies In Brave Rats, William A. Medwid
Master's Theses
Brave Rats is a short game with simple rules, yet establishing a comprehensive strategy is very challenging without extensive computation. After explaining the rules, this paper begins by calculating the optimal strategy by recursively solving each turn’s Minimax strategy. It then provides summary statistics about the complex, branching Minimax solution. Next, we examine six other strategy models and evaluate their performance against each other. These models’ flaws highlight the key elements that contribute to the effectiveness of the Minimax strategy and offer insight into simpler strategies that human players could mimic. Finally, we analyze 123 games of human data collected …
Unraveling The History Of Deforestation In The Amazon Rainforest With Statistical Modeling, Ryan Destefano
Unraveling The History Of Deforestation In The Amazon Rainforest With Statistical Modeling, Ryan Destefano
Master's Theses
The Amazon rainforest, a vital ecosystem of immense biodiversity and global climate significance, faces the ongoing threat of deforestation driven by agricultural expansion. This thesis employs remote sensing techniques, focusing on the Enhanced Vegetation Index (EVI) derived from Landsat satellite imagery, to track land cover dynamics within the Amazon. The study examines historical land cover changes in current plantations in Peru and Brazil, regions where the exact timing of deforestation is uncertain. By analyzing EVI measurements dating back to 1984, inflection points indicative of deforestation events preceding plantation establishment are identified. Statistical modeling techniques, including spline fitting to analyze time …
Using Probability Theory To Calculate The Odds That Either Candidate Wins The 2024 Presidential Election, Andrew Ruggero
Using Probability Theory To Calculate The Odds That Either Candidate Wins The 2024 Presidential Election, Andrew Ruggero
Department of Applied Mathematics & Statistics Faculty Publications
In the U.S., where the electoral college is used to determine the votes of an election, calculating the odds of a president winning is not as simple as looking at the total vote percentages. With each state not exactly having a proportionally linear amount of votes per its population, the total percentage doesn’t mean much. As such, in order to calculate these odds, we must look at a variety of winning combinations per each candidate. First, we must take into account that swing states are the only ones that matter. Defined as a 5% difference between the two main candidates, …
"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson
"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson
Senior Honors Theses
The authorship of Hebrews has been a point of contention for scholars for the past two millennia. While the epistle is traditionally attributed to Paul, many scholars assert that it carries thematic, structural, and stylistic differences from the remainder of his extant epistles; therefore, many other possible authors have been proposed. Of these, only Luke has other New Testament writings. Therefore, this project conducts a statistical comparison of Hebrews to the Pauline and Lukan corpora using stylometric authorial analysis methods. This analysis demonstrates that Hebrews is stylistically closer to Lukan literature than Pauline (but not to a significant degree), and …
Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky
Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky
Williams Honors College, Honors Research Projects
This project will examine the impact of using second-order terms in regression. For illustration, we use an example of regression where a baseball player's three by three heat map, including the height and distance from inside to outside of the pitch, are variables used to predict batting average. We find that second-order terms are crucial in discovering nonlinear relationships and interaction effects in regression models, and maintain that the common practice of using first-order additive models is insufficient.
To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis
To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis
Williams Honors College, Honors Research Projects
Regression to the mean is a statistical phenomenon that can hide important characteristics of what is truly happening in a research study. Caused by statistical randomness, regression to the mean occurs when extreme values, high or low, are followed by less extreme values. To correctly deal with it, one must understand what it is and how to distinguish its effect on conclusions made from the data. This paper provides examples of regression to the mean in both a medical and academic performance study and explains simple identifiers one can observe. Those are then followed up by the introduction of the …
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Ensemble Classification: An Analysis Of The Random Forest Model, Jarod Korn
Williams Honors College, Honors Research Projects
The random forest model proposed by Dr. Leo Breiman in 2001 is an ensemble machine learning method for classification prediction and regression. In the following paper, we will conduct an analysis on the random forest model with a focus on how the model works, how it is applied in software, and how it performs on a set of data. To fully understand the model, we will introduce the concept of decision trees, give a summary of the CART model, explain in detail how the random forest model operates, discuss how the model is implemented in software, demonstrate the model by …
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
SMU Data Science Review
Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …
Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang
Bayesian Statistical Modeling Of Spatially Resolved Transcriptomics Data, Xi Jiang
Statistical Science Theses and Dissertations
Spatially resolved transcriptomics (SRT) quantifies expression levels at different spatial locations, providing a new and powerful tool to investigate novel biological insights. As experimental technologies enhance both in capacity and efficiency, there arises a growing demand for the development of analytical methodologies.
One question in SRT data analysis is to identify genes whose expressions exhibit spatially correlated patterns, called spatially variable (SV) genes. Most current methods to identify SV genes are built upon the geostatistical model with Gaussian process, which could limit the models' ability to identify complex spatial patterns. In order to overcome this challenge and capture more types …
Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz
Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz
Beyond: Undergraduate Research Journal
When it comes to registering to vote, Hispanic voters can only register as “Hispanic” in the “Race/Ethnicity” category, causing difficulties when analyzing voting trends amongst the Hispanic community. Upon the recent idea that not all Hispanic Groups vote the same, the goal is to create a model that can possibly identify a voter’s Hispanic Group with the information provided on the public Florida voter file. This is accomplished using name and zip code data for all voters in Palm Beach, Florida. This paper will explore the model implemented, its findings and limitations. Palm Beach, Florida, is met with low confidence …
Sentiment Analysis Before And During The Covid-19 Pandemic, Emily Musgrove
Sentiment Analysis Before And During The Covid-19 Pandemic, Emily Musgrove
Mathematics Summer Fellows
This study examines the change in connotative language use before and during the Covid-19 pandemic. By analyzing news articles from several major US newspapers, we found that there is a statistically significant correlation between the sentiment of the text and the publication period. Specifically, we document a large, systematic, and statistically significant decline in the overall sentiment of articles published in major news outlets. While our results do not directly gauge the sentiment of the population, our findings have important implications regarding the social responsibility of journalists and media outlets especially in times of crisis.