Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistics and Probability (27)
- Applied Statistics (15)
- Statistical Models (12)
- Computer Sciences (10)
- Biostatistics (9)
-
- Social and Behavioral Sciences (9)
- Categorical Data Analysis (5)
- Mathematics (5)
- Life Sciences (4)
- Probability (4)
- Applied Mathematics (3)
- Business (3)
- Statistical Methodology (3)
- Statistical Theory (3)
- Artificial Intelligence and Robotics (2)
- Arts and Humanities (2)
- Bioinformatics (2)
- Economics (2)
- Engineering (2)
- Experimental Analysis of Behavior (2)
- Medicine and Health Sciences (2)
- Numerical Analysis and Scientific Computing (2)
- Operations Research, Systems Engineering and Industrial Engineering (2)
- Other Computer Sciences (2)
- Other Statistics and Probability (2)
- Physics (2)
- Psychology (2)
- Public Health (2)
- Institution
-
- Southern Methodist University (8)
- The University of Akron (2)
- Belmont University (1)
- Bowling Green State University (1)
- California Polytechnic State University, San Luis Obispo (1)
-
- Central Bank of Nigeria (1)
- Chapman University (1)
- City University of New York (CUNY) (1)
- Claremont Colleges (1)
- Clemson University (1)
- Embry-Riddle Aeronautical University (1)
- Loyola Marymount University and Loyola Law School (1)
- Minnesota State University, Mankato (1)
- University of Arkansas, Fayetteville (1)
- University of Central Florida (1)
- University of Georgia School of Law (1)
- University of Kentucky (1)
- University of Lynchburg (1)
- University of Montana (1)
- University of Nebraska - Lincoln (1)
- University of New Hampshire (1)
- University of New Mexico (1)
- University of South Carolina (1)
- University of South Florida (1)
- Ursinus College (1)
- Washington University in St. Louis (1)
- West Virginia University (1)
- Western Michigan University (1)
- Publication
-
- Statistical Science Theses and Dissertations (5)
- SMU Data Science Review (2)
- Williams Honors College, Honors Research Projects (2)
- All Graduate Theses, Dissertations, and Other Capstone Projects (1)
- All Theses (1)
-
- Beyond: Undergraduate Research Journal (1)
- Civil and Environmental Engineering Theses and Dissertations (1)
- Computational and Data Sciences (MS) Theses (1)
- Data Science Undergraduate Honors Theses (1)
- Dissertations (1)
- Economic and Financial Review (1)
- Generative AI Teaching Activities (1)
- Graduate Student Theses, Dissertations, & Professional Papers (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Honors Program: Senior Projects (Public) (1)
- Honors Projects (1)
- Honors Theses and Capstones (1)
- Honors Thesis (1)
- Honors Undergraduate Theses (1)
- Journal of Humanistic Mathematics (1)
- Master's Theses (1)
- Mathematics & Statistics ETDs (1)
- Mathematics, Computer Science & Statistics Presentations (1)
- Numeracy (1)
- Open Educational Resources (1)
- Presentations (1)
- SPARK Symposium Presentations (1)
- Senior Theses (1)
- Theses and Dissertations--Statistics (1)
- Undergraduate Theses and Capstone Projects (1)
- Publication Type
- File Type
Articles 1 - 30 of 36
Full-Text Articles in Data Science
Ai For Regression Analysis And More, Eli Snir
Ai For Regression Analysis And More, Eli Snir
Generative AI Teaching Activities
Students use Copilot and NotebookLM to create a dataset and develop statistical analyses including regression.
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
An Analytical Framework For Quantifying Urban And Community Resilience To Natural Hazards From Cell-Phone Gps-Location And Traffic-Flow Data, Georgios Chatzikyriakidis
Civil and Environmental Engineering Theses and Dissertations
Urban areas are increasingly exposed to natural hazards while accommodating a growing share of the global population, yet a consistent science-based framework for quantifying urban and community resilience remains lacking. This dissertation develops a physics-based analytical framework grounded in statistical mechanics and the quantitative theory of Brownian motion. A city is conceptualized as a complex medium in which citizens move analogously to Brownian particles within a viscoelastic environment, influenced by socioeconomic interactions and infrastructure functionality.
A central premise is that urban resilience, interpreted as engineering resilience (an outcome), can be quantified through a single metric: the mean-square displacement MSD=⟨r²(t)⟩, of …
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Bayesian Spatiotemporal Model For Counterfactual Estimation In Socioeconomic Studies, Duwani W. Gonzalez
Statistical Science Theses and Dissertations
Impact evaluations of regional development programs often require estimating counterfactual outcomes for a small number of treated regions using survey-based areal data. In practice, evaluators typically rely on two-group quasi-experimental methods such as propensity score matching (PSM) and Difference-in-Differences (DiD). These approaches perform poorly when only a few regions receive treatment, and when the set of observed covariates is limited or only partially relevant. Moreover, they typically do not explicitly exploit the spatial and temporal dependence present in survey-based areal data such as in ACS (American Community Survey). This dissertation develops a family of Bayesian spatial predictive models for directly …
A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout
A Spatial Analysis Of Streetlights In The City Of Sugar Land, Samuel J. Trout
Data Science Undergraduate Honors Theses
The purpose of this paper is to analyze patterns between public safety and streetlighting for the City of Sugar Land, TX so that they may better protect their citizens. The data involved come from the City of Sugar Land’s public works division and include type and location for all the attributes. The method of doing so involved visualizing the patterns of streetlights and their closest light readings to visualize which streetlights are underperforming using the Shiny package in R. Statistical tests were also used to quantify the association between lighting, crime occurrence, and crosswalks. From this, and the literature review, …
Predicting Criminal Behavior In Major Us Cities, Madison A. Price
Predicting Criminal Behavior In Major Us Cities, Madison A. Price
SPARK Symposium Presentations
In recent years, especially post pandemic, there has been a decrease in crime in the United States. Unfortunately, the country’s violent crime rates are still significantly higher compared to similar high-income countries, so what predicts crime in major American cities? There is tons of research to support the idea that demographics can offer some insight into predicting crime. There are countless online resources that seek to identify major crime centrals in the United States (Petrino, 2025). In the late 1990s, researchers noticed that crime rates in cities had a downward slope due to an important contributor: demographic change (Fox & …
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Beyond The Lace Index: Benchmarking Machine Learning Architectures And Explaining 30-Day Hospital Readmission Risk With Shap Analysis, Carl E. Hughes Iii
Williams Honors College, Honors Research Projects
Unplanned 30-day hospital readmission remains a fundamental challenge in US healthcare, associated with increased risk to patient recovery and representing an estimated $52.4 billion in annual expenses (Beauvais et al., 2022). While the rigorously validated LACE index serves as the clinical standard for readmission modeling, its linear structure and four explanatory variables lack the complexity to capture the high-dimensional and interactive nature of patient risk. This study utilizes an admission granularity level cohort of the MIMIC-IV database to develop and compare machine learning architectures against the baseline LACE index. Due to the imbalanced prevalence of readmission, the penalized logistic regression, …
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins
Honors Undergraduate Theses
With their deserts, castles, and ghost houses, the environments of Super Mario games are colorful, whimsical, and charming, but why are they so compelling, and what happens when our analysis of these environments extends beyond individual levels to expansive game worlds? Drawing on Cresswell’s theory of place (2014) and recent work on musical place-building in Mario Kart 8 (Heazlewood-Dale, 2024), I propose a spectrum between localized and globalized scale in games. As game environments become increasingly globalized, the music may be similarly altered to account for this shift in scale. Consequently, players may then encounter a broader, less musically congruent …
Smarter Disease Detection From Electronic Health Record Data: An End-To-End Ai-Augmented Pipeline For Computable Phenotyping, Dylan Owens
Statistical Science Theses and Dissertations
Electronic Health Records (EHR) contain a wealth of structured and unstructured patient data that can be leveraged for computable phenotyping, the process of algorithmically identifying patient cohorts with specific diseases or conditions. Traditional rule-based phenotyping approaches, while interpretable, often struggle with scalability, portability across institutions, and effective use of unstructured clinical narratives. Recent advances in large language models (LLMs) present new opportunities for synthesizing complex free-text information into concise, clinically meaningful representations. However, integrating LLMs into phenotyping workflows requires careful design to maintain transparency, interpretability, and measurable uncertainty—features essential for clinical adoption and downstream applications such as decision support.
We …
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Mat 301 - Applied Statistics And Data Analysis, Eric Aragundi
Open Educational Resources
Data analysis using standard statistical methods and relevant computer software. Emphasis on real-world data, interpretation, and misinterpretation of computer output.
This syllabus contains open source notebook about data analysis content.
Rewriting War: Improving The Mlb’S Go-To Advanced Metric, Kolin Atwood
Rewriting War: Improving The Mlb’S Go-To Advanced Metric, Kolin Atwood
Honors Projects
Traditional Wins Above Replacement (WAR) metrics have long served as a cornerstone of player evaluation in Major League Baseball, offering a context-neutral summary of offensive, defensive, and baserunning contributions. However, this neutrality often overlooks critical factors such as game situation, lineup strength, and advanced baserunning impact. This project proposes an enhanced model, WAR-PC (Wins Above Replacement – Plus Context), that integrates three key improvements: context-dependent batting value (RE24), clutch performance (Win Probability Added, WPA), and Statcast-based baserunning metrics. Using R, player logs, and modern baseball data sources, WAR-PC was calculated for eight players from the 2023 MLB season. The revised …
A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn
A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn
Mathematics, Computer Science & Statistics Presentations
The purpose of this project was to discover similarities between sentiments in Old Testament and New Testament books of the Bible, track emotional valence and find the most common words and sentiments in the books. Text analysis was performed on Genesis, Exodus, Matthew and Luke. Word clouds were also created for these texts.
Predicting Heart Disease Using Machine Learning Models, Zeynep Cetin
Predicting Heart Disease Using Machine Learning Models, Zeynep Cetin
Williams Honors College, Honors Research Projects
Heart disease remains the leading cause of death in the United States, particularly among the elderly population. The growing availability of large-scale health data and the advancement of machine learning tools present an opportunity to create more accurate and individualized predictive models. This study utilizes a subset of the 2020 Behavioral Risk Factor Surveillance System (BRFSS) dataset, focusing on individuals aged 70 and above, to explore predictive modeling using logistic regression, random forests, and XGBoost. The models were evaluated using key performance metrics, including sensitivity, specificity, accuracy, and the area under the ROC curve (AUC). The findings suggest that while …
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
Exploring Healthcare Chatbot Information Presentation: Applying Hierarchical Bayesian Regression And Inductive Thematic Analysis In A Mixed Methods Study, Samuel Nelson Koscelny
All Theses
High blood pressure, also known as hypertension, significantly increases the risk of heart disease and stroke, which are leading causes of death in the United States. While contributing to over 691,000 deaths in 2021 alone in the United States (U.S.), it also imposes immense economic burden on the healthcare system, costing approximately $131 billion annually. One way to address this issue is for increased self-care behaviors and medication adherence, both of which require sufficient health literacy. Despite the importance of health literacy, 90% of U.S. adults struggle with health-related subjects. Overcoming the issues associated with health literacy requires addressing the …
Book Review: How To Expect The Unexpected: The Science Of Making Predictions -- And The Art Of Knowing When Not To By Kit Yates, Mark Huber
Journal of Humanistic Mathematics
Humans think about the future all the time. Prediction is a part of how we prepare for the coming of both good and bad events in our lives. Kit Yates' book, How to expect the unexpected, concentrates primarily on the question of why prediction is difficult, and what mental shortcuts people take in prediction that can lead to incorrect results. Unfortunately, a lack of concern for details and several omissions undermine the quality of the book.
Interpretable Word-Level Sentiment Analysis With Attention-Based Multiple Instance Classification Models, Chenyu Yang
Interpretable Word-Level Sentiment Analysis With Attention-Based Multiple Instance Classification Models, Chenyu Yang
Statistical Science Theses and Dissertations
In this study, our main objective is to tackle the black-box nature of popular machine learning models in sentiment analysis and enhance model interpretability. We aim to gain more insight into the decision-making process of sentiment analysis models, which is often obscure in those complex models. To achieve this goal, we introduce two word-level sentiment analysis models.
The first model is called the attention-based multiple instance classification (AMIC) model. It combines the transparent model structure of multiple instance classification and the self-attention mechanism in deep learning to incorporate the contextual information from documents. As demonstrated by a wine review dataset …
Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang
Deep Learning For Microbiome-Based Integrative Modeling And Microbial Biomarkers Identification, Sen Yang
Statistical Science Theses and Dissertations
The human microbiome, comprising trillions of microorganisms, plays a pivotal role in modulating host physiology via molecular and metabolite exchanges. One of the major challenges in this field lies in the effective integration of microbiome and metabolomics data, an achievement that holds the promise of substantially enhancing the precision of disease prediction. However, many datasets prioritize microbiome data while neglecting paired metabolome information. Additionally, the prevalent analytical tools face challenges in effectively merging these intricate datasets, leading to possible misinterpretations and reduced prediction accuracies.
To address these challenges, the first part of this research introduces the Microbiome-based Supervised Contrastive Learning …
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
Differentiation Of Human, Dog, And Cat Hair Fibers Using Dart Tofms And Machine Learning, Laura Ahumada, Erin R. Mcclure-Price, Chad Kwong, Edgard O. Espinoza, John Santerre
SMU Data Science Review
Hair is found in over 90% of crime scenes and has long been analyzed as trace evidence. However, recent reviews of traditional hair fiber analysis techniques, primarily morphological examination, have cast doubt on its reliability. To address these concerns, this study employed machine learning algorithms, specifically Linear Discriminant Analysis (LDA) and Random Forest, on Direct Analysis in Real Time time-of-flight mass spectra collected from human, cat, and dog hair samples. The objective was to develop a chemistry- and statistics-based classification method for unbiased taxonomic identification of hair. The results of the study showed that LDA and Random Forest were highly …
A Prompt Engineering Approach To Creating Automated Commentary For Microsoft Self-Help Documentation Metric Reports Using Chatgpt, Ryan Herrin, Luke Stodgel, Brian Raffety
A Prompt Engineering Approach To Creating Automated Commentary For Microsoft Self-Help Documentation Metric Reports Using Chatgpt, Ryan Herrin, Luke Stodgel, Brian Raffety
SMU Data Science Review
Microsoft collects an immense amount of data from the users of their product-self-help documentation. Employees use this data to identify these self-help articles' performance trends and measure their impact on business Key Performance Indicators (KPIs). Microsoft uses various tools like Power BI and Python to analyze this data. The problem is that their analysis and findings are summarized manually. Therefore, this research will improve upon their current analysis methods by applying the latest prompt engineering practices and the power of ChatGPT's large language models (LLMs). Using VBA code, Microsoft Excel, and the ChatGPT API as an Excel add-in, this research …
Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz
Reu-Deim Classification Of Hispanic Voters In Hispanic Groups Using Name And Zip Code Data In Palm Beach, Florida, Kamila Soto-Ortiz
Beyond: Undergraduate Research Journal
When it comes to registering to vote, Hispanic voters can only register as “Hispanic” in the “Race/Ethnicity” category, causing difficulties when analyzing voting trends amongst the Hispanic community. Upon the recent idea that not all Hispanic Groups vote the same, the goal is to create a model that can possibly identify a voter’s Hispanic Group with the information provided on the public Florida voter file. This is accomplished using name and zip code data for all voters in Palm Beach, Florida. This paper will explore the model implemented, its findings and limitations. Palm Beach, Florida, is met with low confidence …
High-Dimensional Variable Selection Via Knockoffs Using Gradient Boosting, Amr Essam Mohamed
High-Dimensional Variable Selection Via Knockoffs Using Gradient Boosting, Amr Essam Mohamed
Dissertations
As data continue to grow rapidly in size and complexity, efficient and effective statistical methods are needed to detect the important variables/features. Variable selection is one of the most crucial problems in statistical applications. This problem arises when one wants to model the relationship between the response and the predictors. The goal is to reduce the number of variables to a minimal set of explanatory variables that are truly associated with the response of interest to improve the model accuracy. Effectively choosing the true influential variables and controlling the False Discovery Rate (FDR) without sacrificing power has been a challenge …
Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi
Applications Of Machine Learning Algorithms In Materials Science And Bioinformatics, Mohammed Quazi
Mathematics & Statistics ETDs
The piezoelectric response has been a measure of interest in density functional theory (DFT) for micro-electromechanical systems (MEMS) since the inception of MEMS technology. Piezoelectric-based MEMS devices find wide applications in automobiles, mobile phones, healthcare devices, and silicon chips for computers, to name a few. Piezoelectric properties of doped aluminum nitride (AlN) have been under investigation in materials science for piezoelectric thin films because of its wide range of device applicability. In this research using rigorous DFT calculations, high throughput ab-initio simulations for 23 AlN alloys are generated.
This research is the first to report strong enhancements of piezoelectric properties …
Generating A Dataset For Comparing Linear Vs. Non-Linear Prediction Methods In Education Research, Jack Mauro, Elena Martinez, Anna Bargagliotti
Generating A Dataset For Comparing Linear Vs. Non-Linear Prediction Methods In Education Research, Jack Mauro, Elena Martinez, Anna Bargagliotti
Honors Thesis
Machine learning is often used to build predictive models by extracting patterns from large data sets. Such techniques are increasingly being utilized to predict outcomes in the social sciences. One such application is predicting student success. Machine learning can be applied to predicting student acceptance and success in academia. Using these tools for education-related data analysis, may enable the evaluation of programs, resources and curriculum. Currently, research is needed to examine application, admissions, and retention data in order to address equity in college computer science programs. However, most student-level data sets contain sensitive data that cannot be made public. To …
Causalmodels: An R Library For Estimating Causal Effects, Joshua Wolff Anderson
Causalmodels: An R Library For Estimating Causal Effects, Joshua Wolff Anderson
Computational and Data Sciences (MS) Theses
Free and open source software for statistical modeling and machine learning have advanced productivity in data science significantly. Packages such as SciPy in Python and caret in R provide fundamental tools for statistical modeling and machine learning in the two most popular programming languages used by data scientists. Unfortunately, robust tools similar to these are limited in terms of causal inference. The tools in R that exist lack consistent and standardized methodologies and inputs. R lacks a comprehensive package that offers traditional causal inference methods such as standardization, IP weighting, G-estimation, outcome regression, and propensity matching in one common package. …
Split Classification Model For Complex Clustered Data, Katherine Gerot
Split Classification Model For Complex Clustered Data, Katherine Gerot
Honors Program: Senior Projects (Public)
Classification in high-dimensional data has generated tremendous interest in a multitude of fields. Data in higher dimensions often tend to reside in non-Euclidean metric space. This prevents Euclidean-based classification methodologies, such as regression, from reliably modeling the data. Many proposed models rely on computationally-complex embedding to convert the data to a more usable format. Others, namely the Support Vector Machine, rely on kernel manipulation to implicitly describe the "feature space" to arrive at a non-linear decision boundary. The proposed methodology in this paper seeks to classify complex data in a relatively computationally-simple and explainable manner.
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Finding The Best Predictors For Foot Traffic In Us Seafood Restaurants, Isabel Paige Beaulieu
Honors Theses and Capstones
COVID-19 caused state and nation-wide lockdowns, which altered human foot traffic, especially in restaurants. The seafood sector in particular suffered greatly as there was an increase in illegal fishing, it is made up of perishable goods, it is seasonal in some places, and imports and exports were slowed. Foot traffic data is useful for business owners to have to know how much to order, how many employees to schedule, etc. One issue is that the data is very expensive, hard to get, and not available until months after it is recorded. Our goal is to not only find covariates that …
Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su
Theses and Dissertations--Statistics
When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …
A Monte Carlo Simulation Of Rat Choice Behavior With Interdependent Outcomes, Michelle A. Frankot
A Monte Carlo Simulation Of Rat Choice Behavior With Interdependent Outcomes, Michelle A. Frankot
Graduate Theses, Dissertations, and Problem Reports (ETD)
Preclinical behavioral neuroscience often uses choice paradigms to capture psychiatric symptoms. In particular, the subfield of operant research produces nested datasets with many discrete choices in a session. The standard analytic practice is to aggregate choice into a continuous variable and analyze using ANOVA or linear regression. However, choice data often have multiple interdependent outcomes of interest, violating an assumption of general linear models. The aim of the current study was to quantify the accuracy of linear mixed-effects regression (LMER) for analyzing data from a 4-choice operant task called the Rodent Gambling Task (RGT), which measures decision-making in the context …
The Classification Of Basket Neural Cells In The Mammalian Neocortex, Sreya Pudi
The Classification Of Basket Neural Cells In The Mammalian Neocortex, Sreya Pudi
Senior Theses
Basket neuronal cells of the mammalian neocortex have been classically categorized into two or more groups. Originally, it was thought that the large and small types are the naturally occurring groups that emerge from reasons that relate to neurobiological function and anatomical position. Later, a study based on anatomical and physiological features of these neurons introduced a third type, the net basket cell which is intermediate in size as compared to the large and small types. In this study, multivariate analysis was used to test the hypothesis that the large and small types are morphologically distinct groups. The results of …
An Introduction To Calling Bullshit: Learning To Think Outside The Black Box, Jevin D. West, Carl T. Bergstrom
An Introduction To Calling Bullshit: Learning To Think Outside The Black Box, Jevin D. West, Carl T. Bergstrom
Numeracy
Bergstrom, Carl T. and Jevin D. West. 2020. Calling Bullshit: The Art of Skepticism in a Data-Driven World. (New York: Random House) 336 pp. ISBN 978-0525509202.
While statistical methods receive greater attention, the art of critically evaluating information in everyday life more commonly depends on thinking outside the black box of the algorithm. In this piece we introduce readers to our book and associated online teaching materials—for readers who want to more capably call “bullshit” or to teach their students to do the same.
Statistical Analysis Of 2017-18 Premier League Match Statistics Using A Regression Analysis In R, Bergen Campbell
Statistical Analysis Of 2017-18 Premier League Match Statistics Using A Regression Analysis In R, Bergen Campbell
Undergraduate Theses and Capstone Projects
This thesis analyzes the correlation between a team’s statistics and the success of their performances, and develops a predictive model that can be used to forecast final season results for that team. Data from the 2017-2018 Premier League season is to be gathered and broken down within R to highlight what factors and variables are largely contributing to the success or downfall of a team. A multiple linear regression model and stepwise selection process is then used to include any factors that are significant in predicting in match results.
The predictions about the 17-18 season results based on the model …