Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Regression

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 91

Full-Text Articles in Statistics and Probability

Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo Jan 2026

Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo

Williams Honors College, Honors Research Projects

This paper attempts to find the biggest factors and traits that influence the cost of housing. This will include the lot size, type of street, utilities, neighborhood, year built, heating, electrical, yard size, number of different rooms, age, condition, and others. I will attempt to answer the question of whether the prices of houses have changed within the last 5 to 10 years, and obviously this is an easy question to answer. However, the bigger question beyond this is are the main factors affecting housing prices all important in explaining this relationship? Is one factor more important than the rest …


From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins Jan 2026

From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins

Honors Undergraduate Theses

With their deserts, castles, and ghost houses, the environments of Super Mario games are colorful, whimsical, and charming, but why are they so compelling, and what happens when our analysis of these environments extends beyond individual levels to expansive game worlds? Drawing on Cresswell’s theory of place (2014) and recent work on musical place-building in Mario Kart 8 (Heazlewood-Dale, 2024), I propose a spectrum between localized and globalized scale in games. As game environments become increasingly globalized, the music may be similarly altered to account for this shift in scale. Consequently, players may then encounter a broader, less musically congruent …


Seemingly Unrelated Exponetiated Exponential Geometric Regression Model, Oluwaseun Michael Famoni, Bamidele Mustapha Oseni Dec 2025

Seemingly Unrelated Exponetiated Exponential Geometric Regression Model, Oluwaseun Michael Famoni, Bamidele Mustapha Oseni

Al-Bahir

In the class of seemingly unrelated regression models, the dispersion nature of the dependent variable can greatly impact the efficiency and reliability of the parameter estimates for the model. Despite this, the seemingly unrelated Poisson regression model and seemingly unrelated negative binomial model are two most commonly used count data models for these class of regression models. This study introduces the seemingly unrelated exponentiated exponential geometric regression (SUEEGR) for modelling count data which might be equi-, under, or over-dispersed. Parameters estimation for the model was carried out using the method of maximum likelihood. A simulation study was carried out to …


Comparative Evaluation Of Estimation Techniques For Purchasing Power Parity In African Countries Using The Country-Product-Dummy Regression Framework, Rokibat Adeola Tijani, Taiwo Abideen Lasisi, Dahud Kehinde Shangodoyin, Olasunkanmi James Oladapo Sep 2025

Comparative Evaluation Of Estimation Techniques For Purchasing Power Parity In African Countries Using The Country-Product-Dummy Regression Framework, Rokibat Adeola Tijani, Taiwo Abideen Lasisi, Dahud Kehinde Shangodoyin, Olasunkanmi James Oladapo

Al-Bahir

Purchasing Power Parity (PPP) is a popular macroeconomic analysis metric used to compare economic productivity and standards of living between countries. This study examines the estimation of PPP within the International Comparison Program (ICP) at Basic Heading (BH) level stage and leverages on the data from the 2011 ICP round. Focusing on five BHs out of 12 BHs across 50 Africa countries, to empirically evaluate the validity of the classical Ordinary Least Square (OLS) assumptions in the estimation of Country Product Dummy (CPD) regressions. Given the widespread use of OLS for BH level PPP computation, a rigorous examination of these …


Regression With Atypical Data: Measurement Error, Periodicity, And Non-Normality, Nicholas W. Woolsey Jul 2025

Regression With Atypical Data: Measurement Error, Periodicity, And Non-Normality, Nicholas W. Woolsey

Theses and Dissertations

Regression is a ubiquitous and fundamental method that can be found in any ele- mentary statistics course. The simplicity and self evidently useful nature of linear regression beguiles a non-negligible portion of researchers into disrespecting assump- tions required by these models, namely in terms of accuracy of covariates and the underlying nature of the data. This disregard can at best lead to meaningless results and at worse cause significant misunderstandings in scientific pursuit.

In this dissertation we strive propose remedies to violations of specific assump- tions. Namely the assumptions that covariates are either observed without measure- ment error or they …


Profiting On The Kentucky Derby, Bailey Korfhage Apr 2025

Profiting On The Kentucky Derby, Bailey Korfhage

Undergraduate Theses

This paper analyzes the quantitative data of horses that ran in the Kentucky Derby to recognize statistically significant variables to predict the horse that comes in first or in-the-money. This analysis is specific to the post-implementation of the points system that began for the 2013 Kentucky Derby. Churchill Downs, the host of the Kentucky Derby, changed the methodology of qualification for a horse to enter the race; instead of qualifying with highest earnings in lifetime starts, the institution implemented a points system that awarded different proportions of points depending on the value of various prep races leading up to the …


A Bayesian Complex-Valued Latent Variable Model Applied To Functional Magnetic Resonance Imaging, Chase J. Sakitis, D. Andrew Brown, Daniel B. Rowe Jan 2025

A Bayesian Complex-Valued Latent Variable Model Applied To Functional Magnetic Resonance Imaging, Chase J. Sakitis, D. Andrew Brown, Daniel B. Rowe

Mathematical and Statistical Science Faculty Research and Publications

In linear regression, the coefficients are simple to estimate using the least squares method with a known design matrix for the observed measurements. However, real-world applications may encounter complications such as an unknown design matrix and complex-valued parameters. The design matrix can be estimated from prior information but can potentially cause an inverse problem when multiplying by the transpose as it is generally ill-conditioned. This can be combat by adding regularizers to the model but does not always mitigate the issues. Here, we propose our Bayesian approach to a complex-valued latent variable linear model with an application to functional magnetic …


Examining The Interaction Between Calcium Supplement Use, Demographics, And Lifestyle Factors On Bone Health In Women, Vix Talbot Jun 2024

Examining The Interaction Between Calcium Supplement Use, Demographics, And Lifestyle Factors On Bone Health In Women, Vix Talbot

University Honors Theses

Osteoporosis is a condition which poses a significant health threat, particularly among women during the menopause transition, where accelerated bone loss increases fracture risk. Calcium supplementation has been shown to be an important intervention to mitigate bone mineral density (BMD) decline during this and other periods of life. However, the efficacy of calcium supplementation is influenced by various individual factors, including demographics and lifestyle habits. This study investigates the interaction between calcium supplement use, and several interaction terms on bone health in women. Multiple linear regression analysis is employed to assess the impact of these factors on BMD. Data from …


Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles Jan 2024

Defensive Impact Wins: Developing A New Method To Rate Individual Defense In Nba Games, Dylan J. Stiles

Honors Theses and Capstones

With the analytics revolution in sports in the past 20 years, it seems that everything that can be quantified is. In basketball though, trying to break the game down into a set of numbers comes with a unique problem. While we've come up with a good set of advanced numbers to measure offensive efficiency, defense is fundamentally harder to quantify. The game is played five on five, but it has often been popular or convenient to model defense as a set of five one on one games. As defenses became more complex into the 2010s, this methodology became more insignificant. …


Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky Jan 2024

Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky

Williams Honors College, Honors Research Projects

This project will examine the impact of using second-order terms in regression. For illustration, we use an example of regression where a baseball player's three by three heat map, including the height and distance from inside to outside of the pitch, are variables used to predict batting average. We find that second-order terms are crucial in discovering nonlinear relationships and interaction effects in regression models, and maintain that the common practice of using first-order additive models is insufficient.


To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis Jan 2024

To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis

Williams Honors College, Honors Research Projects

Regression to the mean is a statistical phenomenon that can hide important characteristics of what is truly happening in a research study. Caused by statistical randomness, regression to the mean occurs when extreme values, high or low, are followed by less extreme values. To correctly deal with it, one must understand what it is and how to distinguish its effect on conclusions made from the data. This paper provides examples of regression to the mean in both a medical and academic performance study and explains simple identifiers one can observe. Those are then followed up by the introduction of the …


The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña Nov 2023

The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña

Electronic Theses and Dissertations

Since the late 1970s, multiple linear regression has been the preferred method for identifying discrimination in pay. An empirical study on this topic was conducted using quantitative critical methods. A literature review first examined conflicting views on using multiple linear regression in pay equity studies. The review found that multiple linear regression is used so prevalently in pay equity studies because the courts and practitioners have widely accepted it and because of its simplicity and ability to parse multiple sources of variance simultaneously. Commentaries in the literature cautioned about errors in model specification, the use of tainted variables, and the …


Examining Model Complexity's Effects When Predicting Continuous Measures From Ordinal Labels, Mckade S. Thomas May 2023

Examining Model Complexity's Effects When Predicting Continuous Measures From Ordinal Labels, Mckade S. Thomas

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Many real world problems require the prediction of ordinal variables where the values are a set of categories with an ordering to them. However, in many of these cases the categorical nature of the ordinal data is not a desirable outcome. As such, regression models treat ordinal variables as continuous and do not bind their predictions to discrete categories. Prior research has found that these models are capable of learning useful information between the discrete levels of the ordinal labels they are trained on, but complex models may learn ordinal labels too closely, missing the information between levels. In this …


The 2015 Ncaa Cost-Of-Attendance Stipend And Its Effects On Institutional Financial Aid Packages, Sara Greene Apr 2023

The 2015 Ncaa Cost-Of-Attendance Stipend And Its Effects On Institutional Financial Aid Packages, Sara Greene

Honors Theses

In 2015, the National Collegiate Athletic Association (NCAA) allowed “Cost of Attendance” (COA) stipends to be offered to athletic recruits for Division I schools. These stipends are intended to allow schools to grant aid to student-athletes beyond a full-ride scholarship to cover additional costs imposed on student-athletes. These stipends created an opportunity for the “Autonomy” Power 5 programs to utilize a competitive tactic to try to win over the top recruits. There is evidence that these COA stipends have caused an increase in the estimated cost of attendance reported by the university. This paper examines if the COA stipends have …


Analyzing Relationships With Machine Learning, Oscar Ko Feb 2023

Analyzing Relationships With Machine Learning, Oscar Ko

Dissertations, Theses, and Capstone Projects

Procedurally, this project aims to take a dataset, analyze it, and offer insights to the audience in an easy-to-digest format. Conceptually, this project will seek to explore questions like: “Do couples that meet through online dating or dating apps have higher or lower quality relationships?”, “Can any features in this dataset help predict how a subject would rate their relationship quality?”, and “What other insights can I derive from using machine learning for exploratory analysis?” The intended audience for this project is anyone interested in romantic relationships or machine learning.

The dataset is from a Stanford University survey, “How Couples …


Predicting Insulin Pump Therapy Settings, Riccardo L. Ferraro, David Grijalva, Alex Trahan Sep 2022

Predicting Insulin Pump Therapy Settings, Riccardo L. Ferraro, David Grijalva, Alex Trahan

SMU Data Science Review

Millions of people live with diabetes worldwide [7]. To mitigate some of the many symptoms associated with diabetes, an estimated 350,000 people in the United States rely on insulin pumps [17]. For many of these people, how effectively their insulin pump performs is the difference between sleeping through the night and a life threatening emergency treatment at a hospital. Three programmed insulin pump therapy settings governing effective insulin pump function are: Basal Rate (BR), Insulin Sensitivity Factor (ISF), and Carbohydrate Ratio (ICR). For many people using insulin pumps, these therapy settings are often not correct, given their physiological needs. While …


Contributions To Random Forest Variable Importance With Applications In R, Kelvyn K. Bladen Aug 2022

Contributions To Random Forest Variable Importance With Applications In R, Kelvyn K. Bladen

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

A major focus in statistics is building and improving computational algorithms that can use data to predict a response. Two fundamental camps of research arise from such a goal. The first camp is researching ways to get more accurate predictions. Many sophisticated methods, collectively known as machine learning methods, have been developed for this very purpose. One such method that is widely used across industry and many other areas of investigation is called Random Forests.

The second camp of research is that of improving the interpretability of machine learning methods. This is worthy of attention when analysts desire to optimize …


The Short-Term Effects Of Fine Airborne Particulate Matter And Climate On Covid-19 Disease Dynamics, El Hussain Shamsa, Kezhong Zhang Jun 2022

The Short-Term Effects Of Fine Airborne Particulate Matter And Climate On Covid-19 Disease Dynamics, El Hussain Shamsa, Kezhong Zhang

Medical Student Research Symposium

Background: Despite more than 60% of the United States population being fully vaccinated, COVID-19 cases continue to spike in a temporal pattern. These patterns in COVID-19 incidence and mortality may be linked to short-term changes in environmental factors.

Methods: Nationwide, county-wise measurements for COVID-19 cases and deaths, fine-airborne particulate matter (PM2.5), and maximum temperature were obtained from March 20, 2020 to March 20, 2021. Multivariate Linear Regression was used to analyze the association between environmental factors and COVID-19 incidence and mortality rates in each season. Negative Binomial Regression was used to analyze daily fluctuations of COVID-19 cases …


Assessing The Influence Of Health Policy And Population Mobility On Covid-19 Spread In Arkansas, Tayden Barretto May 2022

Assessing The Influence Of Health Policy And Population Mobility On Covid-19 Spread In Arkansas, Tayden Barretto

Industrial Engineering Undergraduate Honors Theses

The outbreak of COVID-19 has created a major crisis across the world since its start in 2019, and its influence on every realm of society is undeniable. Globally, more than 500 million cases have been recorded since March 2020, with almost 6 million deaths. In the wake of this crisis, many governments and health organizations have taken steps and precautions to mitigate its spread. These steps involve public mandates of information, reducing frequency of personal contact, and use of masks to minimize the risk of transmission. Current access to mobility data released from Google detailing population movements has provided a …


Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su Jan 2022

Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su

Theses and Dissertations--Statistics

When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …


Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling Jan 2022

Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling

Williams Honors College, Honors Research Projects

This study uses various statistical analyses to evaluate the justification of rule changes for Major League Baseball that were implemented within the Minor Leagues during the 2021 minor league season. The primary focus of the study is predicting how some of these Minor League rule changes could affect the stolen base success rate and the number of attempts per game within the Major Leagues. A survey was conducted to evaluate how fans feel about stolen bases within the current game and if rules should be altered to increase the number of stolen bases that occur. Additionally, recorded Major and Minor …


(R1239) A New Type Ii Half Logistic-G Family Of Distributions With Properties, Regression Models, System Reliability And Applications, Emrah Altun, Morad Alizadeh, Haitham M. Yousof, Mahdi Rasekhi, G. G. Hamedani Dec 2021

(R1239) A New Type Ii Half Logistic-G Family Of Distributions With Properties, Regression Models, System Reliability And Applications, Emrah Altun, Morad Alizadeh, Haitham M. Yousof, Mahdi Rasekhi, G. G. Hamedani

Applications and Applied Mathematics: An International Journal (AAM)

This study proposes a new family of distributions based on the half logistic distribution. With the new family, the baseline distributions gain flexibility through additional shape parameters. The important statistical properties of the proposed family are derived. A new generalization of the Weibull distribution is used to introduce a location-scale regression model for the censored response variable. The utility of the introduced models is demonstrated in survival analysis and estimation of the system reliability. Three data sets are analyzed. According to the empirical results, it is observed that the proposed family gives better results than other existing models.


Comparison Of Statistical Methods For Modeling Count Data With An Application To Length Of Hospital Stay, Gustavo A. Fernandez Dec 2021

Comparison Of Statistical Methods For Modeling Count Data With An Application To Length Of Hospital Stay, Gustavo A. Fernandez

Theses and Dissertations

Hospital length of stay (LOS) is a key indicator of hospital care management efficiency, cost of care, and hospital planning. Therefore, understanding hospital LOS variability is always an important healthcare focus. Hospital LOS data are count data, with discrete and nonnegative values, typically right-skewed, and often exhibiting excessive zeros. Numerous studies have been conducted to model hospital LOS to identify significant predictors contributing to its variability. Many researchers have used linear regression with or without logarithmic transformation of the outcome variable LOS, or logistic regression on a dichotomized LOS. These regression methods usually violate models’ assumptions and are subject …


Aggregating Twitter Text Through Generalized Linear Regression Models For Tweet Popularity Prediction And Automatic Topic Classification, Chen Mo, Jingjing Yin, Isaac Chun-Hai Fung, Zion Tse Nov 2021

Aggregating Twitter Text Through Generalized Linear Regression Models For Tweet Popularity Prediction And Automatic Topic Classification, Chen Mo, Jingjing Yin, Isaac Chun-Hai Fung, Zion Tse

Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications

Social media platforms have become accessible resources for health data analysis. However, the advanced computational techniques involved in big data text mining and analysis are challenging for public health data analysts to apply. This study proposes and explores the feasibility of a novel yet straightforward method by regressing the outcome of interest on the aggregated influence scores for association and/or classification analyses based on generalized linear models. The method reduces the document term matrix by transforming text data into a continuous summary score, thereby reducing the data dimension substantially and easing the data sparsity issue of the term matrix. To …


Empirical Modeling Of Tilt-Rotor Aerodynamic Performance, Michael C. Stratton Oct 2021

Empirical Modeling Of Tilt-Rotor Aerodynamic Performance, Michael C. Stratton

Mechanical & Aerospace Engineering Theses & Dissertations

There has been increasing interest into the performance of electric vertical takeoff and landing (eVTOL) aircraft. The propellers used for the eVTOL propulsion systems experience a broad range of aerodynamic conditions, not typically experienced by propellers in forward flight, that includes large incidence angles relative to the oncoming airflow. Formal experiment design and analysis techniques featuring response surface methods were applied to a subscale, tilt-rotor wind tunnel test for three, four, five, and six blade, 16-inch diameter, propeller configurations in support of development of the NASA LA-8 aircraft. Investigation of low-speed performance included a maximum speed of 12 m/s and …


Modeling Dynamic Correlation In Zero-Inflated Bivariate Count Data With Applications To Single-Cell Rna Sequencing Data, Zhen Yang, Yen-Yi Ho Mar 2021

Modeling Dynamic Correlation In Zero-Inflated Bivariate Count Data With Applications To Single-Cell Rna Sequencing Data, Zhen Yang, Yen-Yi Ho

Faculty Publications

Interactions between biological molecules in a cell are tightly coordinated and often highly dynamic. As a result of these varying signaling activities, changes in gene coexpression patterns could often be observed. The advancements in next-generation sequencing technologies bring new statistical challenges for studying these dynamic changes of gene coexpression. In recent years, methods have been developed to examine genomic information from individual cells. Single-cell RNA sequencing (scRNA-seq) data are count-based, and often exhibit characteristics such as overdispersion and zero inflation. To explore the dynamic dependence structure in scRNA-seq data and other zero-inflated count data, new approaches are needed. In this …


An Evaluation Of Knot Placement Strategies For Spline Regression, William Klein Jan 2021

An Evaluation Of Knot Placement Strategies For Spline Regression, William Klein

CMC Senior Theses

Regression splines have an established value for producing quality fit at a relatively low-degree polynomial. This paper explores the implications of adopting new methods for knot selection in tandem with established methodology from the current literature. Structural features of generated datasets, as well as residuals collected from sequential iterative models are used to augment the equidistant knot selection process. From analyzing a simulated dataset and an application onto the Racial Animus dataset, I find that a B-spline basis paired with equally-spaced knots remains the best choice when data are evenly distributed, even when structural features of a dataset are known …


A Statistical Learning Regression Model Utilized To Determine Predictive Factors Of Social Distancing During Covid-19 Pandemic, Timothy A. Smith, Albert J. Boquet, Matthew V. Chin Nov 2020

A Statistical Learning Regression Model Utilized To Determine Predictive Factors Of Social Distancing During Covid-19 Pandemic, Timothy A. Smith, Albert J. Boquet, Matthew V. Chin

Publications

In an application of the mathematical theory of statistics, predictive regression modelling can be used to determine if there is a trend to predict the response variable of social distancing in terms of multiple predictor input “predictor” variables. In this study the social distancing is measured as the percentage reduction in average mobility by GPS records, and the mathematical results obtained are interpreted to determine what factors drive that response. This study was done on county level data from the state of Florida during the COVID-19 pandemic, and it is found that the most deterministic predictors are county population density …


A Monte Carlo Analysis Of Ordinary Least Squares Versus Equal Weights, James Brewer Ayres Oct 2020

A Monte Carlo Analysis Of Ordinary Least Squares Versus Equal Weights, James Brewer Ayres

Masters Theses & Specialist Projects

Equal weights are an alternative weighting procedure to the optimal weights offered by ordinary least squares regression analysis. Also called units weights, equal weights are formed by standardizing scores on the predictor variables and averaging these standardized scores to create a composite score. Research is limited regarding the conditions under which equal weights result in cross-validated 𝑅𝑅2 values that meet or exceed optimal weights. In this study, I explored the effect of various predictor-criterion correlations, predictor intercorrelations, and sample sizes to determine the relative performance of equal and optimal weighting schemes upon cross-validation. Results indicated that optimally weighted predictors explained …


Linear Methods For Regression With Small Sample Sizes Relative To The Number Of Variables., Rajesh Sikder Aug 2020

Linear Methods For Regression With Small Sample Sizes Relative To The Number Of Variables., Rajesh Sikder

Electronic Theses and Dissertations

In data sets where there are a small number of observations but a large number of variables observed for each observation, ordinary least squares estimation cannot be used for regression models. There are many alternative including stepwise regression, penalized methods such as ridge regression and the LASSO, and methods based on derived inputs such as principal components regression and partial least squares regression. In this thesis, these five methods are described. K-fold cross validation is also discussed as a way for determining regularization parameters for each method. The performance of these methods in estimation and prediction is also examined through …