Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Regression

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 40

Full-Text Articles in Applied Statistics

Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo Jan 2026

Modeling Housing Prices: Which Features Matter Most?, Alex Ruvolo

Williams Honors College, Honors Research Projects

This paper attempts to find the biggest factors and traits that influence the cost of housing. This will include the lot size, type of street, utilities, neighborhood, year built, heating, electrical, yard size, number of different rooms, age, condition, and others. I will attempt to answer the question of whether the prices of houses have changed within the last 5 to 10 years, and obviously this is an easy question to answer. However, the bigger question beyond this is are the main factors affecting housing prices all important in explaining this relationship? Is one factor more important than the rest …


From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins Jan 2026

From Lap To Map: How Musical Scale, Place, And Play Drive The Interconnected Mario Kart World, Cameron Cummins

Honors Undergraduate Theses

With their deserts, castles, and ghost houses, the environments of Super Mario games are colorful, whimsical, and charming, but why are they so compelling, and what happens when our analysis of these environments extends beyond individual levels to expansive game worlds? Drawing on Cresswell’s theory of place (2014) and recent work on musical place-building in Mario Kart 8 (Heazlewood-Dale, 2024), I propose a spectrum between localized and globalized scale in games. As game environments become increasingly globalized, the music may be similarly altered to account for this shift in scale. Consequently, players may then encounter a broader, less musically congruent …


Comparative Evaluation Of Estimation Techniques For Purchasing Power Parity In African Countries Using The Country-Product-Dummy Regression Framework, Rokibat Adeola Tijani, Taiwo Abideen Lasisi, Dahud Kehinde Shangodoyin, Olasunkanmi James Oladapo Sep 2025

Comparative Evaluation Of Estimation Techniques For Purchasing Power Parity In African Countries Using The Country-Product-Dummy Regression Framework, Rokibat Adeola Tijani, Taiwo Abideen Lasisi, Dahud Kehinde Shangodoyin, Olasunkanmi James Oladapo

Al-Bahir

Purchasing Power Parity (PPP) is a popular macroeconomic analysis metric used to compare economic productivity and standards of living between countries. This study examines the estimation of PPP within the International Comparison Program (ICP) at Basic Heading (BH) level stage and leverages on the data from the 2011 ICP round. Focusing on five BHs out of 12 BHs across 50 Africa countries, to empirically evaluate the validity of the classical Ordinary Least Square (OLS) assumptions in the estimation of Country Product Dummy (CPD) regressions. Given the widespread use of OLS for BH level PPP computation, a rigorous examination of these …


Profiting On The Kentucky Derby, Bailey Korfhage Apr 2025

Profiting On The Kentucky Derby, Bailey Korfhage

Undergraduate Theses

This paper analyzes the quantitative data of horses that ran in the Kentucky Derby to recognize statistically significant variables to predict the horse that comes in first or in-the-money. This analysis is specific to the post-implementation of the points system that began for the 2013 Kentucky Derby. Churchill Downs, the host of the Kentucky Derby, changed the methodology of qualification for a horse to enter the race; instead of qualifying with highest earnings in lifetime starts, the institution implemented a points system that awarded different proportions of points depending on the value of various prep races leading up to the …


Examining The Interaction Between Calcium Supplement Use, Demographics, And Lifestyle Factors On Bone Health In Women, Vix Talbot Jun 2024

Examining The Interaction Between Calcium Supplement Use, Demographics, And Lifestyle Factors On Bone Health In Women, Vix Talbot

University Honors Theses

Osteoporosis is a condition which poses a significant health threat, particularly among women during the menopause transition, where accelerated bone loss increases fracture risk. Calcium supplementation has been shown to be an important intervention to mitigate bone mineral density (BMD) decline during this and other periods of life. However, the efficacy of calcium supplementation is influenced by various individual factors, including demographics and lifestyle habits. This study investigates the interaction between calcium supplement use, and several interaction terms on bone health in women. Multiple linear regression analysis is employed to assess the impact of these factors on BMD. Data from …


Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky Jan 2024

Pitching The Use Of Squared And Interaction Terms In Regression Via Baseball Heat Maps, Lucas Chepelsky

Williams Honors College, Honors Research Projects

This project will examine the impact of using second-order terms in regression. For illustration, we use an example of regression where a baseball player's three by three heat map, including the height and distance from inside to outside of the pitch, are variables used to predict batting average. We find that second-order terms are crucial in discovering nonlinear relationships and interaction effects in regression models, and maintain that the common practice of using first-order additive models is insufficient.


To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis Jan 2024

To Mean Or Not To Mean: An Investigation Of Regression To The Mean, Hunter Ellis

Williams Honors College, Honors Research Projects

Regression to the mean is a statistical phenomenon that can hide important characteristics of what is truly happening in a research study. Caused by statistical randomness, regression to the mean occurs when extreme values, high or low, are followed by less extreme values. To correctly deal with it, one must understand what it is and how to distinguish its effect on conclusions made from the data. This paper provides examples of regression to the mean in both a medical and academic performance study and explains simple identifiers one can observe. Those are then followed up by the introduction of the …


The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña Nov 2023

The Use Of Regularization To Detect Racial Inequities In Pay Equity Studies: An Empirical Study And Reflections On Regulation Methods, Christopher M. Peña

Electronic Theses and Dissertations

Since the late 1970s, multiple linear regression has been the preferred method for identifying discrimination in pay. An empirical study on this topic was conducted using quantitative critical methods. A literature review first examined conflicting views on using multiple linear regression in pay equity studies. The review found that multiple linear regression is used so prevalently in pay equity studies because the courts and practitioners have widely accepted it and because of its simplicity and ability to parse multiple sources of variance simultaneously. Commentaries in the literature cautioned about errors in model specification, the use of tainted variables, and the …


Predicting Insulin Pump Therapy Settings, Riccardo L. Ferraro, David Grijalva, Alex Trahan Sep 2022

Predicting Insulin Pump Therapy Settings, Riccardo L. Ferraro, David Grijalva, Alex Trahan

SMU Data Science Review

Millions of people live with diabetes worldwide [7]. To mitigate some of the many symptoms associated with diabetes, an estimated 350,000 people in the United States rely on insulin pumps [17]. For many of these people, how effectively their insulin pump performs is the difference between sleeping through the night and a life threatening emergency treatment at a hospital. Three programmed insulin pump therapy settings governing effective insulin pump function are: Basal Rate (BR), Insulin Sensitivity Factor (ISF), and Carbohydrate Ratio (ICR). For many people using insulin pumps, these therapy settings are often not correct, given their physiological needs. While …


Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su Jan 2022

Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su

Theses and Dissertations--Statistics

When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …


Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling Jan 2022

Analysis Of Minor League Rule Changes Effect On Stolen Bases, Zachary Houghtaling

Williams Honors College, Honors Research Projects

This study uses various statistical analyses to evaluate the justification of rule changes for Major League Baseball that were implemented within the Minor Leagues during the 2021 minor league season. The primary focus of the study is predicting how some of these Minor League rule changes could affect the stolen base success rate and the number of attempts per game within the Major Leagues. A survey was conducted to evaluate how fans feel about stolen bases within the current game and if rules should be altered to increase the number of stolen bases that occur. Additionally, recorded Major and Minor …


Empirical Modeling Of Tilt-Rotor Aerodynamic Performance, Michael C. Stratton Oct 2021

Empirical Modeling Of Tilt-Rotor Aerodynamic Performance, Michael C. Stratton

Mechanical & Aerospace Engineering Theses & Dissertations

There has been increasing interest into the performance of electric vertical takeoff and landing (eVTOL) aircraft. The propellers used for the eVTOL propulsion systems experience a broad range of aerodynamic conditions, not typically experienced by propellers in forward flight, that includes large incidence angles relative to the oncoming airflow. Formal experiment design and analysis techniques featuring response surface methods were applied to a subscale, tilt-rotor wind tunnel test for three, four, five, and six blade, 16-inch diameter, propeller configurations in support of development of the NASA LA-8 aircraft. Investigation of low-speed performance included a maximum speed of 12 m/s and …


A Monte Carlo Analysis Of Ordinary Least Squares Versus Equal Weights, James Brewer Ayres Oct 2020

A Monte Carlo Analysis Of Ordinary Least Squares Versus Equal Weights, James Brewer Ayres

Masters Theses & Specialist Projects

Equal weights are an alternative weighting procedure to the optimal weights offered by ordinary least squares regression analysis. Also called units weights, equal weights are formed by standardizing scores on the predictor variables and averaging these standardized scores to create a composite score. Research is limited regarding the conditions under which equal weights result in cross-validated 𝑅𝑅2 values that meet or exceed optimal weights. In this study, I explored the effect of various predictor-criterion correlations, predictor intercorrelations, and sample sizes to determine the relative performance of equal and optimal weighting schemes upon cross-validation. Results indicated that optimally weighted predictors explained …


Linear Methods For Regression With Small Sample Sizes Relative To The Number Of Variables., Rajesh Sikder Aug 2020

Linear Methods For Regression With Small Sample Sizes Relative To The Number Of Variables., Rajesh Sikder

Electronic Theses and Dissertations

In data sets where there are a small number of observations but a large number of variables observed for each observation, ordinary least squares estimation cannot be used for regression models. There are many alternative including stepwise regression, penalized methods such as ridge regression and the LASSO, and methods based on derived inputs such as principal components regression and partial least squares regression. In this thesis, these five methods are described. K-fold cross validation is also discussed as a way for determining regularization parameters for each method. The performance of these methods in estimation and prediction is also examined through …


Using Stability To Select A Shrinkage Method, Dean Dustin May 2020

Using Stability To Select A Shrinkage Method, Dean Dustin

Department of Statistics: Dissertations, Theses, and Student Research

Shrinkage methods are estimation techniques based on optimizing expressions to find which variables to include in an analysis, typically a linear regression. The general form of these expressions is the sum of an empirical risk plus a complexity penalty based on the number of parameters. Many shrinkage methods are known to satisfy an ‘oracle’ property meaning that asymptotically they select the correct variables and estimate their coefficients efficiently. In Section 1.2, we show oracle properties in two general settings. The first uses a log likelihood in place of the empirical risk and allows a general class of penalties. The second …


Introduction To Research Statistical Analysis: An Overview Of The Basics, Christian Vandever Apr 2020

Introduction To Research Statistical Analysis: An Overview Of The Basics, Christian Vandever

HCA Healthcare Journal of Medicine

This article covers many statistical ideas essential to research statistical analysis. Sample size is explained through the concepts of statistical significance level and power. Variable types and definitions are included to clarify necessities for how the analysis will be interpreted. Categorical and quantitative variable types are defined, as well as response and predictor variables. Statistical tests described include t-tests, ANOVA and chi-square tests. Multiple regression is also explored for both logistic and linear regression. Finally, the most common statistics produced by these methods are explored.


An Exploration Of Link Functions Used In Ordinal Regression, Thomas J. Smith, David A. Walker, Cornelius M. Mckenna Apr 2020

An Exploration Of Link Functions Used In Ordinal Regression, Thomas J. Smith, David A. Walker, Cornelius M. Mckenna

Journal of Modern Applied Statistical Methods

The purpose of this study is to examine issues involved with choice of a link function in generalized linear models with ordinal outcomes, including distributional appropriateness, link specificity, and palindromic invariance are discussed and an exemplar analysis provided using the Pew Research Center 25th anniversary of the Web Omnibus Survey data. Simulated data are used to compare the relative palindromic invariance of four distinct indices of determination/discrimination, including a newly proposed index by Smith et al. (2017).


Evaluation Of Relationship Between Lead-Dust Loading, Lead-Dust Concentration, And Total Dust Loading Metrics Across Multiple Data Sets, Charles Bevington Dec 2019

Evaluation Of Relationship Between Lead-Dust Loading, Lead-Dust Concentration, And Total Dust Loading Metrics Across Multiple Data Sets, Charles Bevington

Capstone Experience: Master of Public Health

Lead-dust monitoring studies report values as either lead-dust loadings µg/ft2 or as lead-dust concentrations µg/g. It is rare for studies to report both metrics. When only lead-dust loading values are present, professionals require an approach to estimate lead-dust concentration values. A literature search identified five studies that contained raw data for both lead-dust loading and lead-dust concentration. An additional thirty-two studies had summary-statistics available for both lead-dust loading and lead-dust concentration. Studies with raw-data were used to develop an empirically-based loading to concentration statistical relationship. Raw data sets were critically evaluated to determine whether elimination or …


Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia May 2019

Visualization And Machine Learning Techniques For Nasa’S Em-1 Big Data Problem, Antonio P. Garza Iii, Jose Quinonez, Misael Santana, Nibhrat Lohia

SMU Data Science Review

In this paper, we help NASA solve three Exploration Mission-1 (EM-1) challenges: data storage, computation time, and visualization of complex data. NASA is studying one year of trajectory data to determine available launch opportunities (about 90TBs of data). We improve data storage by introducing a cloud-based solution that provides elasticity and server upgrades. This migration will save $120k in infrastructure costs every four years, and potentially avoid schedule slips. Additionally, it increases computational efficiency by 125%. We further enhance computation via machine learning techniques that use the classic orbital elements to predict valid trajectories. Our machine learning model decreases trajectory …


Logistic Regression: An Inferential Method For Identifying The Best Predictors, Rand Wilcox Mar 2019

Logistic Regression: An Inferential Method For Identifying The Best Predictors, Rand Wilcox

Journal of Modern Applied Statistical Methods

When dealing with a logistic regression model, there is a simple method for estimating the strength of the association between the jth covariate and the dependent variable when all covariates are entered into the model. There is the issue of determining whether the jth independent variable has a stronger or weaker association than the kth independent variable. This note describes a method for dealing with this issue that was found to perform reasonably well in simulations.


Yelp’S Review Filtering Algorithm, Yao Yao, Ivelin Angelov, Jack Rasmus-Vorrath, Mooyoung Lee, Daniel W. Engels Aug 2018

Yelp’S Review Filtering Algorithm, Yao Yao, Ivelin Angelov, Jack Rasmus-Vorrath, Mooyoung Lee, Daniel W. Engels

SMU Data Science Review

In this paper, we present an analysis of features influencing Yelp's proprietary review filtering algorithm. Classifying or misclassifying reviews as recommended or non-recommended affects average ratings, consumer decisions, and ultimately, business revenue. Our analysis involves systematically sampling and scraping Yelp restaurant reviews. Features are extracted from review metadata and engineered from metrics and scores generated using text classifiers and sentiment analysis. The coefficients of a multivariate logistic regression model were interpreted as quantifications of the relative importance of features in classifying reviews as recommended or non-recommended. The model classified review recommendations with an accuracy of 78%. We found that reviews …


Examination And Comparison Of The Performance Of Common Non-Parametric And Robust Regression Models, Gregory F. Malek Aug 2017

Examination And Comparison Of The Performance Of Common Non-Parametric And Robust Regression Models, Gregory F. Malek

Electronic Theses and Dissertations

ABSTRACT

Examination and Comparison of the Performance of Common Non-Parametric and Robust Regression Models

By

Gregory Frank Malek

Stephen F. Austin State University, Masters in Statistics Program,

Nacogdoches, Texas, U.S.A.

[email protected]

This work investigated common alternatives to the least-squares regression method in the presence of non-normally distributed errors. An initial literature review identified a variety of alternative methods, including Theil Regression, Wilcoxon Regression, Iteratively Re-Weighted Least Squares, Bounded-Influence Regression, and Bootstrapping methods. These methods were evaluated using a simple simulated example data set, as well as various real data sets, including math proficiency data, Belgian telephone call data, and faculty …


Effective Estimation Strategy Of Finite Population Variance Using Multi-Auxiliary Variables In Double Sampling, Reba Maji, G. N. Singh, Arnab Bandyopadhyay May 2017

Effective Estimation Strategy Of Finite Population Variance Using Multi-Auxiliary Variables In Double Sampling, Reba Maji, G. N. Singh, Arnab Bandyopadhyay

Journal of Modern Applied Statistical Methods

Estimation of population variance in two-phase (double) sampling is considered using information on multiple auxiliary variables. An unbiased estimator is proposed and its properties are studied under two different structures. The superiority of the suggested estimator over some contemporary estimators of population variance was established through empirical studies from a natural and an artificially generated dataset.


Efficient And Unbiased Estimation Procedure Of Population Mean In Two-Phase Sampling, Reba Maji, Arnab Bandyopadhyay, G. N. Singh Nov 2016

Efficient And Unbiased Estimation Procedure Of Population Mean In Two-Phase Sampling, Reba Maji, Arnab Bandyopadhyay, G. N. Singh

Journal of Modern Applied Statistical Methods

In this paper, an unbiased regression-ratio type estimator has been developed for estimating the population mean using two auxiliary variables in double sampling. Its properties are studied under two different cases. Empirical studies and graphical simulation have been done to demonstrate the efficiency of the proposed estimator over other estimators.


Design Optimization Of A Stochastic Multi-Objective Problem: Gaussian Process Regressions For Objective Surrogates, Juan Sebastian Martinez, Piyush Pandita, Rohit K. Tripathy, Ilias Bilionis Aug 2016

Design Optimization Of A Stochastic Multi-Objective Problem: Gaussian Process Regressions For Objective Surrogates, Juan Sebastian Martinez, Piyush Pandita, Rohit K. Tripathy, Ilias Bilionis

The Summer Undergraduate Research Fellowship (SURF) Symposium

Multi-objective optimization (MOO) problems arise frequently in science and engineering situations. In an optimization problem, we want to find the set of input parameters that generate the set of optimal outputs, mathematically known as the Pareto frontier (PF). Solving the MOO problem is a challenge since expensive experiments can be performed only a constrained number of times and there is a limited set of data to work with, e.g. a roll-to-roll microwave plasma chemical vapor deposition (MPCVD) reactor for manufacturing high quality graphene. State-of-the-art techniques, e.g. evolutionary algorithms; particle swarm optimization, require a large amount of observations and do not …


A Spatial Analytical Framework For Examining Road Traffic Crashes, Grace O. Korter May 2016

A Spatial Analytical Framework For Examining Road Traffic Crashes, Grace O. Korter

Journal of Modern Applied Statistical Methods

A number of different modeling techniques have been used to examine road traffic crashes for analytic and predictive purposes. Map-based spatial analysis is introduced. Applications are given which show the power in a combination of existing exploratory and statistical methods.


Contrails: Causal Inference Using Propensity Scores, Dean S. Barron Nov 2015

Contrails: Causal Inference Using Propensity Scores, Dean S. Barron

Journal of Modern Applied Statistical Methods

Contrails are clouds caused by airplane exhausts, which geologists contend decrease daily temperature ranges on Earth. Following the 2001 World Trade Center attack, cancelled domestic flights triggered the first absence of contrails in decades. Resultant exceptional data capacitated causal inference analysis by propensity score matching. Estimated contrail effect was 6.8981°F.


Predicting Successful Long-Term Weight Loss From Short-Term Weight-Loss Outcomes: New Insights From A Dynamic Energy Balance Model (The Pounds Lost Study), Diana Thomas, W Andrada Ivanescu, Corby K. Martin, Steven B. Heymsfield, Kaitlyn Marshall, Victoria E. Bodrato, Donald Williamson, Stephen Anton, Frank M. Sacks, Donna Ryan, George A. Bray Mar 2015

Predicting Successful Long-Term Weight Loss From Short-Term Weight-Loss Outcomes: New Insights From A Dynamic Energy Balance Model (The Pounds Lost Study), Diana Thomas, W Andrada Ivanescu, Corby K. Martin, Steven B. Heymsfield, Kaitlyn Marshall, Victoria E. Bodrato, Donald Williamson, Stephen Anton, Frank M. Sacks, Donna Ryan, George A. Bray

Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works

Background: Currently, early weight-loss predictions of long-term weight-loss success rely on fixed percent-weight-loss thresholds.

Objective: The objective was to develop thresholds during the first 3 mo of intervention that include the influence of age, sex, baseline weight, percent weight loss, and deviations from expected weight to predict whether a participant is likely to lose 5% or more body weight by year 1.

Design: Data consisting of month 1, 2, 3, and 12 treatment weights were obtained from the 2-y Preventing Obesity Using Novel Dietary Strategies (POUNDS Lost) intervention. Logistic regression models that included covariates of age, height, sex, baseline weight, …


Implementation And Application Of The Curds And Whey Algorithm To Regression Problems, John Kidd May 2014

Implementation And Application Of The Curds And Whey Algorithm To Regression Problems, John Kidd

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

A common statistical problem is trying to predict two or more variables using a set of predictor variables. The simplest model for this situation is called multivariate linear regression. This method uses each set of predictor variables to predict each of the response variables separately. This approach seems counter-intuitive as any possible relationship between the variables being predicted is ignored.

Breiman and Friedman found a way to take advantage of relationships among the response variables to increase the accuracy of the predictions for each of the predicted variables with an algorithm they called Curds and
Whey. It uses other statistical …


Revising Common Core Georgia Performance Standards Statistics Lesson Plans To Better Align With Statistical Practice, Rachel Bonilla Jan 2013

Revising Common Core Georgia Performance Standards Statistics Lesson Plans To Better Align With Statistical Practice, Rachel Bonilla

College of Graduate Studies: Theses & Dissertations

In this thesis, lesson plans provided by the Georgia Department of Education are revised to give students better exposure and practice working with real-life data. Three learning tasks and a performance task are presented covering a unit lesson on statistical regression. The development of Georgia statistics curriculum standards are reviewed and presented.