Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Regression

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 61 - 90 of 91

Full-Text Articles in Statistics and Probability

A Likelihood Ratio Test Approach To Profile Monitoring In Tourism Industry, R. Noorossana, H. Izadbakhsh, M. R. Nayebpour Dec 2014

A Likelihood Ratio Test Approach To Profile Monitoring In Tourism Industry, R. Noorossana, H. Izadbakhsh, M. R. Nayebpour

Applications and Applied Mathematics: An International Journal (AAM)

A new statistical profile monitoring technique to monitor and detect changes in logistic profiles with an application in the tourism industry is presented in this paper. In the statistical process control literature, profile is usually referred to as a relationship between a response variable and one or more explanatory variables. In the tourism case study presented in this paper, time is considered as the explanatory variable and tourism satisfaction as the response variable. The Likelihood ratio test is used as a vehicle to detect any changes in the satisfaction profile in phase II of profile monitoring. The performance of the …


Convergence Of A Reinforcement Learning Algorithm In Continuous Domains, Stephen Carden Aug 2014

Convergence Of A Reinforcement Learning Algorithm In Continuous Domains, Stephen Carden

All Dissertations

In the field of Reinforcement Learning, Markov Decision Processes with a finite number of states and actions have been well studied, and there exist algorithms capable of producing a sequence of policies which converge to an optimal policy with probability one. Convergence guarantees for problems with continuous states also exist. Until recently, no online algorithm for continuous states and continuous actions has been proven to produce optimal policies. This Dissertation contains the results of research into reinforcement learning algorithms for problems in which both the state and action spaces are continuous. The problems to be solved are introduced formally as …


Implementation And Application Of The Curds And Whey Algorithm To Regression Problems, John Kidd May 2014

Implementation And Application Of The Curds And Whey Algorithm To Regression Problems, John Kidd

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

A common statistical problem is trying to predict two or more variables using a set of predictor variables. The simplest model for this situation is called multivariate linear regression. This method uses each set of predictor variables to predict each of the response variables separately. This approach seems counter-intuitive as any possible relationship between the variables being predicted is ignored.

Breiman and Friedman found a way to take advantage of relationships among the response variables to increase the accuracy of the predictions for each of the predicted variables with an algorithm they called Curds and
Whey. It uses other statistical …


Hidden Trends In Nfl Data, Scott Santor Apr 2014

Hidden Trends In Nfl Data, Scott Santor

Statistics

This is an analysis on National Football League (NFL) data for the 2013-2014 regular season. The main goal is to find hidden trends in game data that can ultimately determine which factors are statistically significant to award a team with their ultimate objective, a win.

The main response variable to be examined is total wins throughout the regular season, and an alternative dependent variable is spread; the difference between a team’s points scored, and points against. Spread is analyzed to provide a different quantitative response variable that can be both positive and negative.

Game data was gathered from ESPN.com box …


Robust Regression Methods For Massively Decayed Intelligence Data, Akiva Joachim Lorenz Jan 2014

Robust Regression Methods For Massively Decayed Intelligence Data, Akiva Joachim Lorenz

Wayne State University Dissertations

Homeland Security, sponsored by governmental initiatives, has become a vibrant academic research field. However, most efforts were placed with the recognition of threats (e.g. theory) and response options. Less effort was placed in the analysis of the collected data through statistical modeling. In a field that collects more than 20 terabyte of information per minute though diverse overt and covert means and indexes it for future research, understanding how different statistical models behave when it comes to massively decayed data is of vital importance.

Using Monte Carlo methods, three regression techniques (ordinary least squares, least-trimmed, and maximum likelihood) were tested …


On The Causal Interpretation Of Race In Regressions Adjusting For Confounding And Mediating Variables, Tyler J. Vanderweele, Whitney Robinson Nov 2013

On The Causal Interpretation Of Race In Regressions Adjusting For Confounding And Mediating Variables, Tyler J. Vanderweele, Whitney Robinson

Harvard University Biostatistics Working Paper Series

We consider different possible interpretations of the “effect of race” when regressions are run with race as an exposure variable, controlling also for various confounding and mediating variables. When adjustment is made for socioeconomic status early in a person's life, we discuss under what contexts the regression coefficients for race can be interpreted as corresponding to the extent to which a racial disparity would remain if various socioeconomic distributions early in life across racial groups could be equalized. When adjustment is also made for adult socioeconomic status, we note how the overall disparity can be decomposed into the portion that …


Curds And Whey: Little Miss Muffit's Contribution To Multivariate Linear Regression, John Cameron Kidd Jan 2013

Curds And Whey: Little Miss Muffit's Contribution To Multivariate Linear Regression, John Cameron Kidd

Undergraduate Honors Capstone Projects

A common multivariate statistical problem is the prediction of two or more response variables using two or more predictor variables. The simplest model for this situation is the multivariate linear regression model. The standard least squares estimation for this model involves regressing each response variable separately on all the predictor variables. Breiman and Friedman [1] show how to take advantage of correlations among the response variables to increase the predictive accuracy for each of the response variable with an algorithm they call Curds and Whey. In this report, I describe an implementation of the Curds and Whey algorithm in …


Revising Common Core Georgia Performance Standards Statistics Lesson Plans To Better Align With Statistical Practice, Rachel Bonilla Jan 2013

Revising Common Core Georgia Performance Standards Statistics Lesson Plans To Better Align With Statistical Practice, Rachel Bonilla

College of Graduate Studies: Theses & Dissertations

In this thesis, lesson plans provided by the Georgia Department of Education are revised to give students better exposure and practice working with real-life data. Three learning tasks and a performance task are presented covering a unit lesson on statistical regression. The development of Georgia statistics curriculum standards are reviewed and presented.


Statistical Topics Applied To Pressure And Temperature Readings In The Gulf Of Mexico, Malena Kathleen Allison Jan 2013

Statistical Topics Applied To Pressure And Temperature Readings In The Gulf Of Mexico, Malena Kathleen Allison

USF Tampa Graduate Theses and Dissertations

The field of statistical research in weather allows for the application of old and new methods, some of which may describe relationships between certain variables better such as temperatures and pressure. The objective of this study was to apply a variety of traditional and novel statistical methods to analyze data from the National Data Buoy Center, which records among other variables barometric pressure, atmospheric temperature, water temperature and dew point temperature. The analysis included attempts to better describe and model the data as well as to make estimations for certain variables. The following statistical methods were utilized: linear regression, non-response …


Tracking Atlantic Hurricanes Using Statistical Methods, Elizabeth Caitlin Miller Jan 2013

Tracking Atlantic Hurricanes Using Statistical Methods, Elizabeth Caitlin Miller

USF Tampa Graduate Theses and Dissertations

Creating an accurate hurricane location forecasting model is of the utmost importance because of the safety measures that need to occur in the days and hours leading up to a storm's landfall. Hurricanes can be incredibly deadly and costly, but if people are given adequate warning, many lives can be spared. This thesis seeks to develop an accurate model for predicting storm location based on previous location, previous wind speed, and previous pressure. The models are developed using hurricane data from 1980-2009.


An Economic Alternative To The C Chart, Ryan William Black Dec 2012

An Economic Alternative To The C Chart, Ryan William Black

Graduate Theses and Dissertations

Because the probability of Type I error is not evenly distributed beyond upper and lower three-sigma limits the c chart is theoretically inappropriate for a monitor of Poisson distributed phenomena. Furthermore, the normal approximation to the Poisson is of little use when c is small. These practical and theoretical concerns should motivate the computation of true error rates associated with individuals control assuming the Poisson distribution. An economic alternative to the c chart is described as a statistical model of upward shift from c0 to c1 and the two charts are compared in theory. For a range of c chart …


Improved Estimator In The Presence Of Multicollinearity, Ghadban Khalaf May 2012

Improved Estimator In The Presence Of Multicollinearity, Ghadban Khalaf

Journal of Modern Applied Statistical Methods

The performances of two biased estimators for the general linear regression model under conditions of collinearity are examined and a new proposed ridge parameter is introduced. Using Mean Square Error (MSE) and Monte Carlo simulation, the resulting estimator’s performance is evaluated and compared with the Ordinary Least Square (OLS) estimator and the Hoerl and Kennard (1970a) estimator. Results of the simulation study indicate that, with respect to MSE criteria, in all cases investigated the proposed estimator outperforms both the OLS and the Hoerl and Kennard estimators.


Analysis Of Roms Estimated Posterior Error Utilizing 4dvar Data Assimilation, Joseph Patrick Horton Jun 2011

Analysis Of Roms Estimated Posterior Error Utilizing 4dvar Data Assimilation, Joseph Patrick Horton

Mathematics

The appropriateness of the approximate error calculated by the Regional Ocean Modeling System (ROMS) is analyzed using Four-Dimensional Data Assimilation (4DVAR) performed on a numerical model of the San Luis Obispo Bay. An effective method of sampling data to minimize the actual error associated with the assimilated numerical model is explored by using different data sampling methods. An idealized state of the SLO bay region ("Real Run") is created to be used as the real ocean, then a numerical model of this region is created approximating this Real Run; this is known as the "Simulated State". By taking samples from …


Number Of Replications Required In Monte Carlo Simulation Studies: A Synthesis Of Four Studies, Daniel J. Mundform, Jay Schaffer, Myoung-Jin Kim, Dale Shaw, Ampai Thongteeraparp, Pornsin Supawan May 2011

Number Of Replications Required In Monte Carlo Simulation Studies: A Synthesis Of Four Studies, Daniel J. Mundform, Jay Schaffer, Myoung-Jin Kim, Dale Shaw, Ampai Thongteeraparp, Pornsin Supawan

Journal of Modern Applied Statistical Methods

Monte Carlo simulations are used extensively to study the performance of statistical tests and control charts. Researchers have used various numbers of replications, but rarely provide justification for their choice. Currently, no empirically-based recommendations regarding the required number of replications exist. Twenty-two studies were re-analyzed to determine empirically-based recommendations.


Investigating The Ironwood Tree (Casuarina Equisetifolia) Decline On Guam Using Applied Multinomial Modeling, Karl Anthony Schlub Jan 2010

Investigating The Ironwood Tree (Casuarina Equisetifolia) Decline On Guam Using Applied Multinomial Modeling, Karl Anthony Schlub

LSU Master's Theses

The ironwood tree (Casuarina equisetifolia), a protector of coastlines of the sub-tropical and tropical Western Pacific, is in decline on the island of Guam where aggressive data collection and efforts to mitigate the problem are underway. For each sampled tree the level of decline was measured on an ordinal scale consisting of five categories ranging from healthy to near dead. Several predictors were also measured including tree diameter, fire damage, typhoon damage, presence or absence of termites, presence or absence of basidiocarps, and various geographical or cultural factors. The five decline response levels can be viewed as categories of a …


The Em Algorithm For Group Testing Regression Models Under Matrix Pooling, Christopher R. Bilder, Boan Zhang Oct 2009

The Em Algorithm For Group Testing Regression Models Under Matrix Pooling, Christopher R. Bilder, Boan Zhang

Department of Statistics: Faculty Publications

No abstract provided.


Least Squares Percentage Regression, Chris Tofallis Nov 2008

Least Squares Percentage Regression, Chris Tofallis

Journal of Modern Applied Statistical Methods

In prediction, the percentage error is often felt to be more meaningful than the absolute error. We therefore extend the method of least squares to deal with percentage errors, for both simple and multiple regression. Exact expressions are derived for the coefficients, and we show how such models can be estimated using standard software. When the relative error is normally distributed, least squares percentage regression is shown to provide maximum likelihood estimates. The multiplicative error model is linked to least squares percentage regression in the same way that the standard additive error model is linked to ordinary least squares regression.


Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan Jan 2008

Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Regression models are often used to test for cause-effect relationships from data collected in randomized trials or experiments. This practice has deservedly come under heavy scrutiny, since commonly used models such as linear and logistic regression will often not capture the actual relationships between variables, and incorrectly specified models potentially lead to incorrect conclusions. In this paper, we focus on hypothesis test of whether the treatment given in a randomized trial has any effect on the mean of the primary outcome, within strata of baseline variables such as age, sex, and health status. Our primary concern is ensuring that such …


Regression By Data Segments Via Discriminant Analysis, Stan Lipovetsky, Michael Conklin May 2005

Regression By Data Segments Via Discriminant Analysis, Stan Lipovetsky, Michael Conklin

Journal of Modern Applied Statistical Methods

It is known that two-group linear discriminant function can be constructed via binary regression. In this article, it is shown that the opposite relation is also relevant – it is possible to present multiple regression as a linear combination of a main part, based on the pooled variance, and Fisher discriminators by data segments. Presenting regression as an aggregate of the discriminators allows one to decompose coefficients of the model into sum of several vectors related to segments. Using this technique provides an understanding of how the total regression model is composed of the regressions by the segments with possible …


On Time Series Analysis Of Public Health And Biomedical Data, Scott L. Zeger, Rafael A. Irizarry, Roger D. Peng Sep 2004

On Time Series Analysis Of Public Health And Biomedical Data, Scott L. Zeger, Rafael A. Irizarry, Roger D. Peng

Johns Hopkins University, Dept. of Biostatistics Working Papers

A time series is a sequence of observations made over time. Examples in public health include daily ozone concentrations, weekly admissions to an emergency department or annual expenditures on health care in the United States. Time series models are used to describe the dependence of the response at each time on predictor variables including covariates and possibly previous values in the series. Time series methods are necessary to account for the correlation among repeated responses over time. This paper gives an overview of time series ideas and methods used in public health research.


The Cross-Validated Adaptive Epsilon-Net Estimator, Mark J. Van Der Laan, Sandrine Dudoit, Aad W. Van Der Vaart Feb 2004

The Cross-Validated Adaptive Epsilon-Net Estimator, Mark J. Van Der Laan, Sandrine Dudoit, Aad W. Van Der Vaart

U.C. Berkeley Division of Biostatistics Working Paper Series

Suppose that we observe a sample of independent and identically distributed realizations of a random variable. Assume that the parameter of interest can be defined as the minimizer, over a suitably defined parameter space, of the expectation (with respect to the distribution of the random variable) of a particular (loss) function of a candidate parameter value and the random variable. Examples of commonly used loss functions are the squared error loss function in regression and the negative log-density loss function in density estimation. Minimizing the empirical risk (i.e., the empirical mean of the loss function) over the entire parameter space …


To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little Nov 2003

To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little

The University of Michigan Department of Biostatistics Working Paper Series

Finite population sampling is perhaps the only area of statistics where the primary mode of analysis is based on the randomization distribution, rather than on statistical models for the measured variables. This article reviews the debate between design and model-based inference. The basic features of the two approaches are illustrated using the case of inference about the mean from stratified random samples. Strengths and weakness of design-based and model-based inference for surveys are discussed. It is suggested that models that take into account the sample design and make weak parametric assumptions can produce reliable and efficient inferences in surveys settings. …


Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan Feb 2003

Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Risk estimation is an important statistical question for the purposes of selecting a good estimator (i.e., model selection) and assessing its performance (i.e., estimating generalization error). This article introduces a general framework for cross-validation and derives distributional properties of cross-validated risk estimators in the context of estimator selection and performance assessment. Arbitrary classes of estimators are considered, including density estimators and predictors for both continuous and polychotomous outcomes. Results are provided for general full data loss functions (e.g., absolute and squared error, indicator, negative log density). A broad definition of cross-validation is used in order to cover leave-one-out cross-validation, V-fold …


Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe Jan 2003

Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe

UW Biostatistics Working Paper Series

Accurate disease diagnosis is critical for health care. New diagnostic and screening tests must be evaluated for their abilities to discriminate disease from non-diseased states. The partial area under the ROC curve (partial AUC) is a measure of diagnostic test accuracy. We present an interpretation of the partial AUC that gives rise to a new non-parametric estimator. This estimator is more robust than existing estimators, which make parametric assumptions. We show that the robustness is gained with only a moderate loss in efficiency. We describe a regression modelling framework for making inference about covariate effects on the partial AUC. Such …


Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins Sep 2002

Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins

U.C. Berkeley Division of Biostatistics Working Paper Series

In biostatistics applications interest often focuses on the estimation of the distribution of a time-variable T. If one only observes whether or not T exceeds an observed monitoring time C, then the data structure is called current status data, also known as interval censored data, case I. We consider this data structure extended to allow the presence of both time-independent covariates and time-dependent covariate processes that are observed until the monitoring time. We assume that the monitoring process satisfies coarsening at random.

Our goal is to estimate the regression parameter beta of the regression model T = Z*beta+epsilon where the …


Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell Sep 2002

Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell

U.C. Berkeley Division of Biostatistics Working Paper Series

In many applications, it is often of interest to estimate a bivariate distribution of two survival random variables. Complete observation of such random variables is often incomplete. If one only observes whether or not each of the individual survival times exceeds a common observed monitoring time C, then the data structure is referred to as bivariate current status data (Wang and Ding, 2000). For such data, we show that the identifiable part of the joint distribution is represented by three univariate cumulative distribution functions, namely the two marginal cumulative distribution functions, and the bivariate cumulative distribution function evaluated on the …


Self-Consistency: A Fundamental Concept In Statistics, Thaddeus Tarpey, Bernard Flury Aug 1996

Self-Consistency: A Fundamental Concept In Statistics, Thaddeus Tarpey, Bernard Flury

Mathematics and Statistics Faculty Publications

The term ''self-consistency'' was introduced in 1989 by Hastie and Stuetzle to describe the property that each point on a smooth curve or surface is the mean of all points that project orthogonally onto it. We generalize this concept to self-consistent random vectors: a random vector Y is self-consistent for X if E[X|Y] = Y almost surely. This allows us to construct a unified theoretical basis for principal components, principal curves and surfaces, principal points, principal variables, principal modes of variation and other statistical methods. We provide some general results on self-consistent random variables, give …


Generating Unbiased Ratio And Regression Estimators, William (Bill) H. Williams Jun 1991

Generating Unbiased Ratio And Regression Estimators, William (Bill) H. Williams

Publications and Research

Standard ratio and regression are only conditionally unbiased. The paper uses split sample techniques to develop unbiased versions.


Linear Regression Of The Poisson Mean, Duane Steven Brown May 1982

Linear Regression Of The Poisson Mean, Duane Steven Brown

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

The purpose of this thesis was to compare two estimation procedures, the method of least squares and the method of maximum likelihood, on sample data obtained from a Poisson distribution. Point estimates of the slope and intercept of the regression line and point estimates of the mean squared error for both the slope and intercept were obtained. It is shown that least squares, the preferred method due to its simplicity, does yield results as good as maximum likelihood.

Also, confidence intervals were computed by Monte Carlo techniques and then were tested for accuracy. For the method of least squares, confidence …


Multicollinearity And The Estimation Of Regression Coefficients, John Charles Teed May 1978

Multicollinearity And The Estimation Of Regression Coefficients, John Charles Teed

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

The precision of the estimates of the regression coefficients in a regression analysis is affected by multicollinearity. The effect of certain factors on multicollinearity and the estimates was studied. The response variables were the standard error of the regression coefficients and a standarized statistic that measures the deviation of the regression coefficient from the population parameter.

The estimates are not influenced by any one factor in particular, but rather some combination of factors. The larger the sample size, the better the precision of the estimates no matter how "bad" the other factors may be.

The standard error of the regression …