Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (40)
- Statistical Models (28)
- Social and Behavioral Sciences (19)
- Statistical Methodology (18)
- Statistical Theory (17)
-
- Engineering (10)
- Applied Mathematics (9)
- Medicine and Health Sciences (9)
- Biostatistics (8)
- Data Science (8)
- Multivariate Analysis (8)
- Mathematics (7)
- Categorical Data Analysis (6)
- Other Statistics and Probability (6)
- Computer Sciences (5)
- Public Health (5)
- Economics (4)
- Numerical Analysis and Computation (4)
- Operations Research, Systems Engineering and Industrial Engineering (4)
- Probability (4)
- Aerospace Engineering (3)
- Arts and Humanities (3)
- Econometrics (3)
- Industrial Engineering (3)
- Longitudinal Data Analysis and Time Series (3)
- Oceanography and Atmospheric Sciences and Meteorology (3)
- Survival Analysis (3)
- Artificial Intelligence and Robotics (2)
- Institution
-
- Wayne State University (12)
- COBRA (9)
- Utah State University (7)
- The University of Akron (5)
- University of South Carolina (5)
-
- City University of New York (CUNY) (4)
- Southern Methodist University (4)
- University of Arkansas, Fayetteville (3)
- California Polytechnic State University, San Luis Obispo (2)
- Embry-Riddle Aeronautical University (2)
- Georgia Southern University (2)
- Louisiana State University (2)
- Old Dominion University (2)
- Prairie View A&M University (2)
- University of Alkafeel (2)
- University of Central Florida (2)
- University of Nebraska - Lincoln (2)
- University of South Florida (2)
- Bellarmine University (1)
- Claremont Colleges (1)
- Clemson University (1)
- East Tennessee State University (1)
- HCA Healthcare (1)
- Marquette University (1)
- Montclair State University (1)
- Northern Michigan University (1)
- Portland State University (1)
- Purdue University (1)
- Stephen F. Austin State University (1)
- University of Denver (1)
- Publication Year
- Publication
-
- Journal of Modern Applied Statistical Methods (10)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (5)
- Theses and Dissertations (5)
- U.C. Berkeley Division of Biostatistics Working Paper Series (5)
- Williams Honors College, Honors Research Projects (5)
-
- Electronic Theses and Dissertations (4)
- SMU Data Science Review (4)
- Al-Bahir (2)
- Applications and Applied Mathematics: An International Journal (AAM) (2)
- Dissertations, Theses, and Capstone Projects (2)
- Graduate Theses and Dissertations (2)
- LSU Master's Theses (2)
- USF Tampa Graduate Theses and Dissertations (2)
- All Dissertations (1)
- All Graduate Plan B and other Reports, Spring 1920 to Spring 2023 (1)
- All NMU Master's Theses (1)
- Arts & Sciences Graduate Student Theses and Dissertations (1)
- Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications (1)
- Business and Economics Presentations (1)
- CMC Senior Theses (1)
- Capstone Experience: Master of Public Health (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Department of Applied Mathematics and Statistics Faculty Scholarship and Creative Works (1)
- Department of Statistics: Dissertations, Theses, and Student Research (1)
- Department of Statistics: Faculty Publications (1)
- Dissertations and Theses (1)
- Engineering Management & Systems Engineering Theses & Dissertations (1)
- Faculty Publications (1)
- HCA Healthcare Journal of Medicine (1)
- HIM 1990-2015 (1)
- Publication Type
- File Type
Articles 61 - 90 of 91
Full-Text Articles in Statistics and Probability
A Likelihood Ratio Test Approach To Profile Monitoring In Tourism Industry, R. Noorossana, H. Izadbakhsh, M. R. Nayebpour
A Likelihood Ratio Test Approach To Profile Monitoring In Tourism Industry, R. Noorossana, H. Izadbakhsh, M. R. Nayebpour
Applications and Applied Mathematics: An International Journal (AAM)
A new statistical profile monitoring technique to monitor and detect changes in logistic profiles with an application in the tourism industry is presented in this paper. In the statistical process control literature, profile is usually referred to as a relationship between a response variable and one or more explanatory variables. In the tourism case study presented in this paper, time is considered as the explanatory variable and tourism satisfaction as the response variable. The Likelihood ratio test is used as a vehicle to detect any changes in the satisfaction profile in phase II of profile monitoring. The performance of the …
Convergence Of A Reinforcement Learning Algorithm In Continuous Domains, Stephen Carden
Convergence Of A Reinforcement Learning Algorithm In Continuous Domains, Stephen Carden
All Dissertations
In the field of Reinforcement Learning, Markov Decision Processes with a finite number of states and actions have been well studied, and there exist algorithms capable of producing a sequence of policies which converge to an optimal policy with probability one. Convergence guarantees for problems with continuous states also exist. Until recently, no online algorithm for continuous states and continuous actions has been proven to produce optimal policies. This Dissertation contains the results of research into reinforcement learning algorithms for problems in which both the state and action spaces are continuous. The problems to be solved are introduced formally as …
Implementation And Application Of The Curds And Whey Algorithm To Regression Problems, John Kidd
Implementation And Application Of The Curds And Whey Algorithm To Regression Problems, John Kidd
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
A common statistical problem is trying to predict two or more variables using a set of predictor variables. The simplest model for this situation is called multivariate linear regression. This method uses each set of predictor variables to predict each of the response variables separately. This approach seems counter-intuitive as any possible relationship between the variables being predicted is ignored.
Breiman and Friedman found a way to take advantage of relationships among the response variables to increase the accuracy of the predictions for each of the predicted variables with an algorithm they called Curds and
Whey. It uses other statistical …
Hidden Trends In Nfl Data, Scott Santor
Hidden Trends In Nfl Data, Scott Santor
Statistics
This is an analysis on National Football League (NFL) data for the 2013-2014 regular season. The main goal is to find hidden trends in game data that can ultimately determine which factors are statistically significant to award a team with their ultimate objective, a win.
The main response variable to be examined is total wins throughout the regular season, and an alternative dependent variable is spread; the difference between a team’s points scored, and points against. Spread is analyzed to provide a different quantitative response variable that can be both positive and negative.
Game data was gathered from ESPN.com box …
Robust Regression Methods For Massively Decayed Intelligence Data, Akiva Joachim Lorenz
Robust Regression Methods For Massively Decayed Intelligence Data, Akiva Joachim Lorenz
Wayne State University Dissertations
Homeland Security, sponsored by governmental initiatives, has become a vibrant academic research field. However, most efforts were placed with the recognition of threats (e.g. theory) and response options. Less effort was placed in the analysis of the collected data through statistical modeling. In a field that collects more than 20 terabyte of information per minute though diverse overt and covert means and indexes it for future research, understanding how different statistical models behave when it comes to massively decayed data is of vital importance.
Using Monte Carlo methods, three regression techniques (ordinary least squares, least-trimmed, and maximum likelihood) were tested …
On The Causal Interpretation Of Race In Regressions Adjusting For Confounding And Mediating Variables, Tyler J. Vanderweele, Whitney Robinson
On The Causal Interpretation Of Race In Regressions Adjusting For Confounding And Mediating Variables, Tyler J. Vanderweele, Whitney Robinson
Harvard University Biostatistics Working Paper Series
We consider different possible interpretations of the “effect of race” when regressions are run with race as an exposure variable, controlling also for various confounding and mediating variables. When adjustment is made for socioeconomic status early in a person's life, we discuss under what contexts the regression coefficients for race can be interpreted as corresponding to the extent to which a racial disparity would remain if various socioeconomic distributions early in life across racial groups could be equalized. When adjustment is also made for adult socioeconomic status, we note how the overall disparity can be decomposed into the portion that …
Curds And Whey: Little Miss Muffit's Contribution To Multivariate Linear Regression, John Cameron Kidd
Curds And Whey: Little Miss Muffit's Contribution To Multivariate Linear Regression, John Cameron Kidd
Undergraduate Honors Capstone Projects
A common multivariate statistical problem is the prediction of two or more response variables using two or more predictor variables. The simplest model for this situation is the multivariate linear regression model. The standard least squares estimation for this model involves regressing each response variable separately on all the predictor variables. Breiman and Friedman [1] show how to take advantage of correlations among the response variables to increase the predictive accuracy for each of the response variable with an algorithm they call Curds and Whey. In this report, I describe an implementation of the Curds and Whey algorithm in …
Revising Common Core Georgia Performance Standards Statistics Lesson Plans To Better Align With Statistical Practice, Rachel Bonilla
Revising Common Core Georgia Performance Standards Statistics Lesson Plans To Better Align With Statistical Practice, Rachel Bonilla
College of Graduate Studies: Theses & Dissertations
In this thesis, lesson plans provided by the Georgia Department of Education are revised to give students better exposure and practice working with real-life data. Three learning tasks and a performance task are presented covering a unit lesson on statistical regression. The development of Georgia statistics curriculum standards are reviewed and presented.
Statistical Topics Applied To Pressure And Temperature Readings In The Gulf Of Mexico, Malena Kathleen Allison
Statistical Topics Applied To Pressure And Temperature Readings In The Gulf Of Mexico, Malena Kathleen Allison
USF Tampa Graduate Theses and Dissertations
The field of statistical research in weather allows for the application of old and new methods, some of which may describe relationships between certain variables better such as temperatures and pressure. The objective of this study was to apply a variety of traditional and novel statistical methods to analyze data from the National Data Buoy Center, which records among other variables barometric pressure, atmospheric temperature, water temperature and dew point temperature. The analysis included attempts to better describe and model the data as well as to make estimations for certain variables. The following statistical methods were utilized: linear regression, non-response …
Tracking Atlantic Hurricanes Using Statistical Methods, Elizabeth Caitlin Miller
Tracking Atlantic Hurricanes Using Statistical Methods, Elizabeth Caitlin Miller
USF Tampa Graduate Theses and Dissertations
Creating an accurate hurricane location forecasting model is of the utmost importance because of the safety measures that need to occur in the days and hours leading up to a storm's landfall. Hurricanes can be incredibly deadly and costly, but if people are given adequate warning, many lives can be spared. This thesis seeks to develop an accurate model for predicting storm location based on previous location, previous wind speed, and previous pressure. The models are developed using hurricane data from 1980-2009.
An Economic Alternative To The C Chart, Ryan William Black
An Economic Alternative To The C Chart, Ryan William Black
Graduate Theses and Dissertations
Because the probability of Type I error is not evenly distributed beyond upper and lower three-sigma limits the c chart is theoretically inappropriate for a monitor of Poisson distributed phenomena. Furthermore, the normal approximation to the Poisson is of little use when c is small. These practical and theoretical concerns should motivate the computation of true error rates associated with individuals control assuming the Poisson distribution. An economic alternative to the c chart is described as a statistical model of upward shift from c0 to c1 and the two charts are compared in theory. For a range of c chart …
Improved Estimator In The Presence Of Multicollinearity, Ghadban Khalaf
Improved Estimator In The Presence Of Multicollinearity, Ghadban Khalaf
Journal of Modern Applied Statistical Methods
The performances of two biased estimators for the general linear regression model under conditions of collinearity are examined and a new proposed ridge parameter is introduced. Using Mean Square Error (MSE) and Monte Carlo simulation, the resulting estimator’s performance is evaluated and compared with the Ordinary Least Square (OLS) estimator and the Hoerl and Kennard (1970a) estimator. Results of the simulation study indicate that, with respect to MSE criteria, in all cases investigated the proposed estimator outperforms both the OLS and the Hoerl and Kennard estimators.
Analysis Of Roms Estimated Posterior Error Utilizing 4dvar Data Assimilation, Joseph Patrick Horton
Analysis Of Roms Estimated Posterior Error Utilizing 4dvar Data Assimilation, Joseph Patrick Horton
Mathematics
The appropriateness of the approximate error calculated by the Regional Ocean Modeling System (ROMS) is analyzed using Four-Dimensional Data Assimilation (4DVAR) performed on a numerical model of the San Luis Obispo Bay. An effective method of sampling data to minimize the actual error associated with the assimilated numerical model is explored by using different data sampling methods. An idealized state of the SLO bay region ("Real Run") is created to be used as the real ocean, then a numerical model of this region is created approximating this Real Run; this is known as the "Simulated State". By taking samples from …
Number Of Replications Required In Monte Carlo Simulation Studies: A Synthesis Of Four Studies, Daniel J. Mundform, Jay Schaffer, Myoung-Jin Kim, Dale Shaw, Ampai Thongteeraparp, Pornsin Supawan
Number Of Replications Required In Monte Carlo Simulation Studies: A Synthesis Of Four Studies, Daniel J. Mundform, Jay Schaffer, Myoung-Jin Kim, Dale Shaw, Ampai Thongteeraparp, Pornsin Supawan
Journal of Modern Applied Statistical Methods
Monte Carlo simulations are used extensively to study the performance of statistical tests and control charts. Researchers have used various numbers of replications, but rarely provide justification for their choice. Currently, no empirically-based recommendations regarding the required number of replications exist. Twenty-two studies were re-analyzed to determine empirically-based recommendations.
Investigating The Ironwood Tree (Casuarina Equisetifolia) Decline On Guam Using Applied Multinomial Modeling, Karl Anthony Schlub
Investigating The Ironwood Tree (Casuarina Equisetifolia) Decline On Guam Using Applied Multinomial Modeling, Karl Anthony Schlub
LSU Master's Theses
The ironwood tree (Casuarina equisetifolia), a protector of coastlines of the sub-tropical and tropical Western Pacific, is in decline on the island of Guam where aggressive data collection and efforts to mitigate the problem are underway. For each sampled tree the level of decline was measured on an ordinal scale consisting of five categories ranging from healthy to near dead. Several predictors were also measured including tree diameter, fire damage, typhoon damage, presence or absence of termites, presence or absence of basidiocarps, and various geographical or cultural factors. The five decline response levels can be viewed as categories of a …
The Em Algorithm For Group Testing Regression Models Under Matrix Pooling, Christopher R. Bilder, Boan Zhang
The Em Algorithm For Group Testing Regression Models Under Matrix Pooling, Christopher R. Bilder, Boan Zhang
Department of Statistics: Faculty Publications
No abstract provided.
Least Squares Percentage Regression, Chris Tofallis
Least Squares Percentage Regression, Chris Tofallis
Journal of Modern Applied Statistical Methods
In prediction, the percentage error is often felt to be more meaningful than the absolute error. We therefore extend the method of least squares to deal with percentage errors, for both simple and multiple regression. Exact expressions are derived for the coefficients, and we show how such models can be estimated using standard software. When the relative error is normally distributed, least squares percentage regression is shown to provide maximum likelihood estimates. The multiplicative error model is linked to least squares percentage regression in the same way that the standard additive error model is linked to ordinary least squares regression.
Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan
Using Regression Models To Analyze Randomized Trials: Asymptotically Valid Hypothesis Tests Despite Incorrectly Specified Models, Michael Rosenblum, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Regression models are often used to test for cause-effect relationships from data collected in randomized trials or experiments. This practice has deservedly come under heavy scrutiny, since commonly used models such as linear and logistic regression will often not capture the actual relationships between variables, and incorrectly specified models potentially lead to incorrect conclusions. In this paper, we focus on hypothesis test of whether the treatment given in a randomized trial has any effect on the mean of the primary outcome, within strata of baseline variables such as age, sex, and health status. Our primary concern is ensuring that such …
Regression By Data Segments Via Discriminant Analysis, Stan Lipovetsky, Michael Conklin
Regression By Data Segments Via Discriminant Analysis, Stan Lipovetsky, Michael Conklin
Journal of Modern Applied Statistical Methods
It is known that two-group linear discriminant function can be constructed via binary regression. In this article, it is shown that the opposite relation is also relevant – it is possible to present multiple regression as a linear combination of a main part, based on the pooled variance, and Fisher discriminators by data segments. Presenting regression as an aggregate of the discriminators allows one to decompose coefficients of the model into sum of several vectors related to segments. Using this technique provides an understanding of how the total regression model is composed of the regressions by the segments with possible …
On Time Series Analysis Of Public Health And Biomedical Data, Scott L. Zeger, Rafael A. Irizarry, Roger D. Peng
On Time Series Analysis Of Public Health And Biomedical Data, Scott L. Zeger, Rafael A. Irizarry, Roger D. Peng
Johns Hopkins University, Dept. of Biostatistics Working Papers
A time series is a sequence of observations made over time. Examples in public health include daily ozone concentrations, weekly admissions to an emergency department or annual expenditures on health care in the United States. Time series models are used to describe the dependence of the response at each time on predictor variables including covariates and possibly previous values in the series. Time series methods are necessary to account for the correlation among repeated responses over time. This paper gives an overview of time series ideas and methods used in public health research.
The Cross-Validated Adaptive Epsilon-Net Estimator, Mark J. Van Der Laan, Sandrine Dudoit, Aad W. Van Der Vaart
The Cross-Validated Adaptive Epsilon-Net Estimator, Mark J. Van Der Laan, Sandrine Dudoit, Aad W. Van Der Vaart
U.C. Berkeley Division of Biostatistics Working Paper Series
Suppose that we observe a sample of independent and identically distributed realizations of a random variable. Assume that the parameter of interest can be defined as the minimizer, over a suitably defined parameter space, of the expectation (with respect to the distribution of the random variable) of a particular (loss) function of a candidate parameter value and the random variable. Examples of commonly used loss functions are the squared error loss function in regression and the negative log-density loss function in density estimation. Minimizing the empirical risk (i.e., the empirical mean of the loss function) over the entire parameter space …
To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little
To Model Or Not To Model? Competing Modes Of Inference For Finite Population Sampling, Rod Little
The University of Michigan Department of Biostatistics Working Paper Series
Finite population sampling is perhaps the only area of statistics where the primary mode of analysis is based on the randomization distribution, rather than on statistical models for the measured variables. This article reviews the debate between design and model-based inference. The basic features of the two approaches are illustrated using the case of inference about the mean from stratified random samples. Strengths and weakness of design-based and model-based inference for surveys are discussed. It is suggested that models that take into account the sample design and make weak parametric assumptions can produce reliable and efficient inferences in surveys settings. …
Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan
Asymptotics Of Cross-Validated Risk Estimation In Estimator Selection And Performance Assessment, Sandrine Dudoit, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Risk estimation is an important statistical question for the purposes of selecting a good estimator (i.e., model selection) and assessing its performance (i.e., estimating generalization error). This article introduces a general framework for cross-validation and derives distributional properties of cross-validated risk estimators in the context of estimator selection and performance assessment. Arbitrary classes of estimators are considered, including density estimators and predictors for both continuous and polychotomous outcomes. Results are provided for general full data loss functions (e.g., absolute and squared error, indicator, negative log density). A broad definition of cross-validation is used in order to cover leave-one-out cross-validation, V-fold …
Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe
Partial Auc Estimation And Regression, Lori E. Dodd, Margaret S. Pepe
UW Biostatistics Working Paper Series
Accurate disease diagnosis is critical for health care. New diagnostic and screening tests must be evaluated for their abilities to discriminate disease from non-diseased states. The partial area under the ROC curve (partial AUC) is a measure of diagnostic test accuracy. We present an interpretation of the partial AUC that gives rise to a new non-parametric estimator. This estimator is more robust than existing estimators, which make parametric assumptions. We show that the robustness is gained with only a moderate loss in efficiency. We describe a regression modelling framework for making inference about covariate effects on the partial AUC. Such …
Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins
Locally Efficient Estimation Of Regression Parameters Using Current Status Data, Chris Andrews, Mark J. Van Der Laan, James M. Robins
U.C. Berkeley Division of Biostatistics Working Paper Series
In biostatistics applications interest often focuses on the estimation of the distribution of a time-variable T. If one only observes whether or not T exceeds an observed monitoring time C, then the data structure is called current status data, also known as interval censored data, case I. We consider this data structure extended to allow the presence of both time-independent covariates and time-dependent covariate processes that are observed until the monitoring time. We assume that the monitoring process satisfies coarsening at random.
Our goal is to estimate the regression parameter beta of the regression model T = Z*beta+epsilon where the …
Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell
Bivariate Current Status Data, Mark J. Van Der Laan, Nicholas P. Jewell
U.C. Berkeley Division of Biostatistics Working Paper Series
In many applications, it is often of interest to estimate a bivariate distribution of two survival random variables. Complete observation of such random variables is often incomplete. If one only observes whether or not each of the individual survival times exceeds a common observed monitoring time C, then the data structure is referred to as bivariate current status data (Wang and Ding, 2000). For such data, we show that the identifiable part of the joint distribution is represented by three univariate cumulative distribution functions, namely the two marginal cumulative distribution functions, and the bivariate cumulative distribution function evaluated on the …
Self-Consistency: A Fundamental Concept In Statistics, Thaddeus Tarpey, Bernard Flury
Self-Consistency: A Fundamental Concept In Statistics, Thaddeus Tarpey, Bernard Flury
Mathematics and Statistics Faculty Publications
The term ''self-consistency'' was introduced in 1989 by Hastie and Stuetzle to describe the property that each point on a smooth curve or surface is the mean of all points that project orthogonally onto it. We generalize this concept to self-consistent random vectors: a random vector Y is self-consistent for X if E[X|Y] = Y almost surely. This allows us to construct a unified theoretical basis for principal components, principal curves and surfaces, principal points, principal variables, principal modes of variation and other statistical methods. We provide some general results on self-consistent random variables, give …
Generating Unbiased Ratio And Regression Estimators, William (Bill) H. Williams
Generating Unbiased Ratio And Regression Estimators, William (Bill) H. Williams
Publications and Research
Standard ratio and regression are only conditionally unbiased. The paper uses split sample techniques to develop unbiased versions.
Linear Regression Of The Poisson Mean, Duane Steven Brown
Linear Regression Of The Poisson Mean, Duane Steven Brown
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
The purpose of this thesis was to compare two estimation procedures, the method of least squares and the method of maximum likelihood, on sample data obtained from a Poisson distribution. Point estimates of the slope and intercept of the regression line and point estimates of the mean squared error for both the slope and intercept were obtained. It is shown that least squares, the preferred method due to its simplicity, does yield results as good as maximum likelihood.
Also, confidence intervals were computed by Monte Carlo techniques and then were tested for accuracy. For the method of least squares, confidence …
Multicollinearity And The Estimation Of Regression Coefficients, John Charles Teed
Multicollinearity And The Estimation Of Regression Coefficients, John Charles Teed
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
The precision of the estimates of the regression coefficients in a regression analysis is affected by multicollinearity. The effect of certain factors on multicollinearity and the estimates was studied. The response variables were the standard error of the regression coefficients and a standarized statistic that measures the deviation of the regression coefficient from the population parameter.
The estimates are not influenced by any one factor in particular, but rather some combination of factors. The larger the sample size, the better the precision of the estimates no matter how "bad" the other factors may be.
The standard error of the regression …