Logistic Regression: An Inferential Method For Identifying The Best Predictors,
2019
University of Southern California
Logistic Regression: An Inferential Method For Identifying The Best Predictors, Rand Wilcox
Journal of Modern Applied Statistical Methods
When dealing with a logistic regression model, there is a simple method for estimating the strength of the association between the jth covariate and the dependent variable when all covariates are entered into the model. There is the issue of determining whether the jth independent variable has a stronger or weaker association than the kth independent variable. This note describes a method for dealing with this issue that was found to perform reasonably well in simulations.
Non Parametric Test For Testing Exponentiality Against Exponential Better Than Used In Laplace Transform Order,
2019
Al-Azhar University
Non Parametric Test For Testing Exponentiality Against Exponential Better Than Used In Laplace Transform Order, Mahmoud Mansour, M A W Mahmoud Prof.
Basic Science Engineering
In this paper, the test statistic for testing exponentiality against exponential better than used in Laplace transform order (EBUL) based on the Laplace transform technique is proposed. Pitman’s asymptotic efficiency of our test is calculated and compared with other tests. The percentiles of this test are tabulated. The powers of the test are estimated for famously used distributions in aging problems. In the case of censored data, our test is applied and the percentiles are also calculated and tabulated. Finally, real examples in different areas are utilized as practical applications for the proposed test.
Neural Shrubs: Using Neural Networks To Improve Decision Trees,
2019
SDSMT
Neural Shrubs: Using Neural Networks To Improve Decision Trees, Kyle Caudle, Randy Hoover, Aaron Alphonsus
SDSU Data Science Symposium
Decision trees are a method commonly used in machine learning to either predict a categorical response or a continuous response variable. Once the tree partitions the space, the response is either determined by the majority vote – classification trees, or by averaging the response values – regression trees. This research builds a standard regression tree and then instead of averaging the responses, we train a neural network to determine the response value. We have found that our approach typically increases the predicative capability of the decision tree. We have 2 demonstrations of this approach that we wish to present as …
Counting And Coloring Sudoku Graphs,
2019
Portland State University
Counting And Coloring Sudoku Graphs, Kyle Oddson
Mathematics and Statistics Dissertations, Theses, and Final Project Papers
A sudoku puzzle is most commonly a 9 × 9 grid of 3 × 3 boxes wherein the puzzle player writes the numbers 1 - 9 with no repetition in any row, column, or box. We generalize the notion of the n2 × n2 sudoku grid for all n ∈ ℤ≥2 and codify the empty sudoku board as a graph. In the main section of this paper we prove that sudoku boards and sudoku graphs exist for all such n; we prove the equivalence of [3]'s construction using unions and products of graphs to the definition of …
Controlling For Confounding Via Propensity Score Methods Can Result In Biased Estimation Of The Conditional Auc: A Simulation Study,
2019
Old Dominion University
Controlling For Confounding Via Propensity Score Methods Can Result In Biased Estimation Of The Conditional Auc: A Simulation Study, Hadiza I. Galadima, Donna K. Mcclish
Community & Environmental Health Faculty Publications
In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not random in observational studies, comparisons of outcomes between exposed and nonexposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of conditional odds ratio and hazard ratio. However, research is lacking on the performance of propensity score methods for covariate adjustment when estimating the …
Some New Generalized Distribution Via Lindley-Weibuli And Lindley-Log-Logistic Distributions With Applications,
2019
Georgia Southern University
Some New Generalized Distribution Via Lindley-Weibuli And Lindley-Log-Logistic Distributions With Applications, Soliu A. Raheem
College of Graduate Studies: Theses & Dissertations
In this thesis, new generalized distributions, namely Beta Lindley-Log-Logistic (BLLLoG) distribution, Marshall-Olkin Lindley-Weibull (MOLW) distribution, and Gamma LindleyWeibull (GLW) distribution as well as related sub-distributions are proposed. Series expansion of the densities are obtained. Statistical properties of these distributions, including hazard function, reverse hazard function, moments, reliability, quantile function, mean deviations, Bonferroni and Lorenz curves, entropy and Fisher information are derived. Method of maximum likelihood is used to estimate the parameters of the new distributions. Monte Carlo simulation is employed to examine the performance of the proposed distributions. Applications of the generalized distributions to real lifetime data are presented to …
Modeling Stochastically Intransitive Relationships In Paired Comparison Data,
2019
Southern Methodist University
Modeling Stochastically Intransitive Relationships In Paired Comparison Data, Ryan Patrick Alexander Mcshane
Statistical Science Theses and Dissertations
If the Warriors beat the Rockets and the Rockets beat the Spurs, does that mean that the Warriors are better than the Spurs? Sophisticated fans would argue that the Warriors are better by the transitive property, but could Spurs fans make a legitimate argument that their team is better despite this chain of evidence?
We first explore the nature of intransitive (rock-scissors-paper) relationships with a graph theoretic approach to the method of paired comparisons framework popularized by Kendall and Smith (1940). Then, we focus on the setting where all pairs of items, teams, players, or objects have been compared to …
A Flexible Zero-Inflated Poisson Regression Model,
2019
University of Kentucky
A Flexible Zero-Inflated Poisson Regression Model, Eric S. Roemmele
Theses and Dissertations--Statistics
A practical problem often encountered with observed count data is the presence of excess zeros. Zero-inflation in count data can easily be handled by zero-inflated models, which is a two-component mixture of a point mass at zero and a discrete distribution for the count data. In the presence of predictors, zero-inflated Poisson (ZIP) regression models are, perhaps, the most commonly used. However, the fully parametric ZIP regression model could sometimes be restrictive, especially with respect to the mixing proportions. Taking inspiration from some of the recent literature on semiparametric mixtures of regressions models for flexible mixture modeling, we propose a …
Statistical Designs For Network A/B Testing,
2019
Virginia Commonwealth University
Statistical Designs For Network A/B Testing, Victoria V. Pokhilko
Theses and Dissertations
A/B testing refers to the statistical procedure of experimental design and analysis to compare two treatments, A and B, applied to different testing subjects. It is widely used by technology companies such as Facebook, LinkedIn, and Netflix, to compare different algorithms, web-designs, and other online products and services. The subjects participating in these online A/B testing experiments are users who are connected in different scales of social networks. Two connected subjects are similar in terms of their social behaviors, education and financial background, and other demographic aspects. Hence, it is only natural to assume that their reactions to online products …
Statistical Modeling Of Influenza-Like-Illness In Montana Using Spatial And Temporal Methods,
2019
University of Montana, Missoula
Statistical Modeling Of Influenza-Like-Illness In Montana Using Spatial And Temporal Methods, Benjamin A. Stark
Graduate Student Theses, Dissertations, & Professional Papers
Studying air pollution and public health has been a historically important question in science. It has long been hypothesized that severe air pollution conditions lead to negative implications in basic human health. Primarily, areas thats are prone to severe degrees of human pollution are the focus of such studies. Such research relating to less populated areas are scarce, and this scarcity raises the question of how such pollution dynamics (human-made and natural) influence human health in more rural areas.
The aim of this study is to explore this hole in research; in particular we explore possible links between air pollution …
A Proficient Two-Stage Stratified Randomized Response Strategy,
2018
Islamic University of Science and Technology, Awantipora, India
A Proficient Two-Stage Stratified Randomized Response Strategy, Tanveer A. Tarray, Housila P. Singh
Journal of Modern Applied Statistical Methods
A stratified randomized response model based on R. Singh, Singh, Mangat, and Tracy (1995) improved two-stage randomized response strategy is proposed. It has an optimal allocation and large gain in precision. Conditions are obtained under which the proposed model is more efficient than R. Singh et al. (1995) and H. P. Singh and Tarray (2015) models. Numerical illustrations are also given in support of the present study.
Simple Unbalanced Ranked Set Sampling For Mean Estimation Of Response Variable Of Developmental Programs,
2018
Indian Council of Forestry Research and Education
Simple Unbalanced Ranked Set Sampling For Mean Estimation Of Response Variable Of Developmental Programs, Girish Chandra, Dinesh S. Bhoj, Rajiv Pandey
Journal of Modern Applied Statistical Methods
An unbalanced ranked set sampling (RSS) procedure on the skewed survey variable is proposed to estimate the population mean of a response variable from the area of developmental programs which are generally implemented under different phases. It is based on the unbalanced RSS under linear impacts of the program and is compared with the estimators based on simple random sampling (SRS) and balanced RSS. It is shown that the relative precision of the proposed estimator is higher than those of the estimators based on SRS and balanced RSS for three chosen skewed distributions of survey variables.
Extended Method For Several Dichotomous Covariates To Estimate The Instantaneous Risk Function Of The Aalen Additive Model,
2018
Federal University of São João del Rei
Extended Method For Several Dichotomous Covariates To Estimate The Instantaneous Risk Function Of The Aalen Additive Model, Luciane Teixeira Passos Giarola, Mario Javier Ferrua Vivanco, Marcelo Angelo Cirillo, Fortunato Silva Menezes
Journal of Modern Applied Statistical Methods
The instantaneous risk function of Aalen’s model is estimated considering dichotomous covariates, using parametric accumulated risk functions to smooth cumulative risk of Aalen by grouping the individuals into sets named parcels. This methodology can be used for data with dichotomous covariates.
The Impact Of Sample Size In Cross-Classified Multiple Membership Multilevel Models,
2018
Chungnam National University
The Impact Of Sample Size In Cross-Classified Multiple Membership Multilevel Models, Hyewon Chung, Jiseon Kim, Ryoungsun Park, Hyeonjeong Jean
Journal of Modern Applied Statistical Methods
A simulation study was conducted to examine parameter recovery in a cross-classified multiple membership multilevel model. No substantial relative bias was identified for the fixed effect or level-one variance component estimates. However, the level-two cross-classification multiple membership factor variance components were substantially biased with relatively fewer groups.
Using Cyclical Components To Improve The Forecasts Of The Stock Market And Macroeconomic Variables,
2018
Curtin University Malaysia
Using Cyclical Components To Improve The Forecasts Of The Stock Market And Macroeconomic Variables, Kenneth R. Szulczyk, Shibley Sadique
Journal of Modern Applied Statistical Methods
Economic variables such as stock market indices, interest rates, and national output measures contain cyclical components. Forecasting methods excluding these cyclical components yield inaccurate out-of-sample forecasts. Accordingly, a three-stage procedure is developed to estimate a vector autoregression (VAR) with cyclical components. A Monte Carlo simulation shows the procedure estimates the parameters accurately. Subsequently, a VAR with cyclical components improves the root-mean-square error of out-of-sample forecasts by 50% for a stock market model with macroeconomic variables.
Comparison Of Multiple Imputation Methods For Categorical Survey Items With High Missing Rates: Application To The Family Life, Activity, Sun, Health And Eating (Flashe) Study,
2018
National Cancer Institute
Comparison Of Multiple Imputation Methods For Categorical Survey Items With High Missing Rates: Application To The Family Life, Activity, Sun, Health And Eating (Flashe) Study, Benmei Liu, Erin Hennessy, April Oh, Laura A. Dwyer, Linda Nebeling
Journal of Modern Applied Statistical Methods
Two multiple imputation methods, the Sequential Regression Multivariate Imputation Algorithm and the Cox-Lannacchione Weighted Sequential Hotdeck, were examined and compared to impute highly missing categorical variables from the Family Life, Activity, Sun, Health and Eating (FLASHE) study. This paper describes the imputation approaches and results from the study.
Dealing With Sensitive Quantitative Variables: A Comparison Of Sampling Designs For The Procedure Of Gupta And Thornton,
2018
University of Havana
Dealing With Sensitive Quantitative Variables: A Comparison Of Sampling Designs For The Procedure Of Gupta And Thornton, Carlos Narciso Bouza Herrera, Prayas Sharma
Journal of Modern Applied Statistical Methods
The use of randomized response procedures allows diminishing the number of non-responses and increasing the accuracy of the responses. A new sampling strategy is developed where the reports are scrambled using the procedure of Gupta and Thornton. The estimator of the mean as well as the errors are developed for the Rao-Hartley-Cochran and Ranked Sets Sampling designs. The proposals are compared with the original model based on the use of simple random sampling.
Bayesian And Semi-Bayesian Estimation Of The Parameters Of Generalized Inverse Weibull Distribution,
2018
Panjab University, Chandigarh, India
Bayesian And Semi-Bayesian Estimation Of The Parameters Of Generalized Inverse Weibull Distribution, Kamaljit Kaur, Kalpana K. Mahajan, Sangeeta Arora
Journal of Modern Applied Statistical Methods
Bayesian and semi-Bayesian estimators of parameters of the generalized inverse Weibull distribution are obtained using Jeffreys’ prior and informative prior under specific assumptions of loss function. Using simulation, the relative efficiency of the proposed estimators is obtained under different set-ups. A real life example is also given.
Overcoming Small Data Limitations In Heart Disease Prediction By Using Surrogate Data,
2018
Southern Methodist University
Overcoming Small Data Limitations In Heart Disease Prediction By Using Surrogate Data, Alfeo Sabay, Laurie Harris, Vivek Bejugama, Karen Jaceldo-Siegl
SMU Data Science Review
In this paper, we present a heart disease prediction use case showing how synthetic data can be used to address privacy concerns and overcome constraints inherent in small medical research data sets. While advanced machine learning algorithms, such as neural networks models, can be implemented to improve prediction accuracy, these require very large data sets which are often not available in medical or clinical research. We examine the use of surrogate data sets comprised of synthetic observations for modeling heart disease prediction. We generate surrogate data, based on the characteristics of original observations, and compare prediction accuracy results achieved from …
Minimizing The Perceived Financial Burden Due To Cancer,
2018
Southern Methodist University
Minimizing The Perceived Financial Burden Due To Cancer, Hassan Azhar, Zoheb Allam, Gino Varghese, Daniel W. Engels, Sajiny John
SMU Data Science Review
In this paper, we present a regression model that predicts perceived financial burden that a cancer patient experiences in the treatment and management of the disease. Cancer patients do not fully understand the burden associated with the cost of cancer, and their lack of understanding can increase the difficulties associated with living with the disease, in particular coping with the cost. The relationship between demographic characteristics and financial burden were examined in order to better understand the characteristics of a cancer patient and their burden, while all subsets regression was used to determine the best predictors of financial burden. Age, …
