Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

2020

Discipline
Institution
Keyword
Publication
Publication Type

Articles 31 - 60 of 77

Full-Text Articles in Statistical Theory

Regression: Determining Which Of P Independent Variables Has The Largest Or Smallest Correlation With The Dependent Variable, Plus Results On Ordering The Correlations Winsorized, Rand Wilcox Jul 2020

Regression: Determining Which Of P Independent Variables Has The Largest Or Smallest Correlation With The Dependent Variable, Plus Results On Ordering The Correlations Winsorized, Rand Wilcox

Journal of Modern Applied Statistical Methods

In a regression context, consider p independent variables and a single dependent variable. The paper addresses two goals. The first is to determine the extent it is reasonable to make a decision about whether the largest estimate of the Winsorized correlations corresponds to the independent variable that has the largest population Winsorized correlation. The second is to determine the extent it is reasonable to decide that the order of the estimates of the Winsorized correlations correctly reflects the true ordering. Both goals are addressed by testing relevant hypotheses. Results in Wilcox (in press a) suggest using a multiple comparisons procedure …


Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris Jul 2020

Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris

Journal of Modern Applied Statistical Methods

Re-sampling based statistical tests are known to be computationally heavy, but reliable when small sample sizes are available. Despite their nice theoretical properties not much effort has been put to make them efficient. Computationally efficient method for calculating permutation-based p-values for the Pearson correlation coefficient and two independent samples t-test are proposed. The method is general and can be applied to other similar two sample mean or two mean vectors cases.


Empirical Comparison Of Tests For One-Factor Anova Under Heterogeneity And Non-Normality: A Monte Carlo Study, Diep Nguyen, Eunsook Kim, Yan Wang, Thanh Vinh Pham, Yi-Hsin Chen, Jeffrey D. Kromrey Jul 2020

Empirical Comparison Of Tests For One-Factor Anova Under Heterogeneity And Non-Normality: A Monte Carlo Study, Diep Nguyen, Eunsook Kim, Yan Wang, Thanh Vinh Pham, Yi-Hsin Chen, Jeffrey D. Kromrey

Journal of Modern Applied Statistical Methods

Although the Analysis of Variance (ANOVA) F test is one of the most popular statistical tools to compare group means, it is sensitive to violations of the homogeneity of variance (HOV) assumption. This simulation study examines the performance of thirteen tests in one-factor ANOVA models in terms of their Type I error rate and statistical power under numerous (82,080) conditions. The results show that when HOV was satisfied, the ANOVA F or the Brown-Forsythe test outperformed the other methods in terms of both Type I error control and statistical power even under non-normality. When HOV was violated, the Structured Means …


Joint Models Of Longitudinal Outcomes And Informative Time, Jangdong Seo Jun 2020

Joint Models Of Longitudinal Outcomes And Informative Time, Jangdong Seo

Journal of Modern Applied Statistical Methods

Longitudinal data analyses commonly assume that time intervals are predetermined and have no information regarding the outcomes. However, there might be irregular time intervals and informative time. Presented are joint models and asymptotic behaviors of the parameter estimates. Also, the models are applied for real data sets.


Comparison Of Scale Identification Methods In Mixture Irt Models, Youn-Jeng Choi, Allan S. Cohen Jun 2020

Comparison Of Scale Identification Methods In Mixture Irt Models, Youn-Jeng Choi, Allan S. Cohen

Journal of Modern Applied Statistical Methods

The effects of three scale identification constraints in mixture IRT models were studied. A simulation study found no constraint effect on the mixture Rasch and mixture 2PL models, but the item anchoring constraint was the only one that worked well on selecting correct model with the mixture 3PL model.


Comparing Means Under Heteroscedasticity And Nonnormality: Further Exploring Robust Means Modeling, Alyssa Counsell, Robert Philip Chalmers, Robert A. Cribbie Jun 2020

Comparing Means Under Heteroscedasticity And Nonnormality: Further Exploring Robust Means Modeling, Alyssa Counsell, Robert Philip Chalmers, Robert A. Cribbie

Journal of Modern Applied Statistical Methods

Comparing the means of independent groups is a concern when the assumptions of normality and variance homogeneity are violated. Robust means modeling (RMM) was proposed as an alternative to ANOVA-type procedures when the assumptions of normality and variance homogeneity are violated. The purpose of this study is to compare the Type I error and power rates of RMM to the trimmed Welch procedure. A Monte Carlo study was used to investigate RMM and the trimmed Welch procedure under several conditions of nonnormality and variance heterogeneity. The results suggest that the trimmed Welch provides a better balance of Type I error …


A Note On Inferences About The Probability Of Success, Rand Wilcox Jun 2020

A Note On Inferences About The Probability Of Success, Rand Wilcox

Journal of Modern Applied Statistical Methods

There is an extensive literature dealing with inferences about the probability of success. A minor goal in this note is to point out when certain recommended methods can be unsatisfactory when the sample size is small. The main goal is to report results on the two-sample case. Extant results suggest using one of four methods. The results indicate when computing a 0.95 confidence interval, two of these methods can be more satisfactory when dealing with small sample sizes.


Inferences About The Probability Of Success, Given The Value Of A Covariate, Using A Nonparametric Smoother, Rand Wilcox Jun 2020

Inferences About The Probability Of Success, Given The Value Of A Covariate, Using A Nonparametric Smoother, Rand Wilcox

Journal of Modern Applied Statistical Methods

For a binary random variable Y, let p(x) = P(Y = 1 | X = x) for some covariate X. The goal of computing a confidence interval for p(x) is considered. In the logistic regression model, even a slight departure difficult to detect via a goodness-of-fit test can yield inaccurate results. The accuracy of a confidence interval can deteriorate as the sample size increases. The goal is to suggest an alternative approach based on a smoother, which provides a more flexible approximation of p(x).


On Statistical Significance Of Discriminant Function Coefficients, Tolulope T. Sajobi, Gordon H. Fick, Lisa M. Lix May 2020

On Statistical Significance Of Discriminant Function Coefficients, Tolulope T. Sajobi, Gordon H. Fick, Lisa M. Lix

Journal of Modern Applied Statistical Methods

Discriminant function coefficients are useful for describing group differences and identifying variables that distinguish between groups. Test procedures were compared based on asymptotically approximations, empirical, and exact distributions for testing hypotheses about discriminant function coefficients. These tests are useful for assessing variable importance in multivariate group designs.


Sensitivity Analysis For Incomplete Data And Causal Inference, Heng Chen May 2020

Sensitivity Analysis For Incomplete Data And Causal Inference, Heng Chen

Statistical Science Theses and Dissertations

In this dissertation, we explore sensitivity analyses under three different types of incomplete data problems, including missing outcomes, missing outcomes and missing predictors, potential outcomes in \emph{Rubin causal model (RCM)}. The first sensitivity analysis is conducted for the \emph{missing completely at random (MCAR)} assumption in frequentist inference; the second one is conducted for the \emph{missing at random (MAR)} assumption in likelihood inference; the third one is conducted for one novel assumption, the ``sixth assumption'' proposed for the robustness of instrumental variable estimand in causal inference.


A New Exponential Approach For Reducing The Mean Squared Errors Of The Estimators Of Population Mean Using Conventional And Non-Conventional Location Parameters, Housila P. Singh, Anita Yadav May 2020

A New Exponential Approach For Reducing The Mean Squared Errors Of The Estimators Of Population Mean Using Conventional And Non-Conventional Location Parameters, Housila P. Singh, Anita Yadav

Journal of Modern Applied Statistical Methods

Classes of ratio-type estimators t (say) and ratio-type exponential estimators te (say) of the population mean are proposed, and their biases and mean squared errors under large sample approximation are presented. It is the class of ratio-type exponential estimators te provides estimators more efficient than the ratio-type estimators.


Recurrence Relations For Marginal And Joint Moment Generating Functions Of Topp-Leone Generated Exponential Distribution Based On Record Values And Its Characterization, Zaki Anwar, Neetu Gupta, Mohd Akram Raza Khan, Qazi Azhad Jamal May 2020

Recurrence Relations For Marginal And Joint Moment Generating Functions Of Topp-Leone Generated Exponential Distribution Based On Record Values And Its Characterization, Zaki Anwar, Neetu Gupta, Mohd Akram Raza Khan, Qazi Azhad Jamal

Journal of Modern Applied Statistical Methods

The exact expressions and some recurrence relations are derived for marginal and joint moment generating functions of kth lower record values from Topp-Leone Generated (TLG) Exponential distribution. This distribution is characterized by using the recurrence relation of the marginal moment generating function of kth lower record values.


An Improved Two Independent-Samples Randomization Test For Single-Case Ab-Type Intervention Designs: A 20-Year Journey, Joel R. Levin, John M. Ferron, Boris S. Gafurov May 2020

An Improved Two Independent-Samples Randomization Test For Single-Case Ab-Type Intervention Designs: A 20-Year Journey, Joel R. Levin, John M. Ferron, Boris S. Gafurov

Journal of Modern Applied Statistical Methods

Detailed is a 20-year arduous journey to develop a statistically viable two-phase (AB) single-case two independent-samples randomization test procedure. The test is designed to compare the effectiveness of two different interventions that are randomly assigned to cases. In contrast to the unsatisfactory simulation results produced by an earlier proposed randomization test, the present test consistently exhibited acceptable Type I error control under various design and effect-type configurations, while at the same time possessing adequate power to detect moderately sized intervention-difference effects. Selected issues, applications, and a multiple-baseline extension of the two-sample test are discussed.


Support Vector Machine-Based Modified Sp Statistic For Subset Selection With Non-Normal Error Terms, Shivaji Shripati Desai, D N. Kashid May 2020

Support Vector Machine-Based Modified Sp Statistic For Subset Selection With Non-Normal Error Terms, Shivaji Shripati Desai, D N. Kashid

Journal of Modern Applied Statistical Methods

Support vector machine (SVM) is used for estimation of regression parameters to modify the sum of cross products (Sp). It works well for some nonnormal error distributions. The performance of existing robust methods and the modified Sp is evaluated through simulated and real data. The results show the performance of the modified Sp is good.


Using Saddlepoint Approximations And Likelihood-Based Methods To Conduct Statistical Inference For The Mean Of The Beta Distribution, Bryn Brakefield May 2020

Using Saddlepoint Approximations And Likelihood-Based Methods To Conduct Statistical Inference For The Mean Of The Beta Distribution, Bryn Brakefield

Electronic Theses and Dissertations

The prevalence of conducting statistical inference for the mean of the beta distribution has been rising in various fields of academic research, such as in immunology that analyzes proportions of rare cell population subsets. For our purposes, we will address this statistical inference problem by using likelihood-based applications to hypothesis testing, along with a relatively new statistical method called saddlepoint approximations. Through simulation work, we will compare the performance of these statistical procedures and provide both the statistical and scientific communities with recommendations on best practices.


Using Stability To Select A Shrinkage Method, Dean Dustin May 2020

Using Stability To Select A Shrinkage Method, Dean Dustin

Department of Statistics: Dissertations, Theses, and Student Research

Shrinkage methods are estimation techniques based on optimizing expressions to find which variables to include in an analysis, typically a linear regression. The general form of these expressions is the sum of an empirical risk plus a complexity penalty based on the number of parameters. Many shrinkage methods are known to satisfy an ‘oracle’ property meaning that asymptotically they select the correct variables and estimate their coefficients efficiently. In Section 1.2, we show oracle properties in two general settings. The first uses a log likelihood in place of the empirical risk and allows a general class of penalties. The second …


On Arnold–Villasenor Conjectures For Characterizaing Exponential Distribution Based On Sample Of Size Three, George Yanev May 2020

On Arnold–Villasenor Conjectures For Characterizaing Exponential Distribution Based On Sample Of Size Three, George Yanev

School of Mathematical & Statistical Sciences Faculty Publications

Arnold and Villasenor [4] obtain a series of characterizations of the exponential distribution based on random samples of size two. These results were already applied in constructing goodness-of-fit tests. Extending the techniques from [4], we prove some of Arnold and Villasenor’s conjectures for samples of size three. An example with simulated data is discussed.


Logistic Growth Modeling With Markov Chain Monte Carlo Estimation, Jaehwa Choi, Jinsong Chen, Jeffrey R. Harring Apr 2020

Logistic Growth Modeling With Markov Chain Monte Carlo Estimation, Jaehwa Choi, Jinsong Chen, Jeffrey R. Harring

Journal of Modern Applied Statistical Methods

A new growth modeling approach is proposed to can fit inherently nonlinear (i.e., logistic) function without constraint nor reparameterization. A simulation study is employed to investigate the feasibility and performance of a Markov chain Monte Carlo method within Bayesian estimation framework to estimate a fully random version of a logistic growth curve model under manipulated conditions such as the number and timing of measurement occasions and sample sizes.


A Simulation Study On Increasing Capture Periods In Bayesian Closed Population Capture-Recapture Models With Heterogeneity, Ross M. Gosky, Joel Sanqui Apr 2020

A Simulation Study On Increasing Capture Periods In Bayesian Closed Population Capture-Recapture Models With Heterogeneity, Ross M. Gosky, Joel Sanqui

Journal of Modern Applied Statistical Methods

Capture-Recapture models are useful in estimating unknown population sizes. A common modeling challenge for closed population models involves modeling unequal animal catchability in each capture period, referred to as animal heterogeneity. Inference about population size N is dependent on the assumed distribution of animal capture probabilities in the population, and that different models can fit a data set equally well but provide contradictory inferences about N. Three common Bayesian Capture-Recapture heterogeneity models are studied with simulated data to study the prevalence of contradictory inferences is in different population sizes with relatively low capture probabilities, specifically at different numbers of …


An Exploration Of Link Functions Used In Ordinal Regression, Thomas J. Smith, David A. Walker, Cornelius M. Mckenna Apr 2020

An Exploration Of Link Functions Used In Ordinal Regression, Thomas J. Smith, David A. Walker, Cornelius M. Mckenna

Journal of Modern Applied Statistical Methods

The purpose of this study is to examine issues involved with choice of a link function in generalized linear models with ordinal outcomes, including distributional appropriateness, link specificity, and palindromic invariance are discussed and an exemplar analysis provided using the Pew Research Center 25th anniversary of the Web Omnibus Survey data. Simulated data are used to compare the relative palindromic invariance of four distinct indices of determination/discrimination, including a newly proposed index by Smith et al. (2017).


Personal Foul: How Head Trauma And The Insurance Industry Are Threatening Sports, Zachary Cooler Apr 2020

Personal Foul: How Head Trauma And The Insurance Industry Are Threatening Sports, Zachary Cooler

Senior Honors Theses

This thesis will investigate the growing problem of head trauma in contact sports like football, hockey, and soccer through medical studies, implications to the insurance industry, and ongoing litigation. The thesis will investigate medical studies that are finding more evidence to support the claim that contact sports players are more likely to receive head trauma symptoms such as memory loss, mood swings, and even Lou Gehrig’s disease in extreme cases. The thesis will also demonstrate that these medical symptoms and monetary losses from medical claims are convincing insurance companies to withdraw insurance coverage for sports leagues, which they are justifying …


Quasi-Likelihood Ratio Tests For Homoscedasticity In Linear Regression, Lili Yu, Varadan Sevilimedu, Robert Vogel, Hani Samawi Apr 2020

Quasi-Likelihood Ratio Tests For Homoscedasticity In Linear Regression, Lili Yu, Varadan Sevilimedu, Robert Vogel, Hani Samawi

Journal of Modern Applied Statistical Methods

Two quasi-likelihood ratio tests are proposed for the homoscedasticity assumption in the linear regression models. They require few assumptions than the existing tests. The properties of the tests are investigated through simulation studies. An example is provided to illustrate the usefulness of the new proposed tests.


Investigating The Performance Of Propensity Score Approaches For Differential Item Functioning Analysis, Yan Liu, Chanmin Kim, Amrey D. Wu, Paul Gustafson, Edward Kroc, Bruno D. Zumbo Apr 2020

Investigating The Performance Of Propensity Score Approaches For Differential Item Functioning Analysis, Yan Liu, Chanmin Kim, Amrey D. Wu, Paul Gustafson, Edward Kroc, Bruno D. Zumbo

Journal of Modern Applied Statistical Methods

To evaluate the performance of propensity score approaches for differential item functioning analysis, this simulation study was conducted to assess bias, mean square error, Type I error, and power under different levels of effect size and a variety of model misspecification conditions, including different types and missing patterns of covariates.


A Glossary On Building Longitudinal, Population-Based Data Linkages To Explore Children’S Developmental Trajectories, Jennifer E. V. Lloyd, Jacqui Boonstra, Lisa Chen, Barry Forer, Ruth Hershler, Constance Milbrath, Brenda T. Poon, Neda Razaz, Pippa Rowcliffe, Kimberly Schonert-Reichl Apr 2020

A Glossary On Building Longitudinal, Population-Based Data Linkages To Explore Children’S Developmental Trajectories, Jennifer E. V. Lloyd, Jacqui Boonstra, Lisa Chen, Barry Forer, Ruth Hershler, Constance Milbrath, Brenda T. Poon, Neda Razaz, Pippa Rowcliffe, Kimberly Schonert-Reichl

Journal of Modern Applied Statistical Methods

Population-based, person-specific, longitudinal child and youth health and developmental data linkages involve connecting combinations of specially-collected data and administrative data for longitudinal population research purposes. This glossary provides definitions of key terms and concepts related to their theoretical basis, research infrastructure, research methodology, statistical analysis, and knowledge translation.


Robust Confidence Intervals For The Population Mean Alternatives To The Student-T Confidence Interval, Moustafa Omar Ahmed Abu-Shawiesh, Aamir Saghir Apr 2020

Robust Confidence Intervals For The Population Mean Alternatives To The Student-T Confidence Interval, Moustafa Omar Ahmed Abu-Shawiesh, Aamir Saghir

Journal of Modern Applied Statistical Methods

In this paper, three robust confidence intervals are proposed as alternatives to the Student‑t confidence interval. The performance of these intervals was compared through a simulation study shows that Qn-t confidence interval performs the best and it is as good as Student’s‑t confidence interval. Real-life data was used for illustration and performing a comparison that support the findings obtained from the simulation study.


Using Spss To Analyze Complex Survey Data: A Primer, Danjie Zou, Jennifer E. V. Lloyd, Jennifer L. Baumbusch Apr 2020

Using Spss To Analyze Complex Survey Data: A Primer, Danjie Zou, Jennifer E. V. Lloyd, Jennifer L. Baumbusch

Journal of Modern Applied Statistical Methods

An introduction to using SPSS to analyze complex survey data is given. Key features of complex survey design are described briefly, including stratification, clustering, multiple stages, and weights. Then, annotated SPSS syntax for complex survey data analysis is presented to demonstrate the step-by-step process using real complex samples data.


On The Authentic Notion, Relevance, And Solution Of The Jeffreys-Lindley Paradox In The Zettabyte Era, Miodrag M. Lovric Apr 2020

On The Authentic Notion, Relevance, And Solution Of The Jeffreys-Lindley Paradox In The Zettabyte Era, Miodrag M. Lovric

Journal of Modern Applied Statistical Methods

The Jeffreys-Lindley paradox is the most quoted divergence between the frequentist and Bayesian approaches to statistical inference. It is embedded in the very foundations of statistics and divides frequentist and Bayesian inference in an irreconcilable way. This paradox is the Gordian Knot of statistical inference and Data Science in the Zettabyte Era. If statistical science is ready for revolution confronted by the challenges of massive data sets analysis, the first step is to finally solve this anomaly. For more than sixty years, the Jeffreys-Lindley paradox has been under active discussion and debate. Many solutions have been proposed, none entirely satisfactory. …


A Monte Carlo Analysis Of Standard Error-Based Methods For Computing Confidence Intervals, Elayna Wichert Apr 2020

A Monte Carlo Analysis Of Standard Error-Based Methods For Computing Confidence Intervals, Elayna Wichert

Masters Theses & Specialist Projects

The objective of this study is to empirically test existing techniques to calculate the likely range of values for a Classical Test Theory true score given an observed score. The traditional method for forming these confidence intervals has used the standard error of measurement (SEM) as the basis for this confidence interval. An alternate equation, the standard error of estimate (SEE), has been recommended in place of the SEM for this purpose, yet it remains overlooked in the field of psychometrics. It is important that the correct equation be used in various applications in personnel psychology. Monte Carlo analyses were …


Conflicts In Bayesian Statistics Between Inference Based On Credible Intervals And Bayes Factors, Miodrag M. Lovric Apr 2020

Conflicts In Bayesian Statistics Between Inference Based On Credible Intervals And Bayes Factors, Miodrag M. Lovric

Journal of Modern Applied Statistical Methods

In frequentist statistics, point-null hypothesis testing based on significance tests and confidence intervals are harmonious procedures and lead to the same conclusion. This is not the case in the domain of the Bayesian framework. An inference made about the point-null hypothesis using Bayes factor may lead to an opposite conclusion if it is based on the Bayesian credible interval. Bayesian suggestions to test point-nulls using credible intervals are misleading and should be dismissed. A null hypothesized value may be outside a credible interval but supported by Bayes factor (a Type I conflict), or contrariwise, the null value may be inside …


A New Liu Type Of Estimator For The Restricted Sur Estimator, Kristofer Månsson, B. M. Golam Kibria, Ghazi Shukur Mar 2020

A New Liu Type Of Estimator For The Restricted Sur Estimator, Kristofer Månsson, B. M. Golam Kibria, Ghazi Shukur

Journal of Modern Applied Statistical Methods

A new Liu type of estimator for the seemingly unrelated regression (SUR) models is proposed that may be used when estimating the parameters vector in the presence of multicollinearity if the it is suspected to belong to a linear subspace. The dispersion matrices and the mean squared error (MSE) are derived. The new estimator may have a lower MSE than the traditional estimators. It was shown using simulation techniques the new shrinkage estimator outperforms the commonly used estimators in the presence of multicollinearity.