Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

2,918 Full-Text Articles 4,214 Authors 3,508,612 Downloads 170 Institutions

All Articles in Applied Statistics

Faceted Search

2,918 full-text articles. Page 38 of 101.

Identifying Which Of J Independent Binomial Distributions Has The Largest Probability Of Success, Rand Wilcox 2020 University of Southern California

Identifying Which Of J Independent Binomial Distributions Has The Largest Probability Of Success, Rand Wilcox

Journal of Modern Applied Statistical Methods

Let p1,…, pJ denote the probability of a success for J independent random variables having a binomial distribution and let p(1) ≤ … ≤ p(J) denote these probabilities written in ascending order. The goal is to make a decision about which group has the largest probability of a success, p(J). Let p̂1,…, p̂J denote estimates of p1,…,pJ, respectively. The strategy is to test J − 1 hypotheses comparing the group with the largest estimate to each of the J − 1 …


Jmasm 53: Miccerird, Michael Lance 2020 ExceLance, LLC

Jmasm 53: Miccerird, Michael Lance

Journal of Modern Applied Statistical Methods

Fortran 77 and 90 modules (REALPOPS.lib) exist for invoking the 8 distributions estimated by Micceri (1989). These respective modules were created by Sawilowsky et al. (1990) and Sawilowsky and Fahoome (2003). The MicceriRD (Micceri’s Real Distributions) Python package was created because Python is increasingly used for data analysis and, in some cases, Monte Carlo simulations.


Bayesian Analysis Of Extended Cox Model With Time-Varying Covariates Using Bootstrap Prior, Oyebayo R. Olaniran, Mohd Asrul A. Abdullah 2020 University of Ilorin

Bayesian Analysis Of Extended Cox Model With Time-Varying Covariates Using Bootstrap Prior, Oyebayo R. Olaniran, Mohd Asrul A. Abdullah

Journal of Modern Applied Statistical Methods

A new Bayesian estimation procedure for extended cox model with time varying covariate was presented. The prior was determined using bootstrapping technique within the framework of parametric empirical Bayes. The efficiency of the proposed method was observed using Monte Carlo simulation of extended Cox model with time varying covariates under varying scenarios. Validity of the proposed method was also ascertained using real life data set of Stanford heart transplant. Comparison of the proposed method with its competitor established appreciable supremacy of the method.


Regression: Determining Which Of P Independent Variables Has The Largest Or Smallest Correlation With The Dependent Variable, Plus Results On Ordering The Correlations Winsorized, Rand Wilcox 2020 University of Southern California

Regression: Determining Which Of P Independent Variables Has The Largest Or Smallest Correlation With The Dependent Variable, Plus Results On Ordering The Correlations Winsorized, Rand Wilcox

Journal of Modern Applied Statistical Methods

In a regression context, consider p independent variables and a single dependent variable. The paper addresses two goals. The first is to determine the extent it is reasonable to make a decision about whether the largest estimate of the Winsorized correlations corresponds to the independent variable that has the largest population Winsorized correlation. The second is to determine the extent it is reasonable to decide that the order of the estimates of the Winsorized correlations correctly reflects the true ordering. Both goals are addressed by testing relevant hypotheses. Results in Wilcox (in press a) suggest using a multiple comparisons procedure …


Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris 2020 University of Crete, Greece

Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris

Journal of Modern Applied Statistical Methods

Re-sampling based statistical tests are known to be computationally heavy, but reliable when small sample sizes are available. Despite their nice theoretical properties not much effort has been put to make them efficient. Computationally efficient method for calculating permutation-based p-values for the Pearson correlation coefficient and two independent samples t-test are proposed. The method is general and can be applied to other similar two sample mean or two mean vectors cases.


Empirical Comparison Of Tests For One-Factor Anova Under Heterogeneity And Non-Normality: A Monte Carlo Study, Diep Nguyen, Eunsook Kim, Yan Wang, Thanh Vinh Pham, Yi-Hsin Chen, Jeffrey D. Kromrey 2020 University of South Florida

Empirical Comparison Of Tests For One-Factor Anova Under Heterogeneity And Non-Normality: A Monte Carlo Study, Diep Nguyen, Eunsook Kim, Yan Wang, Thanh Vinh Pham, Yi-Hsin Chen, Jeffrey D. Kromrey

Journal of Modern Applied Statistical Methods

Although the Analysis of Variance (ANOVA) F test is one of the most popular statistical tools to compare group means, it is sensitive to violations of the homogeneity of variance (HOV) assumption. This simulation study examines the performance of thirteen tests in one-factor ANOVA models in terms of their Type I error rate and statistical power under numerous (82,080) conditions. The results show that when HOV was satisfied, the ANOVA F or the Brown-Forsythe test outperformed the other methods in terms of both Type I error control and statistical power even under non-normality. When HOV was violated, the Structured Means …


Maximum Likelihood Estimation Of Species Trees And Anomaly Zone Detection Using Ranked Gene Trees, Anastasiia Kim 2020 University of New Mexico

Maximum Likelihood Estimation Of Species Trees And Anomaly Zone Detection Using Ranked Gene Trees, Anastasiia Kim

Mathematics & Statistics ETDs

A phylogenetic tree represents the evolutionary relationships among a set of organisms. Gene trees can be used to reconstruct phylogenetic trees. The methods in this dissertation focus on the gene tree topologies with emphasis on ranked gene tree topologies. A ranked tree depicts the order in which nodes appear in the tree together with topological relationships among gene lineages. One challenge that arises during phylogenetic inference is the existence of the anomaly zones, the regions of branch-length space in the species tree that can produce gene trees that have topologies differing from the species tree topology but are more probable …


Improving The Quality And Design Of Retrospective Clinical Outcome Studies That Utilize Electronic Health Records, Oliwier Dziadkowiec, Jeffery Durbin, Vignesh Jayaraman Muralidharan, Megan Novak, Brendon Cornett 2020 HCA Healthcare Mountain MidAmerica and Continental Divisions

Improving The Quality And Design Of Retrospective Clinical Outcome Studies That Utilize Electronic Health Records, Oliwier Dziadkowiec, Jeffery Durbin, Vignesh Jayaraman Muralidharan, Megan Novak, Brendon Cornett

HCA Healthcare Journal of Medicine

Electronic health records (EHRs) are an excellent source for secondary data analysis. Studies based on EHR-derived data, if designed properly, can answer previously unanswerable clinical research questions. In this paper we will highlight the benefits of large retrospective studies from secondary sources such as EHRs, examine retrospective cohort and case-control study design challenges, as well as methodological and statistical adjustment that can be made to overcome some of the inherent design limitations, in order to increase the generalizability, validity and reliability of the results obtained from these studies.


Assessing Differential Item Functioning In The Perceived Stress Scale, Nana Amma Berko Asamoah 2020 University of Arkansas, Fayetteville

Assessing Differential Item Functioning In The Perceived Stress Scale, Nana Amma Berko Asamoah

Graduate Theses and Dissertations

When an item on a test functions differently for subgroups of respondents with respect to an exogenous variable (or covariate) after conditioning on the latent variable of interest, the item is said to exhibit Differential Item Functioning (DIF). The 10-item Perceived Stress Scale (PSS10) is administered to respondents via MTurk to quantify “perceived stress” and identify if items on the scale function differently for specific subgroups defined by age, sex, race, marital status, number of children, employment status and social media usage.

The purpose of this study was to compare traditional DIF detection approaches (Mantel-Haenszel, logistic regression, likelihood ratio test …


Eco 230 / Mgt 230 Introduction To Economic And Managerial Statistics, George Vachadze 2020 CUNY College of Staten Island

Eco 230 / Mgt 230 Introduction To Economic And Managerial Statistics, George Vachadze

Open Educational Resources

Development and application of modern statistical methods, including such elements of descriptive statistics and statistical inference as correlation and regression analysis, probability theory, sampling procedures, normal distribution and binomial distribution, estimation, and testing of hypotheses.


Working Children On Java Island 2017, Yuniarti 2020 Syracuse University

Working Children On Java Island 2017, Yuniarti

International Programs

Children's wellbeing has currently become a global concern as many of them are engaged in the labor force. A small area estimation (SAE) technique, EBLUP under Fey Herriot model, is employed to reveal their number in regencies of Java Island. Statistics have been disaggregated by geographical location (urban/rural) and gender. These statistics are required by the government as the basis for policy making.


Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker 2020 University of Arkansas, Fayetteville

Learning Networks With Categorical Data Using Distance Correlation, And A Novel Graph-Based Multivariate Test, Jian Tinker

Graduate Theses and Dissertations

We study the use of distance correlation for statistical inference on categorical data, especially the induction of probability networks. Szekely et al. first defined distance correlation for continuous variables in [42], and Zhang translated the concept into the categorical setting in [57] by defining dCor(X,Y) for categorical variables X = (x1,...,xI) and Y = (y1,...,yJ) where P(X=xi)=[pi]i and P(Y=yi)=[pi]j with the formula [Please open the document]

Part I of the dissertation covers the background we need to understand this formula, and prepares us to analyze the properties and performance of its applications.

Part II then presents the main results of …


Effect Of Predictor Dependence On Variable Selection For Linear And Log-Linear Regression, Apu Chandra Das 2020 University of Arkansas, Fayetteville

Effect Of Predictor Dependence On Variable Selection For Linear And Log-Linear Regression, Apu Chandra Das

Graduate Theses and Dissertations

We propose a Bayesian approach to the Dirichlet-Multinomial (DM) regression model, which uses horseshoe, Laplace, and horseshoe plus priors for shrinkage and selection. The Dirichlet-Multinomial model can be used to find the significant association between a set of available covariates and taxa for a microbiome sample. We incorporate the covariates in a log-linear regression framework. We design a simulation study to make a comparison among the performance of the three shrinkage priors in terms of estimation accuracy and the ability to detect true signals. Our results have clearly separated the performance of the three priors and indicated that the horseshoe …


Combining Machine Learning And Empirical Engineering Methods Towards Improving Oil Production Forecasting, Andrew J. Allen 2020 California Polytechnic State University, San Luis Obispo

Combining Machine Learning And Empirical Engineering Methods Towards Improving Oil Production Forecasting, Andrew J. Allen

Master's Theses

Current methods of production forecasting such as decline curve analysis (DCA) or numerical simulation require years of historical production data, and their accuracy is limited by the choice of model parameters. Unconventional resources have proven challenging to apply traditional methods of production forecasting because they lack long production histories and have extremely variable model parameters. This research proposes a data-driven alternative to reservoir simulation and production forecasting techniques. We create a proxy-well model for predicting cumulative oil production by selecting statistically significant well completion parameters and reservoir information as independent predictor variables in regression-based models. Then, principal component analysis (PCA) …


Joint Models Of Longitudinal Outcomes And Informative Time, JangDong Seo 2020 Indiana University, Bloomington

Joint Models Of Longitudinal Outcomes And Informative Time, Jangdong Seo

Journal of Modern Applied Statistical Methods

Longitudinal data analyses commonly assume that time intervals are predetermined and have no information regarding the outcomes. However, there might be irregular time intervals and informative time. Presented are joint models and asymptotic behaviors of the parameter estimates. Also, the models are applied for real data sets.


Comparison Of Scale Identification Methods In Mixture Irt Models, Youn-Jeng Choi, Allan S. Cohen 2020 University of Alabama

Comparison Of Scale Identification Methods In Mixture Irt Models, Youn-Jeng Choi, Allan S. Cohen

Journal of Modern Applied Statistical Methods

The effects of three scale identification constraints in mixture IRT models were studied. A simulation study found no constraint effect on the mixture Rasch and mixture 2PL models, but the item anchoring constraint was the only one that worked well on selecting correct model with the mixture 3PL model.


Comparing Means Under Heteroscedasticity And Nonnormality: Further Exploring Robust Means Modeling, Alyssa Counsell, Robert Philip Chalmers, Robert A. Cribbie 2020 York University, Toronto

Comparing Means Under Heteroscedasticity And Nonnormality: Further Exploring Robust Means Modeling, Alyssa Counsell, Robert Philip Chalmers, Robert A. Cribbie

Journal of Modern Applied Statistical Methods

Comparing the means of independent groups is a concern when the assumptions of normality and variance homogeneity are violated. Robust means modeling (RMM) was proposed as an alternative to ANOVA-type procedures when the assumptions of normality and variance homogeneity are violated. The purpose of this study is to compare the Type I error and power rates of RMM to the trimmed Welch procedure. A Monte Carlo study was used to investigate RMM and the trimmed Welch procedure under several conditions of nonnormality and variance heterogeneity. The results suggest that the trimmed Welch provides a better balance of Type I error …


A Note On Inferences About The Probability Of Success, Rand Wilcox 2020 University of Southern California

A Note On Inferences About The Probability Of Success, Rand Wilcox

Journal of Modern Applied Statistical Methods

There is an extensive literature dealing with inferences about the probability of success. A minor goal in this note is to point out when certain recommended methods can be unsatisfactory when the sample size is small. The main goal is to report results on the two-sample case. Extant results suggest using one of four methods. The results indicate when computing a 0.95 confidence interval, two of these methods can be more satisfactory when dealing with small sample sizes.


Inferences About The Probability Of Success, Given The Value Of A Covariate, Using A Nonparametric Smoother, Rand Wilcox 2020 University of Southern California

Inferences About The Probability Of Success, Given The Value Of A Covariate, Using A Nonparametric Smoother, Rand Wilcox

Journal of Modern Applied Statistical Methods

For a binary random variable Y, let p(x) = P(Y = 1 | X = x) for some covariate X. The goal of computing a confidence interval for p(x) is considered. In the logistic regression model, even a slight departure difficult to detect via a goodness-of-fit test can yield inaccurate results. The accuracy of a confidence interval can deteriorate as the sample size increases. The goal is to suggest an alternative approach based on a smoother, which provides a more flexible approximation of p(x).


Dividend Maximization Under A Set Ruin Probability Target In The Presence Of Proportional And Excess-Of-Loss Reinsurance, Christian Kasumo, Juma Kasozi, Dmitry Kuznetsov 2020 Nelson Mandela African Institution of Science and Technology

Dividend Maximization Under A Set Ruin Probability Target In The Presence Of Proportional And Excess-Of-Loss Reinsurance, Christian Kasumo, Juma Kasozi, Dmitry Kuznetsov

Applications and Applied Mathematics: An International Journal (AAM)

We study dividend maximization with set ruin probability targets for an insurance company whose surplus is modelled by a diffusion perturbed classical risk process. The company is permitted to enter into proportional or excess-of-loss reinsurance arrangements. By applying stochastic control theory, we derive Volterra integral equations and solve numerically using block-by-block methods. In each of the models, we have established the optimal barrier to use for paying dividends provided the ruin probability does not exceed a predetermined target. Numerical examples involving the use of both light- and heavy-tailed distributions are given. The results show that ruin probability targets result in …


Digital Commons powered by bepress