Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,820 Full-Text Articles 23,917 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,820 full-text articles. Page 207 of 487.

On The Noisy Gradient Descent That Generalizes As Sgd, Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, Zhanxing Zhu 2020 Missouri University of Science and Technology

On The Noisy Gradient Descent That Generalizes As Sgd, Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, Zhanxing Zhu

Mathematics and Statistics Faculty Research & Creative Works

The gradient noise of SGD is considered to play a central role in the observed strong generalization abilities of deep learning. While past studies confirm that the magnitude and covariance structure of gradient noise are critical for regularization, it remains unclear whether or not the class of noise distributions is important. In this work we provide negative results by showing that noises in classes different from the SGD noise can also effectively regularize gradient descent. Our finding is based on a novel observation on the structure of the SGD noise: it is the multiplication of the gradient matrix and a …


A Generalized Family Of Lifetime Distributions And Survival Models, Mahmoud Aldeni, Felix Famoye, Carl Lee 2020 Western Carolina University

A Generalized Family Of Lifetime Distributions And Survival Models, Mahmoud Aldeni, Felix Famoye, Carl Lee

Journal of Modern Applied Statistical Methods

In lifetime data, the hazard function is a common technique for describing the characteristics of lifetime distribution. Monotone increasing or decreasing, and unimodal are relatively simple hazard function shapes, which can be modeled by many parametric lifetime distributions. However, fewer distributions are capable of modeling diverse and more complicated shapes such as N-shaped, reflected N-shaped, W-shaped, and M-shaped hazard rate functions. A generalized family of lifetime distributions, the uniform-R{generalized lambda} (U-R{GL}) are introduced and the corresponding survival models are derived, and applied to two lifetime data sets. The survival model is applied to a right censored lifetime data set.


A Revised Logic Model For Educational Program Evaluation, Zsa-Zsa Booker 2020 Wayne State University School of Medicine

A Revised Logic Model For Educational Program Evaluation, Zsa-Zsa Booker

Journal of Modern Applied Statistical Methods

The logic model is an evaluation tool popularly used for obtaining grant funding. Its limitations make it unlike other theory driven evaluation methods. A critical examination of the logic model leads to the construction of an enriched revised logic model.


Forward And Backward Continuation Ratio Models For Ordinal Response Variables, Xing Liu, Haiyan Bai 2020 Eastern Connecticut State University

Forward And Backward Continuation Ratio Models For Ordinal Response Variables, Xing Liu, Haiyan Bai

Journal of Modern Applied Statistical Methods

There are different types of continuation ratio (CR) models for ordinal response variables. The different model equations, corresponding parameterizations, and nonequivalent results are confusing. The purpose of this study is to introduce different types of forward and backward CR models, demonstrate how to implement these models using Stata, and compare the results using data from the Educational Longitudinal Study of 2002 (ELS:2002).


Identifying Which Of J Independent Binomial Distributions Has The Largest Probability Of Success, Rand Wilcox 2020 University of Southern California

Identifying Which Of J Independent Binomial Distributions Has The Largest Probability Of Success, Rand Wilcox

Journal of Modern Applied Statistical Methods

Let p1,…, pJ denote the probability of a success for J independent random variables having a binomial distribution and let p(1) ≤ … ≤ p(J) denote these probabilities written in ascending order. The goal is to make a decision about which group has the largest probability of a success, p(J). Let p̂1,…, p̂J denote estimates of p1,…,pJ, respectively. The strategy is to test J − 1 hypotheses comparing the group with the largest estimate to each of the J − 1 …


Jmasm 53: Miccerird, Michael Lance 2020 ExceLance, LLC

Jmasm 53: Miccerird, Michael Lance

Journal of Modern Applied Statistical Methods

Fortran 77 and 90 modules (REALPOPS.lib) exist for invoking the 8 distributions estimated by Micceri (1989). These respective modules were created by Sawilowsky et al. (1990) and Sawilowsky and Fahoome (2003). The MicceriRD (Micceri’s Real Distributions) Python package was created because Python is increasingly used for data analysis and, in some cases, Monte Carlo simulations.


Bayesian Analysis Of Extended Cox Model With Time-Varying Covariates Using Bootstrap Prior, Oyebayo R. Olaniran, Mohd Asrul A. Abdullah 2020 University of Ilorin

Bayesian Analysis Of Extended Cox Model With Time-Varying Covariates Using Bootstrap Prior, Oyebayo R. Olaniran, Mohd Asrul A. Abdullah

Journal of Modern Applied Statistical Methods

A new Bayesian estimation procedure for extended cox model with time varying covariate was presented. The prior was determined using bootstrapping technique within the framework of parametric empirical Bayes. The efficiency of the proposed method was observed using Monte Carlo simulation of extended Cox model with time varying covariates under varying scenarios. Validity of the proposed method was also ascertained using real life data set of Stanford heart transplant. Comparison of the proposed method with its competitor established appreciable supremacy of the method.


Regression: Determining Which Of P Independent Variables Has The Largest Or Smallest Correlation With The Dependent Variable, Plus Results On Ordering The Correlations Winsorized, Rand Wilcox 2020 University of Southern California

Regression: Determining Which Of P Independent Variables Has The Largest Or Smallest Correlation With The Dependent Variable, Plus Results On Ordering The Correlations Winsorized, Rand Wilcox

Journal of Modern Applied Statistical Methods

In a regression context, consider p independent variables and a single dependent variable. The paper addresses two goals. The first is to determine the extent it is reasonable to make a decision about whether the largest estimate of the Winsorized correlations corresponds to the independent variable that has the largest population Winsorized correlation. The second is to determine the extent it is reasonable to decide that the order of the estimates of the Winsorized correlations correctly reflects the true ordering. Both goals are addressed by testing relevant hypotheses. Results in Wilcox (in press a) suggest using a multiple comparisons procedure …


Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris 2020 University of Crete, Greece

Jmasm 52: Extremely Efficient Permutation And Bootstrap Hypothesis Tests Using R, Christina Chatzipantsiou, Marios Dimitriadis, Manos Papadakis, Michail Tsagris

Journal of Modern Applied Statistical Methods

Re-sampling based statistical tests are known to be computationally heavy, but reliable when small sample sizes are available. Despite their nice theoretical properties not much effort has been put to make them efficient. Computationally efficient method for calculating permutation-based p-values for the Pearson correlation coefficient and two independent samples t-test are proposed. The method is general and can be applied to other similar two sample mean or two mean vectors cases.


Empirical Comparison Of Tests For One-Factor Anova Under Heterogeneity And Non-Normality: A Monte Carlo Study, Diep Nguyen, Eunsook Kim, Yan Wang, Thanh Vinh Pham, Yi-Hsin Chen, Jeffrey D. Kromrey 2020 University of South Florida

Empirical Comparison Of Tests For One-Factor Anova Under Heterogeneity And Non-Normality: A Monte Carlo Study, Diep Nguyen, Eunsook Kim, Yan Wang, Thanh Vinh Pham, Yi-Hsin Chen, Jeffrey D. Kromrey

Journal of Modern Applied Statistical Methods

Although the Analysis of Variance (ANOVA) F test is one of the most popular statistical tools to compare group means, it is sensitive to violations of the homogeneity of variance (HOV) assumption. This simulation study examines the performance of thirteen tests in one-factor ANOVA models in terms of their Type I error rate and statistical power under numerous (82,080) conditions. The results show that when HOV was satisfied, the ANOVA F or the Brown-Forsythe test outperformed the other methods in terms of both Type I error control and statistical power even under non-normality. When HOV was violated, the Structured Means …


Methods Of Uncertainty Quantification For Physical Parameters, Kellin Rumsey 2020 University of New Mexico

Methods Of Uncertainty Quantification For Physical Parameters, Kellin Rumsey

Mathematics & Statistics ETDs

Uncertainty Quantification (UQ) is an umbrella term referring to a broad class of methods which typically involve the combination of computational modeling, experimental data and expert knowledge to study a physical system. A parameter, in the usual statistical sense, is said to be physical if it has a meaningful interpretation with respect to the physical system. Physical parameters can be viewed as inherent properties of a physical process and have a corresponding true value. Statistical inference for physical parameters is a challenging problem in UQ due to the inadequacy of the computer model. In this thesis, we provide a comprehensive …


Maximum Likelihood Estimation Of Species Trees And Anomaly Zone Detection Using Ranked Gene Trees, Anastasiia Kim 2020 University of New Mexico

Maximum Likelihood Estimation Of Species Trees And Anomaly Zone Detection Using Ranked Gene Trees, Anastasiia Kim

Mathematics & Statistics ETDs

A phylogenetic tree represents the evolutionary relationships among a set of organisms. Gene trees can be used to reconstruct phylogenetic trees. The methods in this dissertation focus on the gene tree topologies with emphasis on ranked gene tree topologies. A ranked tree depicts the order in which nodes appear in the tree together with topological relationships among gene lineages. One challenge that arises during phylogenetic inference is the existence of the anomaly zones, the regions of branch-length space in the species tree that can produce gene trees that have topologies differing from the species tree topology but are more probable …


Assessing The Validity Of Sentiment Analysis Measures Through Polychoric Correlation, Kelli N. Kasper 2020 University of New Mexico

Assessing The Validity Of Sentiment Analysis Measures Through Polychoric Correlation, Kelli N. Kasper

Mathematics & Statistics ETDs

Sentiment analysis methods extract the attitude of a text via systematic algorithms. To evaluate the validity of common sentiment analysis methods, we use polychoric correlation to compare computer-mediated methods and human-rated analogues. Our main topics of interest are the internal consistency of the raters' scores, the level of consensus among raters, and how well raters' scores correlate with those given by sentiment analysis methods for randomly collected Twitter data.

Our analysis found that there is good validity for methods that measure negative and positive sentiments in short texts, both in terms of inter-rater consistency and when comparing raters to computer-mediated …


An Investigation Of Gene Regulatory Network State Space Variability, Sara Faye Liesman 2020 Illinois State University

An Investigation Of Gene Regulatory Network State Space Variability, Sara Faye Liesman

Theses and Dissertations

Genes are segments of DNA that provide a blueprint for cells and organisms to effectively control processes and regulations within individuals. There have been many attempts to quantify these processes, as a greater understanding of how genes operate could have large impacts on both personalized and precision medicine. Gene interactions are of particular interest, however, current biological methods can not easily reveal the details of these interactions. Therefore, we infer networks of interactions from gene expression data which we call a gene regulatory network, or GRN. Due to the robust behavior of genes and the inherent variability within interactions, models …


Epidemiology Of Cancers In Men Who Have Sex With Men (Msm): A Protocol For Umbrella Review Of Systematic Reviews, Manoj Kumar Honaryar, Yelena Tarasenko, Maribel Almonte, Vitaly Smelov 2020 Prevention and Implementation Group (PRI), International Agency for Research on Cancer (IARC), World Health Organization (WHO)

Epidemiology Of Cancers In Men Who Have Sex With Men (Msm): A Protocol For Umbrella Review Of Systematic Reviews, Manoj Kumar Honaryar, Yelena Tarasenko, Maribel Almonte, Vitaly Smelov

Biostatistics, Epidemiology & Environmental Health Sciences: Faculty Publications

While earlier studies on men having sex with men (MSM) tended to examine infection-related cancers, an increasing number of studies have been focusing on effects of sexual orientation on other cancers and social and cultural causes for cancer disparities. As a type of tertiary research, this umbrella review (UR) aims to synthesize findings from existing review studies on the effects of sexual orientation on cancer. Relevant peer-reviewed systematic reviews (SRs) will be identified without date or language restrictions using MEDLINE, Cochrane Database of Systematic Reviews, and the International Prospective Register for Systematic Reviews, among others. The research team members will …


An Improved Method For Spectroscopic Quality Classification, Elizabeth G. Mayer 2020 University of New Mexico

An Improved Method For Spectroscopic Quality Classification, Elizabeth G. Mayer

Mathematics & Statistics ETDs

Spectral quality classification is a vital step in data cleaning before the

analysis of magnetic resonance spectroscopy (MRS) data can be done. This

analysis compares five methods of quality classification; three of these are

legacy methods, Maudsley et al. (2006), Zhang et al. (2018), and

Bustillo et al. (2020), and two newly created methods that used a random forests

classifier (RFC) to inform their classifications. We found that the random forest

classifier was the most accurate at predicting spectra quality (balanced

accuracy for RF of 88% vs legacy of 70%, 72%, or 72%). A

Random-Forests-Informed Filtering method (RFIFM) for quality …


A Study Of The Efficacy Of Machine Learning For Diagnosing Obstructive Coronary Artery Disease In Non-Diabetic Patients, Demond Larae Handley 2020 Illinois State University

A Study Of The Efficacy Of Machine Learning For Diagnosing Obstructive Coronary Artery Disease In Non-Diabetic Patients, Demond Larae Handley

Theses and Dissertations

According to the Centers for Disease Control and Prevention, about 18.2 million adults age 20 and older have Coronary Artery Disease in the United States. Early diagnosis is therefore of crucial importance to help prevent debilitating consequences, and principally death for many patients. In this study we use data containing gene expression values from peripheral blood samples in 198 non-diabetic patients, with the goal of developing an age and sex gene expression model for diagnosis of Coronary Artery Disease. We employ machine learning methods to obtain a classification based on genetic information, age and sex. Our implementation uses feed forward …


Improving The Quality And Design Of Retrospective Clinical Outcome Studies That Utilize Electronic Health Records, Oliwier Dziadkowiec, Jeffery Durbin, Vignesh Jayaraman Muralidharan, Megan Novak, Brendon Cornett 2020 HCA Healthcare Mountain MidAmerica and Continental Divisions

Improving The Quality And Design Of Retrospective Clinical Outcome Studies That Utilize Electronic Health Records, Oliwier Dziadkowiec, Jeffery Durbin, Vignesh Jayaraman Muralidharan, Megan Novak, Brendon Cornett

HCA Healthcare Journal of Medicine

Electronic health records (EHRs) are an excellent source for secondary data analysis. Studies based on EHR-derived data, if designed properly, can answer previously unanswerable clinical research questions. In this paper we will highlight the benefits of large retrospective studies from secondary sources such as EHRs, examine retrospective cohort and case-control study design challenges, as well as methodological and statistical adjustment that can be made to overcome some of the inherent design limitations, in order to increase the generalizability, validity and reliability of the results obtained from these studies.


Playfair's Introduction Of Time Series To Represent Data, Diana White, Joshua Eastes, Negar Janani, River Bond 2020 University of Colorado Denver

Playfair's Introduction Of Time Series To Represent Data, Diana White, Joshua Eastes, Negar Janani, River Bond

Statistics and Probability

No abstract provided.


Playfair's Novel Visual Displays Of Data, Diana White, River Bond, Joshua Eastes, Negar Janani 2020 University of Colorado Denver

Playfair's Novel Visual Displays Of Data, Diana White, River Bond, Joshua Eastes, Negar Janani

Statistics and Probability

No abstract provided.


Digital Commons powered by bepress