Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,812 Full-Text Articles 23,893 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,812 full-text articles. Page 324 of 486.

A New Test For Correlation On Bivariate Nonnormal Distributions, Ping Wang, Ping Sa 2016 Great Basin College

A New Test For Correlation On Bivariate Nonnormal Distributions, Ping Wang, Ping Sa

Journal of Modern Applied Statistical Methods

A new method to conduct a right-tailed test for the correlation on bivariate non-normal distribution is proposed. The comparative simulation study shows that the new test controls the type I error rates well for all the distributions considered. An investigation of the power performance is also provided.


Jmasm42: An Alternative Algorithm And Programming Implementation For Least Absolute Deviation Estimator Of The Linear Regression Models (R), Suraju Olaniyi Ogundele, J. I. Mbegbu, C. R. Nwosu 2016 Federal University of Petroleum Resources, Effurun, Delta State, Nigeria.

Jmasm42: An Alternative Algorithm And Programming Implementation For Least Absolute Deviation Estimator Of The Linear Regression Models (R), Suraju Olaniyi Ogundele, J. I. Mbegbu, C. R. Nwosu

Journal of Modern Applied Statistical Methods

We propose a least absolute deviation estimation method that produced a least absolute deviation estimator of parameter of the linear regression model. The method is as accurate as existing method.


Stationary Points For Parametric Stochastic Frontier Models, William C. Horrace, Ian A. Wright 2016 Syracuse University

Stationary Points For Parametric Stochastic Frontier Models, William C. Horrace, Ian A. Wright

Center for Policy Research

The results of Waldman (1982) on the Normal-Half Normal stochastic frontier model are generalized using the theory of the Dirac delta (Dirac, 1930), and distribution-free conditions are established to ensure a stationary point in the likelihood as the variance of the inefficiency distribution goes to zero. Stability of the stationary point and "wrong skew" results are derived or simulated for common parametric assumptions on the model. Identification is discussed.


Evaluating The Efficiency Of Treatment Comparison In Crossover Design By Allocating Subjects Based On Ranked Auxiliary Variable, Yisong Huang, Hani Samawi, Robert Vogel, Jingjing Yin, Worlanyo E. Gato, Daniel Linder 2016 Georgia Southern University

Evaluating The Efficiency Of Treatment Comparison In Crossover Design By Allocating Subjects Based On Ranked Auxiliary Variable, Yisong Huang, Hani Samawi, Robert Vogel, Jingjing Yin, Worlanyo E. Gato, Daniel Linder

Biostatistics: Faculty Publications

The validity of statistical inference depends on proper randomization methods. However, even with proper randomization, we can have imbalanced with respect to important characteristics. In this paper, we introduce a method based on ranked auxiliary variables for treatment allocation in crossover designs using Latin squares models. We evaluate the improvement of the efficiency in treatment comparisons using the proposed method. Our simulation study reveals that our proposed method provides a more powerful test compared to simple randomization with the same sample size. The proposed method is illustrated by conducting an experiment to compare two different concentrations of titanium dioxide nanofiber …


A Browser-Based Ide For The Muzecs Platform, Omokolade Hunpatin, Casey O'Hare, Ryan Thomas, Dennis Brylow 2016 Marquette University

A Browser-Based Ide For The Muzecs Platform, Omokolade Hunpatin, Casey O'Hare, Ryan Thomas, Dennis Brylow

Mathematics, Statistics and Computer Science Faculty Research and Publications

We report on a scalable, portable, and secure visual development environment for programming embedded Arduino platforms with Chromebooks in a successful secondary school computer science curriculum. Our web-based environment is part of the larger MUzECS project, an inexpensive replacement module for the Exploring Computer Science (ECS) course being widely deployed in United States high schools. Students use MUzECS to gain a deeper understanding of computing, through a set of blocks which provide appropriate abstractions for working with low-level hardware.

MUzECS improves upon the existing curriculum module by reducing the hardware cost by an order of magnitude, while still preserving the …


On Characterizations And Infinite Divisibility Of Recently Introduced Distributions, Gholamhossein G. Hamedani 2016 Marquette University

On Characterizations And Infinite Divisibility Of Recently Introduced Distributions, Gholamhossein G. Hamedani

Mathematics, Statistics and Computer Science Faculty Research and Publications

We present here characterizations of the most recently introduced continuous univariate distributions based on: (i) a simple relationship between two truncated moments; (ii) truncated moments of certain functions of the 1th order statistic; (iii) truncated moments of certain functions of the nth order statistic; (iv) truncated moment of certain function of the random variable. We like to mention that the characterization (i) which is expressed in terms of the ratio of truncated moments is stable in the sense of weak convergence. We will also point out that some …


New Classes Of Univariate Continuous Exponential Power Series Distributions, M. Ahsanullah, Gholamhossein G. Hamedani, M. Shakil, B.M. Golam Kibria, F. George 2016 Rider University

New Classes Of Univariate Continuous Exponential Power Series Distributions, M. Ahsanullah, Gholamhossein G. Hamedani, M. Shakil, B.M. Golam Kibria, F. George

Mathematics, Statistics and Computer Science Faculty Research and Publications

Recently, many researchers have developed various classes of continuous probability distributions which can be generated via the generalized Pearson differential equation and other techniques. In this paper, motivated by the importance of the power series in probability theory and its applications, we derive some new classes of univariate exponential power series distributions for a realvalued continuous random variable, which we call exponential power series distributions. Various mathematical properties of the proposed classes of distributions are discussed. Based on these distributional properties, we have established some characterizations of these distributions as well. It is hoped that the findings of the paper …


Hidden Markov Chain Analysis: Impact Of Misclassification On Effect Of Covariates In Disease Progression And Regression, Haritha Polisetti 2016 University of South Florida

Hidden Markov Chain Analysis: Impact Of Misclassification On Effect Of Covariates In Disease Progression And Regression, Haritha Polisetti

USF Tampa Graduate Theses and Dissertations

Most of the chronic diseases have a well-known natural staging system through which the disease progression is interpreted. It is well established that the transition rates from one stage of disease to other stage can be modeled by multi state Markov models. But, it is also well known that the screening systems used to diagnose disease states may subject to error some times. In this study, a simulation study is conducted to illustrate the importance of addressing for misclassification in multi-state Markov models by evaluating and comparing the estimates for the disease progression Markov model with misclassification opposed to disease …


Censoring Unbiased Regression Trees And Ensembles, Jon Arni Steingrimsson, Liqun Diao, Robert L. Strawderman 2016 Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health

Censoring Unbiased Regression Trees And Ensembles, Jon Arni Steingrimsson, Liqun Diao, Robert L. Strawderman

Johns Hopkins University, Dept. of Biostatistics Working Papers

This paper proposes a novel approach to building regression trees and ensemble learning in survival analysis. By first extending the theory of censoring unbiased transformations, we construct observed data estimators of full data loss functions in cases where responses can be right censored. This theory is used to construct two specific classes of methods for building regression trees and regression ensembles that respectively make use of Buckley-James and doubly robust estimating equations for a given full data risk function. For the particular case of squared error loss, we further show how to implement these algorithms using existing software (e.g., CART, …


Prevalence Of And Risk Factors For Adolescent Obesity In Tennessee Using The 2010 Youth Risk Behavior Survey (Yrbs) Data: An Analysis Using Weighted Hierarchical Logistic Regression, Shimin Zheng, Nicole Holt, Jodi L. Southerland, Yan Cao, Trevor Taylor, Deborah L. Slawson, Mark Bloodworth 2016 East Tennessee State University

Prevalence Of And Risk Factors For Adolescent Obesity In Tennessee Using The 2010 Youth Risk Behavior Survey (Yrbs) Data: An Analysis Using Weighted Hierarchical Logistic Regression, Shimin Zheng, Nicole Holt, Jodi L. Southerland, Yan Cao, Trevor Taylor, Deborah L. Slawson, Mark Bloodworth

ETSU Faculty Works

Background: The rate of adolescent overweight and obesity has more than quadrupled over the past few decades, and has become a major public health problem [1]. In 2011, 55% of 12-19 year olds in the United States (U.S.) were overweight or obese [2]. Adolescence is a pivotal time in which many health risk behaviors such as tobacco, alcohol, and drug use are initiated. Such health risk behaviors have been significantly associated with overweight and obesity among adolescents.

Objective: The purpose of this study is to evaluate the relationship between obesity and the health risk behaviors most commonly associated with premature …


High-Throughput Allele-Specific Expression Across 250 Environmental Conditions, Gregory A. Moyerbrailean, Allison L. Richards, Daniel Kurtz, Cynthia A. Kalita, Gordon O. Davis, Chris T. Harvey, Adnan Alazizi, Donovan Watza, Yoram Sorokin, Nancy J. Hauff, Xiang Zhou, Xiaoquan Wen, Roger Pique-Regi, Francesca Luca 2016 Wayne State Center for Molecular Medicine and Genetics, Wayne State University

High-Throughput Allele-Specific Expression Across 250 Environmental Conditions, Gregory A. Moyerbrailean, Allison L. Richards, Daniel Kurtz, Cynthia A. Kalita, Gordon O. Davis, Chris T. Harvey, Adnan Alazizi, Donovan Watza, Yoram Sorokin, Nancy J. Hauff, Xiang Zhou, Xiaoquan Wen, Roger Pique-Regi, Francesca Luca

Center for Molecular Medicine and Genetics

Gene-by-environment (GxE) interactions determine common disease risk factors and biomedically relevant complex traits. However, quantifying how the environment modulates genetic effects on human quantitative phenotypes presents unique challenges. Environmental covariates are complex and difficult to measure and control at the organismal level, as found in GWAS and epidemiological studies. An alternative approach focuses on the cellular environment using in vitro treatments as a proxy for the organismal environment. These cellular environments simplify the organism-level environmental exposures to provide a tractable influence on subcellular phenotypes, such as gene expression. Expression quantitative trait loci (eQTL) mapping studies identified GxE interactions in response …


Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. van der Laan 2016 Division of Biostatistics, University of California, Berkeley

Online Cross-Validation-Based Ensemble Learning, David Benkeser, Samuel D. Lendle, Cheng Ju, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Online estimators update a current estimate with a new incoming batch of data without having to revisit past data thereby providing streaming estimates that are scalable to big data. We develop flexible, ensemble-based online estimators of an infinite-dimensional target parameter, such as a regression function, in the setting where data are generated sequentially by a common conditional data distribution given summary measures of the past. This setting encompasses a wide range of time-series models and as special case, models for independent and identically distributed data. Our estimator considers a large library of candidate online estimators and uses online cross-validation to …


Doubly-Robust Nonparametric Inference On The Average Treatment Effect, David Benkeser, Marco Carone, Mark J. van der Laan, Peter Gilbert 2016 Division of Biostatistics, University of California, Berkeley

Doubly-Robust Nonparametric Inference On The Average Treatment Effect, David Benkeser, Marco Carone, Mark J. Van Der Laan, Peter Gilbert

U.C. Berkeley Division of Biostatistics Working Paper Series

Doubly-robust estimators are widely used to draw inference about the average effect of a treatment. Such estimators are consistent for the effect of interest if either one of two nuisance parameters is consistently estimated. However, if flexible, data-adaptive estimators of these nuisance parameters are used, double-robustness does not readily extend to inference. We present a general theoretical study of the behavior of doubly-robust estimators of an average treatment effect when one of the nuisance parameters is inconsistently estimated. We contrast different approaches for constructing such estimators and investigate the extent to which they may be modified to also allow doubly-robust …


On Combining Family- And Population- Based Sequencing Data, Yuriko Katsumata, David W. Fardo 2016 University of Kentucky

On Combining Family- And Population- Based Sequencing Data, Yuriko Katsumata, David W. Fardo

Biostatistics Faculty Publications

Several statistical group-based approaches have been proposed to detect effects of variation within a gene for each of the population- and family-based designs. However, unified tests to combine gene-phenotype associations obtained from these 2 study designs are not yet well established. In this study, we investigated the efficient combination of population-based and family-based sequencing data to evaluate best practices using the Genetic Analysis Workshop 19 (GAW19) data set. Because one design employed whole genome sequencing and the other whole exome sequencing, we examined variants overlapping both data sets. We used the family-based sequence kernel association test (famSKAT) to analyze the …


Causal Effect Estimation In Sequencing Studies: A Bayesian Method To Account For Confounder Adjustment Uncertainty, Chi Wang, Jinpeng Liu, David W. Fardo 2016 University of Kentucky

Causal Effect Estimation In Sequencing Studies: A Bayesian Method To Account For Confounder Adjustment Uncertainty, Chi Wang, Jinpeng Liu, David W. Fardo

Biostatistics Faculty Publications

Estimating the causal effect of a single nucleotide variant (SNV) on clinical phenotypes is of interest in many genetic studies. The effect estimation may be confounded by other SNVs as a result of linkage disequilibrium as well as demographic and clinical characteristics. Because a large number of these other variables, which we call potential confounders, are collected, it is challenging to select and adjust for the variables that truly confound the causal effect. The Bayesian adjustment for confounding (BAC) method has been proposed as a general method to estimate the average causal effect in the presence of a large number …


Comparing Performance Of Non-Tree-Based And Tree-Based Association Mapping Methods, Katherine L. Thompson, David W. Fardo 2016 University of Kentucky

Comparing Performance Of Non-Tree-Based And Tree-Based Association Mapping Methods, Katherine L. Thompson, David W. Fardo

Statistics Faculty Publications

A central goal in the biomedical and biological sciences is to link variation in quantitative traits to locations along the genome (single nucleotide polymorphisms). Sequencing technology has rapidly advanced in recent decades, along with the statistical methodology to analyze genetic data. Two classes of association mapping methods exist: those that account for the evolutionary relatedness among individuals, and those that ignore the evolutionary relationships among individuals. While the former methods more fully use implicit information in the data, the latter methods are more flexible in the types of data they can handle. This study presents a comparison of the 2 …


Human Exposure Modeling Using Sheds, Luther Smith, William Graham Glen 2016 Alion Science & Technology Inc

Human Exposure Modeling Using Sheds, Luther Smith, William Graham Glen

Annual Symposium on Biomathematics and Ecology Education and Research

No abstract provided.


Nondestructive Testing And Structural Health Monitoring Based On Adams And Svm Techniques, Gang Jiang, Yi Ming Deng, Ji Tai Niu 2016 Southwest University of Science and Technology

Nondestructive Testing And Structural Health Monitoring Based On Adams And Svm Techniques, Gang Jiang, Yi Ming Deng, Ji Tai Niu

The 8th International Conference on Physical and Numerical Simulation of Materials Processing

No abstract provided.


Performance-Constrained Binary Classification Using Ensemble Learning: An Application To Cost-Efficient Targeted Prep Strategies, Wenjing Zheng, Laura Balzer, Maya L. Petersen, Mark J. van der Laan 2016 Division of Biostatistics, School of Public Health, University of California, Berkeley

Performance-Constrained Binary Classification Using Ensemble Learning: An Application To Cost-Efficient Targeted Prep Strategies, Wenjing Zheng, Laura Balzer, Maya L. Petersen, Mark J. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Binary classifications problems are ubiquitous in health and social science applications. In many cases, one wishes to balance two conflicting criteria for an optimal binary classifier. For instance, in resource-limited settings, an HIV prevention program based on offering Pre-Exposure Prophylaxis (PrEP) to select high-risk individuals must balance the sensitivity of the binary classifier in detecting future seroconverters (and hence offering them PrEP regimens) with the total number of PrEP regimens that is financially and logistically feasible for the program to deliver. In this article, we consider a general class of performance-constrained binary classification problems wherein the objective function and the …


Matching The Efficiency Gains Of The Logistic Regression Estimator While Avoiding Its Interpretability Problems, In Randomized Trials, Michael Rosenblum, Jon Arni Steingrimsson 2016 Johns Hopkins Bloomberg School of Public Health, Department of Biostatistics

Matching The Efficiency Gains Of The Logistic Regression Estimator While Avoiding Its Interpretability Problems, In Randomized Trials, Michael Rosenblum, Jon Arni Steingrimsson

Johns Hopkins University, Dept. of Biostatistics Working Papers

Adjusting for prognostic baseline variables can lead to improved power in randomized trials. For binary outcomes, a logistic regression estimator is commonly used for such adjustment. This has resulted in substantial efficiency gains in practice, e.g., gains equivalent to reducing the required sample size by 20-28% were observed in a recent survey of traumatic brain injury trials. Robinson and Jewell (1991) proved that the logistic regression estimator is guaranteed to have equal or better asymptotic efficiency compared to the unadjusted estimator (which ignores baseline variables). Unfortunately, the logistic regression estimator has the following dangerous vulnerabilities: it is only interpretable when …


Digital Commons powered by bepress