Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

2020

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 121 - 150 of 169

Full-Text Articles in Applied Statistics

Bivariate Analogs Of The Wilcoxon–Mann–Whitney Test And The Patel–Hoel Method For Interactions, Rand Wilcox Feb 2020

Bivariate Analogs Of The Wilcoxon–Mann–Whitney Test And The Patel–Hoel Method For Interactions, Rand Wilcox

Journal of Modern Applied Statistical Methods

A fundamental way of characterizing how two independent compares compare is in terms of the probability that a randomly sampled observation from the first group is less than a randomly sampled observation from the second group. The paper suggests a bivariate analog and investigates methods for computing confidence intervals. An interaction for a two-by-two design is investigated as well.


Assessing The Accuracy Of Approximate Confidence Intervals Proposed For The Mean Of Poisson Distribution, Alireza Shirvani, Malek Fathizadeh Feb 2020

Assessing The Accuracy Of Approximate Confidence Intervals Proposed For The Mean Of Poisson Distribution, Alireza Shirvani, Malek Fathizadeh

Journal of Modern Applied Statistical Methods

The Poisson distribution is applied as an appropriate standard model to analyze count data. Because this distribution is known as a discrete distribution, representation of accurate confidence intervals for its distribution mean is extremely difficult. Approximate confidence intervals were presented for the Poisson distribution mean. The purpose of this study is to simultaneously compare several confidence intervals presented, according to the average coverage probability and accurate confidence coefficient and the average confidence interval length criteria.


Analytical Closed-Form Solution For General Factor With Many Variables, Stan Lipovetsky, Vladimir Manewitsch Feb 2020

Analytical Closed-Form Solution For General Factor With Many Variables, Stan Lipovetsky, Vladimir Manewitsch

Journal of Modern Applied Statistical Methods

The factor analytic triad method of one-factor solution gives the explicit analytical form for a common latent factor built by three variables. The current work considers analytical presentation of a general latent factor constructed in a closed-form solution for multivariate case. The results can be supportive to theoretical description and practical application of latent variable modeling, especially for big data because the analytical closed-form solution is not prone to data dimensionality.


Regression Modeling And Prediction By Individual Observations Versus Frequency, Stan Lipovetsky Feb 2020

Regression Modeling And Prediction By Individual Observations Versus Frequency, Stan Lipovetsky

Journal of Modern Applied Statistical Methods

A regression model built by a dataset could sometimes demonstrate a low quality of fit and poor predictions of individual observations. However, using the frequencies of possible combinations of the predictors and the outcome, the same models with the same parameters may yield a high quality of fit and precise predictions for the frequencies of the outcome occurrence. Linear and logistical regressions are used to make an explicit exposition of the results of regression modeling and prediction.


Analysis Of An Agent-Based Model For Predicting The Behavior Of Bighead Carp (Hypophthalmichthys Nobilis) Under The Influence Of Acoustic Deterrence, Craig Garzella, Joseph Gaudy, Karl R. B. Schmitt, Arezu Mansuri Feb 2020

Analysis Of An Agent-Based Model For Predicting The Behavior Of Bighead Carp (Hypophthalmichthys Nobilis) Under The Influence Of Acoustic Deterrence, Craig Garzella, Joseph Gaudy, Karl R. B. Schmitt, Arezu Mansuri

Spora: A Journal of Biomathematics

Bighead carp (Hypophthalmichthys nobilis) are an invasive, voracious, highly fecund species threatening the ecological integrity of the Great Lakes. This agent-based model and analysis explore bighead carp behavior in response to acoustic deterrence in an effort to discover properties that increase likelihood of deterrence system failure. Results indicate the most significant (p < 0.05) influences on barrier failure are the quantity of detritus and plankton behind the barrier, total number of bighead carp successfully deterred by the barrier, and number of native fishes freely moving throughout the simulation. Quantity of resources behind the barrier influence bighead carp to penetrate when populations are resource deprived. When native fish populations are low, an accumulation of phytoplankton can occur, increasing the likelihood of an algal bloom occurrence. Findings of this simulation suggest successful implementation with proper maintenance of an acoustic deterrence system has potential of abating the threat of bighead carp on ecological integrity of the Great Lakes.


Session 11 - Methods: Bootstrap Control Chart For Pareto Percentiles, Ruth Burkhalter Feb 2020

Session 11 - Methods: Bootstrap Control Chart For Pareto Percentiles, Ruth Burkhalter

SDSU Data Science Symposium

Lifetime percentile is an important indicator of product reliability. However, the sampling distribution of a percentile estimator for any lifetime distribution is not a bell shaped one. As a result, the well-known Shewhart-type control chart cannot be applied to monitor the product lifetime percentiles. In this presentation, Bootstrap control charts based on maximum likelihood estimator (MLE) are proposed for monitoring Pareto percentiles. An intensive simulation study is conducted to compare the performance among the proposed MLE Bootstrap control chart and Shewhart-type control chart.


An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone Jan 2020

An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone

Published and Grey Literature from PhD Candidates

Data mining techniques have numerous applications in bankcard response modeling. Logistic regression has been used as the standard modeling tool in the financial industry because of its almost always desirable performance and its interpretability. In this paper, we propose a hybrid bankcard response model, which integrates decision tree-based chi-square automatic interaction detection (CHAID) into logistic regression. In the first stage of the hybrid model, CHAID analysis is used to detect the possible potential variable interactions. Then in the second stage, these potential interactions are served as the additional input variables in logistic regression. The motivation of the proposed hybrid model …


A Two-Stage Hybrid Model By Using Artificial Neural Networks As Feature Construction Algorithms, Yan Wang, Sherry Ni, Brian Stone Jan 2020

A Two-Stage Hybrid Model By Using Artificial Neural Networks As Feature Construction Algorithms, Yan Wang, Sherry Ni, Brian Stone

Published and Grey Literature from PhD Candidates

We propose a two-stage hybrid approach with neural networks as the new feature construction algorithms for bankcard response classifications. The hybrid model uses a very simple neural network structure as the new feature construction tool in the first stage, then the newly created features are used as the additional input variables in logistic regression in the second stage. The model is compared with the traditional one-stage model in credit customer response classification. It is observed that the proposed two-stage model outperforms the one-stage model in terms of accuracy, the area under the ROC curve, and KS statistic. By creating new …


Predicting Class-Imbalanced Business Risk Using Resampling, Regularization, And Model Ensembling Algorithms, Yan Wang, Sherry Ni Jan 2020

Predicting Class-Imbalanced Business Risk Using Resampling, Regularization, And Model Ensembling Algorithms, Yan Wang, Sherry Ni

Published and Grey Literature from PhD Candidates

We aim at developing and improving the imbalanced business risk modeling via jointly using proper evaluation criteria, resampling, cross-validation, classifier regularization, and ensembling techniques. Area Under the Receiver Operating Characteristic Curve (AUC of ROC) is used for model comparison based on 10-fold cross-validation. Two undersampling strategies including random undersampling (RUS) and cluster centroid undersampling (CCUS), as well as two oversampling methods including random oversampling (ROS) and Synthetic Minority Oversampling Technique (SMOTE), are applied. Three highly interpretable classifiers, including logistic regression without regularization (LR), L1-regularized LR (L1LR), and decision tree (DT) are implemented. Two ensembling techniques, including Bagging and Boosting, are …


A Xgboost Risk Model Via Feature Selection And Bayesian Hyper-Parameter Optimization, Yan Wang, Sherry Ni Jan 2020

A Xgboost Risk Model Via Feature Selection And Bayesian Hyper-Parameter Optimization, Yan Wang, Sherry Ni

Published and Grey Literature from PhD Candidates

This paper aims to explore models based on the extreme gradient boosting (XGBoost) approach for business risk classification. Feature selection (FS) algorithms and hyper-parameter optimizations are simultaneously considered during model training. The five most commonly used FS methods including weight by Gini, weight by Chi-square, hierarchical variable clustering, weight by correlation, and weight by information are applied to alleviate the effect of redundant features. Two hyper-parameter optimization approaches, random search (RS) and Bayesian tree-structuredParzen Estimator (TPE), are applied in XGBoost. The effect of different FS and hyper-parameter optimization methods on the model performance are investigated by the Wilcoxon Signed Rank …


Pair-A-Dice Lost: Experiments In Dice Control, Robert H. Scott Iii, Donald R. Smith Jan 2020

Pair-A-Dice Lost: Experiments In Dice Control, Robert H. Scott Iii, Donald R. Smith

UNLV Gaming Research & Review Journal

This paper presents our findings from experiments designed to test whether we could use a custom-made dice throwing machine applying common dice control methods to produce dice rolls that differ from random. In earlier research we calculated the percentages of control a craps player needs to break even or beat the house (Smith and Scott, 2018). Using the most common practices of dice control in craps, we established how dice should be configured (i.e., set) and thrown to achieve certain outcomes such as not rolling a seven in the point cycle. We decided to run experiments to see if a …


The Author’S Reflections On No B.S. (Bad Stats): Black People Need People Who Believe In Black People Enough Not To Believe Every Bad Thing They Hear About Black People, Ivory A. Toldson Jan 2020

The Author’S Reflections On No B.S. (Bad Stats): Black People Need People Who Believe In Black People Enough Not To Believe Every Bad Thing They Hear About Black People, Ivory A. Toldson

Numeracy

Toldson, Ivory. A. 2019. No BS (Bad Stats): Black People Need People Who Believe in Black People Enough Not to Believe Every Bad Thing They Hear About Black People (Boston, MA: Brill-Sense) 194 pp. ISBN 978-9004397026.

This essay provides an introduction to No BS (Bad Stats): Black People Need People Who Believe in Black People Enough Not to Believe Every Bad Thing They Hear About Black People. In the essay, the author discusses how cynical views about the educational potential of Black children motivated him to write a book that challenges negative statistics. The essay also outlines the harmful …


Quantitative Model For Setting Manufacturer's Suggested Retail Price, Peter Byrd, Jonathan Knowles, Dmitry Andreev, Jacob Turner, Brian Mente, Laroux Wallace Jan 2020

Quantitative Model For Setting Manufacturer's Suggested Retail Price, Peter Byrd, Jonathan Knowles, Dmitry Andreev, Jacob Turner, Brian Mente, Laroux Wallace

SMU Data Science Review

In this paper, we present a quantitative approach to model the manufacturer’s suggested retail price (MSRP) for children’s doll- houses and establish relationships among key features that contribute most to establishing MSRP. Determination of the MSRP is a critical step in how consumers respond with their wallets when purchasing an item. KidKraft, a global leader in toys and juvenile products, sets MSRP subjectively using product experts. The process is arduous and time consuming requiring the focus of specialized resources and knowledge of the interaction between key attributes and their impact on consumer value. An accurate prediction of MSRP during the …


Power Analysis On A Pilot Study Of The Caloric Intake Of Children Helping Prepare Meals Versus Children Not, Danielle Clifford Jan 2020

Power Analysis On A Pilot Study Of The Caloric Intake Of Children Helping Prepare Meals Versus Children Not, Danielle Clifford

Student Research Poster Presentations 2020

The purpose of this analysis is to determine the sample size needed for a study that will be used to discover if there is a difference in the caloric intake of children who help with meal preparation and children who do not help with meal preparation.


Internship With Alison's Homemade Welsh Cookies, Heather Harvey Jan 2020

Internship With Alison's Homemade Welsh Cookies, Heather Harvey

Student Research Poster Presentations 2020

During the Spring 2020 semester, I have been undergoing an internship with Alison's Homemade Welsh Cookies. They have allowed me full access to their data, in order to perform an analysis. Using the sames records from the 2019 year, each location has been identified on a map and the amounts sold tell which ares had the most and the least sales in 2019.


Predicting Diabetes Diagnoses, Sarah Netchert Jan 2020

Predicting Diabetes Diagnoses, Sarah Netchert

Student Research Poster Presentations 2020

This study explored the traits and health state of African Americans in central Virginia in order to determine what traits put people at a higher probability of being diagnosed with diabetes. We also want to know which traits will generate the highest probability a person will be diagnosed with diabetes. Traits that were included and used in this study were cholesterol, stabilized glucose, high density lipoprotein levels, age(years), gender, height(inches), weight(pounds), systolic blood pressure, diastolic blood pressure, waist size(inches), and hip size(inches). There were 403 individuals included in study since they were only ones screened for diabetes out of 1,046 …


An Analysis Of The Success Of Farmers Markets In Kentucky Using Logistic Regression And Support Vector Machines, Jeron Russell Jan 2020

An Analysis Of The Success Of Farmers Markets In Kentucky Using Logistic Regression And Support Vector Machines, Jeron Russell

Mahurin Honors College Capstone Experience/Thesis Projects

The purpose of this research is to look at the relationship that market-specific, economic, and demographic variables have with the success of farmers markets in Kentucky. It additionally seeks to build a tool for predicting farmers market success that could be used by policy makers to aid in decision-making processes concerning farmers markets. Logistic regression and Support Vector Machines (SVMs) are used on data acquired from the Kentucky Department of Agriculture and the American Community Survey in order to analyze the data in a traditional statistical approach as well as a machine learning approach. The results included an SVM model …


Theory Of Principal Components For Applications In Exploratory Crime Analysis And Clustering, Daniel Silva Jan 2020

Theory Of Principal Components For Applications In Exploratory Crime Analysis And Clustering, Daniel Silva

All Graduate Theses, Dissertations, and Other Capstone Projects

The purpose of this paper is to develop the theory of principal components analysis succinctly from the fundamentals of matrix algebra and multivariate statistics. Principal components analysis is sometimes used as a descriptive technique to explain the variance-covariance or correlation structure of a dataset. However, most often, it is used as a dimensionality reduction technique to visualize a high dimensional dataset in a lower dimensional space. Principal components analysis accomplishes this by using the first few principal components, provided that they account for a substantial proportion of variation in the original dataset. In the same way, the first few principal …


Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown Jan 2020

Evaluating An Ordinal Output Using Data Modeling, Algorithmic Modeling, And Numerical Analysis, Martin Keagan Wynne Brown

Murray State Theses and Dissertations

Data and algorithmic modeling are two different approaches used in predictive analytics. The models discussed from these two approaches include the proportional odds logit model (POLR), the vector generalized linear model (VGLM), the classification and regression tree model (CART), and the random forests model (RF). Patterns in the data were analyzed using trigonometric polynomial approximations and Fast Fourier Transforms. Predictive modeling is used frequently in statistics and data science to find the relationship between the explanatory (input) variables and a response (output) variable. Both approaches prove advantageous in different cases depending on the data set. In our case, the data …


Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero Jan 2020

Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero

Theses and Dissertations

Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …


Unitary And Symmetric Structure In Deep Neural Networks, Kehelwala Dewage Gayan Maduranga Jan 2020

Unitary And Symmetric Structure In Deep Neural Networks, Kehelwala Dewage Gayan Maduranga

Theses and Dissertations--Mathematics

Recurrent neural networks (RNNs) have been successfully used on a wide range of sequential data problems. A well-known difficulty in using RNNs is the vanishing or exploding gradient problem. Recently, there have been several different RNN architectures that try to mitigate this issue by maintaining an orthogonal or unitary recurrent weight matrix. One such architecture is the scaled Cayley orthogonal recurrent neural network (scoRNN), which parameterizes the orthogonal recurrent weight matrix through a scaled Cayley transform. This parametrization contains a diagonal scaling matrix consisting of positive or negative one entries that can not be optimized by gradient descent. Thus the …


Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich Jan 2020

Orthogonal Recurrent Neural Networks And Batch Normalization In Deep Neural Networks, Kyle Eric Helfrich

Theses and Dissertations--Mathematics

Despite the recent success of various machine learning techniques, there are still numerous obstacles that must be overcome. One obstacle is known as the vanishing/exploding gradient problem. This problem refers to gradients that either become zero or unbounded. This is a well known problem that commonly occurs in Recurrent Neural Networks (RNNs). In this work we describe how this problem can be mitigated, establish three different architectures that are designed to avoid this issue, and derive update schemes for each architecture. Another portion of this work focuses on the often used technique of batch normalization. Although found to be successful …


Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li Jan 2020

Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li

Theses and Dissertations--Statistics

Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …


Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui Jan 2020

Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui

Theses and Dissertations--Statistics

In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.

In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …


Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu Jan 2020

Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu

Theses and Dissertations--Statistics

A common problem in regression analysis (linear or nonlinear) is assessing the lack-of-fit. Existing methods make parametric or semi-parametric assumptions to model the conditional mean or covariance matrices. In this dissertation, we propose fully nonparametric methods that make only additive error assumptions. Our nonparametric approach relies on ideas from nonparametric smoothing to reduce the test of association (lack-of-fit) problem into a nonparametric multivariate analysis of variance. A major problem that arises in this approach is that the key assumptions of independence and constant covariance matrix among the groups will be violated. As a result, the standard asymptotic theory is not …


Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang Jan 2020

Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang

Theses and Dissertations--Statistics

Kinetic modeling of the time dependence of metabolite concentrations including the unstable isotope labeled species is an important approach to simulate metabolic pathway dynamics. It is also essential for quantitative metabolic flux analysis using tracer data. However, as the metabolic networks are complex including extensive compartmentation and interconnections, the parameter estimation for enzymes that catalyze individual reactions needed for kinetic modeling is challenging. As the pa- rameter space is large and multi-dimensional while kinetic data are comparatively sparse, the estimation procedure (especially the point estimation methods) often en- counters multiple local maximum such that standard maximum likelihood methods may yield …


Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou Jan 2020

Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou

Theses and Dissertations--Statistics

Statistical intervals (e.g., confidence, prediction, or tolerance) are widely used to quantify uncertainty, but complex settings can create challenges to obtain such intervals that possess the desired properties. My thesis will address diverse data settings and approaches that are shown empirically to have good performance. We first introduce a focused treatment on using a single-layer bootstrap calibration to improve the coverage probabilities of two-sided parametric tolerance intervals for non-normal distributions. We then turn to zero-inflated data, which are commonly found in, among other areas, pharmaceutical and quality control applications. However, the inference problem often becomes difficult in the presence of …


Measuring Change: Prediction Of Early Onset Sepsis, Aric Schadler Jan 2020

Measuring Change: Prediction Of Early Onset Sepsis, Aric Schadler

Theses and Dissertations--Statistics

Sepsis occurs in a patient when an infection enters into the blood stream and spreads throughout the body causing a cascading response from the immune system. Sepsis is one of the leading causes of morbidity and mortality in today’s hospitals. This is despite published and accepted guidelines for timely and appropriate interventions for septic patients. The largest barrier to applying these interventions is the early identification of septic patients. Early identification and treatment leads to better outcomes, shorter lengths of stay, and financial savings for healthcare institutions. In order to increase the lead time in recognizing patients trending towards septicemia …


How We Can Extend The Standard Deviation Notion With Neutrosophic Interval And Quadruple Neutrosophic Numbers, Victor Christianto, Florentin Smarandache, Muhammad Aslam Jan 2020

How We Can Extend The Standard Deviation Notion With Neutrosophic Interval And Quadruple Neutrosophic Numbers, Victor Christianto, Florentin Smarandache, Muhammad Aslam

Branch Mathematics and Statistics Faculty and Staff Publications

During scientific demonstrating of genuine specialized framework we can meet any sort and rate model vulnerability. Its reasons can be incognizance of modelers or information mistake. In this way, characterization of vulnerabilities, as for their sources, recognizes aleatory and epistemic ones. The aleatory vulnerability is an inalienable information variety related with the researched framework or its condition. Epistemic one is a vulnerability that is because of an absence of information on amounts or procedures of the framework or the earth [7]. Right now, we examine fourfold neutrosophic numbers and their potential application for practical displaying of physical frameworks, particularly in …


Nonparametric False Discovery Rate Control For Identifying Simultaneous Signals, Sihai Dave Zhao, Yet Tian Nguyen Jan 2020

Nonparametric False Discovery Rate Control For Identifying Simultaneous Signals, Sihai Dave Zhao, Yet Tian Nguyen

Mathematics & Statistics Faculty Publications

It is frequently of interest to identify simultaneous signals, defined as features that exhibit statistical significance across each of several independent experiments. For example, genes that are consistently differentially expressed across experiments in different animal species can reveal evolutionarily conserved biological mechanisms. However, in some problems the test statistics corresponding to these features can have complicated or unknown null distributions. This paper proposes a novel nonparametric false discovery rate control procedure that can identify simultaneous signals even without knowing these null distributions. The method is shown, theoretically and in simulations, to asymptotically control the false discovery rate. It was also …