Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2020

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 421 - 450 of 618

Full-Text Articles in Statistics and Probability

Entomological Index And Home Environment Contribution­ ­To Dengue Hemorrhagic Fever In Mataram City, Indonesia, Tri Baskoro Tunggul Satoto, Nur Alvira Pascawati, Tri Wibawa, Roger Frutos, Sylvie Maguin, I Kadek Mulyawan, Ali Wardana Feb 2020

Entomological Index And Home Environment Contribution­ ­To Dengue Hemorrhagic Fever In Mataram City, Indonesia, Tri Baskoro Tunggul Satoto, Nur Alvira Pascawati, Tri Wibawa, Roger Frutos, Sylvie Maguin, I Kadek Mulyawan, Ali Wardana

Kesmas

Indonesia is a member of Southeast Asia Regional Office (SEARO) ranked the first in dengue hemorrhagic fever (DHF) problem based on incidence rate (IR) and case fatality rate (CFR). Several provinces in Indonesia experience an outbreak, one of which is the Mataram City in West Nusa Tenggara Province. Mataram City is an endemic area of DHF because the DHF cases are always found in three consecutive years with the number of cases that fluctuate and tend to increase. This study aimed to obtain factors that could be used to improve early warning systems in controlling DHF. This study used a …


Utilization Of Family Planning Contraceptives Among Women Inthe Coastal Area Of South Buru District, Maluku, 2017, Christiana Rialine Titaley, Ninik Sallatalohy Feb 2020

Utilization Of Family Planning Contraceptives Among Women Inthe Coastal Area Of South Buru District, Maluku, 2017, Christiana Rialine Titaley, Ninik Sallatalohy

Kesmas

Maluku Province is one among provinces in Indonesia with a contraceptive prevalence rate (CPR) lower than the national average. This study aimed toexamine factors associated with the utilization of family planning contraceptives among women of reproductive age living in the coastal area of South BuruDistrict, Maluku, Indonesia. Data were derived from a household health survey conducted in five subdistricts in South Buru, e.g., Namrole, Leksula, Waesama,Kapala Madan and Ambalau Subdistricts on November 2017 by the Faculty of Medicine, Pattimura University in Ambon. Information on contraceptive usewere collected from 390 married women aged 20 - 49 years. Bivariate and multivariate logistic …


Determinants Of Stunted Children In Indonesia: A Multilevel Analysis At The Individual, Household, And Community Levels, Febri Wicaksono, Titik Harsanti Feb 2020

Determinants Of Stunted Children In Indonesia: A Multilevel Analysis At The Individual, Household, And Community Levels, Febri Wicaksono, Titik Harsanti

Kesmas

This study aimed to examine the risk factors of childhood undernutrition in Indonesia. Determinants of childhood stunting were examined by using the 2013 Indonesia Basic Health Research Survey dataset. A total of 76,165 children aged under 5 years were included in this study. The analysis used multivariate multilevel logistic regression to determine adjusted odds ratios (aORs). The prevalence of stunting in the sample population was 36.7%. The odds of stunting increased significantly among the under-five boys, children living in slum area, and the increase of household member (aOR = 1.11, 95 %CI: 1.06–1.15; 1.09, 95%CI: 1.04–1.15; and 1.03, 95%CI: 1.02–1.04 …


Health Risk Behaviors: Smoking, Alcohol, Drugs, And Dating Among Youths In Rural Central Java, Zahroh Shaluhiyah, Syamsulhuda Budi Musthofa, Ratih Indraswari, Aditya Kusumawati Feb 2020

Health Risk Behaviors: Smoking, Alcohol, Drugs, And Dating Among Youths In Rural Central Java, Zahroh Shaluhiyah, Syamsulhuda Budi Musthofa, Ratih Indraswari, Aditya Kusumawati

Kesmas

Adolescents are more likely to adopt risky health behaviors, such as smoking, alcohol use, and sexual activity. This study examined the links betweensmoking, alcohol use, and risky dating behavior and analyzed how these factors influenced risky dating and other behaviors. It is expected that this studywould be used as a foundation for developing appropriate integrated intervention for multiple risk behaviors among youths. This study was an explanatory research study with a cross-sectional approach. It involved 160 youths aged 15-24 years randomly selected from purposive villages. Participants completedself-administrated questionnaires with an enumerator present. Data were analyzed using univariate, chi-square, and multiple …


Evaluation Of Program For Overcoming Intestinal Worm Infections Among Children, Henny Febriyanti, Haerawati Idris Feb 2020

Evaluation Of Program For Overcoming Intestinal Worm Infections Among Children, Henny Febriyanti, Haerawati Idris

Kesmas

Prevalence of intestinal worm infection in generall is extremely high in Indonesia among the poor population with poor sanitation. One of the government programs to address this problem is the distribution of medicines to prevent intestinal worm infections. However, the coverage of the achievement for this program is still low in several areas of public health centers in Palembang. Therefore, this study was conducted to evaluate the efficacy of the national program for preventing intestinal worm infections. The qualitative research design used evaluation model approach Context, Input, Process, and Product (CIPP) model. This study was conducted in one of health …


Small Area Estimation On Zero-Inflated Data Using Frequentist And Bayesian Approach, Kusman Sadik, Rahma Anisa, Euis Aqmaliyah Feb 2020

Small Area Estimation On Zero-Inflated Data Using Frequentist And Bayesian Approach, Kusman Sadik, Rahma Anisa, Euis Aqmaliyah

Journal of Modern Applied Statistical Methods

The most commonly used method of small area estimation (SAE) is the empirical best linear unbiased prediction method based on a linear mixed model. However, it is not appropriate in the case of the zero-inflated target variable with a mixture of zeros and continuously distributed positive values. Therefore, various model-based SAE methods for zero-inflated data are developed, such as the Frequentist approach and the Bayesian approach. Both approaches are compared with the survey regression (SR) method which ignores the presence of zero-inflation in the data. The results show that the two SAE approaches for zero-inflated data are capable to yield …


The Importance Of Type I Error Rates When Studying Bias In Monte Carlo Studies In Statistics, Michael Harwell Feb 2020

The Importance Of Type I Error Rates When Studying Bias In Monte Carlo Studies In Statistics, Michael Harwell

Journal of Modern Applied Statistical Methods

Two common outcomes of Monte Carlo studies in statistics are bias and Type I error rate. Several versions of bias statistics exist but all employ arbitrary cutoffs for deciding when bias is ignorable or non-ignorable. This article argues Type I error rates should be used when assessing bias.


Dynamic Conditional Correlation Garch: A Multivariate Time Series Novel Using A Bayesian Approach, Diego Nascimento, Cleber Xavier, Israel Felipe, Francisco Louzada Neto Feb 2020

Dynamic Conditional Correlation Garch: A Multivariate Time Series Novel Using A Bayesian Approach, Diego Nascimento, Cleber Xavier, Israel Felipe, Francisco Louzada Neto

Journal of Modern Applied Statistical Methods

The Dynamic Conditional Correlation GARCH (DCC-GARCH) mutation model is considered using a Monte Carlo approach via Markov chains in the estimation of parameters, time-dependence variation is visually demonstrated. Fifteen indices were analyzed from the main financial markets of developed and developing countries from different continents. The performances of indices are similar, with a joint evolution. Most index returns, especially SPX and NDX, evolve over time with a higher positive correlation.


Regression When There Are Two Covariates: Some Practical Reasons For Considering Quantile Grids, Rand Wilcox Feb 2020

Regression When There Are Two Covariates: Some Practical Reasons For Considering Quantile Grids, Rand Wilcox

Journal of Modern Applied Statistical Methods

When dealing with the association between some random variable and two covariates, extensive experience with smoothers indicates that often a linear model poorly reflects the nature of the association. A simple approach via quantile grids that reflects the nature of the association is given. The two main goals are to illustrate this approach can make a practical difference, and to describe R functions for applying it. Included are comments on dealing with more than two covariates.


Bivariate Analogs Of The Wilcoxon–Mann–Whitney Test And The Patel–Hoel Method For Interactions, Rand Wilcox Feb 2020

Bivariate Analogs Of The Wilcoxon–Mann–Whitney Test And The Patel–Hoel Method For Interactions, Rand Wilcox

Journal of Modern Applied Statistical Methods

A fundamental way of characterizing how two independent compares compare is in terms of the probability that a randomly sampled observation from the first group is less than a randomly sampled observation from the second group. The paper suggests a bivariate analog and investigates methods for computing confidence intervals. An interaction for a two-by-two design is investigated as well.


Teaching A University Course On The Mathematics Of Gambling, Stewart N. Ethier, Fred M. Hoppe Feb 2020

Teaching A University Course On The Mathematics Of Gambling, Stewart N. Ethier, Fred M. Hoppe

UNLV Gaming Research & Review Journal

Courses on the mathematics of gambling have been offered by a number of colleges and universities, and for a number of reasons. In the past 15 years, at least seven potential textbooks for such a course have been published. In this article we objectively compare these books for their probability content, their gambling content, and their mathematical level, to see which ones might be most suitable, depending on student interests and abilities. This is not a book review (e.g., none of the books is recommended over others) but rather an essay offering advice about which topics to include in a …


Assessing The Accuracy Of Approximate Confidence Intervals Proposed For The Mean Of Poisson Distribution, Alireza Shirvani, Malek Fathizadeh Feb 2020

Assessing The Accuracy Of Approximate Confidence Intervals Proposed For The Mean Of Poisson Distribution, Alireza Shirvani, Malek Fathizadeh

Journal of Modern Applied Statistical Methods

The Poisson distribution is applied as an appropriate standard model to analyze count data. Because this distribution is known as a discrete distribution, representation of accurate confidence intervals for its distribution mean is extremely difficult. Approximate confidence intervals were presented for the Poisson distribution mean. The purpose of this study is to simultaneously compare several confidence intervals presented, according to the average coverage probability and accurate confidence coefficient and the average confidence interval length criteria.


Analytical Closed-Form Solution For General Factor With Many Variables, Stan Lipovetsky, Vladimir Manewitsch Feb 2020

Analytical Closed-Form Solution For General Factor With Many Variables, Stan Lipovetsky, Vladimir Manewitsch

Journal of Modern Applied Statistical Methods

The factor analytic triad method of one-factor solution gives the explicit analytical form for a common latent factor built by three variables. The current work considers analytical presentation of a general latent factor constructed in a closed-form solution for multivariate case. The results can be supportive to theoretical description and practical application of latent variable modeling, especially for big data because the analytical closed-form solution is not prone to data dimensionality.


Regression Modeling And Prediction By Individual Observations Versus Frequency, Stan Lipovetsky Feb 2020

Regression Modeling And Prediction By Individual Observations Versus Frequency, Stan Lipovetsky

Journal of Modern Applied Statistical Methods

A regression model built by a dataset could sometimes demonstrate a low quality of fit and poor predictions of individual observations. However, using the frequencies of possible combinations of the predictors and the outcome, the same models with the same parameters may yield a high quality of fit and precise predictions for the frequencies of the outcome occurrence. Linear and logistical regressions are used to make an explicit exposition of the results of regression modeling and prediction.


Measuring Localization Confidence For Quantifying Accuracy And Heterogeneity In Single-Molecule Super-Resolution Microscopy, Hesam Mazidi, Tianben Ding, Arye Nehorai, Matthew D. Lew Feb 2020

Measuring Localization Confidence For Quantifying Accuracy And Heterogeneity In Single-Molecule Super-Resolution Microscopy, Hesam Mazidi, Tianben Ding, Arye Nehorai, Matthew D. Lew

Electrical & Systems Engineering Publications and Presentations

We present a computational method, termed Wasserstein-induced flux (WIF), to robustly quantify the accuracy of individual localizations within a single-molecule localization microscopy (SMLM) dataset without ground- truth knowledge of the sample. WIF relies on the observation that accurate localizations are stable with respect to an arbitrary computational perturbation. Inspired by optimal transport theory, we measure the stability of individual localizations and develop an efficient optimization algorithm to compute WIF. We demonstrate the advantage of WIF in accurately quantifying imaging artifacts in high-density reconstruction of a tubulin network. WIF represents an advance in quantifying systematic errors with unknown and complex distributions, …


Analysis Of An Agent-Based Model For Predicting The Behavior Of Bighead Carp (Hypophthalmichthys Nobilis) Under The Influence Of Acoustic Deterrence, Craig Garzella, Joseph Gaudy, Karl R. B. Schmitt, Arezu Mansuri Feb 2020

Analysis Of An Agent-Based Model For Predicting The Behavior Of Bighead Carp (Hypophthalmichthys Nobilis) Under The Influence Of Acoustic Deterrence, Craig Garzella, Joseph Gaudy, Karl R. B. Schmitt, Arezu Mansuri

Spora: A Journal of Biomathematics

Bighead carp (Hypophthalmichthys nobilis) are an invasive, voracious, highly fecund species threatening the ecological integrity of the Great Lakes. This agent-based model and analysis explore bighead carp behavior in response to acoustic deterrence in an effort to discover properties that increase likelihood of deterrence system failure. Results indicate the most significant (p < 0.05) influences on barrier failure are the quantity of detritus and plankton behind the barrier, total number of bighead carp successfully deterred by the barrier, and number of native fishes freely moving throughout the simulation. Quantity of resources behind the barrier influence bighead carp to penetrate when populations are resource deprived. When native fish populations are low, an accumulation of phytoplankton can occur, increasing the likelihood of an algal bloom occurrence. Findings of this simulation suggest successful implementation with proper maintenance of an acoustic deterrence system has potential of abating the threat of bighead carp on ecological integrity of the Great Lakes.


Session 11 - Methods: Bootstrap Control Chart For Pareto Percentiles, Ruth Burkhalter Feb 2020

Session 11 - Methods: Bootstrap Control Chart For Pareto Percentiles, Ruth Burkhalter

SDSU Data Science Symposium

Lifetime percentile is an important indicator of product reliability. However, the sampling distribution of a percentile estimator for any lifetime distribution is not a bell shaped one. As a result, the well-known Shewhart-type control chart cannot be applied to monitor the product lifetime percentiles. In this presentation, Bootstrap control charts based on maximum likelihood estimator (MLE) are proposed for monitoring Pareto percentiles. An intensive simulation study is conducted to compare the performance among the proposed MLE Bootstrap control chart and Shewhart-type control chart.


Evaluation Of Text Mining Techniques Using Twitter Data For Hurricane Disaster Resilience, Joshua Eason, Sathish Kumar Feb 2020

Evaluation Of Text Mining Techniques Using Twitter Data For Hurricane Disaster Resilience, Joshua Eason, Sathish Kumar

SDSU Data Science Symposium

Data obtained from social media microblogging websites such as Twitter provide the unique ability to collect and analyze conversations of the public in order to gain perspective on the thoughts and feelings of the general public. Sentiment and volume analysis techniques were applied to the dataset in order to gain an understanding of the amount and level of sentiment associated with certain disaster-related tweets, including a topical analysis of specific terms. This study showed that disaster-type events such as a hurricane can cause some strong negative sentiment in the period of time directly preceding the event, but ultimately returns quickly …


Asymptotic Simultaneous Estimations For Contrasts Of Quantiles, Lawrence Sethor Segbehoe, Frank Schaarschmidt, Gemechis Dilba Djira Feb 2020

Asymptotic Simultaneous Estimations For Contrasts Of Quantiles, Lawrence Sethor Segbehoe, Frank Schaarschmidt, Gemechis Dilba Djira

SDSU Data Science Symposium

Although the expected value is popular, many researches in the health and social sciences involve skewed distributions and inferences concerning quantiles. Most standard multiple comparison procedures require the normality assumption. For example, few methods exist for comparing the medians of independent samples or quantiles of several distributions in general. To our knowledge, there is no general-purpose method for constructing simultaneous confidence intervals for multiple contrasts of quantiles. In this paper, we develop an asymptotic method for constructing such intervals and extend the idea to that of time-to-event data in survival analysis. Small-sample performance of the proposed method is assessed in …


Exploration Of Factors Associated With Perceptions Of Community Safety Among Youth In Hillsborough County, Florida: A Convergent Parallel Mixed-Methods Approach, Yingwei Yang Feb 2020

Exploration Of Factors Associated With Perceptions Of Community Safety Among Youth In Hillsborough County, Florida: A Convergent Parallel Mixed-Methods Approach, Yingwei Yang

USF Tampa Graduate Theses and Dissertations

Introduction: Youth perceived safety is not only linked to crime and violence in a neighborhood but is also associated with health risk behaviors and certain neighborhood characteristics. The purpose of this mixed-methods study was to measure the co-occurring effects of individual and community risk factors by conducting a secondary data analysis using structural equation modeling (SEM) and to explore reasons for youth feeling safe/unsafe in their community using photovoice methodology.

Methods: Syndemic theory/model served as the theoretical framework to guide this mixed-methods study with a convergent parallel design. The quantitative strand (first manuscript) utilized an existing dataset collected from middle …


Comparison Of Upper Extremity Function In Women With And Women Without A History Of Breast Cancer, Mary Insana Fisher, Gilson J. Capilouto, Terry Malone, Heather M. Bush, Timothy L. Uhl Feb 2020

Comparison Of Upper Extremity Function In Women With And Women Without A History Of Breast Cancer, Mary Insana Fisher, Gilson J. Capilouto, Terry Malone, Heather M. Bush, Timothy L. Uhl

Communication Sciences and Disorders Faculty Publications

Background

Breast cancer treatments often result in upper extremity functional limitations in both the short and long term. Current evidence makes comparisons against a baseline or contralateral limb, but does not consider changes in function associated with aging.

Objective

The objective of this study was to compare upper extremity function between women treated for breast cancer more than 12 months in the past and women without cancer.

Design

This was an observational cross-sectional study.

Methods

Women who were diagnosed with breast cancer and had a mean post-surgical treatment time of 51 months (range = 12–336 months) were compared with women …


Bayesian Reliability Analysis Of The Power Law Process And Statistical Modeling Of Computer And Network Vulnerabilities With Cybersecurity Application, Freeh N. Alenezi Feb 2020

Bayesian Reliability Analysis Of The Power Law Process And Statistical Modeling Of Computer And Network Vulnerabilities With Cybersecurity Application, Freeh N. Alenezi

USF Tampa Graduate Theses and Dissertations

As most of mankind now lives in an era of high dependence on multiple technologies and complex systems to store and manage sensitive information, researchers are constantly urged to obtain and improve measurements and methodologies that have the ability to evaluate systems reliability and security. The objectives of the present dissertation are to improve the Bayesian reliability estimation of a software package where the Power Law Process, also known as Non-Homogeneous Poisson Process, is the underlying failure model and to develop a set of statistical models evaluating computer operating systems vulnerabilities. Furthermore, we develop a reliability function of a computer …


Hierarchical Clustering Analyses Of Plasma Proteins In Subjects With Cardiovascular Risk Factors Identify Informative Subsets Based On Differential Levels Of Angiogenic And Inflammatory Biomarkers, Zachary Winder, Tiffany L. Sudduth, David W. Fardo, Qiang Cheng, Larry B. Goldstein, Peter T. Nelson, Frederick A. Schmitt, Gregory A. Jicha, Donna M. Wilcock Feb 2020

Hierarchical Clustering Analyses Of Plasma Proteins In Subjects With Cardiovascular Risk Factors Identify Informative Subsets Based On Differential Levels Of Angiogenic And Inflammatory Biomarkers, Zachary Winder, Tiffany L. Sudduth, David W. Fardo, Qiang Cheng, Larry B. Goldstein, Peter T. Nelson, Frederick A. Schmitt, Gregory A. Jicha, Donna M. Wilcock

Sanders-Brown Center on Aging Faculty Publications

Agglomerative hierarchical clustering analysis (HCA) is a commonly used unsupervised machine learning approach for identifying informative natural clusters of observations. HCA is performed by calculating a pairwise dissimilarity matrix and then clustering similar observations until all observations are grouped within a cluster. Verifying the empirical clusters produced by HCA is complex and not well studied in biomedical applications. Here, we demonstrate the comparability of a novel HCA technique with one that was used in previous biomedical applications while applying both techniques to plasma angiogenic (FGF, FLT, PIGF, Tie-2, VEGF, VEGF-D) and inflammatory (MMP1, MMP3, MMP9, IL8, TNFα) protein data to …


Informal Professional Development On Twitter: Exploring The Online Communities Of Mathematics Educators, Jaymie Ruddock Feb 2020

Informal Professional Development On Twitter: Exploring The Online Communities Of Mathematics Educators, Jaymie Ruddock

SMU Journal of Undergraduate Research

Professional development in its most traditional form is a classroom setting with a lecturer and an overwhelming amount of information. It is no surprise, then, that informal professional development away from institutions and on the teacher's own terms is a growing phenomenon due to an increased presence of educators on social media. These communities of educators use hashtags to broadcast to each other, with general hashtags such as #edchat having the broadest audience. However, many math educators usethe hashtags #ITeachMath and #MTBoS, communities I was interested in learning more about. I built a python script that used Tweepy to connect …


Mathematical Modelling Of The Tuberculosis Epidemiology, Ally Yeketi Ayinla Feb 2020

Mathematical Modelling Of The Tuberculosis Epidemiology, Ally Yeketi Ayinla

Student Works (2020-2029)

This project analyses the tuberculosis (TB) epidemic mathematically using compartmental modelling approach. Three models are presented to discuss drug susceptible and multi-drug resistant TB. The first model presented has 4 compartments; susceptible, exposed, infectious and recovered. The relevance of the exposed class in managing TB is analysed and found to be useful in delaying the eventual onset of the infection. Compared to previous researches, our results significantly show that when efforts are made such that no infected individual bypasses the exposed class and progresses directly to the infectious, the TB epidemic is successfully combatted. Also, the model is used to …


Sufficient Dimension Folding In Regression Via Distance Covariance For Matrix‐Valued Predictors, Wenhui Sheng, Qingcong Yuan Feb 2020

Sufficient Dimension Folding In Regression Via Distance Covariance For Matrix‐Valued Predictors, Wenhui Sheng, Qingcong Yuan

Mathematical and Statistical Science Faculty Research and Publications

In modern data, when predictors are matrix/array‐valued, building a reasonable model is much more difficult due to the complicate structure. However, dimension folding that reduces the predictor dimensions while keeps its structure is critical in helping to build a useful model. In this paper, we develop a new sufficient dimension folding method using distance covariance for regression in such a case. The method works efficiently without strict assumptions on the predictors. It is model‐free and nonparametric, but neither smoothing techniques nor selection of tuning parameters is needed. Moreover, it works for both univariate and multivariate response cases. In addition, we …


An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone Jan 2020

An Automatic Interaction Detection Hybrid Model For Bankcard Response Classification, Yan Wang, Sherry Ni, Brian Stone

Published and Grey Literature from PhD Candidates

Data mining techniques have numerous applications in bankcard response modeling. Logistic regression has been used as the standard modeling tool in the financial industry because of its almost always desirable performance and its interpretability. In this paper, we propose a hybrid bankcard response model, which integrates decision tree-based chi-square automatic interaction detection (CHAID) into logistic regression. In the first stage of the hybrid model, CHAID analysis is used to detect the possible potential variable interactions. Then in the second stage, these potential interactions are served as the additional input variables in logistic regression. The motivation of the proposed hybrid model …


A Two-Stage Hybrid Model By Using Artificial Neural Networks As Feature Construction Algorithms, Yan Wang, Sherry Ni, Brian Stone Jan 2020

A Two-Stage Hybrid Model By Using Artificial Neural Networks As Feature Construction Algorithms, Yan Wang, Sherry Ni, Brian Stone

Published and Grey Literature from PhD Candidates

We propose a two-stage hybrid approach with neural networks as the new feature construction algorithms for bankcard response classifications. The hybrid model uses a very simple neural network structure as the new feature construction tool in the first stage, then the newly created features are used as the additional input variables in logistic regression in the second stage. The model is compared with the traditional one-stage model in credit customer response classification. It is observed that the proposed two-stage model outperforms the one-stage model in terms of accuracy, the area under the ROC curve, and KS statistic. By creating new …


Predicting Class-Imbalanced Business Risk Using Resampling, Regularization, And Model Ensembling Algorithms, Yan Wang, Sherry Ni Jan 2020

Predicting Class-Imbalanced Business Risk Using Resampling, Regularization, And Model Ensembling Algorithms, Yan Wang, Sherry Ni

Published and Grey Literature from PhD Candidates

We aim at developing and improving the imbalanced business risk modeling via jointly using proper evaluation criteria, resampling, cross-validation, classifier regularization, and ensembling techniques. Area Under the Receiver Operating Characteristic Curve (AUC of ROC) is used for model comparison based on 10-fold cross-validation. Two undersampling strategies including random undersampling (RUS) and cluster centroid undersampling (CCUS), as well as two oversampling methods including random oversampling (ROS) and Synthetic Minority Oversampling Technique (SMOTE), are applied. Three highly interpretable classifiers, including logistic regression without regularization (LR), L1-regularized LR (L1LR), and decision tree (DT) are implemented. Two ensembling techniques, including Bagging and Boosting, are …


A Xgboost Risk Model Via Feature Selection And Bayesian Hyper-Parameter Optimization, Yan Wang, Sherry Ni Jan 2020

A Xgboost Risk Model Via Feature Selection And Bayesian Hyper-Parameter Optimization, Yan Wang, Sherry Ni

Published and Grey Literature from PhD Candidates

This paper aims to explore models based on the extreme gradient boosting (XGBoost) approach for business risk classification. Feature selection (FS) algorithms and hyper-parameter optimizations are simultaneously considered during model training. The five most commonly used FS methods including weight by Gini, weight by Chi-square, hierarchical variable clustering, weight by correlation, and weight by information are applied to alleviate the effect of redundant features. Two hyper-parameter optimization approaches, random search (RS) and Bayesian tree-structuredParzen Estimator (TPE), are applied in XGBoost. The effect of different FS and hyper-parameter optimization methods on the model performance are investigated by the Wilcoxon Signed Rank …