Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (51)
- Statistical Models (24)
- Engineering (21)
- Life Sciences (17)
- Statistical Methodology (16)
-
- Social and Behavioral Sciences (13)
- Biostatistics (12)
- Operations Research, Systems Engineering and Industrial Engineering (12)
- Categorical Data Analysis (11)
- Business (10)
- Data Science (10)
- Industrial Engineering (10)
- Computer Sciences (8)
- Longitudinal Data Analysis and Time Series (8)
- Mathematics (7)
- Medicine and Health Sciences (7)
- Multivariate Analysis (7)
- Probability (7)
- Education (6)
- Statistical Theory (6)
- Applied Mathematics (5)
- Geography (5)
- Operational Research (5)
- Sociology (5)
- Business Analytics (4)
- Educational Assessment, Evaluation, and Research (4)
- Software Engineering (4)
- Survival Analysis (4)
- Keyword
-
- Pure sciences (14)
- Statistics (9)
- Bayesian (6)
- Applied sciences (4)
- Arkansas (3)
-
- Bayesian Statistics (3)
- Differential Item Functioning (3)
- Horseshoe (3)
- Machine Learning (3)
- Regression (3)
- Survival analysis (3)
- Artificial Intelligence (2)
- Bayesian Analysis (2)
- Biological sciences (2)
- Clinical trials (2)
- Conditional autoregressive prior (2)
- Data Science (2)
- Distance Correlation (2)
- Item response theory (2)
- Landsat time series (2)
- MCMC (2)
- Machine learning (2)
- Markov Chain Monte Carlo (2)
- Measurement invariance (2)
- Ovarian cancer (2)
- Poisson (2)
- Prediction (2)
- Psychometrics (2)
- Quality control (2)
- Statistical Learning (2)
- Publication Year
- Publication
-
- Graduate Theses and Dissertations (84)
- Industrial Engineering Undergraduate Honors Theses (5)
- Computer Science and Computer Engineering Undergraduate Honors Theses (3)
- Data Science Undergraduate Honors Theses (3)
- Journal of the Arkansas Academy of Science (3)
-
- Finance Undergraduate Honors Theses (2)
- Information Systems Undergraduate Honors Theses (2)
- Biological and Agricultural Engineering Undergraduate Honors Theses (1)
- Chemical Engineering Faculty Publications and Presentations (1)
- Chemical Engineering Undergraduate Honors Theses (1)
- Economics Undergraduate Honors Theses (1)
- Electrical Engineering and Computer Science Undergraduate Honors Theses (1)
- Management Undergraduate Honors Theses (1)
- Marketing Faculty Publications and Presentations (1)
- Mathematical Sciences Faculty Publications and Presentations (1)
- Mathematical Sciences Undergraduate Honors Theses (1)
- Political Science Undergraduate Honors Theses (1)
- Sociology and Criminology Faculty Publications and Presentations (1)
- Publication Type
Articles 1 - 30 of 113
Full-Text Articles in Statistics and Probability
Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder
Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder
Electrical Engineering and Computer Science Undergraduate Honors Theses
Splines are used for representing complex functions. In statistics, splines can be used for distributional shapes that are difficult to model by traditional parametric approaches. Ramsay (1) uses M-Spline bases to estimate continuous distributions. Semi-Parametric Quantile Regression (SPQR), developed by Xu and Reich (2), models conditional distributions where a neural network is used to estimate the basis function weights that depend on covariates. (3) implements a package for SPQR in R. We build on this by implementing a version of SPQR in Python with PyTorch. By using PyTorch, we can use more sophisticated deep learning architectures than those available in …
Do I Have To Know Someone? The Impact Of Factors Beyond The Playing Field That Affect Mlb Hall Of Fame Voting., Jackson Wollscheid
Do I Have To Know Someone? The Impact Of Factors Beyond The Playing Field That Affect Mlb Hall Of Fame Voting., Jackson Wollscheid
Economics Undergraduate Honors Theses
With the creation of WAR (Wins above Replacement), it has created an index to measure the player’s value using a variety of statistics relevant to that player’s position. The higher a player’s career WAR then the greater the impact that player had on the game during his career. However, players are not simply inducted into the Hall of Fame on this measure of value but on voting from members of the BBWAA (Baseball Writers’ Association of America). This research investigates external factors outside of standard playing statistics’ influence on Hall of Fame voting, including team affiliation, career length, and other …
Survival Patterns Among Adult And Pediatric Bone Cancer Patients, Ethan Estes
Survival Patterns Among Adult And Pediatric Bone Cancer Patients, Ethan Estes
Mathematical Sciences Undergraduate Honors Theses
Recently noted, Huang et al. (2023), machine learning (ML) models, while offering great advantages over traditional statistical predictive modeling methods, are less explored in the analysis of survival and other similar time-to-event predictive data modeling. ML methods such as neural networks offer a great deal of promise but need to be further explored to investigate their comparative power in predicting survival outcomes. Focusing specifically on survival analysis in adult and pediatric bone cancer patients, traditional methods, like shown in Emmert-Streib and Dehmer (2019), will be shown with machine learning models using methods in Hothorn, Hornik, and Zeileis (2006). In this …
Explainability In Deep Learning For Density Regression, Dalton James Oxford
Explainability In Deep Learning For Density Regression, Dalton James Oxford
Graduate Theses and Dissertations
Classical statistical methods focus on explainability and inferential power. Machine learning and deep learning can handle non-linear, high-dimensional data better than traditional methods. In modeling, a clear understanding and interpretation are essential to decision-making. Recent work in quantile regression and extreme modeling has begun to use deep learning due to its performance on high-dimensional, non-linear data. Semi-Parametric Quantile Regression (SPQR) is a nonparametric spline-based approach to quantile regression that estimates the conditional PDF and CDF of the response. Semi-Parametric Quantile Regression for Extremes (SPQRx) is a recent extension of SPQR that provides two features: out-of-sample estimation and accurate extreme-tailed estimation. …
Two Topics In Survival Analysis: Restricted Distance Covariance Test For Non-Proportional Hazard And A New Estimation For Dropout Rate, Ruizhe Yin
Graduate Theses and Dissertations
When treatment effects change over time, standard statistical methods, such as the log-rank test and the Cox proportional hazards model, may give misleading results. This dissertation presents the restricted distance covariance (rdcov) test, a nonparametric method that compares survival curves between groups within a chosen study period [0, τ] using right censored data. The statistic measures the dependence between pre-specified group labels and survival times using pairwise distances from Kaplan-Meier estimates. Our method does not rely on the proportional hazards assumption, and it equals zero only when survival functions are identical across groups. Thus, this test can be applied to …
Tree-Based Differential Item Functioning Detection Methods: Exploring Their Performance In Diverse Measurement Scenarios, Nana Amma Berko Asamoah
Tree-Based Differential Item Functioning Detection Methods: Exploring Their Performance In Diverse Measurement Scenarios, Nana Amma Berko Asamoah
Graduate Theses and Dissertations
Despite the availability of numerous methods for detecting differential item functioning (DIF), the continued development and evaluation of innovative, data-driven approaches remains essential. Tree-based methods, in particular, represent a significant advancement in DIF detection. Unlike some traditional techniques, they can simultaneously screen multiple variables for DIF without discretizing continuous variables, and do not require the pre-specification of focal and reference groups; capabilities that are especially valuable in today’s diverse and multifaceted assessment contexts. However, research systematically examining the performance of these methods under realistic measurement conditions is limited. This dissertation, in three simulation studies, critically examines the robustness and practical …
Using Gaussian Process Regression To Learn Thermodynamic Equations Of State With Uncertainty Quantification, Austen T. Lee
Using Gaussian Process Regression To Learn Thermodynamic Equations Of State With Uncertainty Quantification, Austen T. Lee
Chemical Engineering Undergraduate Honors Theses
This study investigates the use of derivative-informed Gaussian Process (GP) models to estimate thermodynamic behavior across temperature and density by building a Helmholtz-based equation of state. Argon, a stable monatomic gas, was chosen as a case study within the vapor region. The GP model was trained using values of experimentally measurable properties found by taking first and second derivatives of the original potential function. Results show that while the GP model offered uncertainty quantification and informed thermodynamic behavior, it predicted values that deviated from the ground truth depending on the property. The model exhibited high confidence in regions with substantial …
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Computer Vision In Soccer: Yolov11 Analytics Engine For Quantifying Game Strategy, Connor S. Maurer
Data Science Undergraduate Honors Theses
Single-shot object detection capabilities significantly reduce computational overhead for real-time computer vision in sports analytics at 60 FPS. YOLO11’s lightweight CNN gives promising accuracy while meeting the low-latency demand of dynamic soccer matches. As data-driven approaches take over the sport of soccer, efficient player tracking systems become critical for informing coach’s strategies. I prototype the ETL (Extract, Transform, Load) process of data collected from a single- shot detection program and evaluate its viability for estimating player fatigue. YOLO11 detects players, the ball, and other characteristics, with the output transformed by homography to estimate the positions in the real world. These …
Exploring Latent Mediation Through Bayesian Regularization Methods Of Lasso, Ridge, Horseshoe, Spike-And-Slab, Ethan Harris
Exploring Latent Mediation Through Bayesian Regularization Methods Of Lasso, Ridge, Horseshoe, Spike-And-Slab, Ethan Harris
Graduate Theses and Dissertations
Regularization is a powerful tool to combat overfitting and drive sparsity in complex models. Regularization was initially applied in regression modeling but has been increasingly utilized in structural equation modeling where its utility in identifying the essential components has helped improve modeling. As structural equation models have increased in complexity both in the number of indicators but also the number of latent factors, researchers have begun to investigate how applying Bayesian regularization to these systems can further push the limits on modeling complex models with limited sample sizes. One area where research is limited is the application of Bayesian regularizations …
Application Of Ordinal Regression Models To Acquired Stress Resistance In Wild Strains Of Saccharomyces Cerevisiae, Carson Stacy
Application Of Ordinal Regression Models To Acquired Stress Resistance In Wild Strains Of Saccharomyces Cerevisiae, Carson Stacy
Graduate Theses and Dissertations
This thesis explores the application of ordinal regression to the analysis of semi-quantitative growth assays often used when comparing fitness for different strains of the model yeast Saccharomyces cerevisiae. For stress survival assays, yeast stress resistance is measured using an ordered survival score that ranges from 0 (no growth) to 4 (confluent growth). Traditional approaches to analyze this type of data either treats data as a nominal categorical variable or as a continuous numerical variable. These approaches risk loss of information or violation of testing assumptions. In contrast, cumulative logit ordinal regression uses the information contained in the order …
Nonparametric Methods For Bayesian Community Detection In Complex Networks, Kedran Young
Nonparametric Methods For Bayesian Community Detection In Complex Networks, Kedran Young
Graduate Theses and Dissertations
Network analysis is becoming an increasingly popular interdisciplinary area of study, with emerging interest in fields like sociology, biology, economics, and ecology. Within the niche of network analysis, capturing the community structure of a network is one important achievement that many statisticians have been working toward over recent decades. The most popular modeling technique for latent community detection is the Stochastic Block Model (SBM), which falls into the category of latent variable models and will serve as the baseline model throughout this thesis. SBM is widely regarded as the most effective community detection method as it detects latent community membership …
Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo
Latent Variable Dyadic Regression Models For Predicting Over/Under Bets In Sports Betting, Alexcia Trejo
Graduate Theses and Dissertations
This thesis explores the use of latent factor models to uncover hidden structures in pair wise outcomes derived from Over/Under betting markets in sports betting. Specifically, we implement and evaluate the Eigen model, a latent space model that represents dyadic data using node-specific vectors whose inner product govern edge probabilities. By modeling relationships between teams as adjacency matrices of binary outcomes, we investigate the extent to which the Eigen model captures both homophily, the tendency of similar teams to yield consistent betting results, and stochastic equivalence, where different teams exhibit indistinguishable patterns of Over/Under outcomes. A Bayesian formulation of the …
Comparison Of Statistical And Machine Learning Genomic Prediction Methods In Plant Breeding: Case Studies In Maize And Soybean, Igor Kuivjogi Fernandes
Comparison Of Statistical And Machine Learning Genomic Prediction Methods In Plant Breeding: Case Studies In Maize And Soybean, Igor Kuivjogi Fernandes
Graduate Theses and Dissertations
Plant breeding is essential to increase genetic gain and food production worldwide. This study was conducted to evaluate new ways to use machine learning (ML) to tackle plant breeding challenges, where two ideas were tested — the first chapter focuses on how to combine genetic and environmental data using ML to improve the prediction of maize grain yield in multi-environment trials, while the second chapter centers on how to couple feature selection of molecular markers with ML to enhance prediction of yield in soybean, and, in both cases, ML approaches were compared to well-established statistical methods greatly adopted by the …
Gan With Skip Patch Discriminator For Biological Electron Microscopy Image Generation, Nishith Ranjon Roy
Gan With Skip Patch Discriminator For Biological Electron Microscopy Image Generation, Nishith Ranjon Roy
Graduate Theses and Dissertations
GAN models have been successfully used for image generation in various sections such as real-life objects like human faces, cars, animal faces, landscapes, etc. This work focuses on biological electron microscopy (EM) image generation. Unlike other real-life objects, biological EM images are obtained through electron microscopy techniques to study biological specimens. Electron microscopy offers high resolution and magnification capabilities, making it a powerful tool for visualizing biological structures at the nanoscale. However, using GAN models for biological EM image generation poses challenges due to the complex and unique arrangements of biological structures and the sparse and asymmetrical patterns in EM …
Sparse Neural Network To Enhance Performance Under Limited Parameter Constraints., Nailah Rawnaq
Sparse Neural Network To Enhance Performance Under Limited Parameter Constraints., Nailah Rawnaq
Graduate Theses and Dissertations
Over the past decade, the widespread adoption of deep neural networks has been a breakthrough driven by significant computational advancements. Additionally, the number of parameters of those models is exponentially increasing for performing complex tasks and achieving better performance. However, in most practical cases, often there are constraints in the number of parameters due to limited resources in storage size and computational cost. Network pruning can lead to an optimal solution to this problem. In this thesis, I present supporting evidence to the hypothesis that higher sparsity leads to better performance for a convolution-based neural network. I perform performance studies …
Effects Of Measurement Error In Student Pre-Post Test Score On The Recovery Of The Estimates Of Teachers Value-Added Scores, Merlin J. Kamgue
Effects Of Measurement Error In Student Pre-Post Test Score On The Recovery Of The Estimates Of Teachers Value-Added Scores, Merlin J. Kamgue
Graduate Theses and Dissertations
Abstract Background: Value-added models (VAMs) are statistical tools used to gauge a teacher’s impact on student performance by analyzing standardized test scores. These models project students’ future performance based on past scores and compare the projection to actual outcomes, accounting for differences in student backgrounds. However, the standard error of measurement (SEM) inherent in all measurement tools is often overlooked in VAMs. Aims and Objectives: This study aims to investigate the impact of test reliability on teacher and school score estimates within a Bayesian framework. We will precisely manipulate the reliability of standardized tests by adjusting the standard error of …
Hierarchical Spatial Abundance Models For Migratory Shorebirds, Md Shahbaz Alam
Hierarchical Spatial Abundance Models For Migratory Shorebirds, Md Shahbaz Alam
Graduate Theses and Dissertations
Predicting the distribution and abundance of migratory shorebirds is crucial for effective conservation planning. This research applies hierarchical spatial models to predict counts and spatial variations of three shorebird species: Semipalmated sandpiper (sesa), Ruddy turnstone (rutu), and Whimbrel (whim). Different versions of the Poisson, Negative Binomial, and Hurdle regression models are employed to tackle specific data characteristics, such as overdispersion and excess zeros. Model comparisons are performed in terms of likelihood measures and cross-validation. The Hurdle model for sesa and rutu and the Negative Binomial model for whim effectively captured spatial patterns, highlighting potential hotspots. Mean predictive count further emphasized …
A Two-Part Validation Study Of The Sexual Minority Identity Emotion Scale, Henrietta Kadi Tettey-Tawiah
A Two-Part Validation Study Of The Sexual Minority Identity Emotion Scale, Henrietta Kadi Tettey-Tawiah
Graduate Theses and Dissertations
A number of existing studies indicate that there is some correlation between pride and shame and various behavioral health outcomes like anxiety, depression, and self-harm in the general population. Research also confirms that this relationship is true among sexual minorities as well. This study sets itself up to interrogate this assertion vis-à-vis sexual minority adolescents (SMAs). Consequently, confirmatory factor analysis (CFA) and differential item functioning (DIF) detection analysis were performed on the 35-item Sexual Minority Identity Emotion Scale (SMIES) for SMAs. A tree-based modeling technique, specifically conditional inference trees (CIT), was then employed to explore and make predictions about the …
The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi
The Quantitative Analysis And Visualization Of Nfl Passing Routes, Sandeep Chitturi
Computer Science and Computer Engineering Undergraduate Honors Theses
The strategic planning of offensive passing plays in the NFL incorporates numerous variables, including defensive coverages, player positioning, historical data, etc. This project develops an application using an analytical framework and an interactive model to simulate and visualize an NFL offense's passing strategy under varying conditions. Using R-programming and data management, the model dynamically represents potential passing routes in response to different defensive schemes. The system architecture integrates data from historical NFL league years to generate quantified route scores through designed mathematical equations. This allows for the prediction of potential passing routes for offensive skill players in response to the …
Comparing North American Professional Sports League Season Formats Using Monte Carlo Simulation, Lathan Gregg
Comparing North American Professional Sports League Season Formats Using Monte Carlo Simulation, Lathan Gregg
Industrial Engineering Undergraduate Honors Theses
Each NFL, NBA, and MLB season consists of a regular season, in which teams play a set number of scheduled games and a playoff, in which qualifying teams compete for a championship. At the conclusion of each season, teams are ranked based on their performance throughout the season. This study aims to investigate the ability of each league's season format to accurately rank teams using Monte Carlo simulation. Matches between two teams are simulated by using the team’s assigned strength ranks to calculate a winning probability for each team. The winning probabilities are simulated with different skill values, dictating how …
Automatic Appraisals Of Houses, Sloan Scroggin
Automatic Appraisals Of Houses, Sloan Scroggin
Graduate Theses and Dissertations
Multiple hedonic models and an automatic appraiser model were used to create a residential house’s estimated sales price. The goal is to use the limited data available to a REALTOR® to estimate the future sales price of a residential home without the aid of pictures of the property or viewing the physical property. The first model automates some of the actions of an appraiser by finding comparable sales based on proximity, based both on distance between houses and characteristics of the houses, and then calculating a weighted average price for an estimated sales price of future sales. If the model …
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Spatiotemporal Negative Inventory Outlier Decomposition For Supply Chain Applications In Consumer-Packaged Goods (Cpg), Hayden Mcdonald
Data Science Undergraduate Honors Theses
Coca-Cola is a popular soft drink brand with sales occurring in every Walmart store across the world, which generates large quantities of data and requires a robust supply chain system. However, the company does not currently have a sophisticated, automated, and/or prescriptive system for detecting where, when, and why inventory outages occur and applying preventative measures to avoid loss of revenue from the absence of inventory on store shelves. This thesis proposes and applies a novel, prescriptive system for this purpose. An inventory outage can be seen as a ‘negative’ statistical outlier in a time series of inventory for an …
Examining Award Compliance To Inform Resource Allocation, Jacob Haarala
Examining Award Compliance To Inform Resource Allocation, Jacob Haarala
Data Science Undergraduate Honors Theses
This project focuses on JB Hunt Transport Inc's intermodal business unit (JBI) by focusing on the challenges associated with Published Pricing and Contractual Pricing. The primary issue revolves around the variance between the awarded freight volumes in Requests for Pricing (RFPs) and the actual volumes realized when the freight is shipped. This discrepancy poses challenges for effective sales planning, revenue goals, and optimal freight network management within JBI. Reporting tools, such as PowerBI, are currently used by JBI to provide insights into award compliance on a weekly basis. However, our goal with this project was to provide a deeper understanding …
A Nonparametric Test For Comparing Survival Functions Based On Restricted Distance Correlation, Quinyang Zhang
A Nonparametric Test For Comparing Survival Functions Based On Restricted Distance Correlation, Quinyang Zhang
Mathematical Sciences Faculty Publications and Presentations
In this article, we propose an omnibus test for comparing two survival functions under non-proportional hazards. The test statistic is based on a product-limit estimate of the restricted distance correlation, which is closely related to the L2 distance between survival curves. The strong consistency is established under mild regularity conditions. Our simulation studies show that the new test has satisfactory power under proportional hazard and various non-proportional hazards settings including delayed treatment effect, diminishing effect, and crossing survival curves; therefore, it can be a competitive alternative to the existing omnibus tests such as Kolmogorov-Smirnov test, Cramer-von Mises test, two-stage …
Bayesian Learning Of Spatiotemporal Source Distribution For Beached Microplastic In The Gulf Of Mexico, David Pojunas
Bayesian Learning Of Spatiotemporal Source Distribution For Beached Microplastic In The Gulf Of Mexico, David Pojunas
Graduate Theses and Dissertations
Over the last several decades, plastic waste has gradually accumulated while slowly degrading in terrestrial and oceanic environments. Recently, there has been an increased effort to identify the possible sources of plastic to understand how they affect vulnerable beaches. This issue is of particular concern in the Gulf of Mexico due to the presence of oil, natural gas, and plastic production. In this thesis, we expand upon existing Bayesian plastic attribution models and develop a rigorous statistical framework to map observed beached microplastics to their sources. Within this framework, we combine Lagrangian backtracking simulations of floating particles using nurdle beaching …
Comparative Analysis Of Teacher Effects Parameters In Models Used For Assessing School Effectiveness: Value-Added Models & Persistence, Merlin J. Kamgue
Comparative Analysis Of Teacher Effects Parameters In Models Used For Assessing School Effectiveness: Value-Added Models & Persistence, Merlin J. Kamgue
Graduate Theses and Dissertations
Longitudinal measures for students have become increasingly popular to estimate the effects of individual teachers and schools. Value-added models are one of the approaches using longitudinal data to evaluate teachers and schools. In the value-added model (VAM) literature, many statistical approaches have been developed and used to estimate teacher or school effects on student learning. This study opted to use a Bayesian multivariate model for evaluating teacher effects. The generalized persistence models can handle longitudinal data, not vertically scaled, allowing for a below-par teacher’s effects correlation across test administrations. This study first generated longitudinal students’ test score data and used …
Analyses Of Effect Indices Across Single-Case Research Designs In Counseling, Cian L. Brown
Analyses Of Effect Indices Across Single-Case Research Designs In Counseling, Cian L. Brown
Graduate Theses and Dissertations
Single case research design (SCRD) is a common methodology used across clinical disciplines to determine treatments effectiveness by comparing treatment conditions to baseline conditions in individual cases, usually among researchers working with smaller samples. Although popular within behavioral disciplines such as special education and behavioral analysis, studies have begun to emerge in counseling. However, guidance and current understanding of the use of SCRD in counseling is limited. A content analysis of counseling journals from 2003 to 2014 yielded only 7 studies using SCRD. In 2015, the flagship counseling journal, Journal of Counseling and Development, published a special issue on the …
Comparing Predictive Performance Of Garch And Stochastic Volatility Models, Swapnaneel Nath
Comparing Predictive Performance Of Garch And Stochastic Volatility Models, Swapnaneel Nath
Graduate Theses and Dissertations
This paper compares the predictive performance of two commonly used financial models, the Generalized Auto-Regressive Conditional Heteroskedasticity (GARCH) model, and the Stochastic Volatility model. Both techniques are used in the finance literature to model returns on an asset; the main difference between the two is that the former holds volatility as deterministic, whereas the latter treats it as a stochastic component. Three 10-year periods (2006-15, 2008-17, and 2010-19) of returns of the S&P-500 Index are used to train the two models. The parameter estimation is done using Hamiltonian Monte Carlo. Then, using Sequential Monte Carlo updates, returns for 2016, 2018, …
A Comparative Study Of Techniques For Non-Monotonic Dependence With Emphasis On Sensitivity To Sample Size, Noise Level And Computational Attributes, Fariha Tasnim
Graduate Theses and Dissertations
Evaluating association between variables is often of interest by many researchers. To serve this purpose, different association measures have been developed. However, type of relation between variables affects the degree of relationship. Hence, detection of the rela- tionship between variables is germane to measuring the correlation coefficient. With that mindset, here we explored six non-monotonic measure of association techniques and com- pared them with three classical approaches. Due to inconsistency in definition and range of different techniques, it is not feasible to compare the correlation estimates as their nature of variability differ. Therefore, we used permutation test based on Monte …
The Appropriateness Of Outlier Exclusion Approaches Depends On The Expected Contamination: Commentary On André (2022), Daniel Villanova
The Appropriateness Of Outlier Exclusion Approaches Depends On The Expected Contamination: Commentary On André (2022), Daniel Villanova
Marketing Faculty Publications and Presentations
In a recent article, André (2022) addressed the decision to exclude outliers using a threshold across conditions or within conditions and offered a clear recommendation to avoid within-conditions exclusions because of the possibility for large false-positive inflation. In this commentary, I note that André’s simulations did not include the situation for which within-conditions exclusion has previously been recommended—when across-conditions exclusion would exacerbate selection bias. Examining test performance in this situation confirms the recommendation for within-conditions exclusion in such a circumstance. Critically, the suitability of exclusion criteria must be considered in relationship to assumptions about data-generating mechanisms.