Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

12,812 Full-Text Articles 23,893 Authors 9,922,835 Downloads 282 Institutions

All Articles in Statistics and Probability

Faceted Search

12,812 full-text articles. Page 102 of 486.

Single-Index Multinomial Model For Analyzing Crime Data, Kwabena Gyamfi Duodu 2023 University of Texas at El Paso

Single-Index Multinomial Model For Analyzing Crime Data, Kwabena Gyamfi Duodu

Open Access Theses & Dissertations

We develop a flexible single-index multinomial model for analyzing crime data. In additionto the number of crimes reported, the data also includes covariates such as location, time of day, weather, and other demographic factors. We provide an estimation algorithm and develop R code for the single-index multinomial model. Using simulations, we evaluate the performance of the proposed estimation algorithm. When applied to crime data, the single-index multinomial model provides important insights into crime trends and risk variables, assisting in the development of tailored crime prevention programs. Policymakers and law enforcement organizations can use the model's projections to more efficiently allocate …


Robust Penalized Density Power Divergence Regression With Scad Penalty For High Dimensional Data Analysis, Maxwell Kwesi Mac-Ocloo 2023 University of Texas at El Paso

Robust Penalized Density Power Divergence Regression With Scad Penalty For High Dimensional Data Analysis, Maxwell Kwesi Mac-Ocloo

Open Access Theses & Dissertations

Amidst the exponential surge in big data, managing high-dimensional datasets across diverse fields and industries has emerged as a significant challenge. Conventional statistical methods struggle to handle their complexity, making analysis intricate. In response, we've formulated a robust estimator tailored to counter outliers and heavy-tailed errors. Our approach integrates the SCAD penalty into the Density Power Divergence method, effectively reducing insignificant coefficients to zero. This enhances analysis precision and result reliability.We benchmark our robust and penalized model against existing techniques like Huber, Tukey, LASSO, LAD, and LAD-LASSO. Employing both simulated and UCI machine learning repository datasets, we assess method performance …


Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu 2023 University of Texas at El Paso

Comparative Study Of Supervised Classification Techniques With A Modified Knn Algorithm, Noah Owusu

Open Access Theses & Dissertations

The goal of classification is to develop a model that can be used to accurately assign new observations to labeled classes based on the patterns learned from the training data. K-nearest Neighbors algorithm (KNN) is a popular and widely used algorithm for classification, however, its performance can be adversely affected by the presence of outliers in a dataset. In this study we have modified this existing KNN algorithm that can alleviate the effect of outliers in a dataset, thereby improving the performance of the KNN algorithm. We compared the performances of the Modified KNN method and the Existing KNN algorithm …


Robust Mahalanobis K-Means Algorithm In Comparison With Other Existing Clustering Methods., Eleazer Tabi Serebour 2023 University of Texas at El Paso

Robust Mahalanobis K-Means Algorithm In Comparison With Other Existing Clustering Methods., Eleazer Tabi Serebour

Open Access Theses & Dissertations

This study enhances K-means Mahalanobis clustering using Density Power Divergence (DPD) for outlier handling and detection. Through the utilization of simulations and the analysis of real-world data, our approach consistently outperforms standard K-means, Mahalanobis K-means, Fuzzy C-means, and others in clustering datasets with outliers. While our method performs similarly to others on spherical datasets, it ranks second to DBSCAN for arbitrary shapes. We showcase its superiority on real-life datasets (Iris flower and wheat seed), demonstrating resilient outlier identification. By navigating various structures and cluster characteristics, our Modified Mahalanobis K-means method proves adaptable and robust, offering insights into diverse clustering scenarios. …


Weighted Mean Difference Statistics For Paired Data In The Presence Of Missing Values, Yuntong Li, Brent J. Shelton, William St Clair, Heidi L. Weiss, John L. Villano, Arnold Stromberg, Chi Wang, Li Chen 2023 University of Kentucky

Weighted Mean Difference Statistics For Paired Data In The Presence Of Missing Values, Yuntong Li, Brent J. Shelton, William St Clair, Heidi L. Weiss, John L. Villano, Arnold Stromberg, Chi Wang, Li Chen

Markey Cancer Center Faculty Publications

Missing data is a common issue in many biomedical studies. Under a paired design, some subjects may have missing values in either one or both of the conditions due to loss of follow-up, insufficient biological samples, etc. Such partially paired data complicate statistical comparison of the distribution of the variable of interest between the two conditions. In this article, we propose a general class of test statistics based on the difference in weighted sample means without imposing any distributional or model assumption. An optimal weight is derived from this class of tests. Simulation studies show that our proposed test with …


Geometric Morphometric Analysis Of Modern Viperid Vertebrae Facilitates Identification Of Fossil Specimens, Lance D. Jessee 2023 East Tennessee State University

Geometric Morphometric Analysis Of Modern Viperid Vertebrae Facilitates Identification Of Fossil Specimens, Lance D. Jessee

Electronic Theses and Dissertations

Snake vertebrae are common in the fossil record, whereas cranial remains are generally fragile and rare. Consequently, vertebrae are the most commonly studied fossil element of snakes. However, identification of snake vertebrae can be problematic due to extensive variation. This study utilizes 2-D geometric morphometrics and canonical variates analysis to 1) reveal variation between genera and species and 2) classify vertebrae of modern and fossil eastern North American Agkistrodon and Crotalus. The results show that vertebrae of Agkistrodon and Crotalus can reliably be classified to genus and species using these methods. Based on the statistical analyses, four of the …


Statistical Modeling Approaches For The Inference Of Cancer Mechanisms, Licai Huang 2023 The University of Texas MD Anderson Cancer Center UTHealth Graduate School of Biomedical Sciences

Statistical Modeling Approaches For The Inference Of Cancer Mechanisms, Licai Huang

Dissertations and Theses (Open Access)

The aim of this study was to explore the potential of integrating multi-platform genomic datasets to improve our understanding of the biological mechanisms behind cancer. By merging clinical outcomes with the data obtained from multi-platform genomic studies, we can gain insight into the biological mechanisms behind a patient’s response to treatment. Additionally, the evaluation of the correlations between genetic variations and gene expression provides a better understanding of the functional significance of these variations. Such knowledge has the potential to revolutionize cancer diagnosis and treatment. This thesis describes methods developed to address two related aims. Aim 1: We have developed …


Statistical Approaches To Estimate Bidirectional And Time-Varying Causal Effects Using Mendelian Randomization, Jinhao Zou 2023 The University of Texas MD Anderson Cancer Center UTHealth Houston Graduate School of Biomedical Sciences

Statistical Approaches To Estimate Bidirectional And Time-Varying Causal Effects Using Mendelian Randomization, Jinhao Zou

Dissertations and Theses (Open Access)

Mendelian Randomization (MR) is an epidemiological framework using genetic variants as instrumental variables (IVs) to examine the causal effect of an exposure on an outcome. It is widely used to detect causal factors of diseases and provide insight into the biological pathway of diseases. Current methods under the MR framework are built to estimate the unidirectional causal effects of exposures on outcomes and neglect the potential bidirectional causal effects. However, a bidirectional causal effect creates a feedback loop that biases the casual inference in MR studies. Furthermore, current MR methods estimate the causal effect as a single value using cross-sectional …


Development Of A Metapgs For Accurate Prediction Of Osteoporotic Fracture, Xiangxue Xiao 2023 University of Nevada, Las Vegas

Development Of A Metapgs For Accurate Prediction Of Osteoporotic Fracture, Xiangxue Xiao

UNLV Theses, Dissertations, Professional Papers, and Capstones

Introduction: Early identification of individuals at high-risk for osteoporotic fractures who may benefit from preventive intervention is essential. However, the predictive accuracy of the currently used fracture risk assessment tool remains suboptimal. The first aim of this research is to construct genome-wide polygenic scores for the femoral neck (PGS_FNBMDidpred) and total body BMD (PGS_TBBMDidpred) and to estimate their potential in identifying individuals with a high risk of osteoporotic fractures. The second aim is to validate the predictive performance of two previously established PGSs (PGS_FNBMDidpred and PGS_TBBMDidpred) in an external cohort …


Stressor: An R Package For Benchmarking Machine Learning Models, Samuel A. Haycock 2023 Utah State University

Stressor: An R Package For Benchmarking Machine Learning Models, Samuel A. Haycock

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Many discipline specific researchers need a way to quickly compare the accuracy of their predictive models to other alternatives. However, many of these researchers are not experienced with multiple programming languages. Python has recently been the leader in machine learning functionality, which includes the PyCaret library that allows users to develop high-performing machine learning models with only a few lines of code. The goal of the stressor package is to help users of the R programming language access the advantages of PyCaret without having to learn Python. This allows the user to leverage R’s powerful data analysis workflows, while simultaneously …


Statistical Graph Quality Analysis Of Utah State University Master Of Science Thesis Reports, Ragan Astle 2023 Utah State University

Statistical Graph Quality Analysis Of Utah State University Master Of Science Thesis Reports, Ragan Astle

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Graphical software packages have become increasingly popular in our modern world, but there are concerns within the statistical visualization field about the default settings provided by these packages, which can make it challenging to create good quality graphs that align with standard graph principles. In this thesis, we investigate whether the quality of graphs from Utah State University (USU) Plan A Master of Science (MS) thesis reports from the years 1930 to 2019 was affected by the rise of graphical software packages. We collected all data stored on the USU Digital Commons website since November 2021 to determine the specific …


Using Natural Language Processing To Quantify The Efficacy Of Language Simplification As A Communication Strategy, Brian Nalley 2023 Utah State University

Using Natural Language Processing To Quantify The Efficacy Of Language Simplification As A Communication Strategy, Brian Nalley

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

People with communication disorders often experience difficulties being understood by unfamiliar listeners or in noisy environments. A common strategy for effectively communicating in these scenarios is to use simpler and more predictable language. Despite the prevalence of this strategy, there has been little to no research to date focused on the effectiveness of language simplification as a communication strategy. This study seeks to begin filling that gap by using natural language processing to determine whether speakers with early-stage Parkinson’s disease and age-matched neurotypical speakers are able to successfully simplify their language while still maintaining the original message.

Simplification was measured …


An Interval-Valued Random Forests, Paul Gaona Partida 2023 Utah State University

An Interval-Valued Random Forests, Paul Gaona Partida

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

There is a growing demand for the development of new statistical models and the refinement of established methods to accommodate different data structures. This need arises from the recognition that traditional statistics often assume the value of each observation to be precise, which may not hold true in many real-world scenarios. Factors such as the collection process and technological advancements can introduce imprecision and uncertainty into the data.

For example, consider data collected over a long period of time, where newer measurement tools may offer greater accuracy and provide more information than previous methods. In such cases, it becomes crucial …


Copula Based Models For Bivariate Zero-Inflated Count Time Series Data, Dimuthu Fernando 2023 Old Dominion University

Copula Based Models For Bivariate Zero-Inflated Count Time Series Data, Dimuthu Fernando

Mathematics & Statistics Theses & Dissertations

Count time series data have multiple applications. The applications can be found in areas of finance, climate, public health and crime data analyses. In most scenarios, time is an important part of the data. Time series counts then come as multivariate vectors that exhibit not only serial dependence within each time series but also with cross-correlation among the series. When considering these observed counts, and when a value, say zero, occurs more often than usual, analysis presents crucial challenges. There is presence of zeroinflation in the data. The literature on bivariate or multivariate count time series, as well as zero-inflated …


Comparing Predictive Performance Of Garch And Stochastic Volatility Models, Swapnaneel Nath 2023 University of Arkansas, Fayetteville

Comparing Predictive Performance Of Garch And Stochastic Volatility Models, Swapnaneel Nath

Graduate Theses and Dissertations

This paper compares the predictive performance of two commonly used financial models, the Generalized Auto-Regressive Conditional Heteroskedasticity (GARCH) model, and the Stochastic Volatility model. Both techniques are used in the finance literature to model returns on an asset; the main difference between the two is that the former holds volatility as deterministic, whereas the latter treats it as a stochastic component. Three 10-year periods (2006-15, 2008-17, and 2010-19) of returns of the S&P-500 Index are used to train the two models. The parameter estimation is done using Hamiltonian Monte Carlo. Then, using Sequential Monte Carlo updates, returns for 2016, 2018, …


A Comparative Study Of Techniques For Non-Monotonic Dependence With Emphasis On Sensitivity To Sample Size, Noise Level And Computational Attributes, Fariha Tasnim 2023 University of Arkansas, Fayetteville

A Comparative Study Of Techniques For Non-Monotonic Dependence With Emphasis On Sensitivity To Sample Size, Noise Level And Computational Attributes, Fariha Tasnim

Graduate Theses and Dissertations

Evaluating association between variables is often of interest by many researchers. To serve this purpose, different association measures have been developed. However, type of relation between variables affects the degree of relationship. Hence, detection of the rela- tionship between variables is germane to measuring the correlation coefficient. With that mindset, here we explored six non-monotonic measure of association techniques and com- pared them with three classical approaches. Due to inconsistency in definition and range of different techniques, it is not feasible to compare the correlation estimates as their nature of variability differ. Therefore, we used permutation test based on Monte …


Editorial, Al Asyary 2023 Department of Environmental Health Faculty of Public Health Universitas Indonesia

Editorial, Al Asyary

Kesmas

No abstract provided.


Sentiment Analysis Before And During The Covid-19 Pandemic, Emily Musgrove 2023 Ursinus College

Sentiment Analysis Before And During The Covid-19 Pandemic, Emily Musgrove

Mathematics Summer Fellows

This study examines the change in connotative language use before and during the Covid-19 pandemic. By analyzing news articles from several major US newspapers, we found that there is a statistically significant correlation between the sentiment of the text and the publication period. Specifically, we document a large, systematic, and statistically significant decline in the overall sentiment of articles published in major news outlets. While our results do not directly gauge the sentiment of the population, our findings have important implications regarding the social responsibility of journalists and media outlets especially in times of crisis.


Correlation Analysis Between External Damages And Transverse Load Distribution Capacit Y Of Hollow Slab Bridges, CAO Hao, SUN Dunhua, WANG Zichen, ZHAO Fuli, SUN Haipeng, XIONG Wen 2023 1. Anhui Communications Holding Group Co., Ltd., Hefei, Anhui 230088, China;

Correlation Analysis Between External Damages And Transverse Load Distribution Capacit Y Of Hollow Slab Bridges, Cao Hao, Sun Dunhua, Wang Zichen, Zhao Fuli, Sun Haipeng, Xiong Wen

Journal of China & Foreign Highway

To correctly evaluate the service status of prefabricated hollow slab girder bridge, based on the measured data of apparent disease and load test of hollow slab girder bridge in reconstruction and expansion project, this paper first analyzed the data characteristics of the beam body, hinge joints, and support disease samples and obtained their statistical distribution rules. Then, according to the typical diseases (transverse cracks and hinge cracks at the bottom of the beam), the BP neural network was used to build a correlation model for them, and the close correlation between the two typical diseases of hollow slab girder bridge …


Research Of Disease Analysis And Reinforcement Technology For Typical Orthotropic Bridge Deck, GUO Fukuan, ZHOU Shangmeng 2023 1. State Key Laboratory of Health and Safety of Bridge Structures, Wuhan, Hubei 430034, China; 2. China Railway Bridge Research Institute Co., Ltd., Wuhan, Hubei 430034, China

Research Of Disease Analysis And Reinforcement Technology For Typical Orthotropic Bridge Deck, Guo Fukuan, Zhou Shangmeng

Journal of China & Foreign Highway

In order to effectively solve the disease problem of orthotropic steel bridge decks, this paper relied on a steel box girder cable-stayed bridge deck reinforcement project to carry out steel bridge deck disease analysis and reinforcement technology research. The results show that the main disease of steel bridge decks is fatigue cracking. The fatigue cracks are numerous, short, and deep, and they continuously grow and cause more severe disease. The excessive fatigue stress amplitude is the main factor leading to the fatigue cracking of steel bridge decks. Once the crack occurs, the stress amplitude of the crack tip increases continuously, …


Digital Commons powered by bepress