Open Access. Powered by Scholars. Published by Universities.®

Statistical Theory Commons

Open Access. Powered by Scholars. Published by Universities.®

2024

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 18 of 18

Full-Text Articles in Statistical Theory

Learning Problems Related To Stochastic Differential Equations, Jinpu Zhou Nov 2024

Learning Problems Related To Stochastic Differential Equations, Jinpu Zhou

LSU Doctoral Dissertations

Stochastic differential equations (SDEs) are essential for modeling systems influenced by both deterministic dynamics and random fluctuations, with applications in a wide variety of disciplines. This thesis develops a Bayesian framework for nonparametric learning in SDEs, addressing key challenges in inference, particularly when dealing with complex systems and incomplete data. The thesis begins by establishing a theoretical foundation in optimization over Hilbert spaces, including a generalized representer theorem to address infinite-dimensional optimization problems encountered in nonparametric inference. Building on this, we introduce a Bayesian framework with shrinkage priors to learn drift functions from high-frequency data. Bayesian approach incorporates low-cost sparse …


A Uniformly Most Powerful Test For The Mean Of A Beta Distribution, Richard Ntiamoah Kyei Aug 2024

A Uniformly Most Powerful Test For The Mean Of A Beta Distribution, Richard Ntiamoah Kyei

Electronic Theses and Dissertations

The beta distribution is used in numerous real-world applications, including areas such as manufacturing (quality control) and analyzing patient outcomes in health care. It also plays a key role in statistical theory, including multivariate analysis of variance (MANOVA) and Bayesian statistics. It is a flexible distribution that can account for many different characteristics of real data. To our surprise, there has been very little work or discussion on performing statistical hypothesis testing for the mean when it is reasonable to assume that the population is beta distributed. Many analysts conduct traditional analyses using a t-test or nonparametric approach, try transformations, …


Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh Aug 2024

Simulation Study On Confidence Interval Estimation For Standard Deviation With Non-Normal Distributions, Theophilus Oppong Kyeremeh

Electronic Theses and Dissertations

This study explores innovative approaches to constructing confidence intervals for the population standard deviation, σ, in non-normal data scenarios. While the sample standard deviation, s, is widely used, its reliability is compromised when dealing with skewed or heavy-tailed distributions and exhibits sensitivity to outliers. Our research addresses these limitations by investigating alternative estimation methods that offer greater robustness and accuracy.


Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han Aug 2024

Bayesian Approaches In Multi-State Markov Models And High Dimensional Time-To-Event Data., Yuchen Han

Electronic Theses and Dissertations

This dissertation consists of two projects. The first one involves nonparametric methods on Continuous Time Markov Chains (CTMCs). The second one is centered around Bayesian shrinkage models for detecting prognostic and predictive biomarkers in high-dimensional clinical data. Both these projects build on methods from across the frequentist and Bayesian paradigm to offer novel solutions. In the first project, we aim to model the nonlinear effects of continuous variables within multistate framework in a non-parametrically by appealing to the rich mathematical framework of Reproducing Kernel Hilbert Spaces (RKHS). Then we adapted the classical Representer Theorem to penalized (squared norm) log-likelihood which …


Contributions To Nonparametric Testing In Clustered Data, Hasika Kalani Wickrama Senevirathne Jul 2024

Contributions To Nonparametric Testing In Clustered Data, Hasika Kalani Wickrama Senevirathne

Mathematics & Statistics Theses & Dissertations

Clustered data refers to a specific kind of correlated data where units within the same cluster are correlated while units from different clusters are independent. The number of units in each cluster, known as the cluster size, can be associated with the cluster’s outcome. This is known as the informative cluster size (ICS) and affects the inference drawn from clustered data. Recently, a hypothesis testing method has been developed to detect the presence of ICS. However, considering ICS alone may not be sufficient when comparing outcomes across multiple groups of units within clustered data. The size of a group within …


Comparison Of Value At Risk Using Historical And Monte Carlo Methods On Pt Xyz Stock Portofolio, Eka Fitriani, Yulial Hikmah, Ira Rosianal Hikmah Jun 2024

Comparison Of Value At Risk Using Historical And Monte Carlo Methods On Pt Xyz Stock Portofolio, Eka Fitriani, Yulial Hikmah, Ira Rosianal Hikmah

Jurnal Administrasi Bisnis Terapan

One way to achieve profits in a company is through investment activities. However, everything has risks. Investing can also be risky. Therefore, the relationship between risk and investment is important because it will influence the determination of investment selection. The problem faced by investors is choosing an efficient portfolio, or a portfolio that provides the smallest risk. This risk can be done by measuring risk, one of which is using the Value at Risk (VaR) measure. Measurement using Value at Risk has several methods that are quite popular, namely the Historical Method, Variance-Covariance, and Monte Carlo. In this research, the …


Comparing The Preparation Of Youth Services Librarians To Their On-The-Ground Experiences: A Grounded Theory Study Incorporating Criticism And Connoisseurship, Anne Holland Jun 2024

Comparing The Preparation Of Youth Services Librarians To Their On-The-Ground Experiences: A Grounded Theory Study Incorporating Criticism And Connoisseurship, Anne Holland

Electronic Theses and Dissertations

The purpose of this study was to better understand the on-the-ground preparation of youth services librarians, in contrast to their professional training in Master’s of Library Science (MLIS) programs. Classic Grounded Theory was the predominant methodology for this qualitative study, and elements of Criticism and Connoisseurship were also utilized. Document review, interviews, and journaling activities with ten participants were the primary methods of data collection. Key findings from this dissertation include a grounded theory explaining the current state of preparation for youth services librarianship, and multiple avenues for further study.


Trade Liberalization, Non-Oil Export And Economic Growth In Nigeria, Jerome T. Andohol, Terhemen Tarzoor, Dennis T. Nomor Jun 2024

Trade Liberalization, Non-Oil Export And Economic Growth In Nigeria, Jerome T. Andohol, Terhemen Tarzoor, Dennis T. Nomor

CBN Journal of Applied Statistics (JAS)

The study examines the impact of trade liberalization and non-oil exports on economic growth in Nigeria from 1986 to 2021. The study utilizes an autoregressive distributed lag model and found the combined effect of trade liberalization and non-oil exports to be positive and statistical significant. While trade liberalization alone may have negative consequences, its synergy with a robust non-oil export can drive sustainable economic growth. The study recommends that strategies to enhance non-oil exports should be encouraged to support the effectiveness of trade liberalization in promoting growth.


Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov Jun 2024

Stock Market Volatility In The United Kingdom: Simulating Post-Covid-19 Recovery, Bala A. Dahiru, Mohammed Shuaibu, Najibullah Hassanov

CBN Journal of Applied Statistics (JAS)

This paper investigates the time it would take for the FTSE-100 index to reach its post-COVID-19 peak. The paper utilises an exponential generalised autoregressive conditional heteroscedasticity (EGARCH) model that accounts for leverage effect and asymmetries. The preferred models amongst competing variants was the Autoregressive Moving Average (ARMA)-EGARCH(2,1) specification and was used to predict daily FTSE-100 data from 5th January 2000 to 21st June 2024. The empirical exercise showed that the COVID-19-induced financial crisis negatively affected the United Kingdom’s stock market performance. The results show that the FTSE100 index could reach its post-pandemic peak around 27th August, 2024 (two months after …


"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson Apr 2024

"Who Wrote The Epistle, God Only Knows": A Statistical Authorial Analysis Of Hebrews In Comparison With Pauline And Lukan Literature, Benjamin J. Erickson

Senior Honors Theses

The authorship of Hebrews has been a point of contention for scholars for the past two millennia. While the epistle is traditionally attributed to Paul, many scholars assert that it carries thematic, structural, and stylistic differences from the remainder of his extant epistles; therefore, many other possible authors have been proposed. Of these, only Luke has other New Testament writings. Therefore, this project conducts a statistical comparison of Hebrews to the Pauline and Lukan corpora using stylometric authorial analysis methods. This analysis demonstrates that Hebrews is stylistically closer to Lukan literature than Pauline (but not to a significant degree), and …


Uconn Baseball Reliever Lane Optimization Tool, Jason Bartholomew Apr 2024

Uconn Baseball Reliever Lane Optimization Tool, Jason Bartholomew

Honors Scholar Theses

The building of a tool to be utilized by UConn’s Division I baseball team that will generate a game plan for when different relievers should be used against different parts of the opponent’s lineup to achieve the lowest total expected value of runs allowed for the remainder of the game based on game situations and matchup probabilities. The tool will also examine and determine situations that may be vital enough to the outcome of the game to bring in a better reliever normally saved for later in the game.


Latent Classification Of Multi-Model Non-Homogeneous Continuous-Time Markov Chains, Joonha Chang Apr 2024

Latent Classification Of Multi-Model Non-Homogeneous Continuous-Time Markov Chains, Joonha Chang

Dissertations and Theses (Open Access)

The continuous-time Markov chain (CTMC) model and latent clustering models are commonly used to study longitudinal measures of categorical outcomes. Because of its simple but powerful Markovian property, CTMC models have been widely used in medical and public health researches. Due to limitations in the standard CTMC model, there have been some studies on non-homogeneous continuous-time Markov chain (NH-CTMC) that utilized time-dependent rates, but the progresses have been limited. NH-CTMC can be more powerful than CTMC by its default nature of time-dependent rate that can be fitted to a wider range of applications in medical studies. In this study, we …


Model Selection Through Cross-Validation For Supervised Learning Tasks With Manifold Data, Derek Brown Jan 2024

Model Selection Through Cross-Validation For Supervised Learning Tasks With Manifold Data, Derek Brown

The Journal of Purdue Undergraduate Research

No abstract provided.


Sensitivity Analysis Of Prior Distributions In Regression Model Estimation, Ayoade I Adewole, Oluwatoyin K. Bodunwa Jan 2024

Sensitivity Analysis Of Prior Distributions In Regression Model Estimation, Ayoade I Adewole, Oluwatoyin K. Bodunwa

Al-Bahir

Bayesian inferences depend solely on specification and accuracy of likelihoods and prior distributions of the observed data. The research delved into Bayesian estimation method of regression models to reduce the impact of some of the problems, posed by convectional method of estimating regression models, such as handling complex models, availability of small sample sizes and inclusion of background information in the estimation procedure. Posterior distributions are based on prior distributions and the data accuracy, which is the fundamental principles of Bayesian statistics to produce accurate final model estimates. Sensitivity analysis is an essential part of mathematical model validation in obtaining …


Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe Jan 2024

Predicting Superconducting Critical Temperature Using Regression Analysis, Roland Fiagbe

Data Science and Data Mining

This project estimates a regression model to predict the superconducting critical temperature based on variables extracted from the superconductor’s chemical formula. The regression model along with the stepwise variable selection gives a reasonable and good predictive model with a lower prediction error (MSE). Variables extracted based on atomic radius, valence, atomic mass and thermal conductivity appeared to have the most contribution to the predictive model.


Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe Jan 2024

Machine Learning Approaches For Cyberbullying Detection, Roland Fiagbe

Data Science and Data Mining

Cyberbullying refers to the act of bullying using electronic means and the internet. In recent years, this act has been identifed to be a major problem among young people and even adults. It can negatively impact one’s emotions and lead to adverse outcomes like depression, anxiety, harassment, and suicide, among others. This has led to the need to employ machine learning techniques to automatically detect cyberbullying and prevent them on various social media platforms. In this study, we want to analyze the combination of some Natural Language Processing (NLP) algorithms (such as Bag-of-Words and TFIDF) with some popular machine learning …


Accounting For Variability Due To Resampling Using Bootstrapping, Dipendra Phuyal Jan 2024

Accounting For Variability Due To Resampling Using Bootstrapping, Dipendra Phuyal

College of Graduate Studies: Theses & Dissertations

Bradley Efron (1979) introduced bootrapping. Typically a researcher is interested in studying a process which generates individuals. The collection of individuals the process has(actual) or could have (conceptual) generated is the population. The collection of conceptual members of the population is an uncountable collection. Hence, the population is anuncountable collection of individuals. The collection of individuals the process has generated (actual individuals) is representative of what the process can generate and will bereferred to as the representative sample. The size of this sample is a nonnegative integervalued random variable N which may be a constant random variable such as in …


The Distribution Of The Significance Level, Paul O. Monnu Jan 2024

The Distribution Of The Significance Level, Paul O. Monnu

College of Graduate Studies: Theses & Dissertations

Reporting the p-value is customary when conducting a test of hypothesis or significance. The likelihood of getting a fictitious second sample and presuming the null hypothesis is correct is the p-value. The significance level is a statistic that interests us to investigate. Being a statistic, it has a distribution. For the F-test in a one-way ANOVA and the t-tests for population means, we define the significance level, its observed value, and the observed significance level. It is possible to derive the significance level distribution. The t-test and the F-test are not without controversy. Specifically, we demonstrate that as sample size …