Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Open Access Theses & Dissertations

Discipline
Keyword
Publication Year

Articles 31 - 60 of 121

Full-Text Articles in Statistics and Probability

Outlier Detection In Multivariate And High-Dimensional Datasets, Yuanhong Wu May 2023

Outlier Detection In Multivariate And High-Dimensional Datasets, Yuanhong Wu

Open Access Theses & Dissertations

Accurate detection of outliers is crucial in the field of statistical analysis. Using classical statisticalmodels without considering the presence of outliers in the data can lead to misleading outcomes. There exist a myriad of procedures to detect outliers in statistics. We concentrate on the statistical techniques that can robustly identify outliers in data sets. To this end, we pursue two aims. First, we give an extensive overview of robust statistical methods which are still popular in recent years for outlier detection. We provide the definitions, algorithms and also discuss some important properties of these methods. Second, two real examples are …


Spatially Adaptive Estimation Of Spectrum, Yi Xie May 2023

Spatially Adaptive Estimation Of Spectrum, Yi Xie

Open Access Theses & Dissertations

A time series may be analyzed either in the time or in the frequency domain. When working in the frequency domain, the main objective is to estimate the underlying spectrum. Various approaches have been proposed to this end, but most are based on smoothing the periodogram using a single smoothing parameter across all Fourier frequencies. Such a global smoothing parameter may result in a biased estimate. To improve the estimation, in this paper, we smooth the log periodogram by placing a dynamic shrinkage prior, such that varying degrees of smoothing may be applied to different regions of the Fourier frequencies, …


Flexible Models For The Estimation Of Treatment Effect, Habeeb Abolaji Bashir May 2023

Flexible Models For The Estimation Of Treatment Effect, Habeeb Abolaji Bashir

Open Access Theses & Dissertations

Estimation of treatment effect is an important problem which is well studied in the literature. While the regression models are one of the most commonly used techniques for the estimation of treatment effect, they are prone to model misspecification. To minimize the model misspecification bias, flexible nonparametric models are introduced for the estimation. Continuing this line of research, we propose two flexible nonparametric models that allow the treatment effect to vary across different levels of covariates. We provide estimation algorithms for both these models. Using simulations and data analysis, we illustrate the usefulness of the proposed methods.


Performance Classification Of Ornstein-Uhlenbeck-Type Models Using Fractal Analysis Of Time Series Data., Peter Kwadwo Asante May 2023

Performance Classification Of Ornstein-Uhlenbeck-Type Models Using Fractal Analysis Of Time Series Data., Peter Kwadwo Asante

Open Access Theses & Dissertations

This dissertation aims to assess the performance of Ornstein-Uhlenbeck-type models by examining the fractal characteristics of time series data from various sources, including finance, volcanic and earthquake events, US COVID-19 reported cases and deaths, and two simulated time series with differing properties. The time series data is categorized as either a Gaussian or a Lévy process (Lévy walk or Lévy flight) by using three scaling methods: Rescaled range analysis, Detrended fluctuation analysis, and Diffusion entropy analysis. The outcomes of this analysis indicate that the financial indices are classified as Lévy walks, while the volcanic, earthquake, and COVID-19 data are classified …


Nonparametric Estimation Of Elliptical Copulas, Panfeng Liang May 2023

Nonparametric Estimation Of Elliptical Copulas, Panfeng Liang

Open Access Theses & Dissertations

Elliptical copulas provide flexibility in modeling the dependence structure of a random vector. They are often parameterized with a correlation matrix and a scalar function, called generator. The estimation of the generator can be challenging, because it is a functional parameter. In this dissertation, we provide a rigorous approach to estimating the generator in a Bayesian framework, which is simpler, more robust, and outperforms existing estimation methods in the literature. Based on the proposed framework in this dissertation, other researchers may modify the model for other types of generators in their own research.


Developing A Risk Assessment Instrument For Immigration Cases Under Federal Supervision, Mayra Eydie Pacheco May 2023

Developing A Risk Assessment Instrument For Immigration Cases Under Federal Supervision, Mayra Eydie Pacheco

Open Access Theses & Dissertations

No abstract provided.


Evaluation Of Effect Of Preprocessing Algorithms On Resting State Fmri Data, Hortencia Josefina Hernandez Dec 2022

Evaluation Of Effect Of Preprocessing Algorithms On Resting State Fmri Data, Hortencia Josefina Hernandez

Open Access Theses & Dissertations

Graph theory modeling is a common modeling approach in neurobiology research studies. These models are useful since they describe patterns of connection for regions of interest in the brain using resting state fMRI images. The standard rule of thumb is to threshold the observed activation levels prior to model building. It is reasonable to assume that the use of this threshold affects the statistical distribution of commonly reported centrality metrics from the graph theory model, such as degree, betweenness, and closeness. In this study we examine the differential effect of using the standard approaches versus alternative direct thresholds and incorporation …


A Computationally Efficient Wald Test In M-Estimation, Denisse Urenda Castañeda Aug 2022

A Computationally Efficient Wald Test In M-Estimation, Denisse Urenda Castañeda

Open Access Theses & Dissertations

Under the maximum likelihood framework, three asymptotic overall tests have been well developed in generalized linear models (GLM) for testing the single null hypothesis H0 : θ = θ0, namely, the Wald test, Likelihood Ratio Test (LRT) and Score test also known as the Lagrange Multiplier test (LM). Modified versions of Wald, LR and LM tests can also be found for testing the significance of a portion of the parameter θ, i.e., if θ = (θ T 1 , θ T 2 ) T it is of interest to test H0 : θ2 = 0. However, with the constant increase …


Efficient Approaches To Steady State Detection In Multivariate Systems, Honglun Xu Aug 2022

Efficient Approaches To Steady State Detection In Multivariate Systems, Honglun Xu

Open Access Theses & Dissertations

Steady state detection is critically important in many engineering fields such as fault detection and diagnosis, process monitoring and control. However, most of the existing methods are designed for univariate signals. In this dissertation, we proposed an efficient online steady state detection method for multivariate systems through a sequential Bayesian partitioning approach. The signal is modeled by a Bayesian piecewise constant mean and covariance model, and a recursive updating method is developed to calculate the posterior distributions analytically. The duration of the current segment is utilized to test the steady state. Insightful guidance is provided for hyperparameter selection. The effectiveness …


A Machine Learning Approach To Stochastic Optimal Control, Pablo Ever Avalos May 2022

A Machine Learning Approach To Stochastic Optimal Control, Pablo Ever Avalos

Open Access Theses & Dissertations

Merton's portfolio optimization problem is a well-renowned problem in financial mathematics which seeks to optimize the investment decision for an investor. In the simplest situation, the market consists of a risk-less asset (i.e. a bond) that pays back a relatively low interest rate, and a risky asset (i.e. a stock) that follows a geometric Brownian motion. The optimal allocation strategy of the investor's wealth is found by optimizing the expected utility along the stochastic evolution of the market. This thesis focuses on several different applications of this optimization problem. We look at pre-constructed analytical solutions and showcase the results. We …


Developing And Applying Computational Algorithms To Reveal Health-Related Biomolecular Interactions, Yixin Xie May 2022

Developing And Applying Computational Algorithms To Reveal Health-Related Biomolecular Interactions, Yixin Xie

Open Access Theses & Dissertations

Computational biology is an interdisciplinary area that applies computational approaches in biological big data, including protein amino acid sequences, genetic sequences, etc., which is widely used to analyze protein-protein interactions, make predictions in drug discovery, develop vaccines, etc. Popular methods include mathematical modeling, molecular dynamics simulations, data science mythology, etc. With the help of computational algorithms and applications, drug development is much faster than traditional processes, as it reduces risks early on in a drug discovery process and helps researchers select target candidates that have the highest potential for success. In my doctoral research, I applied multi-scale computational approaches to …


A New Algorithm For Robust Affine-Invariant Clustering, Andrews Tawiah Anum Dec 2021

A New Algorithm For Robust Affine-Invariant Clustering, Andrews Tawiah Anum

Open Access Theses & Dissertations

Cluster analysis is an unsupervised machine learning technique commonly employed to partition a dataset into distinct categories referred to as clusters. The k-means algorithm is a prominent distance-based clustering method. Despite its overwhelming popularity, the algorithm is not invariant under non-singular linear transformations and is not robust, i.e., can be unduly influenced by outliers. To address these deficiencies, we propose an alternative clustering procedure based on minimizing a “trimmed” variant of the negative log-likelihood function. We develop a “concentration step”, vaguely reminiscent of the classical Lloyd’s algorithm, that can iteratively reduce the objective function. Multiple real and synthetic datasets are …


The Physiological Factors Of Diabetes And Their Effect On The Cognitive And Emotional Functioning In Older Populations: A Secondary Data Analysis, Celeste Anahi Alvidrez Dec 2021

The Physiological Factors Of Diabetes And Their Effect On The Cognitive And Emotional Functioning In Older Populations: A Secondary Data Analysis, Celeste Anahi Alvidrez

Open Access Theses & Dissertations

Background: The rates of Type 2 Diabetes (T2D) have increased over the past 20 years in all age groups. The physiological factors that underlie T2D could have impact on specific brain pathways that support cognitive and emotional functioning. Aims and Objective: The goal of this study was to examine whether older Mexican American individuals with a history of T2D were more likely to develop later cognitive impairment and/or depression. Hypotheses: It was predicted that elderly participants (mean age at time of interview = 87.87 years) with a history of T2D onset prior to age 65, are more likely to have …


Statistical Analysis Of Genetic Sequence Variants In Whole Exome Sequencing Data From Patients With Prostate Cancer, Kelvin Ofori-Minta Aug 2021

Statistical Analysis Of Genetic Sequence Variants In Whole Exome Sequencing Data From Patients With Prostate Cancer, Kelvin Ofori-Minta

Open Access Theses & Dissertations

A single variation in the genetic sequence within the DNA of an organism could easily lead to beneficial, detrimental or neutral effects. Most often than not, these effects are detrimental than beneficial. While many biomedical and bioinformatics studies have been conducted to determine the genetic cause of prostate cancer (PrCa) which is still the second leading cause of cancer related death among men in the United States. An appreciable effort in statistical bioinformatics researches has been directed towards this aim. Through statistical analyses of a set of whole exome sequencing data from patients with PrCa obtained via The Cancer Genome …


The Hybridizing Ions Treatment (Hit) Method Development And Computational Study On Sars-Cov-2 E Protein., Shengjie Sun May 2021

The Hybridizing Ions Treatment (Hit) Method Development And Computational Study On Sars-Cov-2 E Protein., Shengjie Sun

Open Access Theses & Dissertations

Fast and accurate calculations of the electrostatic features for highly charged biomolecules such as DNA, RNA, highly charged proteins, are crucial but challenging tasks. Traditional implicit solvent methods calculate the electrostatic features fast, but they are not able to balance the high net charges in the biomolecules effectively. Explicit solvent methods add unbalanced ions to neutralize the highly charged biomolecules in molecular dynamic simulations, which require more expensive computing resources. Here we developed a novel method, the Hybridizing Ions Treatment (HIT) method, which hybridizes the implicit solvent method with the explicit method to realistically calculate the electrostatic potential for highly …


Robust Variable Selection In Multiple Linear Regression Via Penalized Least Trimmed Squares., Reagan Kesseku May 2021

Robust Variable Selection In Multiple Linear Regression Via Penalized Least Trimmed Squares., Reagan Kesseku

Open Access Theses & Dissertations

Variable selection has been studied using different approaches. Its growing importance lies in numerous applications to high-dimensional data from experiments and natural phenomena. Often, models are to be constructed from such data based on significant variables for estimation or prediction purposes. This demands not just any variable selectionmethod, but one that is robust, computationally efficient and with other desirable statistical properties. Besides the high-dimensionality of such data, the presence of outliers is common due to heterogeneous sources. Though outliers often contain useful information, they can unduly influence non-robust estimators to produce misleading results. This is the case for ordinary least …


High-Dimensional Random Forests, Roland Fiagbe May 2021

High-Dimensional Random Forests, Roland Fiagbe

Open Access Theses & Dissertations

The significant advances in technology have enabled easy collection and management of high-dimensional data in many fields, however, the process of modeling these data imposes a huge problem in the field of data science. Dealing with high-dimensional data is one of the significant challenges that degenerate the performance and precision of most classification and regression algorithms, e.g., random forests. Random Forest (RF) is among the few methods that can be extended to model high-dimensional data; nevertheless, its performance and precision, like others, are highly affected by high dimensions, especially when the dataset contains a huge number of noise or noninformative …


Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil May 2021

Gene Selection And Classification In High-Throughput Biological Data With Integrated Machine Learning Algorithms And Bioinformatics Approaches, Abhijeet R Patil

Open Access Theses & Dissertations

With the rise of high throughput technologies in biomedical research, large volumes of expression profiling, methylation profiling, and RNA-sequencing data are being generated. These high-dimensional data have large number of features with small number of samples, a characteristic called the "curse of dimensionality." The selection of optimal features, which largely affects the performance of classification algorithms in machine learning models, has led to challenging problems in bioinformatics analyses of such high-dimensional datasets. In this work, I focus on the design of two-stage frameworks of feature selection and classification and their applications in multiple sets of colorectal cancer data. The first …


Making Valid Inferences With Decision Tree, George Ekow Quaye May 2021

Making Valid Inferences With Decision Tree, George Ekow Quaye

Open Access Theses & Dissertations

HypoThesis testing and Confidence Interval (CI) estimates are key statistics in predicting future values in data analysis. Most often, CI estimates are directly obtained from the summary statistics of a particular statistical methodology output. However, when it comes to the summary of decision tree outputs, these CI estimates are not directly obtained. So a na\"{i}ve way of making node-level inference is to construct a $(1-\alpha) \times 100\%$ confidence interval for a node mean $\bar{y}_t$ using the relation: $\bar{y}_t \, \pm \, z_{1-\alpha/2} \, \frac{s_t}{\sqrt{n_t}}$, where $\bar{y}_t$ is the node mean and $s_t$ is the standard deviation estimates from the decision …


Refined Moderation Analysis With Binary Outcomes, Eric Anto May 2021

Refined Moderation Analysis With Binary Outcomes, Eric Anto

Open Access Theses & Dissertations

With the growing interest in personalized or precision medicine, it is indispensable thatmoderation analysis which is primarily related to the study of differential treatment effects among patients with different characteristics, also serves as the bedrock for precision medicine is taken more seriously. Concerning moderation analysis with binary outcomes, we start with an interesting observation, which shows that heterogeneous treatment effects could be equivalently estimated via a role exchange between the outcome and the treatment variable. The result holds for both experimental data and observational data, yet with an important difference in interpretation. Two estimators of moderating effects corresponding to two …


Lévy Processes: Characterizing Volcanic And Financial Time Series, Peter Kwadwo Asante Jan 2020

Lévy Processes: Characterizing Volcanic And Financial Time Series, Peter Kwadwo Asante

Open Access Theses & Dissertations

In this work, we use the Diffusion Entropy Analysis (DEA) to analyze and detect the scaling properties of time series from both emerging and well established markets as well as volcanic eruptions recorded by a seismic station, both financial and volcanic time series data are known to have high frequencies (i.e they are collected at an extremely fine scale). The objective is to determine the characterization i.e whether they follow a Gaussian or Lévy distribution. If they do follow a Lévy distribution we are then interested in finding if they are characterized by a Lévy walk which has a finite …


Predicting Stochastic Volatility For Extreme Fluctuations In High Frequency Time Series, Md Al Masum Bhuiyan Jan 2020

Predicting Stochastic Volatility For Extreme Fluctuations In High Frequency Time Series, Md Al Masum Bhuiyan

Open Access Theses & Dissertations

This work is devoted to the study of modeling high frequency time series including extreme fluctuations. As the high frequency data are collected at extremely fine scales, the fluctuations can capture the dynamics of data that evolve over time. A class of volatility models with time-varying parameters is used to forecast the volatility in a stationary condition at different lags. The modeling of stationary time series with consistent properties facilitates prediction with much certainty.

A large set of high frequency financial returns, closing prices of stock markets, high magnitudes of seismograms generated by the natural earthquakes, and the mining explosions …


Applications Of Ornstein-Uhlenbeck Type Stochastic Differential Equations, Osei Kofi Tweneboah Jan 2020

Applications Of Ornstein-Uhlenbeck Type Stochastic Differential Equations, Osei Kofi Tweneboah

Open Access Theses & Dissertations

In this Dissertation, we show with plausible arguments that the Stochastic Differential Equations (SDEs) arising on the superposition and coupling system of independent Ornstein-Uhlenbeck process is a new method available in modern literature that takes the properties and behavior of the data into consideration when performing the statistical analysis of the time series.

The time series to be analyzed is thought of as a source of fluctuations, and thus we need a model that takes this behavior into consideration when performing such analysis. Most of the standard methods fail to take into account the physical behavior of the time series, …


Robust Estimation And Inference For Multivariate Financial Data, Afua Kwakyewaa Amoako Dadey Jan 2020

Robust Estimation And Inference For Multivariate Financial Data, Afua Kwakyewaa Amoako Dadey

Open Access Theses & Dissertations

Predicting and forecasting are routine day-to-day activities that guide us in making the best possible choices. They play an integral role in financial analysis. A lot of work has been done on one dimensional geometric Brownian motion (GBM) in stock price prediction. In this line of work, we focus mainly on how to use the one dimensional geometric Brownian motion and the multidimensional geometric Brownian motion in predicting future stock prices. There are several stock prices in the financial market and the multidimensional geometric Brownian motion gives a more realistic prediction compared to the one dimensional GBM. The reason being …


General Penalized Logistic Regression For Gene Selection In High-Dimensional Microarray Data Classification, Derrick Kwesi Bonney Jan 2020

General Penalized Logistic Regression For Gene Selection In High-Dimensional Microarray Data Classification, Derrick Kwesi Bonney

Open Access Theses & Dissertations

High-dimensional data has become a major research area in the field of genetics, bioinformatics and bio-statistics due to advancement of technologies. Some common issues of modeling high-dimensional gene expression data are that many of the genes may not be relevant. Also, reducing the dimensions of the data using penalized logistic regression is one of the major challenges when there exists a high correlation among genes. High-dimension data correspond to the situation where the number of variables is greater or larger than the number of observations. Gene selection proved to be an effective way to improve the results of many classification …


Stochastic Modeling Of Earthquakes And Option Pricing Using Bns-Gamma-Ou Model, Mandela Bright Quashie Jan 2020

Stochastic Modeling Of Earthquakes And Option Pricing Using Bns-Gamma-Ou Model, Mandela Bright Quashie

Open Access Theses & Dissertations

High frequency data are becoming increasingly popular these days. They are fundamental in basically every facet of people’s lives. They are the determining factors in hedging in the field of finance. In geology, they help in the accurate prediction of earthquakes’ magnitude which goes along way to help save lives and properties.

High frequency data are also used more and more frequently for speculations. For this reason, it is important not only for scientists to apply models allowing correct quantification of these data, but also to improve the eciency of these models.

The Black-Scholes model, which is widely used because …


Spatially Adaptive Estimation Of Spectrum, Yi None Xie Jan 2020

Spatially Adaptive Estimation Of Spectrum, Yi None Xie

Open Access Theses & Dissertations

When analyzing a stationary time series, one of the questions we are often interested in is how to estimate its spectrum. Many approaches have been proposed to this end. Most are focused on smoothing the periodogram using a single smoothing parameter across all Fourier frequencies. In this paper, we smooth the log periodogram by placing a spatially adaptive prior called the dynamic shrinkage prior, so that varying degrees of smoothing may be applied to different intervals of Fourier frequencies, resulting in less biased estimates of the spectrum. Further research will extend this approach to spectral estimation for nonstationary time series.


Time-Reflective Text Representations For Semantic Evolution Tracking And Trend Analytics, Roberto Camacho Barranco Jan 2019

Time-Reflective Text Representations For Semantic Evolution Tracking And Trend Analytics, Roberto Camacho Barranco

Open Access Theses & Dissertations

The extraction of significant, relevant, and useful trends from massive document collections, such as a streaming newswire or scientific publications, is a challenging and significant problem in many different fields, including intelligence analysis, recommendation systems, and scientific research. However, techniques that tackle trend analytics of such large text corpora are limited because research that addresses the temporal nature of these publications is still in its early stages. In this work, we first show that it is possible to capture the evolution of a story (or trend) by connecting the dots between different documents in a text corpus. The observed results …


Combination Of Resampling Based Lasso Feature Selection And Ensembles Of Regularized Regression Models, Abhijeet R. Patil Jan 2019

Combination Of Resampling Based Lasso Feature Selection And Ensembles Of Regularized Regression Models, Abhijeet R. Patil

Open Access Theses & Dissertations

In high-dimensional data, the performance of various classiers is largely dependent on the selection of important features. Most of the individual classiers using existing feature selection (FS) methods do not perform well for highly correlated data. Obtaining important

features using the FS method and selecting the best performing classier is a challenging task in high throughput data. In this research, we propose a combination of resampling based least absolute shrinkage and selection operator (LASSO) feature selection (RLFS)

and ensembles of regularized regression models (ERRM) capable of handling data with the high correlation structures. The ERRM boosts the prediction accuracy with …


Modeling Correlated Data Via Copulas, Panfeng Liang Jan 2019

Modeling Correlated Data Via Copulas, Panfeng Liang

Open Access Theses & Dissertations

Copulas are widely used to model the dependency structure among components of multi- variate data sets. Elliptical copulas, such as Gaussian copula, are most popular copulas being used since many data sets follow elliptical distributions or meta-elliptical distribu- tions (Fang et al. (2002)). However, today's approaches and software packages require us to assume the specific category, such as Gaussian or Student's T, of the elliptical cop- ula before estimating it. In this Thesis, we will propose a Bayesian method using Markov chain Monte Carlo (MCMC) methods to estimate the density function of elliptical copulas without specifying it is the copula …