Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

University of Texas at El Paso

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 61 - 90 of 130

Full-Text Articles in Statistics and Probability

Modeling Correlated Data Via Copulas, Panfeng Liang Jan 2019

Modeling Correlated Data Via Copulas, Panfeng Liang

Open Access Theses & Dissertations

Copulas are widely used to model the dependency structure among components of multi- variate data sets. Elliptical copulas, such as Gaussian copula, are most popular copulas being used since many data sets follow elliptical distributions or meta-elliptical distribu- tions (Fang et al. (2002)). However, today's approaches and software packages require us to assume the specific category, such as Gaussian or Student's T, of the elliptical cop- ula before estimating it. In this Thesis, we will propose a Bayesian method using Markov chain Monte Carlo (MCMC) methods to estimate the density function of elliptical copulas without specifying it is the copula …


Bayesian Analysis Of Variable-Stress Accelerated Life Testing, Richard Okine Jan 2019

Bayesian Analysis Of Variable-Stress Accelerated Life Testing, Richard Okine

Open Access Theses & Dissertations

Several authors have over the years studied the art of modeling data from accelerated life testing and making inferences from such data. In this study, we consider a continuously varying stress accelerated life testing procedure which is the limiting case of the multiple stress-level discussed by Doksum and H´oyland [1]. We derive the likelihood function for the life distribution of the continuously increasing stress accelerated life testing model and consequently the Fisher's Information Matrix. We propose a Bayesian analysis for this distribution using the Gibbs Sampling Procedure. We conduct simulation studies and real data analysis to demonstrate the efficiency of …


Application Of Urinary Metabolites For Cancer Detection, Qin Gao Jan 2019

Application Of Urinary Metabolites For Cancer Detection, Qin Gao

Open Access Theses & Dissertations

Prostate cancer (PCa) is the 3rd most common cause of male cancer mortality in the US. Early diagnosis and treatment of PCa will improve the quality of care and reduce mortality. The prostate specific antigen (PSA) is commonly used in the current PCa screening, but its limitation has resulted in an intense search for more reliable biomarkers. Studies showed that dogs could differentiate PCa patients from negative control by sniffing their urine. As the odor profiles are generated by volatile organic compounds (VOCs), the finding suggests that particular VOCs could be linked to PCa, PCa risk levels and other cancers. …


Confidence Intervals For The Expected P-Value, Emmanuel Kofi Abrefa Jan 2019

Confidence Intervals For The Expected P-Value, Emmanuel Kofi Abrefa

Open Access Theses & Dissertations

The p-value is widely used in many application fields. In common practice, a scientific finding is deemed statistically significant if its resultant p-value is less than a pre-specified significance level, for example α = 0.05, albeit many statistically significant results are not reproducible in new studies. Mixed reasons including misuses, abuses, misunderstanding and misinterpretation arouse intensive debates and conservatives around the p-value from time to time over the years. Yet no reasonable solutions have been proposed. In this research, we make efforts to close the gap by advocating the use of confidence level for the expected p-value p0. This allows …


Robust Statistical Inference For The Gaussian Distribution, Andrews Tawiah Anum Jan 2019

Robust Statistical Inference For The Gaussian Distribution, Andrews Tawiah Anum

Open Access Theses & Dissertations

The aim of robust statistics is to develop statistical procedures which are not unduly influenced by outliers or observations that are not representative of the underlying "true" data generating process. This thesis focuses on an estimator with this characteristic. The divergence function is introduced in Chapter 2 with the sole aim of taking the function f to be the univariate normal distribution and α - [0, 1]. The estimator fails when we rely on the classic Newton's method to converge to the minimum of the density power divergence (MDPD) function. There is a tendency of such estimator never to approach …


Forecasting Crashes, Credit Card Default, And Imputation Analysis On Missing Values By The Use Of Neural Networks, Jazmin Quezada Jan 2019

Forecasting Crashes, Credit Card Default, And Imputation Analysis On Missing Values By The Use Of Neural Networks, Jazmin Quezada

Open Access Theses & Dissertations

A neural network is a system of hardware and/or software patterned after the operation of neurons in the human brain. Neural networks,- also called Artificial Neural Networks - are a variety of deep learning technology, which also falls under the umbrella of artificial intelligence, or AI. Recent studies shows that Artificial Neural Network has the highest coefficient of determination (i.e. measure to assess how well a model explains and predicts future outcomes.) in comparison to the K-nearest neighbor classifiers, logistic regression, discriminant analysis, naive Bayesian classifier, and classification trees. In this work, the theoretical description of the neural network methodology …


On The Performance Of Variable Selection And Classification Via Rank-Based Classifier, Md Showaib Rahman None Sarker Jan 2019

On The Performance Of Variable Selection And Classification Via Rank-Based Classifier, Md Showaib Rahman None Sarker

Open Access Theses & Dissertations

In high-dimensional gene expression data analysis, the accuracy and reliability of cancer classification and selection of important genes play a very crucial role. To identify these important genes and predict future outcomes (tumor vs. non-tumor), various methods have been proposed in the literature. But only few of them take into account correlation patterns and grouping effects among the genes. In this article, we propose a rank-based modification of the popular penalized logistic regression procedure based on a combination of l1 and l2 penalties capable of handling possible correlation among genes in different groups. While the l1 penalty maintains sparsity, the …


A Bayesian Model For Spectral Density Estimation, Yi Xie Jan 2018

A Bayesian Model For Spectral Density Estimation, Yi Xie

Open Access Theses & Dissertations

When we analyze a stationary time series, one of the questions we often meet is how to estimate its spectral density. Many approaches have been proposed to this end. In this paper we estimate the spectral density of a stationary time series nonparametrically. We fit a nonparametric regression model to the log periodogram and use third-degree B-spline functions as basis functions. Since the the number of basis functions is relatively large, we place priors such as random-walk and regularized horseshoe on the coefficients of the basis functions to avoid over-fitting and smooth the log periodogram.


Integrated Statistical And Machine Learning Algorithms For Predicting And Classifying G Protein-Coupled Receptors, Fredrick Ayivor Jan 2018

Integrated Statistical And Machine Learning Algorithms For Predicting And Classifying G Protein-Coupled Receptors, Fredrick Ayivor

Open Access Theses & Dissertations

G protein-coupled receptors (GPCRs) are transmembrane proteins with important functions in signal transduction and often serve as drug targets. With increasing availability of protein sequence information, there is much interest in computationally predicting GPCRs and classifying them according to their biological roles. Such predictions are cost-efficient and can be valuable guides for designing wet lab experiments to help elucidate signaling pathways and expedite drug discovery. There are existing computational tools of GPCR prediction that involve principal component analysis (PCA), intimate sorting (IS), support vector machine, and random forest (RF) techniques using various sequence derived features. While accuracies of over 90\% …


An Efficient Method For Online Identification Of Steady State For Multivariate System, Honglun None Xu Jan 2018

An Efficient Method For Online Identification Of Steady State For Multivariate System, Honglun None Xu

Open Access Theses & Dissertations

Most of the existing steady state detection approaches are designed for univariate signals. For multivariate signals, the univariate approach is often applied to each process variable and the system is claimed to be steady once all signals are steady, which is computationally inefficient and also not accurate. The article proposes an efficient online method for multivariate steady state detection. It estimates the covariance matrices using two different approaches, namely, the mean-squared-deviation and mean-squared-successive-difference. To avoid the usage of a moving window, the process means and the two covariance matrices are calculated recursively through exponentially weighted moving average. A likelihood ratio …


Matroid - Based Variable Selection For Complex Data Structures, Wimarsha Thathsarani Jayanetti Jan 2018

Matroid - Based Variable Selection For Complex Data Structures, Wimarsha Thathsarani Jayanetti

Open Access Theses & Dissertations

This research project has the objective to extend use of the matroid algorithm using statistically based criteria, Joint/Multivariate Cumulants (Speed, 1983) and Effective Dependence (Pena & Rodriguez, 2003) to capture linear as well as non-linear higher order dependencies. We also improve variable selection for complex data structures using the proposed matroid algorithm. The limiting distribution of the joint cumulant was defined using U-statistics theory by Hoeffding (1948). U-statistics variance as theorized by Hoeffding provide a lower bound for the estimated variance, and our simulation results justify the use of Hoeffding U-statistic variance for determining a threshold for joint cumulants deviation …


Using Data Mining To Model Student Achievement On The 4th Grade Timss 2015 Mathematics Assessment: A Five Nation Sudy, Annette M. Siemssen Jan 2018

Using Data Mining To Model Student Achievement On The 4th Grade Timss 2015 Mathematics Assessment: A Five Nation Sudy, Annette M. Siemssen

Open Access Theses & Dissertations

Data mining has been successfully used by financial and retail companies since the mid-1960's to create predictive models and reveal unexpected relationships. However, it remains underutilized as a tool in educational research. Large-scale standardized assessment programs such as the Trends in International Mathematics and Science Study (TIMSS) provide vast amounts of data with the potential for providing new insights in education. Five nations, the Republic of Korea, the United States, Germany, Kuwait, and Kazakhstan were selected based on General Response Style theory to represent a spectrum of cultural backgrounds, from acquiescent to midpoint to individualistic (Hastedt, D. & van de …


Estimating The Optimal Cutoff Point For Logistic Regression, Zheng Zhang Jan 2018

Estimating The Optimal Cutoff Point For Logistic Regression, Zheng Zhang

Open Access Theses & Dissertations

Binary classification is one of the main themes of supervised learning. This research is concerned about determining the optimal cutoff point for the continuous-scaled outcomes (e.g., predicted probabilities) resulting from a classifier such as logistic regression. We make note of the fact that the cutoff point obtained from various methods is a statistic, which can be unstable with substantial variation. Nevertheless, due partly to complexity involved in estimating the cutpoint, there has been no formal study on the variance or standard error of the estimated cutoff point.

In this Thesis, a bootstrap aggregation method is put forward to estimate the …


Hierarchical Multiplicity Control Methods For Linear Models, Dimuthu Dilshan Fernando Jan 2018

Hierarchical Multiplicity Control Methods For Linear Models, Dimuthu Dilshan Fernando

Open Access Theses & Dissertations

HypoThesis testing is a commonly used statistical inference technique on which a statement of the population is investigated through the evidence from a representative sample of the population. With simultaneous testing of more than one null hypotheses need for an appropriate multiple comparison method is essential. With motivation from the study of Bogomolov et al. (2017) we have modified a multiple comparison tree structure to build the required comparisons and focus on controlling the FWER (Family Wise Error Rate) using the Bonferroni procedure. The proposed method has advantages such as controlling the global error rates separately at each level, families …


Backward Elimination Algorithm For High Dimensional Variable Screening, Sophia Korkor Foli Jan 2018

Backward Elimination Algorithm For High Dimensional Variable Screening, Sophia Korkor Foli

Open Access Theses & Dissertations

In recent times, variable selection in high-dimensional data has become a challenging prob- lem. We investigate here a popular but classical variable screening method, the Back- ward Elimination (BE) in a high dimensional setup (small-n-large P). The BE method as a variable screening method reduces the dimension of small-n-large P data into a lower dimensional data and then established shrinkage methods such as: LASSO, SCAD and MCP can be applied directly. To overcome the problems in high dimensional data, Chen and Chen (2008) recently developed a family of Extended Bayesian Information Criterion (EBIC) which is consistent with finite sample properties …


Extraction Of Fiber Morphology From Sem Images For Quality Control Of Fiber Reinforced Composites Manufacturing, Md Fashiar Rahman Jan 2018

Extraction Of Fiber Morphology From Sem Images For Quality Control Of Fiber Reinforced Composites Manufacturing, Md Fashiar Rahman

Open Access Theses & Dissertations

The morphology of fibers (e.g. spatial uniformity, orientation, and length) plays a decisive role in determining the material properties or fabrication quality of fiber-reinforced nanocomposites. Hence, determining the morphology becomes a very critical issue in the field of nanocomposite quality control. The conventional way of quality inspection is to take the scanning electron microscopic (SEM) images of the cross-section of composite material and do the visual checking of these SEM images to evaluate the nanofiber alignment and length distribution. But this type of inspection is often subjective, inaccurate and time consuming. Moreover, the extremely small size of nanofibers makes the …


Propagation Of Probabilistic Uncertainty: The Simplest Case (A Brief Pedagogical Introduction), Olga Kosheleva, Vladik Kreinovich Nov 2017

Propagation Of Probabilistic Uncertainty: The Simplest Case (A Brief Pedagogical Introduction), Olga Kosheleva, Vladik Kreinovich

Departmental Technical Reports (CS)

The main objective of this text is to provide a brief introduction to formulas describing the simplest case of propagation of probabilistic uncertainty -- for students who have not yet taken a probability course.


Is It Legitimate Statistics Or Is It Sexism: Why Discrimination Is Not Rational, Martha Osegueda Escobar, Vladik Kreinovich, Thach N. Nguyen Aug 2017

Is It Legitimate Statistics Or Is It Sexism: Why Discrimination Is Not Rational, Martha Osegueda Escobar, Vladik Kreinovich, Thach N. Nguyen

Departmental Technical Reports (CS)

While in the ideal world, everyone should have the same chance to succeed in a given profession, in reality, often the probability of success is different for people of different gender and/or ethnicity. For example, in the US, the probability of a female undergraduate student in computer science to get a PhD is lower than a similar probability for a male student. At first glance, it may seem that in such a situation, if we try to maximize our gain and we have a limited amount of resources, it is reasonable to concentrate on students with the higher probability of …


How To Estimate Statistical Characteristics Based On A Sample: Nonparametric Maximum Likelihood Approach Leads To Sample Mean, Sample Variance, Etc., Vladik Kreinovich, Thongchai Dumrongpokaphan Jun 2017

How To Estimate Statistical Characteristics Based On A Sample: Nonparametric Maximum Likelihood Approach Leads To Sample Mean, Sample Variance, Etc., Vladik Kreinovich, Thongchai Dumrongpokaphan

Departmental Technical Reports (CS)

In many practical situations, we need to estimate different statistical characteristics based on a sample. In some cases, we know that the corresponding probability distribution belongs to a known finite-parametric family of distributions. In such cases, a reasonable idea is to use the Maximum Likelihood method to estimate the corresponding parameters, and then to compute the value of the desired statistical characteristic for the distribution with these parameters.

In some practical situations, we do not know any family containing the unknown distribution. We show that in such nonparametric cases, the Maximum Likelihood approach leads to the use of sample mean, …


The Nonparametric Estimation Of Elliptical Distributions, Panfeng Liang Jan 2017

The Nonparametric Estimation Of Elliptical Distributions, Panfeng Liang

Open Access Theses & Dissertations

In practice, many multivariate datasets have identical marginal distributions. Elliptical distributions can be used to model many of those datasets. In this Thesis, we will propose a Bayesian method using Markov chain Monte Carlo (MCMC) methods to estimate the density function underlying multivariate datasets assuming it is an elliptical distribution.


Predicting Individualized Treatment Effects Via Random Forests Of Interaction Trees, Annette Pena Franco Jan 2017

Predicting Individualized Treatment Effects Via Random Forests Of Interaction Trees, Annette Pena Franco

Open Access Theses & Dissertations

Abstract Not Available


Analysis Of Bias-Corrected And Exact Estimators For Binomial Generalized Linear Model Parameters, Hamna Hannan Jan 2017

Analysis Of Bias-Corrected And Exact Estimators For Binomial Generalized Linear Model Parameters, Hamna Hannan

Open Access Theses & Dissertations

Typically, small samples have always been a problem for binomial generalized linear models. Though generalized linear models are widely popular in public health, social sciences etc. In small sample scenarios the non-existence of the maximum likelihood (ML) estimators is very common as well as separation occurs in the data. In logistic regression the maximum likelihood estimates are found to have biased away from origin. My work examines the bias-reduced and exact estimators that have been used to estimate the slope parameters and standard errors of the estimated slope parameters as compared to the traditional ML method.

The present work is …


Sample Size Estimation For Linear Mixed Models With Dependent End Points, Michael Nsiah-Nimo Jan 2017

Sample Size Estimation For Linear Mixed Models With Dependent End Points, Michael Nsiah-Nimo

Open Access Theses & Dissertations

The primary objective is sample size estimation in linear mixed model settings. Sample size estimation is an important component of planning a well thought out scientific experiment. Whenever sample size estimation is performed, taking into account a priori model based inferences will provide a sample size estimate that will achieve the desired power without inflating the type I error rate of the study.

One common practice is a traditional approach cited in the literature that uses the largest sample size after you Bonferroni the type I error rate to estimate sample sizes as such. We are going to take into …


Evaluating Binary Splits On Nominal Inputs, Isaac Xoese Ocloo Jan 2017

Evaluating Binary Splits On Nominal Inputs, Isaac Xoese Ocloo

Open Access Theses & Dissertations

The maximally selected statistic approach in building tree models is shown to be a cause of variable selection bias. In this study we propose three methods to solve this problem in building regression trees with nominal predictor variables. Out of the three methods

proposed we explored only two in detail and defer one for further research. We developed an exact method to compute the p-value corresponding to the maximized splitting statistic in regression trees for nominal predictor variables with at most 10 distinct levels and a

method to estimate the best cutoff point as a parameter in a parametric nonlinear …


Estimating The Coefficients Of A Linear Differential Operator, Maria Ivette Barraza Jan 2017

Estimating The Coefficients Of A Linear Differential Operator, Maria Ivette Barraza

Open Access Theses & Dissertations

Principal Differential Analysis (PDA; Ramsay, 1996) is used to obtain low dimensional representations of functional data, where each observation is represented as a curve. PDA seeks to identify a Linear Differential Operator (LDO) L = ω0I + ω 1D + ... + ωmDm, where I denotes the identity function and D j the jth derivative, that satisfies as closely as possible that Lx = 0 for each functional observation x. A theorem from analysis establishes that the coefficients of the LDO are in the Sobolev space, and thus can be approximated by B-splines. Current PDA software used to estimate the …


A New Test For The Mean Vector In High Dimensional Setting, Behzad Aalipur Hafshejani Jan 2016

A New Test For The Mean Vector In High Dimensional Setting, Behzad Aalipur Hafshejani

Open Access Theses & Dissertations

Traditional statistical data analysis mostly includes methods and techniques to deal with problems in which there are many observations but a few variables. Nonetheless, the current inclination is toward more observations but also, toward more variables. Today's observations gathered on individuals are images, curves, or even movies. Unfortunately many traditional methods do not work well in high dimensional settings. As an example Hotelling's test which is well known and widely used in the literature does not work when it comes to high dimensional problems. Consequently statisticians are making an effort to find remedies or new approaches to multivariate mean testing. …


Sample Size Estimation For Genomics Experiments With Dependent End Points, Desmond Koomson Jan 2016

Sample Size Estimation For Genomics Experiments With Dependent End Points, Desmond Koomson

Open Access Theses & Dissertations

In typical genomics studies involving numerous association tests of gene mutations with a disease, error rate control via multiplicity adjustment is paramount because even if all genes were to be non-differentially associated, we would still make some false positives. Many methods exist that incorporate the control of multiplicity for normally distributed endpoints in sample size estimation, but none addresses the issue for non-normally correlated endpoints.

One common practice in the literature is to assume an equal correlation among all differentially associated or expressed genes, thereby using the generalized binomial or beta-binomial model to compute the comparison-wise power of detecting these …


Bayesian Parameter Estimation For The Birnbaum-Saunders Distribution And Its Extension, Tun Lee Ng Jan 2016

Bayesian Parameter Estimation For The Birnbaum-Saunders Distribution And Its Extension, Tun Lee Ng

Open Access Theses & Dissertations

We utilize the Bayesian approach to estimate the parameters of the Birnbaum-Saunders (BS) distribution devised by Birnbaum and Saunders (1969a), as well as the Generalized Birnbaum-Saunders (GBS) distribution obtained by Owen (2006), in the presence of random right censored data. We also derive the classical MLE expressions for the observed Information matrix of the GBS distribution, in order to illustrate the fact that no closed form expressions are available for the MLE, and numerical approximations are required to obtain the point estimates and asymptotic confidence intervals. Where Bayesian approach is concerned, new sets of priors are considered based on the …


Development Of Efficient Simultaneous Confidence Bounds For Linear Mixed Models With Applications In Alcohol Research, Emmanuel Joseph Sequeira Jan 2016

Development Of Efficient Simultaneous Confidence Bounds For Linear Mixed Models With Applications In Alcohol Research, Emmanuel Joseph Sequeira

Open Access Theses & Dissertations

Multiplicity corrections are necessary to ensure the accuracy of conclusions made in studies that carry out multiple inferences simultaneously. This Thesis uses the methodology derived by Hunter and Worsley to obtain improved simultaneous confidence bounds (SCBs) that are less conservative than the highly used Bonferroni SCBs, for studies using linear mixed modeling. Empirical coverage rates were obtained for data that was generated using simulations, to compare the accuracy of the Hunter-Worsley SCBs with that of the Bonferroni SCBs. The bounds were also applied to data in the field of alcohol research, where comparisons were made to determine the moderating effect …


Pre-Tuned Principal Component Regression And Several Variants, Pei Wang Jan 2016

Pre-Tuned Principal Component Regression And Several Variants, Pei Wang

Open Access Theses & Dissertations

The regression coecient estimates from ordinary least squares (OLS) have a low probability of being close to the real value when there is a multicollinearity problem in the design matrix. In order to combat this problem, many regularized methods have been introduced. Principal components regression (PCR) is an important analysis tool for dealing with multicollinearity and high-dimensionality. In conventional PCR, the rst step is to change the original predictors to orthogonal principal components (PC's) by a linear transformation. These PC's correspond to the eigenvalues which are sorted in a decreasing order. The next step is to regress the response on …