Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Statistical Models (20)
- Statistical Methodology (19)
- Biostatistics (10)
- Statistical Theory (6)
- Data Science (3)
-
- Multivariate Analysis (3)
- Categorical Data Analysis (1)
- Computational Biology (1)
- Engineering (1)
- Genetics and Genomics (1)
- Life Sciences (1)
- Longitudinal Data Analysis and Time Series (1)
- Microarrays (1)
- Other Engineering (1)
- Other Statistics and Probability (1)
- Survival Analysis (1)
- Vital and Health Statistics (1)
- Keyword
-
- EM algorithm (3)
- Missing data (2)
- Model selection (2)
- Tolerance interval (2)
- Variable Selection (2)
-
- Adjustment (1)
- Bayesian Analysis (1)
- Bayesian analysis (1)
- Bayesian modeling (1)
- Big Data (1)
- Biomarker (1)
- Bootstrap Calibration (1)
- Bootstrapping (1)
- Carcinogenesis (1)
- Categorical variable (1)
- Cause-specific hazard (1)
- Change point detection (1)
- Clustered data (1)
- Coefficient of partial determination (1)
- Competing event (1)
- Compound estimation (1)
- Confidence Interval (1)
- Confidence intervals (1)
- Contamination Beta mixture model (1)
- Continuous Phenotype (1)
- Continuous Trait (1)
- Convolutional Neural Network (1)
- Copy number variation (1)
- Count data (1)
- Criterion ABC (1)
Articles 1 - 30 of 32
Full-Text Articles in Applied Statistics
Variable Selection For High-Dimensional Data With Interaction Effects: Methods, Applications, And Inferences, Leiyue Li
Theses and Dissertations--Statistics
For high-dimensional data where the number of variables greatly exceeds the number of observations, selecting important variables while maintaining the required heredity conditions can be challenging. This dissertation is structured into three interconnected parts. In the first part, we propose a variable selection method by implementing a well-known optimization technique, the Genetic Algorithm. An R package was developed to simplify the implementation and usage of the proposed method. We then propose another variable selection method by extending the study from the Genetic Algorithm to a different but related optimization technique, Simulated Annealing. We consider three different hierarchical structures in both …
From Non-Parametric Methods To Self-Supervised Learning: Applications In Edge Detection And Image Denoising, Jiacheng Xu
From Non-Parametric Methods To Self-Supervised Learning: Applications In Edge Detection And Image Denoising, Jiacheng Xu
Theses and Dissertations--Statistics
This dissertation explores advanced methodologies for edge detection and image denoising through the application of both traditional non-parametric methods and modern self-supervised deep learning techniques. Beginning with non-parametric approaches, we refine surface fitting and jump detection criteria to enhance the detection of discontinuous regression surfaces in grayscale images. These foundational techniques are extended to color images, with analyses across RGB and CIELAB color spaces to improve edge detection accuracy. We then introduce a self-supervised neural network model that integrates Masked Modeling into the Bi-Directional Cascade Network (BDCN) framework. This approach shows the potential of reducing the dependency on annotated data …
Bar-Code Variable: A Novel Approach To Efficiently Find Interaction Effects, Lee Sak Park
Bar-Code Variable: A Novel Approach To Efficiently Find Interaction Effects, Lee Sak Park
Theses and Dissertations--Statistics
This paper introduces the bar-code variable, a novel method for processing a sequence of binary explanatory variables efficiently in the linear regression modeling framework. Represented as an integer or a sequence of bits, the bar-code variable captures infor- mation on original binary variables and their potential interaction effects. Utilizing the bar-code variable, the study explores streamlined feature selection in linear re- gression modeling with binary explanatory variables. The paper demonstrates how the bar-code variable, through re-parameterization, facilitates the transition from cell means estimates, µ̂, in the cell-means ANOVA model to coefficient estimates, β̂, in the linear regression model, and vice …
Statistical Tolerance Regions For Flexible Modeling Paradigms, Yafan Guo
Statistical Tolerance Regions For Flexible Modeling Paradigms, Yafan Guo
Theses and Dissertations--Statistics
Tolerance intervals in a regression setting allow the user to quantify, with a specified degree of confidence, bounds for a specified proportion of the sampled population when conditioned on a set of covariate values. While methods are available for tolerance intervals in fully-parametric regression settings, the construction of tolerance intervals for semiparametric regression models has been treated in a limited capacity. The first project fills this gap and develops likelihood-based approaches for the construction of pointwise one-sided and two-sided tolerance intervals for semiparametric regression models. A numerical approach is also presented for constructing simultaneous tolerance intervals. An appealing facet of …
Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan
Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan
Theses and Dissertations--Statistics
Neural networks have experienced widespread adoption and have become integral in cutting-edge domains like computer vision, natural language processing, and various contemporary fields. However, addressing the statistical aspects of neural networks has been a persistent challenge, with limited satisfactory results. In my research, I focused on exploring statistical intervals applied to neural networks, specifically confidence intervals and tolerance intervals. I employed variance estimation methods, such as direct estimation and resampling, to assess neural networks and their performance under outlier scenarios. Remarkably, when outliers were present, the resampling method with infinitesimal jackknife estimation yielded confidence intervals that closely aligned with nominal …
Probabilistic Methods For Inferring The Order Of Pathway Alterations During Carcinogenesis And Cancer Subtype Classification, Menghan Wang
Probabilistic Methods For Inferring The Order Of Pathway Alterations During Carcinogenesis And Cancer Subtype Classification, Menghan Wang
Theses and Dissertations--Statistics
Carcinogenesis is a complex process involving somatic mutations in a number of key biological pathways. Studying cancer evolution is an important task which contributes to better understanding of cancer biology and facilitates identification of new therapeutic targets. We focus on two important questions in cancer evolution. The first question is to delineating the temporal order of pathway mutations during tumorigenesis. And the other question is to cluster patients into biologically meaningful cancer subtypes. We present new statistical methods to 1)leverage functional annotations of mutations to enhance estimation of the order of pathway mutations during carcinogenesis, 2) incorporate intra-tumoral heterogeneity information …
Tolerance Intervals For Various Regression Models, Xitong Zhou
Tolerance Intervals For Various Regression Models, Xitong Zhou
Theses and Dissertations--Statistics
Among statistical intervals, confidence intervals and prediction intervals are well-known and commonly used. In many applications, the problem becomes finding an interval that covers at least a certain proportion $P$ of the population for a characteristic of interest with a specified confidence level $(1-\alpha)$. And such interval is named a $P$-content, $(1-\alpha)$-confidence Tolerance Interval (TI). The topic of the dissertation is the utility of tolerance intervals for various regression models. We begin with a discussion of tolerance intervals for linear and nonlinear regression models. We then propose a bootstrap method of constructing TIs for Tobit regression to deal with censored …
Novel Modelling And Inference Considerations Involving The Exponentially-Modified Gaussian Distribution, Yanxi Li
Theses and Dissertations--Statistics
The exponentially-modified Gaussian (EMG) distribution is well-suited for analyzing data with positive skewness due to its characteristic positive skew from the exponential component. Despite its popularity in various fields, the EMG distribution has only been analyzed for univariate data without any regression settings. To address this limitation, we developed a generalized EMG regression model with covariates by assigning parametric functional forms to some or all of the parameters in the EMG distribution that vary with values of the covariates. To further perform data-clustering on observation points, we propose a competing regression model where the error structure is assumed to be …
Methodologies And Computational Tools For Zero-Inflated Discrete Weibull Models, Peng Yeh
Methodologies And Computational Tools For Zero-Inflated Discrete Weibull Models, Peng Yeh
Theses and Dissertations--Statistics
Count data with excess zeros is common in many fields, such as ecology, healthcare, and insurance. Excess zeros data are often causing the inaccurate fit from the count models. While zero-inflated models have been developing for over two decades, one should also consider a more flexible model that can handle the excess zeros and further over- or under-dispersion. In this talk, we discuss zero-inflated discrete Weibull model and some novel computational contributions. The flexibility and competitiveness of the ZIDW model are illustrated by simulation studies and a real data analysis. We also investigate the performance of the proposed model through …
Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su
Theses and Dissertations--Statistics
When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …
Deriving The Distributions And Developing Methods Of Inference For R2-Type Measures, With Applications To Big Data Analysis, Gregory S. Hawk
Deriving The Distributions And Developing Methods Of Inference For R2-Type Measures, With Applications To Big Data Analysis, Gregory S. Hawk
Theses and Dissertations--Statistics
As computing capabilities and cloud-enhanced data sharing has accelerated exponentially in the 21st century, our access to Big Data has revolutionized the way we see data around the world, from healthcare to investments to manufacturing to retail and supply-chain. In many areas of research, however, the cost of obtaining each data point makes more than just a few observations impossible. While machine learning and artificial intelligence (AI) are improving our ability to make predictions from datasets, we need better statistical methods to improve our ability to understand and translate models into meaningful and actionable insights.
A central goal in the …
Beta Mixture And Contaminated Model With Constraints And Application With Micro-Array Data, Ya Qi
Beta Mixture And Contaminated Model With Constraints And Application With Micro-Array Data, Ya Qi
Theses and Dissertations--Statistics
This dissertation research is concentrated on the Contaminated Beta(CB) model and its application in micro-array data analysis. Modified Likelihood Ratio Test (MLRT) introduced by [Chen et al., 2001] is used for testing the omnibus null hypothesis of no contamination of Beta(1,1)([Dai and Charnigo, 2008]). We design constraints for two-component CB model, which put the mode toward the left end of the distribution to reflect the abundance of small p-values of micro-array data, to increase the test power. A three-component CB model might be useful when distinguishing high differentially expressed genes and moderate differentially expressed genes. If the null hypothesis above …
Dimension Reduction Techniques In Regression, Pei Wang
Dimension Reduction Techniques In Regression, Pei Wang
Theses and Dissertations--Statistics
Because of the advances of modern technology, the size of the collected data nowadays is larger and the structure is more complex. To deal with such kinds of data, sufficient dimension reduction (SDR) and reduced rank (RR) regression are two powerful tools. This dissertation focuses on these two tools and it is composed of three projects. In the first project, we introduce a new SDR method through a novel approach of feature filter to recover the central mean subspace exhaustively along with a method to determine the dimension, two variable selection methods, and extensions to multivariate response and large p …
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Theses and Dissertations--Statistics
Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …
Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui
Nonparametric Analysis Of Clustered And Multivariate Data, Yue Cui
Theses and Dissertations--Statistics
In this dissertation, we investigate three distinct but interrelated problems for nonparametric analysis of clustered data and multivariate data in pre-post factorial design.
In the first project, we propose a nonparametric approach for one-sample clustered data in pre-post intervention design. In particular, we consider the situation where for some clusters all members are only observed at either pre or post intervention but not both. This type of clustered data is referred to us as partially complete clustered data. Unlike most of its parametric counterparts, we do not assume specific models for data distributions, intra-cluster dependence structure or variability, in effect …
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Theses and Dissertations--Statistics
A common problem in regression analysis (linear or nonlinear) is assessing the lack-of-fit. Existing methods make parametric or semi-parametric assumptions to model the conditional mean or covariance matrices. In this dissertation, we propose fully nonparametric methods that make only additive error assumptions. Our nonparametric approach relies on ideas from nonparametric smoothing to reduce the test of association (lack-of-fit) problem into a nonparametric multivariate analysis of variance. A major problem that arises in this approach is that the key assumptions of independence and constant covariance matrix among the groups will be violated. As a result, the standard asymptotic theory is not …
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Theses and Dissertations--Statistics
Kinetic modeling of the time dependence of metabolite concentrations including the unstable isotope labeled species is an important approach to simulate metabolic pathway dynamics. It is also essential for quantitative metabolic flux analysis using tracer data. However, as the metabolic networks are complex including extensive compartmentation and interconnections, the parameter estimation for enzymes that catalyze individual reactions needed for kinetic modeling is challenging. As the pa- rameter space is large and multi-dimensional while kinetic data are comparatively sparse, the estimation procedure (especially the point estimation methods) often en- counters multiple local maximum such that standard maximum likelihood methods may yield …
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Theses and Dissertations--Statistics
Statistical intervals (e.g., confidence, prediction, or tolerance) are widely used to quantify uncertainty, but complex settings can create challenges to obtain such intervals that possess the desired properties. My thesis will address diverse data settings and approaches that are shown empirically to have good performance. We first introduce a focused treatment on using a single-layer bootstrap calibration to improve the coverage probabilities of two-sided parametric tolerance intervals for non-normal distributions. We then turn to zero-inflated data, which are commonly found in, among other areas, pharmaceutical and quality control applications. However, the inference problem often becomes difficult in the presence of …
Measuring Change: Prediction Of Early Onset Sepsis, Aric Schadler
Measuring Change: Prediction Of Early Onset Sepsis, Aric Schadler
Theses and Dissertations--Statistics
Sepsis occurs in a patient when an infection enters into the blood stream and spreads throughout the body causing a cascading response from the immune system. Sepsis is one of the leading causes of morbidity and mortality in today’s hospitals. This is despite published and accepted guidelines for timely and appropriate interventions for septic patients. The largest barrier to applying these interventions is the early identification of septic patients. Early identification and treatment leads to better outcomes, shorter lengths of stay, and financial savings for healthcare institutions. In order to increase the lead time in recognizing patients trending towards septicemia …
Serial Testing For Detection Of Multilocus Genetic Interactions, Zaid T. Al-Khaledi
Serial Testing For Detection Of Multilocus Genetic Interactions, Zaid T. Al-Khaledi
Theses and Dissertations--Statistics
A method to detect relationships between disease susceptibility and multilocus genetic interactions is the Multifactor-Dimensionality Reduction (MDR) technique pioneered by Ritchie et al. (2001). Since its introduction, many extensions have been pursued to deal with non-binary outcomes and/or account for multiple interactions simultaneously. Studying the effects of multilocus genetic interactions on continuous traits (blood pressure, weight, etc.) is one case that MDR does not handle. Culverhouse et al. (2004) and Gui et al. (2013) proposed two different methods to analyze such a case. In their research, Gui et al. (2013) introduced the Quantitative Multifactor-Dimensionality Reduction (QMDR) that uses the overall …
Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis
Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis
Theses and Dissertations--Statistics
I consider statistical modelling of data gathered by photographic identification in mark-recapture studies and propose a new method that incorporates the inherent uncertainty of photographic identification in the estimation of abundance, survival and recruitment. A hierarchical model is proposed which accepts scores assigned to pairs of photographs by pattern recognition algorithms as data and allows for uncertainty in matching photographs based on these scores. The new models incorporate latent capture histories that are treated as unknown random variables informed by the data, contrasting past models having the capture histories being fixed. The methods properly account for uncertainty in the matching …
The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie
The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie
Theses and Dissertations--Statistics
When scientists know in advance that some features (variables) are important in modeling a data, then these important features should be kept in the model. How can we utilize this prior information to effectively find other important features? This dissertation is to provide a solution, using such prior information. We propose the Conditional Adaptive Lasso (CAL) estimates to exploit this knowledge. By choosing a meaningful conditioning set, namely the prior information, CAL shows better performance in both variable selection and model estimation. We also propose Sufficient Conditional Adaptive Lasso Variable Screening (SCAL-VS) and Conditioning Set Sufficient Conditional Adaptive Lasso Variable …
Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang
Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang
Theses and Dissertations--Statistics
Finite Mixture model has been studied for a long time, however, traditional methods assume that the variables are measured without error. Mixtures-of-regression model with measurement error imposes challenges to the statisticians, since both the mixture structure and the existence of measurement error can lead to inconsistent estimate for the regression coefficients. In order to solve the inconsistency, We propose series of methods to estimate the mixture likelihood of the mixtures-of-regressions model when there is measurement error, both in the responses and predictors. Different estimators of the parameters are derived and compared with respect to their relative efficiencies. The simulation results …
Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu
Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu
Theses and Dissertations--Statistics
Firstly, we reviewed some popular nonparameteric regression methods during the past several decades. Then we extended the compound estimation (Charnigo and Srinivasan [2011]) to adapt random design points and heteroskedasticity and proposed a modified Cp criteria for tuning parameter selection. Moreover, we developed a DCp criteria for tuning paramter selection problem in general nonparametric derivative estimation. This extends GCp criteria in Charnigo, Hall and Srinivasan [2011] with random design points and heteroskedasticity. Next, we proposed a change point detection method via compound estimation for both fixed design and random design case, the adaptation of heteroskedasticity was considered for the method. …
Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan
Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan
Theses and Dissertations--Statistics
We introduce a new class of measures for testing independence between two random vectors, which uses expected difference of conditional and marginal characteristic functions. By choosing a particular weight function in the class, we propose a new index for measuring independence and study its property. Two empirical versions are developed, their properties, asymptotics, connection with existing measures and applications are discussed. Implementation and Monte Carlo results are also presented.
We propose a two-stage sufficient variable selections method based on the new index to deal with large p small n data. The method does not require model specification and especially focuses …
Multi-State Models With Missing Covariates, Wenjie Lou
Multi-State Models With Missing Covariates, Wenjie Lou
Theses and Dissertations--Statistics
Multi-state models have been widely used to analyze longitudinal event history data obtained in medical studies. The tools and methods developed recently in this area require the complete observed datasets. While, in many applications measurements on certain components of the covariate vector are missing on some study subjects. In this dissertation, several likelihood-based methodologies were proposed to deal with datasets with different types of missing covariates efficiently when applying multi-state models.
Firstly, a maximum observed data likelihood method was proposed when the data has a univariate missing pattern and the missing covariate is a categorical variable. The construction of the …
Statistical Methods For Handling Intentional Inaccurate Responders, Kristen J. Mcquerry
Statistical Methods For Handling Intentional Inaccurate Responders, Kristen J. Mcquerry
Theses and Dissertations--Statistics
In self-report data, participants who provide incorrect responses are known as intentional inaccurate responders. This dissertation provides statistical analyses for address intentional inaccurate responses in the data.
Previous work with adolescent self-report, labeled survey participants who intentionally provide inaccurate answers as mischievous responders. This phenomenon also occurs in clinical research. For example, pregnant women who smoke may report that they are nonsmokers. Our advantage is that we do not solely have self-report answers and can verify responses with lab values. Currently, there is no clear method for handling these intentional inaccurate respondents when it comes to making statistical inferences.
We …
Empirical Likelihood And Differentiable Functionals, Zhiyuan Shen
Empirical Likelihood And Differentiable Functionals, Zhiyuan Shen
Theses and Dissertations--Statistics
Empirical likelihood (EL) is a recently developed nonparametric method of statistical inference. It has been shown by Owen (1988,1990) and many others that empirical likelihood ratio (ELR) method can be used to produce nice confidence intervals or regions. Owen (1988) shows that -2logELR converges to a chi-square distribution with one degree of freedom subject to a linear statistical functional in terms of distribution functions. However, a generalization of Owen's result to the right censored data setting is difficult since no explicit maximization can be obtained under constraint in terms of distribution functions. Pan and Zhou (2002), instead, study the …
Multi-State Models For Interval Censored Data With Competing Risk, Shaoceng Wei
Multi-State Models For Interval Censored Data With Competing Risk, Shaoceng Wei
Theses and Dissertations--Statistics
Multi-state models are often used to evaluate the effect of death as a competing event to the development of dementia in a longitudinal study of the cognitive status of elderly subjects. In this dissertation, both multi-state Markov model and semi-Markov model are used to characterize the flow of subjects from intact cognition to dementia with mild cognitive impairment and global impairment as intervening transient, cognitive states and death as a competing risk.
Firstly, a multi-state Markov model with three transient states: intact cognition, mild cognitive impairment (M.C.I.) and global impairment (G.I.) and one absorbing state: dementia is used to model …
New Results In Ell_1 Penalized Regression, Edward A. Roualdes
New Results In Ell_1 Penalized Regression, Edward A. Roualdes
Theses and Dissertations--Statistics
Here we consider penalized regression methods, and extend on the results surrounding the l1 norm penalty. We address a more recent development that generalizes previous methods by penalizing a linear transformation of the coefficients of interest instead of penalizing just the coefficients themselves. We introduce an approximate algorithm to fit this generalization and a fully Bayesian hierarchical model that is a direct analogue of the frequentist version. A number of benefits are derived from the Bayesian persepective; most notably choice of the tuning parameter and natural means to estimate the variation of estimates – a notoriously difficult task for the …