Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (25)
- Statistical Methodology (22)
- Biostatistics (15)
- Multivariate Analysis (7)
- Environmental Sciences (6)
-
- Statistical Theory (6)
- Earth Sciences (5)
- Hydrology (5)
- Data Science (4)
- Life Sciences (4)
- Medicine and Health Sciences (4)
- Engineering (3)
- Longitudinal Data Analysis and Time Series (3)
- Microarrays (3)
- Social and Behavioral Sciences (3)
- Survival Analysis (3)
- Water Resource Management (3)
- Bioinformatics (2)
- Civil and Environmental Engineering (2)
- Geography (2)
- Public Health (2)
- Soil Science (2)
- Transportation Engineering (2)
- Animal Sciences (1)
- Biochemistry (1)
- Biochemistry, Biophysics, and Structural Biology (1)
- Biology (1)
- Keyword
-
- EM algorithm (4)
- Count data (3)
- Bayesian Analysis (2)
- Big Data (2)
- Bootstrap (2)
-
- Data dispersion (2)
- EM Algorithm (2)
- Generalized linear models (2)
- Interaction (2)
- Latent Class Analysis (2)
- Linear model (2)
- Multivariate Data (2)
- Normalization (2)
- ODE (2)
- Spatial-temporal process (2)
- Statistics (2)
- Sufficient dimension reduction (2)
- Variable Screening (2)
- Variable Selection (2)
- 3-D Highway Geometric Design (1)
- Adjustment (1)
- Algorithm (1)
- Alzheimer's Disease (1)
- Aquifers (1)
- Asymptotic distribution (1)
- Asymptotic properties (1)
- At-Fault (1)
- Autologistic regression models (1)
- Autoregressive models (1)
- Average Causal Effect (1)
- Publication Year
- Publication
- Publication Type
Articles 1 - 30 of 58
Full-Text Articles in Statistical Models
Using Camera-Based Unmarked Spatial Capture-Recapture Modeling To Estimate Reintroduced Elk (Cervus Canadensis) Population Parameters And Distribution In Southeastern Kentucky, Claire Marie Muia
Theses and Dissertations--Forestry and Natural Resources
Estimation of population parameters is important for wildlife management decisions. Elk reintroduced to southeastern Kentucky experienced early irruptive population growth and are currently monitored using a statewide harvest-based statistical population reconstruction model (SPR) across the Kentucky Elk Restoration Zone (KERZ). Because the SPR model is spatially coarse and difficult to scale to the smaller management units comprising the KERZ, we conducted a spatially explicit capture-recapture study using a clustered camera-trapping array deployed for 10 weeks from June–August 2024 to estimate elk population parameters within Management Unit 4. Due to a lack of resights of GPS-marked elk, population parameters were estimated …
Variable Selection For High-Dimensional Data With Interaction Effects: Methods, Applications, And Inferences, Leiyue Li
Theses and Dissertations--Statistics
For high-dimensional data where the number of variables greatly exceeds the number of observations, selecting important variables while maintaining the required heredity conditions can be challenging. This dissertation is structured into three interconnected parts. In the first part, we propose a variable selection method by implementing a well-known optimization technique, the Genetic Algorithm. An R package was developed to simplify the implementation and usage of the proposed method. We then propose another variable selection method by extending the study from the Genetic Algorithm to a different but related optimization technique, Simulated Annealing. We consider three different hierarchical structures in both …
Patterns Into Pathways For Improving Safety Culture: Refined Latent Class Analysis Informs Tailored Decision Support For South Carolina Dss Safety Culture Improvements, Michaela Voit
Theses and Dissertations--Public Health (M.P.H. & Dr.P.H.)
The rising prevalence of exposures to adverse childhood experiences (ACEs) demands a coordinated public health response, as a significant body of research details the cumulative impact of ACEs on chronic morbidities contributing to reduced life expectancy. Child welfare workers (CWW) are embedded in this public health effort, tasked with preventing and mitigating the impacts of ACEs through family and prevention services. The National Partnership for Child Safety (NPCS) may improve the wellbeing of CWWs and the effectiveness of Child Welfare (CW) services by improving the quality of safety culture within CW organizations. To inform NPCS quality improvement efforts, our project …
Difs And Bayescluster: Novel Methods For Single_Cell Rna Sequencing Analysis, Kun Liu
Difs And Bayescluster: Novel Methods For Single_Cell Rna Sequencing Analysis, Kun Liu
Theses and Dissertations--Statistics
Single-cell RNA sequencing (scRNA-seq) has transformed our understanding of cellular heterogeneity and gene expression dynamics. Despite its potential, the inherent noise and sparsity of scRNA-seq data pose significant challenges in clustering cells into biologically meaningful groups. This dissertation addresses these challenges through two novel methodologies aimed at enhancing the accuracy and robustness of scRNA-seq data analysis.
First, we introduce the Differential Feature Selection (DIFS) framework, designed to improve the identification of differential features in scRNA-seq data. DIFS employs a two-stage marker identification process. In the first stage, a modified Dip Test is used to filter and identify genes with significant …
Statistical Tolerance Regions For Flexible Modeling Paradigms, Yafan Guo
Statistical Tolerance Regions For Flexible Modeling Paradigms, Yafan Guo
Theses and Dissertations--Statistics
Tolerance intervals in a regression setting allow the user to quantify, with a specified degree of confidence, bounds for a specified proportion of the sampled population when conditioned on a set of covariate values. While methods are available for tolerance intervals in fully-parametric regression settings, the construction of tolerance intervals for semiparametric regression models has been treated in a limited capacity. The first project fills this gap and develops likelihood-based approaches for the construction of pointwise one-sided and two-sided tolerance intervals for semiparametric regression models. A numerical approach is also presented for constructing simultaneous tolerance intervals. An appealing facet of …
Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan
Statistical Intervals For Neural Network And Its Relationship With Generalized Linear Model, Sheng Yuan
Theses and Dissertations--Statistics
Neural networks have experienced widespread adoption and have become integral in cutting-edge domains like computer vision, natural language processing, and various contemporary fields. However, addressing the statistical aspects of neural networks has been a persistent challenge, with limited satisfactory results. In my research, I focused on exploring statistical intervals applied to neural networks, specifically confidence intervals and tolerance intervals. I employed variance estimation methods, such as direct estimation and resampling, to assess neural networks and their performance under outlier scenarios. Remarkably, when outliers were present, the resampling method with infinitesimal jackknife estimation yielded confidence intervals that closely aligned with nominal …
High Dimensional Data Analysis: Variable Screening And Inference, Lei Fang
High Dimensional Data Analysis: Variable Screening And Inference, Lei Fang
Theses and Dissertations--Statistics
This dissertation focuses on the problem of high dimensional data analysis, which arises in many fields including genomics, finance, and social sciences. In such settings, the number of features or variables is much larger than the number of observations, posing significant challenges to traditional statistical methods.
To address these challenges, this dissertation proposes novel methods for variable screening and inference. The first part of the dissertation focuses on variable screening, which aims to identify a subset of important variables that are strongly associated with the response variable. Specifically, we propose a robust nonparametric screening method to effectively select the predictors …
Novel Modelling And Inference Considerations Involving The Exponentially-Modified Gaussian Distribution, Yanxi Li
Theses and Dissertations--Statistics
The exponentially-modified Gaussian (EMG) distribution is well-suited for analyzing data with positive skewness due to its characteristic positive skew from the exponential component. Despite its popularity in various fields, the EMG distribution has only been analyzed for univariate data without any regression settings. To address this limitation, we developed a generalized EMG regression model with covariates by assigning parametric functional forms to some or all of the parameters in the EMG distribution that vary with values of the covariates. To further perform data-clustering on observation points, we propose a competing regression model where the error structure is assumed to be …
Finite Mixtures Of Mean-Parameterized Conway-Maxwell-Poisson Models, Dongying Zhan
Finite Mixtures Of Mean-Parameterized Conway-Maxwell-Poisson Models, Dongying Zhan
Theses and Dissertations--Statistics
For modeling count data, the Conway-Maxwell-Poisson (CMP) distribution is a popular generalization of the Poisson distribution due to its ability to characterize data over- or under-dispersion. While the classic parameterization of the CMP has been well-studied, its main drawback is that it is does not directly model the mean of the counts. This is mitigated by using a mean-parameterized version of the CMP distribution. In this work, we are concerned with the setting where count data may be comprised of subpopulations, each possibly having varying degrees of data dispersion. Thus, we propose a finite mixture of mean-parameterized CMP distributions. An …
Methodologies And Computational Tools For Zero-Inflated Discrete Weibull Models, Peng Yeh
Methodologies And Computational Tools For Zero-Inflated Discrete Weibull Models, Peng Yeh
Theses and Dissertations--Statistics
Count data with excess zeros is common in many fields, such as ecology, healthcare, and insurance. Excess zeros data are often causing the inaccurate fit from the count models. While zero-inflated models have been developing for over two decades, one should also consider a more flexible model that can handle the excess zeros and further over- or under-dispersion. In this talk, we discuss zero-inflated discrete Weibull model and some novel computational contributions. The flexibility and competitiveness of the ZIDW model are illustrated by simulation studies and a real data analysis. We also investigate the performance of the proposed model through …
Potential Alzheimer's Disease Plasma Biomarkers, Taylor Estepp
Potential Alzheimer's Disease Plasma Biomarkers, Taylor Estepp
Theses and Dissertations--Epidemiology and Biostatistics
In this series of studies, we examined the potential of a variety of blood-based plasma biomarkers for the identification of Alzheimer's disease (AD) progression and cognitive decline. With the end goal of studying these biomarkers via mixture modeling, we began with a literature review of the methodology. An examination of the biomarkers with demographics and other health factors found evidence of minimal risk of confounding along the causal pathway from biomarkers to cognitive performance. Further study examined the usefulness of linear combinations of biomarkers, achieved via partial least squares (PLS) analysis, as predictors of various cognitive assessment scores and clinical …
Statistical Theory For Specialized Linear Regression Adjustment Methods Compared To Multiple Linear Regression In The Presence And Absence Of Interaction Effects, Leon Su
Theses and Dissertations--Statistics
When building models to investigate outcomes and variables of interest, researchers often want to adjust for other variables. There is a variety of ways that these adjustments are performed. In this work, we will consider four approaches to adjustment utilized by researchers in various fields. We will compare the efficacy of these methods to what we call the ”true model method”, fitting a multiple linear regression model in which adjustment variables are model covariates. Our goal is to show that these adjustment methods have inferior performance to the true model method by comparing model parameter estimates, power, type I error, …
Deriving The Distributions And Developing Methods Of Inference For R2-Type Measures, With Applications To Big Data Analysis, Gregory S. Hawk
Deriving The Distributions And Developing Methods Of Inference For R2-Type Measures, With Applications To Big Data Analysis, Gregory S. Hawk
Theses and Dissertations--Statistics
As computing capabilities and cloud-enhanced data sharing has accelerated exponentially in the 21st century, our access to Big Data has revolutionized the way we see data around the world, from healthcare to investments to manufacturing to retail and supply-chain. In many areas of research, however, the cost of obtaining each data point makes more than just a few observations impossible. While machine learning and artificial intelligence (AI) are improving our ability to make predictions from datasets, we need better statistical methods to improve our ability to understand and translate models into meaningful and actionable insights.
A central goal in the …
Beta Mixture And Contaminated Model With Constraints And Application With Micro-Array Data, Ya Qi
Beta Mixture And Contaminated Model With Constraints And Application With Micro-Array Data, Ya Qi
Theses and Dissertations--Statistics
This dissertation research is concentrated on the Contaminated Beta(CB) model and its application in micro-array data analysis. Modified Likelihood Ratio Test (MLRT) introduced by [Chen et al., 2001] is used for testing the omnibus null hypothesis of no contamination of Beta(1,1)([Dai and Charnigo, 2008]). We design constraints for two-component CB model, which put the mode toward the left end of the distribution to reflect the abundance of small p-values of micro-array data, to increase the test power. A three-component CB model might be useful when distinguishing high differentially expressed genes and moderate differentially expressed genes. If the null hypothesis above …
Sexual Behaviors Associated With Online Partner-Seeking Among Men Who Have Sex With Men From Small/Midsized Towns Or Rural Areas In Kentucky, Vira Pravosud
Theses and Dissertations--Epidemiology and Biostatistics
The HIV epidemic remains one of the most significant public health issues in the United States, particularly among men who have sex with men (MSM). New avenues for partner-seeking have emerged over the past three decades, including through the Internet, social media, and geosocial networking applications. Consisting of three cross-sectional studies, this dissertation research aimed to determine associations between the use of various online tools for partner-seeking (hereafter collectively referred to as “apps”) and HIV-related sexual behaviors among 252 young adult MSM residing in small/midsized towns or rural areas in Central Kentucky, a group that has been under-represented in the …
Dimension Reduction Techniques In Regression, Pei Wang
Dimension Reduction Techniques In Regression, Pei Wang
Theses and Dissertations--Statistics
Because of the advances of modern technology, the size of the collected data nowadays is larger and the structure is more complex. To deal with such kinds of data, sufficient dimension reduction (SDR) and reduced rank (RR) regression are two powerful tools. This dissertation focuses on these two tools and it is composed of three projects. In the first project, we introduce a new SDR method through a novel approach of feature filter to recover the central mean subspace exhaustively along with a method to determine the dimension, two variable selection methods, and extensions to multivariate response and large p …
Novel Methods For Characterizing Conditional Quantiles In Zero-Inflated Count Regression Models, Xuan Shi
Novel Methods For Characterizing Conditional Quantiles In Zero-Inflated Count Regression Models, Xuan Shi
Theses and Dissertations--Statistics
Despite its popularity in diverse disciplines, quantile regression methods are primarily designed for the continuous response setting and cannot be directly applied to the discrete (or count) response setting. There can also be challenges when modeling count responses, such as the presence of excess zero counts, formally known as zero-inflation. To address the aforementioned challenges, we propose a comprehensive model-aware strategy that synthesizes quantile regression methods with estimation of zero-inflated count regression models. Various competing computational routines are examined, while residual analysis and model selection procedures are included to validate our method. The performance of these methods is characterized through …
Estimating And Testing Treatment Effects With Misclassified Multivariate Data, Zi Ye
Estimating And Testing Treatment Effects With Misclassified Multivariate Data, Zi Ye
Theses and Dissertations--Statistics
Clinical trials are often used to assess drug efficacy and safety. Participants are sometimes pre-stratified into different groups by diagnostic tools. However, these diagnostic tools are fallible. The traditional method ignores this problem and assumes the diagnostic devices are perfect. This assumption will lead to inefficient and biased estimators. In this era of personalized medicine and measurement-based care, the issues of bias and efficiency are of paramount importance. Despite the prominence, only few researches evaluated the treatment effect in the presence of misclassifications in some special cases and most others focus on assessing the accuracy of the diagnostic devices. In …
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Semiparametric And Nonparametric Methods For Comparing Biomarker Levels Between Groups, Yuntong Li
Theses and Dissertations--Statistics
Comparing the distribution of biomarker measurements between two groups under either an unpaired or paired design is a common goal in many biomarker studies. However, analyzing biomarker data is sometimes challenging because the data may not be normally distributed and contain a large fraction of zero values or missing values. Although several statistical methods have been proposed, they either require data normality assumption, or are inefficient. We proposed a novel two-part semiparametric method for data under an unpaired setting and a nonparametric method for data under a paired setting. The semiparametric method considers a two-part model, a logistic regression for …
Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu
Estimation Of The Treatment Effect With Bayesian Adjustment For Covariates, Li Xu
Theses and Dissertations--Statistics
The Bayesian adjustment for confounding (BAC) is a Bayesian model averaging method to select and adjust for confounding factors when evaluating the average causal effect of an exposure on a certain outcome. We extend the BAC method to time-to-event outcomes. Specifically, the posterior distribution of the exposure effect on a time-to-event outcome is calculated as a weighted average of posterior distributions from a number of candidate proportional hazards models, weighing each model by its ability to adjust for confounding factors. The Bayesian Information Criterion based on the partial likelihood is used to compare different models and approximate the Bayes factor. …
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Nonparametric Tests Of Lack Of Fit For Multivariate Data, Yan Xu
Theses and Dissertations--Statistics
A common problem in regression analysis (linear or nonlinear) is assessing the lack-of-fit. Existing methods make parametric or semi-parametric assumptions to model the conditional mean or covariance matrices. In this dissertation, we propose fully nonparametric methods that make only additive error assumptions. Our nonparametric approach relies on ideas from nonparametric smoothing to reduce the test of association (lack-of-fit) problem into a nonparametric multivariate analysis of variance. A major problem that arises in this approach is that the key assumptions of independence and constant covariance matrix among the groups will be violated. As a result, the standard asymptotic theory is not …
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Bayesian Kinetic Modeling For Tracer-Based Metabolomic Data, Xu Zhang
Theses and Dissertations--Statistics
Kinetic modeling of the time dependence of metabolite concentrations including the unstable isotope labeled species is an important approach to simulate metabolic pathway dynamics. It is also essential for quantitative metabolic flux analysis using tracer data. However, as the metabolic networks are complex including extensive compartmentation and interconnections, the parameter estimation for enzymes that catalyze individual reactions needed for kinetic modeling is challenging. As the pa- rameter space is large and multi-dimensional while kinetic data are comparatively sparse, the estimation procedure (especially the point estimation methods) often en- counters multiple local maximum such that standard maximum likelihood methods may yield …
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Statistical Intervals For Various Distributions Based On Different Inference Methods, Yixuan Zou
Theses and Dissertations--Statistics
Statistical intervals (e.g., confidence, prediction, or tolerance) are widely used to quantify uncertainty, but complex settings can create challenges to obtain such intervals that possess the desired properties. My thesis will address diverse data settings and approaches that are shown empirically to have good performance. We first introduce a focused treatment on using a single-layer bootstrap calibration to improve the coverage probabilities of two-sided parametric tolerance intervals for non-normal distributions. We then turn to zero-inflated data, which are commonly found in, among other areas, pharmaceutical and quality control applications. However, the inference problem often becomes difficult in the presence of …
The Use Of 3-D Highway Differential Geometry In Crash Prediction Modeling, Kiriakos Amiridis
The Use Of 3-D Highway Differential Geometry In Crash Prediction Modeling, Kiriakos Amiridis
Theses and Dissertations--Civil Engineering
The objective of this research is to evaluate and introduce a new methodology regarding rural highway safety. Current practices rely on crash prediction models that utilize specific explanatory variables, whereas the depository of knowledge for past research is the Highway Safety Manual (HSM). Most of the prediction models in the HSM identify the effect of individual geometric elements on crash occurrence and consider their combination in a multiplicative manner, where each effect is multiplied with others to determine their combined influence. The concepts of 3-dimesnional (3-D) representation of the roadway surface have also been explored in the past aiming to …
Transforms In Sufficient Dimension Reduction And Their Applications In High Dimensional Data, Jiaying Weng
Transforms In Sufficient Dimension Reduction And Their Applications In High Dimensional Data, Jiaying Weng
Theses and Dissertations--Statistics
The big data era poses great challenges as well as opportunities for researchers to develop efficient statistical approaches to analyze massive data. Sufficient dimension reduction is such an important tool in modern data analysis and has received extensive attention in both academia and industry.
In this dissertation, we introduce inverse regression estimators using Fourier transforms, which is superior to the existing SDR methods in two folds, (1) it avoids the slicing of the response variable, (2) it can be readily extended to solve the high dimensional data problem. For the ultra-high dimensional problem, we investigate both eigenvalue decomposition and minimum …
Composite Nonparametric Tests In High Dimension, Alejandro G. Villasante Tezanos
Composite Nonparametric Tests In High Dimension, Alejandro G. Villasante Tezanos
Theses and Dissertations--Statistics
This dissertation focuses on the problem of making high-dimensional inference for two or more groups. High-dimensional means both the sample size (n) and dimension (p) tend to infinity, possibly at different rates. Classical approaches for group comparisons fail in the high-dimensional situation, in the sense that they have incorrect sizes and low powers. Much has been done in recent years to overcome these problems. However, these recent works make restrictive assumptions in terms of the number of treatments to be compared and/or the distribution of the data. This research aims to (1) propose and investigate refined …
A Flexible Zero-Inflated Poisson Regression Model, Eric S. Roemmele
A Flexible Zero-Inflated Poisson Regression Model, Eric S. Roemmele
Theses and Dissertations--Statistics
A practical problem often encountered with observed count data is the presence of excess zeros. Zero-inflation in count data can easily be handled by zero-inflated models, which is a two-component mixture of a point mass at zero and a discrete distribution for the count data. In the presence of predictors, zero-inflated Poisson (ZIP) regression models are, perhaps, the most commonly used. However, the fully parametric ZIP regression model could sometimes be restrictive, especially with respect to the mixing proportions. Taking inspiration from some of the recent literature on semiparametric mixtures of regressions models for flexible mixture modeling, we propose a …
Automatic 13C Chemical Shift Reference Correction Of Protein Nmr Spectral Data Using Data Mining And Bayesian Statistical Modeling, Xi Chen
Theses and Dissertations--Molecular and Cellular Biochemistry
Nuclear magnetic resonance (NMR) is a highly versatile analytical technique for studying molecular configuration, conformation, and dynamics, especially of biomacromolecules such as proteins. However, due to the intrinsic properties of NMR experiments, results from the NMR instruments require a refencing step before the down-the-line analysis. Poor chemical shift referencing, especially for 13C in protein Nuclear Magnetic Resonance (NMR) experiments, fundamentally limits and even prevents effective study of biomacromolecules via NMR. There is no available method that can rereference carbon chemical shifts from protein NMR without secondary experimental information such as structure or resonance assignment.
To solve this problem, we …
Effect Of Socioeconomic And Demographic Factors On Kentucky Crashes, Aaron Berry Cambron
Effect Of Socioeconomic And Demographic Factors On Kentucky Crashes, Aaron Berry Cambron
Theses and Dissertations--Civil Engineering
The goal of this research was to examine the potential predictive ability of socioeconomic and demographic data for drivers on Kentucky crash occurrence. Identifying unique background characteristics of at-fault drivers that contribute to crash rates and crash severity may lead to improved and more specific interventions to reduce the negative impacts of motor vehicle crashes. The driver-residence zip code was used as a spatial unit to connect five years of Kentucky crash data with socioeconomic factors from the U.S. Census, such as income, employment, education, age, and others, along with terrain and vehicle age. At-fault driver crash counts, normalized over …
Accounting For Spatial Autocorrelation In Modeling The Distribution Of Water Quality Variables, Lorrayne Miralha
Accounting For Spatial Autocorrelation In Modeling The Distribution Of Water Quality Variables, Lorrayne Miralha
Theses and Dissertations--Geography
Several studies in hydrology have reported differences in outcomes between models in which spatial autocorrelation (SAC) is accounted for and those in which SAC is not. However, the capacity to predict the magnitude of such differences is still ambiguous. In this thesis, I hypothesized that SAC, inherently possessed by a response variable, influences spatial modeling outcomes. I selected ten watersheds in the USA and analyzed them to determine whether water quality variables with higher Moran’s I values undergo greater increases in the coefficient of determination (R²) and greater decreases in residual SAC (rSAC) after spatial modeling. I compared non-spatial ordinary …