Open Access. Powered by Scholars. Published by Universities.®

Statistical Models Commons

Open Access. Powered by Scholars. Published by Universities.®

University of Kentucky

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 31 - 58 of 58

Full-Text Articles in Statistical Models

Improved Methods And Selecting Classification Types For Time-Dependent Covariates In The Marginal Analysis Of Longitudinal Data, I-Chen Chen Jan 2018

Improved Methods And Selecting Classification Types For Time-Dependent Covariates In The Marginal Analysis Of Longitudinal Data, I-Chen Chen

Theses and Dissertations--Epidemiology and Biostatistics

Generalized estimating equations (GEE) are popularly utilized for the marginal analysis of longitudinal data. In order to obtain consistent regression parameter estimates, these estimating equations must be unbiased. However, when certain types of time-dependent covariates are presented, these equations can be biased unless an independence working correlation structure is employed. Moreover, in this case regression parameter estimation can be very inefficient because not all valid moment conditions are incorporated within the corresponding estimating equations. Therefore, approaches using the generalized method of moments or quadratic inference functions have been proposed for utilizing all valid moment conditions. However, we have found that …


Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis Jan 2018

Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis

Theses and Dissertations--Statistics

I consider statistical modelling of data gathered by photographic identification in mark-recapture studies and propose a new method that incorporates the inherent uncertainty of photographic identification in the estimation of abundance, survival and recruitment. A hierarchical model is proposed which accepts scores assigned to pairs of photographs by pattern recognition algorithms as data and allows for uncertainty in matching photographs based on these scores. The new models incorporate latent capture histories that are treated as unknown random variables informed by the data, contrasting past models having the capture histories being fixed. The methods properly account for uncertainty in the matching …


The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie Jan 2018

The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie

Theses and Dissertations--Statistics

When scientists know in advance that some features (variables) are important in modeling a data, then these important features should be kept in the model. How can we utilize this prior information to effectively find other important features? This dissertation is to provide a solution, using such prior information. We propose the Conditional Adaptive Lasso (CAL) estimates to exploit this knowledge. By choosing a meaningful conditioning set, namely the prior information, CAL shows better performance in both variable selection and model estimation. We also propose Sufficient Conditional Adaptive Lasso Variable Screening (SCAL-VS) and Conditioning Set Sufficient Conditional Adaptive Lasso Variable …


Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang Jan 2018

Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang

Theses and Dissertations--Statistics

Finite Mixture model has been studied for a long time, however, traditional methods assume that the variables are measured without error. Mixtures-of-regression model with measurement error imposes challenges to the statisticians, since both the mixture structure and the existence of measurement error can lead to inconsistent estimate for the regression coefficients. In order to solve the inconsistency, We propose series of methods to estimate the mixture likelihood of the mixtures-of-regressions model when there is measurement error, both in the responses and predictors. Different estimators of the parameters are derived and compared with respect to their relative efficiencies. The simulation results …


Estimation In Partially Linear Models With Correlated Observations And Change-Point Models, Liangdong Fan Jan 2018

Estimation In Partially Linear Models With Correlated Observations And Change-Point Models, Liangdong Fan

Theses and Dissertations--Statistics

Methods of estimating parametric and nonparametric components, as well as properties of the corresponding estimators, have been examined in partially linear models by Wahba [1987], Green et al. [1985], Engle et al. [1986], Speckman [1988], Hu et al. [2004], Charnigo et al. [2015] among others. These models are appealing due to their flexibility and wide range of practical applications including the electricity usage study by Engle et al. [1986], gum disease study by Speckman [1988], etc., wherea parametric component explains linear trends and a nonparametric part captures nonlinear relationships.

The compound estimator (Charnigo et al. [2015]) has been used to …


Some Dimension Reduction Strategies For The Analysis Of Survey Data, Jiaying Weng, Derek S. Young Dec 2017

Some Dimension Reduction Strategies For The Analysis Of Survey Data, Jiaying Weng, Derek S. Young

Statistics Faculty Publications

In the era of big data, researchers interested in developing statistical models are challenged with how to achieve parsimony. Usually, some sort of dimension reduction strategy is employed. Classic strategies are often in the form of traditional inference procedures, such as hypothesis testing; however, the increase in computing capabilities has led to the development of more sophisticated methods. In particular, sufficient dimension reduction has emerged as an area of broad and current interest. While these types of dimension reduction strategies have been employed for numerous data problems, they are scantly discussed in the context of analyzing survey data. This …


Improving The Computational Efficiency In Bayesian Fitting Of Cormack-Jolly-Seber Models With Individual, Continuous, Time-Varying Covariates, Woodrow Burchett Jan 2017

Improving The Computational Efficiency In Bayesian Fitting Of Cormack-Jolly-Seber Models With Individual, Continuous, Time-Varying Covariates, Woodrow Burchett

Theses and Dissertations--Statistics

The extension of the CJS model to include individual, continuous, time-varying covariates relies on the estimation of covariate values on occasions on which individuals were not captured. Fitting this model in a Bayesian framework typically involves the implementation of a Markov chain Monte Carlo (MCMC) algorithm, such as a Gibbs sampler, to sample from the posterior distribution. For large data sets with many missing covariate values that must be estimated, this creates a computational issue, as each iteration of the MCMC algorithm requires sampling from the full conditional distributions of each missing covariate value. This dissertation examines two solutions to …


Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu Jan 2017

Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu

Theses and Dissertations--Statistics

Firstly, we reviewed some popular nonparameteric regression methods during the past several decades. Then we extended the compound estimation (Charnigo and Srinivasan [2011]) to adapt random design points and heteroskedasticity and proposed a modified Cp criteria for tuning parameter selection. Moreover, we developed a DCp criteria for tuning paramter selection problem in general nonparametric derivative estimation. This extends GCp criteria in Charnigo, Hall and Srinivasan [2011] with random design points and heteroskedasticity. Next, we proposed a change point detection method via compound estimation for both fixed design and random design case, the adaptation of heteroskedasticity was considered for the method. …


An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert Jan 2017

An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert

Theses and Dissertations--Epidemiology and Biostatistics

It is estimated that Periodontal Diseases effects up to 90% of the adult population. Given the complexity of the host environment, many factors contribute to expression of the disease. Age, Gender, Socioeconomic Status, Smoking Status, and Race/Ethnicity are all known risk factors, as well as a handful of known comorbidities. Certain vitamins and minerals have been shown to be protective for the disease, while some toxins and chemicals have been associated with an increased prevalence. The role of toxins, chemicals, vitamins, and minerals in relation to disease is believed to be complex and potentially modified by known risk factors. A …


Development In Normal Mixture And Mixture Of Experts Modeling, Meng Qi Jan 2016

Development In Normal Mixture And Mixture Of Experts Modeling, Meng Qi

Theses and Dissertations--Statistics

In this dissertation, first we consider the problem of testing homogeneity and order in a contaminated normal model, when the data is correlated under some known covariance structure. To address this problem, we developed a moment based homogeneity and order test, and design weights for test statistics to increase power for homogeneity test. We applied our test to microarray about Down’s syndrome. This dissertation also studies a singular Bayesian information criterion (sBIC) for a bivariate hierarchical mixture model with varying weights, and develops a new data dependent information criterion (sFLIC).We apply our model and criteria to birth- weight and gestational …


Multi-State Models With Missing Covariates, Wenjie Lou Jan 2016

Multi-State Models With Missing Covariates, Wenjie Lou

Theses and Dissertations--Statistics

Multi-state models have been widely used to analyze longitudinal event history data obtained in medical studies. The tools and methods developed recently in this area require the complete observed datasets. While, in many applications measurements on certain components of the covariate vector are missing on some study subjects. In this dissertation, several likelihood-based methodologies were proposed to deal with datasets with different types of missing covariates efficiently when applying multi-state models.

Firstly, a maximum observed data likelihood method was proposed when the data has a univariate missing pattern and the missing covariate is a categorical variable. The construction of the …


Topics In Logistic Regression Analysis, Zhiheng Xie Jan 2016

Topics In Logistic Regression Analysis, Zhiheng Xie

Theses and Dissertations--Statistics

Discrete-time Markov chains have been used to analyze the transition of subjects from intact cognition to dementia with mild cognitive impairment and global impairment as intervening transient states, and death as competing risk. A multinomial logistic regression model is used to estimate the probability distribution in each row of the one-step transition matrix that correspond to the transient states. We investigate some goodness of fit tests for a multinomial distribution with covariates to assess the fit of this model to the data. We propose a modified chi-square test statistic and a score test statistic for the multinomial assumption in each …


Developing An Alternative Way To Analyze Nanostring Data, Shu Shen Jan 2016

Developing An Alternative Way To Analyze Nanostring Data, Shu Shen

Theses and Dissertations--Statistics

Nanostring technology provides a new method to measure gene expressions. It's more sensitive than microarrays and able to do more gene measurements than RT-PCR with similar sensitivity. This system produces counts for each target gene and tabulates them. Counts can be normalized by using an Excel macro or nSolver before analysis. Both methods rely on data normalization prior to statistical analysis to identify differentially expressed genes. Alternatively, we propose to model gene expressions as a function of positive controls and reference gene measurements. Simulations and examples are used to compare this model with Nanostring normalization methods. The results show that …


Statistical Inference On Dynamical Systems, Hongyuan Wang Jan 2016

Statistical Inference On Dynamical Systems, Hongyuan Wang

Theses and Dissertations--Statistics

The ordinary differential equation (ODE) is one representative and popular tool in modeling dynamical systems, which are widely implemented in physics, biology, economics, chemistry and biomedical sciences, etc. Because of the importance of dynamical systems in scientific studies, they are the main focuses of my dissertation.

The first chapter of the dissertation is introduction and literature review, which mainly focuses on numerical integration algorithms of ODEs that are difficult to solve analytically, as well as derivative-free optimization algorithms for the so-called inverse problem.

The second chapter is on the estimation method based on numerical solvers of differential equations. We start …


Statistical Methods For Environmental Exposure Data Subject To Detection Limits, Yuchen Yang Jan 2016

Statistical Methods For Environmental Exposure Data Subject To Detection Limits, Yuchen Yang

Theses and Dissertations--Statistics

In this dissertation, we develop unified and efficient nonparametric statistical methods for estimating and comparing environmental exposure distributions in presence of detection limits. In the first part, we propose a kernel-smoothed nonparametric estimator for the exposure distribution without imposing any independence assumption between the exposure level and detection limit. We show that the proposed estimator is consistent and asymptotically normal. Simulation studies demonstrate that the proposed estimator performs well in practical situations. A colon cancer study is provided for illustration. In the second part, we develop a class of test statistics to compare exposure distributions between two groups by using …


Improved Models For Differential Analysis For Genomic Data, Hong Wang Jan 2016

Improved Models For Differential Analysis For Genomic Data, Hong Wang

Theses and Dissertations--Statistics

This paper intend to develop novel statistical methods to improve genomic data analysis, especially for differential analysis. We considered two different data type: NanoString nCounter data and somatic mutation data. For NanoString nCounter data, we develop a novel differential expression detection method. The method considers a generalized linear model of the negative binomial family to characterize count data and allows for multi-factor design. Data normalization is incorporated in the model framework through data normalization parameters, which are estimated from control genes embedded in the nCounter system. For somatic mutation data, we develop beta-binomial model-based approaches to identify highly or lowly …


Continuous Time Multi-State Models For Interval Censored Data, Lijie Wan Jan 2016

Continuous Time Multi-State Models For Interval Censored Data, Lijie Wan

Theses and Dissertations--Statistics

Continuous-time multi-state models are widely used in modeling longitudinal data of disease processes with multiple transient states, yet the analysis is complex when subjects are observed periodically, resulting in interval censored data. Recently, most studies focused on modeling the true disease progression as a discrete time stationary Markov chain, and only a few studies have been carried out regarding non-homogenous multi-state models in the presence of interval-censored data. In this dissertation, several likelihood-based methodologies were proposed to deal with interval censored data in multi-state models.

Firstly, a continuous time version of a homogenous Markov multi-state model with backward transitions was …


New Results In Ell_1 Penalized Regression, Edward A. Roualdes Jan 2015

New Results In Ell_1 Penalized Regression, Edward A. Roualdes

Theses and Dissertations--Statistics

Here we consider penalized regression methods, and extend on the results surrounding the l1 norm penalty. We address a more recent development that generalizes previous methods by penalizing a linear transformation of the coefficients of interest instead of penalizing just the coefficients themselves. We introduce an approximate algorithm to fit this generalization and a fully Bayesian hierarchical model that is a direct analogue of the frequentist version. A number of benefits are derived from the Bayesian persepective; most notably choice of the tuning parameter and natural means to estimate the variation of estimates – a notoriously difficult task for the …


The Initial Phases Of A Consistent Pricing System That Reflects The Online Sale Value Of A Horse, Curran A. Prettyman Jan 2014

The Initial Phases Of A Consistent Pricing System That Reflects The Online Sale Value Of A Horse, Curran A. Prettyman

Lewis Honors College Capstone Collection

Horses are one of the most uniquely priced commodities. This document provides a solution to an industry-wide weakness of inconsistent pricing and confusion. In the following report, an evaluation of the industry flaw is presented, an econometric approach is described in full, and a solution is proposed using insight gained from a regression analysis. This report uses an econometric approach to determine the impact of hunter jumper horse qualities on internet sale prices. Data is compiled from bigeq.com for seventy-eight horses in the states of Illinois, Indiana, Kentucky, Michigan, and Ohio. A linear regression analysis for twelve variables establishes that …


Genetic Association Testing Of Copy Number Variation, Yinglei Li Jan 2014

Genetic Association Testing Of Copy Number Variation, Yinglei Li

Theses and Dissertations--Statistics

Copy-number variation (CNV) has been implicated in many complex diseases. It is of great interest to detect and locate such regions through genetic association testings. However, the association testings are complicated by the fact that CNVs usually span multiple markers and thus such markers are correlated to each other. To overcome the difficulty, it is desirable to pool information across the markers. In this thesis, we propose a kernel-based method for aggregation of marker-level tests, in which first we obtain a bunch of p-values through association tests for every marker and then the association test involving CNV is based on …


Normal Mixture And Contaminated Model With Nuisance Parameter And Applications, Qian Fan Jan 2014

Normal Mixture And Contaminated Model With Nuisance Parameter And Applications, Qian Fan

Theses and Dissertations--Statistics

This paper intend to find the proper hypothesis and test statistic for testing existence of bilaterally contamination when there exists nuisance parameter. The test statistic is based on method of moments estimators. Union-Intersection test is used for testing if the distribution of population can be implemented by a bilaterally contaminated normal model with unknown variance. This paper also developed a hierarchical normal mixture model (HNM) and applied it to birth weight data. EM algorithm is employed for parameter estimation and a singular Bayesian information criterion (sBIC) is applied to choose the number components. We also proposed a singular flexible information …


Analysis Of Spatial Data, Xiang Zhang Jan 2013

Analysis Of Spatial Data, Xiang Zhang

Theses and Dissertations--Statistics

In many areas of the agriculture, biological, physical and social sciences, spatial lattice data are becoming increasingly common. In addition, a large amount of lattice data shows not only visible spatial pattern but also temporal pattern (see, Zhu et al. 2005). An interesting problem is to develop a model to systematically model the relationship between the response variable and possible explanatory variable, while accounting for space and time effect simultaneously.

Spatial-temporal linear model and the corresponding likelihood-based statistical inference are important tools for the analysis of spatial-temporal lattice data. We propose a general asymptotic framework for spatial-temporal linear models and …


Analysis Of Binary Data Via Spatial-Temporal Autologistic Regression Models, Zilong Wang Jan 2012

Analysis Of Binary Data Via Spatial-Temporal Autologistic Regression Models, Zilong Wang

Theses and Dissertations--Statistics

Spatial-temporal autologistic models are useful models for binary data that are measured repeatedly over time on a spatial lattice. They can account for effects of potential covariates and spatial-temporal statistical dependence among the data. However, the traditional parametrization of spatial-temporal autologistic model presents difficulties in interpreting model parameters across varying levels of statistical dependence, where its non-negative autocovariates could bias the realizations toward 1. In order to achieve interpretable parameters, a centered spatial-temporal autologistic regression model has been developed. Two efficient statistical inference approaches, expectation-maximization pseudo-likelihood approach (EMPL) and Monte Carlo expectation-maximization likelihood approach (MCEML), have been proposed. Also, Bayesian …


Stochastic Dynamics Of Gene Transcription, Yan Xie Jan 2011

Stochastic Dynamics Of Gene Transcription, Yan Xie

Theses and Dissertations--Statistics

Gene transcription in individual living cells is inevitably a stochastic and dynamic process. Little is known about how cells and organisms learn to balance the fidelity of transcriptional control and the stochasticity of transcription dynamics. In an effort to elucidate the contribution of environmental signals to this intricate balance, a Three State Model was recently proposed, and the transcription system was assumed to transit among three different functional states randomly.

In this work, we employ this model to demonstrate how the stochastic dynamics of gene transcription can be characterized by the three transition parameters. We compute the probability distribution of …


Modeling Mass Transport In Aquifers: The Distributed Source Problem, Sergio E. Serrano Aug 1990

Modeling Mass Transport In Aquifers: The Distributed Source Problem, Sergio E. Serrano

KWRRI Research Reports

This report presents a new methodology to model the time and space evolution of groundwater variables in a system of aquifers when certain components of the model, such as the geohydrologic information, the boundary conditions, the magnitude and variability of the sources or physical parameters are uncertain and defined in stochastic terms. This facilitates a more realistic statistical representation of groundwater flow and groundwater pollution forecasting for either the saturated or the unsaturated zone. The method is based on applications of modern mathematics to the solution of the resulting stochastic transport equations. This procedure exhibits considerable advantages over the existing …


Improved Methods And Guidelines For Modeling Stormwater Runoff From Surface Coal Mined Lands, Michael E. Meadows, George E. Blandford Sep 1983

Improved Methods And Guidelines For Modeling Stormwater Runoff From Surface Coal Mined Lands, Michael E. Meadows, George E. Blandford

KWRRI Research Reports

The investgations, developments and guidelines for several hydrologic modeling strategies are presented. Investigations were conducted to determine appropriate event curve numbers for surface mined disturbed watersheds; and performance of four synthetic unit hydrograph models (SCS curvilinear, SCS single triangle, Williams and TVA double triangle) on 38 USDA experimental watersheds in 14 physiographic provinces using in excess of 270 events. A second test using only the SCS curvilinear unit hydrograph on 11 small watersheds and 48 events was conducted to investigate the excess rainfall pattern simulated with the curve number model. A procedure for developing a unit hydrograph using the time …


Modeling Surface And Subsurface Stormflow On Steeply-Sloping Forested Watersheds, Patrick G. Sloan, Ian D. Moore, George B. Coltharp, Joseph D. Eigel Jul 1983

Modeling Surface And Subsurface Stormflow On Steeply-Sloping Forested Watersheds, Patrick G. Sloan, Ian D. Moore, George B. Coltharp, Joseph D. Eigel

KWRRI Research Reports

A simple conceptual rainfall-runoff model, based on the variable source area concept, was developed for predicting runoff from small, steep-sloped, forested Appalachian watersheds. Tests of the model showed that the predicted and observed daily discharges were in good agreement. The results demonstrate the ability of the model to simulate the "flashy" hydrologic behavior of these watersheds.

Five subsurface flow models were evaluated by application to existing data measured at Coweeta on a reconstructed homogeneous forest soil. The five models were: Nieber 's 2-D and 1-D finite element models (based on Richards' equation), the kinematic wave equation, and two simple storage …


Opset Program For Computerized Selection Of Watershed Parameter Values For The Stanford Watershed Model, Earnest Yuan-Shang Liou, L. Douglas James Jan 1970

Opset Program For Computerized Selection Of Watershed Parameter Values For The Stanford Watershed Model, Earnest Yuan-Shang Liou, L. Douglas James

KWRRI Research Reports

The advent of high-speed electronic computer made it possible to model complex hydrologic processes by mathematical expressions and thereby simulate streamflows from climatological data. The most widely used program is the Stanford Watershed Model, a digital parametric model of the land phase of the hydrologic cycle based on moisture accounting processes. It can be used to simulate annual or longer flow sequences at hourly time intervals. Due to its capability of simulating historical streamflows from recorded climatological data, it has a great potential in the planning and design of water resources systems. However, widespread use of the Stanford Watershed Model …