Open Access. Powered by Scholars. Published by Universities.®

Applied Statistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

University of Kentucky

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 31 - 54 of 54

Full-Text Articles in Applied Statistics

Modeling And Mapping Location-Dependent Human Appearance, Zachary Bessinger Jan 2018

Modeling And Mapping Location-Dependent Human Appearance, Zachary Bessinger

Theses and Dissertations--Computer Science

Human appearance is highly variable and depends on individual preferences, such as fashion, facial expression, and makeup. These preferences depend on many factors including a person's sense of style, what they are doing, and the weather. These factors, in turn, are dependent upon geographic location and time. In our work, we build computational models to learn the relationship between human appearance, geographic location, and time. The primary contributions are a framework for collecting and processing geotagged imagery of people, a large dataset collected by our framework, and several generative and discriminative models that use our dataset to learn the relationship …


Occurrence And Attributes Of Two Echinoderm-Bearing Faunas From The Upper Mississippian (Chesterian; Lower Serpukhovian) Ramey Creek Member, Slade Formation, Eastern Kentucky, U.S.A., Ann Well Harris Jan 2018

Occurrence And Attributes Of Two Echinoderm-Bearing Faunas From The Upper Mississippian (Chesterian; Lower Serpukhovian) Ramey Creek Member, Slade Formation, Eastern Kentucky, U.S.A., Ann Well Harris

Theses and Dissertations--Earth and Environmental Sciences

Well-preserved echinoderm faunas are rare in the fossil record, and when uncovered, understanding their occurrence can be useful in interpreting other faunas. In this study, two such faunas of the same age from separate localities in the shallow-marine Ramey Creek Member of the Slade Formation in the Upper Mississippian (Chesterian) rocks of eastern Kentucky are examined. Of the more than 5,000 fossil specimens from both localities, only 9–34 percent were echinoderms from 3–5 classes. Nine non-echinoderm (8 invertebrate and one vertebrate) classes occurred at both localities, but of these, bryozoans, brachiopods and sponges dominated. To understand the attributes of both …


Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis Jan 2018

Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis

Theses and Dissertations--Statistics

I consider statistical modelling of data gathered by photographic identification in mark-recapture studies and propose a new method that incorporates the inherent uncertainty of photographic identification in the estimation of abundance, survival and recruitment. A hierarchical model is proposed which accepts scores assigned to pairs of photographs by pattern recognition algorithms as data and allows for uncertainty in matching photographs based on these scores. The new models incorporate latent capture histories that are treated as unknown random variables informed by the data, contrasting past models having the capture histories being fixed. The methods properly account for uncertainty in the matching …


The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie Jan 2018

The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie

Theses and Dissertations--Statistics

When scientists know in advance that some features (variables) are important in modeling a data, then these important features should be kept in the model. How can we utilize this prior information to effectively find other important features? This dissertation is to provide a solution, using such prior information. We propose the Conditional Adaptive Lasso (CAL) estimates to exploit this knowledge. By choosing a meaningful conditioning set, namely the prior information, CAL shows better performance in both variable selection and model estimation. We also propose Sufficient Conditional Adaptive Lasso Variable Screening (SCAL-VS) and Conditioning Set Sufficient Conditional Adaptive Lasso Variable …


Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang Jan 2018

Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang

Theses and Dissertations--Statistics

Finite Mixture model has been studied for a long time, however, traditional methods assume that the variables are measured without error. Mixtures-of-regression model with measurement error imposes challenges to the statisticians, since both the mixture structure and the existence of measurement error can lead to inconsistent estimate for the regression coefficients. In order to solve the inconsistency, We propose series of methods to estimate the mixture likelihood of the mixtures-of-regressions model when there is measurement error, both in the responses and predictors. Different estimators of the parameters are derived and compared with respect to their relative efficiencies. The simulation results …


Improved Standard Error Estimation For Maintaining The Validities Of Inference In Small-Sample Cluster Randomized Trials And Longitudinal Studies, Whitney Ford Tanner Jan 2018

Improved Standard Error Estimation For Maintaining The Validities Of Inference In Small-Sample Cluster Randomized Trials And Longitudinal Studies, Whitney Ford Tanner

Theses and Dissertations--Epidemiology and Biostatistics

Data arising from Cluster Randomized Trials (CRTs) and longitudinal studies are correlated and generalized estimating equations (GEE) are a popular analysis method for correlated data. Previous research has shown that analyses using GEE could result in liberal inference due to the use of the empirical sandwich covariance matrix estimator, which can yield negatively biased standard error estimates when the number of clusters or subjects is not large. Many techniques have been presented to correct this negative bias; However, use of these corrections can still result in biased standard error estimates and thus test sizes that are not consistently at their …


Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu Jan 2017

Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu

Theses and Dissertations--Statistics

Firstly, we reviewed some popular nonparameteric regression methods during the past several decades. Then we extended the compound estimation (Charnigo and Srinivasan [2011]) to adapt random design points and heteroskedasticity and proposed a modified Cp criteria for tuning parameter selection. Moreover, we developed a DCp criteria for tuning paramter selection problem in general nonparametric derivative estimation. This extends GCp criteria in Charnigo, Hall and Srinivasan [2011] with random design points and heteroskedasticity. Next, we proposed a change point detection method via compound estimation for both fixed design and random design case, the adaptation of heteroskedasticity was considered for the method. …


An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert Jan 2017

An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert

Theses and Dissertations--Epidemiology and Biostatistics

It is estimated that Periodontal Diseases effects up to 90% of the adult population. Given the complexity of the host environment, many factors contribute to expression of the disease. Age, Gender, Socioeconomic Status, Smoking Status, and Race/Ethnicity are all known risk factors, as well as a handful of known comorbidities. Certain vitamins and minerals have been shown to be protective for the disease, while some toxins and chemicals have been associated with an increased prevalence. The role of toxins, chemicals, vitamins, and minerals in relation to disease is believed to be complex and potentially modified by known risk factors. A …


Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan Jan 2017

Informational Index And Its Applications In High Dimensional Data, Qingcong Yuan

Theses and Dissertations--Statistics

We introduce a new class of measures for testing independence between two random vectors, which uses expected difference of conditional and marginal characteristic functions. By choosing a particular weight function in the class, we propose a new index for measuring independence and study its property. Two empirical versions are developed, their properties, asymptotics, connection with existing measures and applications are discussed. Implementation and Monte Carlo results are also presented.

We propose a two-stage sufficient variable selections method based on the new index to deal with large p small n data. The method does not require model specification and especially focuses …


Extending The Latent Multinomial Model With Complex Error Processes And Dynamic Markov Bases, Simon J. Bonner, Matthew R. Schofield, Patrik Noren, Steven J. Price Jan 2016

Extending The Latent Multinomial Model With Complex Error Processes And Dynamic Markov Bases, Simon J. Bonner, Matthew R. Schofield, Patrik Noren, Steven J. Price

Forestry and Natural Resources Faculty Publications

The latent multinomial model (LMM) of Link et al. [Biometrics 66 (2010) 178–185] provides a framework for modelling mark-recapture data with potential identification errors. Key is a Markov chain Monte Carlo (MCMC) scheme for sampling configurations of the latent counts of the true capture histories that could have generated the observed data. Assuming a linear map between the observed and latent counts, the MCMC algorithm uses vectors from a basis of the kernel to move between configurations of the latent data. Schofield and Bonner [Biometrics 71 (2015) 1070–1080] shows that this is sufficient for some models within the …


Multi-State Models With Missing Covariates, Wenjie Lou Jan 2016

Multi-State Models With Missing Covariates, Wenjie Lou

Theses and Dissertations--Statistics

Multi-state models have been widely used to analyze longitudinal event history data obtained in medical studies. The tools and methods developed recently in this area require the complete observed datasets. While, in many applications measurements on certain components of the covariate vector are missing on some study subjects. In this dissertation, several likelihood-based methodologies were proposed to deal with datasets with different types of missing covariates efficiently when applying multi-state models.

Firstly, a maximum observed data likelihood method was proposed when the data has a univariate missing pattern and the missing covariate is a categorical variable. The construction of the …


Statistical Methods For Handling Intentional Inaccurate Responders, Kristen J. Mcquerry Jan 2016

Statistical Methods For Handling Intentional Inaccurate Responders, Kristen J. Mcquerry

Theses and Dissertations--Statistics

In self-report data, participants who provide incorrect responses are known as intentional inaccurate responders. This dissertation provides statistical analyses for address intentional inaccurate responses in the data.

Previous work with adolescent self-report, labeled survey participants who intentionally provide inaccurate answers as mischievous responders. This phenomenon also occurs in clinical research. For example, pregnant women who smoke may report that they are nonsmokers. Our advantage is that we do not solely have self-report answers and can verify responses with lab values. Currently, there is no clear method for handling these intentional inaccurate respondents when it comes to making statistical inferences.

We …


Empirical Likelihood And Differentiable Functionals, Zhiyuan Shen Jan 2016

Empirical Likelihood And Differentiable Functionals, Zhiyuan Shen

Theses and Dissertations--Statistics

Empirical likelihood (EL) is a recently developed nonparametric method of statistical inference. It has been shown by Owen (1988,1990) and many others that empirical likelihood ratio (ELR) method can be used to produce nice confidence intervals or regions. Owen (1988) shows that -2logELR converges to a chi-square distribution with one degree of freedom subject to a linear statistical functional in terms of distribution functions. However, a generalization of Owen's result to the right censored data setting is difficult since no explicit maximization can be obtained under constraint in terms of distribution functions. Pan and Zhou (2002), instead, study the …


Multi-State Models For Interval Censored Data With Competing Risk, Shaoceng Wei Jan 2015

Multi-State Models For Interval Censored Data With Competing Risk, Shaoceng Wei

Theses and Dissertations--Statistics

Multi-state models are often used to evaluate the effect of death as a competing event to the development of dementia in a longitudinal study of the cognitive status of elderly subjects. In this dissertation, both multi-state Markov model and semi-Markov model are used to characterize the flow of subjects from intact cognition to dementia with mild cognitive impairment and global impairment as intervening transient, cognitive states and death as a competing risk.

Firstly, a multi-state Markov model with three transient states: intact cognition, mild cognitive impairment (M.C.I.) and global impairment (G.I.) and one absorbing state: dementia is used to model …


New Results In Ell_1 Penalized Regression, Edward A. Roualdes Jan 2015

New Results In Ell_1 Penalized Regression, Edward A. Roualdes

Theses and Dissertations--Statistics

Here we consider penalized regression methods, and extend on the results surrounding the l1 norm penalty. We address a more recent development that generalizes previous methods by penalizing a linear transformation of the coefficients of interest instead of penalizing just the coefficients themselves. We introduce an approximate algorithm to fit this generalization and a fully Bayesian hierarchical model that is a direct analogue of the frequentist version. A number of benefits are derived from the Bayesian persepective; most notably choice of the tuning parameter and natural means to estimate the variation of estimates – a notoriously difficult task for the …


Developments In Nonparametric Regression Methods With Application To Raman Spectroscopy Analysis, Jing Guo Jan 2015

Developments In Nonparametric Regression Methods With Application To Raman Spectroscopy Analysis, Jing Guo

Theses and Dissertations--Epidemiology and Biostatistics

Raman spectroscopy has been successfully employed in the classification of breast pathologies involving basis spectra for chemical constituents of breast tissue and resulted in high sensitivity (94%) and specificity (96%) (Haka et al, 2005). Motivated by recent developments in nonparametric regression, in this work, we adapt stacking, boosting, and dynamic ensemble learning into a nonparametric regression framework with application to Raman spectroscopy analysis for breast cancer diagnosis. In Chapter 2, we apply compound estimation (Charnigo and Srinivasan, 2011) in Raman spectra analysis to classify normal, benign, and malignant breast tissue. We explore both the spectra profiles and their derivatives to …


Statistics In The Billera-Holmes-Vogtmann Treespace, Grady S. Weyenberg Jan 2015

Statistics In The Billera-Holmes-Vogtmann Treespace, Grady S. Weyenberg

Theses and Dissertations--Statistics

This dissertation is an effort to adapt two classical non-parametric statistical techniques, kernel density estimation (KDE) and principal components analysis (PCA), to the Billera-Holmes-Vogtmann (BHV) metric space for phylogenetic trees. This adaption gives a more general framework for developing and testing various hypotheses about apparent differences or similarities between sets of phylogenetic trees than currently exists.

For example, while the majority of gene histories found in a clade of organisms are expected to be generated by a common evolutionary process, numerous other coexisting processes (e.g. horizontal gene transfers, gene duplication and subsequent neofunctionalization) will cause some genes to exhibit a …


The Initial Phases Of A Consistent Pricing System That Reflects The Online Sale Value Of A Horse, Curran A. Prettyman Jan 2014

The Initial Phases Of A Consistent Pricing System That Reflects The Online Sale Value Of A Horse, Curran A. Prettyman

Lewis Honors College Capstone Collection

Horses are one of the most uniquely priced commodities. This document provides a solution to an industry-wide weakness of inconsistent pricing and confusion. In the following report, an evaluation of the industry flaw is presented, an econometric approach is described in full, and a solution is proposed using insight gained from a regression analysis. This report uses an econometric approach to determine the impact of hunter jumper horse qualities on internet sale prices. Data is compiled from bigeq.com for seventy-eight horses in the states of Illinois, Indiana, Kentucky, Michigan, and Ohio. A linear regression analysis for twelve variables establishes that …


Genetic Association Testing Of Copy Number Variation, Yinglei Li Jan 2014

Genetic Association Testing Of Copy Number Variation, Yinglei Li

Theses and Dissertations--Statistics

Copy-number variation (CNV) has been implicated in many complex diseases. It is of great interest to detect and locate such regions through genetic association testings. However, the association testings are complicated by the fact that CNVs usually span multiple markers and thus such markers are correlated to each other. To overcome the difficulty, it is desirable to pool information across the markers. In this thesis, we propose a kernel-based method for aggregation of marker-level tests, in which first we obtain a bunch of p-values through association tests for every marker and then the association test involving CNV is based on …


Fusarium Head Blight Resistance And Agronomic Performance In Soft Red Winter Wheat Populations, Daniela Sarti Dvorjak Jan 2014

Fusarium Head Blight Resistance And Agronomic Performance In Soft Red Winter Wheat Populations, Daniela Sarti Dvorjak

Theses and Dissertations--Plant and Soil Sciences

Fusarium head blight (FHB), caused by Fusarium graminearum Schwabe [telomorph: Gibberella zeae Schwein.(Petch)], is recognized as one of the most destructive diseases of wheat (Triticum aestivum L. and T. durum L.) and barley (Hordeum vulgare L.) worldwide. Breeding for FHB resistance must be accompanied by selection for desirable agronomic traits. Donor parents with two FHB resistance quantitative trait loci (QTL) Fhb1 (chromosome 3BS) and QFhs.nau-2DL (chromosome 2DL) were crossed to four adapted SRW wheat lines to generate backcross and forward cross progeny. F2 individuals were genotyped and assigned to 4 different groups according to presence/ absence of …


The Psychological Impacts Of False Positive Ovarian Cancer Screening: Assessment Via Mixed And Trajectory Modeling, Amanda T. Wiggins Jan 2013

The Psychological Impacts Of False Positive Ovarian Cancer Screening: Assessment Via Mixed And Trajectory Modeling, Amanda T. Wiggins

Theses and Dissertations--Epidemiology and Biostatistics

Ovarian cancer (OC) is the fifth most common cancer among women and has the highest mortality of any cancer of the female reproductive system. The majority (61%) of OC cases are diagnosed at a distant stage. Because diagnoses occur most commonly at a late-stage and prognosis for advanced disease is poor, research focusing on the development of effective OC screening methods to facilitate early detection in high-risk, asymptomatic women is fundamental in reducing OC-specific mortality. Presently, there is no screening modality proven efficacious in reducing OC-mortality. However, transvaginal ultrasonography (TVS) has shown value in early detection of OC. TVS presents …


Bayesian Semiparametric Generalizations Of Linear Models Using Polya Trees, Angela Schoergendorfer Jan 2011

Bayesian Semiparametric Generalizations Of Linear Models Using Polya Trees, Angela Schoergendorfer

University of Kentucky Doctoral Dissertations

In a Bayesian framework, prior distributions on a space of nonparametric continuous distributions may be defined using Polya trees. This dissertation addresses statistical problems for which the Polya tree idea can be utilized to provide efficient and practical methodological solutions.

One problem considered is the estimation of risks, odds ratios, or other similar measures that are derived by specifying a threshold for an observed continuous variable. It has been previously shown that fitting a linear model to the continuous outcome under the assumption of a logistic error distribution leads to more efficient odds ratio estimates. We will show that deviations …


Statistical Methods In Microarray Data Analysis, Liping Huang Jan 2009

Statistical Methods In Microarray Data Analysis, Liping Huang

University of Kentucky Doctoral Dissertations

This dissertation includes three topics. First topic: Regularized estimation in the AFT model with high dimensional covariates. Second topic: A novel application of quantile regression for identification of biomarkers exemplified by equine cartilage microarray data. Third topic: Normalization and analysis of cDNA microarray using linear contrasts.


Empirical Processes And Roc Curves With An Application To Linear Combinations Of Diagnostic Tests, Costel Chirila Jan 2008

Empirical Processes And Roc Curves With An Application To Linear Combinations Of Diagnostic Tests, Costel Chirila

University of Kentucky Doctoral Dissertations

The Receiver Operating Characteristic (ROC) curve is the plot of Sensitivity vs. 1- Specificity of a quantitative diagnostic test, for a wide range of cut-off points c. The empirical ROC curve is probably the most used nonparametric estimator of the ROC curve. The asymptotic properties of this estimator were first developed by Hsieh and Turnbull (1996) based on strong approximations for quantile processes. Jensen et al. (2000) provided a general method to obtain regional confidence bands for the empirical ROC curve, based on its asymptotic distribution.

Since most biomarkers do not have high enough sensitivity and specificity to …