Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (25)
- Statistical Methodology (22)
- Biostatistics (15)
- Multivariate Analysis (7)
- Environmental Sciences (6)
-
- Statistical Theory (6)
- Earth Sciences (5)
- Hydrology (5)
- Data Science (4)
- Life Sciences (4)
- Medicine and Health Sciences (4)
- Engineering (3)
- Longitudinal Data Analysis and Time Series (3)
- Microarrays (3)
- Social and Behavioral Sciences (3)
- Survival Analysis (3)
- Water Resource Management (3)
- Bioinformatics (2)
- Civil and Environmental Engineering (2)
- Geography (2)
- Public Health (2)
- Soil Science (2)
- Transportation Engineering (2)
- Animal Sciences (1)
- Biochemistry (1)
- Biochemistry, Biophysics, and Structural Biology (1)
- Biology (1)
- Keyword
-
- EM algorithm (4)
- Count data (3)
- Bayesian Analysis (2)
- Big Data (2)
- Bootstrap (2)
-
- Data dispersion (2)
- EM Algorithm (2)
- Generalized linear models (2)
- Interaction (2)
- Latent Class Analysis (2)
- Linear model (2)
- Multivariate Data (2)
- Normalization (2)
- ODE (2)
- Spatial-temporal process (2)
- Statistics (2)
- Sufficient dimension reduction (2)
- Variable Screening (2)
- Variable Selection (2)
- 3-D Highway Geometric Design (1)
- Adjustment (1)
- Algorithm (1)
- Alzheimer's Disease (1)
- Aquifers (1)
- Asymptotic distribution (1)
- Asymptotic properties (1)
- At-Fault (1)
- Autologistic regression models (1)
- Autoregressive models (1)
- Average Causal Effect (1)
- Publication Year
- Publication
- Publication Type
Articles 31 - 58 of 58
Full-Text Articles in Statistical Models
Improved Methods And Selecting Classification Types For Time-Dependent Covariates In The Marginal Analysis Of Longitudinal Data, I-Chen Chen
Theses and Dissertations--Epidemiology and Biostatistics
Generalized estimating equations (GEE) are popularly utilized for the marginal analysis of longitudinal data. In order to obtain consistent regression parameter estimates, these estimating equations must be unbiased. However, when certain types of time-dependent covariates are presented, these equations can be biased unless an independence working correlation structure is employed. Moreover, in this case regression parameter estimation can be very inefficient because not all valid moment conditions are incorporated within the corresponding estimating equations. Therefore, approaches using the generalized method of moments or quadratic inference functions have been proposed for utilizing all valid moment conditions. However, we have found that …
Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis
Accounting For Matching Uncertainty In Photographic Identification Studies Of Wild Animals, Amanda R. Ellis
Theses and Dissertations--Statistics
I consider statistical modelling of data gathered by photographic identification in mark-recapture studies and propose a new method that incorporates the inherent uncertainty of photographic identification in the estimation of abundance, survival and recruitment. A hierarchical model is proposed which accepts scores assigned to pairs of photographs by pattern recognition algorithms as data and allows for uncertainty in matching photographs based on these scores. The new models incorporate latent capture histories that are treated as unknown random variables informed by the data, contrasting past models having the capture histories being fixed. The methods properly account for uncertainty in the matching …
The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie
The Family Of Conditional Penalized Methods With Their Application In Sufficient Variable Selection, Jin Xie
Theses and Dissertations--Statistics
When scientists know in advance that some features (variables) are important in modeling a data, then these important features should be kept in the model. How can we utilize this prior information to effectively find other important features? This dissertation is to provide a solution, using such prior information. We propose the Conditional Adaptive Lasso (CAL) estimates to exploit this knowledge. By choosing a meaningful conditioning set, namely the prior information, CAL shows better performance in both variable selection and model estimation. We also propose Sufficient Conditional Adaptive Lasso Variable Screening (SCAL-VS) and Conditioning Set Sufficient Conditional Adaptive Lasso Variable …
Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang
Mixtures-Of-Regressions With Measurement Error, Xiaoqiong Fang
Theses and Dissertations--Statistics
Finite Mixture model has been studied for a long time, however, traditional methods assume that the variables are measured without error. Mixtures-of-regression model with measurement error imposes challenges to the statisticians, since both the mixture structure and the existence of measurement error can lead to inconsistent estimate for the regression coefficients. In order to solve the inconsistency, We propose series of methods to estimate the mixture likelihood of the mixtures-of-regressions model when there is measurement error, both in the responses and predictors. Different estimators of the parameters are derived and compared with respect to their relative efficiencies. The simulation results …
Estimation In Partially Linear Models With Correlated Observations And Change-Point Models, Liangdong Fan
Estimation In Partially Linear Models With Correlated Observations And Change-Point Models, Liangdong Fan
Theses and Dissertations--Statistics
Methods of estimating parametric and nonparametric components, as well as properties of the corresponding estimators, have been examined in partially linear models by Wahba [1987], Green et al. [1985], Engle et al. [1986], Speckman [1988], Hu et al. [2004], Charnigo et al. [2015] among others. These models are appealing due to their flexibility and wide range of practical applications including the electricity usage study by Engle et al. [1986], gum disease study by Speckman [1988], etc., wherea parametric component explains linear trends and a nonparametric part captures nonlinear relationships.
The compound estimator (Charnigo et al. [2015]) has been used to …
Some Dimension Reduction Strategies For The Analysis Of Survey Data, Jiaying Weng, Derek S. Young
Some Dimension Reduction Strategies For The Analysis Of Survey Data, Jiaying Weng, Derek S. Young
Statistics Faculty Publications
In the era of big data, researchers interested in developing statistical models are challenged with how to achieve parsimony. Usually, some sort of dimension reduction strategy is employed. Classic strategies are often in the form of traditional inference procedures, such as hypothesis testing; however, the increase in computing capabilities has led to the development of more sophisticated methods. In particular, sufficient dimension reduction has emerged as an area of broad and current interest. While these types of dimension reduction strategies have been employed for numerous data problems, they are scantly discussed in the context of analyzing survey data. This …
Improving The Computational Efficiency In Bayesian Fitting Of Cormack-Jolly-Seber Models With Individual, Continuous, Time-Varying Covariates, Woodrow Burchett
Improving The Computational Efficiency In Bayesian Fitting Of Cormack-Jolly-Seber Models With Individual, Continuous, Time-Varying Covariates, Woodrow Burchett
Theses and Dissertations--Statistics
The extension of the CJS model to include individual, continuous, time-varying covariates relies on the estimation of covariate values on occasions on which individuals were not captured. Fitting this model in a Bayesian framework typically involves the implementation of a Markov chain Monte Carlo (MCMC) algorithm, such as a Gibbs sampler, to sample from the posterior distribution. For large data sets with many missing covariate values that must be estimated, this creates a computational issue, as each iteration of the MCMC algorithm requires sampling from the full conditional distributions of each missing covariate value. This dissertation examines two solutions to …
Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu
Nonparametric Compound Estimation, Derivative Estimation, And Change Point Detection, Sisheng Liu
Theses and Dissertations--Statistics
Firstly, we reviewed some popular nonparameteric regression methods during the past several decades. Then we extended the compound estimation (Charnigo and Srinivasan [2011]) to adapt random design points and heteroskedasticity and proposed a modified Cp criteria for tuning parameter selection. Moreover, we developed a DCp criteria for tuning paramter selection problem in general nonparametric derivative estimation. This extends GCp criteria in Charnigo, Hall and Srinivasan [2011] with random design points and heteroskedasticity. Next, we proposed a change point detection method via compound estimation for both fixed design and random design case, the adaptation of heteroskedasticity was considered for the method. …
An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert
An Exploratory Statistical Method For Finding Interactions In A Large Dataset With An Application Toward Periodontal Diseases, Joshua Lambert
Theses and Dissertations--Epidemiology and Biostatistics
It is estimated that Periodontal Diseases effects up to 90% of the adult population. Given the complexity of the host environment, many factors contribute to expression of the disease. Age, Gender, Socioeconomic Status, Smoking Status, and Race/Ethnicity are all known risk factors, as well as a handful of known comorbidities. Certain vitamins and minerals have been shown to be protective for the disease, while some toxins and chemicals have been associated with an increased prevalence. The role of toxins, chemicals, vitamins, and minerals in relation to disease is believed to be complex and potentially modified by known risk factors. A …
Development In Normal Mixture And Mixture Of Experts Modeling, Meng Qi
Development In Normal Mixture And Mixture Of Experts Modeling, Meng Qi
Theses and Dissertations--Statistics
In this dissertation, first we consider the problem of testing homogeneity and order in a contaminated normal model, when the data is correlated under some known covariance structure. To address this problem, we developed a moment based homogeneity and order test, and design weights for test statistics to increase power for homogeneity test. We applied our test to microarray about Down’s syndrome. This dissertation also studies a singular Bayesian information criterion (sBIC) for a bivariate hierarchical mixture model with varying weights, and develops a new data dependent information criterion (sFLIC).We apply our model and criteria to birth- weight and gestational …
Multi-State Models With Missing Covariates, Wenjie Lou
Multi-State Models With Missing Covariates, Wenjie Lou
Theses and Dissertations--Statistics
Multi-state models have been widely used to analyze longitudinal event history data obtained in medical studies. The tools and methods developed recently in this area require the complete observed datasets. While, in many applications measurements on certain components of the covariate vector are missing on some study subjects. In this dissertation, several likelihood-based methodologies were proposed to deal with datasets with different types of missing covariates efficiently when applying multi-state models.
Firstly, a maximum observed data likelihood method was proposed when the data has a univariate missing pattern and the missing covariate is a categorical variable. The construction of the …
Topics In Logistic Regression Analysis, Zhiheng Xie
Topics In Logistic Regression Analysis, Zhiheng Xie
Theses and Dissertations--Statistics
Discrete-time Markov chains have been used to analyze the transition of subjects from intact cognition to dementia with mild cognitive impairment and global impairment as intervening transient states, and death as competing risk. A multinomial logistic regression model is used to estimate the probability distribution in each row of the one-step transition matrix that correspond to the transient states. We investigate some goodness of fit tests for a multinomial distribution with covariates to assess the fit of this model to the data. We propose a modified chi-square test statistic and a score test statistic for the multinomial assumption in each …
Developing An Alternative Way To Analyze Nanostring Data, Shu Shen
Developing An Alternative Way To Analyze Nanostring Data, Shu Shen
Theses and Dissertations--Statistics
Nanostring technology provides a new method to measure gene expressions. It's more sensitive than microarrays and able to do more gene measurements than RT-PCR with similar sensitivity. This system produces counts for each target gene and tabulates them. Counts can be normalized by using an Excel macro or nSolver before analysis. Both methods rely on data normalization prior to statistical analysis to identify differentially expressed genes. Alternatively, we propose to model gene expressions as a function of positive controls and reference gene measurements. Simulations and examples are used to compare this model with Nanostring normalization methods. The results show that …
Statistical Inference On Dynamical Systems, Hongyuan Wang
Statistical Inference On Dynamical Systems, Hongyuan Wang
Theses and Dissertations--Statistics
The ordinary differential equation (ODE) is one representative and popular tool in modeling dynamical systems, which are widely implemented in physics, biology, economics, chemistry and biomedical sciences, etc. Because of the importance of dynamical systems in scientific studies, they are the main focuses of my dissertation.
The first chapter of the dissertation is introduction and literature review, which mainly focuses on numerical integration algorithms of ODEs that are difficult to solve analytically, as well as derivative-free optimization algorithms for the so-called inverse problem.
The second chapter is on the estimation method based on numerical solvers of differential equations. We start …
Statistical Methods For Environmental Exposure Data Subject To Detection Limits, Yuchen Yang
Statistical Methods For Environmental Exposure Data Subject To Detection Limits, Yuchen Yang
Theses and Dissertations--Statistics
In this dissertation, we develop unified and efficient nonparametric statistical methods for estimating and comparing environmental exposure distributions in presence of detection limits. In the first part, we propose a kernel-smoothed nonparametric estimator for the exposure distribution without imposing any independence assumption between the exposure level and detection limit. We show that the proposed estimator is consistent and asymptotically normal. Simulation studies demonstrate that the proposed estimator performs well in practical situations. A colon cancer study is provided for illustration. In the second part, we develop a class of test statistics to compare exposure distributions between two groups by using …
Improved Models For Differential Analysis For Genomic Data, Hong Wang
Improved Models For Differential Analysis For Genomic Data, Hong Wang
Theses and Dissertations--Statistics
This paper intend to develop novel statistical methods to improve genomic data analysis, especially for differential analysis. We considered two different data type: NanoString nCounter data and somatic mutation data. For NanoString nCounter data, we develop a novel differential expression detection method. The method considers a generalized linear model of the negative binomial family to characterize count data and allows for multi-factor design. Data normalization is incorporated in the model framework through data normalization parameters, which are estimated from control genes embedded in the nCounter system. For somatic mutation data, we develop beta-binomial model-based approaches to identify highly or lowly …
Continuous Time Multi-State Models For Interval Censored Data, Lijie Wan
Continuous Time Multi-State Models For Interval Censored Data, Lijie Wan
Theses and Dissertations--Statistics
Continuous-time multi-state models are widely used in modeling longitudinal data of disease processes with multiple transient states, yet the analysis is complex when subjects are observed periodically, resulting in interval censored data. Recently, most studies focused on modeling the true disease progression as a discrete time stationary Markov chain, and only a few studies have been carried out regarding non-homogenous multi-state models in the presence of interval-censored data. In this dissertation, several likelihood-based methodologies were proposed to deal with interval censored data in multi-state models.
Firstly, a continuous time version of a homogenous Markov multi-state model with backward transitions was …
New Results In Ell_1 Penalized Regression, Edward A. Roualdes
New Results In Ell_1 Penalized Regression, Edward A. Roualdes
Theses and Dissertations--Statistics
Here we consider penalized regression methods, and extend on the results surrounding the l1 norm penalty. We address a more recent development that generalizes previous methods by penalizing a linear transformation of the coefficients of interest instead of penalizing just the coefficients themselves. We introduce an approximate algorithm to fit this generalization and a fully Bayesian hierarchical model that is a direct analogue of the frequentist version. A number of benefits are derived from the Bayesian persepective; most notably choice of the tuning parameter and natural means to estimate the variation of estimates – a notoriously difficult task for the …
The Initial Phases Of A Consistent Pricing System That Reflects The Online Sale Value Of A Horse, Curran A. Prettyman
The Initial Phases Of A Consistent Pricing System That Reflects The Online Sale Value Of A Horse, Curran A. Prettyman
Lewis Honors College Capstone Collection
Horses are one of the most uniquely priced commodities. This document provides a solution to an industry-wide weakness of inconsistent pricing and confusion. In the following report, an evaluation of the industry flaw is presented, an econometric approach is described in full, and a solution is proposed using insight gained from a regression analysis. This report uses an econometric approach to determine the impact of hunter jumper horse qualities on internet sale prices. Data is compiled from bigeq.com for seventy-eight horses in the states of Illinois, Indiana, Kentucky, Michigan, and Ohio. A linear regression analysis for twelve variables establishes that …
Genetic Association Testing Of Copy Number Variation, Yinglei Li
Genetic Association Testing Of Copy Number Variation, Yinglei Li
Theses and Dissertations--Statistics
Copy-number variation (CNV) has been implicated in many complex diseases. It is of great interest to detect and locate such regions through genetic association testings. However, the association testings are complicated by the fact that CNVs usually span multiple markers and thus such markers are correlated to each other. To overcome the difficulty, it is desirable to pool information across the markers. In this thesis, we propose a kernel-based method for aggregation of marker-level tests, in which first we obtain a bunch of p-values through association tests for every marker and then the association test involving CNV is based on …
Normal Mixture And Contaminated Model With Nuisance Parameter And Applications, Qian Fan
Normal Mixture And Contaminated Model With Nuisance Parameter And Applications, Qian Fan
Theses and Dissertations--Statistics
This paper intend to find the proper hypothesis and test statistic for testing existence of bilaterally contamination when there exists nuisance parameter. The test statistic is based on method of moments estimators. Union-Intersection test is used for testing if the distribution of population can be implemented by a bilaterally contaminated normal model with unknown variance. This paper also developed a hierarchical normal mixture model (HNM) and applied it to birth weight data. EM algorithm is employed for parameter estimation and a singular Bayesian information criterion (sBIC) is applied to choose the number components. We also proposed a singular flexible information …
Analysis Of Spatial Data, Xiang Zhang
Analysis Of Spatial Data, Xiang Zhang
Theses and Dissertations--Statistics
In many areas of the agriculture, biological, physical and social sciences, spatial lattice data are becoming increasingly common. In addition, a large amount of lattice data shows not only visible spatial pattern but also temporal pattern (see, Zhu et al. 2005). An interesting problem is to develop a model to systematically model the relationship between the response variable and possible explanatory variable, while accounting for space and time effect simultaneously.
Spatial-temporal linear model and the corresponding likelihood-based statistical inference are important tools for the analysis of spatial-temporal lattice data. We propose a general asymptotic framework for spatial-temporal linear models and …
Analysis Of Binary Data Via Spatial-Temporal Autologistic Regression Models, Zilong Wang
Analysis Of Binary Data Via Spatial-Temporal Autologistic Regression Models, Zilong Wang
Theses and Dissertations--Statistics
Spatial-temporal autologistic models are useful models for binary data that are measured repeatedly over time on a spatial lattice. They can account for effects of potential covariates and spatial-temporal statistical dependence among the data. However, the traditional parametrization of spatial-temporal autologistic model presents difficulties in interpreting model parameters across varying levels of statistical dependence, where its non-negative autocovariates could bias the realizations toward 1. In order to achieve interpretable parameters, a centered spatial-temporal autologistic regression model has been developed. Two efficient statistical inference approaches, expectation-maximization pseudo-likelihood approach (EMPL) and Monte Carlo expectation-maximization likelihood approach (MCEML), have been proposed. Also, Bayesian …
Stochastic Dynamics Of Gene Transcription, Yan Xie
Stochastic Dynamics Of Gene Transcription, Yan Xie
Theses and Dissertations--Statistics
Gene transcription in individual living cells is inevitably a stochastic and dynamic process. Little is known about how cells and organisms learn to balance the fidelity of transcriptional control and the stochasticity of transcription dynamics. In an effort to elucidate the contribution of environmental signals to this intricate balance, a Three State Model was recently proposed, and the transcription system was assumed to transit among three different functional states randomly.
In this work, we employ this model to demonstrate how the stochastic dynamics of gene transcription can be characterized by the three transition parameters. We compute the probability distribution of …
Modeling Mass Transport In Aquifers: The Distributed Source Problem, Sergio E. Serrano
Modeling Mass Transport In Aquifers: The Distributed Source Problem, Sergio E. Serrano
KWRRI Research Reports
This report presents a new methodology to model the time and space evolution of groundwater variables in a system of aquifers when certain components of the model, such as the geohydrologic information, the boundary conditions, the magnitude and variability of the sources or physical parameters are uncertain and defined in stochastic terms. This facilitates a more realistic statistical representation of groundwater flow and groundwater pollution forecasting for either the saturated or the unsaturated zone. The method is based on applications of modern mathematics to the solution of the resulting stochastic transport equations. This procedure exhibits considerable advantages over the existing …
Improved Methods And Guidelines For Modeling Stormwater Runoff From Surface Coal Mined Lands, Michael E. Meadows, George E. Blandford
Improved Methods And Guidelines For Modeling Stormwater Runoff From Surface Coal Mined Lands, Michael E. Meadows, George E. Blandford
KWRRI Research Reports
The investgations, developments and guidelines for several hydrologic modeling strategies are presented. Investigations were conducted to determine appropriate event curve numbers for surface mined disturbed watersheds; and performance of four synthetic unit hydrograph models (SCS curvilinear, SCS single triangle, Williams and TVA double triangle) on 38 USDA experimental watersheds in 14 physiographic provinces using in excess of 270 events. A second test using only the SCS curvilinear unit hydrograph on 11 small watersheds and 48 events was conducted to investigate the excess rainfall pattern simulated with the curve number model. A procedure for developing a unit hydrograph using the time …
Modeling Surface And Subsurface Stormflow On Steeply-Sloping Forested Watersheds, Patrick G. Sloan, Ian D. Moore, George B. Coltharp, Joseph D. Eigel
Modeling Surface And Subsurface Stormflow On Steeply-Sloping Forested Watersheds, Patrick G. Sloan, Ian D. Moore, George B. Coltharp, Joseph D. Eigel
KWRRI Research Reports
A simple conceptual rainfall-runoff model, based on the variable source area concept, was developed for predicting runoff from small, steep-sloped, forested Appalachian watersheds. Tests of the model showed that the predicted and observed daily discharges were in good agreement. The results demonstrate the ability of the model to simulate the "flashy" hydrologic behavior of these watersheds.
Five subsurface flow models were evaluated by application to existing data measured at Coweeta on a reconstructed homogeneous forest soil. The five models were: Nieber 's 2-D and 1-D finite element models (based on Richards' equation), the kinematic wave equation, and two simple storage …
Opset Program For Computerized Selection Of Watershed Parameter Values For The Stanford Watershed Model, Earnest Yuan-Shang Liou, L. Douglas James
Opset Program For Computerized Selection Of Watershed Parameter Values For The Stanford Watershed Model, Earnest Yuan-Shang Liou, L. Douglas James
KWRRI Research Reports
The advent of high-speed electronic computer made it possible to model complex hydrologic processes by mathematical expressions and thereby simulate streamflows from climatological data. The most widely used program is the Stanford Watershed Model, a digital parametric model of the land phase of the hydrologic cycle based on moisture accounting processes. It can be used to simulate annual or longer flow sequences at hourly time intervals. Due to its capability of simulating historical streamflows from recorded climatological data, it has a great potential in the planning and design of water resources systems. However, widespread use of the Stanford Watershed Model …