Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Applied Statistics (38)
- Life Sciences (17)
- Statistical Models (17)
- Statistical Methodology (13)
- Biostatistics (12)
-
- Engineering (12)
- Categorical Data Analysis (8)
- Industrial Engineering (8)
- Operations Research, Systems Engineering and Industrial Engineering (8)
- Social and Behavioral Sciences (7)
- Education (6)
- Medicine and Health Sciences (6)
- Multivariate Analysis (5)
- Probability (5)
- Statistical Theory (5)
- Educational Assessment, Evaluation, and Research (4)
- Geography (4)
- Longitudinal Data Analysis and Time Series (4)
- Operational Research (4)
- Applied Mathematics (3)
- Cancer Biology (3)
- Cell and Developmental Biology (3)
- Genetics and Genomics (3)
- Medical Specialties (3)
- Plant Sciences (3)
- Remote Sensing (3)
- Vital and Health Statistics (3)
- Agriculture (2)
- Keyword
-
- Pure sciences (14)
- Bayesian (6)
- Statistics (6)
- Applied sciences (4)
- Bayesian Statistics (3)
-
- Differential Item Functioning (3)
- Horseshoe (3)
- Bayesian Analysis (2)
- Biological sciences (2)
- Clinical trials (2)
- Conditional autoregressive prior (2)
- Distance Correlation (2)
- Item response theory (2)
- Landsat time series (2)
- MCMC (2)
- Measurement invariance (2)
- Ovarian cancer (2)
- Poisson (2)
- Psychometrics (2)
- Quality control (2)
- Regression (2)
- Statistical Learning (2)
- Survival analysis (2)
- Variable selection (2)
- 5-fluorouracil (1)
- Aberrant responding (1)
- Adaptation (1)
- Aggregate failure-time data (1)
- Analytics (1)
- Appraisals (1)
Articles 61 - 84 of 84
Full-Text Articles in Statistics and Probability
Budget-Constrained Regression Model Selection Using Mixed Integer Nonlinear Programming, Jingying Zhang
Budget-Constrained Regression Model Selection Using Mixed Integer Nonlinear Programming, Jingying Zhang
Graduate Theses and Dissertations
Regression analysis fits predictive models to data on a response variable and corresponding values for a set of explanatory variables. Often data on the explanatory variables come at a cost from commercial databases, so the available budget may limit which ones are used in the final model.
In this dissertation, two budget-constrained regression models are proposed for continuous and categorical variables respectively using Mixed Integer Nonlinear Programming (MINLP) to choose the explanatory variables to be included in solutions. First, we propose a budget-constrained linear regression model for continuous response variables. Properties such as solvability and global optimality of the proposed …
Spatio-Temporal Reconstruction Of Remote Sensing Observations, Kamrul Khan
Spatio-Temporal Reconstruction Of Remote Sensing Observations, Kamrul Khan
Graduate Theses and Dissertations
The USDA Forest Service aims to use satellite imagery for monitoring and predicting changes in forest conditions over time within the country. We specifically focus on a 230, 400 hectares region in north-central Wisconsin between 2003 - 2012. The auxiliary data collected from the satellite imagery of this region are relatively dense in space and time and can be used to efficiently predict how the forest condition changed over that decade. However, these records have a significant proportion of missing values due to weather conditions and system failures. To fill in these missing values, we build spaciotemporal models based on …
Comparison Of Correlation, Partial Correlation, And Conditional Mutual Information For Interaction Effects Screening In Generalized Linear Models, Ji Li
Graduate Theses and Dissertations
Numerous screening techniques have been developed in recent years for genome-wide association studies (GWASs) (Moore et al., 2010). In this thesis, a novel model-free screening method was developed and validated by an extensive simulation study. Many screening methods were mainly focused on main effects, while very few studies considered the models containing both main effects and interaction effects. In this work, the interaction effects were fully considered and three different methods (Pearson’s Correlation Coefficient, Partial Correlation, and Conditional Mutual Information) were tested and their prediction accuracies were compared.
Pearson’s Correlation Coefficient method, which is a direct interaction screening (DIS) procedure, …
Adapting To Sparsity And Heavy Tailed Data, Mohamed Abdelkader Abba
Adapting To Sparsity And Heavy Tailed Data, Mohamed Abdelkader Abba
Graduate Theses and Dissertations
The Lasso and the Horseshoe, gold-standards in the frequentist and Bayesian paradigms, critically depend on learning the error variance. This causes a lack of scale invariance and adaptability to heavy-tailed data. The √ Lasso [Belloni et al., 2011] attempt to correct this by using the `1 norm on both the likelihood and the penalty for the objective function. In contrast, there is essentially no methods for uncertainty quantification or automatic parameter tuning via a formal Bayesian treatment of an unknown error distribution. On the other hand, Bayesian shrinkage priors lacking a local shrinkage term fails to adapt to the large …
Hierarchical Bayesian Regression With Application In Spatial Modeling And Outlier Detection, Ghadeer Mahdi
Hierarchical Bayesian Regression With Application In Spatial Modeling And Outlier Detection, Ghadeer Mahdi
Graduate Theses and Dissertations
This dissertation makes two important contributions to the development of Bayesian hierarchical models. The first contribution is focused on spatial modeling. Spatial data observed on a group of areal units is common in scientific applications. The usual hierarchical approach for modeling this kind of dataset is to introduce a spatial random effect with an autoregressive prior. However, the usual Markov chain Monte Carlo scheme for this hierarchical framework requires the spatial effects to be sampled from their full conditional posteriors one-by-one resulting in poor mixing. More importantly, it makes the model computationally inefficient for datasets with large number of units. …
Bayesian Model For Detection Of Outliers In Linear Regression With Application To Longitudinal Data, Zahraa Al-Sharea
Bayesian Model For Detection Of Outliers In Linear Regression With Application To Longitudinal Data, Zahraa Al-Sharea
Graduate Theses and Dissertations
Outlier detection is one of the most important challenges with many present-day applications. Outliers can occur due to uncertainty in data generating mechanisms or due to an error in data recording/processing. Outliers can drastically change the study's results and make predictions less reliable. Detecting outliers in longitudinal studies is quite challenging because this kind of study is working with observations that change over time. Therefore, the same subject can produce an outlier at one point in time produce regular observations at all other time points. A Bayesian hierarchical modeling assigns parameters that can quantify whether each observation is an outlier …
A Linear-Linear Growth Model With Individual Change Point And Its Application To Ecls-K Data, Ping Zhang
A Linear-Linear Growth Model With Individual Change Point And Its Application To Ecls-K Data, Ping Zhang
Graduate Theses and Dissertations
The latent growth curve model with piecewise functions is a useful analytics tool to investigate the growth trajectory consisted of distinct phases of development in observed variables. An interesting feature of the growth trajectory is the time point that the trajectory changes from one phase to another one. In this thesis, we propose a simple computational pipeline to locate the change point under the linear-linear piecewise model and apply it to the longitudinal study of reading and math ability in early childhood (from kindergarten to eighth grade). In the first step, we conduct the hypothesis testing to filter out the …
Identifying Three-Way Gene Interactions From Microarray Data Using Kolmogorov-Smirnov And Cross-Match Tests, Shubhashree Khadka
Identifying Three-Way Gene Interactions From Microarray Data Using Kolmogorov-Smirnov And Cross-Match Tests, Shubhashree Khadka
Graduate Theses and Dissertations
Human gene network is much more complex than just pairwise interaction among the genes. Zhang et al. [6] extracted microarray data from International Genomics Consortium (IGC), and presented the detection of three-way gene interactions in their paper using Fisher’s z-transformation test. Three-way gene interactions are closer than pairwise correlations in representing the complex gene structures. Additionally, it was more tractable than assessing four or more gene interactions. In this paper, we are simulating different models where Fisher’s test might not be as effective. Zhang et al.’s approach utilized Pearson’s correlation coefficients and involved detection of linear interactions only. Since gene …
Genomic And Physiological Approaches To Improve Drought Tolerance In Soybean, Avjinder Kaler
Genomic And Physiological Approaches To Improve Drought Tolerance In Soybean, Avjinder Kaler
Graduate Theses and Dissertations
Drought stress is a major global constraint for crop production, and improving crop tolerance to drought is of critical importance. Direct selection of drought tolerance among genotypes for yield is limited because of low heritability, polygenic control, epistasis effects, and genotype by environment interactions. Crop physiology can play a major role for improving drought tolerance through the identification of traits associated with drought tolerance that can be used as indirect selection criteria in a breeding program. Carbon isotope ratio (δ13C, associated with water use efficiency), oxygen isotope ratio (δ18O, associated with transpiration), canopy temperature (CT), canopy wilting, and canopy coverage …
A Bayesian Variable Selection Method With Applications To Spatial Data, Xiahan Tang
A Bayesian Variable Selection Method With Applications To Spatial Data, Xiahan Tang
Graduate Theses and Dissertations
This thesis first describes the general idea behind Bayes Inference, various sampling methods based on Bayes theorem and many examples. Then a Bayes approach to model selection, called Stochastic Search Variable Selection (SSVS) is discussed. It was originally proposed by George and McCulloch (1993). In a normal regression model where the number of covariates is large, only a small subset tend to be significant most of the times. This Bayes procedure specifies a mixture prior for each of the unknown regression coefficient, the mixture prior was originally proposed by Geweke (1996). This mixture prior will be updated as data becomes …
Monte Carlo Methods In Bayesian Inference: Theory, Methods And Applications, Huarui Zhang
Monte Carlo Methods In Bayesian Inference: Theory, Methods And Applications, Huarui Zhang
Graduate Theses and Dissertations
Monte Carlo methods are becoming more and more popular in statistics due to the fast development of efficient computing technologies. One of the major beneficiaries of this advent is the field of Bayesian inference. The aim of this thesis is two-fold: (i) to explain the theory justifying the validity of the simulation-based schemes in a Bayesian setting (why they should work) and (ii) to apply them in several different types of data analysis that a statistician has to routinely encounter. In Chapter 1, I introduce key concepts in Bayesian statistics. Then we discuss Monte Carlo Simulation methods in detail. Our …
Analysis Of Break-Points In Financial Time Series, Jean Remy Habimana
Analysis Of Break-Points In Financial Time Series, Jean Remy Habimana
Graduate Theses and Dissertations
A time series is a set of random values collected at equal time intervals; this randomness makes these types of series not easy to predict because the structure of the series may change at any time. As discussed in previous research, the structure of time series may change at any time due to the change in mean and/or variance of the series. Consequently, based on this structure, it is wise not to assume that these series are stationary. This paper, discusses, a method of analyzing time series by considering the entire series non-stationary, assuming there is random change in unconditional …
Statistical Modeling Of The Temporal Dynamics In A Large Scale-Citation Network, Luis Javier Ek Jr.
Statistical Modeling Of The Temporal Dynamics In A Large Scale-Citation Network, Luis Javier Ek Jr.
Graduate Theses and Dissertations
Citation Networks of papers are vast networks that grow over time. The manner or the form a citation network grows is not entirely a random process, but a preferential attachment relationship; highly cited papers are more likely to be cited by newly published papers. The result is a network whose degree distribution follows a power law. This growth of citation network of papers will be modeled with a negative binomial regression coupled with logistic growth and/or Cauchy distribution curve. Then a Barabasi-Albert model, based on the negative binomial models, and a combination of the Dirichlet distribution and multinomial will be …
Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai
Identification Of Biomarkers For The Overall Survival Of Ovarian Cancer Patients, Kristi Mai
Graduate Theses and Dissertations
Rapid advance in sequencing technology has led to genome-wide analysis of genetic and epigenetic features simultaneously, making it possible to understand the biological mechanisms underlying cancer initiation and progression. However, how to identify important prognostic features poses a great challenge for both statistical modeling and computing. In this thesis, a network-based approach is applied to the Cancer Genome Atlas (TCGA) ovarian cancer data to identify important genes related to the overall survival of ovarian cancer patients. In the first step, a stepwise correlation-based selector is used to reduce the dimensionality of TCGA data, by filtering out a large number of …
Risk Estimation Toward A Natural History Model For Low Grade Glioma Patients, Anh Thi Hoang Pham
Risk Estimation Toward A Natural History Model For Low Grade Glioma Patients, Anh Thi Hoang Pham
Graduate Theses and Dissertations
Glioma is a common type of primary brain tumor that represents 28% of all brain tumors and 80% of malignant tumors. According to a recent study by the Centers for Disease Control and Prevention (CDC), gliomas account for 53%, 35% and 29% of all brain tumors (68%, 74% and 81% of malignant brain tumors) among children (aged 0-14), teenagers (aged 15-19) and young adults, respectively. Gliomas are often diagnosed through radiological imaging and histopathology. There are two main groups of gliomas following World Health Organization’s classification: Low grade gliomas (LGG), or grade I and II gliomas; and high grade gliomas …
Spread Trading In Corn Futures Market, Ryan D. Napier
Spread Trading In Corn Futures Market, Ryan D. Napier
Graduate Theses and Dissertations
The non-linear relationship between old crop – new crop year spreads in corn futures market and stock-to-use (S-U) ratios published by the United States Department of Agriculture is analyzed. Using a non-linear logarithmic smooth transition regression (LSTR) model, we capture asymmetric market behaviors in high and low S-U regimes. Capturing this relationship and understanding the non-linear aspects of the relationship is of interest of grain merchandizers and speculators in the market. A spread trading strategy is simulated for the sample period, January 1985 through April 2015, to determine if the non-linear relationship is a profitable arbitrage opportunity in the market.
Calorimetry And Body Composition Research In Broilers And Broiler Breeders, Justina Victoria Caldas Cueva
Calorimetry And Body Composition Research In Broilers And Broiler Breeders, Justina Victoria Caldas Cueva
Graduate Theses and Dissertations
Indirect calorimetry to study heat production (HP) and dual energy X-ray absorptiometry (DEXA) for body composition (BC) are powerful techniques to study the dynamics of energy and protein utilization in poultry. The first two chapters present the BC (dry matter, lean, protein, and fat, bone mineral, calcium and phosphorus) of modern broilers from 1 – 60 d of age analyzed by chemical analysis and DEXA. DEXA has been validated for precision, standardized for position, and equations and validations developed for chickens under two different feeding levels. These equations are unique to the machine and software in use. Research in broilers …
Probabilistic Graphical Modeling On Big Data, Ming-Hua Chung
Probabilistic Graphical Modeling On Big Data, Ming-Hua Chung
Graduate Theses and Dissertations
The rise of Big Data in recent years brings many challenges to modern statistical analysis and modeling. In toxicogenomics, the advancement of high-throughput screening technologies facilitates the generation of massive amount of biological data, a big data phenomena in biomedical science. Yet, researchers still heavily rely on key word search and/or literature review to navigate the databases and analyses are often done in rather small-scale. As a result, the rich information of a database has not been fully utilized, particularly for the information embedded in the interactive nature between data points that are largely ignored and buried. For the past …
Analytical Comparison Of Contrasting Approaches To Estimating Competing Risks Models, Brian Stephen Rickard
Analytical Comparison Of Contrasting Approaches To Estimating Competing Risks Models, Brian Stephen Rickard
Graduate Theses and Dissertations
Survival analysis is a commonly used tool in many fields but has seen little use in education research despite a common number of research questions for which it is well suited. Researchers often use logistic regression instead; however, this omits useful information. In research on retention and graduation for example, the timing of the event is an important piece of information omitted when using logistic regression. A simulation study was conducted to evaluate four methods of analyzing competing risks survival data, Cox proportional hazards regression, Weibull regression, Fine and Gray's Method, and Cox proportional hazards regression with frailty. College student …
Reliability-Based Design And Acceptance Protocol For Driven Piles, Joseph Jabo
Reliability-Based Design And Acceptance Protocol For Driven Piles, Joseph Jabo
Graduate Theses and Dissertations
The current use of the Arkansas Standard Specifications for Highway Construction Manuals (2003, 2014) for driven pile foundations faces various limitations which result in designs of questionable reliability. These specifications are based on the Allowable Stress Design method (ASD), cover a wide range of uncertainties, do not take into account pile and soil types, and were developed for general use. To overcome these challenges it is deemed necessary to develop a new design and acceptance protocol for driven piles. This new protocol incorporates locally calibrated RLFD resistance factors for accounting for local design and construction experiences and practices, as well …
Online Detection Of Outliers And Structural Breaks Using Sequential Monte Carlo Methods, Richard Wanjohi
Online Detection Of Outliers And Structural Breaks Using Sequential Monte Carlo Methods, Richard Wanjohi
Graduate Theses and Dissertations
Outliers and structural breaks occur quite frequently in time series data. Whereas outliers often contain valuable information
about the process under study, they are known to have serious negative impact on statistical data analysis. Most obvious effect is model misspecification and biased parameter estimation which results in wrong conclusions and inaccurate predictions. Structural time series consist of underlying features such as level, slope, cycles or seasonal components. Structural breaks are permanent disruptions of one or more of these components and might be a signal of serious changes in the observed process.
Detecting outliers and estimating the location of structural breaks …
Poisson Distributed Individuals Control Charts With Optimal Limits, Negin Enayaty Ahangar
Poisson Distributed Individuals Control Charts With Optimal Limits, Negin Enayaty Ahangar
Graduate Theses and Dissertations
The conventional method used in attribute control charts is the Shewhart three sigma limits. The implicit assumption of the Normal distribution in this approach is not appropriate for skewed distributions such as Poisson, Geometric and Negative Binomial. Normal approximations perform poorly in the tail area of the these distributions. In this research, a type of attribute control chart is introduced to monitor the processes that provide count data. The economic objective of this chart is to minimize the cost of its errors which is determined by the designer. This objective is a linear function of type I and II errors. …
An Economic Alternative To The C Chart, Ryan William Black
An Economic Alternative To The C Chart, Ryan William Black
Graduate Theses and Dissertations
Because the probability of Type I error is not evenly distributed beyond upper and lower three-sigma limits the c chart is theoretically inappropriate for a monitor of Poisson distributed phenomena. Furthermore, the normal approximation to the Poisson is of little use when c is small. These practical and theoretical concerns should motivate the computation of true error rates associated with individuals control assuming the Poisson distribution. An economic alternative to the c chart is described as a statistical model of upward shift from c0 to c1 and the two charts are compared in theory. For a range of c chart …
Investigating The Sensitivity Of Goodness-Of-Fit Indices To Detect Measurement Invariance In The Bifactor Model, Jam Khojasteh
Investigating The Sensitivity Of Goodness-Of-Fit Indices To Detect Measurement Invariance In The Bifactor Model, Jam Khojasteh
Graduate Theses and Dissertations
A Monte Carlo simulation study was conducted to evaluate the sensitivities of five commonly used goodness-of-fit indices to detect metric invariance properties of the bifactor model. The fit indices that performed the best in terms of power were Gamma and Mc. In addition, Gamma, Mc, CFI, and RMSEA all held Type I error to a minimum. However, only Gamma and CFI are recommended to use in the bifactor model because the other GOF indices have cutoff values that are too large. For Gamma and CFI values of -.026 to -.045 and -.004 to -.009, respectively indicate a lack of metric …