Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

2015

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 451 - 480 of 490

Full-Text Articles in Statistics and Probability

Combining Semiparametric Regression And Kriging For Prediction Of Pm2.5 Pollutant Levels At Unmonitored Locations With Meterological And Traffic Data, Justin Jonathan Strate Jan 2015

Combining Semiparametric Regression And Kriging For Prediction Of Pm2.5 Pollutant Levels At Unmonitored Locations With Meterological And Traffic Data, Justin Jonathan Strate

Open Access Theses & Dissertations

Particulate matter (PM) is defined by the Texas Commission on Environmental Quality (TCEQ) as "a mixture of solid particles and liquid droplets found in the air". These particles vary widely in size. Those particles that are less than 2.5 micrometers in aerodynamic diameter are known as Particulate Matter 2.5 or PM2.5. These particles are inhaled, and their health effects are still largely being studied. Past studies have assessed PM2.5 exposure of a population, yet individual exposure is more diffcult to assess and may vary widely in a population. Recent studies have combined semiparametric models with kriging (Li et. al [2012]) …


A Bayes Approach In Step-Stress Accelerated Life Testings, Hao Yang Teng Jan 2015

A Bayes Approach In Step-Stress Accelerated Life Testings, Hao Yang Teng

Open Access Theses & Dissertations

A Bayesian analysis for the Weibull proportional hazard (PH) model is presented. A comparison between the Weibull PH model and the Weibull cumulative exposure (CE) model is made graphically and mathematically. The PH model is as flexible as the CE model in fitting step-stress data and the mathematical form of the PH model enables researchers to do Bayesian inferencemuch easier than the CE model. In addition, the PH model has the desirable proportional hazard property. A convex tent prior is used for Bayesian analysis. Markov chain Monte Carlo methods are used for posterior inferences. In this study, we adopt two …


Pre-Tuned Ridge Regression And Its Extension To Generalized Linear Models, Yaa Tawiah Wonkye Jan 2015

Pre-Tuned Ridge Regression And Its Extension To Generalized Linear Models, Yaa Tawiah Wonkye

Open Access Theses & Dissertations

Ridge regression is regularization or shrinkage method and a common approach in dealing with multicollinearity in conventional regression analysis. Ridge regression is widely used by statistical analyst since it is one of the best compared to other regularization methods. Also, the introduction of high dimension and ultra-high dimensional data has become an issue of concern and ridge regression is one way of dealing with such data. One of the key issues associated with ridge regression is the determination of the tuning or ridge parameter. The common practice is to fit ridge regression for a different number of values of tuning …


Towards Analytical Techniques For Optimizing Knowledge Acquisition, Processing, Propagation, And Use In Cyberinfrastructure, Leonardo Octavio Lerma Jan 2015

Towards Analytical Techniques For Optimizing Knowledge Acquisition, Processing, Propagation, And Use In Cyberinfrastructure, Leonardo Octavio Lerma

Open Access Theses & Dissertations

For many decades, there has been a continuous progress in science and engineering applications.

A large part of this progress comes from the new knowledge that researchers acquire, propagate, and use. This new knowledge has revolutionized many aspects of our life, from driving to communications to shopping.

Somewhat surprisingly, there is one area of human activity which is the least impacted by the modern technological progress: the very processes of acquiring, processing, and propagating information. When we decide where to place sensors, which algorithm to use for processing the data – we rely mostly on our own intuition and on …


Evaluation Of The Signature Molecular Descriptor With Blosum62 And An All-Atom Description For Use In Sequence Alignment Of Proteins, Lindsay M. Aichinger Jan 2015

Evaluation Of The Signature Molecular Descriptor With Blosum62 And An All-Atom Description For Use In Sequence Alignment Of Proteins, Lindsay M. Aichinger

Williams Honors College, Honors Research Projects

This Honors Project focused on a few aspects of this topic. The second is comparing the molecular signature kernels to three of the BLOSUM matrices (30, 62, and 90) to test the accuracy of the mathematical model. The kernel matrix was manipulated in order to improve the relationship by focusing on side groups and also by changing how the structure was represented in the matrix by increasing the initial height distance from the central atom (Height 1 and Height 2 included).

There were multiple design constraints for this project. The first was the comparison with the BLOSUM matrices (30, 62, …


Inequalities And Approximations Of Weighted Distributions By Lindley Reliability Measures, And The Lindley-Cox Model With Applications, Broderick O. Oluyede, Macaulay Okwuokenye, Karl E. Peace Jan 2015

Inequalities And Approximations Of Weighted Distributions By Lindley Reliability Measures, And The Lindley-Cox Model With Applications, Broderick O. Oluyede, Macaulay Okwuokenye, Karl E. Peace

Biostatistics: Faculty Publications

In this note, stochastic comparisons and results for weighted and Lindley models are presented. Approximation of weighted distributions via Lindley distribution in the class of increasing failure rate (IFR) and decreasing failure rate (DFR) weighted distributions with monotone weight functions are obtained including approximations via the length-biased Lindley distribution. Some useful bounds and moment-type inequality for weighted life distributions and applications are presented. Incorporation of covariates into Lindley model is considered and an application to illustrate the usefulness and applicability of the proposed Lindley-Cox model is given.


How Long Does That 10-Year Smoke Alarm Really Last? A Survival Analysis Of Smoke Alarms Installed Through The Saife Program In Rural Georgia, Haresh Rochani, Valamar Malika Reagon, Steve Davidson Jan 2015

How Long Does That 10-Year Smoke Alarm Really Last? A Survival Analysis Of Smoke Alarms Installed Through The Saife Program In Rural Georgia, Haresh Rochani, Valamar Malika Reagon, Steve Davidson

Biostatistics: Faculty Publications

Background: When functioning properly, a smoke alarm alerts individuals in the residence that smoke is near the alarm. Smoke alarms serve as a primary prevention mechanism to abate morbidity and mortality related to residential fires.

Methods: Using survival analysis, we examined the length of operability of 10-year lithium battery powered smoke alarms installed through the Georgia Public Health/CDC SAIFE program in Moultrie, Georgia. Attempts were made to reach all homes in the city limits. The premise of the study is that geographic clusters (in the case of Moultrie city quadrants) are associated with decreases in the length of time that …


A Meta-Analysis Of Association Between One-Carbon Metabolism Gene Polymorphisms And Risk Of Prostate Cancer, Mahmood Tazari Jan 2015

A Meta-Analysis Of Association Between One-Carbon Metabolism Gene Polymorphisms And Risk Of Prostate Cancer, Mahmood Tazari

Walden Dissertations and Doctoral Studies

Prostate cancer is the most common cancer among men. The purpose of this quantitative, meta-analysis study was to examine one-carbon metabolism gene polymorphisms in a group of genes to determine their association with prostate cancer risk. The genetic epidemiology theory provided the framework for the study. The data collected were from published articles. From over 2,800 individual studies, 20 articles were retained for results and data abstraction, following the title, abstract screen, and full text screening in the second phase. The data were analyzed by a meta-analysis statistical method, combining the results from selected studies to estimate the overall association. …


Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu Jan 2015

Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu

Theses and Dissertations

Analysis of variance (ANOVA) is a robust test against the normality assumption, but it may be inappropriate when the assumption of homogeneity of variance has been violated. Welch ANOVA and the Kruskal-Wallis test (a non-parametric method) can be applicable for this case. In this study we compare the three methods in empirical type I error rate and power, when heterogeneity of variance occurs and find out which method is the most suitable with which cases including balanced/unbalanced, small/large sample size, and/or with normal/non-normal distributions.


The Subject Librarian Newsletter, Statistics, Fall 2015, Patti Mccall Jan 2015

The Subject Librarian Newsletter, Statistics, Fall 2015, Patti Mccall

Libraries' Newsletters

No abstract provided.


Ranking Interesting Changes In Correlation Coefficient Matrix Results From Varying Data Partitions In Causal Graphic Modeling, Yesica Daniela Bravo Gonzalez Jan 2015

Ranking Interesting Changes In Correlation Coefficient Matrix Results From Varying Data Partitions In Causal Graphic Modeling, Yesica Daniela Bravo Gonzalez

Master's Theses

Problem

In life we need to compare situations in order to select the best solution. The study in this paper is about analyzing data (variables), which is also called data mining. There are situations where it is not enough to compare variables among themselves at one specific moment. Sometimes it is necessary to compare the behavior of variables at different periods of time and know how they behave at different times in order to select the best arrangements for any situation.

Method

To find correlation among variables, traffic intersections were simulated so they could be compared, since the correlation coefficient …


Bayesian Function-On-Function Regression For Multilevel Functional Data, Mark J. Meyer, Brent A. Coull, Francesco Versace, Paul Cinciripini, Jeffrey S. Morris Jan 2015

Bayesian Function-On-Function Regression For Multilevel Functional Data, Mark J. Meyer, Brent A. Coull, Francesco Versace, Paul Cinciripini, Jeffrey S. Morris

Faculty Journal Articles

Medical and public health research increasingly involves the collection of complex and high dimensional data. In particular, functional data—where the unit of observation is a curve or set of curves that are finely sampled over a grid—is frequently obtained. Moreover, researchers often sample multiple curves per person resulting in repeated functional measures. A common question is how to analyze the relationship between two functional variables. We propose a general function-on-function regression model for repeatedly sampled functional data on a fine grid, presenting a simple model as well as a more extensive mixed model framework, and introducing various functional Bayesian inferential …


Understanding Vulnerability In Alaska Fishing Communities: A Validation Methodology For Rapid Assessment Of Well-Being Indices, Conor M. Maguire Jan 2015

Understanding Vulnerability In Alaska Fishing Communities: A Validation Methodology For Rapid Assessment Of Well-Being Indices, Conor M. Maguire

All Master's Theses

Social well-being indices measure how fishing communities are likely to be affected by social-ecological perturbations, and are a significant tool to identify the primary issues influencing communities’ sustained participation in fishing activities. In an attempt to further our understanding of how communities are affected by such perturbations, we have developed a rapid assessment methodology to test the external validity of a set of well-being indices that measure community vulnerability. This methodology informs how well such indices reflect the communities they represent by measuring elements of well-being through field observations, and comparing them to corresponding index components created from secondary data …


Applications Of Monte Carlo Methods In Statistical Inference Using Regression Analysis, Ji Young Huh Jan 2015

Applications Of Monte Carlo Methods In Statistical Inference Using Regression Analysis, Ji Young Huh

CMC Senior Theses

This paper studies the use of Monte Carlo simulation techniques in the field of econometrics, specifically statistical inference. First, I examine several estimators by deriving properties explicitly and generate their distributions through simulations. Here, simulations are used to illustrate and support the analytical results. Then, I look at test statistics where derivations are costly because of the sensitivity of their critical values to the data generating processes. Simulations here establish significance and necessity for drawing statistical inference. Overall, the paper examines when and how simulations are needed in studying econometric theories.


Acceptance-Rejection Sampling With Hierarchical Models, Christian A. Ayala Jan 2015

Acceptance-Rejection Sampling With Hierarchical Models, Christian A. Ayala

CMC Senior Theses

Hierarchical models provide a flexible way of modeling complex behavior. However, the complicated interdependencies among the parameters in the hierarchy make training such models difficult. MCMC methods have been widely used for this purpose, but can often only approximate the necessary distributions. Acceptance-rejection sampling allows for perfect simulation from these often unnormalized distributions by drawing from another distribution over the same support. The efficacy of acceptance-rejection sampling is explored through application to a small dataset which has been widely used for evaluating different methods for inference on hierarchical models. A particular algorithm is developed to draw variates from the posterior …


Statistical Modeling Of Microrna Expression With Human Cancers, Ke-Sheng Wang, Yue Pan, Chun Xu Jan 2015

Statistical Modeling Of Microrna Expression With Human Cancers, Ke-Sheng Wang, Yue Pan, Chun Xu

Health & Biomedical Sciences Faculty Publications

MicroRNAs (miRNAs) are small non-coding RNAs (containing about 22 nucleotides) that regulate gene expression. MiRNAs are involved in many different biological processes such as cell proliferation, differentiation, apoptosis, fat metabolism, and human cancer genes; while miRNAs may function as candidates for diagnostic and prognostic biomarkers and predictors of drug response. This paper emphasizes the statistical methods in the analysis of the associations of miRNA gene expression with human cancers and related clinical phenotypes: 1) simple statistical methods include chi-square test, correlation analysis, t-test and one-way ANOVA; 2) regression models include linear and logistic regression; 3) survival analysis approaches such as …


Gene Expression Changes Reflect Clinical Response In A Placebo-Controlled Randomized Trial Of Abatacept In Patients With Diffuse Cutaneous Systemic Sclerosis, Eliza F. Chakravarty, Viktor Martyanov, David Fiorentino, Tammara A. Wood, David J. Haddon, Justin A. Jarrell, Paul Utz, Mark Genovese, Michael Whitfield, Lorinda Chung Jan 2015

Gene Expression Changes Reflect Clinical Response In A Placebo-Controlled Randomized Trial Of Abatacept In Patients With Diffuse Cutaneous Systemic Sclerosis, Eliza F. Chakravarty, Viktor Martyanov, David Fiorentino, Tammara A. Wood, David J. Haddon, Justin A. Jarrell, Paul Utz, Mark Genovese, Michael Whitfield, Lorinda Chung

Dartmouth Scholarship

Systemic sclerosis is an autoimmune disease characterized by inflammation and fibrosis of the skin and internal organs. We sought to assess the clinical and molecular effects associated with response to intravenous abatacept in patients with diffuse cutaneous systemic.


A Model For Determining Drivers Of Phenology In Western United States Rangelands, Joseph R. St. Peter Jan 2015

A Model For Determining Drivers Of Phenology In Western United States Rangelands, Joseph R. St. Peter

Graduate Student Theses, Dissertations, & Professional Papers

Plant phenology has long been used as an indicator of climate. Recent changes in plant phenology are evidence of the influence of climate change. Modeling plant phenology has become an effective tool to understand the impacts of climate change. Using machine learning techniques I developed a modeling process for accurately predicting phenology across a diverse landscape. This model uses individual site data to set site specific climate thresholds for plant phenology. This model also identifies the limiting factors to vegetation phenology for rangelands in the western United States. NDVI remotely sensed data was used to quantify land surface phenology and …


The Sensitivity Of A Test Based On Spearman's Rho In Cross-Correlation Change Point Problems, Congjian Liu Jan 2015

The Sensitivity Of A Test Based On Spearman's Rho In Cross-Correlation Change Point Problems, Congjian Liu

College of Graduate Studies: Theses & Dissertations

In change point problems, there are three main questions that researchers are interested in. First of all, is there a change point or not? Second, when does the change point occur in a time series? Third, how quickly can we detect the change point? In this thesis, we first explain what a change point is, and what a cross-correlation is. We then discuss prior research in this area. Then we discuss and examine a test based on Spearman's rho, introduced by Wied and Dehling (2011), which tests the null hypothesis of no change point, and compare the change point we …


Bayesian Inference Of The Weibull-Pareto Distribution, James Dow Jan 2015

Bayesian Inference Of The Weibull-Pareto Distribution, James Dow

College of Graduate Studies: Theses & Dissertations

The Weibull distribution has many applications in various topics. Some of these topics include survival analysis, reliability engineering, general insurance, electrical engineering, and industrial engineering. The Weibull distribution was further extended by the Weibull-Pareto distribution. A desirable property this distribution has is its shape can skew being able to better model left or right skewed data. Examples of skewed data include human longevity and actuarial data. In this work a hierarchical Bayesian model was developed using the Weibull-Pareto distribution.


Computational Intelligence Based Complex Adaptive System-Of-Systems Architecture Evolution Strategy, Siddharth Agarwal Jan 2015

Computational Intelligence Based Complex Adaptive System-Of-Systems Architecture Evolution Strategy, Siddharth Agarwal

Doctoral Dissertations

The dynamic planning for a system-of-systems (SoS) is a challenging endeavor. Large scale organizations and operations constantly face challenges to incorporate new systems and upgrade existing systems over a period of time under threats, constrained budget and uncertainty. It is therefore necessary for the program managers to be able to look at the future scenarios and critically assess the impact of technology and stakeholder changes. Managers and engineers are always looking for options that signify affordable acquisition selections and lessen the cycle time for early acquisition and new technology addition. This research helps in analyzing sequential decisions in an evolving …


Investigation Of Robust Optimization And Evidence Theory With Stochastic Expansions For Aerospace Applications Under Mixed Uncertainty, Harsheel R. Shah Jan 2015

Investigation Of Robust Optimization And Evidence Theory With Stochastic Expansions For Aerospace Applications Under Mixed Uncertainty, Harsheel R. Shah

Doctoral Dissertations

One of the primary objectives of this research is to develop a method to model and propagate mixed (aleatory and epistemic) uncertainty in aerospace simulations using DSTE. In order to avoid excessive computational cost associated with large scale applications and the evaluation of Dempster Shafer structures, stochastic expansions are implemented for efficient UQ. The mixed UQ with DSTE approach was demonstrated on an analytical example and high fidelity computational fluid dynamics (CFD) study of transonic flow over a RAE 2822 airfoil.

Another objective is to devise a DSTE based performance assessment framework through the use of quantification of margins and …


Essays On Unit Root Testing In Time Series, Xiao Zhong Jan 2015

Essays On Unit Root Testing In Time Series, Xiao Zhong

Doctoral Dissertations

"Unit root tests are frequently employed by applied time series analysts to determine if the underlying model that generates an empirical process has a component that can be well-described by a random walk. More specifically, when the time series can be modeled using an autoregressive moving average (ARMA) process, such tests aim to determine if the autoregressive (AR) polynomial has one or more unit roots. The effect of economic shocks do not diminish with time when there is one or more unit roots in the AR polynomial, whereas the contribution of shocks decay geometrically when all the roots are outside …


Small Sample Saddlepoint Confidence Intervals In Epidemiology, Pasan Manuranga Edirisinghe Jan 2015

Small Sample Saddlepoint Confidence Intervals In Epidemiology, Pasan Manuranga Edirisinghe

Doctoral Dissertations

"In section 1, we develop a novel method of confidence interval construction for directly standardized rates. These intervals involve saddlepoint approximations to the intractable distribution of a weighted sum of Poisson random variables and the determination of hypothetical Poisson mean values for each of the age groups. Simulation studies show that, in terms of coverage probability and length, the saddlepoint confidence interval (SP) outperforms four competing confidence intervals obtained from the moment matching (M8), gamma-based (G1,G4) and ABC bootstrap (ABC) methods.

In section 2, we first consider Brillinger's classical model for a vital rate estimate with a random denominator. We …


Small Sample Umpu Equivalence Testing Based On Saddlepoint Approximations, Renren Zhao Jan 2015

Small Sample Umpu Equivalence Testing Based On Saddlepoint Approximations, Renren Zhao

Doctoral Dissertations

"In the first section, we consider small sample equivalence tests for exponentiality. Statistical inference in this setting is particularly challenging since equivalence testing procedures typically require a much larger sample size, in comparison to classical "difference tests", to perform well. We make use of Butler's marginal likelihood for the shape parameter of a gamma distribution in our development of equivalence tests for exponentiality. We consider two procedures using the principle of confidence interval inclusion, four Bayesian methods, and the uniformly most powerful unbiased (UMPU) test where a saddlepoint approximation to the intractable distribution of a canonical sufficient statistic is used. …


Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou Jan 2015

Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou

Theses and Dissertations

Flexible incorporation of both geographical patterning and risk effects in cancer survival models is becoming increasingly important, due in part to the recent availability of large cancer registries. The analysis of spatial survival data is challenged by the presence of spatial dependence and censoring for survival times. Accurately modeling the risk factors and geographical pattern that explain the differences in survival is particularly of interest. Within this dissertation, the first chapter reviews commonlyused baseline priors, semiparametric and nonparametric Bayesian survival models and recent approaches for accommodating spatial dependence, both conditional and marginal. The last three chapters contribute three flexible survival …


Individual-Based Modeling: Mountain Pine Beetle Seasonal Biology In Response To Climate, Jacques Regniere, Barbara J. Bentz, James A. Powell, Remi St-Amant Jan 2015

Individual-Based Modeling: Mountain Pine Beetle Seasonal Biology In Response To Climate, Jacques Regniere, Barbara J. Bentz, James A. Powell, Remi St-Amant

Mathematics and Statistics Faculty Publications

Over the past decades, as significant advances were made in the availability and accessibility of computing power, individual-based models (IBM) have become increasingly appealing to ecologists (Grimm 1999). The individual-based modeling approachprovides a convenient framework to incorporate detailed knowledge of individuals and of their interactions within populations (Lomnicki 1999). Variability among individuals is essential to the success of populations that are exposed to changing environments, and because natural selection acts on this variability, it is an essential component of population performance. © Springer International Publishing Switzerland 2015.


Graph-Based Regularization In Machine Learning: Discovering Driver Modules In Biological Networks, Xi Gao Jan 2015

Graph-Based Regularization In Machine Learning: Discovering Driver Modules In Biological Networks, Xi Gao

Theses and Dissertations

Curiosity of human nature drives us to explore the origins of what makes each of us different. From ancient legends and mythology, Mendel's law, Punnett square to modern genetic research, we carry on this old but eternal question. Thanks to technological revolution, today's scientists try to answer this question using easily measurable gene expression and other profiling data. However, the exploration can easily get lost in the data of growing volume, dimension, noise and complexity. This dissertation is aimed at developing new machine learning methods that take data from different classes as input, augment them with knowledge of feature relationships, …


Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe Jan 2015

Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe

Theses and Dissertations

Combining effect sizes from individual studies using random-effects models are commonly applied in high-dimensional gene expression data. However, unknown study heterogeneity can arise from inconsistency of sample qualities and experimental conditions. High heterogeneity of effect sizes can reduce statistical power of the models. We proposed two new methods for random effects estimation and measurements for model variation and strength of the study heterogeneity. We then developed a statistical technique to test for significance of random effects and identify heterogeneous genes. We also proposed another meta-analytic approach that incorporates informative weights in the random effects meta-analysis models. We compared the proposed …


Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima Jan 2015

Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima

Theses and Dissertations

In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not done randomly in observational studies, comparisons of outcomes between exposed and non-exposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of odds ratio and hazard ratio. However, there is a lack of research into the performance of propensity score methods for estimating the …