Open Access. Powered by Scholars. Published by Universities.®

Statistics and Probability Commons™

Open Access. Powered by Scholars. Published by Universities.®

Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 301 - 330 of 565

Full-Text Articles in Statistics and Probability

High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski Jan 2015

High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski

Theses and Dissertations

The advent of high-throughput sequencing has brought about the creation of an unprecedented amount of research data. Analytical methodology has not been able to keep pace with the plethora of data being produced. Two assays, ImmunoSEQ and the cytokinesisblock micronucleus (CBMN), that both produce count data and have few methods available to analyze them are considered.

ImmunoSEQ is a sequencing assay that measures the beta T-cell receptor (TCR) repertoire. The ImmunoSEQ assay was used to describe the TCR repertoires of patients that have undergone hematopoietic stem cell transplantation (HSCT). Several different methods for spectratype analysis were extended to the TCR …


Considerations For Screening Designs And Follow-Up Experimentation, Robert D. Leonard Jan 2015

Considerations For Screening Designs And Follow-Up Experimentation, Robert D. Leonard

Theses and Dissertations

The success of screening experiments hinges on the effect sparsity assumption, which states that only a few of the factorial effects of interest actually have an impact on the system being investigated. The development of a screening methodology to harness this assumption requires careful consideration of the strengths and weaknesses of a proposed experimental design in addition to the ability of an analysis procedure to properly detect the major influences on the response. However, for the most part, screening designs and their complementing analysis procedures have been proposed separately in the literature without clear consideration of their ability to perform …


Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu Jan 2015

Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu

Theses and Dissertations

Analysis of variance (ANOVA) is a robust test against the normality assumption, but it may be inappropriate when the assumption of homogeneity of variance has been violated. Welch ANOVA and the Kruskal-Wallis test (a non-parametric method) can be applicable for this case. In this study we compare the three methods in empirical type I error rate and power, when heterogeneity of variance occurs and find out which method is the most suitable with which cases including balanced/unbalanced, small/large sample size, and/or with normal/non-normal distributions.


Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou Jan 2015

Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou

Theses and Dissertations

Flexible incorporation of both geographical patterning and risk effects in cancer survival models is becoming increasingly important, due in part to the recent availability of large cancer registries. The analysis of spatial survival data is challenged by the presence of spatial dependence and censoring for survival times. Accurately modeling the risk factors and geographical pattern that explain the differences in survival is particularly of interest. Within this dissertation, the first chapter reviews commonlyused baseline priors, semiparametric and nonparametric Bayesian survival models and recent approaches for accommodating spatial dependence, both conditional and marginal. The last three chapters contribute three flexible survival …


Graph-Based Regularization In Machine Learning: Discovering Driver Modules In Biological Networks, Xi Gao Jan 2015

Graph-Based Regularization In Machine Learning: Discovering Driver Modules In Biological Networks, Xi Gao

Theses and Dissertations

Curiosity of human nature drives us to explore the origins of what makes each of us different. From ancient legends and mythology, Mendel's law, Punnett square to modern genetic research, we carry on this old but eternal question. Thanks to technological revolution, today's scientists try to answer this question using easily measurable gene expression and other profiling data. However, the exploration can easily get lost in the data of growing volume, dimension, noise and complexity. This dissertation is aimed at developing new machine learning methods that take data from different classes as input, augment them with knowledge of feature relationships, …


Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe Jan 2015

Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe

Theses and Dissertations

Combining effect sizes from individual studies using random-effects models are commonly applied in high-dimensional gene expression data. However, unknown study heterogeneity can arise from inconsistency of sample qualities and experimental conditions. High heterogeneity of effect sizes can reduce statistical power of the models. We proposed two new methods for random effects estimation and measurements for model variation and strength of the study heterogeneity. We then developed a statistical technique to test for significance of random effects and identify heterogeneous genes. We also proposed another meta-analytic approach that incorporates informative weights in the random effects meta-analysis models. We compared the proposed …


Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima Jan 2015

Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima

Theses and Dissertations

In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not done randomly in observational studies, comparisons of outcomes between exposed and non-exposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of odds ratio and hazard ratio. However, there is a lack of research into the performance of propensity score methods for estimating the …


Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray Dec 2014

Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray

Theses and Dissertations

Functional neuroimaging is a relatively young discipline within the neurosciences that has led to significant advances in our understanding of the human brain and progress in neuroscientific research related to public health. Accurately identifying activated regions in the brain showing a strong association with an outcome of interest is crucial in terms of disease prediction and prevention. Functional magnetic resonance imaging (fMRI) is the most widely used method for this type of study as it has the ability to measure and identify the location of changes in tissue perfusion, blood oxygenation, and blood volume. In practice, the three-dimensional brain locations …


Semiparametric Regression Analysis Of Bivariate Interval-Censored Data, Naichen Wang Dec 2014

Semiparametric Regression Analysis Of Bivariate Interval-Censored Data, Naichen Wang

Theses and Dissertations

Survival analysis is a long-lasting and popular research area and has numerous applications in all fields such as social science, engineering, economics, industry, and public health. Interval-censored data are a special type of survival data, in which the survival time of interest is never exactly observed but is known to fall within some observed interval. Interval-censored data arise commonly in real-life studies, in which subjects are examined at periodical or irregular follow-up visits. In this dissertation, we develop efficient statistical approaches for regression analysis of bivariate intervalcensored data, in which the two survival times of interest are correlated and both …


Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta Dec 2014

Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta

Theses and Dissertations

The effects of scale on the analysis of spatial data, often referred to as the modifiable areal unit problem in spatial studies, is one of the issues often encountered in small area health models. These spatial effects of scale are also seen in the areas of disease mapping where data are usually available in counts. Often there is a need to consider the different scales of aggregation that exist within count data, since inferences based on analyses can vary if we change the definition of the unit of analysis. This thesis provides a framework that describes the distribution of relative …


Non- And Semi-Parametric Bayesian Inference With Recurrent Events And Coherent Systems Data, A. K. M. Fazlur Rahman Aug 2014

Non- And Semi-Parametric Bayesian Inference With Recurrent Events And Coherent Systems Data, A. K. M. Fazlur Rahman

Theses and Dissertations

This dissertation deals with non- and semi-parametric Bayesian inference of gap-time distribution with recurrent event data and simultaneous inference of component and system reliabilities of coherent systems data. Recurrent event data arise from a wide variety of studies/fields such as clinical trials, epidemiology, public health, biomedicine (e.g. repeated heart attack, repeated tumor occurrences of a cancer patient). In Chapter 2 we develop nonparametric Bayes and empirical Bayes estimators of the survivor function \bar{F} = 1 - F, of the gap-time distribution by assigning a Dirichlet process prior on F. We develop a closed form estimator of \bar{F} as well as …


Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello Apr 2014

Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello

Theses and Dissertations

In clinical settings, the diagnosis of medical conditions is often aided by measurement of various serum biomarkers through the use of laboratory tests. These biomarkers provide information about different aspects of a patient’s health and the overall function of different organs. In this dissertation, we develop and validate a weighted composite index that aggregates the information from a variety of health biomarkers covering multiple organ systems. The index can be used for predicting all-cause mortality and could also be used as a holistic measure of overall physiological health status. We refer to it as the Health Status Metric (HSM). Validation …


A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng Apr 2014

A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng

Theses and Dissertations

BACKGROUND: Little is known of the effects of obesity, body size and body composition, and blood pressure (BP) in childhood on hypertension (HBP) and cardiac structure and function in adulthood due to the lack of long-term serial data on these parameters from childhood into adulthood. In the present study, we are poised to analyze these serial data from the Fels Longitudinal Study (FLS) to evaluate the extent to which body size during childhood determines HBP and cardiac structure and function in the same individuals in adulthood through mathematical modeling. METHODS: The data were from 412 males and 403 females in …


Dynamic Bayesian Approaches To The Statistical Calibration Problem, Derick Lorenzo Rivers Jan 2014

Dynamic Bayesian Approaches To The Statistical Calibration Problem, Derick Lorenzo Rivers

Theses and Dissertations

The problem of statistical calibration of a measuring instrument can be framed both in a statistical context as well as in an engineering context. In the first, the problem is dealt with by distinguishing between the "classical" approach and the "inverse" regression approach. Both of these models are static models and are used to estimate "exact" measurements from measurements that are affected by error. In the engineering context, the variables of interest are considered to be taken at the time at which you observe the measurement. The Bayesian time series analysis method of Dynamic Linear Models (DLM) can be used …


Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes Jan 2014

Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes

Theses and Dissertations

In the study of associated discrete variables, limitations on the range of the possible association measures (Pearson correlation, odds ratio, etc.) arise from the form of the joint probability function between the variables. These limitations are known as the Fréchet bounds. The bounds for cases involving associated binary variables are explored in the context of simulating datasets with a desired correlation and set of marginal probabilities. A new method for creating such datasets is compared to an existing method that uses the multivariate probit. A method for simulating associated binary variables using a desired odds ratio and known marginal probabilities …


Oldtimers & Newcomers In Collective Action, Stefanie R. Chamberlain Jan 2014

Oldtimers & Newcomers In Collective Action, Stefanie R. Chamberlain

Theses and Dissertations

Most work on groups facing collective action assumes that group membership is static, or fixed. Yet static membership is rare, with members joining and leaving groups. In this thesis, I propose to explore how the presence of newcomers to groups affects group coordination. Past research has shown an overall negative effect of newcomers on group contributions. The proposed thesis attempts to further establish the effect by determining whether newcomers, oldtimers, or both are responsible for the declining cooperation in groups. While the empirical component is focused solely on establishing who is responsible for driving down cooperation rates in dynamic groups, …


Methods For Integrative Analysis Of Genomic Data, Paul Manser Jan 2014

Methods For Integrative Analysis Of Genomic Data, Paul Manser

Theses and Dissertations

In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …


Applications Of Bayesian Nonparametrics To Reliability And Survival Data, Li Li Jan 2014

Applications Of Bayesian Nonparametrics To Reliability And Survival Data, Li Li

Theses and Dissertations

Reliability and survival data are widely encountered across many common settings. Subjects under investigation often include machines, bioassays, patients, etc.; their reliability or survival distribution, and its association with covariate processes, are commonly of interest. Within this dissertation, the first two chapters focus on reliability data where repairable systems fail and get interventions, e.g. repairs in the event process. It begins with a nonparametric test for the commonly assumed ''good as old'' assumption for minimal repair models and then a semi-parametric regression model is introduced for reliability data using Kijima's effective age. The third chapter focuses on survival data observed …


Methods For Clustering Mixed Data, Jeanmarie L. Hendrickson Jan 2014

Methods For Clustering Mixed Data, Jeanmarie L. Hendrickson

Theses and Dissertations

We give a brief introduction to cluster analysis and then propose and discuss a few methods for clustering mixed data. In particular, a model-based clustering method for mixed data based on Everitt's (1988) work is described, and we use a simulated annealing method to estimate the parameters for Everitt's model. A penalized log likelihood with the simulated annealing method is proposed as a remedy for the parameter estimates being drawn to extremes. Everitt's approach and the proposed method are compared based on their performance in clustering simulated data. We then use the penalized log likelihood method on a heart disease …


Bayesian Analysis Of Continuous Curve Functions, Wen Cheng Jan 2014

Bayesian Analysis Of Continuous Curve Functions, Wen Cheng

Theses and Dissertations

We consider Bayesian analysis of continuous curve functions in 1D, 2D and 3D spaces. A fundamental feature of the analysis is that it is invariant under a simultaneous warping/re-parameterization of all target curves, as well as translation, rotation and scale of each individual if necessary. We introduce Bayesian models based on a special curve representation named Square Root Velocity Function (SRVF) introduced by Srivastava et al. (2011, IEEE PAMI). A Gaussian process model for the SRVFs of curves is proposed, and suitable prior models such as the Dirichlet distribution are employed for modeling the warping function as a cumulative distribution …


Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren Jan 2014

Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren

Theses and Dissertations

Latent variable models (LVMs) are commonly used in the scenario where the outcome of the main interest is an unobservable measure, associated with multiple observed surrogate outcomes, and affected by potential risk factors. This thesis develops an approach of efficient handling missing surrogate outcomes and covariates in two- and three-level latent variable models. However, corresponding statistical methodologies and computational software are lacking efficiently analyzing the LVMs given surrogate outcomes and covariates subject to missingness in the LVMs. We analyze the two-level LVMs for longitudinal data from the National Growth of Health Study where surrogate outcomes and covariates are subject to …


The Total Picture: Multiple Chemical Exposures To Pregnant Women In The Us – An Nhanes Study Of Data From 2003 Through 2010, Teri Cabana Jan 2014

The Total Picture: Multiple Chemical Exposures To Pregnant Women In The Us – An Nhanes Study Of Data From 2003 Through 2010, Teri Cabana

Theses and Dissertations

INTRODUCTION: Chemical exposures to US pregnant women have been shown to have adverse health impacts on both mother and fetus. A prior paper revealed that US pregnant women in 2003-2004 had widespread exposure to multiple chemicals. The goal of this research is to examine how environmental chemical exposures to US pregnant women have changed from 2003 to 2010 and to look further at the extent of simultaneous exposure to multiple chemicals in US pregnant women using biomonitoring data available through NHANES (the National Health and Nutritional Examination Survey). METHODS: Using available NHANES data from the following cycles (2003-2004, 2005-2006, 2007-2008, …


Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou Nov 2013

Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou

Theses and Dissertations

Ordinal scales are commonly used to measure health status and disease related outcomes in hospital settings as well as in translational medical research. Notable examples include cancer staging, which is a five-category ordinal scale indicating tumor size, node involvement, and likelihood of metastasizing. Glasgow Coma Scale (GCS), which gives a reliable and objective assessment of conscious status of a patient, is an ordinal scaled measure. In addition, repeated measurements are common in clinical practice for tracking and monitoring the progression of complex diseases. Classical ordinal modeling methods based on the likelihood approach have contributed to the analysis of data in …


Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri Nov 2013

Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri

Theses and Dissertations

O'Brien and Fleming (1979) proposed a straightforward and useful multiple testing procedure (group sequential testing procedure) for comparing two treatments in clinical trials where subject responses are dichotomous (e.g. success and failure). O'Brien and Fleming stated that their group sequential testing procedure has the same Type I error rate and power as that of a fixed one-stage chi-square test, but gives the opportunity to terminate the trial early when one treatment is clearly performing better than the other. We studied and tested the O'Brien and Fleming procedure specifically by correcting the originally proposed critical values. Furthermore, we updated the O’Brien …


Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks Nov 2013

Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks

Theses and Dissertations

Response adaptive designs intend to allocate more patients to better treatments without undermining the validity and the integrity of the trial. The immediacy of the primary response (e.g. deaths, remission) determines the efficiency of the response adaptive design, which often requires outcomes to be quickly or immediately observed. This presents difficulties for survival studies, which may require long durations to observe the primary endpoint. Therefore, we introduce auxiliary endpoints to assist the adaptation with the primary endpoint, where an auxiliary endpoint is generally defined as any measurement that is positively associated with the primary endpoint. Our proposed design (referred to …


Construction, Analysis, And Data-Driven Augmentation Of Supersaturated Designs, Alex J. Gutman Sep 2013

Construction, Analysis, And Data-Driven Augmentation Of Supersaturated Designs, Alex J. Gutman

Theses and Dissertations

Screening designs are used in the early stages of industrial and computer experiments to find the most important input factors affecting a system's output. They provide an economical way to remove unimportant factors from further, potentially costly, experimentation. However, when an experiment has a large number of control factors and limited number of available runs, it is infeasible to run a traditional screening design. In these situations, experimenters can use supersaturated designs. A supersaturated design is a fractional factorial design that can screen a set of k factors in n runs, where k is greater than n -1. Unfortunately, they …


Bayesian And Classical Regression Models Inference When The Errors Follow Skewed Distributions, Ebtisam Karim Abdulah Aug 2013

Bayesian And Classical Regression Models Inference When The Errors Follow Skewed Distributions, Ebtisam Karim Abdulah

Theses and Dissertations

The Gamma distribution is one of the most popular distributions for reliability and lifetime data and can be used effectively in analyzing positive skewed data. In real life applications, analysts might wish to have distributions for analyzing positive and negative skewed data, therefore, in this century, we have seen a good attention for fitting data using skew distributions, because it allows continuous variations from symmetric to non-symmetric. Some empirical data, especially, finance (prices and returns) and environmental data have peak distributions and involve tail behavior which affects the model assumptions. In this dissertation, we extend two of the symmetric distributions, …


The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk Jul 2013

The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk

Theses and Dissertations

Many continuous medical tests often rely on a threshold for diagnosis. There are two sequential testing strategies of interest: Believe the Positive (BP) and Believe the Negative (BN). BP classifies a patient positive if either the first test is greater than a threshold θ1 or negative on the first test and greater than θ2 on the second test. BN classifies a patient positive if the first test is greater than a threshold θ3 and greater than θ4 on the second test. Threshold pairs θ = (θ1, θ2) or (θ3, θ4), depending on strategy, are defined as optimal if they maximized …


Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon Apr 2013

Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon

Theses and Dissertations

Survival Analysis generally uses the median survival time as a common summary statistic. While the median possesses the desirable characteristic of being unbiased, there are times when it is not the best statistic to describe the data at hand. Royston and Parmar (2011) provide an argument that the restricted mean survival time should be the summary statistic used when the proportional hazards assumption is in doubt. Work in Restricted Means dates back to 1949 when J.O. Irwin developed a calculation for the standard error of the restricted mean using Greenwood’s formula. Since then the development of the restricted mean has …


Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese Apr 2013

Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese

Theses and Dissertations

Batch effects are due to probe-specific systematic variation between groups of samples (batches) resulting from experimental features that are not of biological interest. Principal components analysis (PCA) is commonly used as a visual tool to determine whether batch effects exist after applying a global normalization method. However, PCA yields linear combinations of the variables that contribute maximum variance and thus will not necessarily detect batch effects if they are not the largest source of variability in the data. We present an extension of principal components analysis to quantify the existence of batch effects, called guided PCA (gPCA). We describe a …