Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Biostatistics (162)
- Applied Statistics (91)
- Engineering (86)
- Statistical Models (59)
- Operations Research, Systems Engineering and Industrial Engineering (47)
-
- Design of Experiments and Sample Surveys (41)
- Operational Research (35)
- Statistical Methodology (34)
- Social and Behavioral Sciences (32)
- Life Sciences (29)
- Applied Mathematics (22)
- Medicine and Health Sciences (22)
- Business (20)
- Longitudinal Data Analysis and Time Series (19)
- Data Science (17)
- Multivariate Analysis (17)
- Computer Sciences (16)
- Mathematics (14)
- Survival Analysis (14)
- Environmental Sciences (13)
- Electrical and Computer Engineering (12)
- Aviation (11)
- Psychology (11)
- Oceanography and Atmospheric Sciences and Meteorology (10)
- Arts and Humanities (9)
- Bioinformatics (9)
- Categorical Data Analysis (9)
- Genetics and Genomics (9)
- Institution
- Keyword
-
- Bayesian (21)
- Statistics (20)
- Physical Sciences and Mathematics, Statistics and Probability (12)
- Machine learning (11)
- Simulation (11)
-
- Biostatistics (9)
- Classification (9)
- Survival analysis (8)
- Design of experiments (7)
- Semiparametric (7)
- Bayesian Hierarchical Models (6)
- Bayesian methods (6)
- Goodness-of-fit tests (6)
- Longitudinal data (6)
- Markov chain Monte Carlo (6)
- Maximum likelihood (6)
- Physical Sciences and Mathematics, Statistics (6)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Variable selection (6)
- #antcenter (5)
- Clinical trial (5)
- Clustering (5)
- EM algorithm (5)
- Experimental design (5)
- Gene expression (5)
- Gibbs sampling (5)
- MCMC (5)
- Neural networks (5)
- Optimization (5)
- Parameter estimation (5)
- Publication Year
Articles 301 - 330 of 565
Full-Text Articles in Statistics and Probability
High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski
High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski
Theses and Dissertations
The advent of high-throughput sequencing has brought about the creation of an unprecedented amount of research data. Analytical methodology has not been able to keep pace with the plethora of data being produced. Two assays, ImmunoSEQ and the cytokinesisblock micronucleus (CBMN), that both produce count data and have few methods available to analyze them are considered.
ImmunoSEQ is a sequencing assay that measures the beta T-cell receptor (TCR) repertoire. The ImmunoSEQ assay was used to describe the TCR repertoires of patients that have undergone hematopoietic stem cell transplantation (HSCT). Several different methods for spectratype analysis were extended to the TCR …
Considerations For Screening Designs And Follow-Up Experimentation, Robert D. Leonard
Considerations For Screening Designs And Follow-Up Experimentation, Robert D. Leonard
Theses and Dissertations
The success of screening experiments hinges on the effect sparsity assumption, which states that only a few of the factorial effects of interest actually have an impact on the system being investigated. The development of a screening methodology to harness this assumption requires careful consideration of the strengths and weaknesses of a proposed experimental design in addition to the ability of an analysis procedure to properly detect the major influences on the response. However, for the most part, screening designs and their complementing analysis procedures have been proposed separately in the literature without clear consideration of their ability to perform …
Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu
Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu
Theses and Dissertations
Analysis of variance (ANOVA) is a robust test against the normality assumption, but it may be inappropriate when the assumption of homogeneity of variance has been violated. Welch ANOVA and the Kruskal-Wallis test (a non-parametric method) can be applicable for this case. In this study we compare the three methods in empirical type I error rate and power, when heterogeneity of variance occurs and find out which method is the most suitable with which cases including balanced/unbalanced, small/large sample size, and/or with normal/non-normal distributions.
Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou
Bayesian Semi- And Non-Parametric Analysis For Spatially Correlated Survival Data, Haiming Zhou
Theses and Dissertations
Flexible incorporation of both geographical patterning and risk effects in cancer survival models is becoming increasingly important, due in part to the recent availability of large cancer registries. The analysis of spatial survival data is challenged by the presence of spatial dependence and censoring for survival times. Accurately modeling the risk factors and geographical pattern that explain the differences in survival is particularly of interest. Within this dissertation, the first chapter reviews commonlyused baseline priors, semiparametric and nonparametric Bayesian survival models and recent approaches for accommodating spatial dependence, both conditional and marginal. The last three chapters contribute three flexible survival …
Graph-Based Regularization In Machine Learning: Discovering Driver Modules In Biological Networks, Xi Gao
Graph-Based Regularization In Machine Learning: Discovering Driver Modules In Biological Networks, Xi Gao
Theses and Dissertations
Curiosity of human nature drives us to explore the origins of what makes each of us different. From ancient legends and mythology, Mendel's law, Punnett square to modern genetic research, we carry on this old but eternal question. Thanks to technological revolution, today's scientists try to answer this question using easily measurable gene expression and other profiling data. However, the exploration can easily get lost in the data of growing volume, dimension, noise and complexity. This dissertation is aimed at developing new machine learning methods that take data from different classes as input, augment them with knowledge of feature relationships, …
Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe
Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe
Theses and Dissertations
Combining effect sizes from individual studies using random-effects models are commonly applied in high-dimensional gene expression data. However, unknown study heterogeneity can arise from inconsistency of sample qualities and experimental conditions. High heterogeneity of effect sizes can reduce statistical power of the models. We proposed two new methods for random effects estimation and measurements for model variation and strength of the study heterogeneity. We then developed a statistical technique to test for significance of random effects and identify heterogeneous genes. We also proposed another meta-analytic approach that incorporates informative weights in the random effects meta-analysis models. We compared the proposed …
Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima
Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima
Theses and Dissertations
In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not done randomly in observational studies, comparisons of outcomes between exposed and non-exposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of odds ratio and hazard ratio. However, there is a lack of research into the performance of propensity score methods for estimating the …
Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray
Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray
Theses and Dissertations
Functional neuroimaging is a relatively young discipline within the neurosciences that has led to significant advances in our understanding of the human brain and progress in neuroscientific research related to public health. Accurately identifying activated regions in the brain showing a strong association with an outcome of interest is crucial in terms of disease prediction and prevention. Functional magnetic resonance imaging (fMRI) is the most widely used method for this type of study as it has the ability to measure and identify the location of changes in tissue perfusion, blood oxygenation, and blood volume. In practice, the three-dimensional brain locations …
Semiparametric Regression Analysis Of Bivariate Interval-Censored Data, Naichen Wang
Semiparametric Regression Analysis Of Bivariate Interval-Censored Data, Naichen Wang
Theses and Dissertations
Survival analysis is a long-lasting and popular research area and has numerous applications in all fields such as social science, engineering, economics, industry, and public health. Interval-censored data are a special type of survival data, in which the survival time of interest is never exactly observed but is known to fall within some observed interval. Interval-censored data arise commonly in real-life studies, in which subjects are examined at periodical or irregular follow-up visits. In this dissertation, we develop efficient statistical approaches for regression analysis of bivariate intervalcensored data, in which the two survival times of interest are correlated and both …
Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta
Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta
Theses and Dissertations
The effects of scale on the analysis of spatial data, often referred to as the modifiable areal unit problem in spatial studies, is one of the issues often encountered in small area health models. These spatial effects of scale are also seen in the areas of disease mapping where data are usually available in counts. Often there is a need to consider the different scales of aggregation that exist within count data, since inferences based on analyses can vary if we change the definition of the unit of analysis. This thesis provides a framework that describes the distribution of relative …
Non- And Semi-Parametric Bayesian Inference With Recurrent Events And Coherent Systems Data, A. K. M. Fazlur Rahman
Non- And Semi-Parametric Bayesian Inference With Recurrent Events And Coherent Systems Data, A. K. M. Fazlur Rahman
Theses and Dissertations
This dissertation deals with non- and semi-parametric Bayesian inference of gap-time distribution with recurrent event data and simultaneous inference of component and system reliabilities of coherent systems data. Recurrent event data arise from a wide variety of studies/fields such as clinical trials, epidemiology, public health, biomedicine (e.g. repeated heart attack, repeated tumor occurrences of a cancer patient). In Chapter 2 we develop nonparametric Bayes and empirical Bayes estimators of the survivor function \bar{F} = 1 - F, of the gap-time distribution by assigning a Dirichlet process prior on F. We develop a closed form estimator of \bar{F} as well as …
Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello
Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello
Theses and Dissertations
In clinical settings, the diagnosis of medical conditions is often aided by measurement of various serum biomarkers through the use of laboratory tests. These biomarkers provide information about different aspects of a patient’s health and the overall function of different organs. In this dissertation, we develop and validate a weighted composite index that aggregates the information from a variety of health biomarkers covering multiple organ systems. The index can be used for predicting all-cause mortality and could also be used as a holistic measure of overall physiological health status. We refer to it as the Health Status Metric (HSM). Validation …
A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng
A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng
Theses and Dissertations
BACKGROUND: Little is known of the effects of obesity, body size and body composition, and blood pressure (BP) in childhood on hypertension (HBP) and cardiac structure and function in adulthood due to the lack of long-term serial data on these parameters from childhood into adulthood. In the present study, we are poised to analyze these serial data from the Fels Longitudinal Study (FLS) to evaluate the extent to which body size during childhood determines HBP and cardiac structure and function in the same individuals in adulthood through mathematical modeling. METHODS: The data were from 412 males and 403 females in …
Dynamic Bayesian Approaches To The Statistical Calibration Problem, Derick Lorenzo Rivers
Dynamic Bayesian Approaches To The Statistical Calibration Problem, Derick Lorenzo Rivers
Theses and Dissertations
The problem of statistical calibration of a measuring instrument can be framed both in a statistical context as well as in an engineering context. In the first, the problem is dealt with by distinguishing between the "classical" approach and the "inverse" regression approach. Both of these models are static models and are used to estimate "exact" measurements from measurements that are affected by error. In the engineering context, the variables of interest are considered to be taken at the time at which you observe the measurement. The Bayesian time series analysis method of Dynamic Linear Models (DLM) can be used …
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Theses and Dissertations
In the study of associated discrete variables, limitations on the range of the possible association measures (Pearson correlation, odds ratio, etc.) arise from the form of the joint probability function between the variables. These limitations are known as the Fréchet bounds. The bounds for cases involving associated binary variables are explored in the context of simulating datasets with a desired correlation and set of marginal probabilities. A new method for creating such datasets is compared to an existing method that uses the multivariate probit. A method for simulating associated binary variables using a desired odds ratio and known marginal probabilities …
Oldtimers & Newcomers In Collective Action, Stefanie R. Chamberlain
Oldtimers & Newcomers In Collective Action, Stefanie R. Chamberlain
Theses and Dissertations
Most work on groups facing collective action assumes that group membership is static, or fixed. Yet static membership is rare, with members joining and leaving groups. In this thesis, I propose to explore how the presence of newcomers to groups affects group coordination. Past research has shown an overall negative effect of newcomers on group contributions. The proposed thesis attempts to further establish the effect by determining whether newcomers, oldtimers, or both are responsible for the declining cooperation in groups. While the empirical component is focused solely on establishing who is responsible for driving down cooperation rates in dynamic groups, …
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Theses and Dissertations
In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …
Applications Of Bayesian Nonparametrics To Reliability And Survival Data, Li Li
Applications Of Bayesian Nonparametrics To Reliability And Survival Data, Li Li
Theses and Dissertations
Reliability and survival data are widely encountered across many common settings. Subjects under investigation often include machines, bioassays, patients, etc.; their reliability or survival distribution, and its association with covariate processes, are commonly of interest. Within this dissertation, the first two chapters focus on reliability data where repairable systems fail and get interventions, e.g. repairs in the event process. It begins with a nonparametric test for the commonly assumed ''good as old'' assumption for minimal repair models and then a semi-parametric regression model is introduced for reliability data using Kijima's effective age. The third chapter focuses on survival data observed …
Methods For Clustering Mixed Data, Jeanmarie L. Hendrickson
Methods For Clustering Mixed Data, Jeanmarie L. Hendrickson
Theses and Dissertations
We give a brief introduction to cluster analysis and then propose and discuss a few methods for clustering mixed data. In particular, a model-based clustering method for mixed data based on Everitt's (1988) work is described, and we use a simulated annealing method to estimate the parameters for Everitt's model. A penalized log likelihood with the simulated annealing method is proposed as a remedy for the parameter estimates being drawn to extremes. Everitt's approach and the proposed method are compared based on their performance in clustering simulated data. We then use the penalized log likelihood method on a heart disease …
Bayesian Analysis Of Continuous Curve Functions, Wen Cheng
Bayesian Analysis Of Continuous Curve Functions, Wen Cheng
Theses and Dissertations
We consider Bayesian analysis of continuous curve functions in 1D, 2D and 3D spaces. A fundamental feature of the analysis is that it is invariant under a simultaneous warping/re-parameterization of all target curves, as well as translation, rotation and scale of each individual if necessary. We introduce Bayesian models based on a special curve representation named Square Root Velocity Function (SRVF) introduced by Srivastava et al. (2011, IEEE PAMI). A Gaussian process model for the SRVFs of curves is proposed, and suitable prior models such as the Dirichlet distribution are employed for modeling the warping function as a cumulative distribution …
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Theses and Dissertations
Latent variable models (LVMs) are commonly used in the scenario where the outcome of the main interest is an unobservable measure, associated with multiple observed surrogate outcomes, and affected by potential risk factors. This thesis develops an approach of efficient handling missing surrogate outcomes and covariates in two- and three-level latent variable models. However, corresponding statistical methodologies and computational software are lacking efficiently analyzing the LVMs given surrogate outcomes and covariates subject to missingness in the LVMs. We analyze the two-level LVMs for longitudinal data from the National Growth of Health Study where surrogate outcomes and covariates are subject to …
The Total Picture: Multiple Chemical Exposures To Pregnant Women In The Us – An Nhanes Study Of Data From 2003 Through 2010, Teri Cabana
Theses and Dissertations
INTRODUCTION: Chemical exposures to US pregnant women have been shown to have adverse health impacts on both mother and fetus. A prior paper revealed that US pregnant women in 2003-2004 had widespread exposure to multiple chemicals. The goal of this research is to examine how environmental chemical exposures to US pregnant women have changed from 2003 to 2010 and to look further at the extent of simultaneous exposure to multiple chemicals in US pregnant women using biomonitoring data available through NHANES (the National Health and Nutritional Examination Survey). METHODS: Using available NHANES data from the following cycles (2003-2004, 2005-2006, 2007-2008, …
Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou
Theses and Dissertations
Ordinal scales are commonly used to measure health status and disease related outcomes in hospital settings as well as in translational medical research. Notable examples include cancer staging, which is a five-category ordinal scale indicating tumor size, node involvement, and likelihood of metastasizing. Glasgow Coma Scale (GCS), which gives a reliable and objective assessment of conscious status of a patient, is an ordinal scaled measure. In addition, repeated measurements are common in clinical practice for tracking and monitoring the progression of complex diseases. Classical ordinal modeling methods based on the likelihood approach have contributed to the analysis of data in …
Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri
Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri
Theses and Dissertations
O'Brien and Fleming (1979) proposed a straightforward and useful multiple testing procedure (group sequential testing procedure) for comparing two treatments in clinical trials where subject responses are dichotomous (e.g. success and failure). O'Brien and Fleming stated that their group sequential testing procedure has the same Type I error rate and power as that of a fixed one-stage chi-square test, but gives the opportunity to terminate the trial early when one treatment is clearly performing better than the other. We studied and tested the O'Brien and Fleming procedure specifically by correcting the originally proposed critical values. Furthermore, we updated the O’Brien …
Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks
Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks
Theses and Dissertations
Response adaptive designs intend to allocate more patients to better treatments without undermining the validity and the integrity of the trial. The immediacy of the primary response (e.g. deaths, remission) determines the efficiency of the response adaptive design, which often requires outcomes to be quickly or immediately observed. This presents difficulties for survival studies, which may require long durations to observe the primary endpoint. Therefore, we introduce auxiliary endpoints to assist the adaptation with the primary endpoint, where an auxiliary endpoint is generally defined as any measurement that is positively associated with the primary endpoint. Our proposed design (referred to …
Construction, Analysis, And Data-Driven Augmentation Of Supersaturated Designs, Alex J. Gutman
Construction, Analysis, And Data-Driven Augmentation Of Supersaturated Designs, Alex J. Gutman
Theses and Dissertations
Screening designs are used in the early stages of industrial and computer experiments to find the most important input factors affecting a system's output. They provide an economical way to remove unimportant factors from further, potentially costly, experimentation. However, when an experiment has a large number of control factors and limited number of available runs, it is infeasible to run a traditional screening design. In these situations, experimenters can use supersaturated designs. A supersaturated design is a fractional factorial design that can screen a set of k factors in n runs, where k is greater than n -1. Unfortunately, they …
Bayesian And Classical Regression Models Inference When The Errors Follow Skewed Distributions, Ebtisam Karim Abdulah
Bayesian And Classical Regression Models Inference When The Errors Follow Skewed Distributions, Ebtisam Karim Abdulah
Theses and Dissertations
The Gamma distribution is one of the most popular distributions for reliability and lifetime data and can be used effectively in analyzing positive skewed data. In real life applications, analysts might wish to have distributions for analyzing positive and negative skewed data, therefore, in this century, we have seen a good attention for fitting data using skew distributions, because it allows continuous variations from symmetric to non-symmetric. Some empirical data, especially, finance (prices and returns) and environmental data have peak distributions and involve tail behavior which affects the model assumptions. In this dissertation, we extend two of the symmetric distributions, …
The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk
The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk
Theses and Dissertations
Many continuous medical tests often rely on a threshold for diagnosis. There are two sequential testing strategies of interest: Believe the Positive (BP) and Believe the Negative (BN). BP classifies a patient positive if either the first test is greater than a threshold θ1 or negative on the first test and greater than θ2 on the second test. BN classifies a patient positive if the first test is greater than a threshold θ3 and greater than θ4 on the second test. Threshold pairs θ = (θ1, θ2) or (θ3, θ4), depending on strategy, are defined as optimal if they maximized …
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Theses and Dissertations
Survival Analysis generally uses the median survival time as a common summary statistic. While the median possesses the desirable characteristic of being unbiased, there are times when it is not the best statistic to describe the data at hand. Royston and Parmar (2011) provide an argument that the restricted mean survival time should be the summary statistic used when the proportional hazards assumption is in doubt. Work in Restricted Means dates back to 1949 when J.O. Irwin developed a calculation for the standard error of the restricted mean using Greenwood’s formula. Since then the development of the restricted mean has …
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Theses and Dissertations
Batch effects are due to probe-specific systematic variation between groups of samples (batches) resulting from experimental features that are not of biological interest. Principal components analysis (PCA) is commonly used as a visual tool to determine whether batch effects exist after applying a global normalization method. However, PCA yields linear combinations of the variables that contribute maximum variance and thus will not necessarily detect batch effects if they are not the largest source of variability in the data. We present an extension of principal components analysis to quantify the existence of batch effects, called guided PCA (gPCA). We describe a …