Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (16)
- Applied Statistics (15)
- Life Sciences (14)
- Statistical Models (13)
- Applied Mathematics (7)
-
- Longitudinal Data Analysis and Time Series (7)
- Public Health (7)
- Statistical Methodology (7)
- Bioinformatics (6)
- Genetics and Genomics (6)
- Genetics (5)
- Microarrays (4)
- Multivariate Analysis (4)
- Data Science (3)
- Mathematics (3)
- Medical Sciences (3)
- Social and Behavioral Sciences (3)
- Categorical Data Analysis (2)
- Environmental Sciences (2)
- Genomics (2)
- Medical Specialties (2)
- Mental and Social Health (2)
- Natural Resources Management and Policy (2)
- Natural Resources and Conservation (2)
- Ordinary Differential Equations and Applied Dynamics (2)
- Other Genetics and Genomics (2)
- Plant Sciences (2)
- Institution
- Keyword
-
- Biostatistics (8)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Bayesian (4)
- Clinical trial (4)
- Genetics (4)
-
- Genomics (4)
- Longitudinal data (4)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (4)
- Survival Analysis (4)
- Classification (3)
- Cluster analysis (3)
- Cox proportional hazards model (3)
- Environment (3)
- Gene expression (3)
- Power (3)
- Spatial clustering (3)
- Variable selection (3)
- Adaptive design (2)
- Algorithm (2)
- Bayesian inference (2)
- Bioinformatics (2)
- Chemicals (2)
- Clustering (2)
- Copy number variation (2)
- Correlated data (2)
- Count Data (2)
- Covariance (2)
- DNA Methylation (2)
- Data (2)
- Data analysis (2)
Articles 91 - 120 of 162
Full-Text Articles in Biostatistics
Regression Models For Count Data Based On The Double Poisson Distribution, Rebecca Wardrop
Regression Models For Count Data Based On The Double Poisson Distribution, Rebecca Wardrop
Theses and Dissertations
This paper explores the double Poisson distribution. The probability mass function and the difficulties associated with derivative-based optimization for this distribution are discussed. Stata software developed for estimation of double Poisson regression is detailed. Simulations are used to test the software. Data which are over-, under-, and equidispersed relative to the Poisson are generated and the software is utilized to estimate a regression model, a zero-inflated model, and a marginalized zero-inflated model all based on the double Poisson distribution. The estimated power of the test for φ = 1 for the double Poisson models are compared to the power of …
Score Test Derivations And Implementations For Bivariate Probability Mass And Density Functions With An Application To Copula Functions, Roy Bower
Theses and Dissertations
This dissertation is comprised and grounded in statistical theory with an application to solving real world problems. In particular, the development and implementation of multiple score tests under a variety of scenarios are derived, applied, and interpreted. In chapter 2, I propose a score test for independence of the marginals based on Lakshminarayana’s bivariate Poisson distribution. Each marginal distribution of the bivariate model is a univariate Poisson distribution, and the parameters of the bivariate distribution can be estimated using maximum likelihood methods. The simulation study shows that the score test maintains size close to the nominal level. To assess the …
Semiparametric Estimation Methods For Complex Accelerated Failure Time Model, Yinding Wang
Semiparametric Estimation Methods For Complex Accelerated Failure Time Model, Yinding Wang
Theses and Dissertations
The proportional hazards (PH) model and the accelerated failure time (AFT) model are the two most popular survival models in fitting the right-censored data. The AFT model is a useful alternative to the PH model, particularly when the PH assumption is not satisfied. Usually, the linear association is assumed with logarithm of survival time in the AFT model. However, the nonlinear association may exist in practice. The first project aims to handle the nonlinear component in the AFT model, which is called the semiparametric additive partial accelerated failure time (AP-AFT) model. Two estimation methods based on the rank-smooth method and …
A Weighted Gene Co-Expression Network Analysis For Streptococcus Sanguinis Microarray Experiments, Erik C. Dvergsten
A Weighted Gene Co-Expression Network Analysis For Streptococcus Sanguinis Microarray Experiments, Erik C. Dvergsten
Theses and Dissertations
Streptococcus sanguinis is a gram-positive, non-motile bacterium native to human mouths. It is the primary cause of endocarditis and is also responsible for tooth decay. Two-component systems (TCSs) are commonly found in bacteria. In response to environmental signals, TCSs may regulate the expression of virulence factor genes.
Gene co-expression networks are exploratory tools used to analyze system-level gene functionality. A gene co-expression network consists of gene expression profiles represented as nodes and gene connections, which occur if two genes are significantly co-expressed. An adjacency function transforms the similarity matrix containing co-expression similarities into the adjacency matrix containing connection strengths. Gene …
Selecting Spatial Scale Of Area-Level Covariates In Regression Models, Lauren Grant
Selecting Spatial Scale Of Area-Level Covariates In Regression Models, Lauren Grant
Theses and Dissertations
Studies have found that the level of association between an area-level covariate and an outcome can vary depending on the spatial scale (SS) of a particular covariate. However, covariates used in regression models are customarily modeled at the same spatial unit. In this dissertation, we developed four SS model selection algorithms that select the best spatial scale for each area-level covariate. The SS forward stepwise, SS incremental forward stagewise, SS least angle regression (LARS), and SS lasso algorithms allow for the selection of different area-level covariates at different spatial scales, while constraining each covariate to enter at most one spatial …
Meta-Analytic Estimation Techniques For Non-Convergent Repeated-Measure Clustered Data, Aobo Wang
Meta-Analytic Estimation Techniques For Non-Convergent Repeated-Measure Clustered Data, Aobo Wang
Theses and Dissertations
Clustered data often feature nested structures and repeated measures. If coupled with binary outcomes and large samples (>10,000), this complexity can lead to non-convergence problems for the desired model especially if random effects are used to account for the clustering. One way to bypass the convergence problem is to split the dataset into small enough sub-samples for which the desired model convergences, and then recombine results from those sub-samples through meta-analysis. We consider two ways to generate sub-samples: the K independent samples approach where the data are split into k mutually-exclusive sub-samples, and the cluster-based approach where naturally existing …
Proof-Of-Concept Of Environmental Dna Tools For Atlantic Sturgeon Management, Jameson Hinkle
Proof-Of-Concept Of Environmental Dna Tools For Atlantic Sturgeon Management, Jameson Hinkle
Theses and Dissertations
Abstract
The Atlantic Sturgeon (Acipenser oxyrinchus oxyrinchus, Mitchell) is an anadromous species that spawns in tidal freshwater rivers from Canada to Florida. Overfishing, river sedimentation and alteration of the river bottom have decreased Atlantic Sturgeon populations, and NOAA lists the species as endangered. Ecologists sometimes find it difficult to locate individuals of a species that is rare, endangered or invasive. The need for methods less invasive that can create more resolution of cryptic species presence is necessary. Environmental DNA (eDNA) is a non-invasive means of detecting rare, endangered, or invasive species by isolating nuclear or mitochondrial DNA (mtDNA) from the …
High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski
High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski
Theses and Dissertations
The advent of high-throughput sequencing has brought about the creation of an unprecedented amount of research data. Analytical methodology has not been able to keep pace with the plethora of data being produced. Two assays, ImmunoSEQ and the cytokinesisblock micronucleus (CBMN), that both produce count data and have few methods available to analyze them are considered.
ImmunoSEQ is a sequencing assay that measures the beta T-cell receptor (TCR) repertoire. The ImmunoSEQ assay was used to describe the TCR repertoires of patients that have undergone hematopoietic stem cell transplantation (HSCT). Several different methods for spectratype analysis were extended to the TCR …
Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu
Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu
Theses and Dissertations
Analysis of variance (ANOVA) is a robust test against the normality assumption, but it may be inappropriate when the assumption of homogeneity of variance has been violated. Welch ANOVA and the Kruskal-Wallis test (a non-parametric method) can be applicable for this case. In this study we compare the three methods in empirical type I error rate and power, when heterogeneity of variance occurs and find out which method is the most suitable with which cases including balanced/unbalanced, small/large sample size, and/or with normal/non-normal distributions.
Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe
Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe
Theses and Dissertations
Combining effect sizes from individual studies using random-effects models are commonly applied in high-dimensional gene expression data. However, unknown study heterogeneity can arise from inconsistency of sample qualities and experimental conditions. High heterogeneity of effect sizes can reduce statistical power of the models. We proposed two new methods for random effects estimation and measurements for model variation and strength of the study heterogeneity. We then developed a statistical technique to test for significance of random effects and identify heterogeneous genes. We also proposed another meta-analytic approach that incorporates informative weights in the random effects meta-analysis models. We compared the proposed …
Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima
Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima
Theses and Dissertations
In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not done randomly in observational studies, comparisons of outcomes between exposed and non-exposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of odds ratio and hazard ratio. However, there is a lack of research into the performance of propensity score methods for estimating the …
Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray
Methods For Identifying Regions Of Brain Activation Using Fmri Meta-Data, Meredith A. Ray
Theses and Dissertations
Functional neuroimaging is a relatively young discipline within the neurosciences that has led to significant advances in our understanding of the human brain and progress in neuroscientific research related to public health. Accurately identifying activated regions in the brain showing a strong association with an outcome of interest is crucial in terms of disease prediction and prevention. Functional magnetic resonance imaging (fMRI) is the most widely used method for this type of study as it has the ability to measure and identify the location of changes in tissue perfusion, blood oxygenation, and blood volume. In practice, the three-dimensional brain locations …
Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta
Simulation Based Evaluation Of Multiscale Small Area Health Models, Purbasha Dasgupta
Theses and Dissertations
The effects of scale on the analysis of spatial data, often referred to as the modifiable areal unit problem in spatial studies, is one of the issues often encountered in small area health models. These spatial effects of scale are also seen in the areas of disease mapping where data are usually available in counts. Often there is a need to consider the different scales of aggregation that exist within count data, since inferences based on analyses can vary if we change the definition of the unit of analysis. This thesis provides a framework that describes the distribution of relative …
Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello
Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello
Theses and Dissertations
In clinical settings, the diagnosis of medical conditions is often aided by measurement of various serum biomarkers through the use of laboratory tests. These biomarkers provide information about different aspects of a patient’s health and the overall function of different organs. In this dissertation, we develop and validate a weighted composite index that aggregates the information from a variety of health biomarkers covering multiple organ systems. The index can be used for predicting all-cause mortality and could also be used as a holistic measure of overall physiological health status. We refer to it as the Health Status Metric (HSM). Validation …
A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng
A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng
Theses and Dissertations
BACKGROUND: Little is known of the effects of obesity, body size and body composition, and blood pressure (BP) in childhood on hypertension (HBP) and cardiac structure and function in adulthood due to the lack of long-term serial data on these parameters from childhood into adulthood. In the present study, we are poised to analyze these serial data from the Fels Longitudinal Study (FLS) to evaluate the extent to which body size during childhood determines HBP and cardiac structure and function in the same individuals in adulthood through mathematical modeling. METHODS: The data were from 412 males and 403 females in …
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Theses and Dissertations
In the study of associated discrete variables, limitations on the range of the possible association measures (Pearson correlation, odds ratio, etc.) arise from the form of the joint probability function between the variables. These limitations are known as the Fréchet bounds. The bounds for cases involving associated binary variables are explored in the context of simulating datasets with a desired correlation and set of marginal probabilities. A new method for creating such datasets is compared to an existing method that uses the multivariate probit. A method for simulating associated binary variables using a desired odds ratio and known marginal probabilities …
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Theses and Dissertations
In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Theses and Dissertations
Latent variable models (LVMs) are commonly used in the scenario where the outcome of the main interest is an unobservable measure, associated with multiple observed surrogate outcomes, and affected by potential risk factors. This thesis develops an approach of efficient handling missing surrogate outcomes and covariates in two- and three-level latent variable models. However, corresponding statistical methodologies and computational software are lacking efficiently analyzing the LVMs given surrogate outcomes and covariates subject to missingness in the LVMs. We analyze the two-level LVMs for longitudinal data from the National Growth of Health Study where surrogate outcomes and covariates are subject to …
The Total Picture: Multiple Chemical Exposures To Pregnant Women In The Us – An Nhanes Study Of Data From 2003 Through 2010, Teri Cabana
Theses and Dissertations
INTRODUCTION: Chemical exposures to US pregnant women have been shown to have adverse health impacts on both mother and fetus. A prior paper revealed that US pregnant women in 2003-2004 had widespread exposure to multiple chemicals. The goal of this research is to examine how environmental chemical exposures to US pregnant women have changed from 2003 to 2010 and to look further at the extent of simultaneous exposure to multiple chemicals in US pregnant women using biomonitoring data available through NHANES (the National Health and Nutritional Examination Survey). METHODS: Using available NHANES data from the following cycles (2003-2004, 2005-2006, 2007-2008, …
Regularization Methods For Predicting An Ordinal Response Using Longitudinal High-Dimensional Genomic Data, Jiayi Hou
Theses and Dissertations
Ordinal scales are commonly used to measure health status and disease related outcomes in hospital settings as well as in translational medical research. Notable examples include cancer staging, which is a five-category ordinal scale indicating tumor size, node involvement, and likelihood of metastasizing. Glasgow Coma Scale (GCS), which gives a reliable and objective assessment of conscious status of a patient, is an ordinal scaled measure. In addition, repeated measurements are common in clinical practice for tracking and monitoring the progression of complex diseases. Classical ordinal modeling methods based on the likelihood approach have contributed to the analysis of data in …
Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri
Review And Extension For The O’Brien Fleming Multiple Testing Procedure, Hanan Hammouri
Theses and Dissertations
O'Brien and Fleming (1979) proposed a straightforward and useful multiple testing procedure (group sequential testing procedure) for comparing two treatments in clinical trials where subject responses are dichotomous (e.g. success and failure). O'Brien and Fleming stated that their group sequential testing procedure has the same Type I error rate and power as that of a fixed one-stage chi-square test, but gives the opportunity to terminate the trial early when one treatment is clearly performing better than the other. We studied and tested the O'Brien and Fleming procedure specifically by correcting the originally proposed critical values. Furthermore, we updated the O’Brien …
Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks
Response Adaptive Design Using Auxiliary And Primary Outcomes, Shuxian Sinks
Theses and Dissertations
Response adaptive designs intend to allocate more patients to better treatments without undermining the validity and the integrity of the trial. The immediacy of the primary response (e.g. deaths, remission) determines the efficiency of the response adaptive design, which often requires outcomes to be quickly or immediately observed. This presents difficulties for survival studies, which may require long durations to observe the primary endpoint. Therefore, we introduce auxiliary endpoints to assist the adaptation with the primary endpoint, where an auxiliary endpoint is generally defined as any measurement that is positively associated with the primary endpoint. Our proposed design (referred to …
The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk
The Estimation And Evaluation Of Optimal Thresholds For Two Sequential Testing Strategies, Amber R. Wilk
Theses and Dissertations
Many continuous medical tests often rely on a threshold for diagnosis. There are two sequential testing strategies of interest: Believe the Positive (BP) and Believe the Negative (BN). BP classifies a patient positive if either the first test is greater than a threshold θ1 or negative on the first test and greater than θ2 on the second test. BN classifies a patient positive if the first test is greater than a threshold θ3 and greater than θ4 on the second test. Threshold pairs θ = (θ1, θ2) or (θ3, θ4), depending on strategy, are defined as optimal if they maximized …
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Choosing The Cut Point For A Restricted Mean In Survival Analysis, A Data Driven Method, Emily H. Sheldon
Theses and Dissertations
Survival Analysis generally uses the median survival time as a common summary statistic. While the median possesses the desirable characteristic of being unbiased, there are times when it is not the best statistic to describe the data at hand. Royston and Parmar (2011) provide an argument that the restricted mean survival time should be the summary statistic used when the proportional hazards assumption is in doubt. Work in Restricted Means dates back to 1949 when J.O. Irwin developed a calculation for the standard error of the restricted mean using Greenwood’s formula. Since then the development of the restricted mean has …
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Detecting And Correcting Batch Effects In High-Throughput Genomic Experiments, Sarah Reese
Theses and Dissertations
Batch effects are due to probe-specific systematic variation between groups of samples (batches) resulting from experimental features that are not of biological interest. Principal components analysis (PCA) is commonly used as a visual tool to determine whether batch effects exist after applying a global normalization method. However, PCA yields linear combinations of the variables that contribute maximum variance and thus will not necessarily detect batch effects if they are not the largest source of variability in the data. We present an extension of principal components analysis to quantify the existence of batch effects, called guided PCA (gPCA). We describe a …
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Characterization Of A Weighted Quantile Score Approach For Highly Correlated Data In Risk Analysis Scenarios, Caroline Carrico
Theses and Dissertations
In risk evaluation, the effect of mixtures of environmental chemicals on a common adverse outcome is of interest. However, due to the high dimensionality and inherent correlations among chemicals that occur together, the traditional methods (e.g. ordinary or logistic regression) are unsuitable. We extend and characterize a weighted quantile score (WQS) approach to estimating an index for a set of highly correlated components. In the case with environmental chemicals, we use the WQS to identify “bad actors” and estimate body burden. The accuracy of the WQS was evaluated through extensive simulation studies in terms of validity (ability of the WQS …
Accounting For Model Uncertainty In Linear Mixed-Effects Models, Adam Sima
Accounting For Model Uncertainty In Linear Mixed-Effects Models, Adam Sima
Theses and Dissertations
Standard statistical decision-making tools, such as inference, confidence intervals and forecasting, are contingent on the assumption that the statistical model used in the analysis is the true model. In linear mixed-effect models, ignoring model uncertainty results in an underestimation of the residual variance, contributing to hypothesis tests that demonstrate larger than nominal Type-I errors and confidence intervals with smaller than nominal coverage probabilities. A novel utilization of the generalized degrees of freedom developed by Zhang et al. (2012) is used to adjust the estimate of the residual variance for model uncertainty. Additionally, the general global linear approximation is extended to …
Heaped Data In Count Models, Tammy Harris
Heaped Data In Count Models, Tammy Harris
Theses and Dissertations
Heaped data result when subjects who recall the frequency of events prefer for reporting from a limited set of rounded responses or preferred digits over reporting exact counts. These rounded responses and digit preferences (also referred to as data coarsening) could be characterized by reported frequencies (or counts) favoring multiples of 20, reporting counts ending with 0 or 5, or a preference for reporting an even number over an odd number or vice versa. This mixture of values is a type of measurement error (pattern of misreporting) that can lead to biased estimation and imprecision in discrete quantitative data. Sometimes …
Models And Software Development For Interval-Censored Data, Chun Pan
Models And Software Development For Interval-Censored Data, Chun Pan
Theses and Dissertations
Interval-censored time-to-event data occur naturally in studies of diseases where the symptoms are not directly observable, and periodic clinical examinations are required for detection. Due to the lack of well-established procedures, interval-censored data have been conventionally treated as right-censored data, however, this introduces bias at the first place. This dissertation focuses on methodological research and software development for interval-censored data. Specifically, it consists of three projects. The first project is to create an R package for regression analysis and survival curve estimation of interval-censored data based on several published papers by our research team. In the second project, a Bayesian …
A New Method For The Comparison Of Survival Distributions, Jaymie Shanahan
A New Method For The Comparison Of Survival Distributions, Jaymie Shanahan
Theses and Dissertations
The assessment of overall homogeneity of time-to-event curves is a key element in survival analysis in biomedical research. The currently commonly used testing methods, e.g. log-rank test, Wilcoxon test, and Kolmogorov-Smirnov test, may have a significant loss of statistical testing power under certain circumstances. In this thesis we replicate a testing method (Lin & Xu, 2009) that is robust for the comparison of the overall homogeneity of survival curves based on the absolute difference of the area under the survival curves using normal approximation by Greenwood's formula, and propose a new weight component to their test statistic. The weight component …