Open Access. Powered by Scholars. Published by Universities.®

Biostatistics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Theses and Dissertations

Discipline
Institution
Keyword
Publication Year

Articles 61 - 90 of 162

Full-Text Articles in Biostatistics

Randomization Analysis Driven Software, Steph-Yves Louis Apr 2019

Randomization Analysis Driven Software, Steph-Yves Louis

Theses and Dissertations

The application of a method of randomization for a clinical trial frequently summarizes to using Simple Randomization. Even though the latter method provides favorable characteristics, if the collected sample is not large enough, it still presents the highest chance of imbalance both marginally in the treatment groups and locally in terms of the covariates. Methods of Permuted Block Randomization, Urn Randomization, Stratified Permuted Block Randomization, and Minimization represent popular alternative methods that one should consider depending on the goal of the study. A comparison of the previously mentioned methods is carried to evaluate their performance with samples that are not …


Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao Jan 2019

Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao

Theses and Dissertations

In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …


Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace Jan 2019

Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace

Theses and Dissertations

Response-Adaptive (RA) designs are used to adaptively allocate patients in clinical trials. These methods have been generalized to include Covariate-Adjusted Response-Adaptive (CARA) designs, which adjust treatment assignments for a set of covariates while maintaining features of the RA designs. Challenges may arise in multi-center trials if differential treatment responses and/or effects among sites exist. We propose Site-Adjusted Response-Adaptive (SARA) approaches to account for inter-center variability in treatment response and/or effectiveness, including either a fixed site effect or both random site and treatment-by-site interaction effects to calculate conditional probabilities. These success probabilities are used to update assignment probabilities for allocating patients …


Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer Jan 2019

Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer

Theses and Dissertations

As researchers increasingly use web-based surveys, the ease of dropping out in the online setting is a growing issue in ensuring data quality. One theory is that dropout or attrition occurs in phases that can be generalized to phases of high dropout and phases of stable use. In order to detect these phases, several methods are explored. First, existing methods and user-specified thresholds are applied to survey data where significant changes in the dropout rate between two questions is interpreted as the start or end of a high dropout phase. Next, survey dropout is considered as a time-to-event outcome and …


Assessing The Impact Of Incorporating Residential Histories Into The Spatial Analysis Of Cancer Risk, Anny-Claude Joseph Jan 2019

Assessing The Impact Of Incorporating Residential Histories Into The Spatial Analysis Of Cancer Risk, Anny-Claude Joseph

Theses and Dissertations

In many spatial epidemiologic studies, investigators use residential location at diagnosis as a surrogate for unknown environmental exposures or as a geographic basis for assigning measured exposures. Inherently, they make assumptions about the timing and location of pertinent exposures which may prove problematic when studying long latency diseases such as cancer.

In this work we explored how the association between environmental exposures and disease risk for long-latency health outcomes like cancer is affected by residential mobility. We used simulation studies conditioned on real data to evaluate the extent to which the commonly held assumption of no residential mobility 1) affected …


Methods For Joint Normalization And Comparison Of Hi-C Data, John C. Stansfield Jan 2019

Methods For Joint Normalization And Comparison Of Hi-C Data, John C. Stansfield

Theses and Dissertations

The development of chromatin conformation capture technology has opened new avenues of study into the 3D structure and function of the genome. Chromatin structure is known to influence gene regulation, and differences in structure are now emerging as a mechanism of regulation between, e.g., cell differentiation and disease vs. normal states. Hi-C sequencing technology now provides a way to study the 3D interactions of the chromatin over the whole genome. However, like all sequencing technologies, Hi-C suffers from several forms of bias stemming from both the technology and the DNA sequence itself. Several normalization methods have been developed for normalizing …


Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna Jan 2019

Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna

Theses and Dissertations

Widely effective treatment for alcohol use disorder is not yet available, because the exact biological mechanisms that underlie this disorder are not completely understood. One way to gain a better understanding of these mechanisms is to examine the genetic frameworks that contribute to the risk for developing this disorder. This dissertation examines genetic association data in combination with gene expression networks in the brain to identify functional groups of genes associated with alcohol consumption and dependence.

The first study took advantage of the behavioral complexity of human samples, and experimental capabilities provided by mouse models, by co-analyzing gene expression networks …


Spectral Methods For The Detection And Characterization Of Topologically Associated Domains, Kellen Garrison Cresswell Jan 2019

Spectral Methods For The Detection And Characterization Of Topologically Associated Domains, Kellen Garrison Cresswell

Theses and Dissertations

The three-dimensional (3D) structure of the genome plays a crucial role in gene expression regulation. Chromatin conformation capture technologies (Hi-C) have revealed that the genome is organized in a hierarchy of topologically associated domains (TADs), sub-TADs, and chromatin loops which is relatively stable across cell-lines and even across species. These TADs dynamically reorganize during development of disease, and exhibit cell- and conditionspecific differences. Identifying such hierarchical structures and how they change between conditions is a critical step in understanding genome regulation and disease development. Despite their importance, there are relatively few tools for identification of TADs and even fewer for …


Clustering Biological Data With Self-Adjusting High-Dimensional Sieve, Josselyn Gonzalez Apr 2018

Clustering Biological Data With Self-Adjusting High-Dimensional Sieve, Josselyn Gonzalez

Theses and Dissertations

Data classification as a preprocessing technique is a crucial step in the analysis and understanding of numerical data. Cluster analysis, in particular, provides insight into the inherent patterns found in data which makes the interpretation of any follow-up analyses more meaningful. A clustering algorithm groups together data points according to a predefined similarity criterion. This allows the data set to be broken up into segments which, in turn, gives way for a more targeted statistical analysis. Cluster analysis has applications in numerous fields of study and, as a result, countless algorithms have been developed. However, the quantity of options makes …


Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry Jan 2018

Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry

Theses and Dissertations

The Brisbane Longitudinal Twin Study (BLTS) was being conducted in Australia and was funded by the US National Institute on Drug Abuse (NIDA). Adolescent twins were sampled as a part of this study and surveyed about their substance use as part of the Pathways to Cannabis Use, Abuse and Dependence project. The methods developed in this dissertation were designed for the purpose of analyzing a subset of the Pathways data that includes demographics, cannabis use metrics, personality measures, and imputed genotypes (SNPs) for 493 complete twin pairs (986 subjects.) The primary goal was to determine what combination of SNPs and …


Adjusting For Mis-Reporting In Count Data, Gelareh Rahimighazikalayeh Jan 2018

Adjusting For Mis-Reporting In Count Data, Gelareh Rahimighazikalayeh

Theses and Dissertations

Any counting system is prone to recording errors including underreporting and overreporting. Ignoring the misreporting pattern in count data can give rise to bias in the estimation of model parameters. Accordingly, Poisson, negative binomial and generalized Poisson regression have been expanded in some instances to capture reporting biases. However, to our knowledge, no program has been developed to allow users to apply all of these models when needed. In the first part of the dissertation, we review the available models for underreported counts and develop a Stata command to estimate Poisson, negative binomial and generalized Poisson regression models for underreported …


Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard Jan 2018

Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard

Theses and Dissertations

Linear regression is a widely used method for analysis that is well understood across a wide variety of disciplines. In order to use linear regression, a number of assumptions must be met. These assumptions, specifically normality and homoscedasticity of the error distribution can at best be met only approximately with real data. Quantile regression requires fewer assumptions, which offers a potential advantage over linear regression. In this simulation study, we compare the performance of linear (least squares) regression to quantile regression when these assumptions are violated, in order to investigate under what conditions quantile regression becomes the more advantageous method …


Estimation Procedures For Complex Survival Models And Their Applications In Epidemiology Studies, Jie Zhou Jan 2018

Estimation Procedures For Complex Survival Models And Their Applications In Epidemiology Studies, Jie Zhou

Theses and Dissertations

In this dissertation, we aim to address three important questions in practice, which can be solved through complex survival models. The first project focuses on studying the longitudinal fitness effect on cardiovascular disease (CVD) mortality. In the second project, we study the disease-death relation between CVD and all-cause mortality and evaluate important covariate effects on the disease or death transitions. In the third project, we compare antiretroviral treatment (ART) for HIV patients and consider both treatment effect and side effect of the drugs. The first two projects are motivated by the Aerobics Center Longitudinal Study (ACLS) datasets and the third …


Examining The Confirmatory Tetrad Analysis (Cta) As A Solution Of The Inadequacy Of Traditional Structural Equation Modeling (Sem) Fit Indices, Hangcheng Liu Jan 2018

Examining The Confirmatory Tetrad Analysis (Cta) As A Solution Of The Inadequacy Of Traditional Structural Equation Modeling (Sem) Fit Indices, Hangcheng Liu

Theses and Dissertations

Structural Equation Modeling (SEM) is a framework of statistical methods that allows us to represent complex relationships between variables. SEM is widely used in economics, genetics and the behavioral sciences (e.g. psychology, psychobiology, sociology and medicine). Model complexity is defined as a model’s ability to fit different data patterns and it plays an important role in model selection when applying SEM. As in linear regression, the number of free model parameters is typically used in traditional SEM model fit indices as a measure of the model complexity. However, only using number of free model parameters to indicate SEM model complexity …


Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang Jan 2018

Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang

Theses and Dissertations

Modern big data often emerge as tensors. Standard statistical methods are inadequate to deal with datasets of large volume, high dimensionality, and complex structure. Therefore, it is important to develop algorithms such as low-rank tensor decomposition for data compression, dimensionality reduction, and approximation.

With the advancement in technology, high-dimensional images are becoming ubiquitous in the medical field. In lung radiation therapy, the respiratory motion of the lung introduces variabilities during treatment as the tumor inside the lung is moving, which brings challenges to the precise delivery of radiation to the tumor. Several approaches to quantifying this uncertainty propose using a …


The Generalized Monotone Incremental Forward Stagewise Method For Modeling Longitudinal, Clustered, And Overdispersed Count Data: Application Predicting Nuclear Bud And Micronuclei Frequencies, Rebecca Lehman Jan 2017

The Generalized Monotone Incremental Forward Stagewise Method For Modeling Longitudinal, Clustered, And Overdispersed Count Data: Application Predicting Nuclear Bud And Micronuclei Frequencies, Rebecca Lehman

Theses and Dissertations

With the influx of high-dimensional data there is an immediate need for statistical methods that are able to handle situations when the number of predictors greatly exceeds the number of samples. One such area of growth is in examining how environmental exposures to toxins impact the body long term. The cytokinesis-block micronucleus assay can measure the genotoxic effect of exposure as a count outcome. To investigate potential biomarkers, high-throughput assays that assess gene expression and methylation have been developed. It is of interest to identify biomarkers or molecular features that are associated with elevated micronuclei (MN) or nuclear bud (Nbud) …


Longitudinal And Geographical Modeling Of Circular Data With An Application To Sudden Infant Death Syndrome, Xinyan Cai Jan 2017

Longitudinal And Geographical Modeling Of Circular Data With An Application To Sudden Infant Death Syndrome, Xinyan Cai

Theses and Dissertations

The aim of this thesis is to study seasonality of death in U.S. infants who died from SIDS. We also propose to investigate secular trends and geographical patterns of seasonal patterns of mortality. The application of circular statistics is used to describe the seasonality of the month of death in infants who died from SIDS in 1990, 2000 and 2010. The secular trends of seasonal patterns of SIDS mortality are investigated using a circular linear regression model after adjusting for potential confounders. The geographical variation in seasonal patterns of SIDS mortality is explored from the U.S. map and quantified by …


Statistical Methods For Multivariate And Correlated Data, Xinling Xu Jan 2017

Statistical Methods For Multivariate And Correlated Data, Xinling Xu

Theses and Dissertations

A commonly encountered data type in real life is count data, especially in selfreported behavioral studies. One issue of the self-reported count data is the inaccuracy. In the first part of the dissertation, we are going to address one specific type of inaccuracy in bivariate count data–heaping. Copula functions are used for the formulation of the bivariate distribution. Using copula functions for solving data inaccuracy problems is still a new area, which we are going to explore in this dissertation.

We also discuss the methods for variable selection when the explanatory variables are highly correlated. In particular, our method is …


Evaluation Of Goodness-Of-Fit Tests For The Cox Proportional Hazards Model With Time-Varying Covariates, Shanshan Hong Jan 2017

Evaluation Of Goodness-Of-Fit Tests For The Cox Proportional Hazards Model With Time-Varying Covariates, Shanshan Hong

Theses and Dissertations

The proportional hazards (PH) model, proposed by Cox (1972), is one of the most popular survival models for analyzing time-to-event data. To use the PH model properly, one must examine whether the data satisfy the PH assumption. An alternative model should be suggested if the PH assumption is invalid. The main purpose of this thesis is to examine the performance of five existing methods for assessing the PH assumption. Through extensive simulations, the powers of five different existing methods are compared; these methods include the likelihood ratio test, the Schoenfeld residuals test, the scaled Schoenfeld residuals test, Lin et al. …


Marginal Structural Cox Model For Survival Data With Treatment-Confounder Feedback, Yanan Zhang Jan 2017

Marginal Structural Cox Model For Survival Data With Treatment-Confounder Feedback, Yanan Zhang

Theses and Dissertations

In an observational longitudinal study, there can be time-varying exposure/treatment and time-varying confounders. When the confounders affect the exposure and prior exposure also has an impact on levels of confounders, there is treatment confounder feedback. To admit estimation of unbiased causal effects, these conditions need to be hold, exchangeability, positivity, consistency. The traditional method of conditioning on potential confounders does not meet these 3 conditions. Therefore, parameter estimates from traditional Cox model are biased casual effect estimates when the treatment confounder feedback exists. The marginal structural Cox model can be used to address this issue. By calculating and including inverse …


Comparing The Structural Components Variance Estimator And U-Statistics Variance Estimator When Assessing The Difference Between Correlated Aucs With Finite Samples, Anna L. Bosse Jan 2017

Comparing The Structural Components Variance Estimator And U-Statistics Variance Estimator When Assessing The Difference Between Correlated Aucs With Finite Samples, Anna L. Bosse

Theses and Dissertations

Introduction: The structural components variance estimator proposed by DeLong et al. (1988) is a popular approach used when comparing two correlated AUCs. However, this variance estimator is biased and could be problematic with small sample sizes.

Methods: A U-statistics based variance estimator approach is presented and compared with the structural components variance estimator through a large-scale simulation study under different finite-sample size configurations.

Results: The U-statistics variance estimator was unbiased for the true variance of the difference between correlated AUCs regardless of the sample size and had lower RMSE than the structural components variance estimator, providing better type 1 error …


Weighted Quantile Sum Regression For Analyzing Correlated Predictors Acting Through A Mediation Pathway On A Biological Outcome, Bhanu M. Evani Jan 2017

Weighted Quantile Sum Regression For Analyzing Correlated Predictors Acting Through A Mediation Pathway On A Biological Outcome, Bhanu M. Evani

Theses and Dissertations

Abstract

Weighted Quantile Sum Regression for Analyzing Correlated Predictors Acting Through a Mediation Pathway on a Biological Outcome

By

Bhanu M. Evani, Ph.D.

A thesis submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy at Virginia Commonwealth University.

Virginia Commonwealth University, 2017.

Major Director: Robert A. Perera, Asst. Professor, Department of Biostatistics

This work examines mediated effects of a set of correlated predictors using the recently developed Weighted Quantile Sum (WQS) regression method. Traditionally, mediation analysis has been conducted using the multiple regression method, first proposed by Baron and Kenny (1986), which has since …


Novel Methods For Analyzing Longitudinal Data With Measurement Error In The Time Variable, Caroline Munindi Mulatya Jun 2016

Novel Methods For Analyzing Longitudinal Data With Measurement Error In The Time Variable, Caroline Munindi Mulatya

Theses and Dissertations

In some longitudinal studies, the observed time points are often confounded with measurement error due to the sampling conditions, resulting into data with measurement error in the time variable. This type of data occurs mainly in observational studies when the onset of a longitudinal process is unknown or in clinical trials when individual visits do not take place as specified by the study protocol, but are often rounded to coincide with the study protocol. Methodological and inferential implications of error in time varying covariates for both linear and nonlinear models have been studied widely. In this dissertation, we shift attention …


On The Dynamics Of Boolean Gene Regulatory Networks With Stochasticity, Yuezhe Li Mar 2016

On The Dynamics Of Boolean Gene Regulatory Networks With Stochasticity, Yuezhe Li

Theses and Dissertations

Genes are responsible for producing proteins that are essential to the construction of complex biological systems. The mechanisms by which this production is regulated have long been the center of wide spread research efforts. Deterministic Boolean gene regulatory models have been a particularly effective avenue of research in this field. However these models fall short of accounting for variations in the gene functionality due to the uncertain internal or external environmental conditions. One of the recent attempts to overcome this weakness is by (Murrugarra, 2012), in which a probabilistic component is introduced as the fixed activation/degradation propensities at the cellular …


The Reflected-Shifted-Truncated-Gamma Distribution For Negatively Skewed Survival Data With Application To Pediatric Nephrotic Syndrome, Sophia D. Waymyers Jan 2016

The Reflected-Shifted-Truncated-Gamma Distribution For Negatively Skewed Survival Data With Application To Pediatric Nephrotic Syndrome, Sophia D. Waymyers

Theses and Dissertations

Negatively skewed survival data arise occasionally in public health fields and in statistical research. Standard distributions such as the exponential, generalized F, generalized gamma, Gompertz, log-logistic, lognormal, Rayleigh, and Weibull distributions are not always well suited to this data. The primary goal of this dissertation is to find a viable alternative for modeling negatively skewed survival data such as the time to first remission for pediatric patients with frequently relapsing or steroid dependent nephrotic syndrome.

We begin with a brief introduction of survival analysis and the nature of pediatric nephrotic syndrome. A meta-analysis on atopy and pediatric nephrotic syndrome using …


Finding The Cutpoint Of A Continuous Covariate In A Parametric Survival Analysis Model, Kabita Joshi Jan 2016

Finding The Cutpoint Of A Continuous Covariate In A Parametric Survival Analysis Model, Kabita Joshi

Theses and Dissertations

In many clinical studies, continuous variables such as age, blood pressure and cholesterol are measured and analyzed. Often clinicians prefer to categorize these continuous variables into different groups, such as low and high risk groups. The goal of this work is to find the cutpoint of a continuous variable where the transition occurs from low to high risk group. Different methods have been published in literature to find such a cutpoint. We extended the methods of Contal and O’Quigley (1999) which was based on the log-rank test and the methods of Klein and Wu (2004) which was based on the …


Sample Size Calculation For Ph Mixture Cure Model, Yihong Zhan Jan 2016

Sample Size Calculation For Ph Mixture Cure Model, Yihong Zhan

Theses and Dissertations

With the development of advanced medical technology, a significant proportion of patients can be cured of many chronic diseases. Because a substantial fraction of patients have censored information, the standard survival model, such as the proportional hazards (PH) model cannot capture the cured information of patients. Thus PH mixture cure model is developed to handle the survival data with potential cured information. A corresponding sample size formula based on log rank test has been proposed by Wang et al. (2012) and the probability of death in their formula is only contributed by the control arm. However, to calculate the sample …


Parametric Reversed Hazards Model For Left Censored Data With Application To Hiv, Farahnaz Islam Jan 2016

Parametric Reversed Hazards Model For Left Censored Data With Application To Hiv, Farahnaz Islam

Theses and Dissertations

Left censoring is generally a rare type of censoring in time-to-event data, however there are some fields such as HIV related studies where it commonly occurs. Currently, there is no clear recommendation in the literature on the optimal model and distribution to analyze left-censored data. Recommendations can help researchers apply more accurate models for this type of censoring. This study derives the Parametric Reversed Hazards (PRH) Model for a variety of distributions which may be appropriate for left censored data. The performance of these derived PRH models to analyze HIV viral load data are compared using extensive simulations and a …


Modeling Spatially Varying Effects Of Chemical Mixtures, Jenna Czarnota Jan 2016

Modeling Spatially Varying Effects Of Chemical Mixtures, Jenna Czarnota

Theses and Dissertations

Cancer incidence is associated with exposures to multiple environmental chemicals, and geographic variation in cancer rates suggests the importance of accommodating spatially varying effects in the analysis of environmental chemical mixtures and disease risk. Traditional regression methods are challenged by the complex correlation patterns inherent among co-occurring chemicals, and the applicability of geographically weighted regression models is limited in the setting of environmental chemical risk analysis. In comparison to traditional methods, weighted quantile sum (WQS) regression performs well in the identification of important environmental exposures, but is limited by the assumption that effects are fixed over space. We present an …


Spatio-Temporal Analysis Of The Occupational Fatal Victimization Of Law Enforcement Officers In The Us, Xueyi Xing Jan 2016

Spatio-Temporal Analysis Of The Occupational Fatal Victimization Of Law Enforcement Officers In The Us, Xueyi Xing

Theses and Dissertations

The models with constant coefficients of the covariates across space and time are commonly used in spatio-temporal analyses. However, the associations between risk factors and the outcome could have locally differential temporal trends in many cases. In this study, a Bayesian latent cluster modeling strategy is employed to identify potential spatial clusters in which locally specific sets of temporally varying coefficients of covariates are allowed. A state-level panel data of police officers occupational fatal victimization for the years 1979-2010 is used. To accommodate overdisperson and excess zeros, a negative binomial model and zero-inflated Poisson/negative binomial models are also utilized. A …