Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Medicine and Health Sciences (16)
- Applied Statistics (15)
- Life Sciences (14)
- Statistical Models (13)
- Applied Mathematics (7)
-
- Longitudinal Data Analysis and Time Series (7)
- Public Health (7)
- Statistical Methodology (7)
- Bioinformatics (6)
- Genetics and Genomics (6)
- Genetics (5)
- Microarrays (4)
- Multivariate Analysis (4)
- Data Science (3)
- Mathematics (3)
- Medical Sciences (3)
- Social and Behavioral Sciences (3)
- Categorical Data Analysis (2)
- Environmental Sciences (2)
- Genomics (2)
- Medical Specialties (2)
- Mental and Social Health (2)
- Natural Resources Management and Policy (2)
- Natural Resources and Conservation (2)
- Ordinary Differential Equations and Applied Dynamics (2)
- Other Genetics and Genomics (2)
- Plant Sciences (2)
- Institution
- Keyword
-
- Biostatistics (8)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (6)
- Bayesian (4)
- Clinical trial (4)
- Genetics (4)
-
- Genomics (4)
- Longitudinal data (4)
- Physical Sciences and Mathematics, Statistics and Probability, Biostatistics (4)
- Survival Analysis (4)
- Classification (3)
- Cluster analysis (3)
- Cox proportional hazards model (3)
- Environment (3)
- Gene expression (3)
- Power (3)
- Spatial clustering (3)
- Variable selection (3)
- Adaptive design (2)
- Algorithm (2)
- Bayesian inference (2)
- Bioinformatics (2)
- Chemicals (2)
- Clustering (2)
- Copy number variation (2)
- Correlated data (2)
- Count Data (2)
- Covariance (2)
- DNA Methylation (2)
- Data (2)
- Data analysis (2)
Articles 61 - 90 of 162
Full-Text Articles in Biostatistics
Randomization Analysis Driven Software, Steph-Yves Louis
Randomization Analysis Driven Software, Steph-Yves Louis
Theses and Dissertations
The application of a method of randomization for a clinical trial frequently summarizes to using Simple Randomization. Even though the latter method provides favorable characteristics, if the collected sample is not large enough, it still presents the highest chance of imbalance both marginally in the treatment groups and locally in terms of the covariates. Methods of Permuted Block Randomization, Urn Randomization, Stratified Permuted Block Randomization, and Minimization represent popular alternative methods that one should consider depending on the goal of the study. A comparison of the previously mentioned methods is carried to evaluate their performance with samples that are not …
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Theses and Dissertations
In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …
Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace
Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace
Theses and Dissertations
Response-Adaptive (RA) designs are used to adaptively allocate patients in clinical trials. These methods have been generalized to include Covariate-Adjusted Response-Adaptive (CARA) designs, which adjust treatment assignments for a set of covariates while maintaining features of the RA designs. Challenges may arise in multi-center trials if differential treatment responses and/or effects among sites exist. We propose Site-Adjusted Response-Adaptive (SARA) approaches to account for inter-center variability in treatment response and/or effectiveness, including either a fixed site effect or both random site and treatment-by-site interaction effects to calculate conditional probabilities. These success probabilities are used to update assignment probabilities for allocating patients …
Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer
Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer
Theses and Dissertations
As researchers increasingly use web-based surveys, the ease of dropping out in the online setting is a growing issue in ensuring data quality. One theory is that dropout or attrition occurs in phases that can be generalized to phases of high dropout and phases of stable use. In order to detect these phases, several methods are explored. First, existing methods and user-specified thresholds are applied to survey data where significant changes in the dropout rate between two questions is interpreted as the start or end of a high dropout phase. Next, survey dropout is considered as a time-to-event outcome and …
Assessing The Impact Of Incorporating Residential Histories Into The Spatial Analysis Of Cancer Risk, Anny-Claude Joseph
Assessing The Impact Of Incorporating Residential Histories Into The Spatial Analysis Of Cancer Risk, Anny-Claude Joseph
Theses and Dissertations
In many spatial epidemiologic studies, investigators use residential location at diagnosis as a surrogate for unknown environmental exposures or as a geographic basis for assigning measured exposures. Inherently, they make assumptions about the timing and location of pertinent exposures which may prove problematic when studying long latency diseases such as cancer.
In this work we explored how the association between environmental exposures and disease risk for long-latency health outcomes like cancer is affected by residential mobility. We used simulation studies conditioned on real data to evaluate the extent to which the commonly held assumption of no residential mobility 1) affected …
Methods For Joint Normalization And Comparison Of Hi-C Data, John C. Stansfield
Methods For Joint Normalization And Comparison Of Hi-C Data, John C. Stansfield
Theses and Dissertations
The development of chromatin conformation capture technology has opened new avenues of study into the 3D structure and function of the genome. Chromatin structure is known to influence gene regulation, and differences in structure are now emerging as a mechanism of regulation between, e.g., cell differentiation and disease vs. normal states. Hi-C sequencing technology now provides a way to study the 3D interactions of the chromatin over the whole genome. However, like all sequencing technologies, Hi-C suffers from several forms of bias stemming from both the technology and the DNA sequence itself. Several normalization methods have been developed for normalizing …
Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna
Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna
Theses and Dissertations
Widely effective treatment for alcohol use disorder is not yet available, because the exact biological mechanisms that underlie this disorder are not completely understood. One way to gain a better understanding of these mechanisms is to examine the genetic frameworks that contribute to the risk for developing this disorder. This dissertation examines genetic association data in combination with gene expression networks in the brain to identify functional groups of genes associated with alcohol consumption and dependence.
The first study took advantage of the behavioral complexity of human samples, and experimental capabilities provided by mouse models, by co-analyzing gene expression networks …
Spectral Methods For The Detection And Characterization Of Topologically Associated Domains, Kellen Garrison Cresswell
Spectral Methods For The Detection And Characterization Of Topologically Associated Domains, Kellen Garrison Cresswell
Theses and Dissertations
The three-dimensional (3D) structure of the genome plays a crucial role in gene expression regulation. Chromatin conformation capture technologies (Hi-C) have revealed that the genome is organized in a hierarchy of topologically associated domains (TADs), sub-TADs, and chromatin loops which is relatively stable across cell-lines and even across species. These TADs dynamically reorganize during development of disease, and exhibit cell- and conditionspecific differences. Identifying such hierarchical structures and how they change between conditions is a critical step in understanding genome regulation and disease development. Despite their importance, there are relatively few tools for identification of TADs and even fewer for …
Clustering Biological Data With Self-Adjusting High-Dimensional Sieve, Josselyn Gonzalez
Clustering Biological Data With Self-Adjusting High-Dimensional Sieve, Josselyn Gonzalez
Theses and Dissertations
Data classification as a preprocessing technique is a crucial step in the analysis and understanding of numerical data. Cluster analysis, in particular, provides insight into the inherent patterns found in data which makes the interpretation of any follow-up analyses more meaningful. A clustering algorithm groups together data points according to a predefined similarity criterion. This allows the data set to be broken up into segments which, in turn, gives way for a more targeted statistical analysis. Cluster analysis has applications in numerous fields of study and, as a result, countless algorithms have been developed. However, the quantity of options makes …
Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry
Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry
Theses and Dissertations
The Brisbane Longitudinal Twin Study (BLTS) was being conducted in Australia and was funded by the US National Institute on Drug Abuse (NIDA). Adolescent twins were sampled as a part of this study and surveyed about their substance use as part of the Pathways to Cannabis Use, Abuse and Dependence project. The methods developed in this dissertation were designed for the purpose of analyzing a subset of the Pathways data that includes demographics, cannabis use metrics, personality measures, and imputed genotypes (SNPs) for 493 complete twin pairs (986 subjects.) The primary goal was to determine what combination of SNPs and …
Adjusting For Mis-Reporting In Count Data, Gelareh Rahimighazikalayeh
Adjusting For Mis-Reporting In Count Data, Gelareh Rahimighazikalayeh
Theses and Dissertations
Any counting system is prone to recording errors including underreporting and overreporting. Ignoring the misreporting pattern in count data can give rise to bias in the estimation of model parameters. Accordingly, Poisson, negative binomial and generalized Poisson regression have been expanded in some instances to capture reporting biases. However, to our knowledge, no program has been developed to allow users to apply all of these models when needed. In the first part of the dissertation, we review the available models for underreported counts and develop a Stata command to estimate Poisson, negative binomial and generalized Poisson regression models for underreported …
Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard
Comparison Of The Performance Of Simple Linear Regression And Quantile Regression With Non-Normal Data: A Simulation Study, Marjorie Howard
Theses and Dissertations
Linear regression is a widely used method for analysis that is well understood across a wide variety of disciplines. In order to use linear regression, a number of assumptions must be met. These assumptions, specifically normality and homoscedasticity of the error distribution can at best be met only approximately with real data. Quantile regression requires fewer assumptions, which offers a potential advantage over linear regression. In this simulation study, we compare the performance of linear (least squares) regression to quantile regression when these assumptions are violated, in order to investigate under what conditions quantile regression becomes the more advantageous method …
Estimation Procedures For Complex Survival Models And Their Applications In Epidemiology Studies, Jie Zhou
Estimation Procedures For Complex Survival Models And Their Applications In Epidemiology Studies, Jie Zhou
Theses and Dissertations
In this dissertation, we aim to address three important questions in practice, which can be solved through complex survival models. The first project focuses on studying the longitudinal fitness effect on cardiovascular disease (CVD) mortality. In the second project, we study the disease-death relation between CVD and all-cause mortality and evaluate important covariate effects on the disease or death transitions. In the third project, we compare antiretroviral treatment (ART) for HIV patients and consider both treatment effect and side effect of the drugs. The first two projects are motivated by the Aerobics Center Longitudinal Study (ACLS) datasets and the third …
Examining The Confirmatory Tetrad Analysis (Cta) As A Solution Of The Inadequacy Of Traditional Structural Equation Modeling (Sem) Fit Indices, Hangcheng Liu
Theses and Dissertations
Structural Equation Modeling (SEM) is a framework of statistical methods that allows us to represent complex relationships between variables. SEM is widely used in economics, genetics and the behavioral sciences (e.g. psychology, psychobiology, sociology and medicine). Model complexity is defined as a model’s ability to fit different data patterns and it plays an important role in model selection when applying SEM. As in linear regression, the number of free model parameters is typically used in traditional SEM model fit indices as a measure of the model complexity. However, only using number of free model parameters to indicate SEM model complexity …
Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang
Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang
Theses and Dissertations
Modern big data often emerge as tensors. Standard statistical methods are inadequate to deal with datasets of large volume, high dimensionality, and complex structure. Therefore, it is important to develop algorithms such as low-rank tensor decomposition for data compression, dimensionality reduction, and approximation.
With the advancement in technology, high-dimensional images are becoming ubiquitous in the medical field. In lung radiation therapy, the respiratory motion of the lung introduces variabilities during treatment as the tumor inside the lung is moving, which brings challenges to the precise delivery of radiation to the tumor. Several approaches to quantifying this uncertainty propose using a …
The Generalized Monotone Incremental Forward Stagewise Method For Modeling Longitudinal, Clustered, And Overdispersed Count Data: Application Predicting Nuclear Bud And Micronuclei Frequencies, Rebecca Lehman
Theses and Dissertations
With the influx of high-dimensional data there is an immediate need for statistical methods that are able to handle situations when the number of predictors greatly exceeds the number of samples. One such area of growth is in examining how environmental exposures to toxins impact the body long term. The cytokinesis-block micronucleus assay can measure the genotoxic effect of exposure as a count outcome. To investigate potential biomarkers, high-throughput assays that assess gene expression and methylation have been developed. It is of interest to identify biomarkers or molecular features that are associated with elevated micronuclei (MN) or nuclear bud (Nbud) …
Longitudinal And Geographical Modeling Of Circular Data With An Application To Sudden Infant Death Syndrome, Xinyan Cai
Longitudinal And Geographical Modeling Of Circular Data With An Application To Sudden Infant Death Syndrome, Xinyan Cai
Theses and Dissertations
The aim of this thesis is to study seasonality of death in U.S. infants who died from SIDS. We also propose to investigate secular trends and geographical patterns of seasonal patterns of mortality. The application of circular statistics is used to describe the seasonality of the month of death in infants who died from SIDS in 1990, 2000 and 2010. The secular trends of seasonal patterns of SIDS mortality are investigated using a circular linear regression model after adjusting for potential confounders. The geographical variation in seasonal patterns of SIDS mortality is explored from the U.S. map and quantified by …
Statistical Methods For Multivariate And Correlated Data, Xinling Xu
Statistical Methods For Multivariate And Correlated Data, Xinling Xu
Theses and Dissertations
A commonly encountered data type in real life is count data, especially in selfreported behavioral studies. One issue of the self-reported count data is the inaccuracy. In the first part of the dissertation, we are going to address one specific type of inaccuracy in bivariate count data–heaping. Copula functions are used for the formulation of the bivariate distribution. Using copula functions for solving data inaccuracy problems is still a new area, which we are going to explore in this dissertation.
We also discuss the methods for variable selection when the explanatory variables are highly correlated. In particular, our method is …
Evaluation Of Goodness-Of-Fit Tests For The Cox Proportional Hazards Model With Time-Varying Covariates, Shanshan Hong
Evaluation Of Goodness-Of-Fit Tests For The Cox Proportional Hazards Model With Time-Varying Covariates, Shanshan Hong
Theses and Dissertations
The proportional hazards (PH) model, proposed by Cox (1972), is one of the most popular survival models for analyzing time-to-event data. To use the PH model properly, one must examine whether the data satisfy the PH assumption. An alternative model should be suggested if the PH assumption is invalid. The main purpose of this thesis is to examine the performance of five existing methods for assessing the PH assumption. Through extensive simulations, the powers of five different existing methods are compared; these methods include the likelihood ratio test, the Schoenfeld residuals test, the scaled Schoenfeld residuals test, Lin et al. …
Marginal Structural Cox Model For Survival Data With Treatment-Confounder Feedback, Yanan Zhang
Marginal Structural Cox Model For Survival Data With Treatment-Confounder Feedback, Yanan Zhang
Theses and Dissertations
In an observational longitudinal study, there can be time-varying exposure/treatment and time-varying confounders. When the confounders affect the exposure and prior exposure also has an impact on levels of confounders, there is treatment confounder feedback. To admit estimation of unbiased causal effects, these conditions need to be hold, exchangeability, positivity, consistency. The traditional method of conditioning on potential confounders does not meet these 3 conditions. Therefore, parameter estimates from traditional Cox model are biased casual effect estimates when the treatment confounder feedback exists. The marginal structural Cox model can be used to address this issue. By calculating and including inverse …
Comparing The Structural Components Variance Estimator And U-Statistics Variance Estimator When Assessing The Difference Between Correlated Aucs With Finite Samples, Anna L. Bosse
Theses and Dissertations
Introduction: The structural components variance estimator proposed by DeLong et al. (1988) is a popular approach used when comparing two correlated AUCs. However, this variance estimator is biased and could be problematic with small sample sizes.
Methods: A U-statistics based variance estimator approach is presented and compared with the structural components variance estimator through a large-scale simulation study under different finite-sample size configurations.
Results: The U-statistics variance estimator was unbiased for the true variance of the difference between correlated AUCs regardless of the sample size and had lower RMSE than the structural components variance estimator, providing better type 1 error …
Weighted Quantile Sum Regression For Analyzing Correlated Predictors Acting Through A Mediation Pathway On A Biological Outcome, Bhanu M. Evani
Weighted Quantile Sum Regression For Analyzing Correlated Predictors Acting Through A Mediation Pathway On A Biological Outcome, Bhanu M. Evani
Theses and Dissertations
Abstract
Weighted Quantile Sum Regression for Analyzing Correlated Predictors Acting Through a Mediation Pathway on a Biological Outcome
By
Bhanu M. Evani, Ph.D.
A thesis submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy at Virginia Commonwealth University.
Virginia Commonwealth University, 2017.
Major Director: Robert A. Perera, Asst. Professor, Department of Biostatistics
This work examines mediated effects of a set of correlated predictors using the recently developed Weighted Quantile Sum (WQS) regression method. Traditionally, mediation analysis has been conducted using the multiple regression method, first proposed by Baron and Kenny (1986), which has since …
Novel Methods For Analyzing Longitudinal Data With Measurement Error In The Time Variable, Caroline Munindi Mulatya
Novel Methods For Analyzing Longitudinal Data With Measurement Error In The Time Variable, Caroline Munindi Mulatya
Theses and Dissertations
In some longitudinal studies, the observed time points are often confounded with measurement error due to the sampling conditions, resulting into data with measurement error in the time variable. This type of data occurs mainly in observational studies when the onset of a longitudinal process is unknown or in clinical trials when individual visits do not take place as specified by the study protocol, but are often rounded to coincide with the study protocol. Methodological and inferential implications of error in time varying covariates for both linear and nonlinear models have been studied widely. In this dissertation, we shift attention …
On The Dynamics Of Boolean Gene Regulatory Networks With Stochasticity, Yuezhe Li
On The Dynamics Of Boolean Gene Regulatory Networks With Stochasticity, Yuezhe Li
Theses and Dissertations
Genes are responsible for producing proteins that are essential to the construction of complex biological systems. The mechanisms by which this production is regulated have long been the center of wide spread research efforts. Deterministic Boolean gene regulatory models have been a particularly effective avenue of research in this field. However these models fall short of accounting for variations in the gene functionality due to the uncertain internal or external environmental conditions. One of the recent attempts to overcome this weakness is by (Murrugarra, 2012), in which a probabilistic component is introduced as the fixed activation/degradation propensities at the cellular …
The Reflected-Shifted-Truncated-Gamma Distribution For Negatively Skewed Survival Data With Application To Pediatric Nephrotic Syndrome, Sophia D. Waymyers
The Reflected-Shifted-Truncated-Gamma Distribution For Negatively Skewed Survival Data With Application To Pediatric Nephrotic Syndrome, Sophia D. Waymyers
Theses and Dissertations
Negatively skewed survival data arise occasionally in public health fields and in statistical research. Standard distributions such as the exponential, generalized F, generalized gamma, Gompertz, log-logistic, lognormal, Rayleigh, and Weibull distributions are not always well suited to this data. The primary goal of this dissertation is to find a viable alternative for modeling negatively skewed survival data such as the time to first remission for pediatric patients with frequently relapsing or steroid dependent nephrotic syndrome.
We begin with a brief introduction of survival analysis and the nature of pediatric nephrotic syndrome. A meta-analysis on atopy and pediatric nephrotic syndrome using …
Finding The Cutpoint Of A Continuous Covariate In A Parametric Survival Analysis Model, Kabita Joshi
Finding The Cutpoint Of A Continuous Covariate In A Parametric Survival Analysis Model, Kabita Joshi
Theses and Dissertations
In many clinical studies, continuous variables such as age, blood pressure and cholesterol are measured and analyzed. Often clinicians prefer to categorize these continuous variables into different groups, such as low and high risk groups. The goal of this work is to find the cutpoint of a continuous variable where the transition occurs from low to high risk group. Different methods have been published in literature to find such a cutpoint. We extended the methods of Contal and O’Quigley (1999) which was based on the log-rank test and the methods of Klein and Wu (2004) which was based on the …
Sample Size Calculation For Ph Mixture Cure Model, Yihong Zhan
Sample Size Calculation For Ph Mixture Cure Model, Yihong Zhan
Theses and Dissertations
With the development of advanced medical technology, a significant proportion of patients can be cured of many chronic diseases. Because a substantial fraction of patients have censored information, the standard survival model, such as the proportional hazards (PH) model cannot capture the cured information of patients. Thus PH mixture cure model is developed to handle the survival data with potential cured information. A corresponding sample size formula based on log rank test has been proposed by Wang et al. (2012) and the probability of death in their formula is only contributed by the control arm. However, to calculate the sample …
Parametric Reversed Hazards Model For Left Censored Data With Application To Hiv, Farahnaz Islam
Parametric Reversed Hazards Model For Left Censored Data With Application To Hiv, Farahnaz Islam
Theses and Dissertations
Left censoring is generally a rare type of censoring in time-to-event data, however there are some fields such as HIV related studies where it commonly occurs. Currently, there is no clear recommendation in the literature on the optimal model and distribution to analyze left-censored data. Recommendations can help researchers apply more accurate models for this type of censoring. This study derives the Parametric Reversed Hazards (PRH) Model for a variety of distributions which may be appropriate for left censored data. The performance of these derived PRH models to analyze HIV viral load data are compared using extensive simulations and a …
Modeling Spatially Varying Effects Of Chemical Mixtures, Jenna Czarnota
Modeling Spatially Varying Effects Of Chemical Mixtures, Jenna Czarnota
Theses and Dissertations
Cancer incidence is associated with exposures to multiple environmental chemicals, and geographic variation in cancer rates suggests the importance of accommodating spatially varying effects in the analysis of environmental chemical mixtures and disease risk. Traditional regression methods are challenged by the complex correlation patterns inherent among co-occurring chemicals, and the applicability of geographically weighted regression models is limited in the setting of environmental chemical risk analysis. In comparison to traditional methods, weighted quantile sum (WQS) regression performs well in the identification of important environmental exposures, but is limited by the assumption that effects are fixed over space. We present an …
Spatio-Temporal Analysis Of The Occupational Fatal Victimization Of Law Enforcement Officers In The Us, Xueyi Xing
Spatio-Temporal Analysis Of The Occupational Fatal Victimization Of Law Enforcement Officers In The Us, Xueyi Xing
Theses and Dissertations
The models with constant coefficients of the covariates across space and time are commonly used in spatio-temporal analyses. However, the associations between risk factors and the outcome could have locally differential temporal trends in many cases. In this study, a Bayesian latent cluster modeling strategy is employed to identify potential spatial clusters in which locally specific sets of temporally varying coefficients of covariates are allowed. A state-level panel data of police officers occupational fatal victimization for the years 1979-2010 is used. To accommodate overdisperson and excess zeros, a negative binomial model and zero-inflated Poisson/negative binomial models are also utilized. A …