Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Life Sciences (18)
- Medicine and Health Sciences (17)
- Applied Statistics (13)
- Statistical Models (13)
- Applied Mathematics (9)
-
- Genetics and Genomics (9)
- Statistical Methodology (8)
- Longitudinal Data Analysis and Time Series (7)
- Bioinformatics (6)
- Genetics (5)
- Microarrays (5)
- Multivariate Analysis (5)
- Public Health (5)
- Genomics (4)
- Probability (4)
- Social and Behavioral Sciences (4)
- Data Science (3)
- Diseases (3)
- Ecology and Evolutionary Biology (3)
- Medical Sciences (3)
- Ordinary Differential Equations and Applied Dynamics (3)
- Population Biology (3)
- Psychology (3)
- Survival Analysis (3)
- Categorical Data Analysis (2)
- Computational Biology (2)
- Computer Sciences (2)
- Keyword
-
- Biostatistics (6)
- Clinical trial (4)
- Genomics (4)
- Environment (3)
- Gene expression (3)
-
- Longitudinal data (3)
- Power (3)
- Adaptive design (2)
- Algorithm (2)
- Bayesian (2)
- Chemicals (2)
- Classification (2)
- Cluster analysis (2)
- Covariance (2)
- DNA Methylation (2)
- Dropout (2)
- Epidemiology (2)
- Genetics (2)
- Hi-C (2)
- Meta-analysis (2)
- Microarray (2)
- Missing data (2)
- Mixed models (2)
- Model selection (2)
- NHANES (2)
- Neuroscience (2)
- Principal Components Analysis (2)
- ROC curves (2)
- Structural Equation Modeling (2)
- Survival analysis (2)
- Publication Year
- Publication
- Publication Type
Articles 31 - 60 of 108
Full-Text Articles in Biostatistics
Assessing The Impact Of Incorporating Residential Histories Into The Spatial Analysis Of Cancer Risk, Anny-Claude Joseph
Assessing The Impact Of Incorporating Residential Histories Into The Spatial Analysis Of Cancer Risk, Anny-Claude Joseph
Theses and Dissertations
In many spatial epidemiologic studies, investigators use residential location at diagnosis as a surrogate for unknown environmental exposures or as a geographic basis for assigning measured exposures. Inherently, they make assumptions about the timing and location of pertinent exposures which may prove problematic when studying long latency diseases such as cancer.
In this work we explored how the association between environmental exposures and disease risk for long-latency health outcomes like cancer is affected by residential mobility. We used simulation studies conditioned on real data to evaluate the extent to which the commonly held assumption of no residential mobility 1) affected …
Methods For Joint Normalization And Comparison Of Hi-C Data, John C. Stansfield
Methods For Joint Normalization And Comparison Of Hi-C Data, John C. Stansfield
Theses and Dissertations
The development of chromatin conformation capture technology has opened new avenues of study into the 3D structure and function of the genome. Chromatin structure is known to influence gene regulation, and differences in structure are now emerging as a mechanism of regulation between, e.g., cell differentiation and disease vs. normal states. Hi-C sequencing technology now provides a way to study the 3D interactions of the chromatin over the whole genome. However, like all sequencing technologies, Hi-C suffers from several forms of bias stemming from both the technology and the DNA sequence itself. Several normalization methods have been developed for normalizing …
Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna
Genome-Wide Systems Genetics Of Alcohol Consumption And Dependence, Kristin Mignogna
Theses and Dissertations
Widely effective treatment for alcohol use disorder is not yet available, because the exact biological mechanisms that underlie this disorder are not completely understood. One way to gain a better understanding of these mechanisms is to examine the genetic frameworks that contribute to the risk for developing this disorder. This dissertation examines genetic association data in combination with gene expression networks in the brain to identify functional groups of genes associated with alcohol consumption and dependence.
The first study took advantage of the behavioral complexity of human samples, and experimental capabilities provided by mouse models, by co-analyzing gene expression networks …
Spectral Methods For The Detection And Characterization Of Topologically Associated Domains, Kellen Garrison Cresswell
Spectral Methods For The Detection And Characterization Of Topologically Associated Domains, Kellen Garrison Cresswell
Theses and Dissertations
The three-dimensional (3D) structure of the genome plays a crucial role in gene expression regulation. Chromatin conformation capture technologies (Hi-C) have revealed that the genome is organized in a hierarchy of topologically associated domains (TADs), sub-TADs, and chromatin loops which is relatively stable across cell-lines and even across species. These TADs dynamically reorganize during development of disease, and exhibit cell- and conditionspecific differences. Identifying such hierarchical structures and how they change between conditions is a critical step in understanding genome regulation and disease development. Despite their importance, there are relatively few tools for identification of TADs and even fewer for …
Quantitative Electroencephalography For Detecting Concussions, Sara Krehbiel, Kathy Hoke, Joanna Wares
Quantitative Electroencephalography For Detecting Concussions, Sara Krehbiel, Kathy Hoke, Joanna Wares
Biology and Medicine Through Mathematics Conference
No abstract provided.
Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry
Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry
Theses and Dissertations
The Brisbane Longitudinal Twin Study (BLTS) was being conducted in Australia and was funded by the US National Institute on Drug Abuse (NIDA). Adolescent twins were sampled as a part of this study and surveyed about their substance use as part of the Pathways to Cannabis Use, Abuse and Dependence project. The methods developed in this dissertation were designed for the purpose of analyzing a subset of the Pathways data that includes demographics, cannabis use metrics, personality measures, and imputed genotypes (SNPs) for 493 complete twin pairs (986 subjects.) The primary goal was to determine what combination of SNPs and …
Examining The Confirmatory Tetrad Analysis (Cta) As A Solution Of The Inadequacy Of Traditional Structural Equation Modeling (Sem) Fit Indices, Hangcheng Liu
Theses and Dissertations
Structural Equation Modeling (SEM) is a framework of statistical methods that allows us to represent complex relationships between variables. SEM is widely used in economics, genetics and the behavioral sciences (e.g. psychology, psychobiology, sociology and medicine). Model complexity is defined as a model’s ability to fit different data patterns and it plays an important role in model selection when applying SEM. As in linear regression, the number of free model parameters is typically used in traditional SEM model fit indices as a measure of the model complexity. However, only using number of free model parameters to indicate SEM model complexity …
Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang
Estimating The Respiratory Lung Motion Model Using Tensor Decomposition On Displacement Vector Field, Kingston Kang
Theses and Dissertations
Modern big data often emerge as tensors. Standard statistical methods are inadequate to deal with datasets of large volume, high dimensionality, and complex structure. Therefore, it is important to develop algorithms such as low-rank tensor decomposition for data compression, dimensionality reduction, and approximation.
With the advancement in technology, high-dimensional images are becoming ubiquitous in the medical field. In lung radiation therapy, the respiratory motion of the lung introduces variabilities during treatment as the tumor inside the lung is moving, which brings challenges to the precise delivery of radiation to the tumor. Several approaches to quantifying this uncertainty propose using a …
The Generalized Monotone Incremental Forward Stagewise Method For Modeling Longitudinal, Clustered, And Overdispersed Count Data: Application Predicting Nuclear Bud And Micronuclei Frequencies, Rebecca Lehman
Theses and Dissertations
With the influx of high-dimensional data there is an immediate need for statistical methods that are able to handle situations when the number of predictors greatly exceeds the number of samples. One such area of growth is in examining how environmental exposures to toxins impact the body long term. The cytokinesis-block micronucleus assay can measure the genotoxic effect of exposure as a count outcome. To investigate potential biomarkers, high-throughput assays that assess gene expression and methylation have been developed. It is of interest to identify biomarkers or molecular features that are associated with elevated micronuclei (MN) or nuclear bud (Nbud) …
Comparing The Structural Components Variance Estimator And U-Statistics Variance Estimator When Assessing The Difference Between Correlated Aucs With Finite Samples, Anna L. Bosse
Theses and Dissertations
Introduction: The structural components variance estimator proposed by DeLong et al. (1988) is a popular approach used when comparing two correlated AUCs. However, this variance estimator is biased and could be problematic with small sample sizes.
Methods: A U-statistics based variance estimator approach is presented and compared with the structural components variance estimator through a large-scale simulation study under different finite-sample size configurations.
Results: The U-statistics variance estimator was unbiased for the true variance of the difference between correlated AUCs regardless of the sample size and had lower RMSE than the structural components variance estimator, providing better type 1 error …
Weighted Quantile Sum Regression For Analyzing Correlated Predictors Acting Through A Mediation Pathway On A Biological Outcome, Bhanu M. Evani
Weighted Quantile Sum Regression For Analyzing Correlated Predictors Acting Through A Mediation Pathway On A Biological Outcome, Bhanu M. Evani
Theses and Dissertations
Abstract
Weighted Quantile Sum Regression for Analyzing Correlated Predictors Acting Through a Mediation Pathway on a Biological Outcome
By
Bhanu M. Evani, Ph.D.
A thesis submitted in partial fulfillment of the requirements for the degree of Doctor of Philosophy at Virginia Commonwealth University.
Virginia Commonwealth University, 2017.
Major Director: Robert A. Perera, Asst. Professor, Department of Biostatistics
This work examines mediated effects of a set of correlated predictors using the recently developed Weighted Quantile Sum (WQS) regression method. Traditionally, mediation analysis has been conducted using the multiple regression method, first proposed by Baron and Kenny (1986), which has since …
Homeolog Specific Expression Bias, Ronald D. Smith
Homeolog Specific Expression Bias, Ronald D. Smith
Biology and Medicine Through Mathematics Conference
No abstract provided.
Heterogeneous Responses To Viral Infection: Insights From Mathematical Modeling Of Yellow Fever Vaccine, James R. Moore
Heterogeneous Responses To Viral Infection: Insights From Mathematical Modeling Of Yellow Fever Vaccine, James R. Moore
Biology and Medicine Through Mathematics Conference
No abstract provided.
Finding The Cutpoint Of A Continuous Covariate In A Parametric Survival Analysis Model, Kabita Joshi
Finding The Cutpoint Of A Continuous Covariate In A Parametric Survival Analysis Model, Kabita Joshi
Theses and Dissertations
In many clinical studies, continuous variables such as age, blood pressure and cholesterol are measured and analyzed. Often clinicians prefer to categorize these continuous variables into different groups, such as low and high risk groups. The goal of this work is to find the cutpoint of a continuous variable where the transition occurs from low to high risk group. Different methods have been published in literature to find such a cutpoint. We extended the methods of Contal and O’Quigley (1999) which was based on the log-rank test and the methods of Klein and Wu (2004) which was based on the …
Modeling Spatially Varying Effects Of Chemical Mixtures, Jenna Czarnota
Modeling Spatially Varying Effects Of Chemical Mixtures, Jenna Czarnota
Theses and Dissertations
Cancer incidence is associated with exposures to multiple environmental chemicals, and geographic variation in cancer rates suggests the importance of accommodating spatially varying effects in the analysis of environmental chemical mixtures and disease risk. Traditional regression methods are challenged by the complex correlation patterns inherent among co-occurring chemicals, and the applicability of geographically weighted regression models is limited in the setting of environmental chemical risk analysis. In comparison to traditional methods, weighted quantile sum (WQS) regression performs well in the identification of important environmental exposures, but is limited by the assumption that effects are fixed over space. We present an …
A Weighted Gene Co-Expression Network Analysis For Streptococcus Sanguinis Microarray Experiments, Erik C. Dvergsten
A Weighted Gene Co-Expression Network Analysis For Streptococcus Sanguinis Microarray Experiments, Erik C. Dvergsten
Theses and Dissertations
Streptococcus sanguinis is a gram-positive, non-motile bacterium native to human mouths. It is the primary cause of endocarditis and is also responsible for tooth decay. Two-component systems (TCSs) are commonly found in bacteria. In response to environmental signals, TCSs may regulate the expression of virulence factor genes.
Gene co-expression networks are exploratory tools used to analyze system-level gene functionality. A gene co-expression network consists of gene expression profiles represented as nodes and gene connections, which occur if two genes are significantly co-expressed. An adjacency function transforms the similarity matrix containing co-expression similarities into the adjacency matrix containing connection strengths. Gene …
Selecting Spatial Scale Of Area-Level Covariates In Regression Models, Lauren Grant
Selecting Spatial Scale Of Area-Level Covariates In Regression Models, Lauren Grant
Theses and Dissertations
Studies have found that the level of association between an area-level covariate and an outcome can vary depending on the spatial scale (SS) of a particular covariate. However, covariates used in regression models are customarily modeled at the same spatial unit. In this dissertation, we developed four SS model selection algorithms that select the best spatial scale for each area-level covariate. The SS forward stepwise, SS incremental forward stagewise, SS least angle regression (LARS), and SS lasso algorithms allow for the selection of different area-level covariates at different spatial scales, while constraining each covariate to enter at most one spatial …
Meta-Analytic Estimation Techniques For Non-Convergent Repeated-Measure Clustered Data, Aobo Wang
Meta-Analytic Estimation Techniques For Non-Convergent Repeated-Measure Clustered Data, Aobo Wang
Theses and Dissertations
Clustered data often feature nested structures and repeated measures. If coupled with binary outcomes and large samples (>10,000), this complexity can lead to non-convergence problems for the desired model especially if random effects are used to account for the clustering. One way to bypass the convergence problem is to split the dataset into small enough sub-samples for which the desired model convergences, and then recombine results from those sub-samples through meta-analysis. We consider two ways to generate sub-samples: the K independent samples approach where the data are split into k mutually-exclusive sub-samples, and the cluster-based approach where naturally existing …
Proof-Of-Concept Of Environmental Dna Tools For Atlantic Sturgeon Management, Jameson Hinkle
Proof-Of-Concept Of Environmental Dna Tools For Atlantic Sturgeon Management, Jameson Hinkle
Theses and Dissertations
Abstract
The Atlantic Sturgeon (Acipenser oxyrinchus oxyrinchus, Mitchell) is an anadromous species that spawns in tidal freshwater rivers from Canada to Florida. Overfishing, river sedimentation and alteration of the river bottom have decreased Atlantic Sturgeon populations, and NOAA lists the species as endangered. Ecologists sometimes find it difficult to locate individuals of a species that is rare, endangered or invasive. The need for methods less invasive that can create more resolution of cryptic species presence is necessary. Environmental DNA (eDNA) is a non-invasive means of detecting rare, endangered, or invasive species by isolating nuclear or mitochondrial DNA (mtDNA) from the …
High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski
High-Throughput Data Analysis: Application To Micronuclei Frequency And T-Cell Receptor Sequencing, Mateusz Makowski
Theses and Dissertations
The advent of high-throughput sequencing has brought about the creation of an unprecedented amount of research data. Analytical methodology has not been able to keep pace with the plethora of data being produced. Two assays, ImmunoSEQ and the cytokinesisblock micronucleus (CBMN), that both produce count data and have few methods available to analyze them are considered.
ImmunoSEQ is a sequencing assay that measures the beta T-cell receptor (TCR) repertoire. The ImmunoSEQ assay was used to describe the TCR repertoires of patients that have undergone hematopoietic stem cell transplantation (HSCT). Several different methods for spectratype analysis were extended to the TCR …
Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu
Comparing Welch's Anova, A Kruskal-Wallis Test And Traditional Anova In Case Of Heterogeneity Of Variance, Hangcheng Liu
Theses and Dissertations
Analysis of variance (ANOVA) is a robust test against the normality assumption, but it may be inappropriate when the assumption of homogeneity of variance has been violated. Welch ANOVA and the Kruskal-Wallis test (a non-parametric method) can be applicable for this case. In this study we compare the three methods in empirical type I error rate and power, when heterogeneity of variance occurs and find out which method is the most suitable with which cases including balanced/unbalanced, small/large sample size, and/or with normal/non-normal distributions.
Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe
Meta-Analysis Of Gene Expression Studies, Umaporn Siangphoe
Theses and Dissertations
Combining effect sizes from individual studies using random-effects models are commonly applied in high-dimensional gene expression data. However, unknown study heterogeneity can arise from inconsistency of sample qualities and experimental conditions. High heterogeneity of effect sizes can reduce statistical power of the models. We proposed two new methods for random effects estimation and measurements for model variation and strength of the study heterogeneity. We then developed a statistical technique to test for significance of random effects and identify heterogeneous genes. We also proposed another meta-analytic approach that incorporates informative weights in the random effects meta-analysis models. We compared the proposed …
Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima
Controlling For Confounding When Association Is Quantified By Area Under The Roc Curve, Hadiza I. Galadima
Theses and Dissertations
In the medical literature, there has been an increased interest in evaluating association between exposure and outcomes using nonrandomized observational studies. However, because assignments to exposure are not done randomly in observational studies, comparisons of outcomes between exposed and non-exposed subjects must account for the effect of confounders. Propensity score methods have been widely used to control for confounding, when estimating exposure effect. Previous studies have shown that conditioning on the propensity score results in biased estimation of odds ratio and hazard ratio. However, there is a lack of research into the performance of propensity score methods for estimating the …
Realistic Spiking Neuron Statistics In A Population Are Described By A Single Parametric Distribution, Lauren Crow 9370373
Realistic Spiking Neuron Statistics In A Population Are Described By A Single Parametric Distribution, Lauren Crow 9370373
Undergraduate Research Posters
The spiking of activity of neurons throughout the cortex is random and complicated. This complicated activity requires theoretical formulations in order to understand the underlying principles of neural processing. A key aspect of theoretical investigations is characterizing the probability distribution of spiking activity. This study aims to better understand the statistics of the time between spikes, or interspike interval, in both real data and a spiking model with many time scales. Exploration of the interspike intervals of neural network activity can provide a better understanding of neural responses to different stimuli. We consider different parametric distribution fitting techniques to characterize …
Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello
Application And Extension Of Weighted Quantile Sum Regression For The Development Of A Clinical Risk Prediction Tool, Ghalib Bello
Theses and Dissertations
In clinical settings, the diagnosis of medical conditions is often aided by measurement of various serum biomarkers through the use of laboratory tests. These biomarkers provide information about different aspects of a patient’s health and the overall function of different organs. In this dissertation, we develop and validate a weighted composite index that aggregates the information from a variety of health biomarkers covering multiple organ systems. The index can be used for predicting all-cause mortality and could also be used as a holistic measure of overall physiological health status. We refer to it as the Health Status Metric (HSM). Validation …
A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng
A Study Of The Relationship Between Childhood Body Size And Adult Blood Pressure, Cardiovascular Structure And Function, Yangyang Deng
Theses and Dissertations
BACKGROUND: Little is known of the effects of obesity, body size and body composition, and blood pressure (BP) in childhood on hypertension (HBP) and cardiac structure and function in adulthood due to the lack of long-term serial data on these parameters from childhood into adulthood. In the present study, we are poised to analyze these serial data from the Fels Longitudinal Study (FLS) to evaluate the extent to which body size during childhood determines HBP and cardiac structure and function in the same individuals in adulthood through mathematical modeling. METHODS: The data were from 412 males and 403 females in …
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Incorporating Dependence Boundaries In Simulating Associated Discrete Data, Mary E. Haynes
Theses and Dissertations
In the study of associated discrete variables, limitations on the range of the possible association measures (Pearson correlation, odds ratio, etc.) arise from the form of the joint probability function between the variables. These limitations are known as the Fréchet bounds. The bounds for cases involving associated binary variables are explored in the context of simulating datasets with a desired correlation and set of marginal probabilities. A new method for creating such datasets is compared to an existing method that uses the multivariate probit. A method for simulating associated binary variables using a desired odds ratio and known marginal probabilities …
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Methods For Integrative Analysis Of Genomic Data, Paul Manser
Theses and Dissertations
In recent years, the development of new genomic technologies has allowed for the investigation of many regulatory epigenetic marks besides expression levels, on a genome-wide scale. As the price for these technologies continues to decrease, study sizes will not only increase, but several different assays are beginning to be used for the same samples. It is therefore desirable to develop statistical methods to integrate multiple data types that can handle the increased computational burden of incorporating large data sets. Furthermore, it is important to develop sound quality control and normalization methods as technical errors can compound when integrating multiple genomic …
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Latent Variable Models Given Incompletely Observed Surrogate Outcomes And Covariates, Chunfeng Ren
Theses and Dissertations
Latent variable models (LVMs) are commonly used in the scenario where the outcome of the main interest is an unobservable measure, associated with multiple observed surrogate outcomes, and affected by potential risk factors. This thesis develops an approach of efficient handling missing surrogate outcomes and covariates in two- and three-level latent variable models. However, corresponding statistical methodologies and computational software are lacking efficiently analyzing the LVMs given surrogate outcomes and covariates subject to missingness in the LVMs. We analyze the two-level LVMs for longitudinal data from the National Growth of Health Study where surrogate outcomes and covariates are subject to …
The Total Picture: Multiple Chemical Exposures To Pregnant Women In The Us – An Nhanes Study Of Data From 2003 Through 2010, Teri Cabana
Theses and Dissertations
INTRODUCTION: Chemical exposures to US pregnant women have been shown to have adverse health impacts on both mother and fetus. A prior paper revealed that US pregnant women in 2003-2004 had widespread exposure to multiple chemicals. The goal of this research is to examine how environmental chemical exposures to US pregnant women have changed from 2003 to 2010 and to look further at the extent of simultaneous exposure to multiple chemicals in US pregnant women using biomonitoring data available through NHANES (the National Health and Nutritional Examination Survey). METHODS: Using available NHANES data from the following cycles (2003-2004, 2005-2006, 2007-2008, …