Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Life Sciences (18)
- Medicine and Health Sciences (17)
- Applied Statistics (13)
- Statistical Models (13)
- Applied Mathematics (9)
-
- Genetics and Genomics (9)
- Statistical Methodology (8)
- Longitudinal Data Analysis and Time Series (7)
- Bioinformatics (6)
- Genetics (5)
- Microarrays (5)
- Multivariate Analysis (5)
- Public Health (5)
- Genomics (4)
- Probability (4)
- Social and Behavioral Sciences (4)
- Data Science (3)
- Diseases (3)
- Ecology and Evolutionary Biology (3)
- Medical Sciences (3)
- Ordinary Differential Equations and Applied Dynamics (3)
- Population Biology (3)
- Psychology (3)
- Survival Analysis (3)
- Categorical Data Analysis (2)
- Computational Biology (2)
- Computer Sciences (2)
- Keyword
-
- Biostatistics (6)
- Clinical trial (4)
- Genomics (4)
- Environment (3)
- Gene expression (3)
-
- Longitudinal data (3)
- Power (3)
- Adaptive design (2)
- Algorithm (2)
- Bayesian (2)
- Chemicals (2)
- Classification (2)
- Cluster analysis (2)
- Covariance (2)
- DNA Methylation (2)
- Dropout (2)
- Epidemiology (2)
- Genetics (2)
- Hi-C (2)
- Meta-analysis (2)
- Microarray (2)
- Missing data (2)
- Mixed models (2)
- Model selection (2)
- NHANES (2)
- Neuroscience (2)
- Principal Components Analysis (2)
- ROC curves (2)
- Structural Equation Modeling (2)
- Survival analysis (2)
- Publication Year
- Publication
- Publication Type
Articles 1 - 30 of 108
Full-Text Articles in Biostatistics
Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich
Propensity Score Matching Accounting For Longitudinal Trends Before Baseline With Group-Based Trajectory Modeling, Dustin R. Bastaich
Theses and Dissertations
Propensity score matching is used in observational studies to balance baseline attributes between a treatment of interest and a control group. Propensity score matching typically relies on baseline variables, but longitudinal trends in patient characteristics can also influence treatment decisions and subsequent health outcomes. This dissertation extends standard approaches by explicitly incorporating longitudinal trajectories of key variables into the propensity score estimation process.
Trends in a longitudinal variable prior to baseline were characterized using group-based trajectory modeling. A two-step modeling approach was implemented where trajectory groups of a key variable were first estimated and then included as covariates in the …
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Theses and Dissertations
Traditional models in psychiatric research often impose assumptions of causal homogeneity, treating population-level associations as reflective of uniform underlying mechanisms. This dissertation challenges that assumption by introducing statistical and machine learning frameworks designed to detect and model causal heterogeneity in the development of psychopathology. Central to this approach is the advancement of finite mixture structural equation modeling (FM-SEM) to identify latent subgroups characterized by distinct, and sometimes opposing, causal pathways.
The dissertation comprises three integrated empirical studies. The first introduces mixDoC, a finite mixture extension of the classical Direction of Causation (DoC) model applied to twin data, enabling the detection …
An Adaptive Method For Covariate Balancing In Block Randomized Clinical Trials, Ren Rasnick
An Adaptive Method For Covariate Balancing In Block Randomized Clinical Trials, Ren Rasnick
Theses and Dissertations
Clinical trials are randomized in part to limit allocation bias, but also to ensure comparability between treatment arms for a baseline variable of concern. Comparable with regard to a baseline variable of concern is necessary for the validity of statistical methods. However, comparability is not guaranteed for trials of any size and is even more likely in trials with < 200 total participants. We propose a new method for adapting the allocation of participants in a sequentially allocated two-armed study with a small sample size to better ensure comparability.
The proposed method calculates the expected final imbalance (lack of comparability) based on the current participant values. Unlike several other methods, our method ensures the final desired sample size for each treatment arm, utilizes the expected final imbalance, increases comparability between …
The Genetic Architecture Of Cervical Change During Pregnancy: From Modeling To Mechanism — Does The Cervix Mediate Maternal Risk For Spontaneous Preterm Birth?, Hope M. Wolf
Theses and Dissertations
This project leverages clinical data and biospecimens from a prospective longitudinal cohort of pregnant women to study the genetic and phenotypic relationships between cervical shortening and the duration of pregnancy. Sonographic cervical length (CL) was measured throughout pregnancy in a cohort of 5,160 Black/African American women in Detroit, Michigan. Maternal DNA samples were sequenced with a next-generation low-pass whole genome platform. The heritability of cervical change during pregnancy and its genetic correlation with gestational age at delivery (GAD) were estimated using Genome-Wide Complex Trait Analysis. These estimates suggest that cervical change is heritable (h²CL = 51%) and highly polygenic trait. …
A Novel Family Of Chain Binomial Models To Investigate Correlated Vaccination And Infection Rates In Sveirs Epidemic Dynamics, Divine Wanduku
A Novel Family Of Chain Binomial Models To Investigate Correlated Vaccination And Infection Rates In Sveirs Epidemic Dynamics, Divine Wanduku
Biology and Medicine Through Mathematics Conference
No abstract provided.
Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi
Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi
Graduate Research Posters
Background: Head and neck cancer is the 6th most common cancer worldwide with an expected 1.08 million new cases each year. Such cancer data are ultra-high dimensional with thousands of clinical features and gene expressions, making it challenging for the traditional analytical tools to extract the potential biomarker for the cancer survival and control false discoveries. In addition, presence of heavy censoring can affect the screening procedures based on Kaplan-Meier (K-M) survival estimates.
Aim: To propose a model free, ultra-high dimensional feature screening method with two-dimensional survival outcome allowing false discovery rate (FDR) control.
Method: 516 primary tumor patients with …
Variability In Causal Effects On A Binary Outcome And Noncompliance In A Multisite Randomized Trial, Xinxin Sun
Variability In Causal Effects On A Binary Outcome And Noncompliance In A Multisite Randomized Trial, Xinxin Sun
Theses and Dissertations
Noncompliance to treatment assignment is widespread in randomized trials and presents challenges in causal inference. In the presence of noncompliance, the most commonly estimated effect of treatment assignment, also known as intent-to-treat (ITT) effect, is biased. Of interest in this setting is the complier average causal effect (CACE), the ITT effect among compliers. Further complication arises when the outcome variable is partially observed.
My research focuses on estimating the distribution of a site-specific CACE in a multisite randomized controlled trial (MRCT) by maximum likelihood (ML). Assuming compliance missing at random (MAR). We express the likelihood as an integral with respect …
Reassessing Replication: Addressing The Replication Crisis From A Statistical Perspective, Alicia Richards Phd
Reassessing Replication: Addressing The Replication Crisis From A Statistical Perspective, Alicia Richards Phd
Theses and Dissertations
In 2015, Open Science Framework directly replicated 100 psychology studies and found astonishingly low replication rates. Since, researchers have suggested factors that may have influenced the low rates, including the metrics used to assess replications. The definitions used to decide whether a replication study was successful all suffer from flaws. Therefore, we propose a new metric for assessing replication that can estimate the likelihood a study successfully replicated rather than forcing a binary choice and accounts for study design limitations.
Using equivalence study techniques, we first propose a new metric to assess replication, defining a successful replication as one where …
Model-Based Imputation Of Below Detection Limit Missing Data And Group Selection In Bayesian Group Index Regression, Matthew Carli
Model-Based Imputation Of Below Detection Limit Missing Data And Group Selection In Bayesian Group Index Regression, Matthew Carli
Theses and Dissertations
Investigations into the association between chemical exposure and health outcomes are increasingly focused on the role of chemical mixtures, as opposed to individual chemicals. The analysis of chemical mixture data required the development of novel statistical methods, one of these being Bayesian group index regression. A statistical challenge common to all chemical mixture analyses is the ubiquitous presence of below detection limit (BDL) data. We propose an extension of Bayesian group index regression that treats both regression effects and missing BDL observations as parameters in a model estimated through a Markov Chain Monte Carlo algorithm that we refer to as …
Integrative Post-Gwas Analyses Of Psychiatric Disorders: Identifying Putative Risk Genes And Gene Sets Using Transcriptome, Proteome And Methylome Information, Huseyin Gedik
Theses and Dissertations
Genome-wide association studies (GWAS) of psychiatric disorders (PD) yield numerous loci with significant signals, but often they do not implicate specific protein coding genes. Because GWAS risk loci are enriched in expression/protein/methylation quantitative loci (e/p/mQTL, hereafter xQTL), transcriptome/proteome/methylome-wide association studies (T/P/MWAS, hereafter XWAS), which integrate information from GWAS and x-level (mRNA, protein or DNA methylation levels) coming from largest xQTL studies, can link GWAS signals to effects on specific genes. For gene level analyses, researchers use mendelian randomization (MR) methods to fine-map the association between x-levels and trait. However, none of the previous studies ever jointly analyzed XWAS of multiple …
Rewriting The Rules For Diagnostics: Implications Of Probability And Measure Theory For Sars-Cov-2 Testing, Paul Patrone, Anthony Kearsley
Rewriting The Rules For Diagnostics: Implications Of Probability And Measure Theory For Sars-Cov-2 Testing, Paul Patrone, Anthony Kearsley
Biology and Medicine Through Mathematics Conference
No abstract provided.
Long-Read Sequencing Of The Zebrafish Genome Reorganizes Genomic Architecture, Yelena Chernyavskaya, Xiaofei Zhang, Jinze Liu, Jessica Blackburn
Long-Read Sequencing Of The Zebrafish Genome Reorganizes Genomic Architecture, Yelena Chernyavskaya, Xiaofei Zhang, Jinze Liu, Jessica Blackburn
Biostatistics Publications
Background
Nanopore sequencing technology has revolutionized the field of genome biology with its ability to generate extra-long reads that can resolve regions of the genome that were previously inaccessible to short-read sequencing platforms. Over 50% of the zebrafish genome consists of difficult to map, highly repetitive, low complexity elements that pose inherent problems for short-read sequencers and assemblers.
Results
We used long-read nanopore sequencing to generate a de novo assembly of the zebrafish genome and compared our assembly to the current reference genome, GRCz11. The new assembly identified 1697 novel insertions and deletions over one kilobase in length and placed …
Estimating Weighted Panel Sizes For Primary Care Providers: An Assessment Of Clustering And Novel Methods Of Panel Size Estimation On Electronic Medical Records, Martin A. Lavallee
Estimating Weighted Panel Sizes For Primary Care Providers: An Assessment Of Clustering And Novel Methods Of Panel Size Estimation On Electronic Medical Records, Martin A. Lavallee
Theses and Dissertations
Primary Care is on the frontlines of healthcare, thus they see the most diverse set of patients. In order to achieve high functioning primary care, a practice must establish empanelment, the pairing of patients to providers. Enumeration of empanelment, or estimating panel sizes, helps ensure that the demands of the patients demand the supply of providers and optimize the balance of primary care resources to improve quality of care. Further we can adjust panel sizes by using patient-level data on healthcare utilization and complexity extracted from the electronic medial record to determine the amount of care or burden of work …
Statistical Approaches For Estimation And Comparison Of Brain Functional Connectivity, Jifang Zhao
Statistical Approaches For Estimation And Comparison Of Brain Functional Connectivity, Jifang Zhao
Theses and Dissertations
Drug addiction can lead to many health-related problems and social concerns. Functional connectivity obtained from functional magnetic resonance imaging (fMRI) data promotes a variety of fundamental understandings in such association. Due to its complex correlation structure and large dimensionality, the modeling and analysis of the functional connectivity from neuroimage are challenging. By proposing a spatio-temporal model for multi-subject neuroimage data, we incorporate voxel-level spatio-temporal dependencies of whole-brain measurements to improve the accuracy of statistical inference. To tackle large-scale spatio-temporal neuroimage data, we develop a computationally efficient algorithm to estimate the parameters. Our method is used to identify functional connectivity and …
Principal Components Analysis Corrects Collider Bias In Polygenic Risk Score Effect Size Estimation, Nathaniel S. Thomas, Peter B. Barr, Fazil Aliev, Sally I. Kuo, Danielle M. Dick, Jessica E. Salvatore
Principal Components Analysis Corrects Collider Bias In Polygenic Risk Score Effect Size Estimation, Nathaniel S. Thomas, Peter B. Barr, Fazil Aliev, Sally I. Kuo, Danielle M. Dick, Jessica E. Salvatore
Graduate Research Posters
BACKGROUND: Genome-wide polygenic scoring has emerged as a way to predict psychiatric and behavioral outcomes and identify environments that promote the expression of genetic risks. An increasing number of studies demonstrate that the effects of polygenic risk scores (PRS) may be biased by the inclusion of heritable environments as covariates when the environment is influenced by unmeasured confounding variables, an example of collider bias. Inclusion of the principal components of observed confounders as covariates may correct for the effect of unmeasured confounders.
METHODS: A simulation study was conducted to test principal components analysis (PCA) as a correction for collider bias. …
Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis
Methods For Developing A Machine Learning Framework For Precise 3d Domain Boundary Prediction At Base-Level Resolution, Spiro C. Stilianoudakis
Theses and Dissertations
High-throughput chromosome conformation capture technology (Hi-C) has revealed extensive DNA looping and folding into discrete 3D domains. These include Topologically Associating Domains (TADs) and chromatin loops, the 3D domains critical for cellular processes like gene regulation and cell differentiation. The relatively low resolution of Hi-C data (regions of several kilobases in size) prevents precise mapping of domain boundaries by conventional TAD/loop-callers. However, high resolution genomic annotations associated with boundaries, such as CTCF and members of cohesin complex, suggest a computational approach for precise location of domain boundaries.
We developed preciseTAD, an optimized machine learning framework that leverages a random …
Analyzing Electronic Health Records With Time-To-Event Endpoints: Propensity Scores And Semiparametric Approaches, Jonathan W. Yu
Analyzing Electronic Health Records With Time-To-Event Endpoints: Propensity Scores And Semiparametric Approaches, Jonathan W. Yu
Theses and Dissertations
For analyzing large electronic health records (EHR) with time-to-event endpoints, such as in kidney transplantation, a major challenge is to provide an accurate risk analyses, while accounting for a multitude of epidemiological and statistical complexities. Motivated by a right-censored kidney transplantation EHR dataset derived from the United Network of Organ Sharing (UNOS), this dissertation, through a culmination of two interrelated yet distinctly different projects, focuses on developments of novel statistical procedures and methodologies to address some pressing issues arising in EHR-based research. In the first project, we aim to decouple the causal effects of treatments (here, studying subgroups, such as …
Prognostic Modeling Of Recovery Following Stem Cell Transplantation, Brielle A. Forsthoffer
Prognostic Modeling Of Recovery Following Stem Cell Transplantation, Brielle A. Forsthoffer
Theses and Dissertations
Predicting the trajectory of lymphoid recovery following myeloablative hematopoietic stem cell transplantation (SCT) can help guide subsequent therapeutic decisions, since poor recovery has been associated with graft-versus-host disease (GVHD), relapse and mortality. Previous attempts at classifying patients depended on absolute criteria being set prior to modeling absolute lymphocyte counts (ALCs) over time. Having an empirical clinical decision support tool for objectively determining the trajectory an individual might take during their recovery would be advantageous. We propose using growth-based trajectory modeling (GBTM) and growth mixture modeling (GMM), which utilize machine learning algorithms to empirically identify latent groupings of data. Due to …
Bayesian Techniques For Relating Genetic Polymorphisms To Diffusion Tensor Images Of Cocaine Users, Tmader Alballa
Bayesian Techniques For Relating Genetic Polymorphisms To Diffusion Tensor Images Of Cocaine Users, Tmader Alballa
Theses and Dissertations
Past investigations utilizing Diffusion Tensor Imaging (DTI) have demonstrated that cocaine use disorder (CUD) yields white matter changes. We proposed three Bayesian techniques in order to explore the relationship between Fractional Anisotropy (FA), genetic data, and years of cocaine use (YCU). CUD participants exhibit abnormality in different areas of the brain versus non-drug using controls, which is measured by DTI. This dissertation is motivated by a neuroimaging genetic study in cocaine dependence, which found that there were relationships between several genes such as GAD and 5-HT2R and CUD subjects.
In the first chapter, there is background on the …
Statistical Inference Of Adaptation At Multiple Genomic Scales Using Supervised Classification And A Hidden Markov Model, Lauren A. Sugden
Statistical Inference Of Adaptation At Multiple Genomic Scales Using Supervised Classification And A Hidden Markov Model, Lauren A. Sugden
Biology and Medicine Through Mathematics Conference
No abstract provided.
Integrated Multiple Adaptive Design Involving Sample Size Re-Estimation And (Covariate-Adjusted) Response-Adaptive Randomization For Continuous And Binary Outcomes, Christine M. Orndahl
Integrated Multiple Adaptive Design Involving Sample Size Re-Estimation And (Covariate-Adjusted) Response-Adaptive Randomization For Continuous And Binary Outcomes, Christine M. Orndahl
Theses and Dissertations
Historically, clinical trials have been performed based on decisions made prior to the start of the trial. Adaptive designs have been developed to provide increased flexibility, allowing pre-specified changes to occur based on interim data. Each adaptive design addresses a unique pitfall of a non-adaptive design, such as minimizing the chance of an under- or over-powered study by utilizing interim data to update the sample size estimate (sample size re-estimation) or increasing the ethical benefit of a trial by allocating more participants to the better performing treatment group ([covariate-adjusted] response-adaptive randomization). Additional benefit is attainable by combining more than one …
Adjusting For Dropout In Randomized Controlled Clinical Trials, Katharine Stromberg
Adjusting For Dropout In Randomized Controlled Clinical Trials, Katharine Stromberg
Theses and Dissertations
Dropout is a common issue in randomized controlled clinical trials and can negatively impact the internal validity of a study and potentially bias the treatment effect. When subjects discontinue study participation, they are not being given the opportunity to gain from the investigational therapy as if they had remained in the study, defeating one of the main purposes of clinical trials, providing treatment. Specifically in unblinded studies, such as the wait-list control (WLC) design, dropout is often due to group membership. Subjects allocated to the control group often dropout at higher rates than in the treatment group. Adaptive designs have …
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Zero-Inflated Longitudinal Mixture Model For Stochastic Radiographic Lung Compositional Change Following Radiotherapy Of Lung Cancer, Viviana A. Rodríguez Romero
Theses and Dissertations
Compositional data (CD) is mostly analyzed as relative data, using ratios of components, and log-ratio transformations to be able to use known multivariable statistical methods. Therefore, CD where some components equal zero represent a problem. Furthermore, when the data is measured longitudinally, observations are spatially related and appear to come from a mixture population, the analysis becomes highly complex. For this matter, a two-part model was proposed to deal with structural zeros in longitudinal CD using a mixed-effects model. Furthermore, the model has been extended to the case where the non-zero components of the vector might a two component mixture …
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Accounting For The Uncertainty Due To Chemicals Below The Detection Limit In Mixture Analysis, Paul M. Hargarten
Theses and Dissertations
Humans are exposed to multiple chemicals every day. Epidemiological studies have shown that chemical mixtures are associated with cancers, allergies, neurodevelopmental disorders, and other adverse health effects. To assess these associations, investigators are increasingly using chemical mixture approaches like weighted quantile sum (WQS) regression. In these studies, the research objectives are to determine whether a mixture of correlated chemicals is associated with an adverse health outcome and to identify the important chemicals. However, as experimental equipment measures each exposure to a chemical-specific detection limit, the exposures are unknown between zero and the detection limit. Indeed, the number of exposures below …
Estimating Response Status In Sequential Multiple Assignment (Smar)-Like Trials, Keighly Bradbrook
Estimating Response Status In Sequential Multiple Assignment (Smar)-Like Trials, Keighly Bradbrook
Theses and Dissertations
Sequential, multiple assignment, randomized trials (SMARTs) allow investigators to develop and compare experimental treatment regimens in which individuals are successively randomized to different treatments based on some set of predetermined rules. The rules used to make decisions on when and how to switch treatments are based on a chosen set of tailoring variables. Although not always true, intermediate response is commonly used as the primary tailoring variable as it is often predictive of future treatment success. As such, successful implementation depends on identifying patients who respond to treatment, though in some situations such mechanisms may not exist. Further, patient-level covariates …
The Analysis Of Neural Heterogeneity Through Mathematical And Statistical Methods, Kyle Wendling
The Analysis Of Neural Heterogeneity Through Mathematical And Statistical Methods, Kyle Wendling
Theses and Dissertations
Diversity of intrinsic neural attributes and network connections is known to exist in many areas of the brain and is thought to significantly affect neural coding. Recent theoretical and experimental work has argued that in uncoupled networks, coding is most accurate at intermediate levels of heterogeneity. I explore this phenomenon through two distinct approaches: a theoretical mathematical modeling approach and a data-driven statistical modeling approach.
Through the mathematical approach, I examine firing rate heterogeneity in a feedforward network of stochastic neural oscillators utilizing a high-dimensional model. The firing rate heterogeneity stems from two sources: intrinsic (different individual cells) and network …
Characterizing The Permanence And Stationary Distribution For A Family Of Malaria Stochastic Models, Divine Wanduku
Characterizing The Permanence And Stationary Distribution For A Family Of Malaria Stochastic Models, Divine Wanduku
Biology and Medicine Through Mathematics Conference
No abstract provided.
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Bayesian Nonparametric Analysis Of Longitudinal Data With Non-Ignorable Non-Monotone Missingness, Yu Cao
Theses and Dissertations
In longitudinal studies, outcomes are measured repeatedly over time, but in reality clinical studies are full of missing data points of monotone and non-monotone nature. Often this missingness is related to the unobserved data so that it is non-ignorable. In such context, pattern-mixture model (PMM) is one popular tool to analyze the joint distribution of outcome and missingness patterns. Then the unobserved outcomes are imputed using the distribution of observed outcomes, conditioned on missing patterns. However, the existing methods suffer from model identification issues if data is sparse in specific missing patterns, which is very likely to happen with a …
Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace
Site- And Location-Adjusted Approaches To Adaptive Allocation Clinical Trial Designs, Brian S. Di Pace
Theses and Dissertations
Response-Adaptive (RA) designs are used to adaptively allocate patients in clinical trials. These methods have been generalized to include Covariate-Adjusted Response-Adaptive (CARA) designs, which adjust treatment assignments for a set of covariates while maintaining features of the RA designs. Challenges may arise in multi-center trials if differential treatment responses and/or effects among sites exist. We propose Site-Adjusted Response-Adaptive (SARA) approaches to account for inter-center variability in treatment response and/or effectiveness, including either a fixed site effect or both random site and treatment-by-site interaction effects to calculate conditional probabilities. These success probabilities are used to update assignment probabilities for allocating patients …
Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer
Methods For Evaluating Dropout Attrition In Survey Data, Camille J. Hochheimer
Theses and Dissertations
As researchers increasingly use web-based surveys, the ease of dropping out in the online setting is a growing issue in ensuring data quality. One theory is that dropout or attrition occurs in phases that can be generalized to phases of high dropout and phases of stable use. In order to detect these phases, several methods are explored. First, existing methods and user-specified thresholds are applied to survey data where significant changes in the dropout rate between two questions is interpreted as the start or end of a high dropout phase. Next, survey dropout is considered as a time-to-event outcome and …