Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (23)
- University of Kentucky (20)
- Virginia Commonwealth University (9)
- The Texas Medical Center Library (7)
- Dartmouth College (6)
-
- Michigan Technological University (4)
- University of Louisville (4)
- Himmelfarb Health Sciences Library, The George Washington University (3)
- University of Nebraska - Lincoln (3)
- Old Dominion University (2)
- University of Arkansas, Fayetteville (2)
- Wayne State University (2)
- Yale University (2)
- Cal Poly Humboldt (1)
- California Polytechnic State University, San Luis Obispo (1)
- Chapman University (1)
- City University of New York (CUNY) (1)
- Illinois State University (1)
- LSU Health New Orleans (1)
- Liberty University (1)
- Louisiana Tech University (1)
- Missouri University of Science and Technology (1)
- University at Albany, State University of New York (1)
- University of Denver (1)
- University of Massachusetts Boston (1)
- University of Montana (1)
- University of Nebraska Medical Center (1)
- University of South Dakota (1)
- University of Texas at Arlington (1)
- University of Texas at El Paso (1)
- Keyword
-
- Genetics (16)
- Humans (9)
- Bioinformatics (6)
- Genomics (6)
- Algorithms (5)
-
- Female (5)
- GWAS (5)
- Methods (5)
- Feature selection (4)
- Gene expression (4)
- Machine learning (4)
- Diagnosis (3)
- Gene Expression Regulation (3)
- Gene expression profiling (3)
- Genetic (3)
- Longitudinal data (3)
- Male (3)
- Pregnancy (3)
- SNP (3)
- Statistical (3)
- Statistics (3)
- Alzheimer Disease (2)
- Bayesian (2)
- Biomarkers (2)
- Breast cancer (2)
- Chemistry (2)
- Computational Biology (2)
- Computational biology (2)
- Computer simulation (2)
- DNA methylation (2)
- Publication Year
- Publication
-
- Biostatistics Faculty Publications (11)
- Dissertations and Theses (Open Access) (7)
- Harvard University Biostatistics Working Paper Series (7)
- Theses and Dissertations (6)
- COBRA Preprint Series (5)
-
- Dartmouth Scholarship (5)
- Dissertations, Master's Theses and Master's Reports (4)
- Electronic Theses and Dissertations (4)
- Epidemiology and Environmental Health Faculty Publications (4)
- U.C. Berkeley Division of Biostatistics Working Paper Series (4)
- UW Biostatistics Working Paper Series (4)
- Epidemiology Faculty Publications (2)
- Graduate Research Posters (2)
- Graduate Theses and Dissertations (2)
- Internal Medicine Faculty Publications (2)
- Master's Theses (2)
- The University of Michigan Department of Biostatistics Working Paper Series (2)
- Yale Day of Data (2)
- Annual Symposium on Biomathematics and Ecology Education and Research (1)
- Biology Dissertations - Archive (1)
- Biology and Medicine Through Mathematics Conference (1)
- Cal Poly Humboldt theses and projects (1)
- Capstone Experience: Master of Public Health (1)
- Center for Molecular Medicine and Genetics (1)
- Computer Science Faculty Publications (1)
- Dartmouth College Ph.D Dissertations (1)
- Department of Animal Science: Faculty Publications (1)
- Department of Statistics: Dissertations, Theses, and Student Research (1)
- Dissertations and Theses (1)
- Dissertations, Theses, and Capstone Projects (1)
- Publication Type
- File Type
Articles 1 - 30 of 105
Full-Text Articles in Biostatistics
Generating Predictive Gene Expression Signatures For Alzheimer's Disease Using Postmortem Brain Tissue, Ashley Duche
Generating Predictive Gene Expression Signatures For Alzheimer's Disease Using Postmortem Brain Tissue, Ashley Duche
Pharmaceutical Sciences (PhD) Dissertations
Background: Alzheimer’s Disease (AD) is a progressive neurodegenerative disorder characterized by the accumulation of amyloid-beta (Aβ) plaques and tau protein aggregates. These pathological features develop in specific brain regions, but why some areas are more vulnerable to early AD-related changes remains unclear. To address this, predictive gene expression signatures were developed to explore the molecular mechanisms underlying regional susceptibility to AD pathology.
Methods: This was performed using postmortem brain (PMB) tissue from participants in the Religious Orders Study and Memory and Aging Project (ROSMAP), Mayo Clinic, and Mount Sinai Brain Bank (MSBB) to generate gene expression signatures from six brain …
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
A Persistent Homology Framework For Scrna-Seq: Assessing Clustering Robustness And Quantifying Preprocessing And Integration Effects On Topological Features., Jonah Daneshmand
Electronic Theses and Dissertations
As single-cell RNA sequencing (scRNA-seq) data expands, robust methods for integrating diverse datasets are critical. This dissertation applies Persistent Homology (PH), a technique from Topological Data Analysis (TDA), to a collection of scRNA-seq datasets spanning eight tissue types to quantify how data integration affects topological features and biological interpretability. We assessed global topological structure using Betti curves, Euler characteristics, and persistence landscapes across raw, normalized, and integrated data representations. Our analysis revealed a performance inversion: while conventional methods excelled on unintegrated data, high-granularity topological methods, particularly those sensitive to global data structure, became superior after integration. This suggests a synergy …
Benchmarking Dna Foundation Models For Genomic And Genetic Tasks, Haonan Feng, Lang Wu, Bingxin Zhao, Chad Huff, Jianjun Zhang, Jia Wu, Lifeng Lin, Peng Wei, Chong Wu
Benchmarking Dna Foundation Models For Genomic And Genetic Tasks, Haonan Feng, Lang Wu, Bingxin Zhao, Chad Huff, Jianjun Zhang, Jia Wu, Lifeng Lin, Peng Wei, Chong Wu
School of Medicine Faculty Publications
The rapid evolution of DNA foundation models promises to revolutionize genomics, yet comprehensive evaluations are lacking. Here, we present a comprehensive, unbiased benchmark of five models (DNABERT-2, Nucleotide Transformer V2, HyenaDNA, Caduceus-Ph, and GROVER) across diverse genomic and genetic tasks including sequence classification, gene expression prediction, variant effect quantification, and topologically associating domain (TAD) region recognition, using zero-shot embeddings. Our analysis reveals that mean token embedding consistently and significantly improves sequence classification performance, outperforming other pooling strategies. Model performance varies among tasks and datasets; while general purpose DNA foundation models showed competitive performance in pathogenic variant identification, they were less …
Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher
Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher
Master's Theses
Neuronal cell types are categorized by transcriptomic identity, yet their morphological heterogeneity defies this classification. In response, researchers have adopted unsupervised graph representation learning as a tool to reveal morphological variation within single-class transcriptomic types. However, the complex geometry of neuronal morphology—especially long axons and dense dendrites—challenges graph neural networks, which struggle with message propagation across extended structures. To mitigate this, current approaches enforce sub-sampling on neuronal graphs and omit axons entirely, sacrificing critical biological features for computational efficiency. To overcome this trade-off, this thesis introduces TopoDINO, a self-supervised, topology-aware representation learning model designed to preserve the full hierarchical organization …
Emerging Technologies For Forensic Genetic Identification, Lilly Llanos
Emerging Technologies For Forensic Genetic Identification, Lilly Llanos
Senior Honors Theses
There are many new innovations in forensic science that are being developed for the identification of biological evidence. These techniques include next-generation DNA sequencing, DNA phenotyping, and forensic genetic genealogy. This thesis will explore each, as well as newer applications of proteomics. The methodologies, reliability, practicality of cost and training, moral implications, and past research of each will be discussed. Finally, some ideas for future research and steps to drive growth and greater understanding will be suggested. This will encourage further innovations and the increased acceptance of forensic evidence in court. Each method was found to have both advantages and …
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Beyond Homogeneity: Exploring Causal Heterogeneity In Psychopathology, Philip B. Vinh
Theses and Dissertations
Traditional models in psychiatric research often impose assumptions of causal homogeneity, treating population-level associations as reflective of uniform underlying mechanisms. This dissertation challenges that assumption by introducing statistical and machine learning frameworks designed to detect and model causal heterogeneity in the development of psychopathology. Central to this approach is the advancement of finite mixture structural equation modeling (FM-SEM) to identify latent subgroups characterized by distinct, and sometimes opposing, causal pathways.
The dissertation comprises three integrated empirical studies. The first introduces mixDoC, a finite mixture extension of the classical Direction of Causation (DoC) model applied to twin data, enabling the detection …
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Methods In Statistics, Machine Learning, And Deep Learning For Combining Multi-Omics Dataset, Md Mutasim Billah
Dissertations, Master's Theses and Master's Reports
Transcriptome-wide association studies (TWAS) have emerged as a powerful strategy to bridge genome-wide association studies (GWAS) with gene regulatory mechanisms by integrating genotypic data with gene expression data. While early TWAS methods typically rely on linear models and single-tissue expression references, recent advances underscore the need for flexible, multi-tissue approaches that can capture heterogeneous regulatory architectures and tissue-specific expression patterns. This dissertation introduces a three‑part research project that advances multi‑tissue transcriptome‑wide association studies (TWAS) along complementary axes of methodology, statistical power, and modelling flexibility.
In chapter One, TWAS‑CTL introduces a two‑stage cross‑tissue learner that trains any user‑chosen single‑tissue imputers (STLs) …
Reports Of Autosomal Recessive Disease And Consanguineous Mating Within The Human Population, Johnathon L. Schluter
Reports Of Autosomal Recessive Disease And Consanguineous Mating Within The Human Population, Johnathon L. Schluter
Master's Theses
It is anecdotally evident when investigating published reports of autosomal recessive disease that a substantial number of cases are the result of related (consanguineous) mating. This research seeks to quantify the percent of manuscripts describing autosomal recessive diseases published between 2000 and 2020 in which consanguineous mating is indicated. We analyzed 602 peer-reviewed manuscripts to identify the percentage of cases presented in which consanguineous mating was indicated, the underlying genes (novel gene or new mutation) and geographical region. These papers were accessed through a specific set of parameters on the free access PubMed Central (PMC) database. A total of 552 …
Identifying Environmental And Genetic Risk Factors Of Diseases In Case-Control Studies, Siting Li
Identifying Environmental And Genetic Risk Factors Of Diseases In Case-Control Studies, Siting Li
Dartmouth College Ph.D Dissertations
In response to the increasing efforts in disease prevention and treatment, this thesis applies statistical methods to investigate environmental and genetic risk factors associated with two diseases: bladder cancer and amyotrophic lateral sclerosis (ALS). For bladder cancer, we investigated the association between toenail metal mixture and bladder cancer risk, along with gene expression levels associated with bladder cancer risk. For ALS, our investigation involves identifying genetic variants and gene expression levels associated with ALS risk and exploring gene-smoking interactions linked to ALS risk.
In chapter two, we developed an adaptive-mixture-categorization (AMC)-based g-computation method combining g-computation with optimized exposure categorization. We …
Coral Scar Investigation: An Application Of Machine Learning And Computational Biology Methods To Understand Coral Holobiont Response To Various Tissue Loss Diseases, Emily W. Van Buren
Coral Scar Investigation: An Application Of Machine Learning And Computational Biology Methods To Understand Coral Holobiont Response To Various Tissue Loss Diseases, Emily W. Van Buren
Biology Dissertations - Archive
Coral disease is one of the biggest challenges facing coral reefs that actively changes biodiversity resulting in coral decline. With the rising threat of diseases, corals require biomarkers that reflect the immune systems and differences between common coral tissue loss diseases to best assist in coral restoration efforts. To obtain these biomarkers, my dissertation leverages two previously published datasets from two tissue loss disease exposure studies to investigate genes that are relevant for coral immune pathways, disease susceptibility, and classification between the diseases. In Chapter 2, I use comparative computational biology tools and protein assays to identify the melanin cascade …
Assessing The Utility Of Breast Cancer Polygenic Risk Scores And Association With Clinical Factors In A Population Of Breast Cancer Patients, John L. Slunecka
Assessing The Utility Of Breast Cancer Polygenic Risk Scores And Association With Clinical Factors In A Population Of Breast Cancer Patients, John L. Slunecka
Dissertations and Theses
INTRODUCTION: Breast cancer (BC) is the most common cancer among women and is classified as a complex disease. Advances in population genomics have led to the development of polygenic risk scores (PRSs) with the potential to enhance current risk models, but replication is often limited. OBJECTIVE: We sought to assess the predictive capabilities of two high-powered BC PRSs in a sample population selected for breast cancer. In addition, the capacity of the PRSs to predict clinical variables that could improve BC screening and treatments was explored. METHODS: Two published PRS algorithms (313 vs 3820) were used to score female subjects …
The Genetic Architecture Of Cervical Change During Pregnancy: From Modeling To Mechanism — Does The Cervix Mediate Maternal Risk For Spontaneous Preterm Birth?, Hope M. Wolf
Theses and Dissertations
This project leverages clinical data and biospecimens from a prospective longitudinal cohort of pregnant women to study the genetic and phenotypic relationships between cervical shortening and the duration of pregnancy. Sonographic cervical length (CL) was measured throughout pregnancy in a cohort of 5,160 Black/African American women in Detroit, Michigan. Maternal DNA samples were sequenced with a next-generation low-pass whole genome platform. The heritability of cervical change during pregnancy and its genetic correlation with gestational age at delivery (GAD) were estimated using Genome-Wide Complex Trait Analysis. These estimates suggest that cervical change is heritable (h²CL = 51%) and highly polygenic trait. …
Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi
Novel Feature Evaluation In Ultra-High Dimensional Right-Censored Data, With Applications To Head And Neck Cancer, Atika Farzana Urmi
Graduate Research Posters
Background: Head and neck cancer is the 6th most common cancer worldwide with an expected 1.08 million new cases each year. Such cancer data are ultra-high dimensional with thousands of clinical features and gene expressions, making it challenging for the traditional analytical tools to extract the potential biomarker for the cancer survival and control false discoveries. In addition, presence of heavy censoring can affect the screening procedures based on Kaplan-Meier (K-M) survival estimates.
Aim: To propose a model free, ultra-high dimensional feature screening method with two-dimensional survival outcome allowing false discovery rate (FDR) control.
Method: 516 primary tumor patients with …
Integrative Post-Gwas Analyses Of Psychiatric Disorders: Identifying Putative Risk Genes And Gene Sets Using Transcriptome, Proteome And Methylome Information, Huseyin Gedik
Theses and Dissertations
Genome-wide association studies (GWAS) of psychiatric disorders (PD) yield numerous loci with significant signals, but often they do not implicate specific protein coding genes. Because GWAS risk loci are enriched in expression/protein/methylation quantitative loci (e/p/mQTL, hereafter xQTL), transcriptome/proteome/methylome-wide association studies (T/P/MWAS, hereafter XWAS), which integrate information from GWAS and x-level (mRNA, protein or DNA methylation levels) coming from largest xQTL studies, can link GWAS signals to effects on specific genes. For gene level analyses, researchers use mendelian randomization (MR) methods to fine-map the association between x-levels and trait. However, none of the previous studies ever jointly analyzed XWAS of multiple …
Statistical Methods For Gene Selection And Genetic Association Studies, Xuewei Cao
Statistical Methods For Gene Selection And Genetic Association Studies, Xuewei Cao
Dissertations, Master's Theses and Master's Reports
This dissertation includes five Chapters. A brief description of each chapter is organized as follows.
In Chapter One, we propose a signed bipartite genotype and phenotype network (GPN) by linking phenotypes and genotypes based on the statistical associations. It provides a new insight to investigate the genetic architecture among multiple correlated phenotypes and explore where phenotypes might be related at a higher level of cellular and organismal organization. We show that multiple phenotypes association studies by considering the proposed network are improved by incorporating the genetic information into the phenotype clustering.
In Chapter Two, we first illustrate the proposed GPN …
Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury
Bayesian Methods For Graphical Models With Neighborhood Selection., Sagnik Bhadury
Electronic Theses and Dissertations
Graphical models determine associations between variables through the notion of conditional independence. Gaussian graphical models are a widely used class of such models, where the relationships are formalized by non-null entries of the precision matrix. However, in high-dimensional cases, covariance estimates are typically unstable. Moreover, it is natural to expect only a few significant associations to be present in many realistic applications. This necessitates the injection of sparsity techniques into the estimation method. Classical frequentist methods, like GLASSO, use penalization techniques for this purpose. Fully Bayesian methods, on the contrary, are slow because they require iteratively sampling over a quadratic …
Comparing Machine Learning Techniques With State-Of-The-Art Parametric Prediction Models For Predicting Soybean Traits, Susweta Ray
Department of Statistics: Dissertations, Theses, and Student Research
Soybean is a significant source of protein and oil, and also widely used as animal feed. Thus, developing lines that are superior in terms of yield, protein and oil content is important to feed the ever-growing population. As opposed to the high-cost phenotyping, genotyping is both cost and time efficient for breeders while evaluating new lines in different environments (location-year combinations) can be costly. Several Genomic prediction (GP) methods have been developed to use the marker and environment data effectively to predict the yield or other relevant phenotypic traits of crops. Our study compares a conventional GP method (GBLUP), a …
Genomics Of Postprandial Lipidomics In The Genetics Of Lipid-Lowering Drugs And Diet Network Study, Marguerite R. Irvin, May E. Montasser, Tobias Kind, Sili Fan, Dinesh K. Barupal, Amit Patki, Rikki M. Tanner, Nicole D. Armstrong, Kathleen A. Ryan, Steven A. Claas, Jeffrey R. O’Connell, Hemant K. Tiwari, Donna K. Arnett
Genomics Of Postprandial Lipidomics In The Genetics Of Lipid-Lowering Drugs And Diet Network Study, Marguerite R. Irvin, May E. Montasser, Tobias Kind, Sili Fan, Dinesh K. Barupal, Amit Patki, Rikki M. Tanner, Nicole D. Armstrong, Kathleen A. Ryan, Steven A. Claas, Jeffrey R. O’Connell, Hemant K. Tiwari, Donna K. Arnett
Epidemiology and Environmental Health Faculty Publications
Postprandial lipemia (PPL) is an important risk factor for cardiovascular disease. Inter-individual variation in the dietary response to a meal is known to be influenced by genetic factors, yet genes that dictate variation in postprandial lipids are not completely characterized. Genetic studies of the plasma lipidome can help to better understand postprandial metabolism by isolating lipid molecular species which are more closely related to the genome. We measured the plasma lipidome at fasting and 6 h after a standardized high-fat meal in 668 participants from the Genetics of Lipid-Lowering Drugs and Diet Network study (GOLDN) using ultra-performance liquid chromatography coupled …
Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang
Mixture Model Approaches To Integrative Analysis Of Multi-Omics Data And Spatially Correlated Genomic Data, Ziqiao Wang
Dissertations and Theses (Open Access)
Integrative genomic data analysis is a powerful tool to study the complex biological processes behind a disease. Statistical methods can model the interrelationships of the involved gene activities through jointly analyzing multiple types of genomic data from different platforms (vertical integration), or improve the power of a study through aggregating the same type of genomic data across studies (horizontal integration). In this dissertation, we propose statistical methods and strategies for integrative multi-omics data in association analysis of disease phenotypes, with an emphasis on cancer applications.
We develop a new strategy based on horizontal integration by leveraging publicly available datasets into …
Epigenome-Wide Association Study Of Kidney Function Identifies Trans-Ethnic And Ethnic-Specific Loci, Charles E. Breeze, Anna Batorsky, Mi Kyeong Lee, Mindy D. Szeto, Xiaoguang Xu, Daniel L. Mccartney, Rong Jiang, Amit Patki, Holly J. Kramer, James M. Eales, Laura Raffield, Leslie Lange, Ethan Lange, Peter Durda, Yongmei Liu, Russ P. Tracy, David Van Den Berg, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Topmed Mesa Multi-Omics Working Group, Kathryn L. Evans, William E. Kraus, Donna K. Arnett
Epigenome-Wide Association Study Of Kidney Function Identifies Trans-Ethnic And Ethnic-Specific Loci, Charles E. Breeze, Anna Batorsky, Mi Kyeong Lee, Mindy D. Szeto, Xiaoguang Xu, Daniel L. Mccartney, Rong Jiang, Amit Patki, Holly J. Kramer, James M. Eales, Laura Raffield, Leslie Lange, Ethan Lange, Peter Durda, Yongmei Liu, Russ P. Tracy, David Van Den Berg, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Topmed Mesa Multi-Omics Working Group, Kathryn L. Evans, William E. Kraus, Donna K. Arnett
Epidemiology and Environmental Health Faculty Publications
BACKGROUND: DNA methylation (DNAm) is associated with gene regulation and estimated glomerular filtration rate (eGFR), a measure of kidney function. Decreased eGFR is more common among US Hispanics and African Americans. The causes for this are poorly understood. We aimed to identify trans-ethnic and ethnic-specific differentially methylated positions (DMPs) associated with eGFR using an agnostic, genome-wide approach.
METHODS: The study included up to 5428 participants from multi-ethnic studies for discovery and 8109 participants for replication. We tested the associations between whole blood DNAm and eGFR using beta values from Illumina 450K or EPIC arrays. Ethnicity-stratified analyses were performed using linear …
An Ensemble Of The Icluster Method To Analyze Longitudinal Lncrna Expression Data For Psoriasis Patients, Suyan Tian, Chi Wang
An Ensemble Of The Icluster Method To Analyze Longitudinal Lncrna Expression Data For Psoriasis Patients, Suyan Tian, Chi Wang
Internal Medicine Faculty Publications
BACKGROUND: Psoriasis is an immune-mediated, inflammatory disorder of the skin with chronic inflammation and hyper-proliferation of the epidermis. Since psoriasis has genetic components and the diseased tissue of psoriasis is very easily accessible, it is natural to use high-throughput technologies to characterize psoriasis and thus seek targeted therapies. Transcriptional profiles change correspondingly after an intervention. Unlike cross-sectional gene expression data, longitudinal gene expression data can capture the dynamic changes and thus facilitate causal inference.
METHODS: Using the iCluster method as a building block, an ensemble method was proposed and applied to a longitudinal gene expression dataset for psoriasis, with the …
Chromosome Xq23 Is Associated With Lower Atherogenic Lipid Concentrations And Favorable Cardiometabolic Indices, Pradeep Natarajan, Akhil Pampana, Sarah E. Graham, Sanni E. Ruotsalainen, James A. Perry, Paul S. De Vries, Jai G. Broome, James P. Pirruccello, Michael C. Honigberg, Krishna Aragam, Brooke Wolford, Jennifer A. Brody, Lucinda Antonacci-Fulton, Moscati Arden, Stella Aslibekyan, Themistocles L. Assimes, Christie M. Ballantyne, Lawrence F. Bielak, Joshua C. Bis, Brian E. Cade, Donna K. Arnett
Chromosome Xq23 Is Associated With Lower Atherogenic Lipid Concentrations And Favorable Cardiometabolic Indices, Pradeep Natarajan, Akhil Pampana, Sarah E. Graham, Sanni E. Ruotsalainen, James A. Perry, Paul S. De Vries, Jai G. Broome, James P. Pirruccello, Michael C. Honigberg, Krishna Aragam, Brooke Wolford, Jennifer A. Brody, Lucinda Antonacci-Fulton, Moscati Arden, Stella Aslibekyan, Themistocles L. Assimes, Christie M. Ballantyne, Lawrence F. Bielak, Joshua C. Bis, Brian E. Cade, Donna K. Arnett
Epidemiology and Environmental Health Faculty Publications
Autosomal genetic analyses of blood lipids have yielded key insights for coronary heart disease (CHD). However, X chromosome genetic variation is understudied for blood lipids in large sample sizes. We now analyze genetic and blood lipid data in a high-coverage whole X chromosome sequencing study of 65,322 multi-ancestry participants and perform replication among 456,893 European participants. Common alleles on chromosome Xq23 are strongly associated with reduced total cholesterol, LDL cholesterol, and triglycerides (min P = 8.5 × 10−72), with similar effects for males and females. Chromosome Xq23 lipid-lowering alleles are associated with reduced odds for CHD among 42,545 …
Whole-Exome Sequencing And Hipsc Cardiomyocyte Models Identify Myrip, Trappc11, And Slc27a6 Of Potential Importance To Left Ventricular Hypertrophy In An African Ancestry Population, Marguerite R. Irvin, Praful Aggarwal, Steven A. Claas, Lisa De Las Fuentes, Anh N. Do, C. Charles Gu, Andrea Matter, Benjamin S. Olson, Amit Patki, Karen Schwander, Joshua D. Smith, Vinodh Srinivasasainagendra, Hemant K. Tiwari, Amy J. Turner, Deborah A. Nickerson, Dabeeru C. Rao, Ulrich Broeckel, Donna K. Arnett
Whole-Exome Sequencing And Hipsc Cardiomyocyte Models Identify Myrip, Trappc11, And Slc27a6 Of Potential Importance To Left Ventricular Hypertrophy In An African Ancestry Population, Marguerite R. Irvin, Praful Aggarwal, Steven A. Claas, Lisa De Las Fuentes, Anh N. Do, C. Charles Gu, Andrea Matter, Benjamin S. Olson, Amit Patki, Karen Schwander, Joshua D. Smith, Vinodh Srinivasasainagendra, Hemant K. Tiwari, Amy J. Turner, Deborah A. Nickerson, Dabeeru C. Rao, Ulrich Broeckel, Donna K. Arnett
Epidemiology and Environmental Health Faculty Publications
Background: Indices of left ventricular (LV) structure and geometry represent useful intermediate phenotypes related to LV hypertrophy (LVH), a predictor of cardiovascular (CV) disease (CVD) outcomes.
Methods and Results: We conducted an exome-wide association study of LV mass (LVM) adjusted to height2.7, LV internal diastolic dimension (LVIDD), and relative wall thickness (RWT) among 1,364 participants of African ancestry (AAs) in the Hypertension Genetic Epidemiology Network (HyperGEN). Both single-variant and gene-based sequence kernel association tests were performed to examine whether common and rare coding variants contribute to variation in echocardiographic traits in AAs. We then used a data-driven …
Principal Components Analysis Corrects Collider Bias In Polygenic Risk Score Effect Size Estimation, Nathaniel S. Thomas, Peter B. Barr, Fazil Aliev, Sally I. Kuo, Danielle M. Dick, Jessica E. Salvatore
Principal Components Analysis Corrects Collider Bias In Polygenic Risk Score Effect Size Estimation, Nathaniel S. Thomas, Peter B. Barr, Fazil Aliev, Sally I. Kuo, Danielle M. Dick, Jessica E. Salvatore
Graduate Research Posters
BACKGROUND: Genome-wide polygenic scoring has emerged as a way to predict psychiatric and behavioral outcomes and identify environments that promote the expression of genetic risks. An increasing number of studies demonstrate that the effects of polygenic risk scores (PRS) may be biased by the inclusion of heritable environments as covariates when the environment is influenced by unmeasured confounding variables, an example of collider bias. Inclusion of the principal components of observed confounders as covariates may correct for the effect of unmeasured confounders.
METHODS: A simulation study was conducted to test principal components analysis (PCA) as a correction for collider bias. …
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Graduate Student Theses, Dissertations, & Professional Papers
The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …
Construction And Analysis Of Genetic Regulatory Networks With Rna-Seq Data From Arabidopsis Thaliana, Tessa Kriz
Construction And Analysis Of Genetic Regulatory Networks With Rna-Seq Data From Arabidopsis Thaliana, Tessa Kriz
Dissertations, Master's Theses and Master's Reports
Reconstruction of gene regulatory networks (GRNs) is a fundamental aspect of genetic engineering and provides a deeper understanding of the biological processes of an organism. Two methods were implemented to reconstruct the gene regulatory networks of Arabidopsis thaliana under two treatments: methyl jasmonate (MeJa) and salicylic acid (SA). The Joint Reconstruction of multiple Gene Regulatory Networks (JRmGRN) method was utilized to construct a joint network for identifying hub genes common to both conditions in addition to networks specific to each condition. The Differential Network Analysis with False Discover Rate Control method constructed a network of connections unique to only one …
Statistical Methods In Genetic Studies, Cheng Gao
Statistical Methods In Genetic Studies, Cheng Gao
Dissertations, Master's Theses and Master's Reports
This dissertation includes three Chapters. A brief description of each chapter is organized as follows.
In Chapter 1, we proposed a new method, called MF-TOWmuT, for genome-wide association studies with multiple genetic variants and multiple phenotypes using family samples. MF-TOWmuT uses kinship matrix to account for sample relatedness. It is worth mentioning that in simulations, we considered hidden polygenic effects and varied the proportion of variance contributed by it to generate phenotypes. Simulation studies show that MF-TOWmuT can preserve the type I error rates and is more powerful than several existing methods in different simulation scenarios, MFTOWmuT is also quite …
Gene Set Testing By Distance Correlation, Sho-Hsien Su
Gene Set Testing By Distance Correlation, Sho-Hsien Su
Graduate Theses and Dissertations
Pathways are the functional building blocks of complex diseases such as cancers. Pathway-level studies may provide insights on some important biological processes. Gene set test is an important tool to study the differential expression of a gene set between two groups, e.g., cancer vs normal. The differential expression of a gene set could be due to the difference in mean, variability, or both. However, most existing gene set tests only target the mean difference but overlook other types of differential expression. In this thesis, we propose to use the recently developed distance correlation for gene set testing. To assess the …
Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das
Statistical Approaches Of Gene Set Analysis With Quantitative Trait Loci For High-Throughput Genomic Studies., Samarendra Das
Electronic Theses and Dissertations
Recently, gene set analysis has become the first choice for gaining insights into the underlying complex biology of diseases through high-throughput genomic studies, such as Microarrays, bulk RNA-Sequencing, single cell RNA-Sequencing, etc. It also reduces the complexity of statistical analysis and enhances the explanatory power of the obtained results. Further, the statistical structure and steps common to these approaches have not yet been comprehensively discussed, which limits their utility. Hence, a comprehensive overview of the available gene set analysis approaches used for different high-throughput genomic studies is provided. The analysis of gene sets is usually carried out based on …
Machine Learning Applications For Drug Repurposing, Hansaim Lim
Machine Learning Applications For Drug Repurposing, Hansaim Lim
Dissertations, Theses, and Capstone Projects
The cost of bringing a drug to market is astounding and the failure rate is intimidating. Drug discovery has been of limited success under the conventional reductionist model of one-drug-one-gene-one-disease paradigm, where a single disease-associated gene is identified and a molecular binder to the specific target is subsequently designed. Under the simplistic paradigm of drug discovery, a drug molecule is assumed to interact only with the intended on-target. However, small molecular drugs often interact with multiple targets, and those off-target interactions are not considered under the conventional paradigm. As a result, drug-induced side effects and adverse reactions are often neglected …