Open Access. Powered by Scholars. Published by Universities.®

Medical Genetics Commons

Open Access. Powered by Scholars. Published by Universities.®

Physical Sciences and Mathematics

Institution
Keyword
Publication Year
Publication
Publication Type

Articles 61 - 90 of 106

Full-Text Articles in Medical Genetics

Development Of A Chemical Biology Approach To Uncover The Influence Of Sequence Variations On Ces1 Activity In Live Cells, Samuel James Knebel Jan 2023

Development Of A Chemical Biology Approach To Uncover The Influence Of Sequence Variations On Ces1 Activity In Live Cells, Samuel James Knebel

Masters Theses

Drug metabolism is the biochemical process of modifying drugs to detoxify and remove them through enzymatic transformations. These biotransformation’s occur primarily in the liver and are critical to understanding how pharmaceutical compounds are chemically altered inside the human body. Human carboxylesterases (CESs) catalyze the hydrolysis of esters, amides, thioesters, and carbamates. CES-mediated hydrolysis plays an important role in the metabolism of many drugs including the first FDA approved antiviral treatment for COVID-19, remdesivir (Veklury), the seizure control medication rufinamide (Banzel), and the flu antiviral drug oseltamivir (Tamiflu). CES activity is known to be influenced by a variety of factors including …


Privacy-Aware Estimation Of Relatedness In Admixed Populations, Su Wang, Miran Kim, Wentao Li, Xiaoqian Jiang, Han Chen, Arif Harmanci Nov 2022

Privacy-Aware Estimation Of Relatedness In Admixed Populations, Su Wang, Miran Kim, Wentao Li, Xiaoqian Jiang, Han Chen, Arif Harmanci

Faculty, Staff and Student Publications

BACKGROUND: Estimation of genetic relatedness, or kinship, is used occasionally for recreational purposes and in forensic applications. While numerous methods were developed to estimate kinship, they suffer from high computational requirements and often make an untenable assumption of homogeneous population ancestry of the samples. Moreover, genetic privacy is generally overlooked in the usage of kinship estimation methods. There can be ethical concerns about finding unknown familial relationships in third-party databases. Similar ethical concerns may arise while estimating and reporting sensitive population-level statistics such as inbreeding coefficients for the concerns around marginalization and stigmatization.

RESULTS: Here, we present SIGFRIED, which makes …


The Evolving Privacy And Security Concerns For Genomic Data Analysis And Sharing As Observed From The Idash Competition, Tsung-Ting Kuo, Xiaoqian Jiang, Haixu Tang, Xiaofeng Wang, Arif Harmanci, Miran Kim, Kai Post, Diyue Bu, Tyler Bath, Jihoon Kim, Weijie Liu, Hongbo Chen, Lucila Ohno-Machado Nov 2022

The Evolving Privacy And Security Concerns For Genomic Data Analysis And Sharing As Observed From The Idash Competition, Tsung-Ting Kuo, Xiaoqian Jiang, Haixu Tang, Xiaofeng Wang, Arif Harmanci, Miran Kim, Kai Post, Diyue Bu, Tyler Bath, Jihoon Kim, Weijie Liu, Hongbo Chen, Lucila Ohno-Machado

Faculty, Staff and Student Publications

Concerns regarding inappropriate leakage of sensitive personal information as well as unauthorized data use are increasing with the growth of genomic data repositories. Therefore, privacy and security of genomic data have become increasingly important and need to be studied. With many proposed protection techniques, their applicability in support of biomedical research should be well understood. For this purpose, we have organized a community effort in the past 8 years through the integrating data for analysis, anonymization and sharing consortium to address this practical challenge. In this article, we summarize our experience from these competitions, report lessons learned from the events …


Scgwas: Landscape Of Trait-Cell Type Associations By Integrating Single-Cell Transcriptomics-Wide And Genome-Wide Association Studies, Peilin Jia, Ruifeng Hu, Fangfang Yan, Yulin Dai, Zhongming Zhao Oct 2022

Scgwas: Landscape Of Trait-Cell Type Associations By Integrating Single-Cell Transcriptomics-Wide And Genome-Wide Association Studies, Peilin Jia, Ruifeng Hu, Fangfang Yan, Yulin Dai, Zhongming Zhao

Faculty, Staff and Student Publications

BACKGROUND: The rapid accumulation of single-cell RNA sequencing (scRNA-seq) data presents unique opportunities to decode the genetically mediated cell-type specificity in complex diseases. Here, we develop a new method, scGWAS, which effectively leverages scRNA-seq data to achieve two goals: (1) to infer the cell types in which the disease-associated genes manifest and (2) to construct cellular modules which imply disease-specific activation of different processes.

RESULTS: scGWAS only utilizes the average gene expression for each cell type followed by virtual search processes to construct the null distributions of module scores, making it scalable to large scRNA-seq datasets. We demonstrated scGWAS in …


Federated Learning Algorithms For Generalized Mixed-Effects Model (Glmm) On Horizontally Partitioned Data From Distributed Sources, Wentao Li, Jiayi Tong, Md Monowar Anjum, Noman Mohammed, Yong Chen, Xiaoqian Jiang Oct 2022

Federated Learning Algorithms For Generalized Mixed-Effects Model (Glmm) On Horizontally Partitioned Data From Distributed Sources, Wentao Li, Jiayi Tong, Md Monowar Anjum, Noman Mohammed, Yong Chen, Xiaoqian Jiang

Faculty, Staff and Student Publications

OBJECTIVES: This paper developed federated solutions based on two approximation algorithms to achieve federated generalized linear mixed effect models (GLMM). The paper also proposed a solution for numerical errors and singularity issues. And showed the two proposed methods can perform well in revealing the significance of parameter in distributed datasets, comparing to a centralized GLMM algorithm from R package ('lme4') as the baseline model.

METHODS: The log-likelihood function of GLMM is approximated by two numerical methods (Laplace approximation and Gaussian Hermite approximation, abbreviated as LA and GH), which supports federated decomposition of GLMM to bring computation to data. To solve …


Complement Component C4 Structural Variation And Quantitative Traits Contribute To Sex-Biased Vulnerability In Systemic Sclerosis, Martin Kerick, Marialbert Acosta-Herrera, Carmen Pilar Simeón-Aznar, José Luis Callejas, Shervin Assassi, Susanna M Proudman, Mandana Nikpour, Nicolas Hunzelmann, Gianluca Moroncini, Jeska K De Vries-Bouwstra, Gisela Orozco, Anne Barton, Ariane L Herrick, Chikashi Terao, Yannick Allanore, Carmen Fonseca, Marta Eugenia Alarcón-Riquelme, Timothy R D J Radstake, Lorenzo Beretta, Christopher P Denton, Maureen D Mayes, Javier Martin Oct 2022

Complement Component C4 Structural Variation And Quantitative Traits Contribute To Sex-Biased Vulnerability In Systemic Sclerosis, Martin Kerick, Marialbert Acosta-Herrera, Carmen Pilar Simeón-Aznar, José Luis Callejas, Shervin Assassi, Susanna M Proudman, Mandana Nikpour, Nicolas Hunzelmann, Gianluca Moroncini, Jeska K De Vries-Bouwstra, Gisela Orozco, Anne Barton, Ariane L Herrick, Chikashi Terao, Yannick Allanore, Carmen Fonseca, Marta Eugenia Alarcón-Riquelme, Timothy R D J Radstake, Lorenzo Beretta, Christopher P Denton, Maureen D Mayes, Javier Martin

Faculty, Staff and Student Publications

Copy number (CN) polymorphisms of complement C4 play distinct roles in many conditions, including immune-mediated diseases. We investigated the association of C4 CN with systemic sclerosis (SSc) risk. Imputed total C4, C4A, C4B, and HERV-K CN were analyzed in 26,633 individuals and validated in an independent cohort. Our results showed that higher C4 CN confers protection to SSc, and deviations from CN parity of C4A and C4B augmented risk. The protection contributed per copy of C4A and C4B differed by sex. Stronger protection was afforded by C4A in men and by C4B in women. C4 CN correlated well with its …


Svat: Secure Outsourcing Of Variant Annotation And Genotype Aggregation, Miran Kim, Su Wang, Xiaoqian Jiang, Arif Harmanci Oct 2022

Svat: Secure Outsourcing Of Variant Annotation And Genotype Aggregation, Miran Kim, Su Wang, Xiaoqian Jiang, Arif Harmanci

Faculty, Staff and Student Publications

BACKGROUND: Sequencing of thousands of samples provides genetic variants with allele frequencies spanning a very large spectrum and gives invaluable insight into genetic determinants of diseases. Protecting the genetic privacy of participants is challenging as only a few rare variants can easily re-identify an individual among millions. In certain cases, there are policy barriers against sharing genetic data from indigenous populations and stigmatizing conditions.

RESULTS: We present SVAT, a method for secure outsourcing of variant annotation and aggregation, which are two basic steps in variant interpretation and detection of causal variants. SVAT uses homomorphic encryption to encrypt the data at …


Gpu Accelerated Estimation Of A Shared Random Effect Joint Model For Dynamic Prediction, Shikun Wang, Zhao Li, Lan Lan, Jieyi Zhao, W Jim Zheng, Liang Li Oct 2022

Gpu Accelerated Estimation Of A Shared Random Effect Joint Model For Dynamic Prediction, Shikun Wang, Zhao Li, Lan Lan, Jieyi Zhao, W Jim Zheng, Liang Li

Faculty, Staff and Student Publications

In longitudinal cohort studies, it is often of interest to predict the risk of a terminal clinical event using longitudinal predictor data among subjects at risk by the time of the prediction. The at-risk population changes over time; so does the association between predictors and the outcome, as well as the accumulating longitudinal predictor history. The dynamic nature of this prediction problem has received increasing interest in the literature, but computation often poses a challenge. The widely used joint model of longitudinal and survival data often comes with intensive computation and excessive model fitting time, due to numerical optimization and …


Identifying Candidate Genes And Drug Targets For Alzheimer’S Disease By An Integrative Network Approach Using Genetic And Brain Region-Specific Proteomic Data, Andi Liu, Astrid M Manuel, Yulin Dai, Brisa S Fernandes, Nitesh Enduru, Peilin Jia, Zhongming Zhao Sep 2022

Identifying Candidate Genes And Drug Targets For Alzheimer’S Disease By An Integrative Network Approach Using Genetic And Brain Region-Specific Proteomic Data, Andi Liu, Astrid M Manuel, Yulin Dai, Brisa S Fernandes, Nitesh Enduru, Peilin Jia, Zhongming Zhao

Faculty, Staff and Student Publications

Genome-wide association studies (GWAS) have identified more than 75 genetic variants associated with Alzheimer's disease (ad). However, how these variants function and impact protein expression in brain regions remain elusive. Large-scale proteomic datasets of ad postmortem brain tissues have become available recently. In this study, we used these datasets to investigate brain region-specific molecular pathways underlying ad pathogenesis and explore their potential drug targets. We applied our new network-based tool, Edge-Weighted Dense Module Search of GWAS (EW_dmGWAS), to integrate ad GWAS statistics of 472 868 individuals with proteomic profiles from two brain regions from two large-scale ad cohorts [parahippocampal gyrus …


Delineating Covid-19 Immunological Features Using Single-Cell Rna Sequencing, Wendao Liu, Johnathan Jia, Yulin Dai, Wenhao Chen, Guangsheng Pei, Qiheng Yan, Zhongming Zhao Sep 2022

Delineating Covid-19 Immunological Features Using Single-Cell Rna Sequencing, Wendao Liu, Johnathan Jia, Yulin Dai, Wenhao Chen, Guangsheng Pei, Qiheng Yan, Zhongming Zhao

Faculty, Staff and Student Publications

Understanding the molecular mechanisms of coronavirus disease 2019 (COVID-19) pathogenesis and immune response is vital for developing therapies. Single-cell RNA sequencing has been applied to delineate the cellular heterogeneity of the host response toward COVID-19 in multiple tissues and organs. Here, we review the applications and findings from over 80 original COVID-19 single-cell RNA sequencing studies as well as many secondary analysis studies. We describe that single-cell RNA sequencing reveals multiple features of COVID-19 patients with different severity, including cell populations with proportional alteration, COVID-19-induced genes and pathways, severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) infection in single cells, and adaptation …


Evaluation Of Vicinity-Based Hidden Markov Models For Genotype Imputation, Su Wang, Miran Kim, Xiaoqian Jiang, Arif Ozgun Harmanci Aug 2022

Evaluation Of Vicinity-Based Hidden Markov Models For Genotype Imputation, Su Wang, Miran Kim, Xiaoqian Jiang, Arif Ozgun Harmanci

Faculty, Staff and Student Publications

BACKGROUND: The decreasing cost of DNA sequencing has led to a great increase in our knowledge about genetic variation. While population-scale projects bring important insight into genotype-phenotype relationships, the cost of performing whole-genome sequencing on large samples is still prohibitive. In-silico genotype imputation coupled with genotyping-by-arrays is a cost-effective and accurate alternative for genotyping of common and uncommon variants. Imputation methods compare the genotypes of the typed variants with the large population-specific reference panels and estimate the genotypes of untyped variants by making use of the linkage disequilibrium patterns. Most accurate imputation methods are based on the Li-Stephens hidden Markov …


Pancancer Analysis Of A Potential Gene Mutation Model In The Prediction Of Immunotherapy Outcomes, Lishan Yu, Caifeng Gong Aug 2022

Pancancer Analysis Of A Potential Gene Mutation Model In The Prediction Of Immunotherapy Outcomes, Lishan Yu, Caifeng Gong

Faculty, Staff and Student Publications

Background: Immune checkpoint blockade (ICB) represents a promising treatment for cancer, but predictive biomarkers are needed. We aimed to develop a cost-effective signature to predict immunotherapy benefits across cancers.

Methods: We proposed a study framework to construct the signature. Specifically, we built a multivariate Cox proportional hazards regression model with LASSO using 80% of an ICB-treated cohort (n = 1661) from MSKCC. The desired signature named SIGP was the risk score of the model and was validated in the remaining 20% of patients and an external ICB-treated cohort (n = 249) from DFCI.

Results: SIGP was based on …


A Method For Bridging Population-Specific Genotypes To Detect Gene Modules Associated With Alzheimer's Disease, Yulin Dai, Peilin Jia, Zhongming Zhao, Assaf Gottlieb Jul 2022

A Method For Bridging Population-Specific Genotypes To Detect Gene Modules Associated With Alzheimer's Disease, Yulin Dai, Peilin Jia, Zhongming Zhao, Assaf Gottlieb

Faculty, Staff and Student Publications

BACKGROUND: Genome-wide association studies have successfully identified variants associated with multiple conditions. However, generalizing discoveries across diverse populations remains challenging due to large variations in genetic composition. Methods that perform gene expression imputation have attempted to address the transferability of gene discoveries across populations, but with limited success.

METHODS: Here, we introduce a pipeline that combines gene expression imputation with gene module discovery, including a dense gene module search and a gene set variation analysis, to address the transferability issue. Our method feeds association probabilities of imputed gene expression with a selected phenotype into tissue-specific gene-module discovery over protein interaction …


Tissue-Specific Variations In Transcription Factors Elucidate Complex Immune System Regulation, Hengwei Lu, Yi-Ching Tang, Assaf Gottlieb May 2022

Tissue-Specific Variations In Transcription Factors Elucidate Complex Immune System Regulation, Hengwei Lu, Yi-Ching Tang, Assaf Gottlieb

Faculty, Staff and Student Publications

Gene expression plays a key role in health and disease. Estimating the genetic components underlying gene expression can thus help understand disease etiology. Polygenic models termed "transcriptome imputation" are used to estimate the genetic component of gene expression, but these models typically consider only the cis regions of the gene. However, these cis-based models miss large variability in expression for multiple genes. Transcription factors (TFs) that regulate gene expression are natural candidates for looking for additional sources of the missing variability. We developed a hypothesis-driven approach to identify second-tier regulation by variability in TFs. Our approach tested two models …


An Evidence-Based Lexical Pattern Approach For Quality Assurance Of Gene Ontology Relations, Rashmie Abeysinghe, Yuntao Yang, Mason Bartels, W Jim Zheng, Licong Cui May 2022

An Evidence-Based Lexical Pattern Approach For Quality Assurance Of Gene Ontology Relations, Rashmie Abeysinghe, Yuntao Yang, Mason Bartels, W Jim Zheng, Licong Cui

Faculty, Staff and Student Publications

Gene Ontology (GO) is widely used in the biological domain. It is the most comprehensive ontology providing formal representation of gene functions (GO concepts) and relations between them. However, unintentional quality defects (e.g. missing or erroneous relations) in GO may exist due to the large size of GO concepts and complexity of GO structures. Such quality defects would impact the results of GO-based analyses and applications. In this work, we introduce a novel evidence-based lexical pattern approach for quality assurance of GO relations. We leverage two layers of evidence to suggest potentially missing relations in GO as follows. We first …


Prioritization Of Risk Genes In Multiple Sclerosis By A Refined Bayesian Framework Followed By Tissue-Specificity And Cell Type Feature Assessment, Andi Liu, Astrid M Manuel, Yulin Dai, Zhongming Zhao May 2022

Prioritization Of Risk Genes In Multiple Sclerosis By A Refined Bayesian Framework Followed By Tissue-Specificity And Cell Type Feature Assessment, Andi Liu, Astrid M Manuel, Yulin Dai, Zhongming Zhao

Faculty, Staff and Student Publications

BACKGROUND: Multiple sclerosis (MS) is a debilitating immune-mediated disease of the central nervous system that affects over 2 million people worldwide, resulting in a heavy burden to families and entire communities. Understanding the genetic basis underlying MS could help decipher the pathogenesis and shed light on MS treatment. We refined a recently developed Bayesian framework, Integrative Risk Gene Selector (iRIGS), to prioritize risk genes associated with MS by integrating the summary statistics from the largest GWAS to date (n = 115,803), various genomic features, and gene-gene closeness.

RESULTS: We identified 163 MS-associated prioritized risk genes (MS-PRGenes) through the Bayesian framework. …


Fusionai, A Dna-Sequence-Based Deep Learning Protocol Reduces The False Positives Of Human Fusion Gene Prediction, Pora Kim, Hua Tan, Jiajia Liu, Himansu Kumar, Xiaobo Zhou Mar 2022

Fusionai, A Dna-Sequence-Based Deep Learning Protocol Reduces The False Positives Of Human Fusion Gene Prediction, Pora Kim, Hua Tan, Jiajia Liu, Himansu Kumar, Xiaobo Zhou

Faculty, Staff and Student Publications

Even though there were many tool developments of fusion gene prediction from NGS data, too many false positives are still an issue. Wise use of the genomic features around the fusion gene breakpoints will be helpful to identify reliable fusion genes efficiently. For this aim, we developed FusionAI, a deep learning pipeline predicting human fusion gene breakpoints from DNA sequence. FusionAI is freely available via https://compbio.uth.edu/FusionGDB2/FusionAI. For complete details on the use and execution of this protocol, please refer to Kim et al. (2021b).


Fusiongdb 20: Fusion Gene Annotation Updates Aided By Deep Learning, Pora Kim, Hua Tan, Jiajia Liu, Haeseung Lee, Hyesoo Jung, Himanshu Kumar, Xiaobo Zhou Jan 2022

Fusiongdb 20: Fusion Gene Annotation Updates Aided By Deep Learning, Pora Kim, Hua Tan, Jiajia Liu, Haeseung Lee, Hyesoo Jung, Himanshu Kumar, Xiaobo Zhou

Faculty, Staff and Student Publications

A knowledgebase of the systematic functional annotation of fusion genes is critical for understanding genomic breakage context and developing therapeutic strategies. FusionGDB is a unique functional annotation database of human fusion genes and has been widely used for studies with diverse aims. In this study, we report fusion gene annotation updates aided by deep learning (FusionGDB 2.0) available at https://compbio.uth.edu/FusionGDB2/. FusionGDB 2.0 has substantial updates of contents such as up-to-date human fusion genes, fusion gene breakage tendency score with FusionAI deep learning model based on 20 kb DNA sequence around BP, investigation of overlapping between fusion breakpoints with 44 human …


Facilitating Federated Genomic Data Analysis By Identifying Record Correlations While Ensuring Privacy, Leonard Dervishi, Xinyue Wang, Wentao Li, Anisa Halimi, Jaideep Vaidya, Xiaoqian Jiang, Erman Ayday Jan 2022

Facilitating Federated Genomic Data Analysis By Identifying Record Correlations While Ensuring Privacy, Leonard Dervishi, Xinyue Wang, Wentao Li, Anisa Halimi, Jaideep Vaidya, Xiaoqian Jiang, Erman Ayday

Faculty, Staff and Student Publications

With the reduction of sequencing costs and the pervasiveness of computing devices, genomic data collection is continually growing. However, data collection is highly fragmented and the data is still siloed across different repositories. Analyzing all of this data would be transformative for genomics research. However, the data is sensitive, and therefore cannot be easily centralized. Furthermore, there may be correlations in the data, which if not detected, can impact the analysis. In this paper, we take the first step towards identifying correlated records across multiple data repositories in a privacy-preserving manner. The proposed framework, based on random shuffling, synthetic record …


Causal Inference Of Genetic Variants And Genes In Amyotrophic Lateral Sclerosis, Siyu Pan, Xinxuan Liu, Tianzi Liu, Zhongming Zhao, Yulin Dai, Yin-Ying Wang, Peilin Jia, Fan Liu Jan 2022

Causal Inference Of Genetic Variants And Genes In Amyotrophic Lateral Sclerosis, Siyu Pan, Xinxuan Liu, Tianzi Liu, Zhongming Zhao, Yulin Dai, Yin-Ying Wang, Peilin Jia, Fan Liu

Faculty, Staff and Student Publications

Amyotrophic lateral sclerosis (ALS) is a fatal progressive multisystem disorder with limited therapeutic options. Although genome-wide association studies (GWASs) have revealed multiple ALS susceptibility loci, the exact identities of causal variants, genes, cell types, tissues, and their functional roles in the development of ALS remain largely unknown. Here, we reported a comprehensive post-GWAS analysis of the recent large ALS GWAS (n = 80,610), including functional mapping and annotation (FUMA), transcriptome-wide association study (TWAS), colocalization (COLOC), and summary data-based Mendelian randomization analyses (SMR) in extensive multi-omics datasets. Gene property analysis highlighted inhibitory neuron 6, oligodendrocytes, and GABAergic neurons (Gad1/Gad2) as …


An Autoencoder-Based Deep Learning Method For Genotype Imputation, Meng Song, Jonathan Greenbaum, Joseph Luttrell, Weihua Zhou, Chong Wu, Zhe Luo, Chuan Qiu, Lan Juan Zhao, Kuan-Jui Su, Qing Tian, Hui Shen, Huixiao Hong, Ping Gong, Xinghua Shi, Hong-Wen Deng, Chaoyang Zhang Jan 2022

An Autoencoder-Based Deep Learning Method For Genotype Imputation, Meng Song, Jonathan Greenbaum, Joseph Luttrell, Weihua Zhou, Chong Wu, Zhe Luo, Chuan Qiu, Lan Juan Zhao, Kuan-Jui Su, Qing Tian, Hui Shen, Huixiao Hong, Ping Gong, Xinghua Shi, Hong-Wen Deng, Chaoyang Zhang

Faculty, Staff and Student Publications

Genotype imputation has a wide range of applications in genome-wide association study (GWAS), including increasing the statistical power of association tests, discovering trait-associated loci in meta-analyses, and prioritizing causal variants with fine-mapping. In recent years, deep learning (DL) based methods, such as sparse convolutional denoising autoencoder (SCDA), have been developed for genotype imputation. However, it remains a challenging task to optimize the learning process in DL-based methods to achieve high imputation accuracy. To address this challenge, we have developed a convolutional autoencoder (AE) model for genotype imputation and implemented a customized training loop by modifying the training process with a …


The Ratio Method: Addressing Complex Tort Liability In The Fourth Industrial Revolution, Harrison C. Margolin, Grant H. Frazier Oct 2021

The Ratio Method: Addressing Complex Tort Liability In The Fourth Industrial Revolution, Harrison C. Margolin, Grant H. Frazier

St. Mary's Law Journal

Emerging technologies of the Fourth Industrial Revolution show fundamental promise for improving productivity and quality of life, though their misuse may also cause significant social disruption. For example, while artificial intelligence will be used to accelerate society’s processes, it may also displace millions of workers and arm cybercriminals with increasingly powerful hacking capabilities. Similarly, human gene editing shows promise for curing numerous diseases, but also raises significant concerns about adverse health consequences related to the corruption of human and pathogenic genomes.

In most instances, only specialists understand the growing intricacies of these novel technologies. As the complexity and speed of …


The Concurrence Of Dna Methylation And Demethylation Is Associated With Transcription Regulation, Jiejun Shi, Jianfeng Xu, Yiling Elaine Chen, Jason Sheng Li, Ya Cui, Lanlan Shen, Jingyi Jessica Li, Wei Li Sep 2021

The Concurrence Of Dna Methylation And Demethylation Is Associated With Transcription Regulation, Jiejun Shi, Jianfeng Xu, Yiling Elaine Chen, Jason Sheng Li, Ya Cui, Lanlan Shen, Jingyi Jessica Li, Wei Li

Children’s Nutrition Research Center Staff Publications

The mammalian DNA methylome is formed by two antagonizing processes, methylation by DNA methyltransferases (DNMT) and demethylation by ten-eleven translocation (TET) dioxygenases. Although the dynamics of either methylation or demethylation have been intensively studied in the past decade, the direct effects of their interaction on gene expression remain elusive. Here, we quantify the concurrence of DNA methylation and demethylation by the percentage of unmethylated CpGs within a partially methylated read from bisulfite sequencing. After verifying 'methylation concurrence' by its strong association with the co-localization of DNMT and TET enzymes, we observe that methylation concurrence is strongly correlated with gene expression. …


Epigenome-Wide Association Study Of Kidney Function Identifies Trans-Ethnic And Ethnic-Specific Loci, Charles E. Breeze, Anna Batorsky, Mi Kyeong Lee, Mindy D. Szeto, Xiaoguang Xu, Daniel L. Mccartney, Rong Jiang, Amit Patki, Holly J. Kramer, James M. Eales, Laura Raffield, Leslie Lange, Ethan Lange, Peter Durda, Yongmei Liu, Russ P. Tracy, David Van Den Berg, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Topmed Mesa Multi-Omics Working Group, Kathryn L. Evans, William E. Kraus, Donna K. Arnett Apr 2021

Epigenome-Wide Association Study Of Kidney Function Identifies Trans-Ethnic And Ethnic-Specific Loci, Charles E. Breeze, Anna Batorsky, Mi Kyeong Lee, Mindy D. Szeto, Xiaoguang Xu, Daniel L. Mccartney, Rong Jiang, Amit Patki, Holly J. Kramer, James M. Eales, Laura Raffield, Leslie Lange, Ethan Lange, Peter Durda, Yongmei Liu, Russ P. Tracy, David Van Den Berg, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Topmed Mesa Multi-Omics Working Group, Kathryn L. Evans, William E. Kraus, Donna K. Arnett

Epidemiology and Environmental Health Faculty Publications

BACKGROUND: DNA methylation (DNAm) is associated with gene regulation and estimated glomerular filtration rate (eGFR), a measure of kidney function. Decreased eGFR is more common among US Hispanics and African Americans. The causes for this are poorly understood. We aimed to identify trans-ethnic and ethnic-specific differentially methylated positions (DMPs) associated with eGFR using an agnostic, genome-wide approach.

METHODS: The study included up to 5428 participants from multi-ethnic studies for discovery and 8109 participants for replication. We tested the associations between whole blood DNAm and eGFR using beta values from Illumina 450K or EPIC arrays. Ethnicity-stratified analyses were performed using linear …


Understanding The Effect Of Adaptive Mutations On The Three-Dimensional Structure Of Rna, Justin Cook Apr 2021

Understanding The Effect Of Adaptive Mutations On The Three-Dimensional Structure Of Rna, Justin Cook

Undergraduate Research and Scholarship Symposium

Single-nucleotide polymorphisms (SNPs) are variations in the genome where one base pair can differ between individuals.1 SNPs occur throughout the genome and can correlate to a disease-state if they occur in a functional region of DNA.1According to the central dogma of molecular biology, any variation in the DNA sequence will have a direct effect on the RNA sequence and will potentially alter the identity or conformation of a protein product. A single RNA molecule, due to intramolecular base pairing, can acquire a plethora of 3-D conformations that are described by its structural ensemble. One SNP, rs12477830, which …


A Comparison Of Exhaustive And Non-Lattice-Based Methods For Auditing Hierarchical Relations In Gene Ontology, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui Jan 2021

A Comparison Of Exhaustive And Non-Lattice-Based Methods For Auditing Hierarchical Relations In Gene Ontology, Rashmie Abeysinghe, Fengbo Zheng, Licong Cui

Faculty, Staff and Student Publications

Uncovering and fixing errors in biomedical terminologies is essential so that they provide accurate knowledge to downstream applications that rely on them. Non-lattice-based methods have been applied to identify various kinds of inconsistencies in different biomedical terminologies. In previous work, we have introduced two inference-based approaches that were applied in an exhaustive manner to audit hierarchical relations in the Gene Ontology: (1) Lexical-based inference framework, and (2) Subsumption-based sub-term inference framework. However, it is unclear how effective these exhaustive approaches perform compared with their corresponding non-lattice-based approaches. Therefore, in this paper, we implement the non-lattice versions of these two exhaustive …


Gene Selection For Cancer Classification: A New Hybrid Filter-C5.0 Approach For Breast Cancer Risk Prediction, Mohammed Hamim, Ismail El Moudden, Hicham Moutachaouik, Mustapha Hain Jan 2021

Gene Selection For Cancer Classification: A New Hybrid Filter-C5.0 Approach For Breast Cancer Risk Prediction, Mohammed Hamim, Ismail El Moudden, Hicham Moutachaouik, Mustapha Hain

Department of Medicine Faculty Publications

Despite the significant progress made in data mining technologies in recent years, breast cancer risk prediction and diagnosis at an early stage using DNA microarray technology still a real challenging task. This challenge comes especially from the high-dimensionality in gene expression data, i.e., an enormous number of genes versus a few tens of subjects (samples). To overcome this problem of data imbalance, a gene selection phase becomes a crucial step for gene expression data analysis. This study proposes a new Decision Tree model-based attributes (genes) selection strategy, which incorporates two stages: fisher-score-based filter technique and the gene selection ability of …


A Novel Dimensionality Reduction Approach To Improve Microarray Data Classification, Mohammed Hasim, Ismail El Mouden, Mounir Ouzir, Hicham Moutachaouik, Mustapha Hain Jan 2021

A Novel Dimensionality Reduction Approach To Improve Microarray Data Classification, Mohammed Hasim, Ismail El Mouden, Mounir Ouzir, Hicham Moutachaouik, Mustapha Hain

Department of Medicine Faculty Publications

Cancer tumor prediction and diagnosis at an early stage has become a necessity in cancer research, as it provides an increase in the treatment success chances. Recently, DNA microarray technology became a powerful tool for cancer identification, that can analyze the expression level of a different and huge number of genes simultaneously. In microarray data, the large genes number versus a few records may affect the prediction performance. In order to handle this "curse of dimensionality” constraint of microarray dataset while improving the cancer identification performance, a dimensional reduction phase is necessary. In this paper, we proposed a framework that …


Subject Level Clustering Using A Negative Binomial Model For Small Transcriptomic Studies., Qian Li, Janelle R. Noel-Macdonnell, Devin C. Koestler, Ellen L. Goode, Brooke L. Fridley Dec 2018

Subject Level Clustering Using A Negative Binomial Model For Small Transcriptomic Studies., Qian Li, Janelle R. Noel-Macdonnell, Devin C. Koestler, Ellen L. Goode, Brooke L. Fridley

Manuscripts, Articles, Book Chapters and Other Papers

BACKGROUND: Unsupervised clustering represents one of the most widely applied methods in analysis of high-throughput 'omics data. A variety of unsupervised model-based or parametric clustering methods and non-parametric clustering methods have been proposed for RNA-seq count data, most of which perform well for large samples, e.g. N ≥ 500. A common issue when analyzing limited samples of RNA-seq count data is that the data follows an over-dispersed distribution, and thus a Negative Binomial likelihood model is often used. Thus, we have developed a Negative Binomial model-based (NBMB) clustering approach for application to RNA-seq studies.

RESULTS: We have developed a Negative …


Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry Jan 2018

Penalized Mixed-Effects Ordinal Response Models For High-Dimensional Genomic Data In Twins And Families, Amanda E. Gentry

Theses and Dissertations

The Brisbane Longitudinal Twin Study (BLTS) was being conducted in Australia and was funded by the US National Institute on Drug Abuse (NIDA). Adolescent twins were sampled as a part of this study and surveyed about their substance use as part of the Pathways to Cannabis Use, Abuse and Dependence project. The methods developed in this dissertation were designed for the purpose of analyzing a subset of the Pathways data that includes demographics, cannabis use metrics, personality measures, and imputed genotypes (SNPs) for 493 complete twin pairs (986 subjects.) The primary goal was to determine what combination of SNPs and …