Open Access. Powered by Scholars. Published by Universities.®

Computational Biology Commons

Open Access. Powered by Scholars. Published by Universities.®

Humans

Discipline
Institution
Publication Year
Publication

Articles 1 - 30 of 35

Full-Text Articles in Computational Biology

From Fair To Cure: Guidelines For Computational Models Of Biological Systems, Herbert M. Sauro, Eran Agmon, Michael L. Blinov, John H. Gennari, Joseph L. Hellerstein, Adel Heydarabadipour, Bartholomew E. Jardine, Elebeoba May, David P. Nickerson, Lucian P. Smith, Gary D. Bader, Frank T. Bergmann, Patrick M. Boyle, Andreas Dräger, James R. Faeder, Song Feng, Juliana Freire, Fabian Fröhlich, James A. Glazier, Thomas E. Gorochowski, Tomas Helikar, Henning Hermjakob, Stefan Hoops, Peter Hunter, Princess I. Imoukhuede, Sarah M. Keating, Matthias König, Reinhard Laubenbacher, Leslie M. Loew, Carlos F. Lopez, William W. Lytton, Rahuman S. Malik-Sheriff, Andrew Mcculloch, Pedro Mendes, Lealem Mulugeta, Chris J. Myers, Jerry G. Myers, Anna Niarakis, David D. Van Niekerk, Brett G. Olivier, Alexander A. Patrie, Ellen M. Quardokus, Nicole Radde, Johann M. Rohwer, Sven Sahle, James C. Schaff, Falk Schreiber, T. J. Sego, Janis Shin, Jacky L. Snoep, Rajanikanth Vadigepalli, H. Steven Wiley, Dagmar Waltemath, Ion I. Moraru Mar 2026

From Fair To Cure: Guidelines For Computational Models Of Biological Systems, Herbert M. Sauro, Eran Agmon, Michael L. Blinov, John H. Gennari, Joseph L. Hellerstein, Adel Heydarabadipour, Bartholomew E. Jardine, Elebeoba May, David P. Nickerson, Lucian P. Smith, Gary D. Bader, Frank T. Bergmann, Patrick M. Boyle, Andreas Dräger, James R. Faeder, Song Feng, Juliana Freire, Fabian Fröhlich, James A. Glazier, Thomas E. Gorochowski, Tomas Helikar, Henning Hermjakob, Stefan Hoops, Peter Hunter, Princess I. Imoukhuede, Sarah M. Keating, Matthias König, Reinhard Laubenbacher, Leslie M. Loew, Carlos F. Lopez, William W. Lytton, Rahuman S. Malik-Sheriff, Andrew Mcculloch, Pedro Mendes, Lealem Mulugeta, Chris J. Myers, Jerry G. Myers, Anna Niarakis, David D. Van Niekerk, Brett G. Olivier, Alexander A. Patrie, Ellen M. Quardokus, Nicole Radde, Johann M. Rohwer, Sven Sahle, James C. Schaff, Falk Schreiber, T. J. Sego, Janis Shin, Jacky L. Snoep, Rajanikanth Vadigepalli, H. Steven Wiley, Dagmar Waltemath, Ion I. Moraru

Computational Medicine Center Faculty Papers

Guidelines for managing scientific data have been established under the FAIR principles, requiring that data be Findable, Accessible, Interoperable, and Reusable. In many scientific disciplines, especially computational biology, both data and models are key to progress. For this reason, and recognizing that such models are a very special type of "data", we argue that computational models, especially mechanistic models prevalent in medicine, physiology and systems biology, deserve a complementary set of guidelines. We propose the CURE principles, emphasizing that models should be Credible, Understandable, Reproducible, and Extensible. We delve into each principle, discussing verification, validation, and uncertainty quantification for model …


Organism-Specific Sequence Motifs Link Ribosomal Rnas To Brain Disorders, Isidore Rigoutsos, Stepan Nersisyan, Eric Londin, Iliza Nazeraj, Bonnie Dong, Anastasios Vourekas, Phillipe Loher Oct 2025

Organism-Specific Sequence Motifs Link Ribosomal Rnas To Brain Disorders, Isidore Rigoutsos, Stepan Nersisyan, Eric Londin, Iliza Nazeraj, Bonnie Dong, Anastasios Vourekas, Phillipe Loher

Computational Medicine Center Faculty Papers

We report that in humans, mice, fruit flies, and worms, the ribosomal RNAs and the transcribed spacers of 45S are densely packed with organism-specific sequence motifs that are primarily shared with nervous system genes. The human ribosomal RNAs and 45S spacers contain 1,723 such motifs. Specific combinations of these motifs are predominantly found in 3,430 human nervous system genes, of which 1,046 are genes associated with brain disorders, including autism spectrum disorder and schizophrenia. The sequences of the 1,723 motifs and their locations in the introns and exons of nervous system genes are unique to primates. Experimental evidence indicates that …


Quantile Index Predictors Using R Package Hyper.Gam, Tingting Zhan, Misung Yi, Inna Chervoneva Aug 2025

Quantile Index Predictors Using R Package Hyper.Gam, Tingting Zhan, Misung Yi, Inna Chervoneva

Department of Pharmacology, Physiology, and Cancer Biology Faculty Papers

MOTIVATION: Evaluation of single-cell protein expression from immunohistochemistry images is used increasingly in biomedical research. Many proteins are used solely for phenotyping cells in the tumor microenvironment. Other proteins with meaningfully quantitative expression levels provide so-called functional protein biomarkers. There is still a limited number of methods and software tools available for utilizing the entire distributions of single-cell expression levels.

RESULTS: We present the R package hyper.gam, providing a supervised learning framework for deriving biomarkers based on single-cell distribution quantiles. The single-cell data are first converted into sample quantile functions, which are then used as predictors in scalar-on-function regression models …


A Rubric For Assessing Conformance To The Ten Rules For Credible Practice Of Modeling And Simulation In Healthcare, Alexandra Manchel, Ahmet Erdemir, Lealem Mulugeta, Joy Ku, Bruno Rego, Marc Horner, William Lytton, Jerry Myers, Rajanikanth Vadigepalli Jun 2025

A Rubric For Assessing Conformance To The Ten Rules For Credible Practice Of Modeling And Simulation In Healthcare, Alexandra Manchel, Ahmet Erdemir, Lealem Mulugeta, Joy Ku, Bruno Rego, Marc Horner, William Lytton, Jerry Myers, Rajanikanth Vadigepalli

Computational Medicine Center Faculty Papers

The power of computational modeling and simulation (M&S) is realized when the results are credible, and the workflow generates evidence that supports credibility for the context of use. The Committee on Credible Practice of Modeling & Simulation in Healthcare was established to help address the need for processes and procedures to support the credible use of M&S in healthcare and biomedical research. Our community efforts have led to the Ten Rules (TR) for Credible Practice of M&S in life sciences and healthcare. This framework is an outcome of a multidisciplinary investigation from a wide range of stakeholders beginning in 2012. …


Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh Jan 2025

Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh

Computer Science Faculty Publications

Drug–target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug–target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance …


Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang Jan 2025

Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang

Computer Science Faculty Publications

Motivation: Protein-protein interactions (PPIs) are fundamental aspects in understanding biological processes. Accurately predicting the effects of mutations on PPIs remains a critical requirement for drug design and disease mechanistic studies. Recently, deep learning models using protein 3D structures have become predominant for predicting mutation effects. However, significant challenges remain in practical applications, in part due to the considerable disparity in generalization capabilities between easy and hard mutations. Specifically, a hard mutation is defined as one with its maximum TM-score < 0.6 when compared to the training set. Additionally, compared to physics-based approaches, deep learning models may overestimate performance due to potential data leakage.

Results: We propose new training/test splits that mitigate data leakage according to the CATH homologous superfamily. Under the constraints of physical …


A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh Jan 2025

A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh

Computer Science Faculty Publications

Conventional drug discovery is expensive, time-consuming, and prone to failure. Artificial intelligence has become a potent substitute over the last decade, providing strong answers to challenging biological issues in this field. Among these difficulties, drug-target binding (DTB) is a key component of drug discovery techniques. In this context, drug-target affinity and drug–target interaction are complementary and essential frameworks that work together to improve our comprehension of DTB dynamics. In this work, we thoroughly analyze the most recent deep learning models, popular benchmark datasets, and assessment metrics for DTB prediction. We look at the paradigm shift in the development of drug …


The Trnaval Half: A Strong Endogenous Toll-Like Receptor 7 Ligand With A 5′-Terminal Universal Sequence Signature, Kamlesh Ganesh Pawar, Takuya Kawamura, Yohei Kirino May 2024

The Trnaval Half: A Strong Endogenous Toll-Like Receptor 7 Ligand With A 5′-Terminal Universal Sequence Signature, Kamlesh Ganesh Pawar, Takuya Kawamura, Yohei Kirino

Computational Medicine Center Faculty Papers

Toll-like receptors (TLRs) are crucial components of the innate immune system. Endosomal TLR7 recognizes single-stranded RNAs, yet its endogenous ssRNA ligands are not fully understood. We previously showed that extracellular (ex-) 5'-half molecules of tRNAHisGUG (the 5'-tRNAHisGUG half) in extracellular vesicles (EVs) of human macrophages activate TLR7 when delivered into endosomes of recipient macrophages. Here, we fully explored immunostimulatory ex-5'-tRNA half molecules and identified the 5'-tRNAValCAC/AAC half, the most abundant tRNA-derived RNA in macrophage EVs, as another 5'-tRNA half molecule with strong TLR7 activation capacity. Levels of the ex-5'-tRNAValCAC/AAC half were highly up-regulated in macrophage EVs …


Integrated Transcriptomics And Histopathology Approach Identifies A Subset Of Rejected Donor Livers With Potential Suitability For Transplantation, Ankita Srivastava, Alexandra Manchel, John Waters, Manju Ambelil, Benjamin K. Barnhart, Jan B. Hoek, Ashesh P. Shah, Rajanikanth Vadigepalli May 2024

Integrated Transcriptomics And Histopathology Approach Identifies A Subset Of Rejected Donor Livers With Potential Suitability For Transplantation, Ankita Srivastava, Alexandra Manchel, John Waters, Manju Ambelil, Benjamin K. Barnhart, Jan B. Hoek, Ashesh P. Shah, Rajanikanth Vadigepalli

Department of Pathology, Anatomy, and Cell Biology Faculty Papers

BACKGROUND: Liver transplantation is an effective treatment for liver failure. There is a large unmet demand, even as not all donated livers are transplanted. The clinical selection criteria for donor livers based on histopathological evaluation and liver function tests are variable. We integrated transcriptomics and histopathology to characterize donor liver biopsies obtained at the time of organ recovery. We performed RNA sequencing as well as manual and artificial intelligence-based histopathology (10 accepted and 21 rejected for transplantation).

RESULTS: We identified two transcriptomically distinct rejected subsets (termed rejected-1 and rejected-2), where rejected-2 exhibited a near-complete transcriptomic overlap with the accepted livers, …


Identifying New Cancer Genes Based On The Integration Of Annotated Gene Sets Via Hypergraph Neural Networks, Chao Deng, Hong-Dong Li, Li-Shen Zhang, Yiwei Liu, Yaohang Li, Jianxin Wang Jan 2024

Identifying New Cancer Genes Based On The Integration Of Annotated Gene Sets Via Hypergraph Neural Networks, Chao Deng, Hong-Dong Li, Li-Shen Zhang, Yiwei Liu, Yaohang Li, Jianxin Wang

Computer Science Faculty Publications

Motivation

Identifying cancer genes remains a significant challenge in cancer genomics research. Annotated gene sets encode functional associations among multiple genes, and cancer genes have been shown to cluster in hallmark signaling pathways and biological processes. The knowledge of annotated gene sets is critical for discovering cancer genes but remains to be fully exploited.

Results

Here, we present the DIsease-Specific Hypergraph neural network (DISHyper), a hypergraph-based computational method that integrates the knowledge from multiple types of annotated gene sets to predict cancer genes. First, our benchmark results demonstrate that DISHyper outperforms the existing state-of-the-art methods and highlight the advantages of …


Sccad: Cluster Decomposition-Based Anomaly Detection For Rare Cell Identification In Single-Cell Expression Data, Yunpei Xu, Shaokai Wang, Qilong Feng, Jiazhi Xia, Yaohang Li, Hong-Dong Li, Jianxin Wang Jan 2024

Sccad: Cluster Decomposition-Based Anomaly Detection For Rare Cell Identification In Single-Cell Expression Data, Yunpei Xu, Shaokai Wang, Qilong Feng, Jiazhi Xia, Yaohang Li, Hong-Dong Li, Jianxin Wang

Computer Science Faculty Publications

Single-cell RNA sequencing (scRNA-seq) technologies have become essential tools for characterizing cellular landscapes within complex tissues. Large-scale single-cell transcriptomics holds great potential for identifying rare cell types critical to the pathogenesis of diseases and biological processes. Existing methods for identifying rare cell types often rely on one-time clustering using partial or global gene expression. However, these rare cell types may be overlooked during the clustering phase, posing challenges for their accurate identification. In this paper, we propose a Cluster decomposition-based Anomaly Detection method (scCAD), which iteratively decomposes clusters based on the most differential signals in each cluster to effectively separate …


Radiation Exposure Determination In A Secure, Cloud-Based Online Environment, Ben C. Shirley, Eliseos J. Mucaki, Peter Rogan Oct 2022

Radiation Exposure Determination In A Secure, Cloud-Based Online Environment, Ben C. Shirley, Eliseos J. Mucaki, Peter Rogan

Biochemistry Publications

Rapid sample processing and interpretation of estimated exposures will be critical for triaging exposed individuals after a major radiation incident. The dicentric chromosome (DC) assay assesses absorbed radiation using metaphase cells from blood. The Automated Dicentric Chromosome Identifier and Dose Estimator System (ADCI) identifies DCs and determines radiation doses. This study aimed to broaden accessibility and speed of this system, while protecting data and software integrity. ADCI Online is a secure web-streaming platform accessible worldwide from local servers. Cloud-based systems containing data and software are separated until they are linked for radiation exposure estimation. Dose estimates are identical to ADCI …


Rare Coding Variants In 35 Genes Associate With Circulating Lipid Levels-A Multi-Ancestry Analysis Of 170,000 Exomes, George Hindy, Peter Dornbos, Mark D Chaffin, Dajiang J Liu, Minxian Wang, Margaret Sunitha Selvaraj, David Zhang, Joseph Park, Carlos A Aguilar-Salinas, Lucinda Antonacci-Fulton, Diego Ardissino, Donna K Arnett, Stella Aslibekyan, Gil Atzmon, Christie M Ballantyne, Francisco Barajas-Olmos, Nir Barzilai, Lewis C Becker, Lawrence F Bielak, Joshua C Bis, John Blangero, Eric Boerwinkle, Lori L Bonnycastle, Erwin Bottinger, Donald W Bowden, Matthew J Bown, Jennifer A Brody, Jai G Broome, Noël P Burtt, Brian E Cade, Federico Centeno-Cruz, Edmund Chan, Yi-Cheng Chang, Yii-Der I Chen, Ching-Yu Cheng, Won Jung Choi, Rajiv Chowdhury, Cecilia Contreras-Cubas, Emilio J Córdova, Adolfo Correa, L Adrienne Cupples, Joanne E Curran, John Danesh, Paul S De Vries, Ralph A Defronzo, Harsha Doddapaneni, Ravindranath Duggirala, Susan K Dutcher, Patrick T Ellinor, Leslie S Emery, Jose C Florez, Myriam Fornage, Barry I Freedman, Valentin Fuster, Ma Eugenia Garay-Sevilla, Humberto García-Ortiz, Soren Germer, Richard A Gibbs, Christian Gieger, Benjamin Glaser, Clicerio Gonzalez, Maria Elena Gonzalez-Villalpando, Mariaelisa Graff, Sarah E Graham, Niels Grarup, Leif C Groop, Xiuqing Guo, Namrata Gupta, Sohee Han, Craig L Hanis, Torben Hansen, Jiang He, Nancy L Heard-Costa, Yi-Jen Hung, Mi Yeong Hwang, Marguerite R Irvin, Sergio Islas-Andrade, Gail P Jarvik, Hyun Min Kang, Sharon L R Kardia, Tanika Kelly, Eimear E Kenny, Alyna T Khan, Bong-Jo Kim, Ryan W Kim, Young Jin Kim, Heikki A Koistinen, Charles Kooperberg, Johanna Kuusisto, Soo Heon Kwak, Markku Laakso, Leslie A Lange, Jiwon Lee, Juyoung Lee, Seonwook Lee, Donna M Lehman, Rozenn N Lemaitre, Allan Linneberg, Jianjun Liu, Ruth J F Loos, Steven A Lubitz, Valeriya Lyssenko, Ronald C W Ma, Lisa Warsinger Martin, Angélica Martínez-Hernández, Rasika A Mathias, Stephen T Mcgarvey, Ruth Mcpherson, James B Meigs, Thomas Meitinger, Olle Melander, Elvia Mendoza-Caamal, Ginger A Metcalf, Xuenan Mi, Karen L Mohlke, May E Montasser, Jee-Young Moon, Hortensia Moreno-Macías, Alanna C Morrison, Donna M Muzny, Sarah C Nelson, Peter M Nilsson, Jeffrey R O'Connell, Marju Orho-Melander, Lorena Orozco, Colin N A Palmer, Nicholette D Palmer, Cheol Joo Park, Kyong Soo Park, Oluf Pedersen, Juan M Peralta, Patricia A Peyser, Wendy S Post, Michael Preuss, Bruce M Psaty, Qibin Qi, D C Rao, Susan Redline, Alexander P Reiner, Cristina Revilla-Monsalve, Stephen S Rich, Nilesh Samani, Heribert Schunkert, Claudia Schurmann, Daekwan Seo, Jeong-Sun Seo, Xueling Sim, Rob Sladek, Kerrin S Small, Wing Yee So, Adrienne M Stilp, E Shyong Tai, Claudia H T Tam, Kent D Taylor, Yik Ying Teo, Farook Thameem, Brian Tomlinson, Michael Y Tsai, Tiinamaija Tuomi, Jaakko Tuomilehto, Teresa Tusié-Luna, Miriam S Udler, Rob M Van Dam, Ramachandran S Vasan, Karine A Viaud Martinez, Fei Fei Wang, Xuzhi Wang, Hugh Watkins, Daniel E Weeks, James G Wilson, Daniel R Witte, Tien-Yin Wong, Lisa R Yanek, Amp-T2d-Genes, Myocardial Infarction Genetics Consortium, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Nhlbi Topmed Lipids Working Group, Sekar Kathiresan, Daniel J Rader, Jerome I Rotter, Michael Boehnke, Mark I Mccarthy, Cristen J Willer, Pradeep Natarajan, Jason A Flannick, Amit V Khera, Gina M Peloso Jan 2022

Rare Coding Variants In 35 Genes Associate With Circulating Lipid Levels-A Multi-Ancestry Analysis Of 170,000 Exomes, George Hindy, Peter Dornbos, Mark D Chaffin, Dajiang J Liu, Minxian Wang, Margaret Sunitha Selvaraj, David Zhang, Joseph Park, Carlos A Aguilar-Salinas, Lucinda Antonacci-Fulton, Diego Ardissino, Donna K Arnett, Stella Aslibekyan, Gil Atzmon, Christie M Ballantyne, Francisco Barajas-Olmos, Nir Barzilai, Lewis C Becker, Lawrence F Bielak, Joshua C Bis, John Blangero, Eric Boerwinkle, Lori L Bonnycastle, Erwin Bottinger, Donald W Bowden, Matthew J Bown, Jennifer A Brody, Jai G Broome, Noël P Burtt, Brian E Cade, Federico Centeno-Cruz, Edmund Chan, Yi-Cheng Chang, Yii-Der I Chen, Ching-Yu Cheng, Won Jung Choi, Rajiv Chowdhury, Cecilia Contreras-Cubas, Emilio J Córdova, Adolfo Correa, L Adrienne Cupples, Joanne E Curran, John Danesh, Paul S De Vries, Ralph A Defronzo, Harsha Doddapaneni, Ravindranath Duggirala, Susan K Dutcher, Patrick T Ellinor, Leslie S Emery, Jose C Florez, Myriam Fornage, Barry I Freedman, Valentin Fuster, Ma Eugenia Garay-Sevilla, Humberto García-Ortiz, Soren Germer, Richard A Gibbs, Christian Gieger, Benjamin Glaser, Clicerio Gonzalez, Maria Elena Gonzalez-Villalpando, Mariaelisa Graff, Sarah E Graham, Niels Grarup, Leif C Groop, Xiuqing Guo, Namrata Gupta, Sohee Han, Craig L Hanis, Torben Hansen, Jiang He, Nancy L Heard-Costa, Yi-Jen Hung, Mi Yeong Hwang, Marguerite R Irvin, Sergio Islas-Andrade, Gail P Jarvik, Hyun Min Kang, Sharon L R Kardia, Tanika Kelly, Eimear E Kenny, Alyna T Khan, Bong-Jo Kim, Ryan W Kim, Young Jin Kim, Heikki A Koistinen, Charles Kooperberg, Johanna Kuusisto, Soo Heon Kwak, Markku Laakso, Leslie A Lange, Jiwon Lee, Juyoung Lee, Seonwook Lee, Donna M Lehman, Rozenn N Lemaitre, Allan Linneberg, Jianjun Liu, Ruth J F Loos, Steven A Lubitz, Valeriya Lyssenko, Ronald C W Ma, Lisa Warsinger Martin, Angélica Martínez-Hernández, Rasika A Mathias, Stephen T Mcgarvey, Ruth Mcpherson, James B Meigs, Thomas Meitinger, Olle Melander, Elvia Mendoza-Caamal, Ginger A Metcalf, Xuenan Mi, Karen L Mohlke, May E Montasser, Jee-Young Moon, Hortensia Moreno-Macías, Alanna C Morrison, Donna M Muzny, Sarah C Nelson, Peter M Nilsson, Jeffrey R O'Connell, Marju Orho-Melander, Lorena Orozco, Colin N A Palmer, Nicholette D Palmer, Cheol Joo Park, Kyong Soo Park, Oluf Pedersen, Juan M Peralta, Patricia A Peyser, Wendy S Post, Michael Preuss, Bruce M Psaty, Qibin Qi, D C Rao, Susan Redline, Alexander P Reiner, Cristina Revilla-Monsalve, Stephen S Rich, Nilesh Samani, Heribert Schunkert, Claudia Schurmann, Daekwan Seo, Jeong-Sun Seo, Xueling Sim, Rob Sladek, Kerrin S Small, Wing Yee So, Adrienne M Stilp, E Shyong Tai, Claudia H T Tam, Kent D Taylor, Yik Ying Teo, Farook Thameem, Brian Tomlinson, Michael Y Tsai, Tiinamaija Tuomi, Jaakko Tuomilehto, Teresa Tusié-Luna, Miriam S Udler, Rob M Van Dam, Ramachandran S Vasan, Karine A Viaud Martinez, Fei Fei Wang, Xuzhi Wang, Hugh Watkins, Daniel E Weeks, James G Wilson, Daniel R Witte, Tien-Yin Wong, Lisa R Yanek, Amp-T2d-Genes, Myocardial Infarction Genetics Consortium, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Nhlbi Topmed Lipids Working Group, Sekar Kathiresan, Daniel J Rader, Jerome I Rotter, Michael Boehnke, Mark I Mccarthy, Cristen J Willer, Pradeep Natarajan, Jason A Flannick, Amit V Khera, Gina M Peloso

Faculty, Staff and Student Publications

Large-scale gene sequencing studies for complex traits have the potential to identify causal genes with therapeutic implications. We performed gene-based association testing of blood lipid levels with rare (minor allele frequency < 1%) predicted damaging coding variation by using sequence data from >170,000 individuals from multiple ancestries: 97,493 European, 30,025 South Asian, 16,507 African, 16,440 Hispanic/Latino, 10,420 East Asian, and 1,182 Samoan. We identified 35 genes associated with circulating lipid levels; some of these genes have not been previously associated with lipid levels when using rare coding variation from population-based samples. We prioritize 32 genes in array-based genome-wide association study (GWAS) loci based on aggregations of rare coding variants; three (EVI5, …


Analysis Of Subtelomeric Rextal Assemblies Using Quast, Tunazzina Islam, Desh Ranjan, Mohammad Zubair, Eleanor Young, Ming Xiao, Harold Riethman Jan 2021

Analysis Of Subtelomeric Rextal Assemblies Using Quast, Tunazzina Islam, Desh Ranjan, Mohammad Zubair, Eleanor Young, Ming Xiao, Harold Riethman

Computer Science Faculty Publications

Genomic regions of high segmental duplication content and/or structural variation have led to gaps and misassemblies in the human reference sequence, and are refractory to assembly from whole-genome short-read datasets. Human subtelomere regions are highly enriched in both segmental duplication content and structural variations, and as a consequence are both impossible to assemble accurately and highly variable from individual to individual. Recently, we developed a pipeline for improved region-specific assembly called Regional Extension of Assemblies Using Linked-Reads (REXTAL). In this study, we evaluate REXTAL and genome-wide assembly (Supernova) approaches on 10X Genomics linked-reads data sets partitioned and barcoded using the …


Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang Apr 2019

Incorporating Pathway Information Into Feature Selection Towards Better Performed Gene Signatures, Suyan Tian, Chi Wang, Bing Wang

Biostatistics Faculty Publications

To analyze gene expression data with sophisticated grouping structures and to extract hidden patterns from such data, feature selection is of critical importance. It is well known that genes do not function in isolation but rather work together within various metabolic, regulatory, and signaling pathways. If the biological knowledge contained within these pathways is taken into account, the resulting method is a pathway-based algorithm. Studies have demonstrated that a pathway-based method usually outperforms its gene-based counterpart in which no biological knowledge is considered. In this article, a pathway-based feature selection is firstly divided into three major categories, namely, pathway-level selection, …


A Longitudinal Cline Characterizes The Genetic Structure Of Human Populations In The Tibetan Plateau, Choongwon Jeong, Benjamin M. Peter, Buddha Basnyat, Maniraj Neupane, Geoff Childs, Sienna Craig, John Novembre, Anna Di Rienzo Apr 2017

A Longitudinal Cline Characterizes The Genetic Structure Of Human Populations In The Tibetan Plateau, Choongwon Jeong, Benjamin M. Peter, Buddha Basnyat, Maniraj Neupane, Geoff Childs, Sienna Craig, John Novembre, Anna Di Rienzo

Dartmouth Scholarship

Indigenous populations of the Tibetan plateau have attracted much attention for their good performance at extreme high altitude. Most genetic studies of Tibetan adaptations have used genetic variation data at the genome scale, while genetic inferences about their de- mography and population structure are largely based on uniparental markers. To provide genome-wide information on population structure, we analyzed new and published data of 338 individuals from indigenous populations across the plateau in conjunction with world- wide genetic variation data. We found a clear signal of genetic stratification across the east- west axis within Tibetan samples. Samples from more eastern locations …


Discovery And Validation Of Information Theory-Based Transcription Factor And Cofactor Binding Site Motifs., Ruipeng Lu, Eliseos J Mucaki, Peter K Rogan Mar 2017

Discovery And Validation Of Information Theory-Based Transcription Factor And Cofactor Binding Site Motifs., Ruipeng Lu, Eliseos J Mucaki, Peter K Rogan

Biochemistry Publications

Data from ChIP-seq experiments can derive the genome-wide binding specificities of transcription factors (TFs) and other regulatory proteins. We analyzed 765 ENCODE ChIP-seq peak datasets of 207 human TFs with a novel motif discovery pipeline based on recursive, thresholded entropy minimization. This approach, while obviating the need to compensate for skewed nucleotide composition, distinguishes true binding motifs from noise, quantifies the strengths of individual binding sites based on computed affinity and detects adjacent cofactor binding sites that coordinate with the targets of primary, immunoprecipitated TFs. We obtained contiguous and bipartite information theory-based position weight matrices (iPWMs) for 93 sequence-specific TFs, …


Leveraging Global Gene Expression Patterns To Predict Expression Of Unmeasured Genes, James Rudd, René A. Zelaya, Eugene Demidenko, Ellen L. Goode, Casey S. Greene S. Greene, Jennifer A. Doherty Dec 2015

Leveraging Global Gene Expression Patterns To Predict Expression Of Unmeasured Genes, James Rudd, René A. Zelaya, Eugene Demidenko, Ellen L. Goode, Casey S. Greene S. Greene, Jennifer A. Doherty

Dartmouth Scholarship

BackgroundLarge collections of paraffin-embedded tissue represent a rich resource to test hypotheses based on gene expression patterns; however, measurement of genome-wide expression is cost-prohibitive on a large scale. Using the known expression correlation structure within a given disease type (in this case, high grade serous ovarian cancer; HGSC), we sought to identify reduced sets of directly measured (DM) genes which could accurately predict the expression of a maximized number of unmeasured genes.


Loregic: A Method To Characterize The Cooperative Logic Of Regulatory Factors, Daifeng Wang, Koon-Kiu Yan, Cristina Sisu, Chao Cheng, Joel Rozowsky, William Meyerson, Mark B. Gerstein Apr 2015

Loregic: A Method To Characterize The Cooperative Logic Of Regulatory Factors, Daifeng Wang, Koon-Kiu Yan, Cristina Sisu, Chao Cheng, Joel Rozowsky, William Meyerson, Mark B. Gerstein

Dartmouth Scholarship

The topology of the gene-regulatory network has been extensively analyzed. Now, given the large amount of available functional genomic data, it is possible to go beyond this and systematically study regulatory circuits in terms of logic elements. To this end, we present Loregic, a computational method integrating gene expression and regulatory network data, to characterize the cooperativity of regulatory factors. Loregic uses all 16 possible two-input-one-output logic gates (e.g. AND or XOR) to describe triplets of two factors regulating a common target. We attempt to find the gate that best matches each triplet’s observed gene expression pattern across many conditions. …


Machine Learning Methods Enable Predictive Modeling Of Antibody Feature:Function Relationships In Rv144 Vaccinees, Ickwon Choi, Amy W. Chung, Todd J. Suscovich, Supachai Rerks-Ngarm, Punnee Pitisuttithum, Sorachai Nitayapha, Jaranit Kaewkungwal, Robert J. O'Connell, Donald Francis, Merlin L. Robb, Nelson L. Michael, Jerome H. Kim, Galit Alter, Margaret E. Ackerman, Chris Bailey-Kellogg Apr 2015

Machine Learning Methods Enable Predictive Modeling Of Antibody Feature:Function Relationships In Rv144 Vaccinees, Ickwon Choi, Amy W. Chung, Todd J. Suscovich, Supachai Rerks-Ngarm, Punnee Pitisuttithum, Sorachai Nitayapha, Jaranit Kaewkungwal, Robert J. O'Connell, Donald Francis, Merlin L. Robb, Nelson L. Michael, Jerome H. Kim, Galit Alter, Margaret E. Ackerman, Chris Bailey-Kellogg

Dartmouth Scholarship

The adaptive immune response to vaccination or infection can lead to the production of specific antibodies to neutralize the pathogen or recruit innate immune effector cells for help. The non-neutralizing role of antibodies in stimulating effector cell responses may have been a key mechanism of the protection observed in the RV144 HIV vaccine trial. In an extensive investigation of a rich set of data collected from RV144 vaccine recipients, we here employ machine learning methods to identify and model associations between antibody features (IgG subclass and antigen specificity) and effector function activities (antibody dependent cellular phagocytosis, cellular cytotoxicity, and cytokine …


Modeling Neurovascular Coupling From Clustered Parameter Sets For Multimodal Eeg-Nirs, M. Tanveer Talukdar, H. Robert Frost, Solomon G. G. Diamond Feb 2015

Modeling Neurovascular Coupling From Clustered Parameter Sets For Multimodal Eeg-Nirs, M. Tanveer Talukdar, H. Robert Frost, Solomon G. G. Diamond

Dartmouth Scholarship

Despite significant improvements in neuroimaging technologies and analysis methods, the fundamental relationship between local changes in cerebral hemodynamics and the underlying neural activity remains largely unknown. In this study, a data driven approach is proposed for modeling this neurovascular coupling relationship from simultaneously acquired electroencephalographic (EEG) and near-infrared spectroscopic (NIRS) data. The approach uses gamma transfer functions to map EEG spectral envelopes that reflect time-varying power variations in neural rhythms to hemodynamics measured with NIRS during median nerve stimulation. The approach is evaluated first with simulated EEG-NIRS data and then by applying the method to experimental EEG-NIRS data measured from …


A Quick Guide For Building A Successful Bioinformatics Community., Aidan Budd, Manuel Corpas, Michelle D Brazas, Jonathan C Fuller, Jeremy Goecks, Nicola J Mulder, Magali Michaut, B F Francis Ouellette, Aleksandra Pawlik, Niklas Blomberg Feb 2015

A Quick Guide For Building A Successful Bioinformatics Community., Aidan Budd, Manuel Corpas, Michelle D Brazas, Jonathan C Fuller, Jeremy Goecks, Nicola J Mulder, Magali Michaut, B F Francis Ouellette, Aleksandra Pawlik, Niklas Blomberg

Computational Biology Institute

"Scientific community" refers to a group of people collaborating together on scientific-research-related activities who also share common goals, interests, and values. Such communities play a key role in many bioinformatics activities. Communities may be linked to a specific location or institute, or involve people working at many different institutions and locations. Education and training is typically an important component of these communities, providing a valuable context in which to develop skills and expertise, while also strengthening links and relationships within the community. Scientific communities facilitate: (i) the exchange and development of ideas and expertise; (ii) career development; (iii) coordinated funding …


Mapping The Pareto Optimal Design Space For A Functionally Deimmunized Biotherapeutic Candidate, Regina S. Salvat, Andrew S. Parker, Yoonjoo Choi, Chris Bailey-Kellogg, Karl E. Griswold Jan 2015

Mapping The Pareto Optimal Design Space For A Functionally Deimmunized Biotherapeutic Candidate, Regina S. Salvat, Andrew S. Parker, Yoonjoo Choi, Chris Bailey-Kellogg, Karl E. Griswold

Dartmouth Scholarship

The immunogenicity of biotherapeutics can bottleneck development pipelines and poses a barrier to widespread clinical application. As a result, there is a growing need for improved deimmunization technologies. We have recently described algorithms that simultaneously optimize proteins for both reduced T cell epitope content and high-level function. In silico analysis of this dual objective design space reveals that there is no single global optimum with respect to protein deimmunization. Instead, mutagenic epitope deletion yields a spectrum of designs that exhibit tradeoffs between immunogenic potential and molecular function. The leading edge of this design space is the Pareto frontier, i.e. the …


Systems Level Analysis Of Systemic Sclerosis Shows A Network Of Immune And Profibrotic Pathways Connected With Genetic Polymorphisms, J. Matthew Mahoney, Jaclyn Taroni, Viktor Martyanov, Tammara A. A. Wood, Casey S. Greene, Patricia A. Pioli, Monique E. Hinchcliff, Michael L. Whitfield Jan 2015

Systems Level Analysis Of Systemic Sclerosis Shows A Network Of Immune And Profibrotic Pathways Connected With Genetic Polymorphisms, J. Matthew Mahoney, Jaclyn Taroni, Viktor Martyanov, Tammara A. A. Wood, Casey S. Greene, Patricia A. Pioli, Monique E. Hinchcliff, Michael L. Whitfield

Dartmouth Scholarship

Systemic sclerosis (SSc) is a rare systemic autoimmune disease characterized by skin and organ fibrosis. The pathogenesis of SSc and its progression are poorly understood. The SSc intrinsic gene expression subsets (inflammatory, fibroproliferative, normal-like, and limited) are observed in multiple clinical cohorts of patients with SSc. Analysis of longitudinal skin biopsies suggests that a patient's subset assignment is stable over 6-12 months. Genetically, SSc is multi-factorial with many genetic risk loci for SSc generally and for specific clinical manifestations. Here we identify the genes consistently associated with the intrinsic subsets across three independent cohorts, show the relationship between these genes …


Trail-Based High Throughput Screening Reveals A Link Between Trail-Mediated Apoptosis And Glutathione Reductase, A Key Component Of Oxidative Stress Response., Dmitri Rozanov, Anton Cheltsov, Eduard Sergienko, Stefan Vasile, Vladislav Golubkov, Alexander E Aleshin, Trevor Levin, Elie Traer, Byron Hann, Julia Freimuth, Nikita Alexeev, Max A Alekseyev, Sergey P Budko, Hans Peter Bächinger, Paul Spellman Jan 2015

Trail-Based High Throughput Screening Reveals A Link Between Trail-Mediated Apoptosis And Glutathione Reductase, A Key Component Of Oxidative Stress Response., Dmitri Rozanov, Anton Cheltsov, Eduard Sergienko, Stefan Vasile, Vladislav Golubkov, Alexander E Aleshin, Trevor Levin, Elie Traer, Byron Hann, Julia Freimuth, Nikita Alexeev, Max A Alekseyev, Sergey P Budko, Hans Peter Bächinger, Paul Spellman

Computational Biology Institute

A high throughput screen for compounds that induce TRAIL-mediated apoptosis identified ML100 as an active chemical probe, which potentiated TRAIL activity in prostate carcinoma PPC-1 and melanoma MDA-MB-435 cells. Follow-up in silico modeling and profiling in cell-based assays allowed us to identify NSC130362, pharmacophore analog of ML100 that induced 65-95% cytotoxicity in cancer cells and did not affect the viability of human primary hepatocytes. In agreement with the activation of the apoptotic pathway, both ML100 and NSC130362 synergistically with TRAIL induced caspase-3/7 activity in MDA-MB-435 cells. Subsequent affinity chromatography and inhibition studies convincingly demonstrated that glutathione reductase (GSR), a key …


Integrated Assessment Of Predicted Mhc Binding And Cross-Conservation With Self Reveals Patterns Of Viral Camouflage, Lu He, Anne S. De Groot, Andres H. Gutierrez, William D. Martin, Lenny Moise, Chris Bailey-Kellogg Mar 2014

Integrated Assessment Of Predicted Mhc Binding And Cross-Conservation With Self Reveals Patterns Of Viral Camouflage, Lu He, Anne S. De Groot, Andres H. Gutierrez, William D. Martin, Lenny Moise, Chris Bailey-Kellogg

Dartmouth Scholarship

Immune recognition of foreign proteins by T cells hinges on the formation of a ternary complex sandwiching a constituent peptide of the protein between a major histocompatibility complex (MHC) molecule and a T cell receptor (TCR). Viruses have evolved means of "camouflaging" themselves, avoiding immune recognition by reducing the MHC and/or TCR binding of their constituent peptides. Computer-driven T cell epitope mapping tools have been used to evaluate the degree to which articular viruses have used this means of avoiding immune response, but most such analyses focus on MHC-facing ‘agretopes'. Here we set out a new means of evaluating the …


Identifying Potential Cancer Driver Genes By Genomic Data Integration., Yong Chen, Jingjing Hao, Wei Jiang, Tong He, Xuegong Zhang, Tao Jiang, Rui Jiang Dec 2013

Identifying Potential Cancer Driver Genes By Genomic Data Integration., Yong Chen, Jingjing Hao, Wei Jiang, Tong He, Xuegong Zhang, Tao Jiang, Rui Jiang

College of Science & Mathematics Departmental Research

Cancer is a genomic disease associated with a plethora of gene mutations resulting in a loss of control over vital cellular functions. Among these mutated genes, driver genes are defined as being causally linked to oncogenesis, while passenger genes are thought to be irrelevant for cancer development. With increasing numbers of large-scale genomic datasets available, integrating these genomic data to identify driver genes from aberration regions of cancer genomes becomes an important goal of cancer genome analysis and investigations into mechanisms responsible for cancer development. A computational method, MAXDRIVER, is proposed here to identify potential driver genes on the basis …


Pathoscope: Species Identification And Strain Attribution With Unassembled Sequencing Data., Owen E Francis, Matthew Bendall, Solaiappan Manimaran, Changjin Hong, Nathan L Clement, Eduardo Castro-Nallar, Quinn Snell, G Bruce Schaalje, Mark J Clement, Keith A Crandall, W Evan Johnson Oct 2013

Pathoscope: Species Identification And Strain Attribution With Unassembled Sequencing Data., Owen E Francis, Matthew Bendall, Solaiappan Manimaran, Changjin Hong, Nathan L Clement, Eduardo Castro-Nallar, Quinn Snell, G Bruce Schaalje, Mark J Clement, Keith A Crandall, W Evan Johnson

Computational Biology Institute

Emerging next-generation sequencing technologies have revolutionized the collection of genomic data for applications in bioforensics, biosurveillance, and for use in clinical settings. However, to make the most of these new data, new methodology needs to be developed that can accommodate large volumes of genetic data in a computationally efficient manner. We present a statistical framework to analyze raw next-generation sequence reads from purified or mixed environmental or targeted infected tissue samples for rapid species identification and strain attribution against a robust database of known biological agents. Our method, Pathoscope, capitalizes on a Bayesian statistical framework that accommodates information on sequence …


A Unified Framework Integrating Parent-Of-Origin Effects For Association Study, Feifei Xiao, Jianzhong Ma, Christopher I. I. Amos Aug 2013

A Unified Framework Integrating Parent-Of-Origin Effects For Association Study, Feifei Xiao, Jianzhong Ma, Christopher I. I. Amos

Dartmouth Scholarship

Genetic imprinting is the most well-known cause for parent-of-origin effect (POE) whereby a gene is differentially expressed depending on the parental origin of the same alleles. Genetic imprinting is related to several human disorders, including diabetes, breast cancer, alcoholism, and obesity. This phenomenon has been shown to be important for normal embryonic development in mammals. Traditional association approaches ignore this important genetic phenomenon. In this study, we generalize the natural and orthogonal interactions (NOIA) framework to allow for estimation of both main allelic effects and POEs. We develop a statistical (Stat-POE) model that has the orthogonal estimates of parameters including …


Key Genes For Modulating Information Flow Play A Temporal Role As Breast Tumor Coexpression Networks Are Dynamically Rewired By Letrozole, Nadia M. Penrod, Jason H. Moore May 2013

Key Genes For Modulating Information Flow Play A Temporal Role As Breast Tumor Coexpression Networks Are Dynamically Rewired By Letrozole, Nadia M. Penrod, Jason H. Moore

Dartmouth Scholarship

Genes do not act in isolation but instead as part of complex regulatory networks. To understand how breast tumors adapt to the presence of the drug letrozole, at the molecular level, it is necessary to consider how the expression levels of genes in these networks change relative to one another. Using transcriptomic data generated from sequential tumor biopsy samples, taken at diagnosis, following 10-14 days and following 90 days of letrozole treatment, and a pairwise partial orrelation statistic, we build temporal gene coexpression networks. We characterize the structure of each network and identify genes that hold prominent positions for maintaining …