Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Genomics (4)
- Physical Sciences and Mathematics (4)
- Medicine and Health Sciences (3)
- Biostatistics (2)
- Computer Sciences (2)
-
- Diseases (2)
- Genetics (2)
- Medical Genetics (2)
- Medical Sciences (2)
- Statistics and Probability (2)
- Artificial Intelligence and Robotics (1)
- Bioinformatics (1)
- Cancer Biology (1)
- Cell and Developmental Biology (1)
- Databases and Information Systems (1)
- Ecology and Evolutionary Biology (1)
- Evolution (1)
- Immune System Diseases (1)
- Molecular Genetics (1)
- Neoplasms (1)
- Public Health (1)
- Institution
- Publication Year
- Publication
- Publication Type
Articles 1 - 13 of 13
Full-Text Articles in Computational Biology
Creation Of A Digital Storage System For Genome Sequencing Metadata, Jacquelin W. Olexa
Creation Of A Digital Storage System For Genome Sequencing Metadata, Jacquelin W. Olexa
Undergraduate Theses, Professional Papers, and Capstone Artifacts
As the field of computational genomics continues to expand in both potential and application, it is now more imperative than ever to ensure that massive genetic sequencing datasets are properly stored in an accessible manner. This project sought to establish a practical, user-friendly, secure system for a genomics research lab (the Good Lab; thegoodlab.org) at the University of Montana. A MySQL database and connected web application was ruled the best configuration to maximize utility and accessibility for the lab’s researchers. Building the logical framework for the database, creating the server, and sourcing data occurred over several months. The dataset ranged …
Identifying New Cancer Genes Based On The Integration Of Annotated Gene Sets Via Hypergraph Neural Networks, Chao Deng, Hong-Dong Li, Li-Shen Zhang, Yiwei Liu, Yaohang Li, Jianxin Wang
Identifying New Cancer Genes Based On The Integration Of Annotated Gene Sets Via Hypergraph Neural Networks, Chao Deng, Hong-Dong Li, Li-Shen Zhang, Yiwei Liu, Yaohang Li, Jianxin Wang
Computer Science Faculty Publications
Motivation
Identifying cancer genes remains a significant challenge in cancer genomics research. Annotated gene sets encode functional associations among multiple genes, and cancer genes have been shown to cluster in hallmark signaling pathways and biological processes. The knowledge of annotated gene sets is critical for discovering cancer genes but remains to be fully exploited.
Results
Here, we present the DIsease-Specific Hypergraph neural network (DISHyper), a hypergraph-based computational method that integrates the knowledge from multiple types of annotated gene sets to predict cancer genes. First, our benchmark results demonstrate that DISHyper outperforms the existing state-of-the-art methods and highlight the advantages of …
Rare Coding Variants In 35 Genes Associate With Circulating Lipid Levels-A Multi-Ancestry Analysis Of 170,000 Exomes, George Hindy, Peter Dornbos, Mark D Chaffin, Dajiang J Liu, Minxian Wang, Margaret Sunitha Selvaraj, David Zhang, Joseph Park, Carlos A Aguilar-Salinas, Lucinda Antonacci-Fulton, Diego Ardissino, Donna K Arnett, Stella Aslibekyan, Gil Atzmon, Christie M Ballantyne, Francisco Barajas-Olmos, Nir Barzilai, Lewis C Becker, Lawrence F Bielak, Joshua C Bis, John Blangero, Eric Boerwinkle, Lori L Bonnycastle, Erwin Bottinger, Donald W Bowden, Matthew J Bown, Jennifer A Brody, Jai G Broome, Noël P Burtt, Brian E Cade, Federico Centeno-Cruz, Edmund Chan, Yi-Cheng Chang, Yii-Der I Chen, Ching-Yu Cheng, Won Jung Choi, Rajiv Chowdhury, Cecilia Contreras-Cubas, Emilio J Córdova, Adolfo Correa, L Adrienne Cupples, Joanne E Curran, John Danesh, Paul S De Vries, Ralph A Defronzo, Harsha Doddapaneni, Ravindranath Duggirala, Susan K Dutcher, Patrick T Ellinor, Leslie S Emery, Jose C Florez, Myriam Fornage, Barry I Freedman, Valentin Fuster, Ma Eugenia Garay-Sevilla, Humberto García-Ortiz, Soren Germer, Richard A Gibbs, Christian Gieger, Benjamin Glaser, Clicerio Gonzalez, Maria Elena Gonzalez-Villalpando, Mariaelisa Graff, Sarah E Graham, Niels Grarup, Leif C Groop, Xiuqing Guo, Namrata Gupta, Sohee Han, Craig L Hanis, Torben Hansen, Jiang He, Nancy L Heard-Costa, Yi-Jen Hung, Mi Yeong Hwang, Marguerite R Irvin, Sergio Islas-Andrade, Gail P Jarvik, Hyun Min Kang, Sharon L R Kardia, Tanika Kelly, Eimear E Kenny, Alyna T Khan, Bong-Jo Kim, Ryan W Kim, Young Jin Kim, Heikki A Koistinen, Charles Kooperberg, Johanna Kuusisto, Soo Heon Kwak, Markku Laakso, Leslie A Lange, Jiwon Lee, Juyoung Lee, Seonwook Lee, Donna M Lehman, Rozenn N Lemaitre, Allan Linneberg, Jianjun Liu, Ruth J F Loos, Steven A Lubitz, Valeriya Lyssenko, Ronald C W Ma, Lisa Warsinger Martin, Angélica Martínez-Hernández, Rasika A Mathias, Stephen T Mcgarvey, Ruth Mcpherson, James B Meigs, Thomas Meitinger, Olle Melander, Elvia Mendoza-Caamal, Ginger A Metcalf, Xuenan Mi, Karen L Mohlke, May E Montasser, Jee-Young Moon, Hortensia Moreno-Macías, Alanna C Morrison, Donna M Muzny, Sarah C Nelson, Peter M Nilsson, Jeffrey R O'Connell, Marju Orho-Melander, Lorena Orozco, Colin N A Palmer, Nicholette D Palmer, Cheol Joo Park, Kyong Soo Park, Oluf Pedersen, Juan M Peralta, Patricia A Peyser, Wendy S Post, Michael Preuss, Bruce M Psaty, Qibin Qi, D C Rao, Susan Redline, Alexander P Reiner, Cristina Revilla-Monsalve, Stephen S Rich, Nilesh Samani, Heribert Schunkert, Claudia Schurmann, Daekwan Seo, Jeong-Sun Seo, Xueling Sim, Rob Sladek, Kerrin S Small, Wing Yee So, Adrienne M Stilp, E Shyong Tai, Claudia H T Tam, Kent D Taylor, Yik Ying Teo, Farook Thameem, Brian Tomlinson, Michael Y Tsai, Tiinamaija Tuomi, Jaakko Tuomilehto, Teresa Tusié-Luna, Miriam S Udler, Rob M Van Dam, Ramachandran S Vasan, Karine A Viaud Martinez, Fei Fei Wang, Xuzhi Wang, Hugh Watkins, Daniel E Weeks, James G Wilson, Daniel R Witte, Tien-Yin Wong, Lisa R Yanek, Amp-T2d-Genes, Myocardial Infarction Genetics Consortium, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Nhlbi Topmed Lipids Working Group, Sekar Kathiresan, Daniel J Rader, Jerome I Rotter, Michael Boehnke, Mark I Mccarthy, Cristen J Willer, Pradeep Natarajan, Jason A Flannick, Amit V Khera, Gina M Peloso
Rare Coding Variants In 35 Genes Associate With Circulating Lipid Levels-A Multi-Ancestry Analysis Of 170,000 Exomes, George Hindy, Peter Dornbos, Mark D Chaffin, Dajiang J Liu, Minxian Wang, Margaret Sunitha Selvaraj, David Zhang, Joseph Park, Carlos A Aguilar-Salinas, Lucinda Antonacci-Fulton, Diego Ardissino, Donna K Arnett, Stella Aslibekyan, Gil Atzmon, Christie M Ballantyne, Francisco Barajas-Olmos, Nir Barzilai, Lewis C Becker, Lawrence F Bielak, Joshua C Bis, John Blangero, Eric Boerwinkle, Lori L Bonnycastle, Erwin Bottinger, Donald W Bowden, Matthew J Bown, Jennifer A Brody, Jai G Broome, Noël P Burtt, Brian E Cade, Federico Centeno-Cruz, Edmund Chan, Yi-Cheng Chang, Yii-Der I Chen, Ching-Yu Cheng, Won Jung Choi, Rajiv Chowdhury, Cecilia Contreras-Cubas, Emilio J Córdova, Adolfo Correa, L Adrienne Cupples, Joanne E Curran, John Danesh, Paul S De Vries, Ralph A Defronzo, Harsha Doddapaneni, Ravindranath Duggirala, Susan K Dutcher, Patrick T Ellinor, Leslie S Emery, Jose C Florez, Myriam Fornage, Barry I Freedman, Valentin Fuster, Ma Eugenia Garay-Sevilla, Humberto García-Ortiz, Soren Germer, Richard A Gibbs, Christian Gieger, Benjamin Glaser, Clicerio Gonzalez, Maria Elena Gonzalez-Villalpando, Mariaelisa Graff, Sarah E Graham, Niels Grarup, Leif C Groop, Xiuqing Guo, Namrata Gupta, Sohee Han, Craig L Hanis, Torben Hansen, Jiang He, Nancy L Heard-Costa, Yi-Jen Hung, Mi Yeong Hwang, Marguerite R Irvin, Sergio Islas-Andrade, Gail P Jarvik, Hyun Min Kang, Sharon L R Kardia, Tanika Kelly, Eimear E Kenny, Alyna T Khan, Bong-Jo Kim, Ryan W Kim, Young Jin Kim, Heikki A Koistinen, Charles Kooperberg, Johanna Kuusisto, Soo Heon Kwak, Markku Laakso, Leslie A Lange, Jiwon Lee, Juyoung Lee, Seonwook Lee, Donna M Lehman, Rozenn N Lemaitre, Allan Linneberg, Jianjun Liu, Ruth J F Loos, Steven A Lubitz, Valeriya Lyssenko, Ronald C W Ma, Lisa Warsinger Martin, Angélica Martínez-Hernández, Rasika A Mathias, Stephen T Mcgarvey, Ruth Mcpherson, James B Meigs, Thomas Meitinger, Olle Melander, Elvia Mendoza-Caamal, Ginger A Metcalf, Xuenan Mi, Karen L Mohlke, May E Montasser, Jee-Young Moon, Hortensia Moreno-Macías, Alanna C Morrison, Donna M Muzny, Sarah C Nelson, Peter M Nilsson, Jeffrey R O'Connell, Marju Orho-Melander, Lorena Orozco, Colin N A Palmer, Nicholette D Palmer, Cheol Joo Park, Kyong Soo Park, Oluf Pedersen, Juan M Peralta, Patricia A Peyser, Wendy S Post, Michael Preuss, Bruce M Psaty, Qibin Qi, D C Rao, Susan Redline, Alexander P Reiner, Cristina Revilla-Monsalve, Stephen S Rich, Nilesh Samani, Heribert Schunkert, Claudia Schurmann, Daekwan Seo, Jeong-Sun Seo, Xueling Sim, Rob Sladek, Kerrin S Small, Wing Yee So, Adrienne M Stilp, E Shyong Tai, Claudia H T Tam, Kent D Taylor, Yik Ying Teo, Farook Thameem, Brian Tomlinson, Michael Y Tsai, Tiinamaija Tuomi, Jaakko Tuomilehto, Teresa Tusié-Luna, Miriam S Udler, Rob M Van Dam, Ramachandran S Vasan, Karine A Viaud Martinez, Fei Fei Wang, Xuzhi Wang, Hugh Watkins, Daniel E Weeks, James G Wilson, Daniel R Witte, Tien-Yin Wong, Lisa R Yanek, Amp-T2d-Genes, Myocardial Infarction Genetics Consortium, Nhlbi Trans-Omics For Precision Medicine (Topmed) Consortium, Nhlbi Topmed Lipids Working Group, Sekar Kathiresan, Daniel J Rader, Jerome I Rotter, Michael Boehnke, Mark I Mccarthy, Cristen J Willer, Pradeep Natarajan, Jason A Flannick, Amit V Khera, Gina M Peloso
Faculty, Staff and Student Publications
Large-scale gene sequencing studies for complex traits have the potential to identify causal genes with therapeutic implications. We performed gene-based association testing of blood lipid levels with rare (minor allele frequency < 1%) predicted damaging coding variation by using sequence data from >170,000 individuals from multiple ancestries: 97,493 European, 30,025 South Asian, 16,507 African, 16,440 Hispanic/Latino, 10,420 East Asian, and 1,182 Samoan. We identified 35 genes associated with circulating lipid levels; some of these genes have not been previously associated with lipid levels when using rare coding variation from population-based samples. We prioritize 32 genes in array-based genome-wide association study (GWAS) loci based on aggregations of rare coding variants; three (EVI5, …
Leveraging Global Gene Expression Patterns To Predict Expression Of Unmeasured Genes, James Rudd, René A. Zelaya, Eugene Demidenko, Ellen L. Goode, Casey S. Greene S. Greene, Jennifer A. Doherty
Leveraging Global Gene Expression Patterns To Predict Expression Of Unmeasured Genes, James Rudd, René A. Zelaya, Eugene Demidenko, Ellen L. Goode, Casey S. Greene S. Greene, Jennifer A. Doherty
Dartmouth Scholarship
BackgroundLarge collections of paraffin-embedded tissue represent a rich resource to test hypotheses based on gene expression patterns; however, measurement of genome-wide expression is cost-prohibitive on a large scale. Using the known expression correlation structure within a given disease type (in this case, high grade serous ovarian cancer; HGSC), we sought to identify reduced sets of directly measured (DM) genes which could accurately predict the expression of a maximized number of unmeasured genes.
Loregic: A Method To Characterize The Cooperative Logic Of Regulatory Factors, Daifeng Wang, Koon-Kiu Yan, Cristina Sisu, Chao Cheng, Joel Rozowsky, William Meyerson, Mark B. Gerstein
Loregic: A Method To Characterize The Cooperative Logic Of Regulatory Factors, Daifeng Wang, Koon-Kiu Yan, Cristina Sisu, Chao Cheng, Joel Rozowsky, William Meyerson, Mark B. Gerstein
Dartmouth Scholarship
The topology of the gene-regulatory network has been extensively analyzed. Now, given the large amount of available functional genomic data, it is possible to go beyond this and systematically study regulatory circuits in terms of logic elements. To this end, we present Loregic, a computational method integrating gene expression and regulatory network data, to characterize the cooperativity of regulatory factors. Loregic uses all 16 possible two-input-one-output logic gates (e.g. AND or XOR) to describe triplets of two factors regulating a common target. We attempt to find the gate that best matches each triplet’s observed gene expression pattern across many conditions. …
Systems Level Analysis Of Systemic Sclerosis Shows A Network Of Immune And Profibrotic Pathways Connected With Genetic Polymorphisms, J. Matthew Mahoney, Jaclyn Taroni, Viktor Martyanov, Tammara A. A. Wood, Casey S. Greene, Patricia A. Pioli, Monique E. Hinchcliff, Michael L. Whitfield
Systems Level Analysis Of Systemic Sclerosis Shows A Network Of Immune And Profibrotic Pathways Connected With Genetic Polymorphisms, J. Matthew Mahoney, Jaclyn Taroni, Viktor Martyanov, Tammara A. A. Wood, Casey S. Greene, Patricia A. Pioli, Monique E. Hinchcliff, Michael L. Whitfield
Dartmouth Scholarship
Systemic sclerosis (SSc) is a rare systemic autoimmune disease characterized by skin and organ fibrosis. The pathogenesis of SSc and its progression are poorly understood. The SSc intrinsic gene expression subsets (inflammatory, fibroproliferative, normal-like, and limited) are observed in multiple clinical cohorts of patients with SSc. Analysis of longitudinal skin biopsies suggests that a patient's subset assignment is stable over 6-12 months. Genetically, SSc is multi-factorial with many genetic risk loci for SSc generally and for specific clinical manifestations. Here we identify the genes consistently associated with the intrinsic subsets across three independent cohorts, show the relationship between these genes …
Orthoclust: An Orthology-Based Network Framework For Clustering Data Across Multiple Species, Koon-Kiu Yan, Daifeng Wang, Joel Rozowsky, Henry Zheng, Chao Cheng, Mark Gerstein Gerstein
Orthoclust: An Orthology-Based Network Framework For Clustering Data Across Multiple Species, Koon-Kiu Yan, Daifeng Wang, Joel Rozowsky, Henry Zheng, Chao Cheng, Mark Gerstein Gerstein
Dartmouth Scholarship
Increasingly, high-dimensional genomics data are becoming available for many organisms.Here, we develop OrthoClust for simultaneously clustering data across multiple species. OrthoClust is a computational framework that integrates the co-association networks of individual species by utilizing the orthology relationships of genes between species. It outputs optimized modules that are fundamentally cross-species, which can either be conserved or species-specific. We demonstrate the application of OrthoClust using the RNA-Seq expression profiles of Caenorhabditis elegans and Drosophila melanogaster from the modENCODE consortium. A potential application of cross-species modules is to infer putative analogous functions of uncharacterized elements like non-coding RNAs based on guilt-by-association.
A Unified Framework Integrating Parent-Of-Origin Effects For Association Study, Feifei Xiao, Jianzhong Ma, Christopher I. I. Amos
A Unified Framework Integrating Parent-Of-Origin Effects For Association Study, Feifei Xiao, Jianzhong Ma, Christopher I. I. Amos
Dartmouth Scholarship
Genetic imprinting is the most well-known cause for parent-of-origin effect (POE) whereby a gene is differentially expressed depending on the parental origin of the same alleles. Genetic imprinting is related to several human disorders, including diabetes, breast cancer, alcoholism, and obesity. This phenomenon has been shown to be important for normal embryonic development in mammals. Traditional association approaches ignore this important genetic phenomenon. In this study, we generalize the natural and orthogonal interactions (NOIA) framework to allow for estimation of both main allelic effects and POEs. We develop a statistical (Stat-POE) model that has the orthogonal estimates of parameters including …
Cross-Ontology Multi-Level Association Rule Mining In The Gene Ontology., Prashanti Manda, Seval Ozkan, Hui Wang, Fiona M. Mccarthy, Susan M. Bridges
Cross-Ontology Multi-Level Association Rule Mining In The Gene Ontology., Prashanti Manda, Seval Ozkan, Hui Wang, Fiona M. Mccarthy, Susan M. Bridges
BCoE Publications
The Gene Ontology (GO) has become the internationally accepted standard for representing function, process, and location aspects of gene products. The wealth of GO annotation data provides a valuable source of implicit knowledge of relationships among these aspects. We describe a new method for association rule mining to discover implicit co-occurrence relationships across the GO sub-ontologies at multiple levels of abstraction. Prior work on association rule mining in the GO has concentrated on mining knowledge at a single level of abstraction and/or between terms from the same sub-ontology. We have developed a bottom-up generalization procedure called Cross-Ontology Data Mining-Level by …
Additive Functions In Boolean Models Of Gene Regulatory Network Modules, Christian Darabos, Ferdinando Ferdinando Di Cunto, Marco Tomassini, Jason H. Moore
Additive Functions In Boolean Models Of Gene Regulatory Network Modules, Christian Darabos, Ferdinando Ferdinando Di Cunto, Marco Tomassini, Jason H. Moore
Dartmouth Scholarship
Gene-on-gene regulations are key components of every living organism. Dynamical abstract models of genetic regulatory networks help explain the genome’s evolvability and robustness. These properties can be attributed to the structural topology of the graph formed by genes, as vertices, and regulatory interactions, as edges. Moreover, the actual gene interaction of each gene is believed to play a key role in the stability of the structure. With advances in biology, some effort was deployed to develop update functions in Boolean models that include recent knowledge. We combine real-life gene interaction networks with novel update functions in a Boolean model. We …
Micrornas And The Advent Of Vertebrate Morphological Complexity, Alysha M. Heimberg, Lorenzo F. Sempere, Vanessa N. Moy, Phillip C. J. Donoghue, Kevin J. Peterson
Micrornas And The Advent Of Vertebrate Morphological Complexity, Alysha M. Heimberg, Lorenzo F. Sempere, Vanessa N. Moy, Phillip C. J. Donoghue, Kevin J. Peterson
Dartmouth Scholarship
The causal basis of vertebrate complexity has been sought in genome duplication events (GDEs) that occurred during the emergence of vertebrates, but evidence beyond coincidence is wanting. MicroRNAs (miRNAs) have recently been identified as a viable causal factor in increasing organismal complexity through the action of these ≈22-nt noncoding RNAs in regulating gene expression. Because miRNAs are continuously being added to animalian genomes, and, once integrated into a gene regulatory network, are strongly conserved in primary sequence and rarely secondarily lost, their evolutionary history can be accurately reconstructed. Here, using a combination of Northern analyses and genomic searches, we show …
Bounded Search For De Novo Identification Of Degenerate Cis-Regulatory Elements, Jonathan M. Carlson, Arijit Chakravarty, Radhika S. Khetani, Robert H. Gross
Bounded Search For De Novo Identification Of Degenerate Cis-Regulatory Elements, Jonathan M. Carlson, Arijit Chakravarty, Radhika S. Khetani, Robert H. Gross
Dartmouth Scholarship
The identification of statistically overrepresented sequences in the upstream regions of coregulated genes should theoretically permit the identification of potential cis-regulatory elements. However, in practice many cis-regulatory elements are highly degenerate, precluding the use of an exhaustive word-counting strategy for their identification. While numerous methods exist for inferring base distributions using a position weight matrix, recent studies suggest that the independence assumptions inherent in the model, as well as the inability to reach a global optimum, limit this approach.
Principal Component Analysis For Predicting Transcription-Factor Binding Motifs From Array-Derived Data, Yunlong Liu, Matthew P Vincenti, Hiroki Yokota
Principal Component Analysis For Predicting Transcription-Factor Binding Motifs From Array-Derived Data, Yunlong Liu, Matthew P Vincenti, Hiroki Yokota
Dartmouth Scholarship
The responses to interleukin 1 (IL-1) in human chondrocytes constitute a complex regulatory mechanism, where multiple transcription factors interact combinatorially to transcription-factor binding motifs (TFBMs). In order to select a critical set of TFBMs from genomic DNA information and an array-derived data, an efficient algorithm to solve a combinatorial optimization problem is required. Although computational approaches based on evolutionary algorithms are commonly employed, an analytical algorithm would be useful to predict TFBMs at nearly no computational cost and evaluate varying modelling conditions. Singular value decomposition (SVD) is a powerful method to derive primary components of a given matrix. Applying SVD …