Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Genomics (66)
- Biology (62)
- Medicine and Health Sciences (55)
- Bioinformatics (49)
- Molecular Genetics (47)
-
- Medical Sciences (37)
- Medical Genetics (33)
- Biomedical Informatics (32)
- Medical Specialties (31)
- Biochemistry, Biophysics, and Structural Biology (30)
- Genetics (30)
- Molecular Biology (28)
- Biological Phenomena, Cell Phenomena, and Immunity (27)
- Medical Molecular Biology (27)
- Computational Biology (19)
- Ecology and Evolutionary Biology (17)
- Evolution (14)
- Microbiology (12)
- Physical Sciences and Mathematics (12)
- Plant Sciences (12)
- Biochemistry (10)
- Engineering (10)
- Computer Engineering (9)
- Cell and Developmental Biology (7)
- Other Genetics and Genomics (7)
- Computer Sciences (6)
- Animal Sciences (5)
- Biotechnology (5)
- Institution
-
- Augustana College (42)
- The Texas Medical Center Library (36)
- Dartmouth College (17)
- University of Kentucky (17)
- University of South Carolina (9)
-
- Old Dominion University (6)
- Nova Southeastern University (4)
- Aga Khan University (3)
- Mississippi State University (3)
- Central Washington University (2)
- Children's Mercy Kansas City (2)
- Rowan University (2)
- The University of Southern Mississippi (2)
- University of Nebraska - Lincoln (2)
- University of New Hampshire (2)
- Utah State University (2)
- Brigham Young University (1)
- Bucknell University (1)
- City University of New York (CUNY) (1)
- Clemson University (1)
- East Tennessee State University (1)
- James Madison University (1)
- LSU Health New Orleans (1)
- Louisiana State University (1)
- Philadelphia College of Osteopathic Medicine (1)
- Santa Clara University (1)
- Thomas Jefferson University (1)
- Touro College and University System (1)
- University of Connecticut (1)
- University of Montana (1)
- Publication Year
- Publication
-
- Meiothermus ruber Genome Analysis Project (42)
- Faculty, Staff and Students Publications (27)
- Dartmouth Scholarship (17)
- Faculty Publications (9)
- Biology Faculty Publications (8)
-
- Faculty, Staff and Student Publications (8)
- Biology Faculty Articles (4)
- Computer Science Faculty Publications (3)
- Biological Sciences Faculty Publications (2)
- CALS Publications (2)
- College of Science & Mathematics Departmental Research (2)
- Commonwealth Computational Summit (2)
- Entomology Faculty Publications (2)
- Manuscripts, Articles, Book Chapters and Other Papers (2)
- Office of the Provost (2)
- RISK: Health, Safety & Environment (1990-2002) (2)
- Theses and Dissertations (2)
- All Faculty Scholarship for the College of the Sciences (1)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (1)
- All Master's Theses (1)
- Biochemistry Publications (1)
- Biology (1)
- Department of Animal Science: Faculty Publications (1)
- Department of Pathology and Laboratory Medicine (1)
- Department of Pathology, Anatomy, and Cell Biology Faculty Papers (1)
- Dissertations (1)
- Dissertations and Theses (Open Access) (1)
- Electronic Theses and Dissertations (1)
- Faculty Journal Articles (1)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (1)
- Publication Type
Articles 1 - 30 of 170
Full-Text Articles in Genetics and Genomics
From Chromosomes To Precision Therapy: Clinical Cytogenetics And Cytogenomics In The Era Of Genomic Medicine, Jinglan Liu
From Chromosomes To Precision Therapy: Clinical Cytogenetics And Cytogenomics In The Era Of Genomic Medicine, Jinglan Liu
Department of Pathology, Anatomy, and Cell Biology Faculty Papers
No abstract provided.
A Complete Diploid Human Genome Benchmark For Personalized Genomics, Nancy F. Hansen, Nathan Dwarshuis, Hyun Joo Ji, Arang Rhie, Hailey Loucks, Glennis A. Logsdon, Mitchell R. Vallger, Jessica M. Storer, Juhyun Kim, Eleni Adam, Nicolas Alternose, Dmitry Antipov, Mobin Asri, Sofia Barreira, Stephanie C. Bohaczuk, Andrey V. Bzikadze, Sara A. Carioscia, Andrew Carroll, Kuan-Hao Chao, Yanan Chu, Arun Das, Peter Ebert, Adam English, Mark Fleharty, Laura E. Fleming, Giulio Formenti, Andrea Guarracino, Gabrielle A. Hartley, Katharine Jenike, Jenna Kalleberg, Yu Kang, Robert King, Josipa Lipovac, Mira Mastoras, Matthew W. Mitchell, Shloka Negi, Nathan D. Olson, Keisuke K. Oshima, Luis F. Paulin, Brandon D. Pickett, David Porubsky, Jane Ranchalis, Desh Ranjan, Mikko Rautiainen, Harold Riethman, Robert D. Schnabel, Fritz J. Sedlazeck, Kishwar Shafin, Mile Sikic, Steven J. Solar, Alexander P. Sweeten, Winston Timp, Justin Wagner, Dongahn Yoo, Ying Zhou, Erik Garrison, Evan E. Eichler, Michaeel C. Schatz, Andrew B. Stergachis, Rachel J. O'Neill, Karen H. Miga, Steven L. Salzberg, Sergey Koren, Justin M. Zook, Adam M. Phillippy
A Complete Diploid Human Genome Benchmark For Personalized Genomics, Nancy F. Hansen, Nathan Dwarshuis, Hyun Joo Ji, Arang Rhie, Hailey Loucks, Glennis A. Logsdon, Mitchell R. Vallger, Jessica M. Storer, Juhyun Kim, Eleni Adam, Nicolas Alternose, Dmitry Antipov, Mobin Asri, Sofia Barreira, Stephanie C. Bohaczuk, Andrey V. Bzikadze, Sara A. Carioscia, Andrew Carroll, Kuan-Hao Chao, Yanan Chu, Arun Das, Peter Ebert, Adam English, Mark Fleharty, Laura E. Fleming, Giulio Formenti, Andrea Guarracino, Gabrielle A. Hartley, Katharine Jenike, Jenna Kalleberg, Yu Kang, Robert King, Josipa Lipovac, Mira Mastoras, Matthew W. Mitchell, Shloka Negi, Nathan D. Olson, Keisuke K. Oshima, Luis F. Paulin, Brandon D. Pickett, David Porubsky, Jane Ranchalis, Desh Ranjan, Mikko Rautiainen, Harold Riethman, Robert D. Schnabel, Fritz J. Sedlazeck, Kishwar Shafin, Mile Sikic, Steven J. Solar, Alexander P. Sweeten, Winston Timp, Justin Wagner, Dongahn Yoo, Ying Zhou, Erik Garrison, Evan E. Eichler, Michaeel C. Schatz, Andrew B. Stergachis, Rachel J. O'Neill, Karen H. Miga, Steven L. Salzberg, Sergey Koren, Justin M. Zook, Adam M. Phillippy
School of Medical Diagnostics & Translational Sciences Publications
Human genome sequencing typically relies on mapping reads to a reference genome to call variants, but this approach introduces technical biases, excluding duplicated and structurally polymorphic regions of the genome. To overcome this, we present a telomere-to-telomere genome benchmark with near-perfect accuracy across 99.4% of the diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), which were absent from prior benchmarks. We annotated genes and repeats on both haplotypes, including 19,956 protein-coding genes on the maternal haplotype and 19,190 on the paternal haplotype, and developed new methods to measure the accuracy of …
Comparative Analysis Of Six New Chloroplast Genomes In Platanthera (Orchidaceae) Enhances Understanding Of Its Diversification, Lisa E. Wallace, Martin I. Batalla, Matthew Maisonave
Comparative Analysis Of Six New Chloroplast Genomes In Platanthera (Orchidaceae) Enhances Understanding Of Its Diversification, Lisa E. Wallace, Martin I. Batalla, Matthew Maisonave
Biological Sciences Faculty Publications
Platanthera is among the most diverse genera of orchids in north temperate regions, with high species diversity in Asia and North America. Despite many ecological studies in this group, little is known about its genomic diversity. Here, we report the newly annotated chloroplast genomes from six species in subgenus Limnorchis from North America and compare them with plastomes published for Platanthera species from Asia and Europe. We found that the plastomes of species in subg. Limnorchis are among the smallest reported for Platanthera thus far. While major structural differences were not detected in the newly sequenced species, ndh genes were …
Explainable Convolutional Neural Network Model Provides An Alternative Genome-Wide Association Perspective On Mutations In Sars-Cov-2, Parisa C. Hatami, Richard Annan, Luis Miranda, Jane L. Gorman, Mengjun Xie, Letu Qingge, Hong Qin
Explainable Convolutional Neural Network Model Provides An Alternative Genome-Wide Association Perspective On Mutations In Sars-Cov-2, Parisa C. Hatami, Richard Annan, Luis Miranda, Jane L. Gorman, Mengjun Xie, Letu Qingge, Hong Qin
Computer Science Faculty Publications
Identifying informative genomic features in SARS-CoV-2 can help clarify patterns of viral evolution. In this study, we developed an explainable convolutional neural network (CNN) model to classify SARS-CoV-2 genomic sequences into the WHO-designated Variants of Concern (VOCs), Alpha, Beta, Gamma, Delta, and Omicron. Using a balanced dataset of genomes, the classification CNN achieved 99.96% accuracy on the held-out test set. To interpret the model’s predictions, we applied SHapley Additive exPlanations (SHAP) to estimate the contribution of each nucleotide position to VOC-label prediction and compared aggregated attributions with a chi-square GWAS baseline applied to the same categorical labels. SHAP prioritized several …
Complete Genome Sequence Of Rothia Mucilaginosa D1, A Genetically Amenable Strain, Hemendra P. Dhaked, Isabella Andonie, Bo Yang, Vinayka M. Joshi, Zezhang T. Wen
Complete Genome Sequence Of Rothia Mucilaginosa D1, A Genetically Amenable Strain, Hemendra P. Dhaked, Isabella Andonie, Bo Yang, Vinayka M. Joshi, Zezhang T. Wen
School of Dentistry Faculty Publications
Here, we report the complete genome sequence of Rothia mucilaginosa D1. Unlike many other strains that have been sequenced, R. mucilaginosa D1, a clinical isolate, is genetically amenable. The complete sequencing of the genome provides potential for further studies on the genetics and pathophysiology of this emerging pathobiont.
Small Variant Benchmark From A Complete Assembly Of X And Y Chromosomes, Justin Wagner, Nathan D Olson, Jennifer Mcdaniel, Lindsay Harris, Brendan J Pinto, David Jáspez, Adrián Muñoz-Barrera, Luis A Rubio-Rodríguez, José M Lorenzo-Salazar, Carlos Flores, Sayed Mohammad Ebrahim Sahraeian, Giuseppe Narzisi, Marta Byrska-Bishop, Uday S Evani, Chunlin Xiao, Juniper A Lake, Peter Fontana, Craig Greenberg, Donald Freed, Mohammed Faizal Eeman Mootor, Paul C Boutros, Lisa Murray, Kishwar Shafin, Andrew Carroll, Fritz J Sedlazeck, Melissa Wilson, Justin M Zook
Small Variant Benchmark From A Complete Assembly Of X And Y Chromosomes, Justin Wagner, Nathan D Olson, Jennifer Mcdaniel, Lindsay Harris, Brendan J Pinto, David Jáspez, Adrián Muñoz-Barrera, Luis A Rubio-Rodríguez, José M Lorenzo-Salazar, Carlos Flores, Sayed Mohammad Ebrahim Sahraeian, Giuseppe Narzisi, Marta Byrska-Bishop, Uday S Evani, Chunlin Xiao, Juniper A Lake, Peter Fontana, Craig Greenberg, Donald Freed, Mohammed Faizal Eeman Mootor, Paul C Boutros, Lisa Murray, Kishwar Shafin, Andrew Carroll, Fritz J Sedlazeck, Melissa Wilson, Justin M Zook
Faculty, Staff and Students Publications
The sex chromosomes contain complex, important genes impacting medical phenotypes, but differ from the autosomes in their ploidy and large repetitive regions. To enable technology developers along with research and clinical laboratories to evaluate variant detection on male sex chromosomes X and Y, we create a small variant benchmark set with 111,725 variants for the Genome in a Bottle HG002 reference material. We develop an active evaluation approach to demonstrate the benchmark set reliably identifies errors in challenging genomic regions and across short and long read callsets. We show how complete assemblies can expand benchmarks to difficult regions, but highlight …
Leveraging The T2t Assembly To Resolve Rare And Pathogenic Inversions In Reference Genome Gaps, Kristine Bilgrav Saether, Jesper Eisfeldt, Jesse D Bengtsson, Ming Yin Lun, Christopher M Grochowski, Medhat Mahmoud, Hsiao-Tuan Chao, Jill A Rosenfeld, Pengfei Liu, Marlene Ek, Jakob Schuy, Adam Ameur, Hongzheng Dai, Undiagnosed Diseases Network, James Paul Hwang, Fritz J Sedlazeck, Weimin Bi, Ronit Marom, Josephine Wincent, Ann Nordgren, Claudia M B Carvalho, Anna Lindstrand
Leveraging The T2t Assembly To Resolve Rare And Pathogenic Inversions In Reference Genome Gaps, Kristine Bilgrav Saether, Jesper Eisfeldt, Jesse D Bengtsson, Ming Yin Lun, Christopher M Grochowski, Medhat Mahmoud, Hsiao-Tuan Chao, Jill A Rosenfeld, Pengfei Liu, Marlene Ek, Jakob Schuy, Adam Ameur, Hongzheng Dai, Undiagnosed Diseases Network, James Paul Hwang, Fritz J Sedlazeck, Weimin Bi, Ronit Marom, Josephine Wincent, Ann Nordgren, Claudia M B Carvalho, Anna Lindstrand
Faculty, Staff and Students Publications
Chromosomal inversions (INVs) are particularly challenging to detect due to their copy-number neutral state and association with repetitive regions. Inversions represent about 1/20 of all balanced structural chromosome aberrations and can lead to disease by gene disruption or altering regulatory regions of dosage-sensitive genes in cis. Short-read genome sequencing (srGS) can only resolve ∼70% of cytogenetically visible inversions referred to clinical diagnostic laboratories, likely due to breakpoints in repetitive regions. Here, we study 12 inversions by long-read genome sequencing (lrGS) (n = 9) or srGS (n = 3) and resolve nine of them. In four cases, the …
High-Coverage Nanopore Sequencing Of Samples From The 1000 Genomes Project To Build A Comprehensive Catalog Of Human Genetic Variation, Jonas A Gustafson, Sophia B Gibson, Nikhita Damaraju, Miranda P G Zalusky, Kendra Hoekzema, David Twesigomwe, Lei Yang, Anthony A Snead, Phillip A Richmond, Wouter De Coster, Nathan D Olson, Andrea Guarracino, Qiuhui Li, Angela L Miller, Joy Goffena, Zachary B Anderson, Sophie H R Storz, Sydney A Ward, Maisha Sinha, Claudia Gonzaga-Jauregui, Wayne E Clarke, Anna O Basile, André Corvelo, Catherine Reeves, Adrienne Helland, Rajeeva Lochan Musunuri, Mahler Revsine, Karynne E Patterson, Cate R Paschal, Christina Zakarian, Sara Goodwin, Tanner D Jensen, Esther Robb, 1000 Genomes Ont Sequencing Consortium, University Of Washington Center For Rare Disease Research (Uw-Crdr), Genomics Research To Elucidate The Genetics Of Rare Diseases (Gregor) Consortium, William Richard Mccombie, Fritz J Sedlazeck, Justin M Zook, Stephen B Montgomery, Erik Garrison, Mikhail Kolmogorov, Michael C Schatz, Richard N Mclaughlin, Harriet Dashnow, Michael C Zody, Matt Loose, Miten Jain, Evan E Eichler, Danny E Miller
High-Coverage Nanopore Sequencing Of Samples From The 1000 Genomes Project To Build A Comprehensive Catalog Of Human Genetic Variation, Jonas A Gustafson, Sophia B Gibson, Nikhita Damaraju, Miranda P G Zalusky, Kendra Hoekzema, David Twesigomwe, Lei Yang, Anthony A Snead, Phillip A Richmond, Wouter De Coster, Nathan D Olson, Andrea Guarracino, Qiuhui Li, Angela L Miller, Joy Goffena, Zachary B Anderson, Sophie H R Storz, Sydney A Ward, Maisha Sinha, Claudia Gonzaga-Jauregui, Wayne E Clarke, Anna O Basile, André Corvelo, Catherine Reeves, Adrienne Helland, Rajeeva Lochan Musunuri, Mahler Revsine, Karynne E Patterson, Cate R Paschal, Christina Zakarian, Sara Goodwin, Tanner D Jensen, Esther Robb, 1000 Genomes Ont Sequencing Consortium, University Of Washington Center For Rare Disease Research (Uw-Crdr), Genomics Research To Elucidate The Genetics Of Rare Diseases (Gregor) Consortium, William Richard Mccombie, Fritz J Sedlazeck, Justin M Zook, Stephen B Montgomery, Erik Garrison, Mikhail Kolmogorov, Michael C Schatz, Richard N Mclaughlin, Harriet Dashnow, Michael C Zody, Matt Loose, Miten Jain, Evan E Eichler, Danny E Miller
Faculty, Staff and Students Publications
Fewer than half of individuals with a suspected Mendelian or monogenic condition receive a precise molecular diagnosis after comprehensive clinical genetic testing. Improvements in data quality and costs have heightened interest in using long-read sequencing (LRS) to streamline clinical genomic testing, but the absence of control data sets for variant filtering and prioritization has made tertiary analysis of LRS data challenging. To address this, the 1000 Genomes Project (1KGP) Oxford Nanopore Technologies Sequencing Consortium aims to generate LRS data from at least 800 of the 1KGP samples. Our goal is to use LRS to identify a broader spectrum of variation …
The Giab Genomic Stratifications Resource For Human Reference Genomes, Nathan Dwarshuis, Divya Kalra, Jennifer Mcdaniel, Philippe Sanio, Pilar Alvarez Jerez, Bharati Jadhav, Wenyu Eddy Huang, Rajarshi Mondal, Ben Busby, Nathan D Olson, Fritz J Sedlazeck, Justin Wagner, Sina Majidian, Justin M Zook
The Giab Genomic Stratifications Resource For Human Reference Genomes, Nathan Dwarshuis, Divya Kalra, Jennifer Mcdaniel, Philippe Sanio, Pilar Alvarez Jerez, Bharati Jadhav, Wenyu Eddy Huang, Rajarshi Mondal, Ben Busby, Nathan D Olson, Fritz J Sedlazeck, Justin Wagner, Sina Majidian, Justin M Zook
Faculty, Staff and Students Publications
Despite the growing variety of sequencing and variant-calling tools, no workflow performs equally well across the entire human genome. Understanding context-dependent performance is critical for enabling researchers, clinicians, and developers to make informed tradeoffs when selecting sequencing hardware and software. Here we describe a set of “stratifications,” which are BED files that define distinct contexts throughout the genome. We define these for GRCh37/38 as well as the new T2T-CHM13 reference, adding many new hard-to-sequence regions which are critical for understanding performance as the field progresses. Specifically, we highlight the increase in hard-to-map and GC-rich stratifications in CHM13 relative to the …
Stratomod: Predicting Sequencing And Variant Calling Errors With Interpretable Machine Learning, Nathan Dwarshuis, Peter Tonner, Nathan D Olson, Fritz J Sedlazeck, Justin Wagner, Justin M Zook
Stratomod: Predicting Sequencing And Variant Calling Errors With Interpretable Machine Learning, Nathan Dwarshuis, Peter Tonner, Nathan D Olson, Fritz J Sedlazeck, Justin Wagner, Justin M Zook
Faculty, Staff and Students Publications
Despite the variety in sequencing platforms, mappers, and variant callers, no single pipeline is optimal across the entire human genome. Therefore, developers, clinicians, and researchers need to make tradeoffs when designing pipelines for their application. Currently, assessing such tradeoffs relies on intuition about how a certain pipeline will perform in a given genomic context. We present StratoMod, which addresses this problem using an interpretable machine-learning classifier to predict germline variant calling errors in a data-driven manner. We show StratoMod can precisely predict recall using Hifi or Illumina and leverage StratoMod's interpretability to measure contributions from difficult-to-map and homopolymer regions for …
Single-Cell Somatic Copy Number Variants In Brain Using Different Amplification Methods And Reference Genomes, Ester Kalef-Ezra, Zeliha Gozde Turan, Diego Perez-Rodriguez, Ida Bomann, Sairam Behera, Caoimhe Morley, Sonja W Scholz, Zane Jaunmuktane, Jonas Demeulemeester, Fritz J Sedlazeck, Christos Proukakis
Single-Cell Somatic Copy Number Variants In Brain Using Different Amplification Methods And Reference Genomes, Ester Kalef-Ezra, Zeliha Gozde Turan, Diego Perez-Rodriguez, Ida Bomann, Sairam Behera, Caoimhe Morley, Sonja W Scholz, Zane Jaunmuktane, Jonas Demeulemeester, Fritz J Sedlazeck, Christos Proukakis
Faculty, Staff and Students Publications
The presence of somatic mutations, including copy number variants (CNVs), in the brain is well recognized. Comprehensive study requires single-cell whole genome amplification, with several methods available, prior to sequencing. Here we compare PicoPLEX with two recent adaptations of multiple displacement amplification (MDA): primary template-directed amplification (PTA) and droplet MDA, across 93 human brain cortical nuclei. We demonstrate different properties for each, with PTA providing the broadest amplification, PicoPLEX the most even, and distinct chimeric profiles. Furthermore, we perform CNV calling on two brains with multiple system atrophy and one control brain using different reference genomes. We find that 20.6% …
Whole Genomes Of Amazonian Uakari Monkeys Reveal Complex Connectivity And Fast Differentiation Driven By High Environmental Dynamism, Núria Hermosilla-Albala, Felipe Ennes Silva, Sebastián Cuadros-Espinoza, Claudia Fontsere, Alejandro Valenzuela-Seba, Harvinder Pawar, Marta Gut, Joanna L Kelley, Sandra Ruibal-Puertas, Pol Alentorn-Moron, Armida Faella, Esther Lizano, Izeni Farias, Tomas Hrbek, Joao Valsecchi, Ivo G Gut, Jeffrey Rogers, Kyle Kai-How Farh, Lukas F K Kuderna, Tomas Marques-Bonet, Jean P Boubli
Whole Genomes Of Amazonian Uakari Monkeys Reveal Complex Connectivity And Fast Differentiation Driven By High Environmental Dynamism, Núria Hermosilla-Albala, Felipe Ennes Silva, Sebastián Cuadros-Espinoza, Claudia Fontsere, Alejandro Valenzuela-Seba, Harvinder Pawar, Marta Gut, Joanna L Kelley, Sandra Ruibal-Puertas, Pol Alentorn-Moron, Armida Faella, Esther Lizano, Izeni Farias, Tomas Hrbek, Joao Valsecchi, Ivo G Gut, Jeffrey Rogers, Kyle Kai-How Farh, Lukas F K Kuderna, Tomas Marques-Bonet, Jean P Boubli
Faculty, Staff and Students Publications
Despite showing the greatest primate diversity on the planet, genomic studies on Amazonian primates show very little representation in the literature. With 48 geolocalized high coverage whole genomes from wild uakari monkeys, we present the first population-level study on platyrrhines using whole genome data. In a very restricted range of the Amazon rainforest, eight uakari species (Cacajao genus) have been described and categorized into the bald and black uakari groups, based on phenotypic and ecological differences. Despite a slight habitat overlap, we show that posterior to their split 0.92 Mya, bald and black uakaris have remained independent, without gene flow. …
Impact And Characterization Of Serial Structural Variations Across Humans And Great Apes, Wolfram Höps, Tobias Rausch, Michael Jendrusch, Jan O Korbel, Fritz J Sedlazeck
Impact And Characterization Of Serial Structural Variations Across Humans And Great Apes, Wolfram Höps, Tobias Rausch, Michael Jendrusch, Jan O Korbel, Fritz J Sedlazeck
Faculty, Staff and Students Publications
Modern sequencing technology enables the systematic detection of complex structural variation (SV) across genomes. However, extensive DNA rearrangements arising through a series of mutations, a phenomenon we refer to as serial SV (sSV), remain underexplored, posing a challenge for SV discovery. Here, we present NAHRwhals ( https://github.com/WHops/NAHRwhals ), a method to infer repeat-mediated series of SVs in long-read genomic assemblies. Applying NAHRwhals to haplotype-resolved human genomes from 28 individuals reveals 37 sSV loci of various length and complexity. These sSVs explain otherwise cryptic variation in medically relevant regions such as the TPSAB1 gene, 8p23.1, 22q11 and Sotos syndrome regions. Comparisons …
Crispr-Cas9 And Cas12a Target Site Richness Reflects Genomic Diversity In Natural Populations Of Anopheles Gambiae And Aedes Aegypti Mosquitoes, Travis C Collier, Yoosook Lee, Derrick K Mathias, Víctor López Del Amo
Crispr-Cas9 And Cas12a Target Site Richness Reflects Genomic Diversity In Natural Populations Of Anopheles Gambiae And Aedes Aegypti Mosquitoes, Travis C Collier, Yoosook Lee, Derrick K Mathias, Víctor López Del Amo
Faculty, Staff and Student Publications
Due to limitations in conventional disease vector control strategies including the rise of insecticide resistance in natural populations of mosquitoes, genetic control strategies using CRISPR gene drive systems have been under serious consideration. The identification of CRISPR target sites in mosquito populations is a key aspect for developing efficient genetic vector control strategies. While genome-wide Cas9 target sites have been explored in mosquitoes, a precise evaluation of target sites focused on coding sequence (CDS) is lacking. Additionally, target site polymorphisms have not been characterized for other nucleases such as Cas12a, which require a different DNA recognition site (PAM) and would …
Inverted Triplications Formed By Iterative Template Switches Generate Structural Variant Diversity At Genomic Disorder Loci, Christopher M Grochowski, Jesse D Bengtsson, Haowei Du, Mira Gandhi, Ming Yin Lun, Michele G Mehaffey, Kyunghee Park, Wolfram Höps, Eva Benito, Patrick Hasenfeld, Jan O Korbel, Medhat Mahmoud, Luis F Paulin, Shalini N Jhangiani, James Paul Hwang, Sravya V Bhamidipati, Donna M Muzny, Jawid M Fatih, Richard A Gibbs, Matthew Pendleton, Eoghan Harrington, Sissel Juul, Anna Lindstrand, Fritz J Sedlazeck, Davut Pehlivan, James R Lupski, Claudia M B Carvalho
Inverted Triplications Formed By Iterative Template Switches Generate Structural Variant Diversity At Genomic Disorder Loci, Christopher M Grochowski, Jesse D Bengtsson, Haowei Du, Mira Gandhi, Ming Yin Lun, Michele G Mehaffey, Kyunghee Park, Wolfram Höps, Eva Benito, Patrick Hasenfeld, Jan O Korbel, Medhat Mahmoud, Luis F Paulin, Shalini N Jhangiani, James Paul Hwang, Sravya V Bhamidipati, Donna M Muzny, Jawid M Fatih, Richard A Gibbs, Matthew Pendleton, Eoghan Harrington, Sissel Juul, Anna Lindstrand, Fritz J Sedlazeck, Davut Pehlivan, James R Lupski, Claudia M B Carvalho
Faculty, Staff and Students Publications
The duplication-triplication/inverted-duplication (DUP-TRP/INV-DUP) structure is a complex genomic rearrangement (CGR). Although it has been identified as an important pathogenic DNA mutation signature in genomic disorders and cancer genomes, its architecture remains unresolved. Here, we studied the genomic architecture of DUP-TRP/INV-DUP by investigating the DNA of 24 patients identified by array comparative genomic hybridization (aCGH) on whom we found evidence for the existence of 4 out of 4 predicted structural variant (SV) haplotypes. Using a combination of short-read genome sequencing (GS), long-read GS, optical genome mapping, and single-cell DNA template strand sequencing (strand-seq), the haplotype structure was resolved in 18 samples. …
A Harmonized Public Resource Of Deeply Sequenced Diverse Human Genomes, Zan Koenig, Mary T Yohannes, Lethukuthula L Nkambule, Xuefang Zhao, Julia K Goodrich, Heesu Ally Kim, Michael W Wilson, Grace Tiao, Stephanie P Hao, Nareh Sahakian, Katherine R Chao, Mark A Walker, Yunfei Lyu, Heidi L Rehm, Benjamin M Neale, Michael E Talkowski, Mark J Daly, Harrison Brand, Konrad J Karczewski, Elizabeth G Atkinson, Alicia R Martin
A Harmonized Public Resource Of Deeply Sequenced Diverse Human Genomes, Zan Koenig, Mary T Yohannes, Lethukuthula L Nkambule, Xuefang Zhao, Julia K Goodrich, Heesu Ally Kim, Michael W Wilson, Grace Tiao, Stephanie P Hao, Nareh Sahakian, Katherine R Chao, Mark A Walker, Yunfei Lyu, Heidi L Rehm, Benjamin M Neale, Michael E Talkowski, Mark J Daly, Harrison Brand, Konrad J Karczewski, Elizabeth G Atkinson, Alicia R Martin
Faculty, Staff and Students Publications
Underrepresented populations are often excluded from genomic studies owing in part to a lack of resources supporting their analyses. The 1000 Genomes Project (1kGP) and Human Genome Diversity Project (HGDP), which have recently been sequenced to high coverage, are valuable genomic resources because of the global diversity they capture and their open data sharing policies. Here, we harmonized a high-quality set of 4094 whole genomes from 80 populations in the HGDP and 1kGP with data from the Genome Aggregation Database (gnomAD) and identified over 153 million high-quality SNVs, indels, and SVs. We performed a detailed ancestry analysis of this cohort, …
Methphaser: Methylation-Based Long-Read Haplotype Phasing Of Human Genomes, Yilei Fu, Sergey Aganezov, Medhat Mahmoud, John Beaulaurier, Sissel Juul, Todd J Treangen, Fritz J Sedlazeck
Methphaser: Methylation-Based Long-Read Haplotype Phasing Of Human Genomes, Yilei Fu, Sergey Aganezov, Medhat Mahmoud, John Beaulaurier, Sissel Juul, Todd J Treangen, Fritz J Sedlazeck
Faculty, Staff and Students Publications
The assignment of variants across haplotypes, phasing, is crucial for predicting the consequences, interaction, and inheritance of mutations and is a key step in improving our understanding of phenotype and disease. However, phasing is limited by read length and stretches of homozygosity along the genome. To overcome this limitation, we designed MethPhaser, a method that utilizes methylation signals from Oxford Nanopore Technologies to extend Single Nucleotide Variation (SNV)-based phasing. We demonstrate that haplotype-specific methylations extensively exist in Human genomes and the advent of long-read technologies enabled direct report of methylation signals. For ONT R9 and R10 cell line data, we …
Discovery Of Runs-Of-Homozygosity Diplotype Clusters And Their Associations With Diseases In Uk Biobank, Ardalan Naseri, Degui Zhi, Shaojie Zhang
Discovery Of Runs-Of-Homozygosity Diplotype Clusters And Their Associations With Diseases In Uk Biobank, Ardalan Naseri, Degui Zhi, Shaojie Zhang
Faculty, Staff and Student Publications
Runs-of-homozygosity (ROH) segments, contiguous homozygous regions in a genome were traditionally linked to families and inbred populations. However, a growing literature suggests that ROHs are ubiquitous in outbred populations. Still, most existing genetic studies of ROH in populations are limited to aggregated ROH content across the genome, which does not offer the resolution for mapping causal loci. This limitation is mainly due to a lack of methods for the efficient identification of shared ROH diplotypes. Here, we present a new method, ROH-DICE (runs-of-homozygous diplotype cluster enumerator), to find large ROH diplotype clusters, sufficiently long ROHs shared by a sufficient number …
De Novo Genome Assembly For The Coppery Titi Monkey (Plecturocebus Cupreus): An Emerging Nonhuman Primate Model For Behavioral Research, Susanne P Pfeifer, Alexander Baxter, Logan E Savidge, Fritz J Sedlazeck, Karen L Bales
De Novo Genome Assembly For The Coppery Titi Monkey (Plecturocebus Cupreus): An Emerging Nonhuman Primate Model For Behavioral Research, Susanne P Pfeifer, Alexander Baxter, Logan E Savidge, Fritz J Sedlazeck, Karen L Bales
Faculty, Staff and Students Publications
The coppery titi monkey (Plecturocebus cupreus) is an emerging nonhuman primate model system for behavioral and neurobiological research. At the same time, the almost entire absence of genomic resources for the species has hampered insights into the genetic underpinnings of the phenotypic traits of interest. To facilitate future genotype-to-phenotype studies, we here present a high-quality, fully annotated de novo genome assembly for the species with chromosome-length scaffolds spanning the autosomes and chromosome X (scaffold N50 = 130.8 Mb), constructed using data obtained from several orthologous short- and long-read sequencing and scaffolding techniques. With a base-level accuracy of ∼99.99% in chromosome-length …
Critical Assessment Of Variant Prioritization Methods For Rare Disease Diagnosis Within The Rare Genomes Project, Sarah L Stenton, Melanie C O'Leary, Gabrielle Lemire, Grace E Vannoy, Stephanie Ditroia, Vijay S Ganesh, Emily Groopman, Emily O'Heir, Brian Mangilog, Ikeoluwa Osei-Owusu, Lynn S Pais, Jillian Serrano, Moriel Singer-Berk, Ben Weisburd, Michael W Wilson, Christina Austin-Tse, Marwa Abdelhakim, Azza Althagafi, Giulia Babbi, Riccardo Bellazzi, Samuele Bovo, Maria Giulia Carta, Rita Casadio, Pieter-Jan Coenen, Federica De Paoli, Matteo Floris, Manavalan Gajapathy, Robert Hoehndorf, Julius O B Jacobsen, Thomas Joseph, Akash Kamandula, Panagiotis Katsonis, Cyrielle Kint, Olivier Lichtarge, Ivan Limongelli, Yulan Lu, Paolo Magni, Tarun Karthik Kumar Mamidi, Pier Luigi Martelli, Marta Mulargia, Giovanna Nicora, Keith Nykamp, Vikas Pejaver, Yisu Peng, Thi Hong Cam Pham, Maurizio S Podda, Aditya Rao, Ettore Rizzo, Vangala G Saipradeep, Castrense Savojardo, Peter Schols, Yang Shen, Naveen Sivadasan, Damian Smedley, Dorian Soru, Rajgopal Srinivasan, Yuanfei Sun, Uma Sunderam, Wuwei Tan, Naina Tiwari, Xiao Wang, Yaqiong Wang, Amanda Williams, Elizabeth A Worthey, Rujie Yin, Yuning You, Daniel Zeiberg, Susanna Zucca, Constantina Bakolitsa, Steven E Brenner, Stephanie M Fullerton, Predrag Radivojac, Heidi L Rehm, Anne O'Donnell-Luria
Critical Assessment Of Variant Prioritization Methods For Rare Disease Diagnosis Within The Rare Genomes Project, Sarah L Stenton, Melanie C O'Leary, Gabrielle Lemire, Grace E Vannoy, Stephanie Ditroia, Vijay S Ganesh, Emily Groopman, Emily O'Heir, Brian Mangilog, Ikeoluwa Osei-Owusu, Lynn S Pais, Jillian Serrano, Moriel Singer-Berk, Ben Weisburd, Michael W Wilson, Christina Austin-Tse, Marwa Abdelhakim, Azza Althagafi, Giulia Babbi, Riccardo Bellazzi, Samuele Bovo, Maria Giulia Carta, Rita Casadio, Pieter-Jan Coenen, Federica De Paoli, Matteo Floris, Manavalan Gajapathy, Robert Hoehndorf, Julius O B Jacobsen, Thomas Joseph, Akash Kamandula, Panagiotis Katsonis, Cyrielle Kint, Olivier Lichtarge, Ivan Limongelli, Yulan Lu, Paolo Magni, Tarun Karthik Kumar Mamidi, Pier Luigi Martelli, Marta Mulargia, Giovanna Nicora, Keith Nykamp, Vikas Pejaver, Yisu Peng, Thi Hong Cam Pham, Maurizio S Podda, Aditya Rao, Ettore Rizzo, Vangala G Saipradeep, Castrense Savojardo, Peter Schols, Yang Shen, Naveen Sivadasan, Damian Smedley, Dorian Soru, Rajgopal Srinivasan, Yuanfei Sun, Uma Sunderam, Wuwei Tan, Naina Tiwari, Xiao Wang, Yaqiong Wang, Amanda Williams, Elizabeth A Worthey, Rujie Yin, Yuning You, Daniel Zeiberg, Susanna Zucca, Constantina Bakolitsa, Steven E Brenner, Stephanie M Fullerton, Predrag Radivojac, Heidi L Rehm, Anne O'Donnell-Luria
Faculty, Staff and Students Publications
BACKGROUND: A major obstacle faced by families with rare diseases is obtaining a genetic diagnosis. The average "diagnostic odyssey" lasts over five years and causal variants are identified in under 50%, even when capturing variants genome-wide. To aid in the interpretation and prioritization of the vast number of variants detected, computational methods are proliferating. Knowing which tools are most effective remains unclear. To evaluate the performance of computational methods, and to encourage innovation in method development, we designed a Critical Assessment of Genome Interpretation (CAGI) community challenge to place variant prioritization models head-to-head in a real-life clinical diagnostic setting.
METHODS: …
Multilocus Pathogenic Variants Contribute To Intrafamilial Clinical Heterogeneity: A Retrospective Study Of Sibling Pairs With Neurodevelopmental Disorders, Tugce Bozkurt-Yozgatli, Davut Pehlivan, Richard A Gibbs, Ugur Sezerman, Jennifer E Posey, James R Lupski, Zeynep Coban-Akdemir
Multilocus Pathogenic Variants Contribute To Intrafamilial Clinical Heterogeneity: A Retrospective Study Of Sibling Pairs With Neurodevelopmental Disorders, Tugce Bozkurt-Yozgatli, Davut Pehlivan, Richard A Gibbs, Ugur Sezerman, Jennifer E Posey, James R Lupski, Zeynep Coban-Akdemir
Faculty, Staff and Students Publications
BACKGROUND: Multilocus pathogenic variants (MPVs) are genetic changes that affect multiple gene loci or regions of the genome, collectively leading to multiple molecular diagnoses. MPVs may also contribute to intrafamilial phenotypic variability between affected individuals within a nuclear family. In this study, we aim to gain further insights into the influence of MPVs on a disease manifestation in individual research subjects and explore the complexities of the human genome within a familial context.
METHODS: We conducted a systematic reanalysis of exome sequencing data and runs of homozygosity (ROH) regions of 47 sibling pairs previously diagnosed with various neurodevelopmental disorders (NDD). …
Creation Of A Digital Storage System For Genome Sequencing Metadata, Jacquelin W. Olexa
Creation Of A Digital Storage System For Genome Sequencing Metadata, Jacquelin W. Olexa
Undergraduate Theses, Professional Papers, and Capstone Artifacts
As the field of computational genomics continues to expand in both potential and application, it is now more imperative than ever to ensure that massive genetic sequencing datasets are properly stored in an accessible manner. This project sought to establish a practical, user-friendly, secure system for a genomics research lab (the Good Lab; thegoodlab.org) at the University of Montana. A MySQL database and connected web application was ruled the best configuration to maximize utility and accessibility for the lab’s researchers. Building the logical framework for the database, creating the server, and sourcing data occurred over several months. The dataset ranged …
Identification Of Constrained Sequence Elements Across 239 Primate Genomes, Lukas F K Kuderna, Jacob C Ulirsch, Sabrina Rashid, Mohamed Ameen, Laksshman Sundaram, Glenn Hickey, Anthony J Cox, Hong Gao, Arvind Kumar, Francois Aguet, Matthew J Christmas, Hiram Clawson, Maximilian Haeussler, Mareike C Janiak, Martin Kuhlwilm, Joseph D Orkin, Thomas Bataillon, Shivakumara Manu, Alejandro Valenzuela, Juraj Bergman, Marjolaine Rouselle, Felipe Ennes Silva, Lidia Agueda, Julie Blanc, Marta Gut, Dorien De Vries, Ian Goodhead, R Alan Harris, Muthuswamy Raveendran, Axel Jensen, Idriss S Chuma, Julie E Horvath, Christina Hvilsom, David Juan, Peter Frandsen, Joshua G Schraiber, Fabiano R De Melo, Fabrício Bertuol, Hazel Byrne, Iracilda Sampaio, Izeni Farias, João Valsecchi, Malu Messias, Maria N F Da Silva, Mihir Trivedi, Rogerio Rossi, Tomas Hrbek, Nicole Andriaholinirina, Clément J Rabarivola, Alphonse Zaramody, Clifford J Jolly, Jane Phillips-Conroy, Gregory Wilkerson, Christian Abee, Joe H Simmons, Eduardo Fernandez-Duque, Sree Kanthaswamy, Fekadu Shiferaw, Dongdong Wu, Long Zhou, Yong Shao, Guojie Zhang, Julius D Keyyu, Sascha Knauf, Minh D Le, Esther Lizano, Stefan Merker, Arcadi Navarro, Tilo Nadler, Chiea Chuen Khor, Jessica Lee, Patrick Tan, Weng Khong Lim, Andrew C Kitchener, Dietmar Zinner, Ivo Gut, Amanda D Melin, Katerina Guschanski, Mikkel Heide Schierup, Robin M D Beck, Ioannis Karakikes, Kevin C Wang, Govindhaswamy Umapathy, Christian Roos, Jean P Boubli, Adam Siepel, Anshul Kundaje, Benedict Paten, Kerstin Lindblad-Toh, Jeffrey Rogers, Tomas Marques Bonet, Kyle Kai-How Farh
Identification Of Constrained Sequence Elements Across 239 Primate Genomes, Lukas F K Kuderna, Jacob C Ulirsch, Sabrina Rashid, Mohamed Ameen, Laksshman Sundaram, Glenn Hickey, Anthony J Cox, Hong Gao, Arvind Kumar, Francois Aguet, Matthew J Christmas, Hiram Clawson, Maximilian Haeussler, Mareike C Janiak, Martin Kuhlwilm, Joseph D Orkin, Thomas Bataillon, Shivakumara Manu, Alejandro Valenzuela, Juraj Bergman, Marjolaine Rouselle, Felipe Ennes Silva, Lidia Agueda, Julie Blanc, Marta Gut, Dorien De Vries, Ian Goodhead, R Alan Harris, Muthuswamy Raveendran, Axel Jensen, Idriss S Chuma, Julie E Horvath, Christina Hvilsom, David Juan, Peter Frandsen, Joshua G Schraiber, Fabiano R De Melo, Fabrício Bertuol, Hazel Byrne, Iracilda Sampaio, Izeni Farias, João Valsecchi, Malu Messias, Maria N F Da Silva, Mihir Trivedi, Rogerio Rossi, Tomas Hrbek, Nicole Andriaholinirina, Clément J Rabarivola, Alphonse Zaramody, Clifford J Jolly, Jane Phillips-Conroy, Gregory Wilkerson, Christian Abee, Joe H Simmons, Eduardo Fernandez-Duque, Sree Kanthaswamy, Fekadu Shiferaw, Dongdong Wu, Long Zhou, Yong Shao, Guojie Zhang, Julius D Keyyu, Sascha Knauf, Minh D Le, Esther Lizano, Stefan Merker, Arcadi Navarro, Tilo Nadler, Chiea Chuen Khor, Jessica Lee, Patrick Tan, Weng Khong Lim, Andrew C Kitchener, Dietmar Zinner, Ivo Gut, Amanda D Melin, Katerina Guschanski, Mikkel Heide Schierup, Robin M D Beck, Ioannis Karakikes, Kevin C Wang, Govindhaswamy Umapathy, Christian Roos, Jean P Boubli, Adam Siepel, Anshul Kundaje, Benedict Paten, Kerstin Lindblad-Toh, Jeffrey Rogers, Tomas Marques Bonet, Kyle Kai-How Farh
Faculty, Staff and Students Publications
Noncoding DNA is central to our understanding of human gene regulation and complex diseases1,2, and measuring the evolutionary sequence constraint can establish the functional relevance of putative regulatory elements in the human genome3–9. Identifying the genomic elements that have become constrained specifically in primates has been hampered by the faster evolution of noncoding DNA compared to protein-coding DNA10, the relatively short timescales separating primate species11, and the previously limited availability of whole-genome sequences12. Here we construct a whole-genome alignment of 239 species, representing nearly half of …
Improved Sequence Mapping Using A Complete Reference Genome And Lift-Over, Nae-Chyun Chen, Luis F Paulin, Fritz J Sedlazeck, Sergey Koren, Adam M Phillippy, Ben Langmead
Improved Sequence Mapping Using A Complete Reference Genome And Lift-Over, Nae-Chyun Chen, Luis F Paulin, Fritz J Sedlazeck, Sergey Koren, Adam M Phillippy, Ben Langmead
Faculty, Staff and Students Publications
Complete, telomere-to-telomere genome assemblies promise improved analyses and the discovery of new variants, but many essential genomic resources remain associated with older reference genomes. Thus, there is a need to translate genomic features and read alignments between references. Here we describe a new method called levioSAM2 that accounts for reference changes and performs fast and accurate lift-over between assemblies using a whole-genome map. In addition to enabling the use of multiple references, we demonstrate that aligning reads to a high-quality reference (e.g. T2T-CHM13) and lifting to an older reference (e.g. GRCh38) actually improves the accuracy of the resulting variant calls …
Increased Coding Potential Of Bovine Herpesvirus 1, Victoria Jefferson
Increased Coding Potential Of Bovine Herpesvirus 1, Victoria Jefferson
Theses and Dissertations
Bovine respiratory disease (BRD) costs the cattle industry millions of dollars in costs in treatment and loss every year in the United States. A significant pathogen often contributes to BRD is Bovine Herpesvirus 1 (BoHV-1), a double stranded DNA virus with the ability to establish latency in the trigeminal ganglia and neurons. Primary infection with BoHV-1 results in immunosuppression that increases the risk of secondary bacterial infection and pneumonia. Because herpesviruses infect their hosts for life and can be reactivated in times of stress, BoHV-1 can present a recurring risk of BRD. The following research aims to expand the knowledge …
Whole Genome Analysis Of Snv And Indel Polymorphism In Common Marmosets (Callithrix Jacchus), R Alan Harris, Muthuswamy Raveendran, Wes Warren, Hillier W Ladeana, Chad Tomlinson, Tina Graves-Lindsay, Richard E Green, Jenna K Schmidt, Julia C Colwell, Allison T Makulec, Shelley A Cole, Ian H Cheeseman, Corinna N Ross, Saverio Capuano, Evan E Eichler, Jon E Levine, Jeffrey Rogers
Whole Genome Analysis Of Snv And Indel Polymorphism In Common Marmosets (Callithrix Jacchus), R Alan Harris, Muthuswamy Raveendran, Wes Warren, Hillier W Ladeana, Chad Tomlinson, Tina Graves-Lindsay, Richard E Green, Jenna K Schmidt, Julia C Colwell, Allison T Makulec, Shelley A Cole, Ian H Cheeseman, Corinna N Ross, Saverio Capuano, Evan E Eichler, Jon E Levine, Jeffrey Rogers
Faculty, Staff and Students Publications
The common marmoset (Callithrix jacchus) is one of the most widely used nonhuman primate models of human disease. Owing to limitations in sequencing technology, early genome assemblies of this species using short-read sequencing suffered from gaps. In addition, the genetic diversity of the species has not yet been adequately explored. Using long-read genome sequencing and expert annotation, we generated a high-quality genome resource creating a 2.898 Gb marmoset genome in which most of the euchromatin portion is assembled contiguously (contig N50 = 25.23 Mbp, scaffold N50 = 98.2 Mbp). We then performed whole genome sequencing on 84 marmosets …
Complex Evolutionary History With Extensive Ancestral Gene Flow In An African Primate Radiation, Axel Jensen, Frances Swift, Dorien De Vries, Robin M D Beck, Lukas F K Kuderna, Sascha Knauf, Idrissa S Chuma, Julius D Keyyu, Andrew C Kitchener, Kyle Farh, Jeffrey Rogers, Tomas Marques-Bonet, Kate M Detwiler, Christian Roos, Katerina Guschanski
Complex Evolutionary History With Extensive Ancestral Gene Flow In An African Primate Radiation, Axel Jensen, Frances Swift, Dorien De Vries, Robin M D Beck, Lukas F K Kuderna, Sascha Knauf, Idrissa S Chuma, Julius D Keyyu, Andrew C Kitchener, Kyle Farh, Jeffrey Rogers, Tomas Marques-Bonet, Kate M Detwiler, Christian Roos, Katerina Guschanski
Faculty, Staff and Students Publications
Understanding the drivers of speciation is fundamental in evolutionary biology, and recent studies highlight hybridization as an important evolutionary force. Using whole-genome sequencing data from 22 species of guenons (tribe Cercopithecini), one of the world's largest primate radiations, we show that rampant gene flow characterizes their evolutionary history and identify ancient hybridization across deeply divergent lineages that differ in ecology, morphology, and karyotypes. Some hybridization events resulted in mitochondrial introgression between distant lineages, likely facilitated by cointrogression of coadapted nuclear variants. Although the genomic landscapes of introgression were largely lineage specific, we found that genes with immune functions were overrepresented …
Mosaic Chromosomal Alterations In Blood Across Ancestries Using Whole-Genome Sequencing, Yasminka A Jakubek, Ying Zhou, Adrienne Stilp, Jason Bacon, Justin W Wong, Zuhal Ozcan, Donna Arnett, Kathleen Barnes, Joshua C Bis, Eric Boerwinkle, Jennifer A Brody, April P Carson, Daniel I Chasman, Jiawen Chen, Michael Cho, Matthew P Conomos, Nancy Cox, Margaret F Doyle, Myriam Fornage, Xiuqing Guo, Sharon L R Kardia, Joshua P Lewis, Ruth J F Loos, Xiaolong Ma, Mitchell J Machiela, Taralynn M Mack, Rasika A Mathias, Braxton D Mitchell, Josyf C Mychaleckyj, Kari North, Nathan Pankratz, Patricia A Peyser, Michael H Preuss, Bruce Psaty, Laura M Raffield, Ramachandran S Vasan, Susan Redline, Stephen S Rich, Jerome I Rotter, Edwin K Silverman, Jennifer A Smith, Aaron P Smith, Margaret Taub, Kent D Taylor, Jeong Yun, Yun Li, Pinkal Desai, Alexander G Bick, Alexander P Reiner, Paul Scheet, Paul L Auer
Mosaic Chromosomal Alterations In Blood Across Ancestries Using Whole-Genome Sequencing, Yasminka A Jakubek, Ying Zhou, Adrienne Stilp, Jason Bacon, Justin W Wong, Zuhal Ozcan, Donna Arnett, Kathleen Barnes, Joshua C Bis, Eric Boerwinkle, Jennifer A Brody, April P Carson, Daniel I Chasman, Jiawen Chen, Michael Cho, Matthew P Conomos, Nancy Cox, Margaret F Doyle, Myriam Fornage, Xiuqing Guo, Sharon L R Kardia, Joshua P Lewis, Ruth J F Loos, Xiaolong Ma, Mitchell J Machiela, Taralynn M Mack, Rasika A Mathias, Braxton D Mitchell, Josyf C Mychaleckyj, Kari North, Nathan Pankratz, Patricia A Peyser, Michael H Preuss, Bruce Psaty, Laura M Raffield, Ramachandran S Vasan, Susan Redline, Stephen S Rich, Jerome I Rotter, Edwin K Silverman, Jennifer A Smith, Aaron P Smith, Margaret Taub, Kent D Taylor, Jeong Yun, Yun Li, Pinkal Desai, Alexander G Bick, Alexander P Reiner, Paul Scheet, Paul L Auer
Faculty, Staff and Student Publications
Megabase-scale mosaic chromosomal alterations (mCAs) in blood are prognostic markers for a host of human diseases. Here, to gain a better understanding of mCA rates in genetically diverse populations, we analyzed whole-genome sequencing data from 67,390 individuals from the National Heart, Lung, and Blood Institute Trans-Omics for Precision Medicine program. We observed higher sensitivity with whole-genome sequencing data, compared with array-based data, in uncovering mCAs at low mutant cell fractions and found that individuals of European ancestry have the highest rates of autosomal mCAs and the lowest rates of chromosome X mCAs, compared with individuals of African or Hispanic ancestry. …
Genomic Variant Benchmark: If You Cannot Measure It, You Cannot Improve It, Sina Majidian, Daniel Paiva Agustinho, Chen-Shan Chin, Fritz J Sedlazeck, Medhat Mahmoud
Genomic Variant Benchmark: If You Cannot Measure It, You Cannot Improve It, Sina Majidian, Daniel Paiva Agustinho, Chen-Shan Chin, Fritz J Sedlazeck, Medhat Mahmoud
Faculty, Staff and Students Publications
Genomic benchmark datasets are essential to driving the field of genomics and bioinformatics. They provide a snapshot of the performances of sequencing technologies and analytical methods and highlight future challenges. However, they depend on sequencing technology, reference genome, and available benchmarking methods. Thus, creating a genomic benchmark dataset is laborious and highly challenging, often involving multiple sequencing technologies, different variant calling tools, and laborious manual curation. In this review, we discuss the available benchmark datasets and their utility. Additionally, we focus on the most recent benchmark of genes with medical relevance and challenging genomic complexity.
Rare Variants Found In Clinical Gene Panels Illuminate The Genetic And Allelic Architecture Of Orofacial Clefting, Kimberly K Diaz Perez, Sarah W Curtis, Alba Sanchis-Juan, Xuefang Zhao, Taylor Head, Samantha Ho, Bridget Carter, Toby Mchenry, Madison R Bishop, Luz C Valencia-Ramirez, Claudia Restrepo, Jacqueline T Hecht, Lina M Uribe, George Wehby, Seth M Weinberg, Terri H Beaty, Jeffrey C Murray, Eleanor Feingold, Mary L Marazita, David J Cutler, Michael P Epstein, Harrison Brand, Elizabeth J Leslie
Rare Variants Found In Clinical Gene Panels Illuminate The Genetic And Allelic Architecture Of Orofacial Clefting, Kimberly K Diaz Perez, Sarah W Curtis, Alba Sanchis-Juan, Xuefang Zhao, Taylor Head, Samantha Ho, Bridget Carter, Toby Mchenry, Madison R Bishop, Luz C Valencia-Ramirez, Claudia Restrepo, Jacqueline T Hecht, Lina M Uribe, George Wehby, Seth M Weinberg, Terri H Beaty, Jeffrey C Murray, Eleanor Feingold, Mary L Marazita, David J Cutler, Michael P Epstein, Harrison Brand, Elizabeth J Leslie
Faculty, Staff and Student Publications
PURPOSE: Orofacial clefts (OFCs) are common birth defects including cleft lip, cleft lip and palate, and cleft palate. OFCs have heterogeneous etiologies, complicating clinical diagnostics because it is not always apparent if the cause is Mendelian, environmental, or multifactorial. Sequencing is not currently performed for isolated or sporadic OFCs; therefore, we estimated the diagnostic yield for 418 genes in 841 cases and 294 controls.
METHODS: We evaluated 418 genes using genome sequencing and curated variants to assess their pathogenicity using American College of Medical Genetics criteria.
RESULTS: 9.04% of cases and 1.02% of controls had "likely pathogenic" variants (P < .0001), which was almost exclusively driven by heterozygous variants in autosomal genes. Cleft palate (17.6%) and cleft lip and palate (9.09%) cases had the highest yield, whereas cleft lip cases had a 2.80% yield. Out of 39 genes with likely pathogenic variants, 9 genes, including CTNND1 and IRF6, accounted for more than half of the yield (4.64% of cases). Most variants (61.8%) were "variants of uncertain significance", occurring more frequently in cases (P = .004), but no individual gene showed a significant excess of variants of uncertain significance.
CONCLUSION: …