Open Access. Powered by Scholars. Published by Universities.®

Genetic Structures Commons

Open Access. Powered by Scholars. Published by Universities.®

Computational Biology

Publication Year

Articles 1 - 13 of 13

Full-Text Articles in Genetic Structures

Dna Methylation And Machine Learning: Challenges And Perspective Toward Enhanced Clinical Diagnostics, Erfan Aref-Eshghi, Arash B Abadi, Mohammad-Erfan Farhadieh, Amirreza Hooshmand, Fatemeh Ghasemi, Leila Youssefian, Hassan Vahidnezhad, Taylor Martin Kerrins, Xiaonan Zhao, Mahdi Akbarzadeh, Hakon Hakonarson, Amir Hossein Saeidian Oct 2025

Dna Methylation And Machine Learning: Challenges And Perspective Toward Enhanced Clinical Diagnostics, Erfan Aref-Eshghi, Arash B Abadi, Mohammad-Erfan Farhadieh, Amirreza Hooshmand, Fatemeh Ghasemi, Leila Youssefian, Hassan Vahidnezhad, Taylor Martin Kerrins, Xiaonan Zhao, Mahdi Akbarzadeh, Hakon Hakonarson, Amir Hossein Saeidian

Faculty, Staff and Students Publications

DNA methylation is an epigenetic modification that regulates gene expression by adding methyl groups to DNA, affecting cellular function and disease development. Machine learning, a subset of artificial intelligence, analyzes large datasets to identify patterns and make predictions. Over the past two decades, advances in bioinformatics technologies for arrays and sequencing have generated vast amounts of data, leading to the widespread adoption of machine learning methods for analyzing complex biological information for medical problems. This review explores recent advancements in DNA methylation studies that leverage emerging machine learning techniques for more precise, comprehensive, and rapid patient diagnostics based on DNA …


Comprehensive Evaluation Of Phosphoproteomic-Based Kinase Activity Inference, Sophia Müller-Dott, Eric J Jaehnig, Khoi Pham Munchic, Wen Jiang, Tomer M Yaron-Barir, Sara R Savage, Martin Garrido-Rodriguez, Jared L Johnson, Alessandro Lussana, Evangelia Petsalaki, Jonathan T Lei, Aurelien Dugourd, Karsten Krug, Lewis C Cantley, D R Mani, Bing Zhang, Julio Saez-Rodriguez May 2025

Comprehensive Evaluation Of Phosphoproteomic-Based Kinase Activity Inference, Sophia Müller-Dott, Eric J Jaehnig, Khoi Pham Munchic, Wen Jiang, Tomer M Yaron-Barir, Sara R Savage, Martin Garrido-Rodriguez, Jared L Johnson, Alessandro Lussana, Evangelia Petsalaki, Jonathan T Lei, Aurelien Dugourd, Karsten Krug, Lewis C Cantley, D R Mani, Bing Zhang, Julio Saez-Rodriguez

Faculty, Staff and Students Publications

Kinases regulate cellular processes and are essential for understanding cellular function and disease. To investigate the regulatory state of a kinase, numerous methods have been developed to infer kinase activities from phosphoproteomics data using kinase-substrate libraries. However, few phosphorylation sites can be attributed to an upstream kinase in these libraries, limiting the scope of kinase activity inference. Moreover, inferred activities vary across methods, necessitating evaluation for accurate interpretation. Here, we present benchmarKIN, an R package enabling comprehensive evaluation of kinase activity inference methods. Alongside classical perturbation experiments, benchmarKIN introduces a tumor-based benchmarking approach utilizing multi-omics data to identify highly active …


Closing The Gaps, And Improving Somatic Structural Variant Analysis And Benchmarking Using Chm13-T2t, Luis F Paulin, Jeremy Fan, Kieran O'Neill, Erin Pleasance, Vanessa L Porter, Steven J M Jones, Fritz J Sedlazeck Apr 2025

Closing The Gaps, And Improving Somatic Structural Variant Analysis And Benchmarking Using Chm13-T2t, Luis F Paulin, Jeremy Fan, Kieran O'Neill, Erin Pleasance, Vanessa L Porter, Steven J M Jones, Fritz J Sedlazeck

Faculty, Staff and Students Publications

The complexities of cancer genomes are becoming more easily interpreted due to advancements in sequencing technologies and improved bioinformatic analysis. Structural variants (SVs) represent an important subset of somatic events in tumors. While the detection of SVs has been markedly improved by the development of long-read sequencing, somatic variant identification and annotation remain challenging. We hypothesized that the use of a completed human reference genome (CHM13-T2T) would improve somatic SV calling. Our findings in a tumor-normal matched benchmark sample and three patient samples show that the CHM13-T2T improves SV detection accuracy compared to GRCh38 with a notable reduction in false-positive …


A Hitchhiker’S Guide To Long-Read Genomic Analysis, Medhat Mahmoud, Daniel P Agustinho, Fritz J Sedlazeck Apr 2025

A Hitchhiker’S Guide To Long-Read Genomic Analysis, Medhat Mahmoud, Daniel P Agustinho, Fritz J Sedlazeck

Faculty, Staff and Students Publications

Over the past decade, long-read sequencing has evolved into a pivotal technology for uncovering the hidden and complex regions of the genome. Significant cost efficiency, scalability, and accuracy advancements have driven this evolution. Concurrently, novel analytical methods have emerged to harness the full potential of long reads. These advancements have enabled milestones such as the first fully completed human genome, enhanced identification and understanding of complex genomic variants, and deeper insights into the interplay between epigenetics and genomic variation. This mini-review provides a comprehensive overview of the latest developments in long-read DNA sequencing analysis, encompassing reference-based and de novo assembly …


Unraveling The Hidden Complexity Of Cancer Through Long-Read Sequencing, Qiuhui Li, Ayse G Keskus, Justin Wagner, Michal B Izydorczyk, Winston Timp, Fritz J Sedlazeck, Alison P Klein, Justin M Zook, Mikhail Kolmogorov, Michael C Schatz Apr 2025

Unraveling The Hidden Complexity Of Cancer Through Long-Read Sequencing, Qiuhui Li, Ayse G Keskus, Justin Wagner, Michal B Izydorczyk, Winston Timp, Fritz J Sedlazeck, Alison P Klein, Justin M Zook, Mikhail Kolmogorov, Michael C Schatz

Faculty, Staff and Students Publications

Cancer is fundamentally a disease of the genome, characterized by extensive genomic, transcriptomic, and epigenomic alterations. Most current studies predominantly use short-read sequencing, gene panels, or microarrays to explore these alterations; however, these technologies can systematically miss or misrepresent certain types of alterations, especially structural variants, complex rearrangements, and alterations within repetitive regions. Long-read sequencing is rapidly emerging as a transformative technology for cancer research by providing a comprehensive view across the genome, transcriptome, and epigenome, including the ability to detect alterations that previous technologies have overlooked. In this Perspective, we explore the current applications of long-read sequencing for both …


Playbook Workflow Builder: Interactive Construction Of Bioinformatics Workflows, Daniel J B Clarke, John Erol Evangelista, Zhuorui Xie, Giacomo B Marino, Anna I Byrd, Mano R Maurya, Sumana Srinivasan, Keyang Yu, Varduhi Petrosyan, Matthew E Roth, Miroslav Milinkov, Charles Hadley King, Jeet Kiran Vora, Jonathon Keeney, Christopher Nemarich, William Khan, Alexander Lachmann, Nasheath Ahmed, Alexandra Agris, Juncheng Pan, Srinivasan Ramachandran, Eoin Fahy, Emmanuel Esquivel, Aleksandar Mihajlovic, Bosko Jevtic, Vuk Milinovic, Sean Kim, Patrick Mcneely, Tianyi Wang, Eric Wenger, Miguel A Brown, Alexander Sickler, Yuankun Zhu, Sherry L Jenkins, Philip D Blood, Deanne M Taylor, Adam C Resnick, Raja Mazumder, Aleksandar Milosavljevic, Shankar Subramaniam, Avi Ma'ayan Apr 2025

Playbook Workflow Builder: Interactive Construction Of Bioinformatics Workflows, Daniel J B Clarke, John Erol Evangelista, Zhuorui Xie, Giacomo B Marino, Anna I Byrd, Mano R Maurya, Sumana Srinivasan, Keyang Yu, Varduhi Petrosyan, Matthew E Roth, Miroslav Milinkov, Charles Hadley King, Jeet Kiran Vora, Jonathon Keeney, Christopher Nemarich, William Khan, Alexander Lachmann, Nasheath Ahmed, Alexandra Agris, Juncheng Pan, Srinivasan Ramachandran, Eoin Fahy, Emmanuel Esquivel, Aleksandar Mihajlovic, Bosko Jevtic, Vuk Milinovic, Sean Kim, Patrick Mcneely, Tianyi Wang, Eric Wenger, Miguel A Brown, Alexander Sickler, Yuankun Zhu, Sherry L Jenkins, Philip D Blood, Deanne M Taylor, Adam C Resnick, Raja Mazumder, Aleksandar Milosavljevic, Shankar Subramaniam, Avi Ma'ayan

Faculty, Staff and Students Publications

The Playbook Workflow Builder (PWB) is a web-based platform to dynamically construct and execute bioinformatics workflows by utilizing a growing network of input datasets, semantically annotated API endpoints, and data visualization tools contributed by an ecosystem of collaborators. Via a user-friendly user interface, workflows can be constructed from contributed building-blocks without technical expertise. The output of each step of the workflow is added into reports containing textual descriptions, figures, tables, and references. To construct workflows, users can click on cards that represent each step in a workflow, or construct workflows via a chat interface that is assisted by a large …


Evaluating Predictors Of Kinase Activity Of Stk11 Variants Identified In Primary Human Non-Small Cell Lung Cancers, Yile Chen, Kyoungyeul Lee, Junwoo Woo, Dong-Wook Kim, Changwon Keum, Giulia Babbi, Rita Casadio, Pier Luigi Martelli, Castrense Savojardo, Matteo Manfredi, Yang Shen, Yuanfei Sun, Panagiotis Katsonis, Olivier Lichtarge, Vikas Pejaver, David J Seward, Akash Kamandula, Constantina Bakolitsa, Steven E Brenner, Predrag Radivojac, Anne O'Donnell-Luria, Sean D Mooney, Shantanu Jain Mar 2025

Evaluating Predictors Of Kinase Activity Of Stk11 Variants Identified In Primary Human Non-Small Cell Lung Cancers, Yile Chen, Kyoungyeul Lee, Junwoo Woo, Dong-Wook Kim, Changwon Keum, Giulia Babbi, Rita Casadio, Pier Luigi Martelli, Castrense Savojardo, Matteo Manfredi, Yang Shen, Yuanfei Sun, Panagiotis Katsonis, Olivier Lichtarge, Vikas Pejaver, David J Seward, Akash Kamandula, Constantina Bakolitsa, Steven E Brenner, Predrag Radivojac, Anne O'Donnell-Luria, Sean D Mooney, Shantanu Jain

Faculty, Staff and Students Publications

Critical evaluation of computational tools for predicting variant effects is important considering their increased use in disease diagnosis and driving molecular discoveries. In the sixth edition of the Critical Assessment of Genome Interpretation (CAGI) challenge, a dataset of 28 STK11 rare variants (27 missense, 1 single amino acid deletion), identified in primary non-small cell lung cancer biopsies, was experimentally assayed to characterize computational methods from four participating teams and five publicly available tools. Predictors demonstrated a high level of performance on key evaluation metrics, measuring correlation with the assay outputs and separating loss-of-function (LoF) variants from wildtype-like (WT-like) variants. The …


Cagi6 Id Panel Challenge: Assessment Of Phenotype And Variant Predictions In 415 Children With Neurodevelopmental Disorders (Ndds), Maria Cristina Aspromonte, Alessio Del Conte, Shaowen Zhu, Wuwei Tan, Yang Shen, Yexian Zhang, Qi Li, Maggie Haitian Wang, Giulia Babbi, Samuele Bovo, Pier Luigi Martelli, Rita Casadio, Azza Althagafi, Sumyyah Toonsi, Maxat Kulmanov, Robert Hoehndorf, Panagiotis Katsonis, Amanda Williams, Olivier Lichtarge, Su Xian, Wesley Surento, Vikas Pejaver, Sean D Mooney, Uma Sunderam, Rajgopal Srinivasan, Alessandra Murgia, Damiano Piovesan, Silvio C E Tosatto, Emanuela Leonardi Mar 2025

Cagi6 Id Panel Challenge: Assessment Of Phenotype And Variant Predictions In 415 Children With Neurodevelopmental Disorders (Ndds), Maria Cristina Aspromonte, Alessio Del Conte, Shaowen Zhu, Wuwei Tan, Yang Shen, Yexian Zhang, Qi Li, Maggie Haitian Wang, Giulia Babbi, Samuele Bovo, Pier Luigi Martelli, Rita Casadio, Azza Althagafi, Sumyyah Toonsi, Maxat Kulmanov, Robert Hoehndorf, Panagiotis Katsonis, Amanda Williams, Olivier Lichtarge, Su Xian, Wesley Surento, Vikas Pejaver, Sean D Mooney, Uma Sunderam, Rajgopal Srinivasan, Alessandra Murgia, Damiano Piovesan, Silvio C E Tosatto, Emanuela Leonardi

Faculty, Staff and Students Publications

The Genetics of Neurodevelopmental Disorders Lab in Padua provided a new intellectual disability (ID) Panel challenge for computational methods to predict patient phenotypes and their causal variants in the context of the Critical Assessment of the Genome Interpretation, 6th edition (CAGI6). Eight research teams submitted a total of 30 models to predict phenotypes based on the sequences of 74 genes (VCF format) in 415 pediatric patients affected by Neurodevelopmental Disorders (NDDs). NDDs are clinically and genetically heterogeneous conditions, with onset in infant age. Here, we assess the ability and accuracy of computational methods to predict comorbid phenotypes based on clinical …


Predicting The Impact Of Rare Variants On Rna Splicing In Cagi6, Jenny Lord, Carolina Jaramillo Oquendo, Htoo A Wai, Andrew G L Douglas, David J Bunyan, Yaqiong Wang, Zhiqiang Hu, Zishuo Zeng, Daniel Danis, Panagiotis Katsonis, Amanda Williams, Olivier Lichtarge, Yuchen Chang, Richard D Bagnall, Stephen M Mount, Brynja Matthiasardottir, Chiaofeng Lin, Thomas Van Overeem Hansen, Raphael Leman, Alexandra Martins, Claude Houdayer, Sophie Krieger, Constantina Bakolitsa, Yisu Peng, Akash Kamandula, Predrag Radivojac, Diana Baralle Mar 2025

Predicting The Impact Of Rare Variants On Rna Splicing In Cagi6, Jenny Lord, Carolina Jaramillo Oquendo, Htoo A Wai, Andrew G L Douglas, David J Bunyan, Yaqiong Wang, Zhiqiang Hu, Zishuo Zeng, Daniel Danis, Panagiotis Katsonis, Amanda Williams, Olivier Lichtarge, Yuchen Chang, Richard D Bagnall, Stephen M Mount, Brynja Matthiasardottir, Chiaofeng Lin, Thomas Van Overeem Hansen, Raphael Leman, Alexandra Martins, Claude Houdayer, Sophie Krieger, Constantina Bakolitsa, Yisu Peng, Akash Kamandula, Predrag Radivojac, Diana Baralle

Faculty, Staff and Students Publications

Variants which disrupt splicing are a frequent cause of rare disease that have been under-ascertained clinically. Accurate and efficient methods to predict a variant's impact on splicing are needed to interpret the growing number of variants of unknown significance (VUS) identified by exome and genome sequencing. Here, we present the results of the CAGI6 Splicing VUS challenge, which invited predictions of the splicing impact of 56 variants ascertained clinically and functionally validated to determine splicing impact. The performance of 12 prediction methods, along with SpliceAI and CADD, was compared on the 56 functionally validated variants. The maximum accuracy achieved was …


Assessing Predictions On Fitness Effects Of Missense Variants In Hmbs In Cagi6, Jing Zhang, Lisa Kinch, Panagiotis Katsonis, Olivier Lichtarge, Milind Jagota, Yun S Song, Yuanfei Sun, Yang Shen, Nurdan Kuru, Onur Dereli, Ogun Adebali, Muttaqi Ahmad Alladin, Debnath Pal, Emidio Capriotti, Maria Paola Turina, Castrense Savojardo, Pier Luigi Martelli, Giulia Babbi, Rita Casadio, Fabrizio Pucci, Marianne Rooman, Gabriel Cia, Matsvei Tsishyn, Alexey Strokach, Zhiqiang Hu, Warren Van Loggerenberg, Frederick P Roth, Predrag Radivojac, Steven E Brenner, Qian Cong, Nick V Grishin Mar 2025

Assessing Predictions On Fitness Effects Of Missense Variants In Hmbs In Cagi6, Jing Zhang, Lisa Kinch, Panagiotis Katsonis, Olivier Lichtarge, Milind Jagota, Yun S Song, Yuanfei Sun, Yang Shen, Nurdan Kuru, Onur Dereli, Ogun Adebali, Muttaqi Ahmad Alladin, Debnath Pal, Emidio Capriotti, Maria Paola Turina, Castrense Savojardo, Pier Luigi Martelli, Giulia Babbi, Rita Casadio, Fabrizio Pucci, Marianne Rooman, Gabriel Cia, Matsvei Tsishyn, Alexey Strokach, Zhiqiang Hu, Warren Van Loggerenberg, Frederick P Roth, Predrag Radivojac, Steven E Brenner, Qian Cong, Nick V Grishin

Faculty, Staff and Students Publications

This paper presents an evaluation of predictions submitted for the "HMBS" challenge, a component of the sixth round of the Critical Assessment of Genome Interpretation held in 2021. The challenge required participants to predict the effects of missense variants of the human HMBS gene on yeast growth. The HMBS enzyme, critical for the biosynthesis of heme in eukaryotic cells, is highly conserved among eukaryotes. Despite the application of a variety of algorithms and methods, the performance of predictors was relatively similar, with Kendall's tau correlation coefficients between predictions and experimental scores around 0.3 for a majority of submissions. Notably, the …


Genomics And Multiomics In The Age Of Precision Medicine, Srinivasan Mani, Seema R Lalani, Mohan Pammi Mar 2025

Genomics And Multiomics In The Age Of Precision Medicine, Srinivasan Mani, Seema R Lalani, Mohan Pammi

Faculty, Staff and Students Publications

Precision medicine is a transformative healthcare model that utilizes an understanding of a person's genome, environment, lifestyle, and interplay to deliver customized healthcare. Precision medicine has the potential to improve the health and productivity of the population, enhance patient trust and satisfaction in healthcare, and accrue health cost-benefits both at an individual and population level. Through faster and cost-effective genomics data, next-generation sequencing has provided us the impetus to understand the nuances of complex interactions between genes, diet, and lifestyle that are heterogeneous across the population. The emergence of multiomics technologies, including transcriptomics, proteomics, epigenomics, metabolomics, and microbiomics, has enhanced …


Unveiling Microbial Diversity: Harnessing Long-Read Sequencing Technology, Daniel P Agustinho, Yilei Fu, Vipin K Menon, Ginger A Metcalf, Todd J Treangen, Fritz J Sedlazeck Jun 2024

Unveiling Microbial Diversity: Harnessing Long-Read Sequencing Technology, Daniel P Agustinho, Yilei Fu, Vipin K Menon, Ginger A Metcalf, Todd J Treangen, Fritz J Sedlazeck

Faculty, Staff and Students Publications

Long-read sequencing has recently transformed metagenomics, enhancing strain-level pathogen characterization, enabling accurate and complete metagenome-assembled genomes, and improving microbiome taxonomic classification and profiling. These advancements are not only due to improvements in sequencing accuracy, but also happening across rapidly changing analysis methods. In this Review, we explore long-read sequencing's profound impact on metagenomics, focusing on computational pipelines for genome assembly, taxonomic characterization and variant detection, to summarize recent advancements in the field and provide an overview of available analytical methods to fully leverage long reads. We provide insights into the advantages and disadvantages of long reads over short reads and …


Fam20a: A Potential Diagnostic Biomarker For Lung Squamous Cell Carcinoma, Yalin Zhang, Qin Sun, Yangbo Liang, Xian Yang, Hailian Wang, Siyuan Song, Yi Wang, Yong Feng Jan 2024

Fam20a: A Potential Diagnostic Biomarker For Lung Squamous Cell Carcinoma, Yalin Zhang, Qin Sun, Yangbo Liang, Xian Yang, Hailian Wang, Siyuan Song, Yi Wang, Yong Feng

Faculty, Staff and Students Publications

Background: Lung squamous cell carcinoma (LUSC) ranks among the carcinomas with the highest incidence and dismal survival rates, suffering from a lack of effective therapeutic strategies. Consequently, biomarkers facilitating early diagnosis of LUSC could significantly enhance patient survival. This study aims to identify novel biomarkers for LUSC.

Methods: Utilizing the TCGA, GTEx, and CGGA databases, we focused on the gene encoding Family with Sequence Similarity 20, Member A (FAM20A) across various cancers. We then corroborated these bioinformatic predictions with clinical samples. A range of analytical tools, including Kaplan-Meier, MethSurv database, Wilcoxon rank-sum, Kruskal-Wallis tests, Gene Set Enrichment Analysis, …