Open Access. Powered by Scholars. Published by Universities.®

Computational Biology Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 15 of 15

Full-Text Articles in Computational Biology

Toward Interpretable Multi-Omics Multimodal Biomedical Artificial Intelligence, Yanjun Lyu Jan 2026

Toward Interpretable Multi-Omics Multimodal Biomedical Artificial Intelligence, Yanjun Lyu

Computer Science and Engineering Dissertations

The complexity of human disease arises from biological processes that unfold across multiple scales, from molecular variation through cellular function, tissue organisation, brain phenotypes, each of which is associated with distinct measurement modalities, regularities, and characteristic. Contemporary biomedical artificial intelligence has brought the opportunity to reveal the complexity with in; however, its methodological default, in which models are trained on most readily available modality, does not adequately engage with the multi-scale connected structure by which biological meaning is constituted. The research area of multi-omics and multi-modal AI for biomedicine remains at an early exploratory stage, and the work presented in …


Attention-Based Multi-Omics Fusion For Drug Synergy Prediction, Kusal Debnath, Pratip Rana, Preetam Ghosh Jan 2026

Attention-Based Multi-Omics Fusion For Drug Synergy Prediction, Kusal Debnath, Pratip Rana, Preetam Ghosh

Computer Science Faculty Publications

Drug combination therapy in disease management gained popularity in the last few decades. Computational modeling of such combinations is an active area of research in the drug discovery domain. While earlier approaches solely emphasized on the structural features of participating drugs for designing synergistic models, they lack other crucial factors directly linked with drug administration - omics expressions. As differential omics expression is a downstream consequence of the administered drug combinations, utilizing such expressions while designing synergistic models promises robust and dynamic modeling. In this work, we propose SynergyLM that fuses multi-omics features with drug embeddings to build an omics-aware …


Ai-Powered Multi-Omics Integration For Predictive Modeling Of Genotype-Environment-Phenotype Relationships, You Wu Sep 2025

Ai-Powered Multi-Omics Integration For Predictive Modeling Of Genotype-Environment-Phenotype Relationships, You Wu

Dissertations, Theses, and Capstone Projects

This dissertation presents a series of machine learning frameworks for modeling genotype–environment–phenotype relationships through integrative predictive modeling of multi-omics data. The work addresses three major axes of biological complexity: modeling biological information transmission cross-levels from genes to proteins to phenotypes, predicting molecular features cross-scale from cells to tissues to organisms, and translating phenotypes cross-species from model systems to humans. Each proposed method also tackles key machine learning (ML) challenges in the biomedical domain, including data scarcity, domain shift, out-of-distribution (OOD) generalization, and hierarchical modeling. Specifically, this dissertation introduces five novel deep learning algorithms: MultiDCP predicts drug-induced transcriptomic and viability responses …


Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh Apr 2025

Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh

Computer Science ETDs

Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …


Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh Jan 2025

Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh

Computer Science Faculty Publications

Drug–target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug–target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance …


Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang Jan 2025

Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang

Computer Science Faculty Publications

Motivation: Protein-protein interactions (PPIs) are fundamental aspects in understanding biological processes. Accurately predicting the effects of mutations on PPIs remains a critical requirement for drug design and disease mechanistic studies. Recently, deep learning models using protein 3D structures have become predominant for predicting mutation effects. However, significant challenges remain in practical applications, in part due to the considerable disparity in generalization capabilities between easy and hard mutations. Specifically, a hard mutation is defined as one with its maximum TM-score < 0.6 when compared to the training set. Additionally, compared to physics-based approaches, deep learning models may overestimate performance due to potential data leakage.

Results: We propose new training/test splits that mitigate data leakage according to the CATH homologous superfamily. Under the constraints of physical …


A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh Jan 2025

A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh

Computer Science Faculty Publications

Conventional drug discovery is expensive, time-consuming, and prone to failure. Artificial intelligence has become a potent substitute over the last decade, providing strong answers to challenging biological issues in this field. Among these difficulties, drug-target binding (DTB) is a key component of drug discovery techniques. In this context, drug-target affinity and drug–target interaction are complementary and essential frameworks that work together to improve our comprehension of DTB dynamics. In this work, we thoroughly analyze the most recent deep learning models, popular benchmark datasets, and assessment metrics for DTB prediction. We look at the paradigm shift in the development of drug …


Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He Jan 2025

Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He

Computer Science Faculty Publications

DeepSSETracer is a method for segmenting protein secondary structure from medium-resolution (5-10Å) cryogenic electron microscopy (cryo-EM) density maps. We conducted experiments and ablation studies to examine the effects of normalization methods, max-pooling, activation functions, and loss calculation region on DeepSSETracer. By combining multiple technical improvements, the performance of the new version, DeepSSETracer 2.0, was significantly enhanced compared to DeepSSETracer 1.1. On a set of 77 test cases, the weighted average per-voxel F1 score increased from 62.1% to 70.3% for helix detection, and from 47.8% to 62.5% for β-sheet detection. While each of the five modifications in the network enhanced the …


Model-Based Deep Autoencoders For Clustering Single-Cell Rna Sequencing Data With Side Information, Xiang Lin Dec 2023

Model-Based Deep Autoencoders For Clustering Single-Cell Rna Sequencing Data With Side Information, Xiang Lin

Dissertations

Clustering analysis has been conducted extensively in single-cell RNA sequencing (scRNA-seq) studies. scRNA-seq can profile tens of thousands of genes' activities within a single cell. Thousands or tens of thousands of cells can be captured simultaneously in a typical scRNA-seq experiment. Biologists would like to cluster these cells for exploring and elucidating cell types or subtypes. Numerous methods have been designed for clustering scRNA-seq data. Yet, single-cell technologies develop so fast in the past few years that those existing methods do not catch up with these rapid changes and fail to fully fulfil their potential. For instance, besides profiling transcription …


Missing Value Imputation For Single Omics And Multi-Omics Data, Meng Song Jul 2023

Missing Value Imputation For Single Omics And Multi-Omics Data, Meng Song

Dissertations

The integration analyses of multi-omics data have the advantages of extending our understanding of biological system across multiple omics layers, unraveling the functional mechanism of complex disease development, and refining the discovery of novel drug targets. However, multi-omics studies often face challenges such as data heterogeneity, missing values problem, interpretability, and imbalance classes. Among these challenges, the missing values problem is a critical issue for large cohort studies as not all samples will get a complete measurement for all the omics layers. To address the problem of missing values in multi-omics data, I focused on the imputation of completely missing …


An Approach To Developing Benchmark Datasets For Protein Secondary Structure Segmentation From Cryo-Em Density Maps, Thu Nguyen, Yongcheng Mu, Jiangwen Sun, Jing He Jan 2023

An Approach To Developing Benchmark Datasets For Protein Secondary Structure Segmentation From Cryo-Em Density Maps, Thu Nguyen, Yongcheng Mu, Jiangwen Sun, Jing He

Computer Science Faculty Publications

More and more deep learning approaches have been proposed to segment secondary structures from cryo-electron density maps at medium resolution range (5--10Å). Although the deep learning approaches show great potential, only a few small experimental data sets have been used to test the approaches. There is limited understanding about potential factors, in data, that affect the performance of segmentation. We propose an approach to generate data sets with desired specifications in three potential factors - the protein sequence identity, structural contents, and data quality. The approach was implemented and has generated a test set and various training sets to study …


Evaluation Of Deep Neural Network Prospr For Accurate Protein Distance Predictions On Casp14 Targets, Jacob A. Stern, Bryce Eric Hedelius, Olivia Fisher, Wendy M. Billings, Dennis Della Corte Nov 2021

Evaluation Of Deep Neural Network Prospr For Accurate Protein Distance Predictions On Casp14 Targets, Jacob A. Stern, Bryce Eric Hedelius, Olivia Fisher, Wendy M. Billings, Dennis Della Corte

Faculty Publications

The field of protein structure prediction has recently been revolutionized through the introduction of deep learning. The current state-of-the-art tool AlphaFold2 can predict highly accurate structures; however, it has a prohibitively long inference time for applications that require the folding of hundreds of sequences. The prediction of protein structure annotations, such as amino acid distances, can be achieved at a higher speed with existing tools, such as the ProSPr network. Here, we report on important updates to the ProSPr network, its performance in the recent Critical Assessment of Techniques for Protein Structure Prediction (CASP14) competition, and an evaluation of its …


The Whole Is Greater Than Its Parts: Ensembling Improves Protein Contact Prediction, Wendy M. Billings, Connor J. Morris, Dennis Della Corte Apr 2021

The Whole Is Greater Than Its Parts: Ensembling Improves Protein Contact Prediction, Wendy M. Billings, Connor J. Morris, Dennis Della Corte

Faculty Publications

The prediction of amino acid contacts from protein sequence is an important problem, as protein contacts are a vital step towards the prediction of folded protein structures. We propose that a powerful concept from deep learning, called ensembling, can increase the accuracy of protein contact predictions by combining the outputs of different neural network models. We show that ensembling the predictions made by different groups at the recent Critical Assessment of Protein Structure Prediction (CASP13) outperforms all individual groups. Further, we show that contacts derived from the distance predictions of three additional deep neural networks—AlphaFold, trRosetta, and ProSPr—can be substantially …


Deepep: A Deep Learning Framework For Identifying Essential Proteins, Min Zeng, Min Li, Fang-Xiang Wu, Yaohang Li, Yi Pan Dec 2019

Deepep: A Deep Learning Framework For Identifying Essential Proteins, Min Zeng, Min Li, Fang-Xiang Wu, Yaohang Li, Yi Pan

Computer Science Faculty Publications

Background: Essential proteins are crucial for cellular life and thus, identification of essential proteins is an important topic and a challenging problem for researchers. Recently lots of computational approaches have been proposed to handle this problem. However, traditional centrality methods cannot fully represent the topological features of biological networks. In addition, identifying essential proteins is an imbalanced learning problem; but few current shallow machine learning-based methods are designed to handle the imbalanced characteristics. Results: We develop DeepEP based on a deep learning framework that uses the node2vec technique, multi-scale convolutional neural networks and a sampling technique to identify essential proteins. …


Recurrent Neural Networks And Their Applications To Rna Secondary Structure Inference, Devin Willmott Jan 2018

Recurrent Neural Networks And Their Applications To Rna Secondary Structure Inference, Devin Willmott

Theses and Dissertations--Mathematics

Recurrent neural networks (RNNs) are state of the art sequential machine learning tools, but have difficulty learning sequences with long-range dependencies due to the exponential growth or decay of gradients backpropagated through the RNN. Some methods overcome this problem by modifying the standard RNN architecure to force the recurrent weight matrix W to remain orthogonal throughout training. The first half of this thesis presents a novel orthogonal RNN architecture that enforces orthogonality of W by parametrizing with a skew-symmetric matrix via the Cayley transform. We present rules for backpropagation through the Cayley transform, show how to deal with the Cayley …