Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Old Dominion University (23)
- University of Nebraska - Lincoln (8)
- Swarthmore College (6)
- Brigham Young University (3)
- City University of New York (CUNY) (3)
-
- Loyola University Chicago (3)
- University of Kentucky (3)
- University of Montana (3)
- Illinois Math and Science Academy (2)
- LSU New Orleans (2)
- Missouri University of Science and Technology (2)
- New Jersey Institute of Technology (2)
- University of Louisville (2)
- Association of Arab Universities (1)
- COBRA (1)
- California Polytechnic State University, San Luis Obispo (1)
- California State University, San Bernardino (1)
- Clemson University (1)
- Connecticut College (1)
- Dartmouth College (1)
- Duke Law (1)
- Georgia Southern University (1)
- Gonzaga University (1)
- Indiana State University (1)
- Louisiana State University (1)
- Munster Technological University (1)
- Purdue University (1)
- The Texas Medical Center Library (1)
- University of Missouri, St. Louis (1)
- University of Nebraska at Omaha (1)
- Keyword
-
- Bioinformatics (11)
- Deep learning (11)
- Gene expression (6)
- Machine Learning (6)
- Machine learning (6)
-
- Computational biology (5)
- Genomics (5)
- Humans (5)
- Neural networks (5)
- Protein (5)
- Clustering (4)
- Evolution (4)
- Amino acid (3)
- Data mining (3)
- Genetics (3)
- Genome (3)
- RNA-seq (3)
- Secondary structure (3)
- Transcriptome (3)
- Transcriptomics (3)
- Academic subjects (2)
- Algorithm (2)
- Applied computing (2)
- Artificial Intelligence (2)
- Artificial intelligence (2)
- Benchmarking (2)
- Big data (2)
- Computational biology/methods (2)
- Computer (2)
- Computer science (2)
- Publication Year
- Publication
-
- Computer Science Faculty Publications (20)
- Computer Science Faculty Works (3)
- Faculty Publications (3)
- School of Computing: Faculty Publications (3)
- Bioinformatics Faculty Publications (2)
-
- Computer Science Theses & Dissertations (2)
- Digital Humanities Curricular Development (2)
- Dissertations (2)
- Electronic Theses and Dissertations (2)
- Engineering Management and Systems Engineering Faculty Research & Creative Works (2)
- Graduate Student Theses, Dissertations, & Professional Papers (2)
- LSU New Orleans Theses and Dissertations (2)
- Publications and Research (2)
- School of Computing: Conference and Workshop Papers (2)
- School of Computing: Dissertations, Theses, and Student Research (2)
- Student Publications & Research (2)
- Theses and Dissertations--Computer Science (2)
- All Dissertations (1)
- All-Inclusive List of Electronic Theses and Dissertations (1)
- Biological Sciences Faculty Publications (1)
- Biology Faculty Works (1)
- COBRA Preprint Series (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computer Science ETDs (1)
- Computer Science Faculty Scholarship (1)
- Computer Science Honors Papers (1)
- Computer Science Senior Theses (1)
- Computer Science and Engineering Dissertations (1)
- Computer Science: Faculty Publications and Other Works (1)
- Department of Computer Electronics and Engineering: Dissertations, Theses, and Student Research (1)
- Publication Type
Articles 1 - 30 of 86
Full-Text Articles in Computational Biology
Individualized Bayesian Inference Identifies Novel Genetic Variants For Parkinson's Disease, Jin Ren, Yasaman J. Soofi, Md Asad Rahman, Qing Lu, Jinling Liu
Individualized Bayesian Inference Identifies Novel Genetic Variants For Parkinson's Disease, Jin Ren, Yasaman J. Soofi, Md Asad Rahman, Qing Lu, Jinling Liu
Engineering Management and Systems Engineering Faculty Research & Creative Works
Parkinson's disease (PD) is a complex neurodegenerative disorder with a significant genetic component. While genome-wide association studies (GWAS) have been instrumental in identifying genetic variants associated with PD, the reliance on large sample sizes and population-level analyses may overlook variants with lower minor allele frequencies or individual-specific relevance. Individualized Bayesian Inference (IBI) offers a promising method to complement GWAS by identifying and prioritizing candidate genetic markers at both the individual and patients-like-me subgroup levels. This study evaluates the application of IBI to PD genetics, using GWAS as a baseline for comparison. We analyzed genetic data from the Fox Insight online …
Biologically Informed Negative Samplingfor Antibody Chain Pairing Classification, Ishita Singh
Biologically Informed Negative Samplingfor Antibody Chain Pairing Classification, Ishita Singh
Computer Science Senior Theses
Antibody heavy and light chain (H/L) pairing is fundamental to antigen recognition and stability. While single-cell sequencing preserves native pairing information, widely used bulk repertoire and spatial transcriptomics platforms do not, motivating the need for efficient ML methods to infer H/L pairing. Training a binary classifier for this task faces the methodological challenge of a lack of true biological negatives, since natural selection eliminates B cells with incompatible H/L pairs.
In this thesis, I introduce a biologically informed negative sampling strategy for H/L pairing classification, drawing on known V-gene biases in heavy and light chain pairing. Pseudo-negatives are constructed by …
Toward Interpretable Multi-Omics Multimodal Biomedical Artificial Intelligence, Yanjun Lyu
Toward Interpretable Multi-Omics Multimodal Biomedical Artificial Intelligence, Yanjun Lyu
Computer Science and Engineering Dissertations
The complexity of human disease arises from biological processes that unfold across multiple scales, from molecular variation through cellular function, tissue organisation, brain phenotypes, each of which is associated with distinct measurement modalities, regularities, and characteristic. Contemporary biomedical artificial intelligence has brought the opportunity to reveal the complexity with in; however, its methodological default, in which models are trained on most readily available modality, does not adequately engage with the multi-scale connected structure by which biological meaning is constituted. The research area of multi-omics and multi-modal AI for biomedicine remains at an early exploratory stage, and the work presented in …
Privacy-Preserving Federated Learning With Optimized Ensemble Weighting And Knowledge Distillation For Covid-19 Detection From Non-Iid Medical Imaging Data, Richard Annan, Hong Qin, Robert Newman, Madhuri Siddula, Letu Qingge
Privacy-Preserving Federated Learning With Optimized Ensemble Weighting And Knowledge Distillation For Covid-19 Detection From Non-Iid Medical Imaging Data, Richard Annan, Hong Qin, Robert Newman, Madhuri Siddula, Letu Qingge
Computer Science Faculty Publications
Medical imaging enables rapid and accurate diagnosis of COVID-19, with CT scans proving especially effective. However, data privacy concerns limit collaborative model development across hospitals. To address this issue, we introduce a novel federated learning framework. It is referred to as Independent Knowledge Distillation with post-Ensemble Federated Learning (IKDEFL). Differential Privacy (DP) is integrated into the framework to improve privacy guarantees. Three DP mechanisms are evaluated. These include Fixed Gaussian, Gaussian Adaptive, and Tree Adaptive. The evaluation has been conducted on heterogeneous and Non-Independent and Identically Distributed (Non-IID) datasets. These datasets reflect real-world hospital scenarios. Results show that IKDEFL significantly …
Attention-Based Multi-Omics Fusion For Drug Synergy Prediction, Kusal Debnath, Pratip Rana, Preetam Ghosh
Attention-Based Multi-Omics Fusion For Drug Synergy Prediction, Kusal Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Drug combination therapy in disease management gained popularity in the last few decades. Computational modeling of such combinations is an active area of research in the drug discovery domain. While earlier approaches solely emphasized on the structural features of participating drugs for designing synergistic models, they lack other crucial factors directly linked with drug administration - omics expressions. As differential omics expression is a downstream consequence of the administered drug combinations, utilizing such expressions while designing synergistic models promises robust and dynamic modeling. In this work, we propose SynergyLM that fuses multi-omics features with drug embeddings to build an omics-aware …
Chronosort: Revealing Hidden Dynamics In Alphafold3 Structure Predictions, Matthew J. Argyle, William P. Heaps, Corbyn Kubalek, Spencer Gardiner, Bradley C. Bundy, Dennis Della Corte
Chronosort: Revealing Hidden Dynamics In Alphafold3 Structure Predictions, Matthew J. Argyle, William P. Heaps, Corbyn Kubalek, Spencer Gardiner, Bradley C. Bundy, Dennis Della Corte
Faculty Publications
Protein function emerges from dynamic conformational changes, yet structure prediction methods provide only static snapshots. While AlphaFold3 (AF3) predicts protein structures, the potential for extracting dynamic information from its ensemble predictions has remained underexplored. Here, we demonstrate that AF3 structural ensembles contain substantial dynamic information that correlates remarkably well with molecular dynamics simulations (MD). We developed ChronoSort, a novel algorithm that organizes static structure predictions into temporally coherent trajectories by minimizing structural differences between neighboring frames. Through systematic analysis of four diverse protein targets, we show that root-mean-square fluctuations derived from AF3 ensembles can correlate strongly with those from MD …
Ibi-Dt: A Novel Approach Combining Individualized Bayesian Inference And Decision Tree For Identifying Cancer Drivers And Their Interactions, Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu, Jinling Liu
Ibi-Dt: A Novel Approach Combining Individualized Bayesian Inference And Decision Tree For Identifying Cancer Drivers And Their Interactions, Md Asad Rahman, Gregory F. Cooper, Jinying Zhao, Xinghua Lu, Jinling Liu
Engineering Management and Systems Engineering Faculty Research & Creative Works
Cancer is mainly caused by a relatively small portion of somatic genome alterations (SGAs), called cancer drivers. Despite success in identifying a good number of cancer drivers, many more remain to be discovered to explain various cancers. Moreover, limited tools are available to identify potential interactions among cancer drivers for a better understanding of oncogenesis. To tackle these challenges, we have developed a novel approach called individualized Bayesian inference using a decision tree (IBI-DT). IBI-DT recognizes the genetic heterogeneity among cancer patients, where different individuals or patient subgroups of distinct genomic makeup may have different drivers. IBI-DT works by constructing …
Performance Analysis Of Computational Methods For Predicting Protein Function In Rare Diseases, Aichetou Mohamed Sidiya, Hanin Alzaher, Razan Almahdi, Tayeb Brahimi
Performance Analysis Of Computational Methods For Predicting Protein Function In Rare Diseases, Aichetou Mohamed Sidiya, Hanin Alzaher, Razan Almahdi, Tayeb Brahimi
Effat Undergraduate Research Journal
Protein function prediction is crucial for understanding the underlying mechanisms of rare diseases. With the increasing availability of computational methods including machine learning-based approaches, network-based methods, and sequence-based methods, predicting protein functions has become more accessible. However, it is not clear which of these methods performs better or how they compare to each other in terms of accuracy, efficiency, and scalability. In this study, we evaluate several computational methods for predicting protein functions in rare diseases using key performance indicators (KPIs). We analyze the strengths and weaknesses of each method and provide recommendations for researchers and clinicians interested in using …
Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher
Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher
Master's Theses
Neuronal cell types are categorized by transcriptomic identity, yet their morphological heterogeneity defies this classification. In response, researchers have adopted unsupervised graph representation learning as a tool to reveal morphological variation within single-class transcriptomic types. However, the complex geometry of neuronal morphology—especially long axons and dense dendrites—challenges graph neural networks, which struggle with message propagation across extended structures. To mitigate this, current approaches enforce sub-sampling on neuronal graphs and omit axons entirely, sacrificing critical biological features for computational efficiency. To overcome this trade-off, this thesis introduces TopoDINO, a self-supervised, topology-aware representation learning model designed to preserve the full hierarchical organization …
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Computer Science ETDs
Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …
Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh
Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Drug–target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug–target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance …
A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh
A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Conventional drug discovery is expensive, time-consuming, and prone to failure. Artificial intelligence has become a potent substitute over the last decade, providing strong answers to challenging biological issues in this field. Among these difficulties, drug-target binding (DTB) is a key component of drug discovery techniques. In this context, drug-target affinity and drug–target interaction are complementary and essential frameworks that work together to improve our comprehension of DTB dynamics. In this work, we thoroughly analyze the most recent deep learning models, popular benchmark datasets, and assessment metrics for DTB prediction. We look at the paradigm shift in the development of drug …
Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He
Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He
Computer Science Faculty Publications
DeepSSETracer is a method for segmenting protein secondary structure from medium-resolution (5-10Å) cryogenic electron microscopy (cryo-EM) density maps. We conducted experiments and ablation studies to examine the effects of normalization methods, max-pooling, activation functions, and loss calculation region on DeepSSETracer. By combining multiple technical improvements, the performance of the new version, DeepSSETracer 2.0, was significantly enhanced compared to DeepSSETracer 1.1. On a set of 77 test cases, the weighted average per-voxel F1 score increased from 62.1% to 70.3% for helix detection, and from 47.8% to 62.5% for β-sheet detection. While each of the five modifications in the network enhanced the …
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
Genomic Data Science Approaches For Understanding Human Diseases, Snehal Shah
All Dissertations
The intricate interplay of genetic predisposition, environmental influences, and lifestyle acts as the multifactorial landscape of diseases. Understanding this complexity presents a significant challenge. Molecular insights into disease mechanisms, particularly the interactions of DNA, RNA, and proteins with environmental and lifestyle factors, have revolutionized disease diagnosis, prognosis, and treatment. High-throughput technologies, such as next-generation sequencing, generate large amounts of molecular data, holding a wealth of knowledge. These datasets unveil the roles of genes and their interactions with various factors through analysis, shedding light on previously unknown molecular mechanisms underlying disease pathogenesis. Furthermore, they facilitate the discovery of biomarkers crucial for …
A Machine Learning Model Of Perturb-Seq Data For Use In Space Flight Gene Expression Profile Analysis, Liam F. Johnson, James Casaletto, Lauren Sanders, Sylvain Costes
A Machine Learning Model Of Perturb-Seq Data For Use In Space Flight Gene Expression Profile Analysis, Liam F. Johnson, James Casaletto, Lauren Sanders, Sylvain Costes
Graduate Industrial Research Symposium
The genetic perturbations caused by spaceflight on biological systems tend to have a system-wide effect which is often difficult to deconvolute it into individual signals with specific points of origin. Single cell multi-omic data can provide a profile of the perturbational effects, but does not necessarily indicate the initial point of interference within the network. The objective of this project is to take advantage of large scale and genome-wide perturbational datasets by using them to train a tuned machine learning model that is capable of predicting the effects of unseen perturbations in new data. Perturb-Seq datasets are large libraries of …
Creation Of A Digital Storage System For Genome Sequencing Metadata, Jacquelin W. Olexa
Creation Of A Digital Storage System For Genome Sequencing Metadata, Jacquelin W. Olexa
Undergraduate Theses, Professional Papers, and Capstone Artifacts
As the field of computational genomics continues to expand in both potential and application, it is now more imperative than ever to ensure that massive genetic sequencing datasets are properly stored in an accessible manner. This project sought to establish a practical, user-friendly, secure system for a genomics research lab (the Good Lab; thegoodlab.org) at the University of Montana. A MySQL database and connected web application was ruled the best configuration to maximize utility and accessibility for the lab’s researchers. Building the logical framework for the database, creating the server, and sourcing data occurred over several months. The dataset ranged …
Identifying New Cancer Genes Based On The Integration Of Annotated Gene Sets Via Hypergraph Neural Networks, Chao Deng, Hong-Dong Li, Li-Shen Zhang, Yiwei Liu, Yaohang Li, Jianxin Wang
Identifying New Cancer Genes Based On The Integration Of Annotated Gene Sets Via Hypergraph Neural Networks, Chao Deng, Hong-Dong Li, Li-Shen Zhang, Yiwei Liu, Yaohang Li, Jianxin Wang
Computer Science Faculty Publications
Motivation
Identifying cancer genes remains a significant challenge in cancer genomics research. Annotated gene sets encode functional associations among multiple genes, and cancer genes have been shown to cluster in hallmark signaling pathways and biological processes. The knowledge of annotated gene sets is critical for discovering cancer genes but remains to be fully exploited.
Results
Here, we present the DIsease-Specific Hypergraph neural network (DISHyper), a hypergraph-based computational method that integrates the knowledge from multiple types of annotated gene sets to predict cancer genes. First, our benchmark results demonstrate that DISHyper outperforms the existing state-of-the-art methods and highlight the advantages of …
Sccad: Cluster Decomposition-Based Anomaly Detection For Rare Cell Identification In Single-Cell Expression Data, Yunpei Xu, Shaokai Wang, Qilong Feng, Jiazhi Xia, Yaohang Li, Hong-Dong Li, Jianxin Wang
Sccad: Cluster Decomposition-Based Anomaly Detection For Rare Cell Identification In Single-Cell Expression Data, Yunpei Xu, Shaokai Wang, Qilong Feng, Jiazhi Xia, Yaohang Li, Hong-Dong Li, Jianxin Wang
Computer Science Faculty Publications
Single-cell RNA sequencing (scRNA-seq) technologies have become essential tools for characterizing cellular landscapes within complex tissues. Large-scale single-cell transcriptomics holds great potential for identifying rare cell types critical to the pathogenesis of diseases and biological processes. Existing methods for identifying rare cell types often rely on one-time clustering using partial or global gene expression. However, these rare cell types may be overlooked during the clustering phase, posing challenges for their accurate identification. In this paper, we propose a Cluster decomposition-based Anomaly Detection method (scCAD), which iteratively decomposes clusters based on the most differential signals in each cluster to effectively separate …
Machine Learning And Rna Bioinformatics, Jason Rafe Miller
Machine Learning And Rna Bioinformatics, Jason Rafe Miller
Graduate Theses, Dissertations, and Problem Reports (ETD)
The applied science of bioinformatics encompasses computational analysis of molecular biology data. Advances in genomics and DNA sequencing technology have enabled computational analysis of ribonucleic acids (RNAs), which play diverse and critical roles in most cells. To assist the study of human RNA, we trained machine learning models on RNA nucleotide sequences, devoid of domain knowledge. We built models that distinguish long non-coding lncRNA from protein-coding mRNA, and models that predict the cytoplasmic vs. nuclear preferences of lncRNAs. In a review of published lncRNA subcellular localization classifiers, we show that the commonly used validation protocol generates optimistic performance measures, and …
Ai And Ml-Based Risk Assessment Of Chemicals: Predicting Carcinogenic Risk From Chemical-Induced Genomic Instability, Ajay Vikram Singh, Preeti Bhardwaj, Peter Laux, Prachi Pradeep, Madleen Busse, Andreas Luch, Akihiko Hirose, Christopher J. Osgood, Michael W. Stacey
Ai And Ml-Based Risk Assessment Of Chemicals: Predicting Carcinogenic Risk From Chemical-Induced Genomic Instability, Ajay Vikram Singh, Preeti Bhardwaj, Peter Laux, Prachi Pradeep, Madleen Busse, Andreas Luch, Akihiko Hirose, Christopher J. Osgood, Michael W. Stacey
Biological Sciences Faculty Publications
Chemical risk assessment plays a pivotal role in safeguarding public health and environmental safety by evaluating the potential hazards and risks associated with chemical exposures. In recent years, the convergence of artificial intelligence (AI), machine learning (ML), and omics technologies has revolutionized the field of chemical risk assessment, offering new insights into toxicity mechanisms, predictive modeling, and risk management strategies. This perspective review explores the synergistic potential of AI/ML and omics in deciphering clastogen-induced genomic instability for carcinogenic risk prediction. We provide an overview of key findings, challenges, and opportunities in integrating AI/ML and omics technologies for chemical risk assessment, …
Model-Based Deep Autoencoders For Clustering Single-Cell Rna Sequencing Data With Side Information, Xiang Lin
Model-Based Deep Autoencoders For Clustering Single-Cell Rna Sequencing Data With Side Information, Xiang Lin
Dissertations
Clustering analysis has been conducted extensively in single-cell RNA sequencing (scRNA-seq) studies. scRNA-seq can profile tens of thousands of genes' activities within a single cell. Thousands or tens of thousands of cells can be captured simultaneously in a typical scRNA-seq experiment. Biologists would like to cluster these cells for exploring and elucidating cell types or subtypes. Numerous methods have been designed for clustering scRNA-seq data. Yet, single-cell technologies develop so fast in the past few years that those existing methods do not catch up with these rapid changes and fail to fully fulfil their potential. For instance, besides profiling transcription …
Machine Learning And Network Embedding Methods For Gene Co-Expression Networks, Niloofar Aghaieabiane
Machine Learning And Network Embedding Methods For Gene Co-Expression Networks, Niloofar Aghaieabiane
Dissertations
High-throughput technologies such as DNA microarrays and RNA-seq are used to measure the expression levels of large numbers of genes simultaneously. To support the extraction of biological knowledge, individual gene expression levels are transformed into Gene Co-expression Networks (GCNs). GCNs are analyzed to discover gene modules. GCN construction and analysis is a well-studied topic, for nearly two decades. While new types of sequencing and the corresponding data are now available, the software package WGCNA and its most recent variants are still widely used, contributing to biological discovery.
The discovery of biologically significant modules of genes from raw expression data is …
A Computational Analysis Of Hybrid Genome Assembly Strategies, Joseph Walewski
A Computational Analysis Of Hybrid Genome Assembly Strategies, Joseph Walewski
Computer Science Honors Papers
The central dogma of molecular biology states that DNA is transcribed to RNA and then translated into proteins. Since DNA is the starting material for many of biology’s macromolecules, it has been referred to as “nature’s instruction book.” The sum of all DNA in a cell is referred to as the genome, and genome sequencing is how we interpret the DNA.
Due to limitations on currently available technology, it is not possible to retrieve the entire genome in one contiguous set of data. Therefore, genome sequencing is a computer science problem as sequencing “reads” must be stitched together to obtain …
An Approach To Developing Benchmark Datasets For Protein Secondary Structure Segmentation From Cryo-Em Density Maps, Thu Nguyen, Yongcheng Mu, Jiangwen Sun, Jing He
An Approach To Developing Benchmark Datasets For Protein Secondary Structure Segmentation From Cryo-Em Density Maps, Thu Nguyen, Yongcheng Mu, Jiangwen Sun, Jing He
Computer Science Faculty Publications
More and more deep learning approaches have been proposed to segment secondary structures from cryo-electron density maps at medium resolution range (5--10Å). Although the deep learning approaches show great potential, only a few small experimental data sets have been used to test the approaches. There is limited understanding about potential factors, in data, that affect the performance of segmentation. We propose an approach to generate data sets with desired specifications in three potential factors - the protein sequence identity, structural contents, and data quality. The approach was implemented and has generated a test set and various training sets to study …
Cellbrf: A Feature Selection Method For Single-Cell Clustering Using Cell Balance And Random Forest, Yunpei Xu, Hong-Dong Li, Cui-Xiang Lin, Ruiqing Zheng, Yaohang Li, Jinhui Xu, Jianxin Wang
Cellbrf: A Feature Selection Method For Single-Cell Clustering Using Cell Balance And Random Forest, Yunpei Xu, Hong-Dong Li, Cui-Xiang Lin, Ruiqing Zheng, Yaohang Li, Jinhui Xu, Jianxin Wang
Computer Science Faculty Publications
Motivation
Single-cell RNA sequencing (scRNA-seq) offers a powerful tool to dissect the complexity of biological tissues through cell sub-population identification in combination with clustering approaches. Feature selection is a critical step for improving the accuracy and interpretability of single-cell clustering. Existing feature selection methods underutilize the discriminatory potential of genes across distinct cell types. We hypothesize that incorporating such information could further boost the performance of single cell clustering. Results
We develop CellBRF, a feature selection method that considers genes’ relevance to cell types for single-cell clustering. The key idea is to identify genes that are most important for discriminating …
Dfhic: A Dilated Full Convolution Model To Enhance The Resolution Of Hi-C Data, Bin Wang, Kun Liu, Yaohang Li, Jianxin Wang
Dfhic: A Dilated Full Convolution Model To Enhance The Resolution Of Hi-C Data, Bin Wang, Kun Liu, Yaohang Li, Jianxin Wang
Computer Science Faculty Publications
Motivation: Hi-C technology has been the most widely used chromosome conformation capture(3C) experiment that measures the frequency of all paired interactions in the entire genome, which is a powerful tool for studying the 3D structure of the genome. The fineness of the constructed genome structure depends on the resolution of Hi-C data. However, due to the fact that high-resolution Hi-C data require deep sequencing and thus high experimental cost, most available Hi-C data are in low-resolution. Hence, it is essential to enhance the quality of Hi-C data by developing the effective computational methods.
Results: In this work, we propose …
Intergenic Transcription In In Vivo Developed Bovine Oocytes And Pre-Implantation Embryos, Saurav Ranjitkar, Mohammad Shiri, Jiangwen Sun, Xiuchun Tian
Intergenic Transcription In In Vivo Developed Bovine Oocytes And Pre-Implantation Embryos, Saurav Ranjitkar, Mohammad Shiri, Jiangwen Sun, Xiuchun Tian
Computer Science Faculty Publications
Background
Intergenic transcription, either failure to terminate at the transcription end site (TES), or transcription initiation at other intergenic regions, is present in cultured cells and enhanced in the presence of stressors such as viral infection. Transcription termination failure has not been characterized in natural biological samples such as pre-implantation embryos which express more than 10,000 genes and undergo drastic changes in DNA methylation.
Results
Using Automatic Readthrough Transcription Detection (ARTDeco) and data of in vivo developed bovine oocytes and embryos, we found abundant intergenic transcripts that we termed as read-outs (transcribed from 5 to 15 kb after TES) and …
Sequence-Based Bioinformatics Approaches To Predict Virus–Host Relationships In Archaea And Eukaryotes, Yingshan Li
Sequence-Based Bioinformatics Approaches To Predict Virus–Host Relationships In Archaea And Eukaryotes, Yingshan Li
School of Computing: Dissertations, Theses, and Student Research
Viral metagenomics is independent of lab culturing and capable of investigating viromes of virtually any given environmental niches. While numerous sequences of viral genomes have been assembled from metagenomic studies over the past years, the natural hosts for the majority of these viral contigs have not been determined. Different computational approaches have been developed to predict hosts of bacteria phages. Nevertheless, little progress has been made in the virus-host prediction, especially for viruses that infect eukaryotes and archaea. In this study, by analyzing all documented viruses with known eukaryotic and archaeal hosts, we assessed the predictive power of four computational …
Development Of Graphical Models And Statistical Physics Motivated Approaches To Genomic Investigations, Yashwanth Lagisetty
Development Of Graphical Models And Statistical Physics Motivated Approaches To Genomic Investigations, Yashwanth Lagisetty
Dissertations and Theses (Open Access)
Identifying genes involved in disease pathology has been a goal of genomic research since the early days of the field. However, as technology improves and the body of research grows, we are faced with more questions than answers. Among these is the pressing matter of our incomplete understanding of the genetic underpinnings of complex diseases. Many hypotheses offer explanations as to why direct and independent analyses of variants, as done in genome-wide association studies (GWAS), may not fully elucidate disease genetics. These range from pointing out flaws in statistical testing to invoking the complex dynamics of epigenetic processes. In the …
Comparative Analyses Of De Novo Transcriptome Assembly Pipelines For Diploid Wheat, Natasha Pavlovikj
Comparative Analyses Of De Novo Transcriptome Assembly Pipelines For Diploid Wheat, Natasha Pavlovikj
School of Computing: Dissertations, Theses, and Student Research
Gene expression and transcriptome analysis are currently one of the main focuses of research for a great number of scientists. However, the assembly of raw sequence data to obtain a draft transcriptome of an organism is a complex multi-stage process usually composed of pre-processing, assembling, and post-processing. Each of these stages includes multiple steps such as data cleaning, error correction and assembly validation. Different combinations of steps, as well as different computational methods for the same step, generate transcriptome assemblies with different accuracy. Thus, using a combination that generates more accurate assemblies is crucial for any novel biological discoveries. Implementing …