Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Bioinformatics (398)
- Genomics (222)
- Physical Sciences and Mathematics (208)
- Genetics (172)
- Medicine and Health Sciences (153)
-
- Biochemistry, Biophysics, and Structural Biology (127)
- Molecular Genetics (105)
- Biology (104)
- Statistics and Probability (93)
- Ecology and Evolutionary Biology (87)
- Computer Sciences (86)
- Cell and Developmental Biology (81)
- Molecular Biology (68)
- Evolution (57)
- Microbiology (55)
- Biotechnology (51)
- Systems Biology (51)
- Biochemistry (43)
- Cell Biology (41)
- Microarrays (39)
- Structural Biology (39)
- Biostatistics (38)
- Medical Sciences (38)
- Engineering (37)
- Animal Sciences (33)
- Diseases (33)
- Statistical Methodology (32)
- Institution
-
- COBRA (113)
- The Texas Medical Center Library (49)
- Dartmouth College (47)
- Old Dominion University (40)
- University of Kentucky (32)
-
- City University of New York (CUNY) (31)
- Himmelfarb Health Sciences Library, The George Washington University (27)
- University of Nebraska - Lincoln (26)
- Thomas Jefferson University (20)
- Virginia Commonwealth University (19)
- Louisiana State University (13)
- The University of Southern Mississippi (13)
- California Polytechnic State University, San Luis Obispo (12)
- Clemson University (12)
- Illinois State University (9)
- Loyola University Chicago (9)
- University of Arkansas, Fayetteville (9)
- University of Connecticut (9)
- Augustana College (8)
- University of Louisville (8)
- Mississippi State University (7)
- West Virginia University (7)
- Western University (7)
- Chapman University (6)
- Munster Technological University (6)
- Swarthmore College (6)
- University of Montana (6)
- University of Nevada, Las Vegas (6)
- University of New Mexico (6)
- Harrisburg University of Science and Technology (5)
- Keyword
-
- Bioinformatics (70)
- Genetics (52)
- Computational biology (40)
- Genomics (37)
- Humans (35)
-
- Gene expression (32)
- Algorithms (24)
- Genome (19)
- Animals (18)
- Evolution (18)
- Machine learning (18)
- Models (16)
- Deep learning (15)
- Transcriptomics (15)
- Genetic (13)
- Computational Biology (12)
- Protein (12)
- Phylogeny (11)
- Transcriptome (11)
- Annotation (10)
- Computer simulation (10)
- Epigenetics (10)
- Machine Learning (10)
- Metabolism (10)
- Phylogenetics (10)
- Population genetics (10)
- Systems biology (10)
- Transcription factors (10)
- Cancer (9)
- Sequence analysis (9)
- Publication Year
- Publication
-
- Dissertations and Theses (Open Access) (48)
- Dartmouth Scholarship (36)
- Computer Science Faculty Publications (27)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (27)
- Computational Biology Institute (24)
-
- Harvard University Biostatistics Working Paper Series (24)
- Theses and Dissertations (22)
- Dissertations, Theses, and Capstone Projects (18)
- COBRA Preprint Series (17)
- Electronic Theses and Dissertations (13)
- UW Biostatistics Working Paper Series (12)
- LSU Doctoral Dissertations (11)
- Publications and Research (11)
- Theses and Dissertations--Biology (11)
- Bioconductor Project Working Papers (10)
- Dartmouth College Ph.D Dissertations (10)
- Dissertations (10)
- UPenn Biostatistics Working Papers (10)
- All Dissertations (8)
- Annual Symposium on Biomathematics and Ecology Education and Research (8)
- Computational Medicine Center Faculty Papers (8)
- Honors Theses (8)
- Meiothermus ruber Genome Analysis Project (8)
- STAR Program Research Presentations (8)
- Bioinformatics Faculty Publications (7)
- Biology Faculty Publications (7)
- Department of Pathology, Anatomy, and Cell Biology Faculty Papers (7)
- Master's Theses (7)
- U.C. Berkeley Division of Biostatistics Working Paper Series (7)
- Biochemistry Publications (6)
- Publication Type
- File Type
Articles 61 - 90 of 754
Full-Text Articles in Computational Biology
Chromosome Evolution Model Reveals Hidden Variation In Karyotype-Driven Speciation In Ferns, Thomas Buchloh
Chromosome Evolution Model Reveals Hidden Variation In Karyotype-Driven Speciation In Ferns, Thomas Buchloh
All Theses
Because karyotype change commonly generates reproductive isolation between diverging species, rapid karyotype evolution, like that seen in plants, may increase the total rate of diversification. However, few studies have investigated this predicted relationship. Tests of this prediction have identified a positive correlation between the rate of karyotype change (specifically whole genome duplications) and diversification, but tests have been restricted to relatively small and young clades. Fortunately, novel macroevolutionary models have recently become available to investigate patterns of karyotype evolution in large phylogenies, providing new opportunity to investigate variation in karyotype driven diversification. Ferns are one of the most karyotype rich …
Integration Of Multi-Omics Datasets Evaluating Structural And Functional Features Of The Gut Microbiome In People With Hiv (Pwh)., Aakarsha Vijayakumar Rao
Integration Of Multi-Omics Datasets Evaluating Structural And Functional Features Of The Gut Microbiome In People With Hiv (Pwh)., Aakarsha Vijayakumar Rao
Electronic Theses and Dissertations
Gut dysbiosis characterized by reduced abundance of beneficial butyrate-producing bacteria has been independently linked to HIV-1 infection and heavy alcohol drinking. Further, gut dysbiosis results in loss of gut barrier integrity, microbial translocation and host-specific systemic inflammation. Therefore, to evaluate the functional consequences of structural changes in the gut microbiome, integrated data analysis is imperative. In this dissertation, we perform integrated analyses using data from multi-omics platforms to examine the structural and functional features of the gut microbiome of PWH. Gut microbiome composition was evaluated by sequencing V4 region of 16S rDNA, concentrations of metabolites and host-specific immune markers were …
Elucidating The Multi-Omics Of Early-Onset Colorectal Cancer, Jumanah Alshenaifi
Elucidating The Multi-Omics Of Early-Onset Colorectal Cancer, Jumanah Alshenaifi
Dissertations and Theses (Open Access)
The incidence and mortality rates of sporadic early-onset colorectal cancer have increased in recent decades, but there is no clear etiological basis for this trend. EOCRC is commonly defined as colon and rectal cancers diagnosed before the age of 50 years. The rising incidence of EOCRC has made it the second most common cancer and the third leading cause of cancer death in this age group. The rising incidence of EOCRC is also documented internationally in more than 20 countries across different continents. Clinically, EOCRC has a distinct, more aggressive clinical profile than LOCRC. While approximately 15% of EOCRC cases …
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Computer Science ETDs
Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …
Hierarchical Lineage Tracing To Unravel Mechanisms Of Cancer Treatment Resistance, Rachel Danielle Saxe
Hierarchical Lineage Tracing To Unravel Mechanisms Of Cancer Treatment Resistance, Rachel Danielle Saxe
Dartmouth College Ph.D Dissertations
Cancer cells adapt to treatment, leading to the emergence of clones that are more aggressive and resistant to anti-cancer therapies. We have a limited understanding of the development of treatment resistance as we lack technologies to map the evolution of cancer under the selective pressure of treatment. To address this, we developed a hierarchical, dynamic lineage tracing method called FLARE (Following Lineage Adaptation and Resistance Evolution). We use this technique to track the progression of acute myeloid leukemia (AML) cell lines through exposure to Cytarabine (AraC), a front-line treatment in AML, in vitro and in vivo. We map distinct cellular …
Analysis Of Chromatin Accessibility Changes In Endothelial Cells Exposed To Plastic Contaminants, Mikhail Y. Salnikov, Carly Boye, David B. Witonsky, Gabrielle Garlicki, Adnan Alazizi, Francesca Luca, Roger Pique-Regi
Analysis Of Chromatin Accessibility Changes In Endothelial Cells Exposed To Plastic Contaminants, Mikhail Y. Salnikov, Carly Boye, David B. Witonsky, Gabrielle Garlicki, Adnan Alazizi, Francesca Luca, Roger Pique-Regi
Medical Student Research Symposium
Degradation products from everyday plastic products are known to bioaccumulate and have also been shown to contaminate drinking water and food sources. BPA and phthalates are endocrine disrupting chemicals and plastic components that have previously been associated with endothelial cell dysfunction, atherosclerotic and other adverse cardiovascular events. However, there is a limited understanding of the mechanisms underlying these associations, such as genome-wide chromatin accessibility changes in endothelial cells exposed to these compounds. The purpose of this study is to explore genome-wide changes in chromatin accessibility associated with plastic exposure, as well as the discovery of transcription factor binding motifs dysregulated …
Cloning And Expression Of Scamp3 In Escherichia Coli Bl21(De3) With In Silico Sequence-Based Cancer Epitopes Prediction, Selly Setiati Rajagukguk, Sabar Pambudi, Astari Dwiranti, Doddy Irawan Setyo Utomo, Anom Bowolaksono
Cloning And Expression Of Scamp3 In Escherichia Coli Bl21(De3) With In Silico Sequence-Based Cancer Epitopes Prediction, Selly Setiati Rajagukguk, Sabar Pambudi, Astari Dwiranti, Doddy Irawan Setyo Utomo, Anom Bowolaksono
Makara Journal of Science
Secretory carrier membrane protein 3 (SCAMP3) is a crucial membrane protein involved in intracellular vesicle traffick-ing and exocytosis. The SCAMP3 expression has been observed in diverse cancer types, such as melanoma, glioma, hepatocellular and breast cancer. Increased SCAMP3 expression has been reported in certain cancer cells relative to that in normal cells, suggesting the potential role of SCAMP3 in cancer development or progression. In this study, we successfully cloned and expressed SCAMP3 in Escherichia coli strain BL21(DE3). SCAMP3 was amplified and insert-ed directionally into the prokaryotic expression vector pET21d(+). The transformation of recombinant plasmid into E. coli BL21(DE3) cells were …
Screening For Penicillin G Acylase (Pga)-Producing Bacteria And Gene Cloning Using Degenerate Oligonucleotide Primed-Pcr, Masdalifah Masdalifah, Sri Rezeki Wulandari, Gabriela Christy Sabbathini, Maria Ulfah, Dini Achnafani, Ahmad Wibisana, Feronika Heppy Sriherfyna, Is Helianti, Niknik Nurhayati
Screening For Penicillin G Acylase (Pga)-Producing Bacteria And Gene Cloning Using Degenerate Oligonucleotide Primed-Pcr, Masdalifah Masdalifah, Sri Rezeki Wulandari, Gabriela Christy Sabbathini, Maria Ulfah, Dini Achnafani, Ahmad Wibisana, Feronika Heppy Sriherfyna, Is Helianti, Niknik Nurhayati
Makara Journal of Science
The growing concern over antibiotic resistance has driven global efforts to explore innovative solutions, including the use of Penicillin G acylase (PGA) to produce semisynthetic β-lactam antibiotics. This study screened four potential in-tracellular PGA-producing bacteria: Alcaligenes faecalis InaCC B444 (AfPGA), Kluyvera cryocrescens InaCC B850 (KcPGA), Providencia rettgeri InaCC B25 (Pr25PGA), and P. rettgeri InaCC B466 (Pr466PGA). Penicillin G Acylase encoding genes (pgas) were isolated from them using a Degenerate Oligonucleotide Primed-PCR (DOP-PCR) approach and sequenced. Microbiological assays confirmed all tested crude extracts to exhibit inhibitory effects. Penicillin G was used for evaluating hydrolytic activity and 6-Amino Penicillanic Acid (6-APA) coupled …
Identification Of Novel Argonaute Proteins Using A Metagenomic Mining Approach, Lobna Abdallah Ghonaim
Identification Of Novel Argonaute Proteins Using A Metagenomic Mining Approach, Lobna Abdallah Ghonaim
Theses and Dissertations
Gene editing is one of the most promising tools in science. It enables precise modifications of an organism's genetic material. Metagenomics is considered a powerful tool that unlocks the broad genetic potential found in uncultured microbial communities. Exploring the genetic diversity of uncultured microbial communities helps identify novel functional proteins with unique properties and make the best use of these diverse microbial ecosystems.
We developed and employed a metagenomic-based approach to mine more than 1000 metagenomes for prokaryotic argonaute proteins (pAgos), a potential gene editing machinery encoded in bacterial and archaeal genomes. Our workflow involved strict quality control, sequence assembly, …
The Role Of Secondary And Tertiary Structure In The Cap-Independent Translation Of Fgf-9 And Hif-1-Alpha, Amanda Michelle Whittaker
The Role Of Secondary And Tertiary Structure In The Cap-Independent Translation Of Fgf-9 And Hif-1-Alpha, Amanda Michelle Whittaker
Dissertations, Theses, and Capstone Projects
Under normoxic conditions, eukaryotes initiate translation of RNA through eIF4E recognition of the 5’ cap. However, under cellular stress, eukaryotic translation must be initiated through a 4E-independent, or “cap-independent” mechanism, involving eukaryotic initiation factor 4G (eIF4G) binding directly to the 5’ untranslated regions (5’ UTR) of the RNA. eIF4G binding then recruits the ribosome to the transcript. While this mechanism is useful for translation of apoptotic transcripts and transcripts involved in cell survival, cap-independent translation is also utilized by oncogenic RNA for tumorigenesis. Previous work by our lab and others has categorized this recruitment and initiation mechanism as either internal-ribosome-entry-site …
Taming Biological Complexity Through The Use Of Symmetries, Luis A. Álvarez-García
Taming Biological Complexity Through The Use Of Symmetries, Luis A. Álvarez-García
Dissertations, Theses, and Capstone Projects
The study of biological systems is, inherently, the study of very complex systems. This is essentially due to the fact that they are made up of numerous, often very complicated, interactions between an extensive number of components. Often necessitating an abundance of quantitative parameters and details for a precise description. The human brain for ex- ample, consisting of ∼ 80 billion neurons with ∼ 800 to 100 trillion connections between them, each of them depending on a large set of parameters. Even simpler examples such as bacterial organisms, such as E. coli and B. subtilis, which we focus on …
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
Doves And Symbiotic Bacteria: Whole-Genome Assembly And Phylogenomic Inference, Mariam Topchyan
Doves And Symbiotic Bacteria: Whole-Genome Assembly And Phylogenomic Inference, Mariam Topchyan
Theses and Dissertations
Part 1
I examined a potential source of variation in phylogenetic inference, investigating whether handling single-nucleotide polymorphisms (SNPs) through different methods during reference-guided genome assembly of pigeon and dove species impacts phylogenetic reconstruction and inference. Specifically, I created a custom consensus base calling tool that handles allelic variation through two different approaches, either ignoring heterozygous sites by denoting them as an ambiguous character “N,” or using a pseudo-random choice between supported base variants. Through reference-guided assembly of 108 dove species, I generated two datasets of alignments using both consensus calling methods, created two trees from both datasets via supermatrix and …
18s Metabarcode Analyses Of Eukaryotic Species In The Respiratory Microbiomes Of Wild Canids From New Hampshire, Collin Sinclair Blake
18s Metabarcode Analyses Of Eukaryotic Species In The Respiratory Microbiomes Of Wild Canids From New Hampshire, Collin Sinclair Blake
Honors Theses and Capstones
This study is of an exploratory nature and focuses on characterizing the eukaryotic microbiota present in the respiratory tissues of six wild canids and one domestic canine. The contents of this document largely pertain to dry lab analyses of 18S barcodes in bioinformatics programs – primarily QIIME2 running in the GitBash command line, service for which was hosted by the UNH Ron Bioinformatics training server. All procedures listed within the section below were performed by second parties at the UNH Hubbard Center for Genomics Studies (HCGS), the New Hampshire Veterinary Diagnostics Lab (NHVDL), and the UNH Microbial Ecology and Emerging …
Genomic Epidemiology Of Staphylococcus Aureus Sequence-Type 72, Sachitaa Senthilkumar
Genomic Epidemiology Of Staphylococcus Aureus Sequence-Type 72, Sachitaa Senthilkumar
Honors Undergraduate Theses
Staphylococcus aureus (S. aureus, SA) is a gram-positive bacterial colonizer and pathogen commonly carried on the skin and in the anterior nares of humans. SA colonization may lead to a variety of non-invasive (e.g., skin and soft tissue infections, SSTIs) and invasive (e.g., bacteremia and osteomyelitis) infections. The SA population is comprised of multiple lineages that can be delineated through multi-locus sequence typing (MLST), which compares nucleotide sequences within seven housekeeping genes. Strains belonging to an MLST share a common evolutionary history and phenotypic characteristics like antibiotic resistance or virulence. As a result, the epidemiology of lineages can be …
Enhancing Clinical Trial Matching In Molecular Diagnostics: Using Natural Language Processing And Clustering Approaches In Hematological Malignancies, Gillian Fanning
Enhancing Clinical Trial Matching In Molecular Diagnostics: Using Natural Language Processing And Clustering Approaches In Hematological Malignancies, Gillian Fanning
Theses and Dissertations
Clinical trial matching is a critical component of personalized medicine, particularly in the management of hematologic malignancies. At Virginia Commonwealth University (VCU) Health, the Molecular Diagnostics (MDX) Lab produces somatic variant reports and recommends clinical trials based on the presence of clinically significant mutations. However, the current manual trial recommendation process is time-intensive and lacks scalability.
This study introduces a computational framework to streamline and standardize clinical trial matching using natural language processing (NLP) and unsupervised clustering. Trial brief descriptions were analyzed to extract frequent terms, and trials were grouped based on term similarity using joint dimensionality reduction and clustering. …
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
The rapid growth of diverse -omics datasets has made multiomics data integration crucial in cancer research. This study adapts the expectation–maximization routine for the joint latent variable modeling of multiomics patient profiles. By combining this approach with traditional biological feature selection methods, this study optimizes latent distribution, enabling efficient patient clustering from well-studied cancer types with reduced computational expense. The proposed optimization subroutines enhance survival analysis and improve runtime performance. This article presents a framework for distinguishing cancer subtypes and identifying potential biomarkers for breast cancer. Key insights into individual subtype expression and function were obtained through differentially expressed gene …
Foundation Models In Bioinformatics, Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, Jianxin Wang
Foundation Models In Bioinformatics, Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, Jianxin Wang
Computer Science Faculty Publications
With the adoption of foundation models (FMs), artificial intelligence (AI) has become increasingly significant in bioinformatics and has successfully addressed many historical challenges, such as pre-training frameworks, model evaluation and interpretability. FMs demonstrate notable proficiency in managing large-scale, unlabeled datasets, because experimental procedures are costly and labor intensive. In various downstream tasks, FMs have consistently achieved noteworthy results, demonstrating high levels of accuracy in representing biological entities. A new era in computational biology has been ushered in by the application of FMs, focusing on both general and specific biological issues. In this review, we introduce recent advancements in bioinformatics FMs …
Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh
Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Drug–target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug–target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance …
Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang
Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang
Computer Science Faculty Publications
Motivation: Protein-protein interactions (PPIs) are fundamental aspects in understanding biological processes. Accurately predicting the effects of mutations on PPIs remains a critical requirement for drug design and disease mechanistic studies. Recently, deep learning models using protein 3D structures have become predominant for predicting mutation effects. However, significant challenges remain in practical applications, in part due to the considerable disparity in generalization capabilities between easy and hard mutations. Specifically, a hard mutation is defined as one with its maximum TM-score < 0.6 when compared to the training set. Additionally, compared to physics-based approaches, deep learning models may overestimate performance due to potential data leakage.
Results: We propose new training/test splits that mitigate data leakage according to the CATH homologous superfamily. Under the constraints of physical …
A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh
A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Conventional drug discovery is expensive, time-consuming, and prone to failure. Artificial intelligence has become a potent substitute over the last decade, providing strong answers to challenging biological issues in this field. Among these difficulties, drug-target binding (DTB) is a key component of drug discovery techniques. In this context, drug-target affinity and drug–target interaction are complementary and essential frameworks that work together to improve our comprehension of DTB dynamics. In this work, we thoroughly analyze the most recent deep learning models, popular benchmark datasets, and assessment metrics for DTB prediction. We look at the paradigm shift in the development of drug …
Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He
Deepssetracer 2.0: Improved Deep Learning Model Performance For Protein Secondary Structure Segmentation From Cryo-Em Maps, Bryan Hawickhorst, Thu Nguyen, Willy Wriggers, Jiangwen Sun, Jing He
Computer Science Faculty Publications
DeepSSETracer is a method for segmenting protein secondary structure from medium-resolution (5-10Å) cryogenic electron microscopy (cryo-EM) density maps. We conducted experiments and ablation studies to examine the effects of normalization methods, max-pooling, activation functions, and loss calculation region on DeepSSETracer. By combining multiple technical improvements, the performance of the new version, DeepSSETracer 2.0, was significantly enhanced compared to DeepSSETracer 1.1. On a set of 77 test cases, the weighted average per-voxel F1 score increased from 62.1% to 70.3% for helix detection, and from 47.8% to 62.5% for β-sheet detection. While each of the five modifications in the network enhanced the …
Computational Tools For Protein Classification In Metagenomes, Fawad Ullah
Computational Tools For Protein Classification In Metagenomes, Fawad Ullah
Dissertations, Master's Theses and Master's Reports
Advances in genomic sequencing have dramatically increased the amount and the speed at which genomic data is being generated. These technological advances have enabled the ability to profile the genetic information of organisms and communities at unprecedented scales. Many methods have been developed to identify and classify genes within these datasets. However, many generic pipelines for gene annotation struggle to accurately predict specific protein classes that may not be represented in their databases. Two major challenges exist for classification of specific protein classes in metagenomic databases. The first is the fact that many metagenomic assemblies are highly fragmented with many …
Mobula, Bioinspiration, Filter Feeding, Form And Function, J. B. Teeple, S. R. Kahane-Rapport, K. E. Cohen, L. Hamann, J. A. Strother, E. W.M. Paig-Tran
Mobula, Bioinspiration, Filter Feeding, Form And Function, J. B. Teeple, S. R. Kahane-Rapport, K. E. Cohen, L. Hamann, J. A. Strother, E. W.M. Paig-Tran
Biological Sciences Faculty Publications
Mobulas (manta and devil rays) are large-scale ram filter feeders that separate planktonic food particles from large volumes of water with minimal clogging. This contrasts with most human-made filters that can suffer from problematic clogging requiring additional mechanisms for clearing blocked surfaces and maintaining performance. Prior studies have shown that mobulas employ a unique mechanism referred to as ricochet separation to filter feed, whereby captive vortices in filter pores cause particles to bounce off the filter surfaces and away from the filter pores. This mechanism enables the filtration of particles smaller than the pore size and reduced clogging. However, few …
Copula-Based Bayesian Model For Detecting Differential Gene Expression, Prasansha Liyanaarachchi, N. Rao Chaganty
Copula-Based Bayesian Model For Detecting Differential Gene Expression, Prasansha Liyanaarachchi, N. Rao Chaganty
Mathematics & Statistics Faculty Publications
Deoxyribonucleic acid, more commonly known as DNA, is a fundamental genetic material in all living organisms, containing thousands of genes, but only a subset exhibit differential expression and play a crucial role in diseases. Microarray technology has revolutionized the study of gene expression, with two primary types available for expression analysis: spotted cDNA arrays and oligonucleotide arrays. This research focuses on the statistical analysis of data from spotted cDNA microarrays. Numerous models have been developed to identify differentially expressed genes based on the red and green fluorescence intensities measured using these arrays. We propose a novel approach using a Gaussian …
Characterizing Somatic Variants In Nanopore Data With Machine Learning, Shwethal Sayeeram Trikannad
Characterizing Somatic Variants In Nanopore Data With Machine Learning, Shwethal Sayeeram Trikannad
Master's Projects
Oxford Nanopore Technology (ONT) is a popular long-read sequencer in genomics. However, its high base-calling error rate produces several sequencing artifacts. Detection of somatic variants in ONT sequenced tumor-normal samples remains challenging due to low frequencies. In this study, machine learning was applied to a dataset created by benchmarking ClairS output against HCC1395 and colo829 truth sets to classify variants and artifacts. Relevant features were engineered from sequence context and variant site characteristics to model artifact profiles. HistGradientBoostingClassifier achieved 0.876950 accuracy, outperforming all other models. Variant quality was the top predictor with an aggregate accuracy of over 85%. This work …
Phylogenetic Analysis Of Metabolic Enzymes In Hypoxic/Anoxic Conditions In Cetaceans And Cancer, Samerna A. Masih
Phylogenetic Analysis Of Metabolic Enzymes In Hypoxic/Anoxic Conditions In Cetaceans And Cancer, Samerna A. Masih
Electronic Theses & Dissertations (2024 - present)
Cancer cells often exhibit a Warburg-like metabolism, which includes increased fatty acid synthesis and storage that help them survive in hypoxic conditions. This shift is marked by the heightened synthesis and accumulation of fatty acids, acting as a survival mechanism in low-oxygen environments (Baumann et al., 2016). Interestingly, deep-diving whales may utilize a similar metabolic pathway to produce wax esters during their prolonged dives.
Our study focused on exploring the genetic and metabolic similarities between breast cancer cells and deep-diving whales, using computational analyses to investigate lipid metabolism genes. We looked for evolutionarily conserved variations that might reveal adaptive mechanisms …
Hidden Markov Model For Identifying Local Variants In Human Genomes Using Simulated Data, Scott Mccallum
Hidden Markov Model For Identifying Local Variants In Human Genomes Using Simulated Data, Scott Mccallum
Electronic Theses and Dissertations
Identifying adaptive mutations in genetic data is challenging due to the low frequency of occurrence of such events, and because signatures of selection are intertwined with the footprints of various other evolutionary forces that shape our genomes. Even when a larger region appears to be under selection, genomic sites that are linked to adaptive mutations have similar statistical signals, and thus can obfuscate the identification of the actual adaptive mutation. The new method described here uses a Hidden Markov Model that allows for classification of neutral, linked, and sweep (adaptive mutation) genomic sites. This model is general and can be …
From Sampling To Simulating: Single-Cell Multiomics In Systems Pathophysiological Modeling, Alexandra Manchel, Michelle M. Gee, Rajanikanth Vadigepalli
From Sampling To Simulating: Single-Cell Multiomics In Systems Pathophysiological Modeling, Alexandra Manchel, Michelle M. Gee, Rajanikanth Vadigepalli
Department of Pathology, Anatomy, and Cell Biology Faculty Papers
As single-cell omics data sampling and acquisition methods have accumulated at an unprecedented rate, various data analysis pipelines have been developed for the inference of cell types, cell states and their distribution, state transitions, state trajectories, and state interactions. This presents a new opportunity in which single-cell omics data can be utilized to generate high-resolution, high-fidelity computational models. In this review, we discuss how single-cell omics data can be used to build computational models to simulate biological systems at various scales. We propose that single-cell data can be integrated with physiological information to generate organ-specific models, which can then be …
Timigp: A Computational Framework To Determine The Tumor Immune Microenvironment Associated With Prognosis And Immunotherapy Response, Chenyang Li
Dissertations and Theses (Open Access)
Accumulating evidence has suggested that the tumor immune microenvironment (TIME) drastically impacts cancer patients’ clinical outcomes, including prognosis and immunotherapy response. However, understanding TIME remains challenging due to its complexity and heterogeneity. In this dissertation, we introduce TimiGP (Tumor Immune Microenvironment Illustration based on Gene Pairs), a computational framework designed to address this challenge. Leveraging single-cell RNA-seq (scRNA-seq) and bulk gene expression data alongside clinical information, TimiGP constructs a cell-cell interaction network that elucidates the relationship between immune cell function and relevant clinical outcomes, such as prognosis and treatment response. With immunological insights, these cell-cell interactions also facilitate the development …