Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (113)
- The Texas Medical Center Library (50)
- Dartmouth College (47)
- Old Dominion University (40)
- University of Kentucky (32)
-
- City University of New York (CUNY) (31)
- Himmelfarb Health Sciences Library, The George Washington University (27)
- University of Nebraska - Lincoln (26)
- Thomas Jefferson University (23)
- Virginia Commonwealth University (19)
- Louisiana State University (13)
- The University of Southern Mississippi (13)
- California Polytechnic State University, San Luis Obispo (12)
- Clemson University (12)
- Illinois State University (9)
- Loyola University Chicago (9)
- University of Arkansas, Fayetteville (9)
- University of Connecticut (9)
- Augustana College (8)
- University of Louisville (8)
- West Virginia University (8)
- Mississippi State University (7)
- Western University (7)
- Chapman University (6)
- Munster Technological University (6)
- Swarthmore College (6)
- University of Montana (6)
- University of Nevada, Las Vegas (6)
- University of New Mexico (6)
- Harrisburg University of Science and Technology (5)
- Keyword
-
- Bioinformatics (70)
- Genetics (52)
- Computational biology (40)
- Genomics (37)
- Humans (36)
-
- Gene expression (32)
- Algorithms (24)
- Animals (19)
- Genome (19)
- Evolution (18)
- Machine learning (18)
- Models (16)
- Deep learning (15)
- Transcriptomics (15)
- Genetic (13)
- Computational Biology (12)
- Protein (12)
- Transcriptome (12)
- Phylogeny (11)
- Annotation (10)
- Computer simulation (10)
- Epigenetics (10)
- Machine Learning (10)
- Metabolism (10)
- Phylogenetics (10)
- Population genetics (10)
- RNA (10)
- Systems biology (10)
- Transcription factors (10)
- Cancer (9)
- Publication Year
- Publication
-
- Dissertations and Theses (Open Access) (49)
- Dartmouth Scholarship (36)
- Computer Science Faculty Publications (27)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (27)
- Computational Biology Institute (24)
-
- Harvard University Biostatistics Working Paper Series (24)
- Theses and Dissertations (22)
- Dissertations, Theses, and Capstone Projects (18)
- COBRA Preprint Series (17)
- Electronic Theses and Dissertations (13)
- UW Biostatistics Working Paper Series (12)
- Computational Medicine Center Faculty Papers (11)
- LSU Doctoral Dissertations (11)
- Publications and Research (11)
- Theses and Dissertations--Biology (11)
- Bioconductor Project Working Papers (10)
- Dartmouth College Ph.D Dissertations (10)
- Dissertations (10)
- UPenn Biostatistics Working Papers (10)
- All Dissertations (8)
- Annual Symposium on Biomathematics and Ecology Education and Research (8)
- Honors Theses (8)
- Meiothermus ruber Genome Analysis Project (8)
- STAR Program Research Presentations (8)
- Bioinformatics Faculty Publications (7)
- Biology Faculty Publications (7)
- Department of Pathology, Anatomy, and Cell Biology Faculty Papers (7)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (7)
- Master's Theses (7)
- U.C. Berkeley Division of Biostatistics Working Paper Series (7)
- Publication Type
- File Type
Articles 61 - 90 of 761
Full-Text Articles in Genetics and Genomics
Analytical Approaches For Identification Of Essential Genomes Of Plasmodium Knowlesi And Babesia Divergens, Sida Ye
Graduate Doctoral Dissertations
Apicomplexa constitute a large phylum of single-celled, obligate intracellular protozoan parasites. Notably, Plasmodium spp. and Babesia spp. are apicomplexan parasites that infect red blood cells. Plasmodium species are the causative agent of malaria and are transmitted by Anopheles mosquitoes, affecting large human populations, whereas Babesia spp., transmitted through the bite of Ixodes ticks cause babesiosis.
In this dissertation, we investigate the essential genome of these parasites using high-throughput transposon mutagenesis. Identifying the essential genome is key to finding new drug targets and understanding resistance mechanisms, a crucial pursuit given the rising resistance to frontline antimalarial drugs and the challenges …
Integrating Radiogenomics And Machine Learning In Musculoskeletal Oncology Care, Rahul Kumar, Kyle Sporn, Akshay Khanna, Phani Paladugu, Chirag Gowda, Alex Ngo, Ram Jagadeesan, Nasif Zaman, Alireza Tavakkoli
Integrating Radiogenomics And Machine Learning In Musculoskeletal Oncology Care, Rahul Kumar, Kyle Sporn, Akshay Khanna, Phani Paladugu, Chirag Gowda, Alex Ngo, Ram Jagadeesan, Nasif Zaman, Alireza Tavakkoli
Department of Medicine Faculty Papers
Musculoskeletal tumors present a diagnostic challenge due to their rarity, histological diversity, and overlapping imaging features. Accurate characterization is essential for effective treatment planning and prognosis, yet current diagnostic workflows rely heavily on invasive biopsy and subjective radiologic interpretation. This review explores the evolving role of radiogenomics and machine learning in improving diagnostic accuracy for bone and soft tissue tumors. We examine integrating quantitative imaging features from MRI, CT, and PET with genomic and transcriptomic data to enable non-invasive tumor profiling. AI-powered platforms employing convolutional neural networks (CNNs) and radiomic texture analysis show promising results in tumor grading, subtype differentiation …
Network Analysis Of Antimicrobial Resistance In Staphylococcus Aureus: Characterization Of Hub Genes And Their Functional Implications, Md Imran Hasan, Davida Smyth, Jeong Yang, Ashley Teufel
Network Analysis Of Antimicrobial Resistance In Staphylococcus Aureus: Characterization Of Hub Genes And Their Functional Implications, Md Imran Hasan, Davida Smyth, Jeong Yang, Ashley Teufel
Masters Theses (Archived)
Antimicrobial resistance is a major cause of morbidity and mortality in patients with S. aureus infections. In this study, we analyzed genes, molecular mechanisms, and pathways driving drug resistance in S. aureus using network analysis. Using whole-genome sequencing (WGS) data and systems biology approaches, we identified 229 AMR-associated genes and constructed a protein-protein interaction network among these genes. Through network topology and functional enrichment analyses, we not only confirmed their association with resistance, but also highlighted the central roles of these genes in resistance pathways, such as efflux, target replacement, and target protection, which are directly linked to multiple drug …
Investigating The Effects Of Transcription Factor Binding And Genetic Variants In The Striatum Of Post-Mortem Cohorts With Opioid Use Disorder, Rajashree Chakraborty
Investigating The Effects Of Transcription Factor Binding And Genetic Variants In The Striatum Of Post-Mortem Cohorts With Opioid Use Disorder, Rajashree Chakraborty
Theses & Dissertations
The opioid crisis has emerged as one of the most pressing public health challenges of our time, with Opioid Use Disorder (OUD) affecting millions of lives across the globe. Studying OUD is not merely an academic pursuit but a critical necessity in addressing this multifaceted epidemic. The urgency of this research is underscored by the staggering prevalence of OUD, with an estimated 3.7% of U.S. adults requiring treatment in 2022 alone. Despite the availability of effective medications for OUD, a significant treatment gap persists, with only a quarter of those in need receiving these life-saving interventions. The far-reaching consequences of …
Integration Of Multi-Omics Datasets Evaluating Structural And Functional Features Of The Gut Microbiome In People With Hiv (Pwh)., Aakarsha Vijayakumar Rao
Integration Of Multi-Omics Datasets Evaluating Structural And Functional Features Of The Gut Microbiome In People With Hiv (Pwh)., Aakarsha Vijayakumar Rao
Electronic Theses and Dissertations
Gut dysbiosis characterized by reduced abundance of beneficial butyrate-producing bacteria has been independently linked to HIV-1 infection and heavy alcohol drinking. Further, gut dysbiosis results in loss of gut barrier integrity, microbial translocation and host-specific systemic inflammation. Therefore, to evaluate the functional consequences of structural changes in the gut microbiome, integrated data analysis is imperative. In this dissertation, we perform integrated analyses using data from multi-omics platforms to examine the structural and functional features of the gut microbiome of PWH. Gut microbiome composition was evaluated by sequencing V4 region of 16S rDNA, concentrations of metabolites and host-specific immune markers were …
Chromosome Evolution Model Reveals Hidden Variation In Karyotype-Driven Speciation In Ferns, Thomas Buchloh
Chromosome Evolution Model Reveals Hidden Variation In Karyotype-Driven Speciation In Ferns, Thomas Buchloh
All Theses
Because karyotype change commonly generates reproductive isolation between diverging species, rapid karyotype evolution, like that seen in plants, may increase the total rate of diversification. However, few studies have investigated this predicted relationship. Tests of this prediction have identified a positive correlation between the rate of karyotype change (specifically whole genome duplications) and diversification, but tests have been restricted to relatively small and young clades. Fortunately, novel macroevolutionary models have recently become available to investigate patterns of karyotype evolution in large phylogenies, providing new opportunity to investigate variation in karyotype driven diversification. Ferns are one of the most karyotype rich …
Elucidating The Multi-Omics Of Early-Onset Colorectal Cancer, Jumanah Alshenaifi
Elucidating The Multi-Omics Of Early-Onset Colorectal Cancer, Jumanah Alshenaifi
Dissertations and Theses (Open Access)
The incidence and mortality rates of sporadic early-onset colorectal cancer have increased in recent decades, but there is no clear etiological basis for this trend. EOCRC is commonly defined as colon and rectal cancers diagnosed before the age of 50 years. The rising incidence of EOCRC has made it the second most common cancer and the third leading cause of cancer death in this age group. The rising incidence of EOCRC is also documented internationally in more than 20 countries across different continents. Clinically, EOCRC has a distinct, more aggressive clinical profile than LOCRC. While approximately 15% of EOCRC cases …
Developing A Small Molecule To Inhibit Hsf1 Expression In Cancer And Evaluating Natural Genetic Variation In Small Molecule Toxicity., Michaela Kendal Foley
Developing A Small Molecule To Inhibit Hsf1 Expression In Cancer And Evaluating Natural Genetic Variation In Small Molecule Toxicity., Michaela Kendal Foley
Theses and Dissertations
Each year cancer affects nearly 20 million people worldwide and genetic differences across populations can impact cancer onset and progression. Specifically, tumors with high levels of HSF1, the master regulator of the cytoprotective heat shock response (HSR), are correlated with poor patient outcomes in multiple cancers such as prostate, breast, and melanoma. Subsequently, the development of pharmacological inhibitors of HSF1 represents a promising strategy for anticancer therapeutics. Using a luciferase-based transcriptional reporter, two small molecule libraries were screened for inhibitors of HSF1 expression in human embryonic kidney cells, yielding ten compounds that decrease HSF1 expression. To identify if cancer lines …
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh
Computer Science ETDs
Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …
Hierarchical Lineage Tracing To Unravel Mechanisms Of Cancer Treatment Resistance, Rachel Danielle Saxe
Hierarchical Lineage Tracing To Unravel Mechanisms Of Cancer Treatment Resistance, Rachel Danielle Saxe
Dartmouth College Ph.D Dissertations
Cancer cells adapt to treatment, leading to the emergence of clones that are more aggressive and resistant to anti-cancer therapies. We have a limited understanding of the development of treatment resistance as we lack technologies to map the evolution of cancer under the selective pressure of treatment. To address this, we developed a hierarchical, dynamic lineage tracing method called FLARE (Following Lineage Adaptation and Resistance Evolution). We use this technique to track the progression of acute myeloid leukemia (AML) cell lines through exposure to Cytarabine (AraC), a front-line treatment in AML, in vitro and in vivo. We map distinct cellular …
Analysis Of Chromatin Accessibility Changes In Endothelial Cells Exposed To Plastic Contaminants, Mikhail Y. Salnikov, Carly Boye, David B. Witonsky, Gabrielle Garlicki, Adnan Alazizi, Francesca Luca, Roger Pique-Regi
Analysis Of Chromatin Accessibility Changes In Endothelial Cells Exposed To Plastic Contaminants, Mikhail Y. Salnikov, Carly Boye, David B. Witonsky, Gabrielle Garlicki, Adnan Alazizi, Francesca Luca, Roger Pique-Regi
Medical Student Research Symposium
Degradation products from everyday plastic products are known to bioaccumulate and have also been shown to contaminate drinking water and food sources. BPA and phthalates are endocrine disrupting chemicals and plastic components that have previously been associated with endothelial cell dysfunction, atherosclerotic and other adverse cardiovascular events. However, there is a limited understanding of the mechanisms underlying these associations, such as genome-wide chromatin accessibility changes in endothelial cells exposed to these compounds. The purpose of this study is to explore genome-wide changes in chromatin accessibility associated with plastic exposure, as well as the discovery of transcription factor binding motifs dysregulated …
Cloning And Expression Of Scamp3 In Escherichia Coli Bl21(De3) With In Silico Sequence-Based Cancer Epitopes Prediction, Selly Setiati Rajagukguk, Sabar Pambudi, Astari Dwiranti, Doddy Irawan Setyo Utomo, Anom Bowolaksono
Cloning And Expression Of Scamp3 In Escherichia Coli Bl21(De3) With In Silico Sequence-Based Cancer Epitopes Prediction, Selly Setiati Rajagukguk, Sabar Pambudi, Astari Dwiranti, Doddy Irawan Setyo Utomo, Anom Bowolaksono
Makara Journal of Science
Secretory carrier membrane protein 3 (SCAMP3) is a crucial membrane protein involved in intracellular vesicle traffick-ing and exocytosis. The SCAMP3 expression has been observed in diverse cancer types, such as melanoma, glioma, hepatocellular and breast cancer. Increased SCAMP3 expression has been reported in certain cancer cells relative to that in normal cells, suggesting the potential role of SCAMP3 in cancer development or progression. In this study, we successfully cloned and expressed SCAMP3 in Escherichia coli strain BL21(DE3). SCAMP3 was amplified and insert-ed directionally into the prokaryotic expression vector pET21d(+). The transformation of recombinant plasmid into E. coli BL21(DE3) cells were …
Screening For Penicillin G Acylase (Pga)-Producing Bacteria And Gene Cloning Using Degenerate Oligonucleotide Primed-Pcr, Masdalifah Masdalifah, Sri Rezeki Wulandari, Gabriela Christy Sabbathini, Maria Ulfah, Dini Achnafani, Ahmad Wibisana, Feronika Heppy Sriherfyna, Is Helianti, Niknik Nurhayati
Screening For Penicillin G Acylase (Pga)-Producing Bacteria And Gene Cloning Using Degenerate Oligonucleotide Primed-Pcr, Masdalifah Masdalifah, Sri Rezeki Wulandari, Gabriela Christy Sabbathini, Maria Ulfah, Dini Achnafani, Ahmad Wibisana, Feronika Heppy Sriherfyna, Is Helianti, Niknik Nurhayati
Makara Journal of Science
The growing concern over antibiotic resistance has driven global efforts to explore innovative solutions, including the use of Penicillin G acylase (PGA) to produce semisynthetic β-lactam antibiotics. This study screened four potential in-tracellular PGA-producing bacteria: Alcaligenes faecalis InaCC B444 (AfPGA), Kluyvera cryocrescens InaCC B850 (KcPGA), Providencia rettgeri InaCC B25 (Pr25PGA), and P. rettgeri InaCC B466 (Pr466PGA). Penicillin G Acylase encoding genes (pgas) were isolated from them using a Degenerate Oligonucleotide Primed-PCR (DOP-PCR) approach and sequenced. Microbiological assays confirmed all tested crude extracts to exhibit inhibitory effects. Penicillin G was used for evaluating hydrolytic activity and 6-Amino Penicillanic Acid (6-APA) coupled …
Identification Of Novel Argonaute Proteins Using A Metagenomic Mining Approach, Lobna Abdallah Ghonaim
Identification Of Novel Argonaute Proteins Using A Metagenomic Mining Approach, Lobna Abdallah Ghonaim
Theses and Dissertations
Gene editing is one of the most promising tools in science. It enables precise modifications of an organism's genetic material. Metagenomics is considered a powerful tool that unlocks the broad genetic potential found in uncultured microbial communities. Exploring the genetic diversity of uncultured microbial communities helps identify novel functional proteins with unique properties and make the best use of these diverse microbial ecosystems.
We developed and employed a metagenomic-based approach to mine more than 1000 metagenomes for prokaryotic argonaute proteins (pAgos), a potential gene editing machinery encoded in bacterial and archaeal genomes. Our workflow involved strict quality control, sequence assembly, …
The Role Of Secondary And Tertiary Structure In The Cap-Independent Translation Of Fgf-9 And Hif-1-Alpha, Amanda Michelle Whittaker
The Role Of Secondary And Tertiary Structure In The Cap-Independent Translation Of Fgf-9 And Hif-1-Alpha, Amanda Michelle Whittaker
Dissertations, Theses, and Capstone Projects
Under normoxic conditions, eukaryotes initiate translation of RNA through eIF4E recognition of the 5’ cap. However, under cellular stress, eukaryotic translation must be initiated through a 4E-independent, or “cap-independent” mechanism, involving eukaryotic initiation factor 4G (eIF4G) binding directly to the 5’ untranslated regions (5’ UTR) of the RNA. eIF4G binding then recruits the ribosome to the transcript. While this mechanism is useful for translation of apoptotic transcripts and transcripts involved in cell survival, cap-independent translation is also utilized by oncogenic RNA for tumorigenesis. Previous work by our lab and others has categorized this recruitment and initiation mechanism as either internal-ribosome-entry-site …
Taming Biological Complexity Through The Use Of Symmetries, Luis A. Álvarez-García
Taming Biological Complexity Through The Use Of Symmetries, Luis A. Álvarez-García
Dissertations, Theses, and Capstone Projects
The study of biological systems is, inherently, the study of very complex systems. This is essentially due to the fact that they are made up of numerous, often very complicated, interactions between an extensive number of components. Often necessitating an abundance of quantitative parameters and details for a precise description. The human brain for ex- ample, consisting of ∼ 80 billion neurons with ∼ 800 to 100 trillion connections between them, each of them depending on a large set of parameters. Even simpler examples such as bacterial organisms, such as E. coli and B. subtilis, which we focus on …
Doves And Symbiotic Bacteria: Whole-Genome Assembly And Phylogenomic Inference, Mariam Topchyan
Doves And Symbiotic Bacteria: Whole-Genome Assembly And Phylogenomic Inference, Mariam Topchyan
Theses and Dissertations
Part 1
I examined a potential source of variation in phylogenetic inference, investigating whether handling single-nucleotide polymorphisms (SNPs) through different methods during reference-guided genome assembly of pigeon and dove species impacts phylogenetic reconstruction and inference. Specifically, I created a custom consensus base calling tool that handles allelic variation through two different approaches, either ignoring heterozygous sites by denoting them as an ambiguous character “N,” or using a pseudo-random choice between supported base variants. Through reference-guided assembly of 108 dove species, I generated two datasets of alignments using both consensus calling methods, created two trees from both datasets via supermatrix and …
18s Metabarcode Analyses Of Eukaryotic Species In The Respiratory Microbiomes Of Wild Canids From New Hampshire, Collin Sinclair Blake
18s Metabarcode Analyses Of Eukaryotic Species In The Respiratory Microbiomes Of Wild Canids From New Hampshire, Collin Sinclair Blake
Honors Theses and Capstones
This study is of an exploratory nature and focuses on characterizing the eukaryotic microbiota present in the respiratory tissues of six wild canids and one domestic canine. The contents of this document largely pertain to dry lab analyses of 18S barcodes in bioinformatics programs – primarily QIIME2 running in the GitBash command line, service for which was hosted by the UNH Ron Bioinformatics training server. All procedures listed within the section below were performed by second parties at the UNH Hubbard Center for Genomics Studies (HCGS), the New Hampshire Veterinary Diagnostics Lab (NHVDL), and the UNH Microbial Ecology and Emerging …
Computational Tools For Protein Classification In Metagenomes, Fawad Ullah
Computational Tools For Protein Classification In Metagenomes, Fawad Ullah
Dissertations, Master's Theses and Master's Reports
Advances in genomic sequencing have dramatically increased the amount and the speed at which genomic data is being generated. These technological advances have enabled the ability to profile the genetic information of organisms and communities at unprecedented scales. Many methods have been developed to identify and classify genes within these datasets. However, many generic pipelines for gene annotation struggle to accurately predict specific protein classes that may not be represented in their databases. Two major challenges exist for classification of specific protein classes in metagenomic databases. The first is the fact that many metagenomic assemblies are highly fragmented with many …
Mobula, Bioinspiration, Filter Feeding, Form And Function, J. B. Teeple, S. R. Kahane-Rapport, K. E. Cohen, L. Hamann, J. A. Strother, E. W.M. Paig-Tran
Mobula, Bioinspiration, Filter Feeding, Form And Function, J. B. Teeple, S. R. Kahane-Rapport, K. E. Cohen, L. Hamann, J. A. Strother, E. W.M. Paig-Tran
Biological Sciences Faculty Publications
Mobulas (manta and devil rays) are large-scale ram filter feeders that separate planktonic food particles from large volumes of water with minimal clogging. This contrasts with most human-made filters that can suffer from problematic clogging requiring additional mechanisms for clearing blocked surfaces and maintaining performance. Prior studies have shown that mobulas employ a unique mechanism referred to as ricochet separation to filter feed, whereby captive vortices in filter pores cause particles to bounce off the filter surfaces and away from the filter pores. This mechanism enables the filtration of particles smaller than the pore size and reduced clogging. However, few …
Copula-Based Bayesian Model For Detecting Differential Gene Expression, Prasansha Liyanaarachchi, N. Rao Chaganty
Copula-Based Bayesian Model For Detecting Differential Gene Expression, Prasansha Liyanaarachchi, N. Rao Chaganty
Mathematics & Statistics Faculty Publications
Deoxyribonucleic acid, more commonly known as DNA, is a fundamental genetic material in all living organisms, containing thousands of genes, but only a subset exhibit differential expression and play a crucial role in diseases. Microarray technology has revolutionized the study of gene expression, with two primary types available for expression analysis: spotted cDNA arrays and oligonucleotide arrays. This research focuses on the statistical analysis of data from spotted cDNA microarrays. Numerous models have been developed to identify differentially expressed genes based on the red and green fluorescence intensities measured using these arrays. We propose a novel approach using a Gaussian …
Genomic Epidemiology Of Staphylococcus Aureus Sequence-Type 72, Sachitaa Senthilkumar
Genomic Epidemiology Of Staphylococcus Aureus Sequence-Type 72, Sachitaa Senthilkumar
Honors Undergraduate Theses
Staphylococcus aureus (S. aureus, SA) is a gram-positive bacterial colonizer and pathogen commonly carried on the skin and in the anterior nares of humans. SA colonization may lead to a variety of non-invasive (e.g., skin and soft tissue infections, SSTIs) and invasive (e.g., bacteremia and osteomyelitis) infections. The SA population is comprised of multiple lineages that can be delineated through multi-locus sequence typing (MLST), which compares nucleotide sequences within seven housekeeping genes. Strains belonging to an MLST share a common evolutionary history and phenotypic characteristics like antibiotic resistance or virulence. As a result, the epidemiology of lineages can be …
Characterizing Somatic Variants In Nanopore Data With Machine Learning, Shwethal Sayeeram Trikannad
Characterizing Somatic Variants In Nanopore Data With Machine Learning, Shwethal Sayeeram Trikannad
Master's Projects
Oxford Nanopore Technology (ONT) is a popular long-read sequencer in genomics. However, its high base-calling error rate produces several sequencing artifacts. Detection of somatic variants in ONT sequenced tumor-normal samples remains challenging due to low frequencies. In this study, machine learning was applied to a dataset created by benchmarking ClairS output against HCC1395 and colo829 truth sets to classify variants and artifacts. Relevant features were engineered from sequence context and variant site characteristics to model artifact profiles. HistGradientBoostingClassifier achieved 0.876950 accuracy, outperforming all other models. Variant quality was the top predictor with an aggregate accuracy of over 85%. This work …
Enhancing Clinical Trial Matching In Molecular Diagnostics: Using Natural Language Processing And Clustering Approaches In Hematological Malignancies, Gillian Fanning
Enhancing Clinical Trial Matching In Molecular Diagnostics: Using Natural Language Processing And Clustering Approaches In Hematological Malignancies, Gillian Fanning
Theses and Dissertations
Clinical trial matching is a critical component of personalized medicine, particularly in the management of hematologic malignancies. At Virginia Commonwealth University (VCU) Health, the Molecular Diagnostics (MDX) Lab produces somatic variant reports and recommends clinical trials based on the presence of clinically significant mutations. However, the current manual trial recommendation process is time-intensive and lacks scalability.
This study introduces a computational framework to streamline and standardize clinical trial matching using natural language processing (NLP) and unsupervised clustering. Trial brief descriptions were analyzed to extract frequent terms, and trials were grouped based on term similarity using joint dimensionality reduction and clustering. …
Foundation Models In Bioinformatics, Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, Jianxin Wang
Foundation Models In Bioinformatics, Fei Guo, Renchu Guan, Yaohang Li, Qi Liu, Xiaowo Wang, Can Yang, Jianxin Wang
Computer Science Faculty Publications
With the adoption of foundation models (FMs), artificial intelligence (AI) has become increasingly significant in bioinformatics and has successfully addressed many historical challenges, such as pre-training frameworks, model evaluation and interpretability. FMs demonstrate notable proficiency in managing large-scale, unlabeled datasets, because experimental procedures are costly and labor intensive. In various downstream tasks, FMs have consistently achieved noteworthy results, demonstrating high levels of accuracy in representing biological entities. A new era in computational biology has been ushered in by the application of FMs, focusing on both general and specific biological issues. In this review, we introduce recent advancements in bioinformatics FMs …
Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh
Gramseq-Dta: A Grammar-Based Drug-Target Affinity Prediction Approach Fusing Gene Expression Information, Kasul Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Drug–target affinity (DTA) prediction is a critical aspect of drug discovery. The meaningful representation of drugs and targets is crucial for accurate prediction. Using 1D string-based representations for drugs and targets is a common approach that has demonstrated good results in drug–target affinity prediction. However, these approach lacks information on the relative position of the atoms and bonds. To address this limitation, graph-based representations have been used to some extent. However, solely considering the structural aspect of drugs and targets may be insufficient for accurate DTA prediction. Integrating the functional aspect of these drugs at the genetic level can enhance …
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang
Cath-Ddg: Towards Robust Mutation Effect Prediction On Protein-Protein Interactions Out Of Cath Homologous Superfamily, Guanglei Yu, Xuehua Bi, Teng Ma, Yaohang Li, Jianxin Wang
Computer Science Faculty Publications
Motivation: Protein-protein interactions (PPIs) are fundamental aspects in understanding biological processes. Accurately predicting the effects of mutations on PPIs remains a critical requirement for drug design and disease mechanistic studies. Recently, deep learning models using protein 3D structures have become predominant for predicting mutation effects. However, significant challenges remain in practical applications, in part due to the considerable disparity in generalization capabilities between easy and hard mutations. Specifically, a hard mutation is defined as one with its maximum TM-score < 0.6 when compared to the training set. Additionally, compared to physics-based approaches, deep learning models may overestimate performance due to potential data leakage.
Results: We propose new training/test splits that mitigate data leakage according to the CATH homologous superfamily. Under the constraints of physical …
A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh
A Survey On Deep Learning For Drug-Target Binding Prediction: Models, Benchmarks, Evaluation, And Case Studies, Kusal Debnath, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
Conventional drug discovery is expensive, time-consuming, and prone to failure. Artificial intelligence has become a potent substitute over the last decade, providing strong answers to challenging biological issues in this field. Among these difficulties, drug-target binding (DTB) is a key component of drug discovery techniques. In this context, drug-target affinity and drug–target interaction are complementary and essential frameworks that work together to improve our comprehension of DTB dynamics. In this work, we thoroughly analyze the most recent deep learning models, popular benchmark datasets, and assessment metrics for DTB prediction. We look at the paradigm shift in the development of drug …
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Heterogeneous Clustering Of Multiomics Data For Breast Cancer Subgroup Classification And Detection, Joseph Pateras, Musaddiq Lodi, Pratip Rana, Preetam Ghosh
Computer Science Faculty Publications
The rapid growth of diverse -omics datasets has made multiomics data integration crucial in cancer research. This study adapts the expectation–maximization routine for the joint latent variable modeling of multiomics patient profiles. By combining this approach with traditional biological feature selection methods, this study optimizes latent distribution, enabling efficient patient clustering from well-studied cancer types with reduced computational expense. The proposed optimization subroutines enhance survival analysis and improve runtime performance. This article presents a framework for distinguishing cancer subtypes and identifying potential biomarkers for breast cancer. Key insights into individual subtype expression and function were obtained through differentially expressed gene …