Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Bioinformatics (11)
- Physical Sciences and Mathematics (8)
- Medicine and Health Sciences (7)
- Computer Sciences (6)
- Genomics (6)
-
- Artificial Intelligence and Robotics (5)
- Biochemistry, Biophysics, and Structural Biology (4)
- Biochemistry (3)
- Diseases (3)
- Applied Mathematics (2)
- Chemicals and Drugs (2)
- Data Science (2)
- Ecology and Evolutionary Biology (2)
- Engineering (2)
- Medical Specialties (2)
- Numerical Analysis and Scientific Computing (2)
- Oncology (2)
- Other Ecology and Evolutionary Biology (2)
- Plant Sciences (2)
- Agriculture (1)
- Amino Acids, Peptides, and Proteins (1)
- Analytical, Diagnostic and Therapeutic Techniques and Equipment (1)
- Applied Statistics (1)
- Biology (1)
- Biomedical Engineering and Bioengineering (1)
- Biostatistics (1)
- Biotechnology (1)
- Institution
- Publication Year
- Publication
-
- Biochemistry Publications (2)
- Computer Science Faculty Publications (2)
- Dartmouth Scholarship (2)
- Graduate Student Theses, Dissertations, & Professional Papers (2)
- Biological Sciences Faculty Publications (1)
-
- Biology Faculty Publications (1)
- Complex Biosystems Program: Dissertations and Student Research (1)
- Dissertations, Master's Theses and Master's Reports (1)
- KGI Theses and Dissertations (1)
- Research outputs 2014 to 2021 (1)
- Student Publications & Research (1)
- Theses and Dissertations--Mathematics (1)
- Theses and Dissertations--Plant and Soil Sciences (1)
- U.C. Berkeley Division of Biostatistics Working Paper Series (1)
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Computational Biology
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Benchmarking Batch-Effect Correction Methods Towards The Construction Of A Triple-Negative Breast Cancer Cell Atlas, Peter Scheible, Amy H. Tang, Jing He, Jiangwen Sun
Computer Science Faculty Publications
Triple-negative breast cancer (TNBC) requires detailed cellular mapping given its aggressive nature, immense tumor heterogeneity and genetic diversity. We integrated 156,794 cells from six scRNA-seq datasets—including tumors, metastases, and cell lines—to build a TNBC scRNA cell atlas, focusing on batch effect mitigation while maintaining biological and molecular details. Preprocessing f ilters noise, normalizes data, and leverages PCA for integration readiness. We utilized scANVI, a semi-supervised tool, to align datasets, preserving TNBC’s complex tumor heterogeneity via marker annotations [1]. UMAPs demonstrate biological clustering in integrated data, contrasted with datasetdriven unintegrated patterns. Assessments verifying effective batch correction. This method aligns with NASA’s …
From Code To Crops: Harnessing Bioinformatics And Artificial Intelligence (Ai) In Agricultural Omics, Lakshay Anand
From Code To Crops: Harnessing Bioinformatics And Artificial Intelligence (Ai) In Agricultural Omics, Lakshay Anand
Theses and Dissertations--Plant and Soil Sciences
Global agricultural faces numerous challenges, such as climate change, resource limitations, novel pests and diseases, increasing costs, and the ever-increasing human population. To tackle these challenges, we need innovative strategies that combine new technologies and data analytics approaches to enhance agricultural output, promote sustainable methods, and optimize resource allocation. The key to this innovation lies in understanding the complex molecular web within plants that governs their growth, defense, and adaptability mechanisms. By mastering this molecular network, we can cultivate crops that are more resilient, sustainable, and suitable for different climatic terrains. Moreover, studying the symbiotic relationship between plants and microorganisms …
Ai And Ml-Based Risk Assessment Of Chemicals: Predicting Carcinogenic Risk From Chemical-Induced Genomic Instability, Ajay Vikram Singh, Preeti Bhardwaj, Peter Laux, Prachi Pradeep, Madleen Busse, Andreas Luch, Akihiko Hirose, Christopher J. Osgood, Michael W. Stacey
Ai And Ml-Based Risk Assessment Of Chemicals: Predicting Carcinogenic Risk From Chemical-Induced Genomic Instability, Ajay Vikram Singh, Preeti Bhardwaj, Peter Laux, Prachi Pradeep, Madleen Busse, Andreas Luch, Akihiko Hirose, Christopher J. Osgood, Michael W. Stacey
Biological Sciences Faculty Publications
Chemical risk assessment plays a pivotal role in safeguarding public health and environmental safety by evaluating the potential hazards and risks associated with chemical exposures. In recent years, the convergence of artificial intelligence (AI), machine learning (ML), and omics technologies has revolutionized the field of chemical risk assessment, offering new insights into toxicity mechanisms, predictive modeling, and risk management strategies. This perspective review explores the synergistic potential of AI/ML and omics in deciphering clastogen-induced genomic instability for carcinogenic risk prediction. We provide an overview of key findings, challenges, and opportunities in integrating AI/ML and omics technologies for chemical risk assessment, …
Convolutional Neural Network-Based Gene Prediction Using Buffalograss As A Model System, Michael Morikone
Convolutional Neural Network-Based Gene Prediction Using Buffalograss As A Model System, Michael Morikone
Complex Biosystems Program: Dissertations and Student Research
The task of gene prediction has been largely stagnant in algorithmic improvements compared to when algorithms were first developed for predicting genes thirty years ago. Rather than iteratively improving the underlying algorithms in gene prediction tools by utilizing better performing models, most current approaches update existing tools through incorporating increasing amounts of extrinsic data to improve gene prediction performance. The traditional method of predicting genes is done using Hidden Markov Models (HMMs). These HMMs are constrained by having strict assumptions made about the independence of genes that do not always hold true. To address this, a Convolutional Neural Network (CNN) …
An Approach To Developing Benchmark Datasets For Protein Secondary Structure Segmentation From Cryo-Em Density Maps, Thu Nguyen, Yongcheng Mu, Jiangwen Sun, Jing He
An Approach To Developing Benchmark Datasets For Protein Secondary Structure Segmentation From Cryo-Em Density Maps, Thu Nguyen, Yongcheng Mu, Jiangwen Sun, Jing He
Computer Science Faculty Publications
More and more deep learning approaches have been proposed to segment secondary structures from cryo-electron density maps at medium resolution range (5--10Å). Although the deep learning approaches show great potential, only a few small experimental data sets have been used to test the approaches. There is limited understanding about potential factors, in data, that affect the performance of segmentation. We propose an approach to generate data sets with desired specifications in three potential factors - the protein sequence identity, structural contents, and data quality. The approach was implemented and has generated a test set and various training sets to study …
Cbp60-Db: An Alphafold-Predicted Plant Kingdom-Wide Database Of The Calmodulin-Binding Protein 60 (Cbp60) Protein Family With A Novel Structural Clustering Algorithm, Keaun Amani, Vanessa Shivnauth, Christian Castroverde
Cbp60-Db: An Alphafold-Predicted Plant Kingdom-Wide Database Of The Calmodulin-Binding Protein 60 (Cbp60) Protein Family With A Novel Structural Clustering Algorithm, Keaun Amani, Vanessa Shivnauth, Christian Castroverde
Biology Faculty Publications
Molecular genetic analyses in the model species Arabidopsis thaliana have demonstrated the major roles of different CAM-BINDING PROTEIN 60 (CBP60) proteins in growth, stress signaling, and immune responses. Prominently, CBP60g and SARD1 are paralogous CBP60 transcription factors that regulate numerous components of the immune system, such as cell surface and intracellular immune receptors, MAP kinases, WRKY transcription factors, and biosynthetic enzymes for immunity-activating metabolites salicylic acid (SA) and N-hydroxypipecolic acid (NHP). However, their function, regulation and diversification in most species remain unclear. Here we have created CBP60-DB, a structural and bioinformatic database that comprehensively characterized 1052 CBP60 gene homologs …
Applications Of Machine Learning In Microbial Forensics, Ryan B. Ghannam
Applications Of Machine Learning In Microbial Forensics, Ryan B. Ghannam
Dissertations, Master's Theses and Master's Reports
Microbial ecosystems are complex, with hundreds of members interacting with each other and the environment. The intricate and hidden behaviors underlying these interactions make research questions challenging – but can be better understood through machine learning. However, most machine learning that is used in microbiome work is a black box form of investigation, where accurate predictions can be made, but the inner logic behind what is driving prediction is hidden behind nontransparent layers of complexity.
Accordingly, the goal of this dissertation is to provide an interpretable and in-depth machine learning approach to investigate microbial biogeography and to use micro-organisms as …
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Ensemble Protein Inference Evaluation, Kyle Lee Lucke
Graduate Student Theses, Dissertations, & Professional Papers
The Protein inference problem is becoming an increasingly important tool that aids in the characterization of complex proteomes and analysis of complex protein samples. In bottom-up shotgun proteomics experiments the metrics for evaluation (like AUC and calibration error) are based on an often imperfect target-decoy database. These metrics make the inherent assumption that all of the proteins in the target set are present in the sample being analyzed. In general, this is not the case, they are typically a mix of present and absent proteins. To objectively evaluate inference methods, protein standard datasets are used. These datasets are special in …
Pathway-Extended Gene Expression Signatures Integrate Novel Biomarkers That Improve Predictions Of Patient Responses To Kinase Inhibitors, Ashis Jem Bagchee-Clark, Eliseos J. Mucaki, Tyson Whitehead, Peter Rogan
Pathway-Extended Gene Expression Signatures Integrate Novel Biomarkers That Improve Predictions Of Patient Responses To Kinase Inhibitors, Ashis Jem Bagchee-Clark, Eliseos J. Mucaki, Tyson Whitehead, Peter Rogan
Biochemistry Publications
No abstract provided.
Machine Learning Prediction Of Glioblastoma Patient One-Year Survival, Andrew Du '20, Warren Mcgee, Jane Y. Wu
Machine Learning Prediction Of Glioblastoma Patient One-Year Survival, Andrew Du '20, Warren Mcgee, Jane Y. Wu
Student Publications & Research
Glioblastoma (GBM) is a grade IV astrocytoma formed primarily from cancerous astrocytes and sustained by intense angiogenesis. GBM often causes non-specific symptoms, creating difficulty for diagnosis. This study aimed to utilize machine learning techniques to provide an accurate one-year survival prognosis for GBM patients using clinical and genomic data from the Chinese Glioma Genome Atlas. Logistic regression (LR), support vector machines (SVM), random forest (RF), and ensemble models were used to identify and select predictors for GBM survival and to classify patients into those with an overall survival (OS) of less than one year and one year or greater. With …
Transcription Factor Binding Site Clusters Identify Target Genes With Similar Tissue-Wide Expression And Buffer Against Mutations., Peter Rogan, Ruipeng Lu
Transcription Factor Binding Site Clusters Identify Target Genes With Similar Tissue-Wide Expression And Buffer Against Mutations., Peter Rogan, Ruipeng Lu
Biochemistry Publications
Background: The distribution and composition of cis-regulatory modules composed of transcription factor (TF) binding site (TFBS) clusters in promoters substantially determine gene expression patterns and TF targets. TF knockdown experiments have revealed that TF binding profiles and gene expression levels are correlated. We use TFBS features within accessible promoter intervals to predict genes with similar tissue-wide expression patterns and TF targets using Machine Learning (ML). Methods: Bray-Curtis Similarity was used to identify genes with correlated expression patterns across 53 tissues. TF targets from knockdown experiments were also analyzed by this approach to set up the ML framework. TFBSs were …
A Comparative Evaluation Of The Generalised Predictive Ability Of Eight Machine Learning Algorithms Across Ten Clinical Metabolomics Data Sets For Binary Classification, Kevin M. Mendez, Stacey N. Reinke, David I. Broadhurst
A Comparative Evaluation Of The Generalised Predictive Ability Of Eight Machine Learning Algorithms Across Ten Clinical Metabolomics Data Sets For Binary Classification, Kevin M. Mendez, Stacey N. Reinke, David I. Broadhurst
Research outputs 2014 to 2021
Introduction:
Metabolomics is increasingly being used in the clinical setting for disease diagnosis, prognosis and risk prediction. Machine learning algorithms are particularly important in the construction of multivariate metabolite prediction. Historically, partial least squares (PLS) regression has been the gold standard for binary classification. Nonlinear machine learning methods such as random forests (RF), kernel support vector machines (SVM) and artificial neural networks (ANN) may be more suited to modelling possible nonlinear metabolite covariance, and thus provide better predictive models.
Objectives:
We hypothesise that for binary classification using metabolomics data, non-linear machine learning methods will provide superior generalised predictive ability when …
Recurrent Neural Networks And Their Applications To Rna Secondary Structure Inference, Devin Willmott
Recurrent Neural Networks And Their Applications To Rna Secondary Structure Inference, Devin Willmott
Theses and Dissertations--Mathematics
Recurrent neural networks (RNNs) are state of the art sequential machine learning tools, but have difficulty learning sequences with long-range dependencies due to the exponential growth or decay of gradients backpropagated through the RNN. Some methods overcome this problem by modifying the standard RNN architecure to force the recurrent weight matrix W to remain orthogonal throughout training. The first half of this thesis presents a novel orthogonal RNN architecture that enforces orthogonality of W by parametrizing with a skew-symmetric matrix via the Cayley transform. We present rules for backpropagation through the Cayley transform, show how to deal with the Cayley …
Understanding Huntington's Disease Using Machine Learning Approaches, Sonali Lokhande
Understanding Huntington's Disease Using Machine Learning Approaches, Sonali Lokhande
KGI Theses and Dissertations
Huntington’s disease (HD) is a debilitating neurodegenerative disorder with a complex pathophysiology. Despite extensive studies to study the disease, the sequence of events through which mutant Huntingtin (mHtt) protein executes its action still remains elusive. The phenotype of HD is an outcome of numerous processes initiated by the mHtt protein along with other proteins that act as either suppressors or enhancers of the effects of mHtt protein and PolyQ aggregates. Utilizing an integrative systems biology approach, I construct and analyze a Huntington’s disease integrome using human orthologs of protein interactors of wild type and mHtt protein. Analysis of this integrome …
K-Mer Analysis Pipeline For Classification Of Dna Sequences From Metagenomic Samples, Russell Kaehler
K-Mer Analysis Pipeline For Classification Of Dna Sequences From Metagenomic Samples, Russell Kaehler
Graduate Student Theses, Dissertations, & Professional Papers
Biological sequence datasets are increasing at a prodigious rate. The volume of data in these datasets surpasses what is observed in many other fields of science. New developments wherein metagenomic DNA from complex bacterial communities is recovered and sequenced are producing a new kind of data known as metagenomic data, which is comprised of DNA fragments from many genomes. Developing a utility to analyze such metagenomic data and predict the sample class from which it originated has many possible implications for ecological and medical applications. Within this document is a description of a series of analytical techniques used to process …
Detecting Gene-Gene Interactions Using A Permutation-Based Random Forest Method, Jing Li, James D. Malley, Angeline S. Andrew, Margaret R. Karagas, Jason H. Moore
Detecting Gene-Gene Interactions Using A Permutation-Based Random Forest Method, Jing Li, James D. Malley, Angeline S. Andrew, Margaret R. Karagas, Jason H. Moore
Dartmouth Scholarship
Identifying gene-gene interactions is essential to understand disease susceptibility and to detect genetic architectures underlying complex diseases. Here, we aimed at developing a permutation-based methodology relying on a machine learning method, random forest (RF), to detect gene-gene interactions. Our approach called permuted random forest (pRF) which identified the top interacting single nucleotide polymorphism (SNP) pairs by estimating how much the power of a random forest classification model is influenced by removing pairwise interactions.
Machine Learning Methods Enable Predictive Modeling Of Antibody Feature:Function Relationships In Rv144 Vaccinees, Ickwon Choi, Amy W. Chung, Todd J. Suscovich, Supachai Rerks-Ngarm, Punnee Pitisuttithum, Sorachai Nitayapha, Jaranit Kaewkungwal, Robert J. O'Connell, Donald Francis, Merlin L. Robb, Nelson L. Michael, Jerome H. Kim, Galit Alter, Margaret E. Ackerman, Chris Bailey-Kellogg
Machine Learning Methods Enable Predictive Modeling Of Antibody Feature:Function Relationships In Rv144 Vaccinees, Ickwon Choi, Amy W. Chung, Todd J. Suscovich, Supachai Rerks-Ngarm, Punnee Pitisuttithum, Sorachai Nitayapha, Jaranit Kaewkungwal, Robert J. O'Connell, Donald Francis, Merlin L. Robb, Nelson L. Michael, Jerome H. Kim, Galit Alter, Margaret E. Ackerman, Chris Bailey-Kellogg
Dartmouth Scholarship
The adaptive immune response to vaccination or infection can lead to the production of specific antibodies to neutralize the pathogen or recruit innate immune effector cells for help. The non-neutralizing role of antibodies in stimulating effector cell responses may have been a key mechanism of the protection observed in the RV144 HIV vaccine trial. In an extensive investigation of a rich set of data collected from RV144 vaccine recipients, we here employ machine learning methods to identify and model associations between antibody features (IgG subclass and antigen specificity) and effector function activities (antibody dependent cellular phagocytosis, cellular cytotoxicity, and cytokine …
Permutation-Based Pathway Testing Using The Super Learner Algorithm, Paul Chaffee, Alan E. Hubbard, Mark L. Van Der Laan
Permutation-Based Pathway Testing Using The Super Learner Algorithm, Paul Chaffee, Alan E. Hubbard, Mark L. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many diseases and other important phenotypic outcomes are the result of a combination of factors. For example, expression levels of genes have been used as input to various statistical methods for predicting phenotypic outcomes. One particular popular variety is the so-called gene set enrichment analysis (GSEA). This paper discusses an augmentation to an existing strategy to estimate the significance of an associations between a disease outcome and a predetermined combination of biological factors, based on a specific data adaptive regression method (the "Super Learner," van der Laan et al., 2007). The procedure uses an aggressive search procedure, potentially resulting in …