Open Access. Powered by Scholars. Published by Universities.®

Computational Biology Commons™

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 12 of 12

Full-Text Articles in Computational Biology

Large Scale Kmer-Based Proteomic Analysis: An Application Towards Evolutionary Constraint Discovery, Matthew Chak Jun 2026

Large Scale Kmer-Based Proteomic Analysis: An Application Towards Evolutionary Constraint Discovery, Matthew Chak

Master's Theses

Large protein databases now make it possible to study short peptides across natural protein sequence space at unprecedented scale, but exhaustively counting k-mers across billions of protein sequences remains computationally difficult. This thesis develops an exact amino-acid k-mer counting method based on direct addressing, in which fixed-length amino-acid strings are encoded as base-20 integers and updated with a sliding-window recurrence. By avoiding key storage and collision resolution, this approach removes overhead inherent to hash-map-based methods when the k-mer space is sufficiently dense. A memory analysis shows when direct addressing is preferable to open-addressing hash tables, and expected-saturation calculations motivate its …


Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher Jun 2025

Topodino: Self-Supervised Topological Representation Learning For Neuronal Morphologies, Yasser Binbisher

Master's Theses

Neuronal cell types are categorized by transcriptomic identity, yet their morphological heterogeneity defies this classification. In response, researchers have adopted unsupervised graph representation learning as a tool to reveal morphological variation within single-class transcriptomic types. However, the complex geometry of neuronal morphology—especially long axons and dense dendrites—challenges graph neural networks, which struggle with message propagation across extended structures. To mitigate this, current approaches enforce sub-sampling on neuronal graphs and omit axons entirely, sacrificing critical biological features for computational efficiency. To overcome this trade-off, this thesis introduces TopoDINO, a self-supervised, topology-aware representation learning model designed to preserve the full hierarchical organization …


Lotus: A Web-Based Computational Tool For The Preliminary Investigation Of A Novel Mst Method Utilizing A Library Of 16s Rrna Bacteroides Otus, Ginger Dewitte May 2021

Lotus: A Web-Based Computational Tool For The Preliminary Investigation Of A Novel Mst Method Utilizing A Library Of 16s Rrna Bacteroides Otus, Ginger Dewitte

Master's Theses

Microbial Source Tracking (MST) is a field of study that attempts to identify the source of fecal contamination in waterways in order to assist with development of remediation strategies. Biologists at Cal Poly Center for Applications in Biotechnology (CAB) are developing a new MST method using microbes from the genus Bacteroides. Bacteroides species are host-specific microorganisms that can theoretically be used to trace back to a single host species. After fecal samples are collected, biologists use Next-Generation Sequencing (NGS) techniques to obtain only the genetic sequences of microorganisms belonging to the phylum Bacteroidetes. Investigators hypothesize that similar sequences belong …


Optimization Of A Genomic Editing System Using Crispr/Cas9-Induced Site-Specific Gene Integration, Jillian L. Mccool Ms., Nick Hum, Gabriela G. Loots Aug 2016

Optimization Of A Genomic Editing System Using Crispr/Cas9-Induced Site-Specific Gene Integration, Jillian L. Mccool Ms., Nick Hum, Gabriela G. Loots

STAR Program Research Presentations

The CRISPR-Cas system is an adaptive immune system found in bacteria which helps protect against the invasion of other microorganisms. This system induces double stranded breaks at precise genomic loci (1) in which repairs are initiated and insertions of a target are completed in the process. This mechanism can be used in eukaryotic cells in combination with sgRNAs (1) as a tool for genome editing. By using this CRISPR-Cas system, in addition to the “safe harbor locus,” ROSAβ26, the incorporation of a target gene into a site that is not susceptible to gene silencing effects can be achieved through few …


Data Development And Analysis Pathways For Marine Mammals And Turtles: Creating A User Interface, Sarina Fernandez, Warren Asfazadour, Eric Archer, Lisa Komoroske Aug 2016

Data Development And Analysis Pathways For Marine Mammals And Turtles: Creating A User Interface, Sarina Fernandez, Warren Asfazadour, Eric Archer, Lisa Komoroske

STAR Program Research Presentations

A major obstacle in genetic research is developing streamlined methods for analyzing large amounts of data. The statistical computer programming language R provides users with the ability to develop packages containing specific functions in order to create more accessible data analysis pipelines. However, writing code in R can still be intimidating to those with little to no coding experience. Fortunately, the R package shiny provides a framework for developing web applications based on R functions. Using shiny, we developed a user-friendly web application containing functions of the R package strataG. The strataG package contains several functions for summarizing genetic data …


Using Hadoop To Identify False Positives In Bacterial Strain Typing From Dna Fingerprints, Colin C. Adams Jun 2016

Using Hadoop To Identify False Positives In Bacterial Strain Typing From Dna Fingerprints, Colin C. Adams

Computer Science and Software Engineering

Pyroprinting is a novel technique used by the Department of Biological Sciences to obtain “fingerprints” from the DNA of E. coli isolates in order to categorize them into strains. To determine the number of false positives that occur in the pyroprinting process, isolates with the same pyroprints needed to be sequenced to see if their underlying alleles match. If they do match, this shows they are indeed the same strain and are a true positive. If the alleles don’t match, they are different strains and are a false positive. To do this 100 isolates with nucleotide identifiers were sequenced. Over …


Characterization Of Putative Wnt3a-Inducible Enhancers, Katelynn C. Lee, Nicholas Hum, Aimy Sebastian, Gabriela Loots Aug 2015

Characterization Of Putative Wnt3a-Inducible Enhancers, Katelynn C. Lee, Nicholas Hum, Aimy Sebastian, Gabriela Loots

STAR Program Research Presentations

The Wnt signaling pathway has been previously shown to play a major role in regulating bone metabolism and it is emerging as a target for the therapeutic intervention of bone thinning disorders such as osteoporosis. Several Wnt proteins have been shown to be expressed in bone and mutations in Wnt pathway members such as Wnt co-receptor Lrp5 and Wnt inhibitor Sost have been shown to be associated with low or high bone mass disorders, however, very little is known about specific roles played by different Wnt ligands in bone development, repair and remodeling. To identify downstream targets of Wnt signaling …


Estimating Migration Rates Between Populations Of Zostera Marina In The San Francisco Bay, Elizabeth S. Gutierrez Sep 2014

Estimating Migration Rates Between Populations Of Zostera Marina In The San Francisco Bay, Elizabeth S. Gutierrez

STAR Program Research Presentations

Eelgrass (Zostera marina) is a highly clonal marine angiosperm that can also reproduce sexually through flowering and seed formation. In a previous study, Fst values from six microsatellite loci suggested that a perennial San Francisco Bay subpopulation at Point Molate (Richmond, California) was able to recover from a drastic 2006 die-off through seed recruitment from neighboring eelgrass subpopulations, changing its reproductive strategy from clonal to sexual. Although Fst measures continue to be widely used in population genetics, the assumptions under which they operate are not always appropriate given certain circumstances, such small population sizes and/or asymmetrical migration rates. Our summer …


Creating A Package In R, Brit Schneiders, Eric Archer Aug 2013

Creating A Package In R, Brit Schneiders, Eric Archer

STAR Program Research Presentations

In a time of increasingly efficient technology and data production, scientists are producing data faster than it can be analyzed. Therefore, user accessibility to data analysis is becoming more and more critical. In general, researchers have a set of raw data and want an efficient means to their final analysis. A package serves as that means by creating a set of functions and making them accessible to the user. Often, a user has a small piece of code to run (a single R script, for example), and that script requires the use of certain functions, which are contained in a …


Physiologically-Based Pharmacokinetic Modeling For Predicting Caffeine/Theophylline-Ciprofloxacin Interactions, David M. Ng, Ali Navid Aug 2013

Physiologically-Based Pharmacokinetic Modeling For Predicting Caffeine/Theophylline-Ciprofloxacin Interactions, David M. Ng, Ali Navid

STAR Program Research Presentations

Dynamics of interactions between the drugs caffeine, theophylline, and ciprofloxacin are predicted using physiologically-based pharmacokinetic (PBPK) modeling. Pharmacokinetic means the model determines where the drugs are distributed in the body over time. Physiologically-based means the anatomy and physiology of the human body are reflected in the structure and functioning of the model. Multiple drugs can interact to increase or decrease their beneficial and/or undesired effects. This is important because some common substances, such as caffeine in coffee, soft drinks, and energy drinks, are actually drugs that affect the body. Ciprofloxacin is an inhibitor of caffeine and theophylline metabolism; such inhibition …


Physiologically-Based Pharmacokinetic Modeling Of Acetaminophen Metabolism And Toxicity, David M. Ng, Ali Navid Aug 2012

Physiologically-Based Pharmacokinetic Modeling Of Acetaminophen Metabolism And Toxicity, David M. Ng, Ali Navid

STAR Program Research Presentations

Acetaminophen is a common analgesic and antipyretic. Metabolism of acetaminophen and acetaminophen-induced liver necrosis are predicted using physiologically-based pharmacokinetic (PBPK) modeling. Pharmacokinetic means the model determines where the drug is distributed in the body over time. Physiologically-based means the anatomy and physiology of the human body is reflected in the structure and functioning of the model. Acetaminophen is usually safe and effective when taken as recommended, but consumption at higher levels may lead to liver damage. Additionally, other factors such as alcoholic liver disease, smoking, and malnutrition affect the maximum safe dose of acetaminophen.


Physiologically-Based Pharmacokinetic Modeling For Predicting Drug-Drug Interactions, David M. Ng, Ali Navid Aug 2011

Physiologically-Based Pharmacokinetic Modeling For Predicting Drug-Drug Interactions, David M. Ng, Ali Navid

STAR Program Research Presentations

Dynamics of interactions between the drugs caffeine and ciprofloxacin are predicted using physiologically-based pharmacokinetic (PBPK) modeling. Pharmacokinetic means the model determines where the drugs are distributed in the body over time. Physiologically-based means the anatomy and physiology of the human body is reflected in the structure and functioning of the model. Multiple drugs can interact to increase or decrease their beneficial and/or undesired effects. This is important because some common substances, such as caffeine in coffee and soft drinks, are actually drugs that affect the body. By implementing the model as a computer program, it is relatively straightforward to perform …