Open Access. Powered by Scholars. Published by Universities.®

Computational Biology Commons

Open Access. Powered by Scholars. Published by Universities.®

754 Full-Text Articles 2,045 Authors 331,672 Downloads 122 Institutions

All Articles in Computational Biology

Faceted Search

754 full-text articles. Page 25 of 32.

Phagephisher: A Pipeline For The Discovery Of Covert Viral Sequences In Complex Genomic Datasets, Thomas Hatzopoulos, Siobhan C. Watkins, Catherine Putonti 2016 Loyola University Chicago

Phagephisher: A Pipeline For The Discovery Of Covert Viral Sequences In Complex Genomic Datasets, Thomas Hatzopoulos, Siobhan C. Watkins, Catherine Putonti

Bioinformatics Faculty Publications

Obtaining meaningful viral information from large sequencing datasets presents unique challenges distinct from prokaryotic and eukaryotic sequencing efforts. The difficulties surrounding this issue can be ascribed in part to the genomic plasticity of viruses themselves as well as the scarcity of existing information in genomic databases. The open-source software PhagePhisher (http://www.putonti-lab.com/phagephisher) has been designed as a simple pipeline to extract relevant information from complex and mixed datasets, and will improve the examination of bacteriophages, viruses, and virally related sequences, in a range of environments. Key aspects of the software include speed and ease of use; PhagePhisher can be used with …


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang 2016 Fox Chase Cancer Center

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret 2016 University of Washington - Seattle Campus

Models For Hsv Shedding Must Account For Two Levels Of Overdispersion, Amalia Magaret

UW Biostatistics Working Paper Series

We have frequently implemented crossover studies to evaluate new therapeutic interventions for genital herpes simplex virus infection. The outcome measured to assess the efficacy of interventions on herpes disease severity is the viral shedding rate, defined as the frequency of detection of HSV on the genital skin and mucosa. We performed a simulation study to ascertain whether our standard model, which we have used previously, was appropriately considering all the necessary features of the shedding data to provide correct inference. We simulated shedding data under our standard, validated assumptions and assessed the ability of 5 different models to reproduce the …


Ten Simple Rules For Digital Data Storage, E. M. Hart, P. Barmby, D. LeBauer, F. Michonneau, S. Mount, P. Mulrooney, T. Poisot, K. H. Woo, Naupaka B. Zimmerman, J. W. Hollister 2016 University of San Francisco

Ten Simple Rules For Digital Data Storage, E. M. Hart, P. Barmby, D. Lebauer, F. Michonneau, S. Mount, P. Mulrooney, T. Poisot, K. H. Woo, Naupaka B. Zimmerman, J. W. Hollister

Biology Faculty Publications

No abstract provided.


System Genetic Analysis Of Mechanisms Underlying Excessive Alcohol Consumption, Maren L. Smith 2016 Virginia Commonwealth University

System Genetic Analysis Of Mechanisms Underlying Excessive Alcohol Consumption, Maren L. Smith

Theses and Dissertations

Increased alcohol consumption over time is one of the characteristic symptoms of Alcohol Use Disorder (AUD). The molecular mechanisms underlying this escalation in intake is still the subject of study. However, the mesocortical and mesolimbic dopamine pathways, and the extended amygdala, because of their involvement in reward and reinforcement are believed to play key roles in these behavioral changes. Multiple gene expression studies have shown that alcohol affects the expression of thousands of genes in the brain. The studies discussed in this document use the systems biology technique of co-expression network analysis to attempt to find

patterns within genome-wide expression …


A Pipeline For Creation Of Genome-Scale Metabolic Reconstructions, Shaun W. Norris 2016 Virginia Commonwealth University

A Pipeline For Creation Of Genome-Scale Metabolic Reconstructions, Shaun W. Norris

Theses and Dissertations

The decreasing costs of next generation sequencing technologies and the increasing speeds at which they work have lead to an abundance of 'omic datasets. The need for tools and methods to analyze, annotate, and model these datasets to better understand biological systems is growing. Here we present a novel software pipeline to reconstruct the metabolic model of an organism in silico starting from its genome sequence and a novel compilation of biological databases to better serve the generation of metabolic models. We validate these methods using five Gardnerella vaginalis strains and compare the gene annotation results to NCBI and the …


Characterization Of Somatically-Eliminated Genes During Development Of The Sea Lamprey (Petromyzon Marinus), Stephanie A. Bryant 2016 University of Kentucky

Characterization Of Somatically-Eliminated Genes During Development Of The Sea Lamprey (Petromyzon Marinus), Stephanie A. Bryant

Theses and Dissertations--Biology

The sea lamprey (Petromyzon marinus) undergoes programmed genome rearrangements (PGRs) during early development that facilitate the elimination of ~20% of the genome from the somatic cell lineage, resulting in distinct somatic and germline genomes. To improve our understanding of the evolutionary/developmental logic of PGR, we generated computational predictions to identify candidate germline-specific genes within a transcriptomic dataset derived from adult germline and the embryonic stages encompassing PGR. Validation studies identified 44 germline-specific genes and characterized patterns of transcription and DNA loss during early embryogenesis. Expression analyses reveal that several of these genes are differentially expressed during early embryogenesis …


Resolving Gnetum Evolutionary History, Angela McFadden 2016 Central Washington University

Resolving Gnetum Evolutionary History, Angela Mcfadden

All Master's Theses

Gnetum are non-flowering seed plants of the tropics, indigenous to South America, Africa, and Asia. This group of about 40 species is fascinating to botanists because it shares distinctive morphological characteristics with flowering plants, such as broad leaves, woody stems, and flower-like strobili. There are still questions surrounding the relationships within the genus of Gnetum. With that in mind, I focused my work on generating phylogenetic hypotheses, using two molecular data sets: a concatenation of over 60 different chloroplast genes (66,815 base pairs), and the whole chloroplast genome (128,772 base pairs). This allowed me to compare the two phylogenies …


Deep Models For Brain Em Image Segmentation: Novel Insights And Improved Performance, Ahmed Fakhry, Hanchuan Peng, Shuiwang Ji 2016 Old Dominion University

Deep Models For Brain Em Image Segmentation: Novel Insights And Improved Performance, Ahmed Fakhry, Hanchuan Peng, Shuiwang Ji

Computer Science Faculty Publications

Motivation: Accurate segmentation of brain electron microscopy (EM) images is a critical step in dense circuit reconstruction. Although deep neural networks (DNNs) have been widely used in a number of applications in computer vision, most of these models that proved to be effective on image classification tasks cannot be applied directly to EM image segmentation, due to the different objectives of these tasks. As a result, it is desirable to develop an optimized architecture that uses the full power of DNNs and tailored specifically for EM image segmentation.

Results: In this work, we proposed a novel design of DNNs for …


Genomic Prediction Of Gene Bank Wheat Landraces, José Crossa, Diego Jarquin, Jorge Franco, Paulino Pérez-Rodríguez, Juan Burgueño, Carolina Saint-Pierre, Prashant Vikram, Carolina Sansaloni, Cesar Petroli, Denis Akdemir, Clay Sneller, Matthew Reynolds, Maria Tattaris, Thomas Payne, Carlos Guzman, Roberto J. Peña, Peter Wenzl, Sukhwinder Singh 2016 International Maize and Wheat improvement Center (CIMMYT)

Genomic Prediction Of Gene Bank Wheat Landraces, José Crossa, Diego Jarquin, Jorge Franco, Paulino Pérez-Rodríguez, Juan Burgueño, Carolina Saint-Pierre, Prashant Vikram, Carolina Sansaloni, Cesar Petroli, Denis Akdemir, Clay Sneller, Matthew Reynolds, Maria Tattaris, Thomas Payne, Carlos Guzman, Roberto J. Peña, Peter Wenzl, Sukhwinder Singh

Department of Agronomy and Horticulture: Faculty Publications

This study examines genomic prediction within 8416 Mexican landrace accessions and 2403 Iranian landrace accessions stored in gene banks. The Mexican and Iranian collections were evaluated in separate field trials, including an optimum environment for several traits, and in two separate environments (drought, D and heat, H) for the highly heritable traits, days to heading (DTH), and days to maturity (DTM). Analyses accounting and not accounting for population structure were performed. Genomic prediction models include genotype × environment interaction (G × E). Two alternative prediction strategies were studied: (1) random cross-validation of the data in 20% training (TRN) and 80% …


Finding Function In The Unknown, Kelly Boyd, Emma Highland, Amanda Misch, Amber Hu, Sushma Reddy, Catherine Putonti 2015 Loyola University Chicago

Finding Function In The Unknown, Kelly Boyd, Emma Highland, Amanda Misch, Amber Hu, Sushma Reddy, Catherine Putonti

Bioinformatics Faculty Publications

Through high-throughput RNA sequencing (RNAseq), transcriptomes for a single cell, tissue, or organism(s) can be ascertained at a high resolution. While a number of bioinformatic tools have been developed for transcriptome analyses, significant challenges exist for studies of non-model organisms. Without a reference sequence available, raw reads must first be assembled de novo followed by the tedious task of BLAST searches and data mining for functional information. We have created a pipeline, PyRanger, to automate this process. The pipeline includes functionality to assess a single transcriptome and also facilitate comparative transcriptomic studies.


Leveraging Global Gene Expression Patterns To Predict Expression Of Unmeasured Genes, James Rudd, René A. Zelaya, Eugene Demidenko, Ellen L. Goode, Casey S. Greene S. Greene, Jennifer A. Doherty 2015 Dartmouth College

Leveraging Global Gene Expression Patterns To Predict Expression Of Unmeasured Genes, James Rudd, René A. Zelaya, Eugene Demidenko, Ellen L. Goode, Casey S. Greene S. Greene, Jennifer A. Doherty

Dartmouth Scholarship

BackgroundLarge collections of paraffin-embedded tissue represent a rich resource to test hypotheses based on gene expression patterns; however, measurement of genome-wide expression is cost-prohibitive on a large scale. Using the known expression correlation structure within a given disease type (in this case, high grade serous ovarian cancer; HGSC), we sought to identify reduced sets of directly measured (DM) genes which could accurately predict the expression of a maximized number of unmeasured genes.


Identifying Gene-Gene Interactions That Are Highly Associated With Body Mass Index Using Quantitative Multifactor Dimensionality Reduction (Qmdr), Rishika De, Shefali S. Verma, Fotios Drenos, Emily R. Holzinger 2015 Dartmouth College

Identifying Gene-Gene Interactions That Are Highly Associated With Body Mass Index Using Quantitative Multifactor Dimensionality Reduction (Qmdr), Rishika De, Shefali S. Verma, Fotios Drenos, Emily R. Holzinger

Dartmouth Scholarship

Despite heritability estimates of 40–70% for obesity, less than 2% of its variation is explained by Body Mass Index (BMI) associated loci that have been identified so far. Epistasis, or gene-gene interactions are a plausible source to explain portions of the missing heritability of BMI. Using genotypic data from 18,686 individuals across five study cohorts – ARIC, CARDIA, FHS, CHS, MESA – we filtered SNPs (Single Nucleotide Polymorphisms) using two parallel approaches. SNPs were filtered either on the strength of their main effects of association with BMI, or on the number of knowledge sources supporting a specific SNP-SNP interaction in …


The Importance Of Physicochemical Characteristics And Nonlinear Classifiers In Determining Hiv-1 Protease Specificity, Timmy Manning, Paul Walsh 2015 Department of Computer Science, Cork Institute of Technology, Cork, Ireland

The Importance Of Physicochemical Characteristics And Nonlinear Classifiers In Determining Hiv-1 Protease Specificity, Timmy Manning, Paul Walsh

Department of Biological Sciences Publications

This paper reviews recent research relating to the application of bioinformatics approaches to determining HIV-1 protease specificity, outlines outstanding issues, and presents a new approach to addressing these issues. Leading machine learning theory for the problem currently suggests that the direct encoding of the physicochemical properties of the amino acid substrates is not required for optimal performance. A number of amino acid encoding approaches which incorporate potentially relevant physicochemical properties of the substrate are identified, and are evaluated using a nonlinear task decomposition based neuroevolution algorithm. The results are evaluated, and compared against a recent benchmark set on a nonlinear …


A Survey Of The Common Loon (Gavia Immer) Genome Reveals Patterns Of Natural Selection, Zach G. Gayk 2015 Northern Michigan University

A Survey Of The Common Loon (Gavia Immer) Genome Reveals Patterns Of Natural Selection, Zach G. Gayk

All NMU Master's Theses

With rapid advances in Next-Generation Sequencing technology, comparative genomics has become a viable method for studying the adaptation of species to their environment at the genome level. I investigated this in common loons (Gavia immer)—for which molecular adaptation has not been characterized—by finding signatures of positive selection as evidence for genomic adaptation.

I used Illumina short read sequencing data from a single female common loon to produce a fragmented assembly of the common loon (Gavia immer) genome. The resulting assembly had a contig N50 of 814 bp, a total length of 767,326,331 bp, and 45.7 % …


A Polyglot Approach To Bioinformatics Data Integration: A Phylogenetic Analysis Of Hiv-1, Steven Reisman, Thomas Hatzopoulos, Konstantin Laufer, George K. Thiruvathukal, Catherine Putonti 2015 Loyola University Chicago

A Polyglot Approach To Bioinformatics Data Integration: A Phylogenetic Analysis Of Hiv-1, Steven Reisman, Thomas Hatzopoulos, Konstantin Laufer, George K. Thiruvathukal, Catherine Putonti

Bioinformatics Faculty Publications

As sequencing technologies continue to drop in price and increase in throughput, new challenges emerge for the management and accessibility of genomic sequence data. We have developed a pipeline for facilitating the storage, retrieval, and subsequent analysis of molecular data, integrating both sequence and metadata. Taking a polyglot approach involving multiple languages, libraries, and persistence mechanisms, sequence data can be aggregated from publicly available and local repositories. Data are exposed in the form of a RESTful web service, formatted for easy querying, and retrieved for downstream analyses. As a proof of concept, we have developed a resource for annotated HIV-1 …


Bacteriophages Isolated From Lake Michigan Demonstrate Broad Host-Range Across Several Bacterial Phyla, Kema Malki, Alex Kula, Katherine Bruder, Emily Sible, Thomas Hatzopoulos, Stephanie Steidel, Siobhan C. Watkins, Catherine Putonti 2015 Loyola University Chicago

Bacteriophages Isolated From Lake Michigan Demonstrate Broad Host-Range Across Several Bacterial Phyla, Kema Malki, Alex Kula, Katherine Bruder, Emily Sible, Thomas Hatzopoulos, Stephanie Steidel, Siobhan C. Watkins, Catherine Putonti

Biology: Faculty Publications and Other Works

BACKGROUND:

The study of bacteriophages continues to generate key information about microbial interactions in the environment. Many phenotypic characteristics of bacteriophages cannot be examined by sequencing alone, further highlighting the necessity for isolation and examination of phages from environmental samples. While much of our current knowledge base has been generated by the study of marine phages, freshwater viruses are understudied in comparison. Our group has previously conducted metagenomics-based studies samples collected from Lake Michigan - the data presented in this study relate to four phages that were extracted from the same samples.

FINDINGS:

Four phages were extracted from Lake Michigan …


Genome-Wide Detection And Analysis Of Multifunctional Genes, Yuri Pritykin, Dario Ghersi, Mona Singh 2015 Princeton University

Genome-Wide Detection And Analysis Of Multifunctional Genes, Yuri Pritykin, Dario Ghersi, Mona Singh

Interdisciplinary Informatics Faculty Publications

Many genes can play a role in multiple biological processes or molecular functions. Identifying multifunctional genes at the genome-wide level and studying their properties can shed light upon the complexity of molecular events that underpin cellular functioning, thereby leading to a better understanding of the functional landscape of the cell. However, to date, genome-wide analysis of multifunctional genes (and the proteins they encode) has been limited. Here we introduce a computational approach that uses known functional annotations to extract genes playing a role in at least two distinct biological processes. We leverage functional genomics data sets for three organisms—H. sapiens, …


A Computational Model Of The Spread Of Ancient Human Populations Based On Mitochondrial Dna Samples, Peter Revesz 2015 University of Nebraska-Lincoln

A Computational Model Of The Spread Of Ancient Human Populations Based On Mitochondrial Dna Samples, Peter Revesz

School of Computing: Conference and Workshop Papers

The extraction of mitochondrial DNA (mtDNA) from ancient human population samples provides important data for the reconstruction of population influences, spread and evolution from the Neolithic to the present. This paper presents a mtDNA-based similarity measure between pairs of human populations and a computational model for the evolution of human populations. In a computational experiment, the paper studies the mtDNA information from five Neolithic and Bronze Age populations, namely the Andronovo, the Bell Beaker, the Minoan, the Rössen and the Únětice populations. In the past these populations were identified as separate cultural groups based on geographic location, age and the …


Mutations Of Adjacent Amino Acid Pairs Are Not Always Independent, Jyotsna Ramanan, Peter Revesz 2015 University of Nebraska-Lincoln

Mutations Of Adjacent Amino Acid Pairs Are Not Always Independent, Jyotsna Ramanan, Peter Revesz

School of Computing: Conference and Workshop Papers

Evolutionary studies usually assume that the genetic mutations are independent of each other. This paper tests the independence hypothesis for genetic mutations with regard to protein coding regions. According to the new experimental results the independence assumption generally holds, but there are certain exceptions. In particular, the coding regions that represent two adjacent amino acids seem to change in ways that sometimes deviate significantly from the expected theoretical probability under the independence assumption.


Digital Commons powered by bepress