Open Access. Powered by Scholars. Published by Universities.®

Computational Biology Commons

Open Access. Powered by Scholars. Published by Universities.®

2011

Discipline
Institution
Keyword
Publication
Publication Type

Articles 1 - 23 of 23

Full-Text Articles in Computational Biology

Modeling Protein Expression And Protein Signaling Pathways, Donatello Telesca, Peter Muller, Steven Kornblau, Marc Suchard, Yuan Ji Dec 2011

Modeling Protein Expression And Protein Signaling Pathways, Donatello Telesca, Peter Muller, Steven Kornblau, Marc Suchard, Yuan Ji

COBRA Preprint Series

High-throughput functional proteomic technologies provide a way to quantify the expression of proteins of interest. Statistical inference centers on identifying the activation state of proteins and their patterns of molecular interaction formalized as dependence structure. Inference on dependence structure is particularly important when proteins are selected because they are part of a common molecular pathway. In that case inference on dependence structure reveals properties of the underlying pathway. We propose a probability model that represents molecular interactions at the level of hidden binary latent variables that can be interpreted as indicators for active versus inactive states of the proteins. The …


A Study Of Correlations Between The Definition And Application Of The Gene Ontology, Yuji Mo Dec 2011

A Study Of Correlations Between The Definition And Application Of The Gene Ontology, Yuji Mo

Department of Computer Electronics and Engineering: Dissertations, Theses, and Student Research

When using the Gene Ontology (GO), nucleotide and amino acid sequences are annotated by terms in a structured and controlled vocabulary organized into relational graphs. The usage of the vocabulary (GO terms) in the annotation of these sequences may diverge from the relations defined in the ontology. We measure the consistency of the use of GO terms by comparing GO's defined structure to the terms' application. To do this, we first use synthetic data with different characteristics to understand how these characteristics influence the correlation values determined by various similarity measures. Using these results as a baseline, we found that …


Planning Combinatorial Disulfide Cross-Links For Protein Fold Determination, Fei Xiong, Alan M Friedman, Chris Bailey-Kellogg Nov 2011

Planning Combinatorial Disulfide Cross-Links For Protein Fold Determination, Fei Xiong, Alan M Friedman, Chris Bailey-Kellogg

Dartmouth Scholarship

Fold recognition techniques take advantage of the limited number of overall structural organizations, and have become increasingly effective at identifying the fold of a given target sequence. However, in the absence of sufficient sequence identity, it remains difficult for fold recognition methods to always select the correct model. While a native-like model is often among a pool of highly ranked models, it is not necessarily the highest-ranked one, and the model rankings depend sensitively on the scoring function used. Structure elucidation methods can then be employed to decide among the models based on relatively rapid biochemical/biophysical experiments.


Additive Functions In Boolean Models Of Gene Regulatory Network Modules, Christian Darabos, Ferdinando Ferdinando Di Cunto, Marco Tomassini, Jason H. Moore Nov 2011

Additive Functions In Boolean Models Of Gene Regulatory Network Modules, Christian Darabos, Ferdinando Ferdinando Di Cunto, Marco Tomassini, Jason H. Moore

Dartmouth Scholarship

Gene-on-gene regulations are key components of every living organism. Dynamical abstract models of genetic regulatory networks help explain the genome’s evolvability and robustness. These properties can be attributed to the structural topology of the graph formed by genes, as vertices, and regulatory interactions, as edges. Moreover, the actual gene interaction of each gene is believed to play a key role in the stability of the structure. With advances in biology, some effort was deployed to develop update functions in Boolean models that include recent knowledge. We combine real-life gene interaction networks with novel update functions in a Boolean model. We …


Transcriptomic Characterization Of A Synergistic Genetic Interaction During Carpel Margin Meristem Development In Arabidopsis Thaliana, April N. Wynn, Elizabeth E. Rueschhoff, Robert G. Franks Oct 2011

Transcriptomic Characterization Of A Synergistic Genetic Interaction During Carpel Margin Meristem Development In Arabidopsis Thaliana, April N. Wynn, Elizabeth E. Rueschhoff, Robert G. Franks

Biological Sciences Research

In flowering plants the gynoecium is the female reproductive structure. In Arabidopsis thalianaovules initiate within the developing gynoecium from meristematic tissue located along the margins of the floral carpels. When fertilized the ovules will develop into seeds. SEUSS (SEU) and AINTEGUMENTA (ANT) encode transcriptional regulators that are critical for the proper formation of ovules from the carpel margin meristem (CMM). The synergistic loss of ovule initiation observed in the seu ant double mutant suggests that SEU and ANT share overlapping functions during CMM development. However the molecular mechanism underlying this synergistic interaction is unknown. Using …


Gc-Content Normalization For Rna-Seq Data, Davide Risso, Katja Schwartz, Gavin Sherlock, Sandrine Dudoit Aug 2011

Gc-Content Normalization For Rna-Seq Data, Davide Risso, Katja Schwartz, Gavin Sherlock, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Background: Transcriptome sequencing (RNA-Seq) has become the assay of choice for high-throughput studies of gene expression. However, as is the case with microarrays, major technology-related artifacts and biases affect the resulting expression measures. Normalization is therefore essential to ensure accurate inference of expression levels and subsequent analyses thereof.

Results: We focus on biases related to GC-content and demonstrate the existence of strong sample-specific GC-content effects on RNA-Seq read counts, which can substantially bias differential expression analysis. We propose three simple within-lane gene-level GC-content normalization approaches and assess their performance on two different RNA-Seq datasets, involving different species and experimental designs. …


Magnetosome Genes In The Gammaproteobacterium Strain Bw-2, Lucero Rivera, Denis Trubitsyn, Dennis A. Bazylinski Aug 2011

Magnetosome Genes In The Gammaproteobacterium Strain Bw-2, Lucero Rivera, Denis Trubitsyn, Dennis A. Bazylinski

Undergraduate Research Opportunities Program (UROP)

Magnetotactic bacteria (MTB) biomineralize intracellular nanometer-sized, magnetic crystals surrounded by a lipid bilayer membrane known as magnetosomes. These crystals, which consist of magnetite (Fe3O4) or greigite (Fe3S4), causes the cell to align along the geomagnetic field lines as they swim, a phenomenon known as magnetotaxis. Strain BW-2 is a magnetite-producing magnetotactic bacterium isolated from Badwater Basin, Death Valley National Park (California) and is one of only two species of MTB that are known to phylogenetically belong to the Gammaproteobacteria class of the Proteobacteria phylum. The biomineralization of magnetite in magnetotactic bacteria is mediated by a series of genes that include …


Multiple Testing Of Local Maxima For Detection Of Peaks In Chip-Seq Data, Armin Schwartzman, Andrew Jaffe, Yulia Gavrilov, Clifford A. Meyer Aug 2011

Multiple Testing Of Local Maxima For Detection Of Peaks In Chip-Seq Data, Armin Schwartzman, Andrew Jaffe, Yulia Gavrilov, Clifford A. Meyer

Harvard University Biostatistics Working Paper Series

No abstract provided.


Physiologically-Based Pharmacokinetic Modeling For Predicting Drug-Drug Interactions, David M. Ng, Ali Navid Aug 2011

Physiologically-Based Pharmacokinetic Modeling For Predicting Drug-Drug Interactions, David M. Ng, Ali Navid

STAR Program Research Presentations

Dynamics of interactions between the drugs caffeine and ciprofloxacin are predicted using physiologically-based pharmacokinetic (PBPK) modeling. Pharmacokinetic means the model determines where the drugs are distributed in the body over time. Physiologically-based means the anatomy and physiology of the human body is reflected in the structure and functioning of the model. Multiple drugs can interact to increase or decrease their beneficial and/or undesired effects. This is important because some common substances, such as caffeine in coffee and soft drinks, are actually drugs that affect the body. By implementing the model as a computer program, it is relatively straightforward to perform …


A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi Jul 2011

A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi

COBRA Preprint Series

Non-negative matrix factorization (NMF) by the multiplicative updates algorithm is a powerful machine learning method for decomposing a high-dimensional nonnegative matrix V into two matrices, W and H, each with nonnegative entries, V ~ WH. NMF has been shown to have a unique parts-based, sparse representation of the data. The nonnegativity constraints in NMF allow only additive combinations of the data which enables it to learn parts that have distinct physical representations in reality. In the last few years, NMF has been successfully applied in a variety of areas such as natural language processing, information retrieval, image processing, speech recognition …


Evolving Hard Problems: Generating Human Genetics Datasets With A Complex Etiology, Daniel S Himmelstein, Casey S Greene, Jason H Moore Jul 2011

Evolving Hard Problems: Generating Human Genetics Datasets With A Complex Etiology, Daniel S Himmelstein, Casey S Greene, Jason H Moore

Dartmouth Scholarship

BackgroundA goal of human genetics is to discover genetic factors that influence individuals' susceptibility to common diseases. Most common diseases are thought to result from the joint failure of two or more interacting components instead of single component failures. This greatly complicates both the task of selecting informative genetic variants and the task of modeling interactions between them. We and others have previously developed algorithms to detect and model the relationships between these genetic factors and disease. Previously these methods have been evaluated with datasets simulated according to pre-defined genetic models.


Stability Analysis And Application Of A Mathematical Cholera Model, Shu Liao, Jim Wang Jul 2011

Stability Analysis And Application Of A Mathematical Cholera Model, Shu Liao, Jim Wang

Mathematics & Statistics Faculty Publications

In this paper, we conduct a dynamical analysis of the deterministic cholera model proposed in [9]. We study the stability of both the disease-free and endemic equilibria so as to explore the complex epidemic and endemic dynamics of the disease. We demonstrate a real-world application of this model by investigating the recent cholera outbreak in Zimbabwe. Meanwhile, we present numerical simulation results to verify the analytical predictions.


A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg Jun 2011

A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg

COBRA Preprint Series

Identifying differentially expressed (DE) genes associated with a sample characteristic is the primary objective of many microarray studies. As more and more studies are carried out with observational rather than well controlled experimental samples, it becomes important to evaluate and properly control the impact of sample heterogeneity on DE gene finding. Typical methods for identifying DE genes require ranking all the genes according to a pre-selected statistic based on a single model for two or more group comparisons, with or without adjustment for other covariates. Such single model approaches unavoidably result in model misspecification, which can lead to increased error …


Component Extraction Of Complex Biomedical Signal And Performance Analysis Based On Different Algorithm, Hemant Pasusangai Kasturiwale Jun 2011

Component Extraction Of Complex Biomedical Signal And Performance Analysis Based On Different Algorithm, Hemant Pasusangai Kasturiwale

Johns Hopkins University, Dept. of Biostatistics Working Papers

Biomedical signals can arise from one or many sources including heart ,brains and endocrine systems. Multiple sources poses challenge to researchers which may have contaminated with artifacts and noise. The Biomedical time series signal are like electroencephalogram(EEG),electrocardiogram(ECG),etc The morphology of the cardiac signal is very important in most of diagnostics based on the ECG. The diagnosis of patient is based on visual observation of recorded ECG,EEG,etc, may not be accurate. To achieve better understanding , PCA (Principal Component Analysis) and ICA algorithms helps in analyzing ECG signals . The immense scope in the field of biomedical-signal processing Independent Component Analysis( …


Removing Technical Variability In Rna-Seq Data Using Conditional Quantile Normalization, Kasper D. Hansen, Rafael A. Irizarry, Zhijin Wu May 2011

Removing Technical Variability In Rna-Seq Data Using Conditional Quantile Normalization, Kasper D. Hansen, Rafael A. Irizarry, Zhijin Wu

Johns Hopkins University, Dept. of Biostatistics Working Papers

The ability to measure gene expression on a genome-wide scale is one of the most promising accomplishments in molecular biology. Microarrays, the technology that first permitted this, were riddled with problems due to unwanted sources of variability. Many of these problems are now mitigated, after a decade’s worth of statistical methodology development. The recently developed RNA sequencing (RNA-seq) technology has generated much excitement in part due to claims of reduced variability in comparison to microarrays. However, we show RNA-seq data demonstrates unwanted and obscuring variability similar to what was first observed in microarrays. In particular, we find GC-content has a …


Molecular Evolution And Historical Biogeography Of New World Birds, Brian T. Smith May 2011

Molecular Evolution And Historical Biogeography Of New World Birds, Brian T. Smith

UNLV Theses, Dissertations, Professional Papers, and Capstones

Deciphering the patterns of how biodiversity has evolved across time and space has remained a fundamental objective for biologists for the last 200 years. Researchers are faced with the challenge of interpreting the complexity of evolutionary patterns that have been generated over the deep history of the Earth. The advancement of DNA sequencing technology has yielded a new and powerful genetic toolkit that has allowed biologists to address novel evolutionary questions. For my dissertation research, I used molecular genetics and a statistical framework to study the evolution and historical biogeography of birds distributed in North and South America. My dissertation …


Microarray Data Mining And Gene Regulatory Network Analysis, Ying Li May 2011

Microarray Data Mining And Gene Regulatory Network Analysis, Ying Li

Dissertations

The novel molecular biological technology, microarray, makes it feasible to obtain quantitative measurements of expression of thousands of genes present in a biological sample simultaneously. Genome-wide expression data generated from this technology are promising to uncover the implicit, previously unknown biological knowledge. In this study, several problems about microarray data mining techniques were investigated, including feature(gene) selection, classifier genes identification, generation of reference genetic interaction network for non-model organisms and gene regulatory network reconstruction using time-series gene expression data. The limitations of most of the existing computational models employed to infer gene regulatory network lie in that they either suffer …


Preliminary Analysis Of An Agent-Based Model For A Tick-Borne Disease, Holly Gaff Apr 2011

Preliminary Analysis Of An Agent-Based Model For A Tick-Borne Disease, Holly Gaff

Biological Sciences Faculty Publications

Ticks have a unique life history including a distinct set of life stages and a single blood meal per life stage. This makes tick-host interactions more complex from a mathematical perspective. In addition, any model of these interactions must involve a significant degree of stochasticity on the individual tick level. In an attempt to quantify these relationships, I have developed an individual-based model of the interactions between ticks and their hosts as well as the transmission of tick-borne disease between the two populations. The results from this model are compared with those from previously published differential equation based population models. …


Blended Biogeography-Based Optimization For Constrained Optimization, Haiping Ma, Daniel J. Simon Apr 2011

Blended Biogeography-Based Optimization For Constrained Optimization, Haiping Ma, Daniel J. Simon

Electrical and Computer Engineering Faculty Publications

Biogeography-based optimization (BBO) is a new evolutionary optimization method that is based on the science of biogeography. We propose two extensions to BBO. First, we propose a blended migration operator. Benchmark results show that blended BBO outperforms standard BBO. Second, we employ blended BBO to solve constrained optimization problems. Constraints are handled by modifying the BBO immigration and emigration procedures. The approach that we use does not require any additional tuning parameters beyond those that are required for unconstrained problems. The constrained blended BBO algorithm is compared with solutions based on a stud genetic algorithm (SGA) and standard particle swarm …


Statistical Properties Of The Integrative Correlation Coefficient: A Measure Of Cross-Study Gene Reproducibility, Leslie Cope, Giovanni Parmigiani Jan 2011

Statistical Properties Of The Integrative Correlation Coefficient: A Measure Of Cross-Study Gene Reproducibility, Leslie Cope, Giovanni Parmigiani

Harvard University Biostatistics Working Paper Series

No abstract provided.


Linear Methods For Analysis And Quality Control Of Relative Expression Ratios From Quantitative Real-Time Polymerase Chain Reaction Experiments, Robert B. Page, Arnold J. Stromberg Jan 2011

Linear Methods For Analysis And Quality Control Of Relative Expression Ratios From Quantitative Real-Time Polymerase Chain Reaction Experiments, Robert B. Page, Arnold J. Stromberg

Biology Faculty Publications

Relative expression quantitative real-time polymerase chain reaction (RT-qPCR) experiments are a common means of estimating transcript abundances across biological groups and experimental treatments. One of the most frequently used expression measures that results from such experiments is the relative expression ratio (RE), which describes expression in experimental samples (i.e., RNA isolated from organisms, tissues, and/or cells that were exposed to one or more experimental or nonbaseline condition) in terms of fold change relative to calibrator samples (i.e., RNA isolated from organisms, tissues, and/or cells that were exposed to a control or baseline condition). Over the past decade, several …


Computational Network Analysis Of The Anatomical And Genetic Organizations In The Mouse Brain, Shuiwang Ji Jan 2011

Computational Network Analysis Of The Anatomical And Genetic Organizations In The Mouse Brain, Shuiwang Ji

Computer Science Faculty Publications

Motivation: The mammalian central nervous system (CNS) generates high-level behavior and cognitive functions. Elucidating the anatomical and genetic organizations in the CNS is a key step toward understanding the functional brain circuitry. The CNS contains an enormous number of cell types, each with unique gene expression patterns. Therefore, it is of central importance to capture the spatial expression patterns in the brain. Currently, genome-wide atlas of spatial expression patterns in the mouse brain has been made available, and the data are in the form of aligned 3D data arrays. The sheer volume and complexity of these data pose significant challenges …


Weighted Scores Method For Regression Models With Dependent Data, Aristidis K. Nikoloulopoulos, Harry Joe, N. Rao Chaganty Jan 2011

Weighted Scores Method For Regression Models With Dependent Data, Aristidis K. Nikoloulopoulos, Harry Joe, N. Rao Chaganty

Mathematics & Statistics Faculty Publications

There are copula-based statistical models in the literature for regression with dependent data such as clustered and longitudinal overdispersed counts, for which parameter estimation and inference are straightforward. For situations where the main interest is in the regression and other univariate parameters and not the dependence, we propose a "weighted scores method", which is based on weighting score functions of the univariate margins. The weight matrices are obtained initially fitting a discretized multivariate normal distribution, which admits a wide range of dependence. The general methodology is applied to negative binomial regression models. Asymptotic and small-sample efficiency calculations show that our …