Genome Sequence Of The Model Mushroom Schizophyllum Commune,
2010
Sacred Heart University
Genome Sequence Of The Model Mushroom Schizophyllum Commune, Robin A. Ohm, Jan F. De Jong, Luis G. Lugones, Andrea Aerts, Erika Kothe, Jason E. Stajich, Ronald P. De Vries, Eric Record, Anthony Levasseur, Scott E. Baker, Kirk A. Bartholomew, Pedro M. Coutinho, Susann Erdmann, Thomas J. Fowler, Allen C. Gathmen, Vincent Lombard, Bernard Henrissat, Nicole Knabe, Ursula Kues, Walt W. Lily
Biology Faculty Publications
Much remains to be learned about the biology of mushroom-forming fungi, which are an important source of food, secondary metabolites and industrial enzymes. The wood-degrading fungus Schizophyllum commune is both a genetically tractable model for studying mushroom development and a likely source of enzymes capable of efficient degradation of lignocellulosic biomass. Comparative analyses of its 38.5-megabase genome, which encodes 13,210 predicted genes, reveal the species's unique wood-degrading machinery. One-third of the 471 genes predicted to encode transcription factors are differentially expressed during sexual development of S. commune. Whereas inactivation of one of these, fst4, prevented mushroom formation, inactivation of another, …
G-Lattices For An Unrooted Perfect Phylogeny,
2010
Rose-Hulman Institute of Technology
G-Lattices For An Unrooted Perfect Phylogeny, Monica Grigg
Mathematical Sciences Technical Reports (MSTR)
We look at the Pure Parsimony problem and the Perfect Phylogeny Haplotyping problem. From the Pure Parsimony problem we consider structures of genotypes called g-lattices. These structures either provide solutions or give bounds to the pure parsimony problem. In particular, we investigate which of these structures supports an unrooted perfect phylogeny, a condition that adds biological interpretation. By understanding which g-lattices support an unrooted perfect phylogeny, we connect two of the standard biological inference rules used to recreate how genetic diversity propagates across generations.
A Perturbation Method For Inference On Regularized Regression Estimates,
2010
Harvard University
A Perturbation Method For Inference On Regularized Regression Estimates, Jessica Minnier, Lu Tian, Tianxi Cai
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Decision-Theory Approach To Interpretable Set Analysis For High-Dimensional Data,
2010
Johns Hopkins University Bloomberg School of Public Health
A Decision-Theory Approach To Interpretable Set Analysis For High-Dimensional Data, Simina Maria Boca, Hector C. Bravo, Brian Caffo, Jeffrey T. Leek, Giovanni Parmigiani
Johns Hopkins University, Dept. of Biostatistics Working Papers
A ubiquitous problem in igh-dimensional analysis is the identification of pre-defined sets that are enriched for features showing an association of interest. In this situation, inference is performed on sets, not individual features. We propose an approach which focuses on estimating the fraction of non-null features in a set. We search for unions of disjoint sets (atoms), using as the loss function a weighted average of the number of false and missed discoveries. We prove that the solution is equivalent to thresholding the atomic false discovery rate and that our approach results in a more interpretable set analysis.
Improved Ibd Detection Using Incomplete Haplotype Information,
2010
Dartmouth College
Improved Ibd Detection Using Incomplete Haplotype Information, Giulio Genovese, Gregory Leibon, Martin R. Pollak, Daniel N. Rockmore
Dartmouth Scholarship
The availability of high density genetic maps and genotyping platforms has transformed human genetic studies. The use of these platforms has enabled population-based genome-wide association studies. However, in inheritance-based studies, current methods do not take full advantage of the information present in such genotyping analyses. In this paper we describe an improved method for identifying genetic regions shared identical-by-descent (IBD) from recent common ancestors. This method improves existing methods by taking advantage of phase information even if it is less than fully accurate or missing. We present an analysis of how using phase information increases the accuracy of IBD detection …
The Strength Of Statistical Evidence For Composite Hypotheses: Inference To The Best Explanation,
2010
Ottawa Institute of Systems Biology, Department of Biochemistry, Microbiology, and Immunology, Department of Mathematics and Statistics
The Strength Of Statistical Evidence For Composite Hypotheses: Inference To The Best Explanation, David R. Bickel
COBRA Preprint Series
A general function to quantify the weight of evidence in a sample of data for one hypothesis over another is derived from the law of likelihood and from a statistical formalization of inference to the best explanation. For a fixed parameter of interest, the resulting weight of evidence that favors one composite hypothesis over another is the likelihood ratio using the parameter value consistent with each hypothesis that maximizes the likelihood function over the parameter of interest. Since the weight of evidence is generally only known up to a nuisance parameter, it is approximated by replacing the likelihood function with …
Constraint-Based Model Of Shewanella Oneidensis Mr-1 Metabolism: A Tool For Data Analysis And Hypothesis Generation,
2010
Pacific Northwest National Laboratory
Constraint-Based Model Of Shewanella Oneidensis Mr-1 Metabolism: A Tool For Data Analysis And Hypothesis Generation, Grigoriy E. Pinchuk, Eric A. Hill, Oleg V. Geydebrekht, Jessica De Ingeniis, Xiaolin Zhang, Andrei Osterman, James H. Scott
Dartmouth Scholarship
Shewanellae are gram-negative facultatively anaerobic metal-reducing bacteria commonly found in chemically (i.e., redox) stratified environments. Occupying such niches requires the ability to rapidly acclimate to changes in electron donor/acceptor type and availability; hence, the ability to compete and thrive in such environments must ultimately be reflected in the organization and utilization of electron transfer networks, as well as central and peripheral carbon metabolism. To understand how Shewanella oneidensis MR-1 utilizes its resources, the metabolic network was reconstructed. The resulting network consists of 774 reactions, 783 genes, and 634 unique metabolites and contains biosynthesis pathways for all cell constituents. Using constraint-based …
Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies,
2010
The University of North Carolina at Chapel Hill
Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Error Correcting Codes And The Human Genome.,
2010
East Tennessee State University
Error Correcting Codes And The Human Genome., Suzanne Mclean Lyle
Electronic Theses and Dissertations
In this work, we study error correcting codes and generalize the concepts with a view toward a novel application in the study of DNA sequences. The author investigates the possibility that an error correcting linear code could be included in the human genome through application and research. The author finds that while it is an accepted hypothesis that it is reasonable that some kind of error correcting code is used in DNA, no one has actually been able to identify one. The author uses the application to illustrate how the subject of coding theory can provide a teaching enrichment activity …
Optimization Algorithms For Functional Deimmunization Of Therapeutic Proteins,
2010
Dartmouth College
Optimization Algorithms For Functional Deimmunization Of Therapeutic Proteins, Andrew S. Parker, Wei Zheng, Karl E. Griswold, Chris Bailey-Kellogg
Dartmouth Scholarship
To develop protein therapeutics from exogenous sources, it is necessary to mitigate the risks of eliciting an anti-biotherapeutic immune response. A key aspect of the response is the recognition and surface display by antigen-presenting cells of epitopes, short peptide fragments derived from the foreign protein. Thus, developing minimal-epitope variants represents a powerful approach to deimmunizing protein therapeutics. Critically, mutations selected to reduce immunogenicity must not interfere with the protein's therapeutic activity.
Why Genes Evolve Faster On Secondary Chromosomes In Bacteria,
2010
University of New Hampshire
Why Genes Evolve Faster On Secondary Chromosomes In Bacteria, Vaughn S. Cooper, Samuel H. Vohr, Sarah C. Wrocklage, Philip J. Hatcher
Molecular, Cellular & Biomedical Sciences
In bacterial genomes composed of more than one chromosome, one replicon is typically larger, harbors more essential genes than the others, and is considered primary. The greater variability of secondary chromosomes among related taxa has led to the theory that they serve as an accessory genome for specific niches or conditions. By this rationale, purifying selection should be weaker on genes on secondary chromosomes because of their reduced necessity or usage. To test this hypothesis we selected bacterial genomes composed of multiple chromosomes from two genera, Burkholderia and Vibrio, and quantified the evolutionary rates (dN and dS) of all orthologs …
Permutation-Based Pathway Testing Using The Super Learner Algorithm,
2010
Division of Biostatistics, UC Berkeley
Permutation-Based Pathway Testing Using The Super Learner Algorithm, Paul Chaffee, Alan E. Hubbard, Mark L. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Many diseases and other important phenotypic outcomes are the result of a combination of factors. For example, expression levels of genes have been used as input to various statistical methods for predicting phenotypic outcomes. One particular popular variety is the so-called gene set enrichment analysis (GSEA). This paper discusses an augmentation to an existing strategy to estimate the significance of an associations between a disease outcome and a predetermined combination of biological factors, based on a specific data adaptive regression method (the "Super Learner," van der Laan et al., 2007). The procedure uses an aggressive search procedure, potentially resulting in …
Accurate Genome-Scale Percentage Dna Methylation Estimates From Microarray Data,
2010
Sidney Kimmel Comprehensive Cancer Center, Johns Hopkins University
Accurate Genome-Scale Percentage Dna Methylation Estimates From Microarray Data, Martin J. Aryee, Zhijin Wu, Christine Ladd-Acosta, Brian Herb, Andrew P. Feinberg, Srinivasan Yegnasurbramanian, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
DNA methylation is a key regulator of gene function in a multitude of both normal and abnormal biological processes, but tools to elucidate its roles on a genome-wide scale are still in their infancy. Methylation sensitive restriction enzymes and microarrays provide a potential high-throughput, low-cost platform to allow methylation profiling. However, accurate absolute methylation estimates have been elusive due to systematic errors and unwanted variability. Previous microarray pre-processing procedures, mostly developed for expression arrays, fail to adequately normalize methylation-related data since they rely on key assumptions that are violated in the case of DNA methylation. We develop a normalization strategy …
Modeling Dependent Gene Expression,
2010
UCLA School of Public Health
Modeling Dependent Gene Expression, Donatello Telesca, Peter Muller, Giovanni Parmigiani, Ralph S. Freedman
Harvard University Biostatistics Working Paper Series
No abstract provided.
Wavelet Based Functional Models For Transcriptome Analysis With Tiling Arrays,
2010
Ghent University, Belgium
Wavelet Based Functional Models For Transcriptome Analysis With Tiling Arrays, Lieven Clement, Kristof Debeuf, Ciprian Crainiceanu, Olivier Thas, Marnik Vuylsteke, Rafael Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
For a better understanding of the biology of an organism a complete description is needed of all regions of the genome that are actively transcribed. Tiling arrays can be used for this purpose. Such arrays allow the discovery of novel transcripts and the assessment of differential expression between two or more experimental conditions such as genotype, treatment, tissue, etc. Much of the initial methodological efforts were designed for transcript discovery, while more recent developments also focus on differential expression. To our knowledge no methods for tiling arrays are described in the literature that can both assess transcript discovery and identify …
An Integrative -Omics Approach To Identify Functional Sub-Networks In Human Colorectal Cancer,
2010
Case Western Reserve University
An Integrative -Omics Approach To Identify Functional Sub-Networks In Human Colorectal Cancer, Rod K. Nibbe, Mehmet Koyutürk, Mark R. Chance
Faculty Scholarship
Emerging evidence indicates that gene products implicated in human cancers often cluster together in "hot spots" in protein-protein interaction (PPI) networks. Additionally, small sub-networks within PPI networks that demonstrate synergistic differential expression with respect to tumorigenic phenotypes were recently shown to be more accurate classifiers of disease progression when compared to single targets identified by traditional approaches. However, many of these studies rely exclusively on mRNA expression data, a useful but limited measure of cellular activity. Proteomic profiling experiments provide information at the post-translational level, yet they generally screen only a limited fraction of the proteome. Here, we demonstrate that …
Bayesian Methods For Network-Structured Genomics Data,
2010
Cornell
Bayesian Methods For Network-Structured Genomics Data, Stefano Monni, Hongzhe Li
UPenn Biostatistics Working Papers
Graphs and networks are common ways of depicting information. In biology, many different processes are represented by graphs, such as regulatory networks, metabolic pathways and protein-protein interaction networks. This information provides useful supplement to the standard numerical genomic data such as microarray gene expression data. Effectively utilizing such an information can lead to a better identification of biologically relevant genomic features in the context of our prior biological knowledge. In this paper, we present a Bayesian variable selection procedure for network-structured covariates for both Gaussian linear and probit models. The key of our approach is the introduction of a Markov …
