Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- The Texas Medical Center Library (131)
- COBRA (113)
- Augustana College (42)
- University of Kentucky (40)
- University of Nebraska - Lincoln (35)
-
- Dartmouth College (26)
- Virginia Commonwealth University (22)
- University of Nebraska at Omaha (19)
- City University of New York (CUNY) (18)
- University of Connecticut (16)
- Louisiana State University (14)
- University of New Mexico (14)
- Michigan Technological University (12)
- University of Arkansas, Fayetteville (12)
- University of Louisville (12)
- Dordt University (11)
- University of Nebraska Medical Center (11)
- Kennesaw State University (10)
- Loyola University Chicago (10)
- The University of Southern Mississippi (10)
- Western University (10)
- California Polytechnic State University, San Luis Obispo (9)
- Clemson University (9)
- Munster Technological University (9)
- West Virginia University (9)
- Illinois State University (8)
- University of Missouri, St. Louis (8)
- Wayne State University (8)
- Mississippi State University (7)
- University of Montana (7)
- Keyword
-
- Bioinformatics (132)
- Humans (58)
- Genetics (55)
- Genome (49)
- Genomics (46)
-
- Annotation (38)
- Meiothermus ruber (38)
- Gene expression (26)
- GENI-ACT (25)
- Machine learning (20)
- Evolution (19)
- Transcriptome (19)
- Cancer (17)
- Animals (15)
- Algorithms (14)
- Computational biology (14)
- Genome-Wide Association Study (14)
- Phylogenetics (14)
- Population genetics (14)
- Cancer genomics (13)
- Female (13)
- Polymorphism, Single Nucleotide (13)
- Transcriptomics (13)
- Epigenetics (12)
- Biology (11)
- Male (11)
- Polymorphism (11)
- RNA-seq (11)
- Biomarkers (10)
- Computational Biology (10)
- Publication Year
- Publication
-
- Faculty, Staff and Student Publications (72)
- Dissertations and Theses (Open Access) (59)
- Meiothermus ruber Genome Analysis Project (40)
- Theses and Dissertations (31)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (27)
-
- Harvard University Biostatistics Working Paper Series (24)
- Electronic Theses and Dissertations (19)
- COBRA Preprint Series (17)
- Dartmouth Scholarship (16)
- Dissertations, Theses, and Capstone Projects (13)
- Master's Theses (13)
- Graduate Theses and Dissertations (12)
- UW Biostatistics Working Paper Series (12)
- Dissertations, Master's Theses and Master's Reports (11)
- Faculty Work Comprehensive List (11)
- Honors Scholar Theses (11)
- LSU Doctoral Dissertations (11)
- Theses & Dissertations (11)
- Bioconductor Project Working Papers (10)
- Interdisciplinary Informatics Faculty Publications (10)
- UPenn Biostatistics Working Papers (10)
- Biochemistry Publications (9)
- Dartmouth College Ph.D Dissertations (9)
- Bioinformatics Faculty Publications (8)
- Biology ETDs (8)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (8)
- Theses and Dissertations--Biology (8)
- All Dissertations (7)
- Annual Symposium on Biomathematics and Ecology Education and Research (7)
- Dissertations (7)
- Publication Type
- File Type
Articles 751 - 780 of 905
Full-Text Articles in Genetics and Genomics
Gata-Family Transcription Factors In Magnaporthe Oryzae, Cristian F. Quispe
Gata-Family Transcription Factors In Magnaporthe Oryzae, Cristian F. Quispe
Department of Agronomy and Horticulture: Dissertations, Theses, and Student Research
The filamentous fungus, Magnaporthe oryzae, responsible for blast rice disease, destroys around 10-30% of the rice crop annually. Infection begins when the specialized infection structure, the appressorium, generates enormous internal turgor pressure through the accumulation of glycerol. This turgor acts on a penetration peg emerging at the base of the cell, causing it to breach the leaf surface allowing its infection.
The enzyme trehalose-6- phosphate synthase (Tps1) is a central regulator of the transition from appressorium development to infectious hyphal growth. In the first chapter we show that initiation of rice blast disease requires a regulatory mechanism involving an …
A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi
A Unified Approach To Non-Negative Matrix Factorization And Probabilistic Latent Semantic Indexing, Karthik Devarajan, Guoli Wang, Nader Ebrahimi
COBRA Preprint Series
Non-negative matrix factorization (NMF) by the multiplicative updates algorithm is a powerful machine learning method for decomposing a high-dimensional nonnegative matrix V into two matrices, W and H, each with nonnegative entries, V ~ WH. NMF has been shown to have a unique parts-based, sparse representation of the data. The nonnegativity constraints in NMF allow only additive combinations of the data which enables it to learn parts that have distinct physical representations in reality. In the last few years, NMF has been successfully applied in a variety of areas such as natural language processing, information retrieval, image processing, speech recognition …
Evolving Hard Problems: Generating Human Genetics Datasets With A Complex Etiology, Daniel S Himmelstein, Casey S Greene, Jason H Moore
Evolving Hard Problems: Generating Human Genetics Datasets With A Complex Etiology, Daniel S Himmelstein, Casey S Greene, Jason H Moore
Dartmouth Scholarship
BackgroundA goal of human genetics is to discover genetic factors that influence individuals' susceptibility to common diseases. Most common diseases are thought to result from the joint failure of two or more interacting components instead of single component failures. This greatly complicates both the task of selecting informative genetic variants and the task of modeling interactions between them. We and others have previously developed algorithms to detect and model the relationships between these genetic factors and disease. Previously these methods have been evaluated with datasets simulated according to pre-defined genetic models.
Beyond Structural Genomics: Computational Approaches For The Identification Of Ligand Binding Sites In Protein Structures, Dario Ghersi, Roberto Sanchez
Beyond Structural Genomics: Computational Approaches For The Identification Of Ligand Binding Sites In Protein Structures, Dario Ghersi, Roberto Sanchez
Interdisciplinary Informatics Faculty Publications
t Structural genomics projects have revealed structures for a large number of proteins of unknown function. Understanding the interactions between these proteins and their ligands would provide an initial step in their functional characterization. Binding site identification methods are a fast and cost-effective way to facilitate the characterization of functionally important protein regions. In this review we describe our recently developed methods for binding site identification in the context of existing methods. The advantage of energy-based approaches is emphasized, since they provide flexibility in the identifi- cation and characterization of different types of binding sites
A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg
A Bayesian Model Averaging Approach For Observational Gene Expression Studies, Xi Kathy Zhou, Fei Liu, Andrew J. Dannenberg
COBRA Preprint Series
Identifying differentially expressed (DE) genes associated with a sample characteristic is the primary objective of many microarray studies. As more and more studies are carried out with observational rather than well controlled experimental samples, it becomes important to evaluate and properly control the impact of sample heterogeneity on DE gene finding. Typical methods for identifying DE genes require ranking all the genes according to a pre-selected statistic based on a single model for two or more group comparisons, with or without adjustment for other covariates. Such single model approaches unavoidably result in model misspecification, which can lead to increased error …
Component Extraction Of Complex Biomedical Signal And Performance Analysis Based On Different Algorithm, Hemant Pasusangai Kasturiwale
Component Extraction Of Complex Biomedical Signal And Performance Analysis Based On Different Algorithm, Hemant Pasusangai Kasturiwale
Johns Hopkins University, Dept. of Biostatistics Working Papers
Biomedical signals can arise from one or many sources including heart ,brains and endocrine systems. Multiple sources poses challenge to researchers which may have contaminated with artifacts and noise. The Biomedical time series signal are like electroencephalogram(EEG),electrocardiogram(ECG),etc The morphology of the cardiac signal is very important in most of diagnostics based on the ECG. The diagnosis of patient is based on visual observation of recorded ECG,EEG,etc, may not be accurate. To achieve better understanding , PCA (Principal Component Analysis) and ICA algorithms helps in analyzing ECG signals . The immense scope in the field of biomedical-signal processing Independent Component Analysis( …
Removing Technical Variability In Rna-Seq Data Using Conditional Quantile Normalization, Kasper D. Hansen, Rafael A. Irizarry, Zhijin Wu
Removing Technical Variability In Rna-Seq Data Using Conditional Quantile Normalization, Kasper D. Hansen, Rafael A. Irizarry, Zhijin Wu
Johns Hopkins University, Dept. of Biostatistics Working Papers
The ability to measure gene expression on a genome-wide scale is one of the most promising accomplishments in molecular biology. Microarrays, the technology that first permitted this, were riddled with problems due to unwanted sources of variability. Many of these problems are now mitigated, after a decade’s worth of statistical methodology development. The recently developed RNA sequencing (RNA-seq) technology has generated much excitement in part due to claims of reduced variability in comparison to microarrays. However, we show RNA-seq data demonstrates unwanted and obscuring variability similar to what was first observed in microarrays. In particular, we find GC-content has a …
Screening Synteny Blocks In Pairwise Genome Comparisons Through Integer Programming, Haibao Tang, Eric Lyons, Brent S. Pedersen, James C. Schnable, Andrew H. Paterson, Michael Freeling
Screening Synteny Blocks In Pairwise Genome Comparisons Through Integer Programming, Haibao Tang, Eric Lyons, Brent S. Pedersen, James C. Schnable, Andrew H. Paterson, Michael Freeling
Department of Agronomy and Horticulture: Faculty Publications
Background:
It is difficult to accurately interpret chromosomal correspondences such as true orthology and paralogy due to significant divergence of genomes from a common ancestor. Analyses are particularly problematic among lineages that have repeatedly experienced whole genome duplication (WGD) events. To compare multiple “subgenomes” derived from genome duplications, we need to relax the traditional requirements of “one-to-one” syntenic matchings of genomic regions in order to reflect “one-to-many” or more generally “many-to-many” matchings. However this relaxation may result in the identification of synteny blocks that are derived from ancient shared WGDs that are not of interest. For many downstream analyses, we …
Systematic Assessment Of Accuracy Of Comparative Model Of Proteins Belonging To Different Structural Fold Classes, Subrata Chakrabarty, Dario Ghersi, Roberto Sanchez
Systematic Assessment Of Accuracy Of Comparative Model Of Proteins Belonging To Different Structural Fold Classes, Subrata Chakrabarty, Dario Ghersi, Roberto Sanchez
Interdisciplinary Informatics Faculty Publications
In the absence of experimental structures, comparative modeling continues to be the chosen method for retrieving structural information on target proteins. However, models lack the accuracy of experimental structures. Alignment error and structural divergence (between target and template) influence model accuracy the most. Here, we examine the potential additional impact of backbone geometry, as our previous studies have suggested that the structural class (all-α, αβ, all-β) of a protein may influence the accuracy of its model. In the twilight zone (sequence identity ≤ 30%) and at a similar level of target-template divergence, the accuracy of protein models does indeed follow …
Computer Simulations Of Heterologous Immunity: Highlights Of An Interdisciplinary Cooperation, Claudia Calcagno, Roberto Puzone, Yanthe E. Pearson, Yiming Cheng, Dario Ghersi, Liisa K. Selin, Raymond M. Welsh, Franco Celada
Computer Simulations Of Heterologous Immunity: Highlights Of An Interdisciplinary Cooperation, Claudia Calcagno, Roberto Puzone, Yanthe E. Pearson, Yiming Cheng, Dario Ghersi, Liisa K. Selin, Raymond M. Welsh, Franco Celada
Interdisciplinary Informatics Faculty Publications
The relationship between biological research and mathematical modeling is complex, critical, and vital. In this review, we summarize the results of the collaboration between two laboratories, exploring the interaction between mathematical modeling and wet-lab immunology. During this collaboration several aspects of the immune defence against viral infections were investigated, focusing primarily on the subject of heterologous immunity. In this manuscript, we emphasize the topics where computational simulations were applied in conjunction with experiments, such as immune attrition, the growing and shrinking of cross-reactive T cell repertoires following repeated infections, the short and long-term effects of cross-reactive immunological memory, and the …
Statistical Properties Of The Integrative Correlation Coefficient: A Measure Of Cross-Study Gene Reproducibility, Leslie Cope, Giovanni Parmigiani
Statistical Properties Of The Integrative Correlation Coefficient: A Measure Of Cross-Study Gene Reproducibility, Leslie Cope, Giovanni Parmigiani
Harvard University Biostatistics Working Paper Series
No abstract provided.
Linear Methods For Analysis And Quality Control Of Relative Expression Ratios From Quantitative Real-Time Polymerase Chain Reaction Experiments, Robert B. Page, Arnold J. Stromberg
Linear Methods For Analysis And Quality Control Of Relative Expression Ratios From Quantitative Real-Time Polymerase Chain Reaction Experiments, Robert B. Page, Arnold J. Stromberg
Biology Faculty Publications
Relative expression quantitative real-time polymerase chain reaction (RT-qPCR) experiments are a common means of estimating transcript abundances across biological groups and experimental treatments. One of the most frequently used expression measures that results from such experiments is the relative expression ratio (RE), which describes expression in experimental samples (i.e., RNA isolated from organisms, tissues, and/or cells that were exposed to one or more experimental or nonbaseline condition) in terms of fold change relative to calibrator samples (i.e., RNA isolated from organisms, tissues, and/or cells that were exposed to a control or baseline condition). Over the past decade, several …
Inflated Type I Error Rates When Using Aggregation Methods To Analyze Rare Variants In The 1000 Genomes Project Exon Sequencing Data In Unrelated Individuals: Summary Results From Group 7 At Genetic Analysis Workshop 17, Nathan L. Tintle, Hugues Aschard, Inchi Hu, Nora Nock, Haitian Wang, Elizabeth Pugh
Inflated Type I Error Rates When Using Aggregation Methods To Analyze Rare Variants In The 1000 Genomes Project Exon Sequencing Data In Unrelated Individuals: Summary Results From Group 7 At Genetic Analysis Workshop 17, Nathan L. Tintle, Hugues Aschard, Inchi Hu, Nora Nock, Haitian Wang, Elizabeth Pugh
Faculty Work Comprehensive List
As part of Genetic Analysis Workshop 17 (GAW17), our group considered the application of novel and standard approaches to the analysis of genotype-phenotype association in next-generation sequencing data. Our group identified a major issue in the analysis of the GAW17 next-generation sequencing data: type I error and false-positive report probability rates higher than those expected based on empirical type I error levels (as high as 90%). Two main causes emerged: population stratification and long-range correlation (gametic phase disequilibrium) between rare variants. Population stratification was expected because of the diverse sample. Correlation between rare variants was attributable to both random causes …
Identification Of Genetic Association Of Multiple Rare Variants Using Collapsing Methods, Yan V. Sun, Yun Ju Sung, Nathan L. Tintle, Andreas Ziegler
Identification Of Genetic Association Of Multiple Rare Variants Using Collapsing Methods, Yan V. Sun, Yun Ju Sung, Nathan L. Tintle, Andreas Ziegler
Faculty Work Comprehensive List
Next-generation sequencing technology allows investigation of both common and rare variants in humans. Exomes are sequenced on the population level or in families to further study the genetics of human diseases. Genetic Analysis Workshop 17 (GAW17) provided exomic data from the 1000 Genomes Project and simulated phenotypes. These data enabled evaluations of existing and newly developed statistical methods for rare variant sequence analysis for which standard statistical methods fail because of the rareness of the alleles. Various alternative approaches have been proposed that overcome the rareness problem by combining multiple rare variants within a gene. These approaches are termed collapsing …
A Parallel Graph Sampling Algorithm For Analyzing Gene Correlation Networks, Kathryn Dempsey Cooper, Kanimathi Duraisamy, Hesham Ali, Sanjukta Bhowmick
A Parallel Graph Sampling Algorithm For Analyzing Gene Correlation Networks, Kathryn Dempsey Cooper, Kanimathi Duraisamy, Hesham Ali, Sanjukta Bhowmick
Interdisciplinary Informatics Faculty Publications
Effcient analysis of complex networks is often a challenging task due to its large size and the noise inherent in the system. One popular method of overcoming this problem is through graph sampling, that is extracting a representative subgraph from the larger network. The accuracy of the sample is validated by comparing the combinatorial properties of the subgraph and the original network. However, there has been little study in comparing networks based on the applications that they represent. Furthermore, sampling methods are generally applied agnostically, without mapping to the requirements of the underlying analysis. In this paper,we introduce a parallel …
Identifying Modular Function Via Edge Annotation In Gene Correlation Networks Using Gene Ontology Search, Kathryn Dempsey Cooper, Ishwor Thapa, Dhundy Raj Bastola, Hesham Ali
Identifying Modular Function Via Edge Annotation In Gene Correlation Networks Using Gene Ontology Search, Kathryn Dempsey Cooper, Ishwor Thapa, Dhundy Raj Bastola, Hesham Ali
Interdisciplinary Informatics Faculty Proceedings & Presentations
Correlation networks provide a powerful tool for analyzing large sets of biological information. This method of high-throughput data modeling has important implications in uncovering novel knowledge of cellular function. Previous studies on other types of network modeling (protein-protein interaction networks, metabolomes, etc.) have demonstrated the presence of relationships between network structures and organization of cellular function. Studies with correlation network further confirm the existence of such network structure and biological function relationship. However, correlation networks are typically noisy and the identified network structures, such as clusters, must be further investigated to verify actual cellular function. This is traditionally done using …
A Novel Correlation Networks Approach For The Identification Of Gene Targets, Kathryn Dempsey Cooper, Stephen Bonasera, Dhundy Raj Bastola, Hesham Ali
A Novel Correlation Networks Approach For The Identification Of Gene Targets, Kathryn Dempsey Cooper, Stephen Bonasera, Dhundy Raj Bastola, Hesham Ali
Interdisciplinary Informatics Faculty Proceedings & Presentations
Correlation networks are emerging as a powerful tool for modeling temporal mechanisms within the cell. Particularly useful in examining coexpression within microarray data, studies have determined that correlation networks follow a power law degree distribution and thus manifest properties such as the existence of “hub” nodes and semicliques that potentially correspond to critical cellular structures. Difficulty lies in filtering coincidental relationships from causative structures in these large, noise-heavy networks. As such, computational expenses and algorithm availability limit accurate comparison, making it difficult to identify changes between networks. In this vein, we present our work identifying temporal relationships from microarray data …
Evaluating Methods For The Analysis Of Rare Variants In Sequence Data, Alexander Luedtke, Scott Powers, Ashley Petersen, Alexandra Sitarik, Airat Bekmetjev, Nathan L. Tintle
Evaluating Methods For The Analysis Of Rare Variants In Sequence Data, Alexander Luedtke, Scott Powers, Ashley Petersen, Alexandra Sitarik, Airat Bekmetjev, Nathan L. Tintle
Faculty Work Comprehensive List
A number of rare variant statistical methods have been proposed for analysis of the impending wave of next-generation sequencing data. To date, there are few direct comparisons of these methods on real sequence data. Furthermore, there is a strong need for practical advice on the proper analytic strategies for rare variant analysis. We compare four recently proposed rare variant methods (combined multivariate and collapsing, weighted sum, proportion regression, and cumulative minor allele test) on simulated phenotype and next-generation sequencing data as part of Genetic Analysis Workshop 17. Overall, we find that all analyzed methods have serious practical limitations on identifying …
Evaluating Methods For Combining Rare Variant Data In Pathway-Based Tests Of Genetic Association, Ashley Petersen, Alexandra Sitarik, Alexander Luedtke, Scott Powers, Airat Bekmetjev, Nathan L. Tintle
Evaluating Methods For Combining Rare Variant Data In Pathway-Based Tests Of Genetic Association, Ashley Petersen, Alexandra Sitarik, Alexander Luedtke, Scott Powers, Airat Bekmetjev, Nathan L. Tintle
Faculty Work Comprehensive List
Analyzing sets of genes in genome-wide association studies is a relatively new approach that aims to capitalize on biological knowledge about the interactions of genes in biological pathways. This approach, called pathway analysis or gene set analysis, has not yet been applied to the analysis of rare variants. Applying pathway analysis to rare variants offers two competing approaches. In the first approach rare variant statistics are used to generate p-values for each gene (e.g., combined multivariate collapsing [CMC] or weighted-sum [WS]) and the gene-level p-values are combined using standard pathway analysis methods (e.g., gene set enrichment analysis or …
Identifying Rare Variants From Exome Scans: The Gaw17 Experience, Saurabh Ghosh, Heike Bickeboller, Julia Bailey, Joan E. Bailey-Wilson, Rita Cantor, Robert Culverhouse, Warwick Daw, Anita L. Destefano, Corinne D. Engelman, Anthony Hinrichs, Jeanine Houwing-Duistermaat, Inke R. Konig, Jack Kent, Nan Laird, Nathan Pankratz, Andrew Paterson, Elizabeth Pugh, Brian Suarez, Yan Sun, Alun Thomas, Nathan L. Tintle, Xiaofeng Zhu, Andreas Ziegler, Jean W. Maccluer, Laura Almasy
Identifying Rare Variants From Exome Scans: The Gaw17 Experience, Saurabh Ghosh, Heike Bickeboller, Julia Bailey, Joan E. Bailey-Wilson, Rita Cantor, Robert Culverhouse, Warwick Daw, Anita L. Destefano, Corinne D. Engelman, Anthony Hinrichs, Jeanine Houwing-Duistermaat, Inke R. Konig, Jack Kent, Nan Laird, Nathan Pankratz, Andrew Paterson, Elizabeth Pugh, Brian Suarez, Yan Sun, Alun Thomas, Nathan L. Tintle, Xiaofeng Zhu, Andreas Ziegler, Jean W. Maccluer, Laura Almasy
Faculty Work Comprehensive List
Genetic Analysis Workshop 17 (GAW17) provided a platform for evaluating existing statistical genetic methods and for developing novel methods to analyze rare variants that modulate complex traits. In this article, we present an overview of the 1000 Genomes Project exome data and simulated phenotype data that were distributed to GAW17 participants for analyses, the different issues addressed by the participants, and the process of preparation of manuscripts resulting from the discussions during the workshop
Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel
Minimum Description Length Measures Of Evidence For Enrichment, Zhenyu Yang, David R. Bickel
COBRA Preprint Series
In order to functionally interpret differentially expressed genes or other discovered features, researchers seek to detect enrichment in the form of overrepresentation of discovered features associated with a biological process. Most enrichment methods treat the p-value as the measure of evidence using a statistical test such as the binomial test, Fisher's exact test or the hypergeometric test. However, the p-value is not interpretable as a measure of evidence apart from adjustments in light of the sample size. As a measure of evidence supporting one hypothesis over the other, the Bayes factor (BF) overcomes this drawback of the p-value but lacks …
Campylobacter Ureolyticus: An Emerging Gastrointestinal Pathogen?, Susan Bullman, Daniel Corcoran, James O'Leary, Brigid Lucey, Deirdre Byrne, Roy D. Sleator
Campylobacter Ureolyticus: An Emerging Gastrointestinal Pathogen?, Susan Bullman, Daniel Corcoran, James O'Leary, Brigid Lucey, Deirdre Byrne, Roy D. Sleator
Department of Biological Sciences Publications
A total of 7194 faecal samples collected over a 1-year period from patients presenting with diarrhoea were screened for Campylobacter spp. using EntericBios, a multiplex-PCR system. Of 349 Campylobacter-positive samples, 23.8% were shown to be Campylobacter ureolyticus, using a combination of 16S rRNA gene analysis and highly specific primers targeting the HSP60 gene of this organism. This is, to the best of our knowledge, the first report of C. ureolyticus in the faeces of patients presenting with gastroenteritis and may suggest a role for this organism as an emerging enteric pathogen.
Using The R Package Crlmm For Genotyping And Copy Number Estimation, Robert B. Scharpf, Rafael Irizarry, Walter Ritchie, Benilton Carvalho, Ingo Ruczinski
Using The R Package Crlmm For Genotyping And Copy Number Estimation, Robert B. Scharpf, Rafael Irizarry, Walter Ritchie, Benilton Carvalho, Ingo Ruczinski
Johns Hopkins University, Dept. of Biostatistics Working Papers
Genotyping platforms such as Affymetrix can be used to assess genotype-phenotype as well as copy number-phenotype associations at millions of markers. While genotyping algorithms are largely concordant when assessed on HapMap samples, tools to assess copy number changes are more variable and often discordant. One explanation for the discordance is that copy number estimates are susceptible to systematic differences between groups of samples that were processed at different times or by different labs. Analysis algorithms that do not adjust for batch effects are prone to spurious measures of association. The R package crlmm implements a multilevel model that adjusts for …
G-Lattices For An Unrooted Perfect Phylogeny, Monica Grigg
G-Lattices For An Unrooted Perfect Phylogeny, Monica Grigg
Mathematical Sciences Technical Reports (MSTR)
We look at the Pure Parsimony problem and the Perfect Phylogeny Haplotyping problem. From the Pure Parsimony problem we consider structures of genotypes called g-lattices. These structures either provide solutions or give bounds to the pure parsimony problem. In particular, we investigate which of these structures supports an unrooted perfect phylogeny, a condition that adds biological interpretation. By understanding which g-lattices support an unrooted perfect phylogeny, we connect two of the standard biological inference rules used to recreate how genetic diversity propagates across generations.
A Perturbation Method For Inference On Regularized Regression Estimates, Jessica Minnier, Lu Tian, Tianxi Cai
A Perturbation Method For Inference On Regularized Regression Estimates, Jessica Minnier, Lu Tian, Tianxi Cai
Harvard University Biostatistics Working Paper Series
No abstract provided.
A Decision-Theory Approach To Interpretable Set Analysis For High-Dimensional Data, Simina Maria Boca, Hector C. Bravo, Brian Caffo, Jeffrey T. Leek, Giovanni Parmigiani
A Decision-Theory Approach To Interpretable Set Analysis For High-Dimensional Data, Simina Maria Boca, Hector C. Bravo, Brian Caffo, Jeffrey T. Leek, Giovanni Parmigiani
Johns Hopkins University, Dept. of Biostatistics Working Papers
A ubiquitous problem in igh-dimensional analysis is the identification of pre-defined sets that are enriched for features showing an association of interest. In this situation, inference is performed on sets, not individual features. We propose an approach which focuses on estimating the fraction of non-null features in a set. We search for unions of disjoint sets (atoms), using as the loss function a weighted average of the number of false and missed discoveries. We prove that the solution is equivalent to thresholding the atomic false discovery rate and that our approach results in a more interpretable set analysis.
The Strength Of Statistical Evidence For Composite Hypotheses: Inference To The Best Explanation, David R. Bickel
The Strength Of Statistical Evidence For Composite Hypotheses: Inference To The Best Explanation, David R. Bickel
COBRA Preprint Series
A general function to quantify the weight of evidence in a sample of data for one hypothesis over another is derived from the law of likelihood and from a statistical formalization of inference to the best explanation. For a fixed parameter of interest, the resulting weight of evidence that favors one composite hypothesis over another is the likelihood ratio using the parameter value consistent with each hypothesis that maximizes the likelihood function over the parameter of interest. Since the weight of evidence is generally only known up to a nuisance parameter, it is approximated by replacing the likelihood function with …
The Human Oral Microbiome Database: A Web-Accessible Resource For Investigating Oral Microbe Taxonomic And Genomic Information, Tsute Chen, Wen-Han Yu, Jacques Izard, Oxana V. Baranova, Abirami Lakshmanan, Floyd E. Dewhirst
The Human Oral Microbiome Database: A Web-Accessible Resource For Investigating Oral Microbe Taxonomic And Genomic Information, Tsute Chen, Wen-Han Yu, Jacques Izard, Oxana V. Baranova, Abirami Lakshmanan, Floyd E. Dewhirst
Department of Food Science and Technology: Faculty Publications
The human oral microbiome is the most studied human microflora, but 53% of the species have not yet been validly named and 35% remain uncultivated. The uncultivated taxa are known primarily from 16S rRNA sequence information. Sequence information tied solely to obscure isolate or clone numbers, and usually lacking accurate phylogenetic placement, is a major impediment to working with human oral microbiome data. The goal of creating the Human Oral Microbiome Database (HOMD) is to provide the scientific community with a body site-specific comprehensive database for the more than 600 prokaryote species that are present in the human oral cavity …
Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin
Powerful Snp Set Analysis For Case-Control Genome Wide Association Studies, Michael C. Wu, Peter Kraft, Michael P. Epstein, Deanne M. Taylor, Stephen J. Chanock, David J. Hunter, Xihong Lin
Harvard University Biostatistics Working Paper Series
No abstract provided.
Optimization Algorithms For Functional Deimmunization Of Therapeutic Proteins, Andrew S. Parker, Wei Zheng, Karl E. Griswold, Chris Bailey-Kellogg
Optimization Algorithms For Functional Deimmunization Of Therapeutic Proteins, Andrew S. Parker, Wei Zheng, Karl E. Griswold, Chris Bailey-Kellogg
Dartmouth Scholarship
To develop protein therapeutics from exogenous sources, it is necessary to mitigate the risks of eliciting an anti-biotherapeutic immune response. A key aspect of the response is the recognition and surface display by antigen-presenting cells of epitopes, short peptide fragments derived from the foreign protein. Thus, developing minimal-epitope variants represents a powerful approach to deimmunizing protein therapeutics. Critically, mutations selected to reduce immunogenicity must not interfere with the protein's therapeutic activity.