Open Access. Powered by Scholars. Published by Universities.®

Genetics and Genomics Commons™

Open Access. Powered by Scholars. Published by Universities.®

Bioinformatics

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 781 - 810 of 905

Full-Text Articles in Genetics and Genomics

An Entrepreneurial Approach To Librarianship, Flora Shrode, Jennifer Duncan, Wendy Holliday Apr 2010

An Entrepreneurial Approach To Librarianship, Flora Shrode, Jennifer Duncan, Wendy Holliday

Library Faculty & Staff Publications

Librarians from Utah State University explain recent efforts to encourage subject librarians to take a more holistic view of their roles. We are shifting from a traditional emphasis primarily on collection development and refocusing on natural connections between collections, instruction, liaison, and reference service. The poster provides background about Utah State University’s situation and explains our approach to analyzing local needs and culture to inform development of a new organizational structure. We describe our vision of subject librarianship, the process by which we assessed librarians’ ideas and goals for performing as subject librarians, and the actions we are taking to …


Permutation-Based Pathway Testing Using The Super Learner Algorithm, Paul Chaffee, Alan E. Hubbard, Mark L. Van Der Laan Mar 2010

Permutation-Based Pathway Testing Using The Super Learner Algorithm, Paul Chaffee, Alan E. Hubbard, Mark L. Van Der Laan

U.C. Berkeley Division of Biostatistics Working Paper Series

Many diseases and other important phenotypic outcomes are the result of a combination of factors. For example, expression levels of genes have been used as input to various statistical methods for predicting phenotypic outcomes. One particular popular variety is the so-called gene set enrichment analysis (GSEA). This paper discusses an augmentation to an existing strategy to estimate the significance of an associations between a disease outcome and a predetermined combination of biological factors, based on a specific data adaptive regression method (the "Super Learner," van der Laan et al., 2007). The procedure uses an aggressive search procedure, potentially resulting in …


Accurate Genome-Scale Percentage Dna Methylation Estimates From Microarray Data, Martin J. Aryee, Zhijin Wu, Christine Ladd-Acosta, Brian Herb, Andrew P. Feinberg, Srinivasan Yegnasurbramanian, Rafael A. Irizarry Mar 2010

Accurate Genome-Scale Percentage Dna Methylation Estimates From Microarray Data, Martin J. Aryee, Zhijin Wu, Christine Ladd-Acosta, Brian Herb, Andrew P. Feinberg, Srinivasan Yegnasurbramanian, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

DNA methylation is a key regulator of gene function in a multitude of both normal and abnormal biological processes, but tools to elucidate its roles on a genome-wide scale are still in their infancy. Methylation sensitive restriction enzymes and microarrays provide a potential high-throughput, low-cost platform to allow methylation profiling. However, accurate absolute methylation estimates have been elusive due to systematic errors and unwanted variability. Previous microarray pre-processing procedures, mostly developed for expression arrays, fail to adequately normalize methylation-related data since they rely on key assumptions that are violated in the case of DNA methylation. We develop a normalization strategy …


Reconstructability Analysis As A Tool For Identifying Gene-Gene Interactions In Studies Of Human Diseases, Stephen Shervais, Patricia L. Kramer, Shawn K. Westaway, Nancy J. Cox, Martin Zwick Mar 2010

Reconstructability Analysis As A Tool For Identifying Gene-Gene Interactions In Studies Of Human Diseases, Stephen Shervais, Patricia L. Kramer, Shawn K. Westaway, Nancy J. Cox, Martin Zwick

Complex Systems Faculty Publications and Presentations

There are a number of common human diseases for which the genetic component may include an epistatic interaction of multiple genes. Detecting these interactions with standard statistical tools is difficult because there may be an interaction effect, but minimal or no main effect. Reconstructability analysis (RA) uses Shannon’s information theory to detect relationships between variables in categorical datasets. We applied RA to simulated data for five different models of gene-gene interaction, and find that even with heritability levels as low as 0.008, and with the inclusion of 50 non-associated genes in the dataset, we can identify the interacting gene pairs …


Modeling Dependent Gene Expression, Donatello Telesca, Peter Muller, Giovanni Parmigiani, Ralph S. Freedman Feb 2010

Modeling Dependent Gene Expression, Donatello Telesca, Peter Muller, Giovanni Parmigiani, Ralph S. Freedman

Harvard University Biostatistics Working Paper Series

No abstract provided.


Wavelet Based Functional Models For Transcriptome Analysis With Tiling Arrays, Lieven Clement, Kristof Debeuf, Ciprian Crainiceanu, Olivier Thas, Marnik Vuylsteke, Rafael Irizarry Feb 2010

Wavelet Based Functional Models For Transcriptome Analysis With Tiling Arrays, Lieven Clement, Kristof Debeuf, Ciprian Crainiceanu, Olivier Thas, Marnik Vuylsteke, Rafael Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

For a better understanding of the biology of an organism a complete description is needed of all regions of the genome that are actively transcribed. Tiling arrays can be used for this purpose. Such arrays allow the discovery of novel transcripts and the assessment of differential expression between two or more experimental conditions such as genotype, treatment, tissue, etc. Much of the initial methodological efforts were designed for transcript discovery, while more recent developments also focus on differential expression. To our knowledge no methods for tiling arrays are described in the literature that can both assess transcript discovery and identify …


Bayesian Methods For Network-Structured Genomics Data, Stefano Monni, Hongzhe Li Jan 2010

Bayesian Methods For Network-Structured Genomics Data, Stefano Monni, Hongzhe Li

UPenn Biostatistics Working Papers

Graphs and networks are common ways of depicting information. In biology, many different processes are represented by graphs, such as regulatory networks, metabolic pathways and protein-protein interaction networks. This information provides useful supplement to the standard numerical genomic data such as microarray gene expression data. Effectively utilizing such an information can lead to a better identification of biologically relevant genomic features in the context of our prior biological knowledge. In this paper, we present a Bayesian variable selection procedure for network-structured covariates for both Gaussian linear and probit models. The key of our approach is the introduction of a Markov …


Biochemical Profiling Of Histone Binding Selectivity Of The Yeast Bromodomain Family, Qiang Zhang, Suvobrata Chakravarty, Dario Ghersi, Lei Zeng, Alexander N. Plotnikov, Roberto Sanchez, Ming-Ming Zhou Jan 2010

Biochemical Profiling Of Histone Binding Selectivity Of The Yeast Bromodomain Family, Qiang Zhang, Suvobrata Chakravarty, Dario Ghersi, Lei Zeng, Alexander N. Plotnikov, Roberto Sanchez, Ming-Ming Zhou

Interdisciplinary Informatics Faculty Publications

Background: It has been shown that molecular interactions between site-specific chemical modifications such as acetylation and methylation on DNA-packing histones and conserved structural modules present in transcriptional proteins are closely associated with chromatin structural changes and gene activation. Unlike methyl-lysine that can interact with different protein modules including chromodomains, Tudor and MBT domains, as well as PHD fingers, acetyl-lysine (Kac) is known thus far to be recognized only by bromodomains. While histone lysine acetylation plays a crucial role in regulation of chromatin-mediated gene transcription, a high degree of sequence variation of the acetyl-lysine binding site in the bromodomains has …


An Intelligent Data-Centric Approach Toward Identification Of Conserved Motifs In Protein Sequences, Kathryn Dempsey Cooper, Benjamin Currall, Richard Hallworth, Hesham Ali Jan 2010

An Intelligent Data-Centric Approach Toward Identification Of Conserved Motifs In Protein Sequences, Kathryn Dempsey Cooper, Benjamin Currall, Richard Hallworth, Hesham Ali

Interdisciplinary Informatics Faculty Proceedings & Presentations

The continued integration of the computational and biological sciences has revolutionized genomic and proteomic studies. However, efficient collaboration between these fields requires the creation of shared standards. A common problem arises when biological input does not properly fit the expectations of the algorithm, which can result in misinterpretation of the output. This potential confounding of input/output is a drawback especially when regarding motif finding software. Here we propose a method for improving output by selecting input based upon evolutionary distance, domain architecture, and known function. This method improved detection of both known and unknown motifs in two separate case studies. …


Computational Identification Of Micro Rna In Musa Sp Est Sequences., Tan Yew Seong Jan 2010

Computational Identification Of Micro Rna In Musa Sp Est Sequences., Tan Yew Seong

Student Works (2010-2019)

MicroRNAs (miRNAs) are small single-stranded non-protein-coding RNAs of about 22 nucleotides that play an important role in post-transcriptional gene regulation in animals and plants by targeting mRNAs for direct cleavage of mRNAs or repression of mRNA translation. In this study, a homology search based approach has been used for identifying the novel miRNA genes of Musa sp. using the BLAST program with a total of 42,978 Musa sp. expressed sequence tags (EST) were compared to a total of 959 previously known plant miRNA sequences. Forty-seven candidate miRNA with homology to previously known miRNAs were identified. By using the mFold3.1 program, …


Tracking Profiles Of Genomic Instability In Spontaneous Transformation And Tumorigenesis, Lesley Lawrenson Jan 2010

Tracking Profiles Of Genomic Instability In Spontaneous Transformation And Tumorigenesis, Lesley Lawrenson

Wayne State University Dissertations

The dominant paradigm for cancer research focuses on the identification of specific genes for cancer causation and for the discovery of therapeutic targets. Alternatively, the current data emphasize the significance of karyotype heterogeneity in cancer progression over specific gene-based causes of cancer. Variability of a magnitude significant to shift cell populations from homogeneous diploid cells to a mosaic of structural and numerical chromosome alterations reflects the characteristic low-fidelity genome transfer of cancer cell populations. This transition marks the departure from micro-evolutionary gene-level change to macro-evolutionary change that facilitates the generation of many unique karyotypes within a cell population. Considering cancer …


Genbank, Dennis A. Benson, Ilene Karasch-Mizrachi, David J. Lipman, James Ostell, Eric W. Sayers Jan 2010

Genbank, Dennis A. Benson, Ilene Karasch-Mizrachi, David J. Lipman, James Ostell, Eric W. Sayers

Harold W. Manter Laboratory of Parasitology: Library Materials

GenBank(R) is a comprehensive database that contains publicly available nucleotide sequences for more than 380,000 organisms named at the genus level or lower, obtained primarily through submissions from individual laboratories and batch submissions from large-scale sequencing projects, including whole genome shotgun (WGS) and environmental sampling projects. Most submissions are made using the web-based BankIt or standalone Sequin programs, and accession numbers are assigned by GenBank staff upon receipt. Daily data exchange with the European Nucleotide Archive (ENA) and the DNA Data Bank of Japan (DDBJ) ensures worldwide coverage. GenBank is accessible through the NCBI Entrez retrieval system that integrates data …


High-Density Screening Reveals A Different Spectrum Of Genomic Aberrations In Chronic Lymphocytic Leukemia Patients With ‘Stereotyped’ Ighv3-21 And Ighv4-34 B-Cell Receptors, Millaray Marincevic, Nicola Cahill, Rebeqa Gunnarsson, Anders Isaksson, Mahmoud Mansouri, Hanna Göransson, Markus Rasmussen, Mattias Jansson, Fergus Ryan, Karin Karlsson, Hans-Olov Adami, Fred Davi, Jesper Jurlander, Gunnar Juliusson, Kostas Stamatopoulos, Richard Rosenquist Jan 2010

High-Density Screening Reveals A Different Spectrum Of Genomic Aberrations In Chronic Lymphocytic Leukemia Patients With ‘Stereotyped’ Ighv3-21 And Ighv4-34 B-Cell Receptors, Millaray Marincevic, Nicola Cahill, Rebeqa Gunnarsson, Anders Isaksson, Mahmoud Mansouri, Hanna Göransson, Markus Rasmussen, Mattias Jansson, Fergus Ryan, Karin Karlsson, Hans-Olov Adami, Fred Davi, Jesper Jurlander, Gunnar Juliusson, Kostas Stamatopoulos, Richard Rosenquist

Articles

Background The existence of multiple subsets of chronic lymphocytic leukemia expressing ‘stereotyped’ Bcell receptors implies the involvement of antigen(s) in leukemogenesis. Studies also indicate that ‘stereotypy’ may influence the clinical course of patients with chronic lymphocytic leukemia, for example, in subsets with stereotyped IGHV3-21 and IGHV4-34 B-cell receptors; however, little is known regarding the genomic profile of patients in these subsets. Design and Methods We applied 250K single nucleotide polymorphism-arrays to study copy-number aberrations and copy-number neutral loss-of-heterozygosity in patients with stereotyped IGHV3-21 (subset #2, n=29), stereotyped IGHV4-34 (subset #4, n=17; subset #16, n=8) and non-subset #2 IGHV3-21 (n=13) and …


Attempted Cloning Of A Wnt Gene From Botrylloides Violaceus, Manasa Chandra, James Tumulak Dec 2009

Attempted Cloning Of A Wnt Gene From Botrylloides Violaceus, Manasa Chandra, James Tumulak

Biological Sciences

Botrylloides violaceus is a colonial ascidian with the ability to undergo sexual and asexual reproduction as well as regeneration. The canonical pathway starts with the extracellular protein Wnt and ends with β-catenin, a transcription factor, which also functions in cell adhesion. The Wnt signaling pathway is involved in embryogenesis and regeneration in a variety of other species. In our studies we attempt to isolate and sequence both a Wnt gene and from Botrylloides via degenerate primer design and PCR. Using bioinformatic methods we aligned sequences from other organisms, as the Botrylloides genome has not yet been sequenced. Using mouse, Ciona, …


Applications Of Variable Number Tandem Repeat Genotyping In The Validation Of An Animal Medical Model And Gene Flow Studies In Threatened Populations Of Reptiles, Candace D. Smith Dec 2009

Applications Of Variable Number Tandem Repeat Genotyping In The Validation Of An Animal Medical Model And Gene Flow Studies In Threatened Populations Of Reptiles, Candace D. Smith

Graduate Theses and Dissertations

We used variable number tandem repeats (VNTR) to validate the chicken as a human medical model for Pulmonary Arterial Hypertension. We identified seven regions on four chromosomes and interrogated for VNTR markers that significantly associate with Pulmonary Hypertension Syndrome/ascites. In those regions, we identified 7 candidate genes; AGTR1, ACE, p38MAPK, SST, 5HT2B, NET1, and CALM3 for further analysis as significantly contributing QTL for ascites/PHS. We also used variable number tandem repeats to measure gene flow and gather evidence for multiple paternity in a population of Timber rattlesnakes, Crotalus horridus. We were able to verify 1 VNTR that can be used …


Targeted Genomic Signature Profiling With Quasi-Alignment Statistics, Rao Mallik Kotamarti, Douglas W. Raiford, Michael Hahsler, Yuhang Wang, Monnie Mcgee, Maggie Dunham Nov 2009

Targeted Genomic Signature Profiling With Quasi-Alignment Statistics, Rao Mallik Kotamarti, Douglas W. Raiford, Michael Hahsler, Yuhang Wang, Monnie Mcgee, Maggie Dunham

COBRA Preprint Series

Genome databases continue to expand with no change in the basic format of sequence data. The prevalent use of the Classic alignment based search tools like BLAST have significantly pushed the limits of Genome Isolate research. The relatively new frontier of Metagenomic research deals with thousands of diverse genomes with newer demands beyond the current homologue search and analysis. Compressing sequence data into a complex form could facilitate a broader range of sequence analyses. To this end, this research explores reorganizing sequence data as complex Markov signatures also known as Extensible Markov Models. Markov models have found successful application in …


Integrative Clustering Of Multiple Genomic Data Types Using A Joint Latent Variable Model With Application To Breast And Lung Cancer Subtype Analysis, Ronglai Shen, Adam Olshen, Marc Ladanyi Sep 2009

Integrative Clustering Of Multiple Genomic Data Types Using A Joint Latent Variable Model With Application To Breast And Lung Cancer Subtype Analysis, Ronglai Shen, Adam Olshen, Marc Ladanyi

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

The molecular complexity of a tumor manifests itself at the genomic, epigenomic, transcriptomic, and proteomic levels. Genomic profiling at these multiple levels should allow an integrated characterization of tumor etiology. However, there is a shortage of effective statistical and bioinformatic tools for truly integrative data analysis. The standard approach to integrative clustering is separate clustering followed by manual integration. A more statistically powerful approach would incorporate all data types simultaneously and generate a single integrated cluster assignment. We developed a joint latent variable model for integrative clustering. We call the resulting methodology iCluster. iCluster incorporates flexible modeling of the associations …


Model-Based Quality Assessment And Base-Calling For Second-Generation Sequencing Data, Rafael A. Irizarry, Hector Corrada Bravo Sep 2009

Model-Based Quality Assessment And Base-Calling For Second-Generation Sequencing Data, Rafael A. Irizarry, Hector Corrada Bravo

Johns Hopkins University, Dept. of Biostatistics Working Papers

Second-generation sequencing (sec-gen) technology can sequence millions of short fragments of DNA in parallel, and is capable of assembling complex genomes for a small fraction of the price and time of previous technologies. In fact, a recently formed international consortium, the 1,000 Genomes Project, plans to fully sequence the genomes of approximately 1,200 people. The prospect of comparative analysis at the sequence level of a large number of samples across multiple populations may be achieved within the next five years. These data present unprecedented challenges in statistical analysis. For instance, analysis operates on millions of short nucleotide sequences, or reads—strings …


A Classification Model For Distinguishing Copy Number Variants From Cancer-Related Alterations, Irina Ostrovnaya, Gouri Nanjangud, Adam Olshen Aug 2009

A Classification Model For Distinguishing Copy Number Variants From Cancer-Related Alterations, Irina Ostrovnaya, Gouri Nanjangud, Adam Olshen

Memorial Sloan-Kettering Cancer Center, Dept. of Epidemiology & Biostatistics Working Paper Series

Both somatic copy number alterations (CNAs) and germline copy number variants (CNVs) that are prevalent in healthy individuals can appear as recurrent changes in comparative genomic hybridization (CGH) analyses of tumors. In order to identify important cancer genes CNAs and CNVs must be distinguished. Although the Database of Genomic Variants (Iafrate et al., 2004) contains a list of all known CNVs, there is no standard methodology to use the database effectively.

We develop a prediction model that distinguishes CNVs from CNAs based on the information contained in the Database and several other variables, including potential CNV’s length, height, closeness to …


Subset Quantile Normalization Using Negative Control Features, Zhijin Wu Jun 2009

Subset Quantile Normalization Using Negative Control Features, Zhijin Wu

Johns Hopkins University, Dept. of Biostatistics Working Papers

No abstract provided.


Frozen Robust Multi-Array Analysis (Frma), Matthew N. Mccall, Benjamin M. Bolstad, Rafael A. Irizarry May 2009

Frozen Robust Multi-Array Analysis (Frma), Matthew N. Mccall, Benjamin M. Bolstad, Rafael A. Irizarry

Johns Hopkins University, Dept. of Biostatistics Working Papers

Robust Multi-array Analysis (RMA) is the most widely used preprocessing algorithm for Affymetrix and Nimblegen gene-expression microarrays. RMA performs background correction, normalization, and summarization in a modular way. The last two steps require multiple arrays to be analyzed simultaneously. The ability to borrow information across samples provides RMA various advantages. For example, the summarization step fits a parametric model that accounts for probe-effects, assumed to be fixed across arrays, and improves outlier detection. Residuals, obtained from the fitted model, permit the creation of useful quality metrics. However, the dependence on multiple arrays has two drawbacks: (1) RMA can- not be …


Resampling-Based Multiple Hypothesis Testing With Applications To Genomics: New Developments In The R/Bioconductor Package Multtest, Houston N. Gilbert, Katherine S. Pollard, Mark J. Van Der Laan, Sandrine Dudoit Apr 2009

Resampling-Based Multiple Hypothesis Testing With Applications To Genomics: New Developments In The R/Bioconductor Package Multtest, Houston N. Gilbert, Katherine S. Pollard, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

The multtest package is a standard Bioconductor package containing a suite of functions useful for executing, summarizing, and displaying the results from a wide variety of multiple testing procedures (MTPs). In addition to many popular MTPs, the central methodological focus of the multtest package is the implementation of powerful joint multiple testing procedures. Joint MTPs are able to account for the dependencies between test statistics by effectively making use of (estimates of) the test statistics joint null distribution. To this end, two additional bootstrap-based estimates of the test statistics joint null distribution have been developed for use in the …


Sitehound-Web: A Server For Ligand Binding Site Identification In Protein Structures, Marylens Hernandez, Dario Ghersi, Roberto Sanchez Apr 2009

Sitehound-Web: A Server For Ligand Binding Site Identification In Protein Structures, Marylens Hernandez, Dario Ghersi, Roberto Sanchez

Interdisciplinary Informatics Faculty Publications

SITEHOUND-web (http://sitehound.sanchezlab.org) is a binding-site identification server powered by the SITEHOUND program. Given a protein structure in PDB format SITEHOUND-web will identify regions of the protein characterized by favorable interactions with a probe molecule. These regions correspond to putative ligand binding sites. Depending on the probe used in the calculation, sites with preference for different ligands will be identified. Currently, a carbon probe for identification of binding sites for drug-like molecules, and a phosphate probe for phosphorylated ligands (ATP, phoshopeptides, etc.) have been implemented. SITEHOUND-web will display the results in HTML pages including an interactive 3D representation of …


Evaluation Of Statistical Methods For Normalization And Differential Expression In Mrna-Seq Experiments, James H. Bullard, Elizabeth A. Purdom, Kasper D. Hansen, Sandrine Dudoit Apr 2009

Evaluation Of Statistical Methods For Normalization And Differential Expression In Mrna-Seq Experiments, James H. Bullard, Elizabeth A. Purdom, Kasper D. Hansen, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

The focus of this article is on the design and analysis of mRNA-Seq experiments, with the aim of inferring transcript levels and identifying differentially expressed genes. We investigate two mRNA-Seq datasets obtained using Illumina's Genome Analyzer platform to measure transcript levels in reference samples considered in the MicroArray Quality Control (MAQC) Project. We address the following four main issues: (1) exploratory data analysis for mapped reads, relating read counts to variables describing input samples and genomic regions of interest; (2) assessment and quantitation of biological effects (e.g., expression levels in Brain vs. UHR) and nuisance experimental effects (e.g., library preparation, …


Gene Set Enrichment Analysis Made Simple, Rafael A. Irizarry, Chi Wang, Yun Zhou, Terence P. Speed Apr 2009

Gene Set Enrichment Analysis Made Simple, Rafael A. Irizarry, Chi Wang, Yun Zhou, Terence P. Speed

Johns Hopkins University, Dept. of Biostatistics Working Papers

Among the many applications of microarray technology, one of the most popular is the identification of genes that are differentially expressed in two conditions. A common statistical approach is to quantify the interest of each gene with a p-value, adjust these p-values for multiple comparisons, chose an appropriate cut-off, and create a list of candidate genes. This approach has been criticized for ignoring biological knowledge regarding how genes work together. Recently a series of methods, that do incorporate biological knowledge, have been proposed. However, many of these methods seem overly complicated. Furthermore, the most popular method, Gene Set Enrichment Analysis …


Generalized Liquid Association, Yen-Yi Ho, Leslie Cope, Thomas A. Louis, Giovanni Parmigiani Apr 2009

Generalized Liquid Association, Yen-Yi Ho, Leslie Cope, Thomas A. Louis, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

The analysis of interactions among a group of genes is fundamental to fur- ther our understanding of their biological interactions in a cell. Several studies suggested that the co-expression relationship of two genes can be modulated by a third controller gene. These controller genes and the corresponding modulated co-expressed gene pairs are the subjects of interests in this study. This described \controller-modulated genes" three-way interactions is referred as liquid association in the literature. Analysis of gene expression data has suggested that these interactions are present in many biological systems.

To quantify the magnitude of liquid association for a given gene …


Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit Apr 2009

Joint Multiple Testing Procedures For Graphical Model Selection With Applications To Biological Networks, Houston N. Gilbert, Mark J. Van Der Laan, Sandrine Dudoit

U.C. Berkeley Division of Biostatistics Working Paper Series

Gaussian graphical models have become popular tools for identifying relationships between genes when analyzing microarray expression data. In the classical undirected Gaussian graphical model setting, conditional independence relationships can be inferred from partial correlations obtained from the concentration matrix (= inverse covariance matrix) when the sample size n exceeds the number of parameters p which need to estimated. In situations where n < p, another approach to graphical model estimation may rely on calculating unconditional (zero-order) and first-order partial correlations. In these settings, the goal is to identify a lower-order conditional independence graph, sometimes referred to as a ‘0-1 graphs’. For either choice of graph, model selection may involve a multiple testing problem, in which edges in a graph are drawn only after rejecting hypotheses involving (saturated or lower-order) partial correlation parameters. Most multiple testing procedures applied in previously proposed graphical model selection algorithms rely on standard, marginal testing methods which do not take into account the joint distribution of the test statistics derived from (partial) correlations. We propose and implement a multiple testing framework useful when testing for edge inclusion during graphical model selection. Two features of our methodology include (i) a computationally efficient and asymptotically valid test statistics joint null distribution derived from influence curves for correlation-based parameters, and (ii) the application of empirical Bayes joint multiple testing procedures which can effectively control a variety of popular Type I error rates by incorpo- rating joint null distributions such as those described here (Dudoit and van der Laan, 2008). Using a dataset from Arabidopsis thaliana, we observe that the use of more sophisticated, modular approaches to multiple testing allows one to identify greater numbers of edges when approximating an undirected graphical model using a 0-1 graph. Our framework may also be extended to edge testing algorithms for other types of graphical models (e.g., for classical undirected, bidirected, and directed acyclic graphs).


Multifactor Dimensionality Reduction Analysis Identifies Specific Nucleotide Patterns Promoting Genetic Polymorphisms, Eric Arehart, Scott Gleim, Bill White, John Hwa, Jason H. Moore Mar 2009

Multifactor Dimensionality Reduction Analysis Identifies Specific Nucleotide Patterns Promoting Genetic Polymorphisms, Eric Arehart, Scott Gleim, Bill White, John Hwa, Jason H. Moore

Dartmouth Scholarship

The fidelity of DNA replication serves as the nidus for both genetic evolution and genomic instability fostering disease. Single nucleotide polymorphisms (SNPs) constitute greater than 80% of the genetic variation between individuals. A new theory regarding DNA replication fidelity has emerged in which selectivity is governed by base-pair geometry through interactions between the selected nucleotide, the complementary strand, and the polymerase active site. We hypothesize that specific nucleotide combinations in the flanking regions of SNP fragments are associated with mutation.


A Novel Topology For Representing Protein Folds, Mark R. Segal Mar 2009

A Novel Topology For Representing Protein Folds, Mark R. Segal

COBRA Preprint Series

Various topologies for representing three dimensional protein structures have been advanced for purposes ranging from prediction of folding rates to ab initio structure prediction. Examples include relative contact order, Delaunay tessellations, and backbone torsion angle distributions. Here we introduce a new topology based on a novel means for operationalizing three dimensional proximities with respect to the underlying chain. The measure involves first interpreting a rank-based representation of the nearest neighbors of each residue as a permutation, then determining how perturbed this permutation is relative to an unfolded chain. We show that the resultant topology provides improved association with folding and …


Sparse Linear Discriminant Analysis For Simultaneous Testing For The Significance Of A Gene Set/Pathway And Gene Selection, Michael C. Wu, Lingson Zhang, Zhaoxi Wang, David C. Christiani, Xihong Lin Jan 2009

Sparse Linear Discriminant Analysis For Simultaneous Testing For The Significance Of A Gene Set/Pathway And Gene Selection, Michael C. Wu, Lingson Zhang, Zhaoxi Wang, David C. Christiani, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.