Open Access. Powered by Scholars. Published by Universities.®

Computational Biology Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 661 - 690 of 754

Full-Text Articles in Computational Biology

Detection Of Recurrent Copy Number Alterations In The Genome: A Probabilistic Approach, Oscar M. Rueda, Ramon Diaz-Uriarte Nov 2008

Detection Of Recurrent Copy Number Alterations In The Genome: A Probabilistic Approach, Oscar M. Rueda, Ramon Diaz-Uriarte

COBRA Preprint Series

Copy number variation (CNV) in genomic DNA is linked to a variety of human diseases (including cancer, HIV acquisition, autoimmune and neurodegenerative diseases), and array-based CGH (aCGH) is currently the main technology to locate CNVs. Several methods can analyze aCGH data at the single sample level, but disease-critical genes are more likely to be found in regions that are common or recurrent among samples. Unfortunately, defining recurrent CNV regions remains a challenge. Moreover, the heterogeneous nature of many diseases requires that we search for CNVs that affect only some subsets of the samples (without prior knowledge of which regions and …


Finding Recurrent Regions Of Copy Number Variation: A Review, Oscar M. Rueda, Ramon Diaz-Uriarte Nov 2008

Finding Recurrent Regions Of Copy Number Variation: A Review, Oscar M. Rueda, Ramon Diaz-Uriarte

COBRA Preprint Series

Copy number variation (CNV) in genomic DNA is linked to a variety of human diseases, and array-based CGH (aCGH) is currently the main technology to locate CNVs. Although many methods have been developed to analyze aCGH from a single array/subject, disease-critical genes are more likely to be found in regions that are common or recurrent among subjects. Unfortunately, finding recurrent CNV regions remains a challenge. We review existing methods for the identification of recurrent CNV regions. The working definition of ``common'' or ``recurrent'' region differs between methods, leading to approaches that use different types of input (discretized output from a …


The Strength Of Statistical Evidence For Composite Hypotheses With An Application To Multiple Comparisons, David R. Bickel Nov 2008

The Strength Of Statistical Evidence For Composite Hypotheses With An Application To Multiple Comparisons, David R. Bickel

COBRA Preprint Series

The strength of the statistical evidence in a sample of data that favors one composite hypothesis over another may be quantified by the likelihood ratio using the parameter value consistent with each hypothesis that maximizes the likelihood function. Unlike the p-value and the Bayes factor, this measure of evidence is coherent in the sense that it cannot support a hypothesis over any hypothesis that it entails. Further, when comparing the hypothesis that the parameter lies outside a non-trivial interval to the hypotheses that it lies within the interval, the proposed measure of evidence almost always asymptotically favors the correct hypothesis …


A Network-Constrained Empirical Bayes Method For Analysis Of Genomic Data, Caiyan Li, Zhi Wei, Hongzhe Li Oct 2008

A Network-Constrained Empirical Bayes Method For Analysis Of Genomic Data, Caiyan Li, Zhi Wei, Hongzhe Li

UPenn Biostatistics Working Papers

Empirical Bayes methods are widely used in the analysis of microarray gene expression data in order to identify the differentially expressed genes or genes that are associated with other general phenotypes. Available methods often assume that genes are independent. However, genes are expected to function interactively and to form molecular modules to affect the phenotypes. In order to account for regulatory dependency among genes, we propose in this paper a network-constrained empirical Bayes method for analyzing genomic data in the framework of general linear models, where the dependency of genes is modeled by a discrete Markov random field model defined …


Estimation And Testing For The Effect Of A Genetic Pathway On A Disease Outcome Using Logistic Kernel Machine Regression Via Logistic Mixed Models, Dawei Liu, Debashis Ghosh, Xihong Lin Jun 2008

Estimation And Testing For The Effect Of A Genetic Pathway On A Disease Outcome Using Logistic Kernel Machine Regression Via Logistic Mixed Models, Dawei Liu, Debashis Ghosh, Xihong Lin

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Powerful And Flexible Multilocus Association Test For Quantitative Traits, Lydia Coulter Kwee, Dawei Liu, Xihong Lin, Debashis Ghosh, Michael P. Epstein Jun 2008

A Powerful And Flexible Multilocus Association Test For Quantitative Traits, Lydia Coulter Kwee, Dawei Liu, Xihong Lin, Debashis Ghosh, Michael P. Epstein

Harvard University Biostatistics Working Paper Series

No abstract provided.


Model-Based Clustering Of Methylation Array Data: A Recursive-Partitioning Algorithm For High-Dimensional Data Arising As A Mixture Of Beta Distributions, E. Andres Houseman, Brock C. Christensen, Ru-Fang Yeh, Carmen J. Marsit, Margaret R. Karagas, Margaret Wrensch, Heather H. Nelson, Joseph Wiemels, Shichun Zheng, John K. Wiencke, Karl T. Kelsey Jun 2008

Model-Based Clustering Of Methylation Array Data: A Recursive-Partitioning Algorithm For High-Dimensional Data Arising As A Mixture Of Beta Distributions, E. Andres Houseman, Brock C. Christensen, Ru-Fang Yeh, Carmen J. Marsit, Margaret R. Karagas, Margaret Wrensch, Heather H. Nelson, Joseph Wiemels, Shichun Zheng, John K. Wiencke, Karl T. Kelsey

Harvard University Biostatistics Working Paper Series

No abstract provided.


Incorporation Of Genetic Pathway Information Into Analysis Of Multivariate Gene Expression Data, Zhi Wei, Jane E. Minturn, Eric Rappaport, Garrett Brodeur, Hongzhe Li Apr 2008

Incorporation Of Genetic Pathway Information Into Analysis Of Multivariate Gene Expression Data, Zhi Wei, Jane E. Minturn, Eric Rappaport, Garrett Brodeur, Hongzhe Li

UPenn Biostatistics Working Papers

Abstract: Multivariate microarray gene expression data are commonly collected to study the genomic responses under ordered conditions such as over increasing/decreasing dose levels or over time during biological processes. One important question from such multivariate gene expression experiments is to identify genes that show different expression patterns over treatment dosages or over time and pathways that are perturbed during a given biological process. In this paper, we develop a hidden Markov random field model for multivariate expression data in order to identify genes and subnetworks that are related to biological processes, where the dependency of the differential expression patterns of …


Likelihood Estimation Of Conjugacy Relationships In Linear Models With Applications To High-Throughput Genomics, Brian S. Caffo, Liu Dongmei, Robert Scharpf, Giovanni Parmigiani Apr 2008

Likelihood Estimation Of Conjugacy Relationships In Linear Models With Applications To High-Throughput Genomics, Brian S. Caffo, Liu Dongmei, Robert Scharpf, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

In the simultaneous estimation of a large number of related quantities, multilevel models provide a formal mechanism for efficiently making use of the ensemble of information for deriving individual estimates. In this article we investigate the ability of the likelihood to identify the relationship between signal and noise in multilevel linear mixed models. Specifically, we consider the ability of the likelihood to diagnose conjugacy or independence between the signals and noises. Our work was motivated by the analysis of data from high-throughput experiments in genomics. The proposed model leads to a more flexible family. However, we further demonstrate that adequately …


Micrornas And The Advent Of Vertebrate Morphological Complexity, Alysha M. Heimberg, Lorenzo F. Sempere, Vanessa N. Moy, Phillip C. J. Donoghue, Kevin J. Peterson Feb 2008

Micrornas And The Advent Of Vertebrate Morphological Complexity, Alysha M. Heimberg, Lorenzo F. Sempere, Vanessa N. Moy, Phillip C. J. Donoghue, Kevin J. Peterson

Dartmouth Scholarship

The causal basis of vertebrate complexity has been sought in genome duplication events (GDEs) that occurred during the emergence of vertebrates, but evidence beyond coincidence is wanting. MicroRNAs (miRNAs) have recently been identified as a viable causal factor in increasing organismal complexity through the action of these ≈22-nt noncoding RNAs in regulating gene expression. Because miRNAs are continuously being added to animalian genomes, and, once integrated into a gene regulatory network, are strongly conserved in primary sequence and rarely secondarily lost, their evolutionary history can be accurately reconstructed. Here, using a combination of Northern analyses and genomic searches, we show …


Empirical Null And False Discovery Rate Inference For Exponential Families, Armin Schwartzman Feb 2008

Empirical Null And False Discovery Rate Inference For Exponential Families, Armin Schwartzman

Harvard University Biostatistics Working Paper Series

No abstract provided.


Assessing The Role Of Multi-Protein Complexes In Determining Phenotype, Nolwenn Le Meur, Robert Gentleman Jan 2008

Assessing The Role Of Multi-Protein Complexes In Determining Phenotype, Nolwenn Le Meur, Robert Gentleman

Bioconductor Project Working Papers

Understanding regulatory mechanisms in complex biological systems is an important challenge, in particular to understand disease mechanisms, and to discover new therapies and drugs. In this paper, we consider the important question of cellular regulation of phenotype. Using single gene deletion data, we address the problem of linking a phenotype to underlying functional roles in the organism and provide a sound computational and statistical paradigm that can be extended to address more complex experimental settings such as multiple deletions. We apply the proposed approaches to publicly available data sets to demonstrate strong evidence for the involvement of multi-protein complexes in …


Design And Analysis Issues In Genome-Wide Somatic Mutation Studies Of Cancer, Giovanni Parmigiani, Simina Boca, Jimmy Lin, Kenneth W. Kinzler, Victor E. Velculescu, Bert Vogelstein Jan 2008

Design And Analysis Issues In Genome-Wide Somatic Mutation Studies Of Cancer, Giovanni Parmigiani, Simina Boca, Jimmy Lin, Kenneth W. Kinzler, Victor E. Velculescu, Bert Vogelstein

Johns Hopkins University, Dept. of Biostatistics Working Papers

The availability of the human genome sequence and progress in sequencing and bioinformatic technologies have enabled genome-wide investigation of somatic mu- tations in human cancers. This article briefly reviews challenges arising in the statistical analysis of mutational data of this kind. A first challenge is that of designing studies that efficiently allocate sequencing resources. We show that this can be addressed by two-stage designs, and demonstrate via simulations that even relatively small studies can produce lists of candidate cancer genes that are highly informative for future research efforts. A second challenge is to distinguish mutated genes that are selected for …


Advancing Epidemiological Science Through Computational Modeling: A Review With Novel Examples, Scott M. Duke-Sylvester, Eli N. Perencevich, Jon P. Furuno, Leslie A. Real, Holly Gaff Jan 2008

Advancing Epidemiological Science Through Computational Modeling: A Review With Novel Examples, Scott M. Duke-Sylvester, Eli N. Perencevich, Jon P. Furuno, Leslie A. Real, Holly Gaff

Biological Sciences Faculty Publications

Computational models have been successfully applied to a wide variety of research areas including infectious disease epidemiology. Especially for questions that are difficult to examine in other ways, computational models have been used to extend the range of epidemiological issues that can be addressed, advance theoretical understanding of disease processes and help identify specific intervention strategies. We explore each of these contributions to epidemiology research through discussion and examples. We also describe in detail models for raccoon rabies and methicillin-resis-tant Staphylococcus aureus, drawn from our own research, to further illustrate the role of computation in epidemiological modeling.


Network-Constrained Regularization And Variable Selection For Analysis Of Genomic Data, Caiyan Li, Hongzhe Li Dec 2007

Network-Constrained Regularization And Variable Selection For Analysis Of Genomic Data, Caiyan Li, Hongzhe Li

UPenn Biostatistics Working Papers

Graphs or networks are common ways of depicting information. In biology in particular, many different biological processes are represented by graphs, such as regulatory networks or metabolic pathways. This kind of {\it a priori} information gathered over many years of biomedical research is a useful supplement to the standard numerical genomic data such as microarray gene expression data. How to incorporate information encoded by the known biological networks or graphs into analysis of numerical data raises interesting statistical challenges. In this paper, we introduce a network-constrained regularization procedure for linear regression analysis in order to incorporate the information from these …


Vertex Clustering In Random Graphs Via Reversible Jump Markov Chain Monte Carlo, Stefano Monni, Hongzhe Li Dec 2007

Vertex Clustering In Random Graphs Via Reversible Jump Markov Chain Monte Carlo, Stefano Monni, Hongzhe Li

UPenn Biostatistics Working Papers

Networks are a natural and effective tool to study relational data, in which observations are collected on pairs of units. The units are represented by nodes and their relations by edges. In biology, for example, proteins and their interactions, and, in social science, people and inter-personal relations may be the nodes and the edges of the network. In this paper we address the question of clustering vertices in networks, as a way to uncover homogeneity patterns in data that enjoy a network representation. We use a mixture model for random graphs and propose a reversible jump Markov chain Monte Carlo …


A Bayesian Model For Cross-Study Differential Gene Expression, Robert B. Scharpf, Hakon Tjelemeland, Giovanni Parmigiani, Andrew B. Nobel Nov 2007

A Bayesian Model For Cross-Study Differential Gene Expression, Robert B. Scharpf, Hakon Tjelemeland, Giovanni Parmigiani, Andrew B. Nobel

Johns Hopkins University, Dept. of Biostatistics Working Papers

In this paper we define a hierarchical Bayesian model for microarray expression data collected from several studies and use it to identify genes that show differential expression between two conditions. Key features include shrinkage across both genes and studies; flexible modeling that allows for interactions between platforms and the estimated effect, and for both concordant and discordant differential expression across studies. We evaluated the performance of our model in a comprehensive fashion, using both artificial data, and a "split-sample" validation approach that provides an agnostic assessment of the model's behavior not only under the null hypothesis but also under a …


Assessing Population Level Genetic Instability Via Moving Average, Samuel Mcdaniel, Rebecca Betensky, Tianxi Cai Nov 2007

Assessing Population Level Genetic Instability Via Moving Average, Samuel Mcdaniel, Rebecca Betensky, Tianxi Cai

Harvard University Biostatistics Working Paper Series

No abstract provided.


Statistical Methods For The Analysis Of Cancer Genome Sequencing Data, Giovanni Parmigiani, J. Lin, Simina Boca, T. Sjoblom, K.W. Kinzler, V.E. Velculescu, B. Vogelstein Oct 2007

Statistical Methods For The Analysis Of Cancer Genome Sequencing Data, Giovanni Parmigiani, J. Lin, Simina Boca, T. Sjoblom, K.W. Kinzler, V.E. Velculescu, B. Vogelstein

Johns Hopkins University, Dept. of Biostatistics Working Papers

The purpose of cancer genome sequencing studies is to determine the nature and types of alterations present in a typical cancer and to discover genes mutated at high frequencies. In this article we discuss statistical methods for the analysis of data generated in these studies. We place special emphasis on a two-stage study design introduced by Sjoblom et al.[1]. In this context, we describe statistical methods for constructing scores that can be used to prioritize candidate genes for further investigation and to assess the statistical signicance of the candidates thus identfied.


A Novel Ensemble Learning Method For De Novo Computational Identification Of Dna Binding Sites, Arijit Chakravarty, Jonathan M. Carlson, Radhika S. Khetani, Robert H H. Gross Jul 2007

A Novel Ensemble Learning Method For De Novo Computational Identification Of Dna Binding Sites, Arijit Chakravarty, Jonathan M. Carlson, Radhika S. Khetani, Robert H H. Gross

Dartmouth Scholarship

Despite the diversity of motif representations and search algorithms, the de novo computational identification of transcription factor binding sites remains constrained by the limited accuracy of existing algorithms and the need for user-specified input parameters that describe the motif being sought.ResultsWe present a novel ensemble learning method, SCOPE, that is based on the assumption that transcription factor binding sites belong to one of three broad classes of motifs: non-degenerate, degenerate and gapped motifs. SCOPE employs a unified scoring metric to combine the results from three motif finding algorithms each aimed at the discovery of one of these classes of motifs. …


Assessment Of A Cgh-Based Genetic Instability, David A. Engler, Yiping Shen, J F. Gusella, Rebecca A. Betensky Jul 2007

Assessment Of A Cgh-Based Genetic Instability, David A. Engler, Yiping Shen, J F. Gusella, Rebecca A. Betensky

Harvard University Biostatistics Working Paper Series

No abstract provided.


Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li Jul 2007

Survival Analysis With Large Dimensional Covariates: An Application In Microarray Studies, David A. Engler, Yi Li

Harvard University Biostatistics Working Paper Series

Use of microarray technology often leads to high-dimensional and low- sample size data settings. Over the past several years, a variety of novel approaches have been proposed for variable selection in this context. However, only a small number of these have been adapted for time-to-event data where censoring is present. Among standard variable selection methods shown both to have good predictive accuracy and to be computationally efficient is the elastic net penalization approach. In this paper, adaptation of the elastic net approach is presented for variable selection both under the Cox proportional hazards model and under an accelerated failure time …


The Integrative Correlation Coefficient: A Measure Of Cross-Study Reproducibility For Gene Expressionea Array Data, Leslie M. Cope, Liz Garrett-Mayer, Edward Gabrielson, Giovanni Parmigiani May 2007

The Integrative Correlation Coefficient: A Measure Of Cross-Study Reproducibility For Gene Expressionea Array Data, Leslie M. Cope, Liz Garrett-Mayer, Edward Gabrielson, Giovanni Parmigiani

Johns Hopkins University, Dept. of Biostatistics Working Papers

Multi-study analysis adds value to microarray experiments. However, because of significant technical differences between microarray platforms, and because of differences in study design, it can be difficult to combine data. We have developed a statistical measure of reproducibility that can be applied to individual genes, measured in two different studies. This statistic, which we call the Integrative Correlation Coefficient or Correlation of Correlations, borrows strength across many genes to estimate the strength of the relationship between expression values in the two studies.


What Is The Best Reference Rna? And Other Questions Regarding The Design And Analysis Of Two-Color Microarray Experiments, Kathleen F. Kerr, Kyle A. Serikawa, Caimiao Wei, Mette A. Peters, Roger E. Bumgarner Apr 2007

What Is The Best Reference Rna? And Other Questions Regarding The Design And Analysis Of Two-Color Microarray Experiments, Kathleen F. Kerr, Kyle A. Serikawa, Caimiao Wei, Mette A. Peters, Roger E. Bumgarner

UW Biostatistics Working Paper Series

The reference design is a practical and popular choice for microarray studies using two-color platforms. In the reference design, the reference RNA uses half of all array resources, leading investigators to ask: What is the best reference RNA? We propose a novel method for evaluating reference RNAs and present the results of an experiment that was specially designed to evaluate three common choices of reference RNA. We found no compelling evidence in favor of any particular reference. In particular, a commercial reference showed no advantage in our data. Our experimental design also enabled a new way to test the effectiveness …


A Markov Random Field Model For Network-Based Analysis Of Genomic Data, Zhi Wei, Hongzhe Li Mar 2007

A Markov Random Field Model For Network-Based Analysis Of Genomic Data, Zhi Wei, Hongzhe Li

UPenn Biostatistics Working Papers

A central problem in genomic research is the identification of genes and pathways involved in diseases and other biological processes. The genes identified or the univariate test statistics are often linked to known biological pathways through gene set enrichment analysis in order to identify the pathways involved. However, most of the procedures for identifying differentially expressed genes do not utilize the known pathway information in the phase of identifying such genes. In this paper, we develop a Markov random field (MRF)-based method for identifying genes and subnetworks that are related to diseases. Such a procedure models the dependency of the …


Statistical Methods For Inference Of Genetic Networks And Regulatory Modules, Hongzhe Li Mar 2007

Statistical Methods For Inference Of Genetic Networks And Regulatory Modules, Hongzhe Li

UPenn Biostatistics Working Papers

Large-scale microarray gene expression data, motif data derived from promotor sequences, genome-wide chromatin immunoprecipitation (ChIP-chip) data, DNA polymorphism data and epigenomic data provide the possibility of constructing genetic networks or biological pathways, especially regulatory networks. In this paper, we review some new statistical methods for inference of genetic networks and regulatory modules, including a threshold gradient descent procedure for inference of Gaussian graphical models, a sparse regression mixture modeling approach for inference of regulatory modules, and the varying coefficient model for identifying regulatory subnetworks by integrating microarray time-course gene expression data and motif or ChIP-chip data. We present the statistical …


Conservative Estimation Of Optimal Multiple Testing Procedures, James E. Signorovitch Mar 2007

Conservative Estimation Of Optimal Multiple Testing Procedures, James E. Signorovitch

Harvard University Biostatistics Working Paper Series

No abstract provided.


A Hidden Markov Model For Joint Estimation Of Genotype And Copy Number In High-Throughput Snp Chips, Robert B. Scharpf, Giovanni Parmigiani, Jonathan Pevnser, Ingo Ruczinski Feb 2007

A Hidden Markov Model For Joint Estimation Of Genotype And Copy Number In High-Throughput Snp Chips, Robert B. Scharpf, Giovanni Parmigiani, Jonathan Pevnser, Ingo Ruczinski

Johns Hopkins University, Dept. of Biostatistics Working Papers

Amplifications and deletions of chromosomal DNA, as well as copy-neutral loss of heterozygosity have been associated with diseases processes. High-throughput single nucleotide polymorphism (SNP) arrays are useful for making genome-wide estimates of copy number and genotype calls. Because neighboring SNPs in high throughput SNP arrays are likely to have dependent copy number and genotype due to the underlying haplotype structure and linkage disequilibrium, hidden Markov models (HMM) may be useful for improving genotype calls and copy number estimates that do not incorporate information from nearby SNPs. We improve previous approaches that utilize a HMM framework for inference in high throughput …


Power Boosting In Genome-Wide Studies Via Methods For Multivariate Outcomes, Mary J. Emond Feb 2007

Power Boosting In Genome-Wide Studies Via Methods For Multivariate Outcomes, Mary J. Emond

UW Biostatistics Working Paper Series

Whole-genome studies are becoming a mainstay of biomedical research. Examples include expression array experiments, comparative genomic hybridization analyses and large case-control studies for detecting polymorphism/disease associations. The tactic of applying a regression model to every locus to obtain test statistics is useful in such studies. However, this approach ignores potential correlation structure in the data that could be used to gain power, particularly when a Bonferroni correction is applied to adjust for multiple testing. In this article, we propose using regression techniques for misspecified multivariate outcomes to increase statistical power over independence-based modeling at each locus. Even when the outcome …


Data Quality Assessment Of Ungated Flow Cytometry Data In High, Nolwenn Le Meur, Anthony Rossini, Maura Gasparetto, Clay Smith, Ryan R. Brinkman, Robert Gentleman Feb 2007

Data Quality Assessment Of Ungated Flow Cytometry Data In High, Nolwenn Le Meur, Anthony Rossini, Maura Gasparetto, Clay Smith, Ryan R. Brinkman, Robert Gentleman

Bioconductor Project Working Papers

Background: The recent development of semi-automated techniques for staining and analyzing flow cytometry samples has presented new challenges. Quality control and quality assessment are critical when developing new high throughput technologies and their associated information services. Our experience suggests that significant bottlenecks remain in the development of high throughput flow cytometry methods for data analysis and display. Especially, data quality control and quality assessment are crucial steps in processing and analyzing high throughput flow cytometry data.

Methods: We propose a variety of graphical exploratory data analytic tools for exploring ungated flow cytometry data. We have implemented a number of specialized …