Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- COBRA (113)
- The Texas Medical Center Library (36)
- Dartmouth College (20)
- University of Nebraska - Lincoln (17)
- University of Kentucky (15)
-
- City University of New York (CUNY) (9)
- Virginia Commonwealth University (9)
- Louisiana State University (8)
- The University of Southern Mississippi (8)
- University of Louisville (8)
- Loyola University Chicago (7)
- Augustana College (6)
- Illinois State University (6)
- University of Connecticut (6)
- Western University (6)
- Chapman University (5)
- Mississippi State University (5)
- University of Montana (5)
- California Polytechnic State University, San Luis Obispo (4)
- Clemson University (4)
- Michigan Technological University (4)
- Nova Southeastern University (4)
- University of Nebraska Medical Center (4)
- West Virginia University (4)
- Claremont Colleges (3)
- Duquesne University (3)
- Kennesaw State University (3)
- LSU New Orleans (3)
- Munster Technological University (3)
- Old Dominion University (3)
- Keyword
-
- Bioinformatics (56)
- Genetics (27)
- Gene expression (19)
- Genomics (15)
- Computational biology (13)
-
- Genome (12)
- Machine learning (11)
- Algorithms (9)
- Annotation (8)
- Evolution (8)
- Phylogenetics (8)
- Cancer genomics (7)
- Protein (7)
- Transcriptomics (7)
- Epigenetics (6)
- Machine Learning (6)
- Meiothermus ruber (6)
- RNA-seq (6)
- Systems biology (6)
- Transcriptome (6)
- CRISPR-Cas (5)
- Chemistry (5)
- Deep learning (5)
- Humans (5)
- Population genetics (5)
- Sequencing (5)
- Biomarkers (4)
- Cancer (4)
- Clustering (4)
- Computational Biology (4)
- Publication Year
- Publication
-
- Dissertations and Theses (Open Access) (36)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (27)
- Harvard University Biostatistics Working Paper Series (24)
- COBRA Preprint Series (17)
- Theses and Dissertations (15)
-
- Dartmouth Scholarship (12)
- UW Biostatistics Working Paper Series (12)
- Bioconductor Project Working Papers (10)
- Electronic Theses and Dissertations (10)
- UPenn Biostatistics Working Papers (10)
- Dissertations, Theses, and Capstone Projects (8)
- Dartmouth College Ph.D Dissertations (7)
- U.C. Berkeley Division of Biostatistics Working Paper Series (7)
- Dissertations (6)
- LSU Doctoral Dissertations (6)
- Meiothermus ruber Genome Analysis Project (6)
- Theses and Dissertations--Biology (6)
- Annual Symposium on Biomathematics and Ecology Education and Research (5)
- Biochemistry Publications (5)
- Bioinformatics Faculty Publications (5)
- Graduate Student Theses, Dissertations, & Professional Papers (5)
- All Dissertations (4)
- Dissertations, Master's Theses and Master's Reports (4)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (4)
- Honors Scholar Theses (4)
- Master's Theses (4)
- Theses & Dissertations (4)
- Theses and Dissertations--Computer Science (4)
- All HCAS Student Capstones, Theses, and Dissertations (3)
- Biology Faculty Publications (3)
- Publication Type
Articles 391 - 398 of 398
Full-Text Articles in Computational Biology
Classification Using Generalized Partial Least Squares, Beiying Ding, Robert Gentleman
Classification Using Generalized Partial Least Squares, Beiying Ding, Robert Gentleman
Bioconductor Project Working Papers
The advances in computational biology have made simultaneous monitoring of thousands of features possible. The high throughput technologies not only bring about a much richer information context in which to study various aspects of gene functions but they also present challenge of analyzing data with large number of covariates and few samples. As an integral part of machine learning, classification of samples into two or more categories is almost always of interest to scientists. In this paper, we address the question of classification in this setting by extending partial least squares (PLS), a popular dimension reduction tool in chemometrics, in …
Mixture Models For Assessing Differential Expression In Complex Tissues Using Microarray Data, Debashis Ghosh
Mixture Models For Assessing Differential Expression In Complex Tissues Using Microarray Data, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
The use of DNA microarrays has become quite popular in many scientific and medical disciplines, such as in cancer research. One common goal of these studies is to determine which genes are differentially expressed between cancer and healthy tissue, or more generally, between two experimental conditions. A major complication in the molecular profiling of tumors using gene expression data is that the data represent a combination of tumor and normal cells. Much of the methodology developed for assessing differential expression with microarray data has assumed that tissue samples are homogeneous. In this article, we outline a general framework for determining …
Bioconductor: Open Software Development For Computational Biology And Bioinformatics, Robert C. Gentleman, Vincent J. Carey, Douglas J. Bates, Benjamin M. Bolstad, Marcel Dettling, Sandrine Dudoit, Byron Ellis, Laurent Gautier, Yongchao Ge, Jeff Gentry, Kurt Hornik, Torsten Hothorn, Wolfgang Huber, Stefano Iacus, Rafael Irizarry, Friedrich Leisch, Cheng Li, Martin Maechler, Anthony J. Rossini, Guenther Sawitzki, Colin Smith, Gordon K. Smyth, Luke Tierney, Yee Hwa Yang, Jianhua Zhang
Bioconductor: Open Software Development For Computational Biology And Bioinformatics, Robert C. Gentleman, Vincent J. Carey, Douglas J. Bates, Benjamin M. Bolstad, Marcel Dettling, Sandrine Dudoit, Byron Ellis, Laurent Gautier, Yongchao Ge, Jeff Gentry, Kurt Hornik, Torsten Hothorn, Wolfgang Huber, Stefano Iacus, Rafael Irizarry, Friedrich Leisch, Cheng Li, Martin Maechler, Anthony J. Rossini, Guenther Sawitzki, Colin Smith, Gordon K. Smyth, Luke Tierney, Yee Hwa Yang, Jianhua Zhang
Bioconductor Project Working Papers
The Bioconductor project is an initiative for the collaborative creation of extensible software for computational biology and bioinformatics. We detail some of the design decisions, software paradigms and operational strategies that have allowed a small number of researchers to provide a wide variety of innovative, extensible, software solutions in a relatively short time. The use of an object oriented programming paradigm, the adoption and development of a software package system, designing by contract, distributed development and collaboration with other projects are elements of this project's success. Individually, each of these concepts are useful and important but when combined they have …
Cluster Stability Scores For Microarray Data In Cancer Studies, Mark Smolkin, Debashis Ghosh
Cluster Stability Scores For Microarray Data In Cancer Studies, Mark Smolkin, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will define subtypes of disease. Hierarchical clustering has been the primary analytical tool used to define disease subtypes from microarray experiments in cancer settings. Assessing cluster reliability poses a major complication in analyzing output from these procedures. While much work has been done on assessing the global question of number of clusters in a dataset, relatively little research exists on assessing stability of individual clusters. A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will …
Simple Parallel Statistical Computing In R, Anthony Rossini, Luke Tierney, Na Li
Simple Parallel Statistical Computing In R, Anthony Rossini, Luke Tierney, Na Li
UW Biostatistics Working Paper Series
Theoretically, many modern statistical procedures are trivial to parallelize. However, practical deployment of a parallelized implementation which is robust and reliably runs on different computational cluster configurations and environments is far from trivial. We present a framework for the R statistical computing language that provides a simple yet powerful programming interface to a computational cluster. This interface allows the development of R functions that distribute independent computations across the nodes of the computational cluster. The resulting framework allows statisticians to obtain significant speed-ups for some computations at little additional development cost. The particular implementation can be deployed in heterogeneous computing …
Literate Statistical Practice, Anthony Rossini, Friedrich Leisch
Literate Statistical Practice, Anthony Rossini, Friedrich Leisch
UW Biostatistics Working Paper Series
Literate Statistical Practice (LSP, Rossini, 2001) describes an approach for creating self-documenting statistical results. It applies literate programming (Knuth, 1992) and related techniques in a natural fashion to the practice of statistics. In particular, documentation, specification, and descriptions of results are written concurrently with writing and evaluation of statistical programs. We discuss how and where LSP can be integrated into practice and illustrate this with an example derived from an actual statistical consulting project. The approach is simplified through the use of a comprehensive, open source toolset incorporating Noweb, Emacs Speaks Statistics (ESS), Sweave (Ramsey, 1994; Rossini, et al, 2002; …
A Neural Network Method For Protein Classification, Arie D. Jones
A Neural Network Method For Protein Classification, Arie D. Jones
All-Inclusive List of Electronic Theses and Dissertations
The investigation detailed in this paper attempts to utilize a Leaming Vector Quantization network in order to classify a set of mitochondrial proteins based upon their amino acid sequence. The Learning Vector Quantization network uses a nearest neighbor approach to classification. Input vectors are fed into the network, which produces output vectors. Those output vectors are then matched by means of a distance bias to a corresponding classification vector. The learning and test sets consisted of thirteen similar mitochodrial proteins from seventy-two different species. This provided a pool of over nine hundred proteins to use. Half of the species were …
Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Current methods for analysis of gene expression data are mostly based on clustering and classification of either genes or samples. We offer support for the idea that more complex patterns can be identified in the data if genes and samples are considered simultaneously. We formalize the approach and propose a statistical framework for two-way clustering. A simultaneous clustering parameter is defined as a function of the true data generating distribution, and an estimate is obtained by applying this function to the empirical distribution. We illustrate that a wide range of clustering procedures, including generalized hierarchical methods, can be defined as …