Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Bioinformatics (398)
- Genomics (222)
- Physical Sciences and Mathematics (208)
- Genetics (172)
- Medicine and Health Sciences (153)
-
- Biochemistry, Biophysics, and Structural Biology (127)
- Molecular Genetics (105)
- Biology (104)
- Statistics and Probability (93)
- Ecology and Evolutionary Biology (87)
- Computer Sciences (86)
- Cell and Developmental Biology (81)
- Molecular Biology (68)
- Evolution (57)
- Microbiology (55)
- Biotechnology (51)
- Systems Biology (51)
- Biochemistry (43)
- Cell Biology (41)
- Microarrays (39)
- Structural Biology (39)
- Biostatistics (38)
- Medical Sciences (38)
- Engineering (37)
- Animal Sciences (33)
- Diseases (33)
- Statistical Methodology (32)
- Institution
-
- COBRA (113)
- The Texas Medical Center Library (49)
- Dartmouth College (47)
- Old Dominion University (40)
- University of Kentucky (32)
-
- City University of New York (CUNY) (31)
- Himmelfarb Health Sciences Library, The George Washington University (27)
- University of Nebraska - Lincoln (26)
- Thomas Jefferson University (20)
- Virginia Commonwealth University (19)
- Louisiana State University (13)
- The University of Southern Mississippi (13)
- California Polytechnic State University, San Luis Obispo (12)
- Clemson University (12)
- Illinois State University (9)
- Loyola University Chicago (9)
- University of Arkansas, Fayetteville (9)
- University of Connecticut (9)
- Augustana College (8)
- University of Louisville (8)
- Mississippi State University (7)
- West Virginia University (7)
- Western University (7)
- Chapman University (6)
- Munster Technological University (6)
- Swarthmore College (6)
- University of Montana (6)
- University of Nevada, Las Vegas (6)
- University of New Mexico (6)
- Harrisburg University of Science and Technology (5)
- Keyword
-
- Bioinformatics (70)
- Genetics (52)
- Computational biology (40)
- Genomics (37)
- Humans (35)
-
- Gene expression (32)
- Algorithms (24)
- Genome (19)
- Animals (18)
- Evolution (18)
- Machine learning (18)
- Models (16)
- Deep learning (15)
- Transcriptomics (15)
- Genetic (13)
- Computational Biology (12)
- Protein (12)
- Phylogeny (11)
- Transcriptome (11)
- Annotation (10)
- Computer simulation (10)
- Epigenetics (10)
- Machine Learning (10)
- Metabolism (10)
- Phylogenetics (10)
- Population genetics (10)
- Systems biology (10)
- Transcription factors (10)
- Cancer (9)
- Sequence analysis (9)
- Publication Year
- Publication
-
- Dissertations and Theses (Open Access) (48)
- Dartmouth Scholarship (36)
- Computer Science Faculty Publications (27)
- Johns Hopkins University, Dept. of Biostatistics Working Papers (27)
- Computational Biology Institute (24)
-
- Harvard University Biostatistics Working Paper Series (24)
- Theses and Dissertations (22)
- Dissertations, Theses, and Capstone Projects (18)
- COBRA Preprint Series (17)
- Electronic Theses and Dissertations (13)
- UW Biostatistics Working Paper Series (12)
- LSU Doctoral Dissertations (11)
- Publications and Research (11)
- Theses and Dissertations--Biology (11)
- Bioconductor Project Working Papers (10)
- Dartmouth College Ph.D Dissertations (10)
- Dissertations (10)
- UPenn Biostatistics Working Papers (10)
- All Dissertations (8)
- Annual Symposium on Biomathematics and Ecology Education and Research (8)
- Computational Medicine Center Faculty Papers (8)
- Honors Theses (8)
- Meiothermus ruber Genome Analysis Project (8)
- STAR Program Research Presentations (8)
- Bioinformatics Faculty Publications (7)
- Biology Faculty Publications (7)
- Department of Pathology, Anatomy, and Cell Biology Faculty Papers (7)
- Master's Theses (7)
- U.C. Berkeley Division of Biostatistics Working Paper Series (7)
- Biochemistry Publications (6)
- Publication Type
- File Type
Articles 721 - 750 of 754
Full-Text Articles in Computational Biology
Bayesian Analysis Of Cell-Cycle Gene Expression Data, Chuan Zhou, Jon Wakefield, Linda Breeden
Bayesian Analysis Of Cell-Cycle Gene Expression Data, Chuan Zhou, Jon Wakefield, Linda Breeden
UW Biostatistics Working Paper Series
The study of the cell-cycle is important in order to aid in our understanding of the basic mechanisms of life, yet progress has been slow due to the complexity of the process and our lack of ability to study it at high resolution. Recent advances in microarray technology have enabled scientists to study the gene expression at the genome-scale with a manageable cost, and there has been an increasing effort to identify cell-cycle regulated genes. In this chapter, we discuss the analysis of cell-cycle gene expression data, focusing on a model-based Bayesian approaches. The majority of the models we describe …
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
Optimal Feature Selection For Nearest Centroid Classifiers, With Applications To Gene Expression Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
Nearest centroid classifiers have recently been successfully employed in high-dimensional applications. A necessary step when building a classifier for high-dimensional data is feature selection. Feature selection is typically carried out by computing univariate statistics for each feature individually, without consideration for how a subset of features performs as a whole. For subsets of a given size, we characterize the optimal choice of features, corresponding to those yielding the smallest misclassification rate. Furthermore, we propose an algorithm for estimating this optimal subset in practice. Finally, we investigate the applicability of shrinkage ideas to nearest centroid classifiers. We use gene-expression microarrays for …
A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey
A New Approach To Intensity-Dependent Normalization Of Two-Channel Microarrays, Alan R. Dabney, John D. Storey
UW Biostatistics Working Paper Series
A two-channel microarray measures the relative expression levels of thousands of genes from a pair of biological samples. In order to reliably compare gene expression levels between and within arrays, it is necessary to remove systematic errors that distort the biological signal of interest. The standard for accomplishing this is smoothing "MA-plots" to remove intensity-dependent dye bias and array-specific effects. However, MA methods require strong assumptions. We review these assumptions and derive several practical scenarios in which they fail. The "dye-swap" normalization method has been much less frequently used because it requires two arrays per pair of samples. We show …
Principal Component Analysis For Predicting Transcription-Factor Binding Motifs From Array-Derived Data, Yunlong Liu, Matthew P Vincenti, Hiroki Yokota
Principal Component Analysis For Predicting Transcription-Factor Binding Motifs From Array-Derived Data, Yunlong Liu, Matthew P Vincenti, Hiroki Yokota
Dartmouth Scholarship
The responses to interleukin 1 (IL-1) in human chondrocytes constitute a complex regulatory mechanism, where multiple transcription factors interact combinatorially to transcription-factor binding motifs (TFBMs). In order to select a critical set of TFBMs from genomic DNA information and an array-derived data, an efficient algorithm to solve a combinatorial optimization problem is required. Although computational approaches based on evolutionary algorithms are commonly employed, an analytical algorithm would be useful to predict TFBMs at nearly no computational cost and evaluate varying modelling conditions. Singular value decomposition (SVD) is a powerful method to derive primary components of a given matrix. Applying SVD …
An Introduction To Low-Level Analysis Methods Of Dna Microarray Data, Wolfgang Huber, Anja Von Heydebreck, Martin Vingron
An Introduction To Low-Level Analysis Methods Of Dna Microarray Data, Wolfgang Huber, Anja Von Heydebreck, Martin Vingron
Bioconductor Project Working Papers
This article gives an overview over the methods used in the low--level analysis of gene expression data generated using DNA microarrays. This type of experiment allows to determine relative levels of nucleic acid abundance in a set of tissues or cell populations for thousands of transcripts or loci simultaneously. Careful statistical design and analysis are essential to improve the efficiency and reliability of microarray experiments throughout the data acquisition and analysis process. This includes the design of probes, the experimental design, the image analysis of microarray scanned images, the normalization of fluorescence intensities, the assessment of the quality of microarray …
Simultaneous And Exact Interval Estimates For The Contrast Of Two Groups Based On An Extremely High Dimensional Response Variable: Application To Mass Spec Data Analysis, Yuhyun Park, Sean R. Downing, Cheng Li Dr., William C. Hahn, Philip W. Kantoff, L. J. Wei
Simultaneous And Exact Interval Estimates For The Contrast Of Two Groups Based On An Extremely High Dimensional Response Variable: Application To Mass Spec Data Analysis, Yuhyun Park, Sean R. Downing, Cheng Li Dr., William C. Hahn, Philip W. Kantoff, L. J. Wei
Harvard University Biostatistics Working Paper Series
No abstract provided.
The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey
The Optimal Discovery Procedure: A New Approach To Simultaneous Significance Testing, John D. Storey
UW Biostatistics Working Paper Series
Significance testing is one of the main objectives of statistics. The Neyman-Pearson lemma provides a simple rule for optimally testing a single hypothesis when the null and alternative distributions are known. This result has played a major role in the development of significance testing strategies that are used in practice. Most of the work extending single testing strategies to multiple tests has focused on formulating and estimating new types of significance measures, such as the false discovery rate. These methods tend to be based on p-values that are calculated from each test individually, ignoring information from the other tests. As …
The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek
The Optimal Discovery Procedure For Large-Scale Significance Testing, With Applications To Comparative Microarray Experiments, John D. Storey, James Y. Dai, Jeffrey T. Leek
UW Biostatistics Working Paper Series
As much of the focus of genetics and molecular biology has shifted toward the systems level, it has become increasingly important to accurately extract biologically relevant signal from thousands of related measurements. The common property among these high-dimensional biological studies is that the measured features have a rich and largely unknown underlying structure. One example of much recent interest is identifying differentially expressed genes in comparative microarray experiments. We propose a new approach aimed at optimally performing many hypothesis tests in a high-dimensional study. This approach estimates the Optimal Discovery Procedure (ODP), which has recently been introduced and theoretically shown …
Metaheuristic Applications And Their Solutions Quality, Dr. Zahid Hussain
Metaheuristic Applications And Their Solutions Quality, Dr. Zahid Hussain
International Conference on Information and Communication Technologies
Over the past few decades, a wide variety of classes of combinatorial problems (e.g. the assignment problem, the knapsack problem, the vehicle routing problem, etc.) have emerged - from such areas as management science, telecommunication, AI, VLSI design and many others. Many large combinatorial problems are NP-hard problems because of the combinatorial growth of their solution search space with the problem size. Such problems are commonly solved by some version of a prominent metaheuristic (e.g. Genetic Algorithms, Tabu Search, Simulated Annealing and etc.). These heuristics seek good but approximate solutions at a reasonable computational cost. These heuristics are of stochastic …
Analysis Of Affymetrix Genechip Data Using Amplified Rna, Leslie Cope, Scott M. Hartman, Hinrich W.H. Gohlmann, Jay P. Tiesman, Rafael A. Irizarry
Analysis Of Affymetrix Genechip Data Using Amplified Rna, Leslie Cope, Scott M. Hartman, Hinrich W.H. Gohlmann, Jay P. Tiesman, Rafael A. Irizarry
Johns Hopkins University, Dept. of Biostatistics Working Papers
The standard method of target synthesis for hybridization to Affymetrix GeneChip® expression microarrays requires a relatively large amount of input total RNA (1-15 micrograms). When small biological samples are collected by microdissection or other methods, amplification techniques are required to provide sufficient target for hybridization to expression arrays. One amplification technique used is to perform two successive rounds of T7-based in vitro transcription. However, the use of random primers required to re-generate cDNA from the first round transcription reaction results in shortened copies of the cDNA, and ultimately the cRNA, transcripts from which the 5' end is missing. In this …
New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski
New Statistical Paradigms Leading To Web-Based Tools For Clinical/Translational Science, Knut M. Wittkowski
COBRA Preprint Series
As the field of functional genetics and genomics is beginning to mature, we become confronted with new challenges. The constant drop in price for sequencing and gene expression profiling as well as the increasing number of genetic and genomic variables that can be measured makes it feasible to address more complex questions. The success with rare diseases caused by single loci or genes has provided us with a proof-of-concept that new therapies can be developed based on functional genomics and genetics.
Common diseases, however, typically involve genetic epistasis, genomic pathways, and proteomic pattern. Moreover, to better understand the underlying biologi-cal …
Cluster Analysis Of Genomic Data With Applications In R, Katherine S. Pollard, Mark J. Van Der Laan
Cluster Analysis Of Genomic Data With Applications In R, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
In this paper, we provide an overview of existing partitioning and hierarchical clustering algorithms in R. We discuss statistical issues and methods in choosing the number of clusters, the choice of clustering algorithm, and the choice of dissimilarity matrix. In particular, we illustrate how the bootstrap can be employed as a statistical method in cluster analysis to establish the reproducibility of the clusters and the overall variability of the followed procedure. We also show how to visualize a clustering result by plotting ordered dissimilarity matrices in R. We present a new R package, hopach, which implements the hybrid clustering method, …
A Digital Atlas To Characterize The Mouse Brain Transcriptome, James P. Carson, Tao Ju, Hui-Chen Lu, Christina Thaller, Mei Xu, Sarah Pallas, Michael C. Crair, Joe Warren, Wah Chiu, Gregor Eichele
A Digital Atlas To Characterize The Mouse Brain Transcriptome, James P. Carson, Tao Ju, Hui-Chen Lu, Christina Thaller, Mei Xu, Sarah Pallas, Michael C. Crair, Joe Warren, Wah Chiu, Gregor Eichele
PCOM Scholarly Works
Massive amounts of data are being generated in an effort to represent for the brain the expression of all genes at cellular resolution. Critical to exploiting this effort is the ability to place these data into a common frame of reference. Here we have developed a computational method for annotating gene expression patterns in the context of a digital atlas to facilitate custom user queries and comparisons of this type of data. This procedure has been applied to 200 genes in the postnatal mouse brain. As an illustration of utility, we identify candidate genes that may be related to Parkinson …
A Brief History Of Bioperl, Colin Crossman, Arti K. Rai
A Brief History Of Bioperl, Colin Crossman, Arti K. Rai
Faculty Scholarship
Large-scale open-source projects face a litany of pitfalls and difficulties. Problems of contribution quality, credit for contributions, project coordination, funding, and mission-creep are ever-present. Of these, long-term funding and project coordination can interact to form a particularly difficult problem for open-source projects in an academic environment.
BioPerl was chosen as an example of a successful academic open-source project. Several of the roadblocks and hurdles encountered and overcome in the development of BioPerl are examined through the telling of the history of the project. Along the way, key points of open-source law are explained, such as license choice and copyright.
The …
Finding Cancer Subtypes In Microarray Data Using Random Projections, Debashis Ghosh
Finding Cancer Subtypes In Microarray Data Using Random Projections, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
One of the benefits of profiling of cancer samples using microarrays is the generation of molecular fingerprints that will define subtypes of disease. Such subgroups have typically been found in microarray data using hierarchical clustering. A major problem in interpretation of the output is determining the number of clusters. We approach the problem of determining disease subtypes using mixture models. A novel estimation procedure of the parameters in the mixture model is developed based on a combination of random projections and the expectation-maximization algorithm. Because the approach is probabilistic, our approach provides a measure for the number of true clusters …
Differential Expression With The Bioconductor Project, Anja Von Heydebreck, Wolfgang Huber, Robert Gentleman
Differential Expression With The Bioconductor Project, Anja Von Heydebreck, Wolfgang Huber, Robert Gentleman
Bioconductor Project Working Papers
A basic, yet challenging task in the analysis of microarray gene expression data is the identification of changes in gene expression that are associated with particular biological conditions. We discuss different approaches to this task and illustrate how they can be applied using software from the Bioconductor Project. A central problem is the high dimensionality of gene expression space, which prohibits a comprehensive statistical analysis without focusing on particular aspects of the joint distribution of the genes expression levels. Possible strategies are to do univariate gene-by-gene analysis, and to perform data-driven nonspecific filtering of genes before the actual statistical analysis. …
Statistical Analyses And Reproducible Research, Robert Gentleman, Duncan Temple Lang
Statistical Analyses And Reproducible Research, Robert Gentleman, Duncan Temple Lang
Bioconductor Project Working Papers
For various reasons, it is important, if not essential, to integrate the computations and code used in data analyses, methodological descriptions, simulations, etc. with the documents that describe and rely on them. This integration allows readers to both verify and adapt the statements in the documents. Authors can easily reproduce them in the future, and they can present the document's contents in a different medium, e.g. with interactive controls. This paper describes a software framework for authoring and distributing these integrated, dynamic documents that contain text, code, data, and any auxiliary content needed to recreate the computations. The documents are …
A Model Based Background Adjustment For Oligonucleotide Expression Arrays, Zhijin Wu, Rafael A. Irizarry, Robert Gentleman, Francisco Martinez Murillo, Forrest Spencer
A Model Based Background Adjustment For Oligonucleotide Expression Arrays, Zhijin Wu, Rafael A. Irizarry, Robert Gentleman, Francisco Martinez Murillo, Forrest Spencer
Johns Hopkins University, Dept. of Biostatistics Working Papers
High density oligonucleotide expression arrays are widely used in many areas of biomedical research. Affymetrix GeneChip arrays are the most popular. In the Affymetrix system, a fair amount of further pre-processing and data reduction occurs following the image processing step. Statistical procedures developed by academic groups have been successful at improving the default algorithms provided by the Affymetrix system. In this paper we present a solution to one of the pre-processing steps, background adjustment, based on a formal statistical framework. Our solution greatly improves the performance of the technology in various practical applications.
Affymetrix GeneChip arrays use short oligonucleotides to …
Reproducible Research: A Bioinformatics Case Study, Robert Gentleman
Reproducible Research: A Bioinformatics Case Study, Robert Gentleman
Bioconductor Project Working Papers
While scientific research and the methodologies involved have gone through substantial technological evolution the technology involved in the publication of the results of these endeavors has remained relatively stagnant. Publication is largely done in the same manner today as it was fifty years ago. Many journals have adopted electronic formats, however, their orientation and style is little different from a printed document. The documents tend to be static and take little advantage of computational resources that might be available. Recent work, Gentleman and Temple Lang (2004), suggests a methodology and basic infrastructure that can be used to publish documents in …
Classification Using Generalized Partial Least Squares, Beiying Ding, Robert Gentleman
Classification Using Generalized Partial Least Squares, Beiying Ding, Robert Gentleman
Bioconductor Project Working Papers
The advances in computational biology have made simultaneous monitoring of thousands of features possible. The high throughput technologies not only bring about a much richer information context in which to study various aspects of gene functions but they also present challenge of analyzing data with large number of covariates and few samples. As an integral part of machine learning, classification of samples into two or more categories is almost always of interest to scientists. In this paper, we address the question of classification in this setting by extending partial least squares (PLS), a popular dimension reduction tool in chemometrics, in …
Mixture Models For Assessing Differential Expression In Complex Tissues Using Microarray Data, Debashis Ghosh
Mixture Models For Assessing Differential Expression In Complex Tissues Using Microarray Data, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
The use of DNA microarrays has become quite popular in many scientific and medical disciplines, such as in cancer research. One common goal of these studies is to determine which genes are differentially expressed between cancer and healthy tissue, or more generally, between two experimental conditions. A major complication in the molecular profiling of tumors using gene expression data is that the data represent a combination of tumor and normal cells. Much of the methodology developed for assessing differential expression with microarray data has assumed that tissue samples are homogeneous. In this article, we outline a general framework for determining …
Bioconductor: Open Software Development For Computational Biology And Bioinformatics, Robert C. Gentleman, Vincent J. Carey, Douglas J. Bates, Benjamin M. Bolstad, Marcel Dettling, Sandrine Dudoit, Byron Ellis, Laurent Gautier, Yongchao Ge, Jeff Gentry, Kurt Hornik, Torsten Hothorn, Wolfgang Huber, Stefano Iacus, Rafael Irizarry, Friedrich Leisch, Cheng Li, Martin Maechler, Anthony J. Rossini, Guenther Sawitzki, Colin Smith, Gordon K. Smyth, Luke Tierney, Yee Hwa Yang, Jianhua Zhang
Bioconductor: Open Software Development For Computational Biology And Bioinformatics, Robert C. Gentleman, Vincent J. Carey, Douglas J. Bates, Benjamin M. Bolstad, Marcel Dettling, Sandrine Dudoit, Byron Ellis, Laurent Gautier, Yongchao Ge, Jeff Gentry, Kurt Hornik, Torsten Hothorn, Wolfgang Huber, Stefano Iacus, Rafael Irizarry, Friedrich Leisch, Cheng Li, Martin Maechler, Anthony J. Rossini, Guenther Sawitzki, Colin Smith, Gordon K. Smyth, Luke Tierney, Yee Hwa Yang, Jianhua Zhang
Bioconductor Project Working Papers
The Bioconductor project is an initiative for the collaborative creation of extensible software for computational biology and bioinformatics. We detail some of the design decisions, software paradigms and operational strategies that have allowed a small number of researchers to provide a wide variety of innovative, extensible, software solutions in a relatively short time. The use of an object oriented programming paradigm, the adoption and development of a software package system, designing by contract, distributed development and collaboration with other projects are elements of this project's success. Individually, each of these concepts are useful and important but when combined they have …
Smart Sequence Similarity Search (S⁴) System, Zhuo Chen
Smart Sequence Similarity Search (S⁴) System, Zhuo Chen
Theses Digitization Project
Sequence similarity searching is commonly used to help clarify the biochemical and physiological features of newly discovered genes or proteins. An efficient similarity search relies on the choice of tools and their associated subprograms and numerous parameter settings. To assist researchers in selecting optimal programs and parameter settings for efficient sequence similarity searches, the web-based expert system, Smart Sequence Similarity Search (S4) was developed.
Computational Protein Biomarker Prediction: A Case Study For Prostate Cancer, Michael Wagner, Dayanand N. Naik, Alex Pothen, Srinivas Kasukurti, Raghu Ram Devineni, Bao-Ling Adam, O. John Semmes, George L. Wright Jr.
Computational Protein Biomarker Prediction: A Case Study For Prostate Cancer, Michael Wagner, Dayanand N. Naik, Alex Pothen, Srinivas Kasukurti, Raghu Ram Devineni, Bao-Ling Adam, O. John Semmes, George L. Wright Jr.
Mathematics & Statistics Faculty Publications
Background: Recent technological advances in mass spectrometry pose challenges in computational mathematics and statistics to process the mass spectral data into predictive models with clinical and biological significance. We discuss several classification-based approaches to finding protein biomarker candidates using protein profiles obtained via mass spectrometry, and we assess their statistical significance. Our overall goal is to implicate peaks that have a high likelihood of being biologically linked to a given disease state, and thus to narrow the search for biomarker candidates.
Results: Thorough cross-validation studies and randomization tests are performed on a prostate cancer dataset with over 300 patients, obtained …
Cluster Stability Scores For Microarray Data In Cancer Studies, Mark Smolkin, Debashis Ghosh
Cluster Stability Scores For Microarray Data In Cancer Studies, Mark Smolkin, Debashis Ghosh
The University of Michigan Department of Biostatistics Working Paper Series
A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will define subtypes of disease. Hierarchical clustering has been the primary analytical tool used to define disease subtypes from microarray experiments in cancer settings. Assessing cluster reliability poses a major complication in analyzing output from these procedures. While much work has been done on assessing the global question of number of clusters in a dataset, relatively little research exists on assessing stability of individual clusters. A potential benefit of profiling of tissue samples using microarrays is the generation of molecular fingerprints that will …
Simple Parallel Statistical Computing In R, Anthony Rossini, Luke Tierney, Na Li
Simple Parallel Statistical Computing In R, Anthony Rossini, Luke Tierney, Na Li
UW Biostatistics Working Paper Series
Theoretically, many modern statistical procedures are trivial to parallelize. However, practical deployment of a parallelized implementation which is robust and reliably runs on different computational cluster configurations and environments is far from trivial. We present a framework for the R statistical computing language that provides a simple yet powerful programming interface to a computational cluster. This interface allows the development of R functions that distribute independent computations across the nodes of the computational cluster. The resulting framework allows statisticians to obtain significant speed-ups for some computations at little additional development cost. The particular implementation can be deployed in heterogeneous computing …
Literate Statistical Practice, Anthony Rossini, Friedrich Leisch
Literate Statistical Practice, Anthony Rossini, Friedrich Leisch
UW Biostatistics Working Paper Series
Literate Statistical Practice (LSP, Rossini, 2001) describes an approach for creating self-documenting statistical results. It applies literate programming (Knuth, 1992) and related techniques in a natural fashion to the practice of statistics. In particular, documentation, specification, and descriptions of results are written concurrently with writing and evaluation of statistical programs. We discuss how and where LSP can be integrated into practice and illustrate this with an example derived from an actual statistical consulting project. The approach is simplified through the use of a comprehensive, open source toolset incorporating Noweb, Emacs Speaks Statistics (ESS), Sweave (Ramsey, 1994; Rossini, et al, 2002; …
A Neural Network Method For Protein Classification, Arie D. Jones
A Neural Network Method For Protein Classification, Arie D. Jones
All-Inclusive List of Electronic Theses and Dissertations
The investigation detailed in this paper attempts to utilize a Leaming Vector Quantization network in order to classify a set of mitochondrial proteins based upon their amino acid sequence. The Learning Vector Quantization network uses a nearest neighbor approach to classification. Input vectors are fed into the network, which produces output vectors. Those output vectors are then matched by means of a distance bias to a corresponding classification vector. The learning and test sets consisted of thirteen similar mitochodrial proteins from seventy-two different species. This provided a pool of over nine hundred proteins to use. Half of the species were …
Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
Statistical Inference For Simultaneous Clustering Of Gene Expression Data, Katherine S. Pollard, Mark J. Van Der Laan
U.C. Berkeley Division of Biostatistics Working Paper Series
Current methods for analysis of gene expression data are mostly based on clustering and classification of either genes or samples. We offer support for the idea that more complex patterns can be identified in the data if genes and samples are considered simultaneously. We formalize the approach and propose a statistical framework for two-way clustering. A simultaneous clustering parameter is defined as a function of the true data generating distribution, and an estimate is obtained by applying this function to the empirical distribution. We illustrate that a wide range of clustering procedures, including generalized hierarchical methods, can be defined as …
Studies On The Formation Of Dna-Cationic Lipid Composite Films And Dna Hybridization In The Composites, Murali Sastry, Vidya Ramakrishnan, Mrunalini Pattarkine, Krishna N. Ganesh
Studies On The Formation Of Dna-Cationic Lipid Composite Films And Dna Hybridization In The Composites, Murali Sastry, Vidya Ramakrishnan, Mrunalini Pattarkine, Krishna N. Ganesh
Faculty Works
The formation of composite films of double-stranded DNA and cationic lipid molecules (octadecylamine, ODA) and the hybridization of complementary single-stranded DNA molecules in such composite films are demonstrated. The immobilization of DNA is accomplished by simple immersion of a thermally evaporated ODA film in the DNA solution at close to physiological pH. The entrapment of the DNA molecules in the cationic lipid film is dominated by attractive electrostatic interaction between the negatively charged phosphate backbone of the DNA molecules and the protonated amine molecules in the thermally evaporated film and has been quantified using quartz crystal microgravimetry (QCM). Fluorescence studies …