Open Access. Powered by Scholars. Published by Universities.®

Articles 691 - 708 of 708

Full-Text Articles in Numerical Analysis and Scientific Computing

Rattle: A Data Mining Gui For R, Graham J. Williams Dec 2009

Rattle: A Data Mining Gui For R, Graham J. Williams

The R Journal

Data mining delivers insights, pat terns, and descriptive and predictive models from the large amounts of data available today in many organisations. The data miner draws heavily on methodologies, techniques and algorithms from statistics, machine learning, and computer science. R increasingly provides a powerful platform for data mining. However, scripting and programming is sometimes a challenge for data analysts moving into data mining. The Rattle package provides a graphical user interface specifically for data mining using R. It also provides a stepping stone toward using R as a programming language for data analysis.


Copas: An R Package For Fitting The Copas Selection Model, J. Carpenter, G. Rücker, G. Schhwarzer Dec 2009

Copas: An R Package For Fitting The Copas Selection Model, J. Carpenter, G. Rücker, G. Schhwarzer

The R Journal

This article describes the R package copas which is an add-on package to the R pack age meta. The R package copas can be used to f it the Copas selection model to adjust for bias in meta-analysis. A clinical example is used to illustrate fitting and interpreting the Copas selection model.


Party On!, Carolin Strobl, Torsten Hothorn, Achim Zeileis Dec 2009

Party On!, Carolin Strobl, Torsten Hothorn, Achim Zeileis

The R Journal

Random forests are one of the most popular statistical learning algorithms, and a variety of methods for fitting random forests and related recursive partitioning approaches is available in R. This paper points out two important features of the random forest implementation cforest available in the party package: The resulting forests are unbiased and thus prefer able to the randomForest implementation avail able in randomForest if predictor variables are of different types. Moreover, a conditional per mutation importance measure has recently been added to the party package, which can help evaluate the importance of correlated predictor variables. The rationale of this …


Aspects Of The Social Organization And Trajectory Of The R Project, John Fox Dec 2009

Aspects Of The Social Organization And Trajectory Of The R Project, John Fox

The R Journal

Based partly on interviews with members of the R Core team, this paper considers the development of the R Project in the context of open-source software development and, more generally, voluntary activities. The paper de scribes aspects of the social organization of the R Project, including the organization of the R Core team; describes the trajectory of the R Project; seeks to identify factors crucial to the success of R; and speculates about the prospects for R.


Asymptest: A Simple R Package For Classical Parametric Statistical Tests And Confidence Intervals In Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye De Micheaux, J.-F. Robineau Dec 2009

Asymptest: A Simple R Package For Classical Parametric Statistical Tests And Confidence Intervals In Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye De Micheaux, J.-F. Robineau

The R Journal

asympTest is an R package implementing large sample tests and confidence intervals. One and two sample mean and variance tests (differences and ratios) are considered. The test statistics are all expressed in the same form as the Student t-test, which facilitates their presentation in the classroom. This contribution also fills the gap of a robust (to non-normality) alternative to the chi-square single variance test for large samples, since no such procedure is implemented in standard statistical software.


Convergenceconcepts: An R Package To Investigate Various Modes Of Convergence, Pierre Lafaye De Micheaux, Benoit Liquet Dec 2009

Convergenceconcepts: An R Package To Investigate Various Modes Of Convergence, Pierre Lafaye De Micheaux, Benoit Liquet

The R Journal

ConvergenceConcepts is an R pack age, built upon the tkrplot, tcltk and lattice packages, designed to investigate the convergence of simulated sequences of random variables. Four classical modes of convergence may be studied, namely: almost sure convergence (a.s.), convergence in probability (P), convergence in law (L) and convergence in r-th mean (r). This investigation is performed through ac curate graphical representations. This package may be used as a pedagogical tool. It may give students a better understanding of these notions and help them to visualize these difficult theoretical concepts. Moreover, …


The R Journal (December 2009) 1(2): Complete Issue, The R Foundation Dec 2009

The R Journal (December 2009) 1(2): Complete Issue, The R Foundation

The R Journal

Contributed Research Articles

Aspects of the Social Organization and Trajectory of the R Project, John Fox

Party on! Carolin Strobl, Torsten Hothorn and Achim Zeileis

ConvergenceConcepts: An R Package to Investigate Various Modes of Convergence, Pierre Lafaye de Micheaux and Benoit Liquet

asympTest: A Simple R Package for Classical Parametric Statistical Tests and Confidence Intervals in Large Samples, J.-F. Coeurjolly, R. Drouilhet, P. Lafaye de Micheaux and J.-F. Robineau

copas: An R package for Fitting the Copas Selection Model, J. Carpenter, G. Rücker and G. Schwarzer

Transitioning to R: Replicating SAS, Stata, and SUDAAN Analysis Techniques in Health Policy Data, …


Sample Size Estimation While Controlling False Discovery Rate For Microarray Experiments Using The Ssize.Fdr Package, Megan Orr, Peng Liu Jun 2009

Sample Size Estimation While Controlling False Discovery Rate For Microarray Experiments Using The Ssize.Fdr Package, Megan Orr, Peng Liu

The R Journal

Microarray experiments are becoming more and more popular and critical in many biological disciplines. As in any statistical experiment, appropriate experimental design is essential for reliable statistical inference, and sample size has a crucial role in experimental design. Because microarray experiments are rather costly, it is important to have an adequate sample size that will achieve a desired power with out wasting resources.

For a given microarray data set, thousands of hypotheses, one for each gene, are simultaneously tested. Storey and Tibshirani (2003) argue that con trolling false discovery rate (FDR) is more reasonable and more powerful than controlling family-wise …


Emd: A Package For Empirical Mode Decomposition And Hilbert Spectrum, Donghoh Kim, Hee-Seok Oh Jun 2009

Emd: A Package For Empirical Mode Decomposition And Hilbert Spectrum, Donghoh Kim, Hee-Seok Oh

The R Journal

The concept of empirical mode decomposition (EMD)and the Hilber tspectrum (HS) has been developed rapidly in many disciplines of science and engineering since Huang et al. (1998) invented EMD. The key feature of EMD is to decompose a signal into so-called intrinsic mode function (IMF). Further more, the Hilbert spectral analysis of intrinsic mode functions provides frequency information evolving with time and quantifies the amount of variation due to oscillation at different time scales and time locations. In this article,we introduce an R package called EMD (KimandOh, 2008) that performs one and two-dimensional EMD and HS.


Modeling Without Data Using Expert Opinion, Vincent Goulet, Michel Jacques, Mathieu Pigeon Jun 2009

Modeling Without Data Using Expert Opinion, Vincent Goulet, Michel Jacques, Mathieu Pigeon

The R Journal

The expert package provides tools to create and manipulate empirical statistical models using expert opinion (or judgment). Here, the latter expression refers to a specific body of techniques to elicit the distribution of a random variable when data is scarce or unavailable. Opinions on the quantiles of the distribution are sought from experts in the field and aggregated into a final estimate. The package supports aggregation by means of the Cooke, Mendel–Sheridan and predefined weights models.

We do not mean to give a complete introduction to the theory and practice of expert opinion elicitation in this paper. However, for the …


Admit, David Ardia, Lennart F. Hoogerheide, Herman K. Van Dijk Jun 2009

Admit, David Ardia, Lennart F. Hoogerheide, Herman K. Van Dijk

The R Journal

This note presents the package AdMit (Ardia et al., 2008, 2009), an R implementation of the adaptive mixture of Student-t distributions (AdMit) procedure developed by Hoogerheide (2006); see also Hoogerheide et al. (2007); Hoogerheide and van Dijk (2008). The AdMit strategy consists of the construction of a mixture of Student-t distributions which approximates a target distribution of interest. The fitting procedure relies only on a kernel of the tar get density, so that the normalizing constant is not required. In a second step, this approximation is used as an importance function in importance sampling or as a candidate density in …


The Hwriter Package: Composing Html Documents With R Objects, Gregoire Pau, Wolfgang Huber Jun 2009

The Hwriter Package: Composing Html Documents With R Objects, Gregoire Pau, Wolfgang Huber

The R Journal

HTML documents are structured documents made of diverse elements such as paragraphs, sections, columns, figures and tables organized in a hierarchical layout. Combination of HTML documents and hyperlinking is useful to report analysis results; for example, in the package array Quality Metrics (Kauffmannetal., 2009), estimating the quality of mi croarray data sets and cellHTS2(Boutrosetal.,2006), performing the analysis of cell-based screens.

There are several tools for exporting data from R into HTML documents. The package R2HTML is able to render a large diversity of R objects in HTML but does not easily support combining them in a structured layout and …


Collaborative Software Development Using R-Forge, Stefan Theußl, Achim Zeileis Jun 2009

Collaborative Software Development Using R-Forge, Stefan Theußl, Achim Zeileis

The R Journal

Open source software (OSS) is typically created in a decentralized self-organizing process by a community of developers having the same or similar interests (see the famous essay by Raymond, 1999). A key factor for the success of OSS over the last two decades is the Internet: Developers who rarely meet face-to-face can employ new means of communication, both for rapidly writing and deploying software (in the spirit of Linus Torvald’s “release early, release often paradigm”). Therefore, many tools emerged that assist a collaborative software development process, including in particular tools for source code management (SCM) and version control.

In the …


Facets Of R, John M. Chambers Jun 2009

Facets Of R, John M. Chambers

The R Journal

We are seeing today a widespread, and welcome, tendency for non-computer-specialists among statisticians and others to write collections of R functions that organize and communicate their work. Along with the flood of software sometimes comes an attitude that one need-only-learn, or teach, a sort of basic how-to-write-the-function level of R programming, beyond which most of the detail is unimportant or can be absorbed without much discussion. As delusions go, this one is not very objectionable if it encourages participation. Nevertheless, a delusion it is. In fact, functions are only one of a variety of important facets that R has acquired …


Easier Parallel Computing In R With Snowfall And Sfcluster, Jochen Knaus, Christine Porzelius, Harald Binder, Guido Schwarzer Jun 2009

Easier Parallel Computing In R With Snowfall And Sfcluster, Jochen Knaus, Christine Porzelius, Harald Binder, Guido Schwarzer

The R Journal

Many statistical analysis tasks in areas such as bioinformatics are computationally very intensive, while lots of them rely on embarrassingly parallel computations (Grama et al., 2003). Multiple computers or even multiple processor cores on standard desktop computers, which are widespread nowadays, can easily contribute to faster analyses.

R itself does not allow parallel execution. There are some existing solutions for R to distribute calculations over many computers — a cluster — for ex ample Rmpi, rpvm, snow, nws or papply. However these solutions require the user to setup and manage the cluster on his own and …


The R Journal (June 2009) 1(1): Complete Issue, The R Foundation Jun 2009

The R Journal (June 2009) 1(1): Complete Issue, The R Foundation

The R Journal

Contributed Research Articles

Facets of R, John M. Chambers

Collaborative Software Development Using R-Forge, Stefan Theußl and Achim Zeileis

Drawing Diagrams with R, Paul Murrell

The hwriter package: Composing HTML Documents with R Objects, Gregoire Pau and Wolfgang Huber

AdMit, David Ardia, Lennart F. Hoogerheide, and Herman K. van Dijk

expert: Modeling Without Data Using Expert Opinion, Vincent Goulet, Michel Jacques, and Mathieu Pigeon

New Numerical Algorithm for Multivariate Normal Probabilities in Package mvtnorm, Xuefei Mi, Tetsuhisa Miwa, and Torsten Hothorn

EMD: A Package for Empirical Mode Decomposition and Hilbert Spectrum, Donghoh Kim, and Hee-Seok Oh

Sample Size Estimation while …


Drawing Diagrams With R, Paul Murrell Jun 2009

Drawing Diagrams With R, Paul Murrell

The R Journal

R provides a number of well-known high-level facilities for producing sophisticated statistical plots, including the “traditional” plots in the graphics pack age (R Development Core Team, 2008), the Trellis style plots provided by lattice (Sarkar, 2008), and the grammar-of-graphics-inspired approach of ggplot2 (Wickham, 2009). However, R also provides a powerful set of low level graphics facilities for drawing basic shapes and, more importantly, for arranging those shapes relative to each other, which can be used to draw a wide variety of graphical images. This article highlights some of R’s low-level graphics facilities by demonstrating their use in the production of …


Pmml: An Open Standard For Sharing Models, Alex Guazzelli, Michael Zeller, Wen-Ching Lin, Graham Williams Jun 2009

Pmml: An Open Standard For Sharing Models, Alex Guazzelli, Michael Zeller, Wen-Ching Lin, Graham Williams

The R Journal

The PMML package exports a variety of predictive and descriptive models from R to the Predictive Model Markup Language (Data Mining Group, 2008). PMML is an XML-based language and has become the de-facto standard to represent not only predictive and descriptive models, but also data pre- and post-processing. In so doing, it allows for the interchange of models among different tools and environments, mostly avoiding proprietary issues and incompatibilities.

The PMML package itself (Williams et al., 2009) was conceived at first as part of Togaware’s data mining toolkit Rattle, the R Analytical Tool To Learn Easily (Williams, 2009). Although it …