Measurement Errors In R,
2018
Universidad Carlos III de Madrid
Measurement Errors In R, Iñaki Ucar, Edzer Pebesma, Arturo Azcorra
The R Journal
This paper presents an R package to handle and represent measurements with errors in a very simple way. We briefly introduce the main concepts of metrology and propagation of uncertainty, and discuss related R packages. Building upon this, we introduce the errors package, which provides a class for associating uncertainty metadata, automated propagation and reporting. Working with errors enables transparent, lightweight, less error-prone handling and convenient representation of measurements with errors. Finally, we discuss the advantages, limitations and future work of computing with errors.
Spatial Uncertainty Propagation Analysis With The Spup R Package,
2018
Wageningen University
Spatial Uncertainty Propagation Analysis With The Spup R Package, Kasia Sawicka, Gerard B.M. Heuvelink, Dennis J.J. Walvoort
The R Journal
Many environmental and geographical models, such as those used in land degradation, agroecological and climate studies, make use of spatially distributed inputs that are known imperfectly. The R package spup provides functions for examining the uncertainty propagation from input data and model parameters onto model outputs via the environmental model. The functions include uncertainty model specification, stochastic simulation and propagation of uncertainty using Monte Carlo (MC) techniques. Uncertain variables are described by probability distributions. Both numerical and categorical data types are handled. The package also accommodates spatial auto-correlation within a variable and cross-correlation between variables. The MC realizations may be …
The Utiml Package: Multi-Label Classification In R,
2018
Federal University of Technology - Parana (UTFPR)
The Utiml Package: Multi-Label Classification In R, Adriano Rivolli, Andre C.P.L.F. De Carvalho
The R Journal
Learning classification tasks in which each instance is associated with one or more labels are known as multi-label learning. The implementation of multi-label algorithms, performed by different researchers, have several specificities, like input/output format, different internal functions, distinct programming language, to mention just some of them. As a result, current machine learning tools include only a small subset of multi-label decomposition strategies. The utiml package is a framework for the application of classification algorithms to multi-label data. Like the well known MULAN used with Weka, it provides a set of multi-label procedures such as sampling methods, transformation strategies, threshold functions, …
Nsroc: An R Package For Non-Standard Roc Curve Analysis,
2018
University of Oviedo
Nsroc: An R Package For Non-Standard Roc Curve Analysis, Sonia Pérez-Fernández, Pablo Martínez-Camblor, Peter Filzmoser, Norberto Corral
The R Journal
The receiver operating characteristic (ROC) curve is a graphical method which has become standard in the analysis of diagnostic markers, that is, in the study of the classification ability of a numerical variable. Most of the commercial statistical software provide routines for the standard ROC curve analysis. Of course, there are also many R packages dealing with the ROC estimation as well as other related problems. In this work we introduce the nsROC package which incorporates some new ROC curve procedures. Particularly: ROC curve comparison based on general distances among functions for both paired and unpaired designs; efficient confidence bands …
Stilt: Easy Emulation Of Time Series Ar(1) Computer Model Output In Multidimensional Parameter Space,
2018
Yonsei University, Center for Climate Physics, Pusan National University
Stilt: Easy Emulation Of Time Series Ar(1) Computer Model Output In Multidimensional Parameter Space, Roman Olson, Kelsey L. Ruckert, Won Chang, Klaus Keller, Murali Haran, Soon-Il An
The R Journal
Statistically approximating or “emulating” time series model output in parameter space is a common problem in climate science and other fields. There are many packages for spatio-temporal modeling. However, they often lack focus on time series, and exhibit statistical complexity. Here, we present the R package stilt designed for simplified AR(1) time series Gaussian process emulation, and provide examples relevant to climate modelling. Notably absent is Markov chain Monte Carlo estimation – a challenging concept to many scientists. We keep the number of user choices to a minimum. Hence, the package can be useful pedagogically, while still applicable to real …
R Foundation News,
2018
Universität Zürich
R Foundation News, Torsten Hothorn
The R Journal
Donations and members
Donations
Supporting benefactors
Supporting institutions
Supporting members
Sarima Analysis And Automated Model Reports With Bets, An R Package,
2018
Instituto Brasileiro de Economia (IBRE)
Sarima Analysis And Automated Model Reports With Bets, An R Package, Talitha F. Speranza, Pedro C. Ferreira, Jonatha A. Da Costa
The R Journal
This article aims to demonstrate how the powerful features of the R package BETS can be applied to SARIMA time series analysis. BETS provides not only thousands of Brazilian economic time series from different institutions, but also a range of analytical tools, and educational resources. In particular, BETS is capable of generating automated model reports for any given time series. These reports rely on a single function call and are able to build three types of models (SARIMA being one of them). The functions need few inputs and output rich content. The output varies according to the inputs and usually …
Profile Likelihood Estimation Of The Correlation Coefficient In The Presence Of Left, Right Or Interval Censoring And Missing Data,
2018
University of Michigan-Ann Arbor
Profile Likelihood Estimation Of The Correlation Coefficient In The Presence Of Left, Right Or Interval Censoring And Missing Data, Yanming Li, Brenda W. Gillespie, Kerby Shedden, John A. Gillespie
The R Journal
We discuss implementation of a profile likelihood method for estimating a Pearson correlation coefficient from bivariate data with censoring and/or missing values. The method is implemented in an R package clikcorr which calculates maximum likelihood estimates of the correlation coefficient when the data are modeled with either a Gaussian or a Student t-distribution, in the presence of left, right, or interval censored and/or missing data. The R package includes functions for conducting inference and also provides graphical functions for visualizing the censored data scatter plot and profile log likelihood function. The performance of clikcorr in a variety of circumstances is …
Dot-Pipe: An S3 Extensible Pipe For R,
2018
Win-Vector LLC
Dot-Pipe: An S3 Extensible Pipe For R, John Mount, Nina Zumel
The R Journal
Pipe notation is popular with a large league of R users, with magrittr being the dominant realization. However, this should not be enough to consider piping in R as a settled topic that is not subject to further discussion, experimentation, or possibility for improvement. To promote innovation opportunities, we describe the wrapr R package and “dot-pipe” notation, a well behaved sequencing operator with S3 extensibility. We include a number of examples of using this pipe to interact with and extend other R packages.
Clustmixtype: User-Friendly Clustering Of Mixed-Type Data In R,
2018
Stralsund University of Applied Sciences
Clustmixtype: User-Friendly Clustering Of Mixed-Type Data In R, Gero Szepannek
The R Journal
Clustering algorithms are designed to identify groups in data where the traditional emphasis has been on numeric data. In consequence, many existing algorithms are devoted to this kind of data even though a combination of numeric and categorical data is more common in most business applications. Recently, new algorithms for clustering mixed-type data have been proposed based on Huang’s k-prototypes algorithm. This paper describes the R package clustMixType which provides an implementation of k-prototypes in R.
Editorial,
2018
R project
Editorial, John Verzani
The R Journal
On behalf of the editorial board, I am pleased to present Volume 10, Issue 2 of the R Journal.
This issue covers a wide range of topics through its 37 articles. As is typical, many of these are related to packages that provide tools for new statistical modeling in R. Examples in this issue include "clustMixType: User-Friendly Clustering of Mixed-Type Data in R" by Szepannek and "BNSP: an R Package for Fitting Bayesian Semiparametric Regression Models and Variable Selection" by Papageorgiou.
Downside Risk Evaluation With The R Package Gas,
2018
University of Neuchâtel, HEC Montréal
Downside Risk Evaluation With The R Package Gas, David Ardia, Kris Boudt, Leopoldo Catania
The R Journal
Financial risk managers routinely use non–linear time series models to predict the downside risk of the capital under management. They also need to evaluate the adequacy of their model using so–called backtesting procedures. The latter involve hypothesis testing and evaluation of loss functions. This paper shows how the R package GAS can be used for both the dynamic prediction and the evaluation of downside risk. Emphasis is given to the two key financial downside risk measures: Value-at-Risk (VaR) and Expected Shortfall (ES). High-level functions for: (i) prediction, (ii) backtesting, and (iii) model comparison are discussed, and code examples are provided. …
Explanations Of Model Predictions With Live And Breakdown Packages,
2018
Warsaw University of Technology
Explanations Of Model Predictions With Live And Breakdown Packages, Mateusz Staniak, Przemysław Biecek
The R Journal
Complex models are commonly used in predictive modeling. In this paper we present R packages that can be used for explaining predictions from complex black box models and attributing parts of these predictions to input features. We introduce two new approaches and corresponding packages for such attribution, namely live and breakDown. We also compare their results with existing implementations of state-of-the-art solutions, namely, lime (Pedersen and Benesty, 2018) which implements Locally Interpretable Model-agnostic Explanations and iml (Molnar et al., 2018) which implements Shapley values.
Sdpt3r: Semidefinite Quadratic Linear Programming In R,
2018
University of Waterloo
Sdpt3r: Semidefinite Quadratic Linear Programming In R, Adam Rahman
The R Journal
We present the package sdpt3r, an R implementation of the Matlab package SDPT3 (Toh et al., 1999). The purpose of the software is to solve semidefinite quadratic linear programming (SQLP) problems, which encompasses problems such as D-optimal experimental design, the nearest correlation matrix problem, and distance weighted discrimination, as well as problems in graph theory such as finding the maximum cut or Lovasz number of a graph.
Current optimization packages in R include Rdsdp, Rcsdp, scs, cccp, and Rmosek. Of these, scs and Rmosek solve a similar suite of problems. In addition to these …
Geospatial Point Density,
2018
United States Military Academy
Geospatial Point Density, Paul F. Evangelista, David Beskow
The R Journal
This paper introduces a spatial point density algorithm designed to be explainable, meaning ful, and efficient. Originally designed for military applications, this technique applies to any spatial point process where there is a desire to clearly understand the measurement of density and maintain fidelity of the point locations. Typical spatial density plotting algorithms, such as kernel density estimation, implement some type of smoothing function that often results in a density value that is difficult to interpret. The purpose of the visualization method in this paper is to understand spatial point activity density with precision and meaning. The temporal tendency of …
Lmridge: A Comprehensive R Package For Ridge Regression,
2018
Bahauddin Zakariya University
Lmridge: A Comprehensive R Package For Ridge Regression, Muhammad Imdad Ullah, Bahauddin Zakariya University Aslam, Saima Atlaf
The R Journal
The ridge regression estimator, one of the commonly used alternatives to the conventional ordinary least squares estimator, avoids the adverse effects in the situations when there exists some considerable degree of multicollinearity among the regressors. There are many software packages available for estimation of ridge regression coefficients. However, most of them display limited methods to estimate the ridge biasing parameters without testing procedures. Our developed package, lmridge can be used to estimate ridge coefficients considering a range of different existing biasing parameters, to test these coefficients with more than 25 ridge related statistics, and to present different graphical displays of …
Feature-Based Transfer Learning In Natural Language Processing,
2018
Singapore Management University
Feature-Based Transfer Learning In Natural Language Processing, Jianfei Yu
Dissertations and Theses Collection (Open Access)
In the past few decades, supervised machine learning approach is one of the most important methodologies in the Natural Language Processing (NLP) community. Although various kinds of supervised learning methods have been proposed to obtain the state-of-the-art performance across most NLP tasks, the bottleneck of them lies in the heavy reliance on the large amount of manually annotated data, which is not always available in our desired target domain/task. To alleviate the data sparsity issue in the target domain/task, an attractive solution is to find sufficient labeled data from a related source domain/task. However, for most NLP applications, due to …
Effectiveness Of Physical Robot Versus Robot Simulator In Teaching Introductory Programming,
2018
Singapore University of Technology and Design
Effectiveness Of Physical Robot Versus Robot Simulator In Teaching Introductory Programming, Oka Kurniawan, Norman Tiong Seng Lee, Subhajit Datta, Nachamma Sockalingam, Pey Lin Leong
Research Collection School Of Computing and Information Systems
This study reports the use of a physical robot and robot simulator in an introductory programming course in a university and measures students' programming background conceptual learning gain and learning experience. One group used physical robots in their lessons to complete programming assignments, while the other group used robot simulators. We are interested in finding out if there is any difference in the learning gain and experiences between those that use physical robots as compared to robot simulators. Our results suggest that there is no significant difference in terms of students' learning between the two approaches. However, the control group …
An Architectural Design And Evaluation Of An Affective Tutoring System For Novice Programmers,
2018
Singapore Management University
An Architectural Design And Evaluation Of An Affective Tutoring System For Novice Programmers, Hua Leong Fwa
Research Collection School Of Computing and Information Systems
Affect is prevalent in learning and it influences students’ learning achievement. This paper details the design and evaluation of an Affective Tutoring System (ATS) that tutors student in computer programming. Although most ATSs are purpose built for a specific domain, making adaptation to another domain difficult, this ATS is architected for adaptability and extensibility. This study also addresses a lack of research exploring the theories and methods of integrating affect and learning within the learning process by proposing methods of regulating the negative affect of students. Both quantitative and qualitative techniques were used for evaluation of the effectiveness of the …
Perflearner: Learning From Bug Reports To Understand And Generate Performance Test Frames,
2018
University of Kentucky
Perflearner: Learning From Bug Reports To Understand And Generate Performance Test Frames, Xue Han, Tingting Yu, David Lo
Research Collection School Of Computing and Information Systems
Software performance is important for ensuring the quality of software products. Performance bugs, defined as programming errors that cause significant performance degradation, can lead to slow systems and poor user experience. While there has been some research on automated performance testing such as test case generation, the main idea is to select workload values to increase the program execution times. These techniques often assume the initial test cases have the right combination of input parameters and focus on evolving values of certain input parameters. However, such an assumption may not hold for highly configurable real-word applications, in which the combinations …
