Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 151 - 180 of 708

Full-Text Articles in Computer Sciences

The R Package Nonprobest For Estimation In Non-Probability Surveys, M. Rueda, R. Ferri-García, L. Castro Jun 2020

The R Package Nonprobest For Estimation In Non-Probability Surveys, M. Rueda, R. Ferri-García, L. Castro

The R Journal

Different inference procedures are proposed in the literature to correct selection bias that might be introduced with non-random sampling mechanisms. The R package NonProbEst enables the estimation of parameters using some of these techniques to correct selection bias in non-probability surveys. The mean and the total of the target variable are estimated using Propensity Score Adjustment, calibration, statistical matching, model-based, model-assisted and model-calibratated techniques. Confidence intervals can also obtained for each method. Machine learning algorithms can be used for estimating the propensities or for predicting the unknown values of the target variable for the non-sampled units. Variance of a given …


Variable Importance Plots: An Introduction To The Vip Package, Brandon M. Bartoszuk, Marek Gagolewski Jun 2020

Variable Importance Plots: An Introduction To The Vip Package, Brandon M. Bartoszuk, Marek Gagolewski

The R Journal

In the era of “big data”, it is becoming more of a challenge to not only build state-of-the-art predictive models, but also gain an understanding of what’s really going on in the data. For example, it is often of interest to know which, if any, of the predictors in a fitted model are relatively influential on the predicted outcome. Some modern algorithms—like random forests (RFs) and gradient boosted decision trees (GBMs)—have a natural way of quantifying the importance or relative influence of each feature. Other algorithms—like naive Bayes classifiers and support vector machines—are not capable of doing so and model-agnostic …


Tsmp: An R Package For Time Series With Matrix Profile, Francisco Bischoff, Pedro Pereira Rodriques Jun 2020

Tsmp: An R Package For Time Series With Matrix Profile, Francisco Bischoff, Pedro Pereira Rodriques

The R Journal

This article describes tsmp, an R package that implements the MP concept for TS. The tsmp package is a toolkit that allows all-pairs similarity joins, motif, discords and chains discovery, semantic segmentation, etc. Here we describe how the tsmp package may be used by showing some of the use-cases from the original articles and evaluate the algorithm speed in the R environment. This package can be downloaded at https://CRAN.R-project.org/package=tsmp.


R Foundation News, Torsten Hothorn Jun 2020

R Foundation News, Torsten Hothorn

The R Journal

Membership fees and donations received between 2020-02-24 and 2020-09-08.


Projectmanagement: An R Package For Managing Projects, Juan Carlos Gonçalves-Dosantos, Ignacio García-Jurado, Julián Costa Jun 2020

Projectmanagement: An R Package For Managing Projects, Juan Carlos Gonçalves-Dosantos, Ignacio García-Jurado, Julián Costa

The R Journal

Project management is an important body of knowledge and practices that comprises the planning, organisation and control of resources to achieve one or more pre-determined objectives. In this paper, we introduce ProjectManagement, a new R package that provides the necessary tools to manage projects in a broad sense, and illustrate its use by examples.


Npordtests: An R Package Of Nonparametric Tests For Equality Of Location Against Ordered Alternatives, Bulent Altunkaynak, Hamza Gamgam Jun 2020

Npordtests: An R Package Of Nonparametric Tests For Equality Of Location Against Ordered Alternatives, Bulent Altunkaynak, Hamza Gamgam

The R Journal

Ordered alternatives are an important statistical problem in many situation such as increased risk of congenital malformation caused by excessive alcohol consumption during pregnancy life test experiments, drug-screening studies, dose-finding studies, the dose-response studies, age-related response. There are numerous other examples of this nature. In this paper, we present the npordtests package to test the equality of locations for ordered alternatives. The package includes the Jonckheere Terpstra, Beier and Buning’s Adaptive, Modified Jonckheere-Terpstra, Terpstra-Magel, Ferdhiana Terpstra-Magel, KTP, S and Gaur’s Gc tests. A simulation study is conducted to determine which test is the most appropriate test for which scenario and …


Spinifex: An R Package For Creating A Manual Tour Of Low-Dimensional Projections Of Multivariate Data, Nicholas Spyrison, Dianne Cook Jun 2020

Spinifex: An R Package For Creating A Manual Tour Of Low-Dimensional Projections Of Multivariate Data, Nicholas Spyrison, Dianne Cook

The R Journal

Dynamic low-dimensional linear projections of multivariate data collectively known as tours provide an important tool for exploring multivariate data and models. The R package tourr provides functions for several types of tours: grand, guided, little, local and frozen. Each of these can be viewed dynamically, or saved into a data object for animation. This paper describes a new package, spinifex, which provides a manual tour of multivariate data where the projection coefficient of a single variable is controlled. The variable is rotated fully into the projection, or completely out of the projection. The resulting sequence of projections can be …


Mistr: A Computational Framework For Mixture And Composite Distributions, Lukas Sablica, Kurt Hornik Jun 2020

Mistr: A Computational Framework For Mixture And Composite Distributions, Lukas Sablica, Kurt Hornik

The R Journal

No abstract provided.


Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis Jun 2020

Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis

The R Journal

In the past 8 months, 1554 new packages were added to the CRAN package repository. 96 packages were unarchived and 843 were archived. The following shows the growth of the number of active packages in the CRAN package repository:


Skew-T Expected Information Matrix Evaluation And Use For Standard Error Calculations, R. Douglas Martin, Chindhanai Uthaisaad, Daniel Z. Xia Jun 2020

Skew-T Expected Information Matrix Evaluation And Use For Standard Error Calculations, R. Douglas Martin, Chindhanai Uthaisaad, Daniel Z. Xia

The R Journal

Skew-t distributions derived from skew-normal distributions, as developed by Azzalini and several co-workers, are popular because of their theoretical foundation and the availability of computational methods in the R package sn. One difficulty with this skew-t family is that the elements of the expected information matrix do not have closed form analytic formulas. Thus, we developed a numerical integration method of computing the expected information matrix in the R package skewtInfo. The accuracy of our expected information matrix calculation method was confirmed by comparing the result with that obtained using an observed information matrix for a very large sample …


Tools For Analyzing R Code The Tidy Way, Lucy D'Agostino Mcgowan, Sean Kross, Jeffrey Leek Jun 2020

Tools For Analyzing R Code The Tidy Way, Lucy D'Agostino Mcgowan, Sean Kross, Jeffrey Leek

The R Journal

With the current emphasis on reproducibility and replicability, there is an increasing need to examine how data analyses are conducted. In order to analyze the between researcher variability in data analysis choices as well as the aspects within the data analysis pipeline that contribute to the variability in results, we have created two R packages: matahari and tidycode. These packages build on methods created for natural language processing; rather than allowing for the processing of natural language, we focus on R code as the substrate of interest. The matahari package facilitates the logging of everything that is typed in the …


Individual-Level Modelling Of Infectious Disease Data: Epiilm, Vineetha Warriyar, Waleed Almutiry, Rob Deardon Jun 2020

Individual-Level Modelling Of Infectious Disease Data: Epiilm, Vineetha Warriyar, Waleed Almutiry, Rob Deardon

The R Journal

In this article we introduce the R package EpiILM, which provides tools for simulation from, and inference for, discrete-time individual-level models of infectious disease transmission proposed by Deardon et al. (2010). The inference is set in a Bayesian framework and is carried out via Metropolis Hastings Markov chain Monte Carlo (MCMC). For its fast implementation, key functions are coded in Fortran. Both spatial and contact network models are implemented in the package and can be set in either susceptible-infected (SI) or susceptible-infected-removed (SIR) compartmental frameworks. Use of the package is demonstrated through examples involving both simulated and real data.


Conference Report: Why R? 2019, Michał Burdukiewicz, Filip Pietluch, Jarosław Chilimoniuk, Katarzyna Sidorczuk, Dominik Rafacz, Leon Eyrich Jessen, Stefan Rödiger, Marcin Kosiński, Piotr Wójcik Jun 2020

Conference Report: Why R? 2019, Michał Burdukiewicz, Filip Pietluch, Jarosław Chilimoniuk, Katarzyna Sidorczuk, Dominik Rafacz, Leon Eyrich Jessen, Stefan Rödiger, Marcin Kosiński, Piotr Wójcik

The R Journal

WhyR?conferences have been the hallmark of the Why R? Foundation (whyr.pl). Our goal has been to establish a series of international R-related events in Poland. After three years, weare happy to announce that our main event, the Why R? conference, has become one of the largest annual R conferences in Central Europe.


Difnlr: Generalized Logistic Regression Models For Dif And Ddf Detection, Adéla Hladká, Patrícia Martinková Jun 2020

Difnlr: Generalized Logistic Regression Models For Dif And Ddf Detection, Adéla Hladká, Patrícia Martinková

The R Journal

Differential item functioning (DIF) and differential distractor functioning (DDF) are impor tant topics in psychometrics, pointing to potential unfairness in items with respect to minorities or different social groups. Various methods have been proposed to detect these issues. The difNLR R package extends DIF methods currently provided in other packages by offering approaches based on generalized logistic regression models that account for possible guessing or inattention, and by pro viding methods to detect DIF and DDF among ordinal and nominal data. In the current paper, we describe implementation of the main functions of the difNLR package, from data generation, through …


Survboost: An R Package For High-Dimensional Variable Selection In The Stratified Proportional Hazards Model Via Gradient Boosting, Emily Morris, Kevin He, Yanming Li, Yi Li, Jian Kang Jun 2020

Survboost: An R Package For High-Dimensional Variable Selection In The Stratified Proportional Hazards Model Via Gradient Boosting, Emily Morris, Kevin He, Yanming Li, Yi Li, Jian Kang

The R Journal

High-dimensional variable selection in the proportional hazards (PH) model has many successful applications in different areas. In practice, data may involve confounding variables that do not satisfy the PH assumption, in which case the stratified proportional hazards (SPH) model can be adopted to control the confounding effects by stratification without directly modeling the confounding effects. However, there is a lack of computationally efficient statistical software for high-dimensional variable selection in the SPH model. In this work an R package, SurvBoost, is developed to implement the gradient boosting algorithm for fitting the SPH model with high-dimensional covariate variables. Simulation studies …


Copulacenr: Copula Based Regression Models For Bivariate Censored Data In R, Tao Sun, Ying Ding Jun 2020

Copulacenr: Copula Based Regression Models For Bivariate Censored Data In R, Tao Sun, Ying Ding

The R Journal

Bivariate time-to-event data frequently arise in research areas such as clinical trials and epidemiological studies, where the occurrence of two events are correlated. In many cases, the exact event times are unknown due to censoring. The copula model is a popular approach for modeling correlated bivariate censored data, in which the two marginal distributions and the between margin dependence are modeled separately. This article presents the R package CopulaCenR, which is designed for modeling and testing bivariate data under right or (general) interval censoring in a regression setting. It provides a variety of Archimedean copula functions including a flexible two-parameter …


Sortedeffects: Sorted Causal Effects In R, Schuowen Chen, Victor Chernozhukov, Iván Fernández-Val, Ye Luo Jun 2020

Sortedeffects: Sorted Causal Effects In R, Schuowen Chen, Victor Chernozhukov, Iván Fernández-Val, Ye Luo

The R Journal

Chernozhukov et al. (2018) proposed the sorted effect method for nonlinear regression models. This method consists of reporting percentiles of the partial effects, the sorted effects, in addition to the average effect commonly used to summarize the heterogeneity in the partial effects. They also propose to use the sorted effects to carry out classification analysis where the observational units are classified as most and least affected if their partial effect are above or below some tail sorted effects. The R package SortedEffects implements the estimation and inference methods therein and provides tools to visualize the results. This vignette serves as …


The R Journal (December 2019) 11(2): Complete Issue, The R Foundation Dec 2019

The R Journal (December 2019) 11(2): Complete Issue, The R Foundation

The R Journal

Editorial, Michael J. Kane

Contributed Research Articles

Using Web Services to Work with Geodata in R, Jan-Philipp Kolb

orthoDr: Semiparametric Dimension Reduction via Orthogonality Constrained Optimization, Ruoqing Zhu, Jiyang Zhang, Ruilin Zhao, Peng Xu, Wenzhuo Zhou, and Xin Zhang

coxed: An R Package for Computing Duration-Based Quantities from the Cox Proportional Hazards Model, Jonathan Kropko and Jeffrey J. Harden

Modeling Regimes with Extremes: The Bayesdfa Package for Identifying and Forecasting Common Trends and Anomalies in Multivariate Time-Series Data, Eric J. Ward, Sean C. Anderson, Luis A. Damiano, Mary E. Hunsicker, and Michael A. Litzow

Fitting Tails by the Empirical Residual …


R Foundation News, Torsten Hothorn Dec 2019

R Foundation News, Torsten Hothorn

The R Journal

Membership fees and donations received between 2019-09-05 and 2020-02-24.


Conference: Report Conectar 2019, Marcela Alfaro Córdoba, Frans Van Dunné, Agustín Gómez Meléndez, Jacob Van Etten Dec 2019

Conference: Report Conectar 2019, Marcela Alfaro Córdoba, Frans Van Dunné, Agustín Gómez Meléndez, Jacob Van Etten

The R Journal

ConectaR 2019: Encuentro de Usuarios R en Latinoamérica, took place during January 24-26, 2019 at the University of Costa Rica, in San José, Costa Rica. It was the first event in Central America endorsed by The R Foundation, and it was held completely in Spanish. The majority of the attendants were from Costa Rica (85%), but we had participants from 12 countries: Costa Rica, Guatemala, Peru, Colombia, Mexico, Argentina, Uruguay, Chile, Spain, the Netherlands, France and the USA. The three-day event consisted of talks, workshops, and poster sessions.


Lpirfs: An R Package To Estimate Impulse Response Functions By Local Projections, Philipp Adämmer Dec 2019

Lpirfs: An R Package To Estimate Impulse Response Functions By Local Projections, Philipp Adämmer

The R Journal

Impulse response analysis is a cornerstone in applied (macro-)econometrics. Estimating impulse response functions using local projections (LPs) has become an appealing alternative to the traditional structural vector autoregressive (SVAR) approach. Despite its growing popularity and applications, however, no R package yet exists that makes this method available. In this paper, I introduce lpirfs, a fast and flexible R package that provides a broad framework to compute and visualize impulse response functions using LPs for a variety of data sets.


Resampling-Based Analysis Of Multivariate Data And Repeated Measures Designs With The R Package Manova.Rm, Sarah Friedrich, Frank Konietschke, Markus Pauly Dec 2019

Resampling-Based Analysis Of Multivariate Data And Repeated Measures Designs With The R Package Manova.Rm, Sarah Friedrich, Frank Konietschke, Markus Pauly

The R Journal

Nonparametric statistical inference methods for a modern and robust analysis of longitudinal and multivariate data in factorial experiments are essential for research. While existing approaches that rely on specific distributional assumptions of the data (multivariate normality and/or equal covariance matrices) are implemented in statistical software packages, there is a need for user-friendly software that can be used for the analysis of data that do not fulfill the aforementioned assumptions and provide accurate p value and confidence interval estimates. Therefore, newly developed nonparametric statistical methods based on bootstrap- and permutation-approaches, which neither assume multivariate normality nor specific covariance matrices, have been …


The Landscape Of R Packages For Automated Exploratory Data Analysis, Mateusz Staniak, Przemysław Biecek Dec 2019

The Landscape Of R Packages For Automated Exploratory Data Analysis, Mateusz Staniak, Przemysław Biecek

The R Journal

The increasing availability of large but noisy data sets with a large number of heterogeneous variables leads to the increasing interest in the automation of common tasks for data analysis. The most time-consuming part of this process is the Exploratory Data Analysis, crucial for better domain understanding, data cleaning, data validation, and feature engineering

There is a growing number of libraries that attempt to automate some of the typical Exploratory Data Analysis tasks to make the search for new insights easier and faster. In this paper, we present a systematic review of existing tools for Automated Exploratory Data Analysis (autoEDA). …


Roahd Package: Robust Analysis Of High Dimensional Data, Francesca Ieva, Anna Maria Paganoni, Juan Romo, Nicholas Tarabelloni Dec 2019

Roahd Package: Robust Analysis Of High Dimensional Data, Francesca Ieva, Anna Maria Paganoni, Juan Romo, Nicholas Tarabelloni

The R Journal

The focus of this paper is on the open-source R package roahd (RObust Analysis of High dimensional Data), see Tarabelloni et al. (2017). roahd has been developed to gather recently proposed statistical methods that deal with the robust inferential analysis of univariate and multivariate functional data. In particular, efficient methods for outlier detection and related graphical tools, methods to represent and simulate functional data, as well as inferential tools for testing differences and dependency among families of curves will be discussed, and the associated functions of the package will be described in details.


Jomo: A Flexible Package For Two-Level Joint Modelling Multiple Imputation, Matteo Quartagno, Simon Grund, James Carpenter Dec 2019

Jomo: A Flexible Package For Two-Level Joint Modelling Multiple Imputation, Matteo Quartagno, Simon Grund, James Carpenter

The R Journal

Multiple imputation is a tool for parameter estimation and inference with partially observed data, which is used increasingly widely in medical and social research. When the data to be imputed are correlated or have a multilevel structure — repeated observations on patients, school children nested in classes within schools within educational districts — the imputation model needs to include this structure. Here we introduce our joint modelling package for multiple imputation of multilevel data, jomo, which uses a multivariate normal model fitted by Markov Chain Monte Carlo (MCMC). Compared to previous packages for multilevel imputation, e.g. pan, jomo adds the …


Cvcrand: A Package For Covariate-Constrained Randomization And The Clustered Permutation Test For Cluster Randomized Trials, Hengshi Yu, Fan Li, John A. Gallis, Elizabeth L. Turner Dec 2019

Cvcrand: A Package For Covariate-Constrained Randomization And The Clustered Permutation Test For Cluster Randomized Trials, Hengshi Yu, Fan Li, John A. Gallis, Elizabeth L. Turner

The R Journal

The cluster randomized trial (CRT) is a randomized controlled trial in which randomization is conducted at the cluster level (e.g., school or hospital) and outcomes are measured for each individual within a cluster. Often, the number of clusters available to randomize is small (≤ 20), which increases the chance of baseline covariate imbalance between comparison arms. Such imbalance is particularly problematic when the covariates are predictive of the outcome because it can threaten the internal validity of the CRT. Pair-matching and stratification are two restricted randomization approaches that are frequently used to ensure balance at the design stage. An alternative, …


Biclustermd: An R Package For Biclustering With Missing Values, John Reisner, Hieu Pham, Sigurdur Olafsson, Stephen Vardeman, Jing Li Dec 2019

Biclustermd: An R Package For Biclustering With Missing Values, John Reisner, Hieu Pham, Sigurdur Olafsson, Stephen Vardeman, Jing Li

The R Journal

Biclustering is a statistical learning technique that attempts to find homogeneous partitions of rows and columns of a data matrix. For example, movie ratings might be biclustered to group both raters and movies. biclust is a current R package allowing users to implement a variety of biclustering algorithms. However, its algorithms do not allow the data matrix to have missing values. We provide a new R package, biclustermd, which allows users to perform biclustering on numeric data even in the presence of missing values.


Modeling Regimes With Extremes: The Bayesdfa Package For Identifying And Forecasting Common Trends And Anomalies In Multivariate Time-Series Data, Eric J. Ward, Sean C. Anderson, Luis A. Damiano, Mary E. Hunsicker, Michael A. Litzow Dec 2019

Modeling Regimes With Extremes: The Bayesdfa Package For Identifying And Forecasting Common Trends And Anomalies In Multivariate Time-Series Data, Eric J. Ward, Sean C. Anderson, Luis A. Damiano, Mary E. Hunsicker, Michael A. Litzow

The R Journal

The bayesdfa package provides a flexible Bayesian modeling framework for applying dynamic factor analysis (DFA) to multivariate time-series data as a dimension reduction tool. The core estimation is done with the Stan probabilistic programming language. In addition to being one of the few Bayesian implementations of DFA, novel features of this model include (1) optionally modeling latent process deviations as drawn from a Student-t distribution to better model extremes, and (2) optionally including autoregressive and moving-average components in the latent trends. Besides estimation, we provide a series of plotting functions to visualize trends, loadings, and model predicted values. A secondary …


Ppci: An R Package For Cluster Identification Using Projection Pursuit, David P. Hofmeyr, Nicos G. Pavlidis Dec 2019

Ppci: An R Package For Cluster Identification Using Projection Pursuit, David P. Hofmeyr, Nicos G. Pavlidis

The R Journal

This paper presents the R package PPCI which implements three recently proposed projection pursuit methods for clustering. The methods are unified by the approach of defining an optimal hyperplane to separate clusters, and deriving a projection index whose optimiser is the vector normal to this separating hyperplane. Divisive hierarchical clustering algorithms that can detect clusters defined in different subspaces are readily obtained by recursively bi-partitioning the data through such hyperplanes. Projecting onto the vector normal to the optimal hyperplane enables visualisations of the data that can be used to validate the partition at each level of the cluster hierarchy. Clustering …


Coxed: An R Package For Computing Duration-Based Quantities From The Cox Proportional Hazards Model, Jonathan Kropko, Jeffrey J. Harden Dec 2019

Coxed: An R Package For Computing Duration-Based Quantities From The Cox Proportional Hazards Model, Jonathan Kropko, Jeffrey J. Harden

The R Journal

The Cox proportional hazards model is one of the most frequently used estimators in duration (survival) analysis. Because it is estimated using only the observed durations’ rank ordering, typical quantities of interest used to communicate results of the Cox model come from the hazard function (e.g., hazard ratios or percentage changes in the hazard rate). These quantities are substantively vague and difficult for many audiences of research to understand. We introduce a suite of methods in the R package coxed to address these problems. The package allows researchers to calculate duration-based quantities from Cox model results, such as the expected …