Open Access. Powered by Scholars. Published by Universities.®

University of Nebraska - Lincoln

Discipline
Keyword
Publication Year
Publication

Articles 181 - 210 of 747

Full-Text Articles in Numerical Analysis and Scientific Computing

Difnlr: Generalized Logistic Regression Models For Dif And Ddf Detection, Adéla Hladká, Patrícia Martinková Jun 2020

Difnlr: Generalized Logistic Regression Models For Dif And Ddf Detection, Adéla Hladká, Patrícia Martinková

The R Journal

Differential item functioning (DIF) and differential distractor functioning (DDF) are impor tant topics in psychometrics, pointing to potential unfairness in items with respect to minorities or different social groups. Various methods have been proposed to detect these issues. The difNLR R package extends DIF methods currently provided in other packages by offering approaches based on generalized logistic regression models that account for possible guessing or inattention, and by pro viding methods to detect DIF and DDF among ordinal and nominal data. In the current paper, we describe implementation of the main functions of the difNLR package, from data generation, through …


Survboost: An R Package For High-Dimensional Variable Selection In The Stratified Proportional Hazards Model Via Gradient Boosting, Emily Morris, Kevin He, Yanming Li, Yi Li, Jian Kang Jun 2020

Survboost: An R Package For High-Dimensional Variable Selection In The Stratified Proportional Hazards Model Via Gradient Boosting, Emily Morris, Kevin He, Yanming Li, Yi Li, Jian Kang

The R Journal

High-dimensional variable selection in the proportional hazards (PH) model has many successful applications in different areas. In practice, data may involve confounding variables that do not satisfy the PH assumption, in which case the stratified proportional hazards (SPH) model can be adopted to control the confounding effects by stratification without directly modeling the confounding effects. However, there is a lack of computationally efficient statistical software for high-dimensional variable selection in the SPH model. In this work an R package, SurvBoost, is developed to implement the gradient boosting algorithm for fitting the SPH model with high-dimensional covariate variables. Simulation studies …


Copulacenr: Copula Based Regression Models For Bivariate Censored Data In R, Tao Sun, Ying Ding Jun 2020

Copulacenr: Copula Based Regression Models For Bivariate Censored Data In R, Tao Sun, Ying Ding

The R Journal

Bivariate time-to-event data frequently arise in research areas such as clinical trials and epidemiological studies, where the occurrence of two events are correlated. In many cases, the exact event times are unknown due to censoring. The copula model is a popular approach for modeling correlated bivariate censored data, in which the two marginal distributions and the between margin dependence are modeled separately. This article presents the R package CopulaCenR, which is designed for modeling and testing bivariate data under right or (general) interval censoring in a regression setting. It provides a variety of Archimedean copula functions including a flexible two-parameter …


Sortedeffects: Sorted Causal Effects In R, Schuowen Chen, Victor Chernozhukov, Iván Fernández-Val, Ye Luo Jun 2020

Sortedeffects: Sorted Causal Effects In R, Schuowen Chen, Victor Chernozhukov, Iván Fernández-Val, Ye Luo

The R Journal

Chernozhukov et al. (2018) proposed the sorted effect method for nonlinear regression models. This method consists of reporting percentiles of the partial effects, the sorted effects, in addition to the average effect commonly used to summarize the heterogeneity in the partial effects. They also propose to use the sorted effects to carry out classification analysis where the observational units are classified as most and least affected if their partial effect are above or below some tail sorted effects. The R package SortedEffects implements the estimation and inference methods therein and provides tools to visualize the results. This vignette serves as …


The R Journal (December 2019) 11(2): Complete Issue, The R Foundation Dec 2019

The R Journal (December 2019) 11(2): Complete Issue, The R Foundation

The R Journal

Editorial, Michael J. Kane

Contributed Research Articles

Using Web Services to Work with Geodata in R, Jan-Philipp Kolb

orthoDr: Semiparametric Dimension Reduction via Orthogonality Constrained Optimization, Ruoqing Zhu, Jiyang Zhang, Ruilin Zhao, Peng Xu, Wenzhuo Zhou, and Xin Zhang

coxed: An R Package for Computing Duration-Based Quantities from the Cox Proportional Hazards Model, Jonathan Kropko and Jeffrey J. Harden

Modeling Regimes with Extremes: The Bayesdfa Package for Identifying and Forecasting Common Trends and Anomalies in Multivariate Time-Series Data, Eric J. Ward, Sean C. Anderson, Luis A. Damiano, Mary E. Hunsicker, and Michael A. Litzow

Fitting Tails by the Empirical Residual …


R Foundation News, Torsten Hothorn Dec 2019

R Foundation News, Torsten Hothorn

The R Journal

Membership fees and donations received between 2019-09-05 and 2020-02-24.


Conference: Report Conectar 2019, Marcela Alfaro Córdoba, Frans Van Dunné, Agustín Gómez Meléndez, Jacob Van Etten Dec 2019

Conference: Report Conectar 2019, Marcela Alfaro Córdoba, Frans Van Dunné, Agustín Gómez Meléndez, Jacob Van Etten

The R Journal

ConectaR 2019: Encuentro de Usuarios R en Latinoamérica, took place during January 24-26, 2019 at the University of Costa Rica, in San José, Costa Rica. It was the first event in Central America endorsed by The R Foundation, and it was held completely in Spanish. The majority of the attendants were from Costa Rica (85%), but we had participants from 12 countries: Costa Rica, Guatemala, Peru, Colombia, Mexico, Argentina, Uruguay, Chile, Spain, the Netherlands, France and the USA. The three-day event consisted of talks, workshops, and poster sessions.


Lpirfs: An R Package To Estimate Impulse Response Functions By Local Projections, Philipp Adämmer Dec 2019

Lpirfs: An R Package To Estimate Impulse Response Functions By Local Projections, Philipp Adämmer

The R Journal

Impulse response analysis is a cornerstone in applied (macro-)econometrics. Estimating impulse response functions using local projections (LPs) has become an appealing alternative to the traditional structural vector autoregressive (SVAR) approach. Despite its growing popularity and applications, however, no R package yet exists that makes this method available. In this paper, I introduce lpirfs, a fast and flexible R package that provides a broad framework to compute and visualize impulse response functions using LPs for a variety of data sets.


Resampling-Based Analysis Of Multivariate Data And Repeated Measures Designs With The R Package Manova.Rm, Sarah Friedrich, Frank Konietschke, Markus Pauly Dec 2019

Resampling-Based Analysis Of Multivariate Data And Repeated Measures Designs With The R Package Manova.Rm, Sarah Friedrich, Frank Konietschke, Markus Pauly

The R Journal

Nonparametric statistical inference methods for a modern and robust analysis of longitudinal and multivariate data in factorial experiments are essential for research. While existing approaches that rely on specific distributional assumptions of the data (multivariate normality and/or equal covariance matrices) are implemented in statistical software packages, there is a need for user-friendly software that can be used for the analysis of data that do not fulfill the aforementioned assumptions and provide accurate p value and confidence interval estimates. Therefore, newly developed nonparametric statistical methods based on bootstrap- and permutation-approaches, which neither assume multivariate normality nor specific covariance matrices, have been …


The Landscape Of R Packages For Automated Exploratory Data Analysis, Mateusz Staniak, Przemysław Biecek Dec 2019

The Landscape Of R Packages For Automated Exploratory Data Analysis, Mateusz Staniak, Przemysław Biecek

The R Journal

The increasing availability of large but noisy data sets with a large number of heterogeneous variables leads to the increasing interest in the automation of common tasks for data analysis. The most time-consuming part of this process is the Exploratory Data Analysis, crucial for better domain understanding, data cleaning, data validation, and feature engineering

There is a growing number of libraries that attempt to automate some of the typical Exploratory Data Analysis tasks to make the search for new insights easier and faster. In this paper, we present a systematic review of existing tools for Automated Exploratory Data Analysis (autoEDA). …


Roahd Package: Robust Analysis Of High Dimensional Data, Francesca Ieva, Anna Maria Paganoni, Juan Romo, Nicholas Tarabelloni Dec 2019

Roahd Package: Robust Analysis Of High Dimensional Data, Francesca Ieva, Anna Maria Paganoni, Juan Romo, Nicholas Tarabelloni

The R Journal

The focus of this paper is on the open-source R package roahd (RObust Analysis of High dimensional Data), see Tarabelloni et al. (2017). roahd has been developed to gather recently proposed statistical methods that deal with the robust inferential analysis of univariate and multivariate functional data. In particular, efficient methods for outlier detection and related graphical tools, methods to represent and simulate functional data, as well as inferential tools for testing differences and dependency among families of curves will be discussed, and the associated functions of the package will be described in details.


Jomo: A Flexible Package For Two-Level Joint Modelling Multiple Imputation, Matteo Quartagno, Simon Grund, James Carpenter Dec 2019

Jomo: A Flexible Package For Two-Level Joint Modelling Multiple Imputation, Matteo Quartagno, Simon Grund, James Carpenter

The R Journal

Multiple imputation is a tool for parameter estimation and inference with partially observed data, which is used increasingly widely in medical and social research. When the data to be imputed are correlated or have a multilevel structure — repeated observations on patients, school children nested in classes within schools within educational districts — the imputation model needs to include this structure. Here we introduce our joint modelling package for multiple imputation of multilevel data, jomo, which uses a multivariate normal model fitted by Markov Chain Monte Carlo (MCMC). Compared to previous packages for multilevel imputation, e.g. pan, jomo adds the …


Cvcrand: A Package For Covariate-Constrained Randomization And The Clustered Permutation Test For Cluster Randomized Trials, Hengshi Yu, Fan Li, John A. Gallis, Elizabeth L. Turner Dec 2019

Cvcrand: A Package For Covariate-Constrained Randomization And The Clustered Permutation Test For Cluster Randomized Trials, Hengshi Yu, Fan Li, John A. Gallis, Elizabeth L. Turner

The R Journal

The cluster randomized trial (CRT) is a randomized controlled trial in which randomization is conducted at the cluster level (e.g., school or hospital) and outcomes are measured for each individual within a cluster. Often, the number of clusters available to randomize is small (≤ 20), which increases the chance of baseline covariate imbalance between comparison arms. Such imbalance is particularly problematic when the covariates are predictive of the outcome because it can threaten the internal validity of the CRT. Pair-matching and stratification are two restricted randomization approaches that are frequently used to ensure balance at the design stage. An alternative, …


Biclustermd: An R Package For Biclustering With Missing Values, John Reisner, Hieu Pham, Sigurdur Olafsson, Stephen Vardeman, Jing Li Dec 2019

Biclustermd: An R Package For Biclustering With Missing Values, John Reisner, Hieu Pham, Sigurdur Olafsson, Stephen Vardeman, Jing Li

The R Journal

Biclustering is a statistical learning technique that attempts to find homogeneous partitions of rows and columns of a data matrix. For example, movie ratings might be biclustered to group both raters and movies. biclust is a current R package allowing users to implement a variety of biclustering algorithms. However, its algorithms do not allow the data matrix to have missing values. We provide a new R package, biclustermd, which allows users to perform biclustering on numeric data even in the presence of missing values.


Modeling Regimes With Extremes: The Bayesdfa Package For Identifying And Forecasting Common Trends And Anomalies In Multivariate Time-Series Data, Eric J. Ward, Sean C. Anderson, Luis A. Damiano, Mary E. Hunsicker, Michael A. Litzow Dec 2019

Modeling Regimes With Extremes: The Bayesdfa Package For Identifying And Forecasting Common Trends And Anomalies In Multivariate Time-Series Data, Eric J. Ward, Sean C. Anderson, Luis A. Damiano, Mary E. Hunsicker, Michael A. Litzow

The R Journal

The bayesdfa package provides a flexible Bayesian modeling framework for applying dynamic factor analysis (DFA) to multivariate time-series data as a dimension reduction tool. The core estimation is done with the Stan probabilistic programming language. In addition to being one of the few Bayesian implementations of DFA, novel features of this model include (1) optionally modeling latent process deviations as drawn from a Student-t distribution to better model extremes, and (2) optionally including autoregressive and moving-average components in the latent trends. Besides estimation, we provide a series of plotting functions to visualize trends, loadings, and model predicted values. A secondary …


Ppci: An R Package For Cluster Identification Using Projection Pursuit, David P. Hofmeyr, Nicos G. Pavlidis Dec 2019

Ppci: An R Package For Cluster Identification Using Projection Pursuit, David P. Hofmeyr, Nicos G. Pavlidis

The R Journal

This paper presents the R package PPCI which implements three recently proposed projection pursuit methods for clustering. The methods are unified by the approach of defining an optimal hyperplane to separate clusters, and deriving a projection index whose optimiser is the vector normal to this separating hyperplane. Divisive hierarchical clustering algorithms that can detect clusters defined in different subspaces are readily obtained by recursively bi-partitioning the data through such hyperplanes. Projecting onto the vector normal to the optimal hyperplane enables visualisations of the data that can be used to validate the partition at each level of the cluster hierarchy. Clustering …


Coxed: An R Package For Computing Duration-Based Quantities From The Cox Proportional Hazards Model, Jonathan Kropko, Jeffrey J. Harden Dec 2019

Coxed: An R Package For Computing Duration-Based Quantities From The Cox Proportional Hazards Model, Jonathan Kropko, Jeffrey J. Harden

The R Journal

The Cox proportional hazards model is one of the most frequently used estimators in duration (survival) analysis. Because it is estimated using only the observed durations’ rank ordering, typical quantities of interest used to communicate results of the Cox model come from the hazard function (e.g., hazard ratios or percentage changes in the hazard rate). These quantities are substantively vague and difficult for many audiences of research to understand. We introduce a suite of methods in the R package coxed to address these problems. The package allows researchers to calculate duration-based quantities from Cox model results, such as the expected …


Orthodr: Semiparametric Dimension Reduction Via Orthogonality Constrained, Ruoqing Zhu, Jiyang Zhang, Ruilin Zhao, Peng Xu, Wenzhuo Zhou, Xin Zhang Dec 2019

Orthodr: Semiparametric Dimension Reduction Via Orthogonality Constrained, Ruoqing Zhu, Jiyang Zhang, Ruilin Zhao, Peng Xu, Wenzhuo Zhou, Xin Zhang

The R Journal

orthoDr is a package in R that solves dimension reduction problems using orthogonality constrained optimization approach. The package serves as a unified framework for many regression and survival analysis dimension reduction models that utilize semiparametric estimating equations. The main computational machinery of orthoDr is a first-order algorithm developed by Wen and Yin (2012) for optimization within the Stiefel manifold. We implement the algorithm through Rcpp and OpenMP for fast computation. In addition, we developed a general-purpose solver for such constrained problems with user-specified objective functions, which works as a drop-in version of optim(). The package also serves as a platform …


Using Web Services To Work With Geodata In R, Jan-Philipp Kolb Dec 2019

Using Web Services To Work With Geodata In R, Jan-Philipp Kolb

The R Journal

Through collaborative mapping, a massive amount of data is accessible. Many individuals contribute information each day. The growing amount of geodata is gathered by volunteers or obtained via crowd-sourcing. One outstanding example of this is the OpenStreetMap (OSM) Project which provides access to big data in geography. Another online mapping service that enables the integration of geodata into the analysis is Google Maps. The expanding content and the availability of geographic information radically changes the perspective on geodata (Chilton 2009). Recently many application programming interfaces (APIs) have been built on OSM and Google Maps. That leads to a point where …


Spgarch: An R-Package For Spatial And Spatiotemporal Arch And Garch Models, Philipp Otto Dec 2019

Spgarch: An R-Package For Spatial And Spatiotemporal Arch And Garch Models, Philipp Otto

The R Journal

In this paper, a general overview on spatial and spatiotemporal ARCH models is provided. In particular, we distinguish between three different spatial ARCH-type models. In addition to the original definition of Otto et al. (2016), we introduce an logarithmic spatial ARCH model in this paper. For this new model, maximum-likelihood estimators for the parameters are proposed. In addition, we consider a new complex-valued definition of the spatial ARCH process. Moreover, spatial GARCH models are briefly discussed. From a practical point of view, the use of the R-package spGARCH is demonstrated. To be precise, we show how the proposed spatial ARCH …


Hcmodelsets: An R Package For Specifying Sets Of Well-Fitting Models In High Dimensions, Henrique Hoeltgebaum, Heather Battey Dec 2019

Hcmodelsets: An R Package For Specifying Sets Of Well-Fitting Models In High Dimensions, Henrique Hoeltgebaum, Heather Battey

The R Journal

In the context of regression with a large number of explanatory variables, Cox and Battey (2017) emphasize that if there are alternative reasonable explanations of the data that are statistically indistinguishable, one should aim to specify as many of these explanations as is feasible. The standard practice, by contrast, is to report a single effective model for prediction. This paper illustrates the R implementation of the new ideas in the package HCmodelSets, using simple reproducible examples and real data. Results of some simulation experiments are also reported.


The R Package Trafo For Transforming Linear Regression Models, Lily Medina, Ann-Kristin Kreutzmann, Natalia Rojas-Perilla, Piedad Castro Dec 2019

The R Package Trafo For Transforming Linear Regression Models, Lily Medina, Ann-Kristin Kreutzmann, Natalia Rojas-Perilla, Piedad Castro

The R Journal

Researchers and data-analysts often use the linear regression model for descriptive, predictive, and inferential purposes. This model relies on a set of assumptions that, when not satisfied, yields biased results and noisy estimates. A common problem that can be solved in many ways – use of less restrictive methods (e.g. generalized linear regression models or non-parametric methods ), variance corrections or transformations of the response variable just to name a few. We focus on the latter option as it allows to keep using the simple and well-known linear regression model. The list of transformations proposed in the literature is long …


Comparing Namedcapture With Other R Packages For Regular Expressions, Toby Dylan Hocking Dec 2019

Comparing Namedcapture With Other R Packages For Regular Expressions, Toby Dylan Hocking

The R Journal

Regular expressions are powerful tools for manipulating non-tabular textual data. For many tasks (visualization, machine learning, etc), tables of numbers must be extracted from such data before processing by other R functions. We present the R package namedCapture, which facilitates such tasks by providing a new user-friendly syntax for defining regular expressions in R code. We begin by describing the history of regular expressions and their usage in R. We then describe the new features of the namedCapture package, and provide detailed comparisons with related R packages (rex, stringr, stringi, tidyr, rematch2, re2r).


News From The Bioconductor Project, Bioconductor Core Team Dec 2019

News From The Bioconductor Project, Bioconductor Core Team

The R Journal

The Bioconductor project provides tools for the analysis and comprehension of high-throughput genomic data. Bioconductor 3.10 was released on 30 October, 2019. It is compatible with R 3.6.1 and consists of 1823 software packages, 384 experiment data packages, 953 up-to-date annotation packages, and 27 workflows. The release announcement includes descriptions of 94 new software packages, and updated NEWS files for many additional packages. Start using Bioconductor by installing the most recent version of R and evaluating the commands


Fitting Tails By The Empirical Residual Coefficient Of Variation: The Ercv Package, Joan Del Castillo, Isabel Serra, Maria Padilla, David Moriña Dec 2019

Fitting Tails By The Empirical Residual Coefficient Of Variation: The Ercv Package, Joan Del Castillo, Isabel Serra, Maria Padilla, David Moriña

The R Journal

This article is a self-contained introduction to the R package ercv and to the methodology on which it is based through the analysis of nine examples. The methodology is simple and trustworthy for the analysis of extreme values and relates the two main existing methodologies. The package contains R functions for visualizing, fitting and validating the distribution of tails. It also provides multiple threshold tests for a generalized Pareto distribution, together with an automatic threshold selection algorithm.


Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis Dec 2019

Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis

The R Journal

In the past 4 months, 632 new packages were added to the CRAN package repository. 27 packages were unarchived and 182 were archived. The following shows the growth of the number of active packages in the CRAN package repository:


Rollmatch: An R Package For Rolling Entry Matching, Kasey Jones, Rob Chew, Allison Witman, Yiyan Liu Dec 2019

Rollmatch: An R Package For Rolling Entry Matching, Kasey Jones, Rob Chew, Allison Witman, Yiyan Liu

The R Journal

The gold standard of experimental research is the randomized control trial. However, interventions are often implemented without a randomized control group for practical or ethical reasons. Propensity score matching (PSM) is a popular method for minimizing the effects of a randomized experiment from observational data by matching members of a treatment group to similar candidates that did not receive the intervention. Traditional PSM is not designed for studies that enroll participants on a rolling basis and does not provide a solution for interventions in which the baseline and intervention period are undefined in the comparison group. Rolling Entry Matching (REM) …


Editorial, Michael J. Kane Dec 2019

Editorial, Michael J. Kane

The R Journal

On behalf of the editorial board, I am pleased to present Volume 11, Issue 2 of the R Journal and my first issue as the Editor in Chief. This year, both Colin Gillespie and Catherine Healey join the Editorial Board, and Norm Matloff will rotate out. The R Journal continues to see increases in impact and popularity and this year we plan on making advances to better serve the community and streamline the publishing process to meet the increase in submissions we have seen over the last few years.


Dr4pl: A Stable Convergence Algorithm For The 4 Parameter Logistic Model, Hyowon An, Justin T. Landis, Aubrey G. Bailey, James S. Marron, Dirk P. Dittmer Dec 2019

Dr4pl: A Stable Convergence Algorithm For The 4 Parameter Logistic Model, Hyowon An, Justin T. Landis, Aubrey G. Bailey, James S. Marron, Dirk P. Dittmer

The R Journal

The 4 Parameter Logistic (4PL) model has been recognized as a major tool to analyze the relationship between doses and responses in pharmacological experiments. A main strength of this model is that each parameter contributes an intuitive meaning enhancing interpretability of a fitted model. However, implementing the 4PL model using conventional statistical software often encounters numerical errors. This paper highlights the issue of convergence failure and presents several causes with solutions. These causes include outliers and a non-logistic data shape, so useful remedies such as robust estimation, outlier diagnostics and constrained optimization are proposed. These features are implemented in a …


Associative Classification In R: Arc, Arulescba, And Rcba, Michael Hahsler, Ian Johnson, Tomáš Kliegr, Jaroslav Kuchař Dec 2019

Associative Classification In R: Arc, Arulescba, And Rcba, Michael Hahsler, Ian Johnson, Tomáš Kliegr, Jaroslav Kuchař

The R Journal

Several methods for creating classifiers based on rules discovered via association rule mining have been proposed in the literature. These classifiers are called associative classifiers and the best-known algorithm is Classification Based on Associations (CBA). Interestingly, only very few implementations are available and, until recently, no implementation was available for R. Now, three packages provide CBA. This paper introduces associative classification, the CBA algorithm, and how it can be used in R. A comparison of the three packages is provided to give the potential user an idea about the advantages of each of the implementations. We also show how the …