Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 91 - 120 of 708

Full-Text Articles in Computer Sciences

R Medicine 2020: The Power Of Going Virtual, Elizabeth J. Atkinson, Peter D. Higgins, Denise Esserman, Michael J. Kane, Steven J. Schwager, Joseph B. Rickert, Daniella Mark, Mara Alexeev, Stephan Kadauke Jun 2021

R Medicine 2020: The Power Of Going Virtual, Elizabeth J. Atkinson, Peter D. Higgins, Denise Esserman, Michael J. Kane, Steven J. Schwager, Joseph B. Rickert, Daniella Mark, Mara Alexeev, Stephan Kadauke

The R Journal

The third annual R/Medicine conference was planned as a physical event to be held in Philadelphia at the end of August 2020. However, a nationwide lockdown induced by the COVID-19 pandemic required a swift transition to a virtual conference. This article describes the challenges and benefits we encountered with this transition and provides an overview of the conference content.


Changes In R 4.0–4.1, Tomas Kalibera, Sebastian Meyer, Kurt Hornik Jun 2021

Changes In R 4.0–4.1, Tomas Kalibera, Sebastian Meyer, Kurt Hornik

The R Journal

We give a selection of the most important changes in R 4.1.0. Some statistics on source code commits and bug tracking activities are also provided.


Bayesspsurv: An R Package To Estimate Bayesian (Spatial) Split-Population Survival Models, Brandon Bolte, Nicolás Schmidt, Sergio Béjar, Nguyen Huynh, Bumba Mukherjee Jun 2021

Bayesspsurv: An R Package To Estimate Bayesian (Spatial) Split-Population Survival Models, Brandon Bolte, Nicolás Schmidt, Sergio Béjar, Nguyen Huynh, Bumba Mukherjee

The R Journal

Survival data often include a fraction of units that are susceptible to an event of interest as well as a fraction of “immune” units. In many applications, spatial clustering in unobserved risk factors across nearby units can also affect their survival rates and odds of becoming immune. To address these methodological challenges, this article introduces our BayesSPsurv R-package, which fits parametric Bayesian Spatial split-population survival (cure) models that can account for spatial autocorrelation in both subpopulations of the user’s time-to-event data. Spatial autocorrelation is modeled with spatially weighted frailties, which are estimated using a conditionally autoregressive prior. The user can …


Gofcopula: Goodness-Of-Fit Tests For Copulae, Ostap Okhrin, Simon Trimborn, Martin Waltz Jun 2021

Gofcopula: Goodness-Of-Fit Tests For Copulae, Ostap Okhrin, Simon Trimborn, Martin Waltz

The R Journal

The last decades show an increased interest in modeling various types of data through copulae. Different copula models have been developed, which lead to the challenge of finding the best fitting model for a particular dataset. From the other side, a strand of literature developed a list of different Goodness-of-Fit (GoF) tests with different powers under different conditions. The usual practice is the selection of the best copula via the p-value of the GoF test. Although this method is not purely correct due to the fact that non-rejection does not imply acception, this strategy is favored by practitioners. Unfortunately, …


Working With Crsp/Compustat In R: Reproducible Empirical Asset Pricing, Majeed Simaan Jun 2021

Working With Crsp/Compustat In R: Reproducible Empirical Asset Pricing, Majeed Simaan

The R Journal

It is common to come across SAS or Stata manuals while working on academic empirical finance research. Nonetheless, given the popularity of open-source programming languages such as R, there are fewer resources in R covering popular databases such as CRSP and COMPUSTAT. The aim of this article is to bridge the gap and illustrate how to leverage R in working with both datasets. As an application, we illustrate how to form size-value portfolios with respect to Fama and French (1993) and study the sensitivity of the results with respect to different inputs. Ultimately, the purpose of the article is to …


Garchx: Flexible And Robust Garch-X Modeling, Genaro Sucarrat Jun 2021

Garchx: Flexible And Robust Garch-X Modeling, Genaro Sucarrat

The R Journal

The garchx package provides a user-friendly, fast, flexible, and robust framework for the estimation and inference of GARCH(p, q,r)-X models, where p is the ARCH order, q is the GARCH order, r is the asymmetry or leverage order, and ’X’ indicates that covariates can be included. Quasi Maximum Likelihood (QML) methods ensure estimates are consistent and standard errors valid, even when the standardized innovations are non-normal or dependent, or both. Zero-coefficient restrictions by omission enable parsimonious specifications, and functions to facilitate the non-standard inference associated with zero-restrictions in the null-hypothesis are provided. Finally, in the formal comparisons of …


Onestep: Le Cam's One-Step Estimation Procedure, Alexandre Brouste, Christophe Dutang, Darel Noutsa Mieniedou Jun 2021

Onestep: Le Cam's One-Step Estimation Procedure, Alexandre Brouste, Christophe Dutang, Darel Noutsa Mieniedou

The R Journal

The OneStep package proposes principally an eponymic function that numerically computes Le Cam’s one-step estimator, which is asymptotically efficient and can be computed faster than the maximum likelihood estimator for large datasets. Monte Carlo simulations are carried out for several examples (discrete and continuous probability distributions) in order to exhibit the performance of Le Cam’s one-step estimation procedure in terms of efficiency and computational cost on observation samples of finite size.


Wide-To-Tall Data Reshaping Using Regular Expressions And The Nc Package, Toby Dylan Hocking Jun 2021

Wide-To-Tall Data Reshaping Using Regular Expressions And The Nc Package, Toby Dylan Hocking

The R Journal

Regular expressions are powerful tools for extracting tables from non-tabular text data. Capturing regular expressions that describe the information to extract from column names can be especially useful when reshaping a data table from wide (few rows with many regularly named columns) to tall (fewer columns with more rows). We present the R package nc (short for named capture), which provides functions for wide-to-tall data reshaping using regular expressions. We describe the main new ideas of nc, and provide detailed comparisons with related R packages (stats, utils, data.table, tidyr, tidyfast, tidyfst, reshape2, cdata).


Stratamatch: Prognostic Score Stratification Using A Pilot Design, Rachael C. Aikens, Joseph Rigdon, Justin Lee, Michael Baiocchi, Andrew B. Goldstone, Peter Chiu, Y Joseph Woo, Jonathan H. Chen Jun 2021

Stratamatch: Prognostic Score Stratification Using A Pilot Design, Rachael C. Aikens, Joseph Rigdon, Justin Lee, Michael Baiocchi, Andrew B. Goldstone, Peter Chiu, Y Joseph Woo, Jonathan H. Chen

The R Journal

Optimal propensity score matching has emerged as one of the most ubiquitous approaches for causal inference studies on observational data. However, outstanding critiques of the statistical properties of propensity score matching have cast doubt on the statistical efficiency of this technique, and the poor scalability of optimal matching to large data sets makes this approach inconvenient if not infeasible for sample sizes that are increasingly commonplace in modern observational data. The stratamatch package provides implementation support and diagnostics for ‘stratified matching designs,’ an approach that addresses both of these issues with optimal propensity score matching for large-sample observational studies. First, …


Conversations In Time: Interactive Visualization To Explore Structured Temporal Data, Earo Wang, Dianne Cook Jun 2021

Conversations In Time: Interactive Visualization To Explore Structured Temporal Data, Earo Wang, Dianne Cook

The R Journal

Temporal data often has a hierarchical structure, defined by categorical variables describing different levels, such as political regions or sales products. The nesting of categorical variables produces a hierarchical structure. The tsibbletalk package is developed to allow a user to interactively explore temporal data, relative to the nested or crossed structures. It can help to discover differences between category levels, and uncover interesting periodic or aperiodic slices. The package implements a shared tsibble object that allows for linked brushing between coordinated views, and a shiny module that aids in wrapping timelines for seasonal patterns. The tools are demonstrated using two …


Automating Reproducible, Collaborative Clinical Trial Document Generation With The Listdown Package, Michael Kane, Xun Jiang, Simon Urbanek Jun 2021

Automating Reproducible, Collaborative Clinical Trial Document Generation With The Listdown Package, Michael Kane, Xun Jiang, Simon Urbanek

The R Journal

the conveyance of clinical trial explorations and analysis results from a statistician to a clinical investigator is a critical component of the drug development and clinical research cycle. Automating the process of generating documents for data descriptions, summaries, exploration, and analysis allows the statistician to provide a more comprehensive view of the information captured by a clinical trial, and efficient generation of these documents allows the statistican to focus more on the conceptual development of a trial or trial analysis and less on the implementation of the summaries and results on which decisions are made. This paper explores the use …


Editorial, Dianne Cook Jun 2021

Editorial, Dianne Cook

The R Journal

First, some news about the journal board. Welcome to Gavin Simpson, who joins as a new Executive Editor! In addition, welcome to our new Associate Editors Nicholas Tierney, Isabella Gollini, Rasmus Bååth, Mark van der Loo, Elizabeth Sweeney, Louis Aslett and Katarina Domijan. With the large volume of submissions, the Associate Editors now play a vital role in processing articles.


Penphcure: Variable Selection In Proportional Hazards Cure Model With Time-Varying Covariates, Alessandro Beretta, Cédric Heuchenne Jun 2021

Penphcure: Variable Selection In Proportional Hazards Cure Model With Time-Varying Covariates, Alessandro Beretta, Cédric Heuchenne

The R Journal

We describe the penPHcure R package, which implements the semiparametric proportional-hazards (PH) cure model of Sy and Taylor (2000) extended to time-varying covariates and the variable selection technique based on its SCAD-penalized likelihood proposed by Beretta and Heuchenne (2019a). In survival analysis, cure models are a useful tool when a fraction of the population is likely to be immune from the event of interest. They can separate the effects of certain factors on the probability of being susceptible and on the time until the occurrence of the event. Moreover, the penPHcure package allows the user to simulate data from a …


The R Journal (December 2020) 12(2): Complete Issue, The R Foundation Dec 2020

The R Journal (December 2020) 12(2): Complete Issue, The R Foundation

The R Journal

Editorial, Michael J. Kane

Contributed Research Articles

The biglasso Package: A Memory- and Computation-Efficient Solver for Lasso Model Fitting with Big Data in R, Yaohui Zeng and Patrick Breheny

Comparing Multiple Survival Functions with Crossing Hazards in R, Hsin-wen Chang, Pei-Yuan Tsai, Jen-Tse Kao, and Guo-You Lan

A Unified Algorithm for the Non-Convex Penalized Estimation: The ncpen Package, Dongshin Kim, Sangin Lee, and Sunghoon Kwon

TULIP: A Toolbox for Linear Discriminant Analysis with Penalties, Yuqing Pan, Qing Mai, and Xin Zhang

fitzRoy: An R Package to Encourage Reproducible Sports Analysis, Robert Nguyen, James Day, David Warton, and Oscar Lane

Assembling …


Changes In R 3.6–4.0, Tomas Kalibera, Sebastian Meyer, Kurt Hornik Dec 2020

Changes In R 3.6–4.0, Tomas Kalibera, Sebastian Meyer, Kurt Hornik

The R Journal

We give a selection of the most important changes in R 4.0.0 and in the R 3.6 release series. Some statistics on source code commits and bug tracking activities are also provided.


Analyzing Basket Trials Under Multisource Exchangeability Assumptions, Michael J. Kane, Nan Chen, Alexander M. Kaizer, Xun Jiang, H Amy Xia, Brian P. Hobbs Dec 2020

Analyzing Basket Trials Under Multisource Exchangeability Assumptions, Michael J. Kane, Nan Chen, Alexander M. Kaizer, Xun Jiang, H Amy Xia, Brian P. Hobbs

The R Journal

Basket designs are prospective clinical trials that are devised with the hypothesis that the presence of selected molecular features determine a patient’s subsequent response to a particular “targeted” treatment strategy. Basket trials are designed to enroll multiple clinical subpopulations to which it is assumed that the therapy in question offers beneficial efficacy in the presence of the targeted molecular profile. The treatment, however, may not offer acceptable efficacy to all subpopulations enrolled. Moreover, for rare disease settings, such as oncology wherein these trials have become popular, marginal measures of statistical evidence are difficult to interpret for sparsely enrolled subpopulations. Consequently, …


Motbfs: An R Package For Learning Hybrid Bayesian Networks Using Mixtures Of Truncated Basis Functions, Inmaculada Pérez-Bernabé, Ana D. Maldonado, Antonio Salmerón, Thomas D. Nielsen Dec 2020

Motbfs: An R Package For Learning Hybrid Bayesian Networks Using Mixtures Of Truncated Basis Functions, Inmaculada Pérez-Bernabé, Ana D. Maldonado, Antonio Salmerón, Thomas D. Nielsen

The R Journal

This paper introduces MoTBFs, an R package for manipulating mixtures of truncated basis functions. This class of functions allows the representation of joint probability distributions involving discrete and continuous variables simultaneously, and includes mixtures of truncated exponentials and mixtures of polynomials as special cases. The package implements functions for learning the parameters of univariate, multivariate, and conditional distributions, and provides support for parameter learning in Bayesian networks with both discrete and continuous variables. Probabilistic inference using forward sampling is also implemented. Part of the functionality of the MoTBFs package relies on the bnlearn package, which includes functions for learning the …


A Graphical Eda Tool With Ggplot2: Brinton, Pere Millán-Martínez, Ramon Oller Dec 2020

A Graphical Eda Tool With Ggplot2: Brinton, Pere Millán-Martínez, Ramon Oller

The R Journal

We present brinton package, which we developed for graphical exploratory data analysis in R. Based on ggplot2, gridExtra and rmarkdown, brinton package introduces wideplot() graphics for exploring the structure of a dataset through a grid of variables and graphic types. It also introduces longplot() graphics, which present the entire catalog of available graphics for representing a particular variable using a grid of graphic types and variations on these types. Finally, it introduces the plotup() function, which complements the previous two functions in that it presents a particular graphic for a specific variable of a dataset. This set of functions is …


Nts: An R Package For Nonlinear Time Series Analysis, Xialu Liu, Rong Chen, Ruey Tsay Dec 2020

Nts: An R Package For Nonlinear Time Series Analysis, Xialu Liu, Rong Chen, Ruey Tsay

The R Journal

Linear time series models are commonly used in analyzing dependent data and in forecasting. On the other hand, real phenomena often exhibit nonlinear behavior and the observed data show nonlinear dynamics. This paper introduces the R package NTS that offers various computational tools and nonlinear models for analyzing nonlinear dependent data. The package fills the gaps of several outstanding R packages for nonlinear time series analysis. Specifically, the NTS package covers the implementation of threshold autoregressive (TAR) models, autoregressive conditional mean models with exogenous variables (ACMx), functional autoregressive models, and state-space models. Users can also evaluate and compare the performance …


Aquadtree: An R Package For Quadtree Anonymization Of Point Data, Raymond Lagonigro, Ramon Oller, Joan Carles Martori Dec 2020

Aquadtree: An R Package For Quadtree Anonymization Of Point Data, Raymond Lagonigro, Ramon Oller, Joan Carles Martori

The R Journal

The demand for precise data for analytical purposes grows rapidly among the research community and decision makers as more geographic information is being collected. Laws protecting data privacy are being enforced to prevent data disclosure. Statistical institutes and agencies need methods to preserve confidentiality while maintaining accuracy when disclosing geographic data. In this paper we present the AQuadtree package, a software intended to produce and deal with official spatial data making data privacy and accuracy compatible. The lack of specific methods in R to anonymize spatial data motivated the development of this package, providing an automatic aggregation tool to anonymize …


Kspm: A Package For Kernel Semi-Parametric Models, Catherine Schramm, Sébastien Jacquemont, Karim Oualkacha, Aurélie Labbe, Celia M. T. Greenwood Dec 2020

Kspm: A Package For Kernel Semi-Parametric Models, Catherine Schramm, Sébastien Jacquemont, Karim Oualkacha, Aurélie Labbe, Celia M. T. Greenwood

The R Journal

Kernel semi-parametric models and their equivalence with linear mixed models provide analysts with the flexibility of machine learning methods and a foundation for inference and tests of hypothesis. These models are not impacted by the number of predictor variables, since the kernel trick transforms them to a kernel matrix whose size only depends on the number of subjects. Hence, methods based on this model are appealing and numerous, however only a few R programs are available and none includes a complete set of features. Here, we present the KSPM package to fit the kernel semi-parametric model and its extensions in …


Ordinalclust: An R Package To Analyze Ordinal Data, Margot Selosse, Julien Jacques, Christophe Biernacki Dec 2020

Ordinalclust: An R Package To Analyze Ordinal Data, Margot Selosse, Julien Jacques, Christophe Biernacki

The R Journal

Ordinal data are used in many domains, especially when measurements are collected from people through observations, tests, or questionnaires. ordinalClust is an innovative R package dedicated to ordinal data that provides tools for modeling, clustering, co-clustering and classifying such data. Ordinal data are modeled using the BOS distribution, which is a model with two meaningful parameters referred to as "position" and "precision". The former indicates the mode of the distribution and the latter describes how scattered the data are around the mode: the user is able to easily interpret the distribution of their data when given these two parameters. The …


A Fast And Scalable Implementation Method For Competing Risks Data With The R Package Fastcmprsk, Eric S. Kawaguchi, Jenny I. Shen, Gang Li, Marc A. Suchard Dec 2020

A Fast And Scalable Implementation Method For Competing Risks Data With The R Package Fastcmprsk, Eric S. Kawaguchi, Jenny I. Shen, Gang Li, Marc A. Suchard

The R Journal

Advancements in medical informatics tools and high-throughput biological experimentation make large-scale biomedical data routinely accessible to researchers. Competing risks data are typical in biomedical studies where individuals are at risk to more than one cause (type of event) which can preclude the others from happening. The Fine and Gray (1999) proportional subdistribution hazards model is a popular and well-appreciated model for competing risks data and is currently implemented in a number of statistical software packages. However, current implementations are not computationally scalable for large-scale competing risks data. We have developed an R package, fastcmprsk, that uses a novel forward-backward scan …


Six Years Of Shiny In Research: Collaborative Development Of Web Tools In R, Peter Kasprzak, Lachlan Mitchell, Olena Kravchuk, Andy Timmins Dec 2020

Six Years Of Shiny In Research: Collaborative Development Of Web Tools In R, Peter Kasprzak, Lachlan Mitchell, Olena Kravchuk, Andy Timmins

The R Journal

The use of Shiny in research publications is investigated over the six and a half years since the appearance of this popular web application framework for R, which has been utilised in many varied research areas. While it is demonstrated that the complexity of Shiny applications is limited by the background architecture, and real security concerns exist for novice app developers, the collaborative benefits are worth attention from the wider research community. Shiny simplifies the use of complex methodologies for people of different specialities, at the level of proficiency appropriate for the end user. This enables a diverse community of …


Assembling Pharmacometric Datasets In R: The Puzzle Package, Mario González-Sales, Olivier Barrière, Pierre Olivier Tremblay, Guillaume Bonnefois, Julie Desrochers, Fahima Nekka Dec 2020

Assembling Pharmacometric Datasets In R: The Puzzle Package, Mario González-Sales, Olivier Barrière, Pierre Olivier Tremblay, Guillaume Bonnefois, Julie Desrochers, Fahima Nekka

The R Journal

Pharmacometric analyses are integral components of the drug development process. The core of each pharmacometric analysis is a dataset. The time required to construct a pharmacometrics dataset can sometimes be higher than the effort required for the modeling per se. To simplify the process, the puzzle R package has been developed aimed at simplifying and facilitating the time consuming and error prone task of assembling pharmacometrics datasets.

Puzzle consist of a series of functions written in R. These functions create, from tabulated files, datasets that are compatible with the formatting requirements of the gold standard non-linear mixed effects modeling …


Fitzroy: An R Package To Encourage Reproducible Sports Analysis, Robert Nguyen, James Day, David Warton, Oscar Lane Dec 2020

Fitzroy: An R Package To Encourage Reproducible Sports Analysis, Robert Nguyen, James Day, David Warton, Oscar Lane

The R Journal

The importance of reproducibility, and the related issue of open access to data, has received a lot of recent attention. Momentum on these issues is gathering in the sports analytics community. While Australian Rules football (AFL) is the leading commercial sport in Australia, unlike popular international sports, there has been no mechanism for the public to access comprehensive statistics on players and teams. Expert commentary currently relies heavily on data that isn’t made readily accessible and this produces an unnecessary barrier for the development of an inclusive sports analytics community. We present the R package fitzRoy to provide easy access …


Comparing Multiple Survival Functions With Crossing Hazards In R, Hsin-Wen Chang, Pei-Yuan Tsai, Jen-Tse Kao, Guo-You Lan Dec 2020

Comparing Multiple Survival Functions With Crossing Hazards In R, Hsin-Wen Chang, Pei-Yuan Tsai, Jen-Tse Kao, Guo-You Lan

The R Journal

It is frequently of interest in time-to-event analysis to compare multiple survival functions nonparametrically. However, when the hazard functions cross, tests in existing R packages do not perform well. To address the issue, we introduce the package survELtest, which provides tests for comparing multiple survival functions with possibly crossing hazards. Due to its powerful likelihood ratio formulation, this is the only R package to date that works when the hazard functions cross. We illustrate the use of the procedures in survELtest by applying them to data from randomized clinical trials and simulated datasets. We show that these methods lead …


The Biglasso Package: A Memory- And Computation-Efficient Solver For Lasso Model Fitting With Big Data In R, Yaohui Zeng, Patrick Breheny Dec 2020

The Biglasso Package: A Memory- And Computation-Efficient Solver For Lasso Model Fitting With Big Data In R, Yaohui Zeng, Patrick Breheny

The R Journal

Penalized regression models such as the lasso have been extensively applied to analyzing high-dimensional data sets. However, due to memory limitations, existing R packages like glmnet and ncvreg are not capable of fitting lasso-type models for ultrahigh-dimensional, multi-gigabyte data sets that are increasingly seen in many areas such as genetics, genomics, biomedical imaging, and high-frequency finance. In this research, we implement an R package called biglasso that tackles this challenge. biglasso utilizes memory-mapped files to store the massive data on the disk, only reading data into memory when necessary during model fitting, and is thus able to handle out-of-core computation …


Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis Dec 2020

Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis

The R Journal

In the past 4 months, 818 new packages were added to the CRAN package repository. 100 packages were archived and 248 were archived. The following shows the growth of the number of active packages in the CRAN package repository


R Foundation News, Torsten Hothorn Dec 2020

R Foundation News, Torsten Hothorn

The R Journal

Membership fees and donations received between 2020-09-09 and 2021-01-28.