Open Access. Powered by Scholars. Published by Universities.®

Programming Languages and Compilers Commons™

Open Access. Powered by Scholars. Published by Universities.®

1,845 Full-Text Articles 3,363 Authors 771,689 Downloads 137 Institutions

All Articles in Programming Languages and Compilers

Faceted Search

1,845 full-text articles. Page 30 of 79.

Similar: R Code Clone And Plagiarism Detection, Maciej Bartoszuk, Marek Gagolewski 2020 Warsaw University of Technology

Similar: R Code Clone And Plagiarism Detection, Maciej Bartoszuk, Marek Gagolewski

The R Journal

Third-party software for assuring source code quality is becoming increasingly popular. Tools that evaluate the coverage of unit tests, perform static code analysis, or inspect run-time memory use are crucial in the software development life cycle. More sophisticated methods allow for performing meta-analyses of large software repositories, e.g., to discover abstract topics they relate to or common design patterns applied by their developers. They may be useful in gaining a better understanding of the component interdependencies, avoiding cloned code as well as detecting plagiarism in programming classes.

Ameaningful measure of similarity of computer programs often forms the basis of such …


Coxphlb: An R Package For Analyzing Length Biased Data Under Cox Model, Chi Hyun Lee, Heng Zhou, Jing Ning, Diane D. Liu, Yu Shen 2020 University of Massachusetts Amherst

Coxphlb: An R Package For Analyzing Length Biased Data Under Cox Model, Chi Hyun Lee, Heng Zhou, Jing Ning, Diane D. Liu, Yu Shen

The R Journal

Data subject to length-biased sampling are frequently encountered in various applications including prevalent cohort studies and are considered as a special case of left-truncated data under the stationarity assumption. Many semiparametric regression methods have been proposed for length biased data to model the association between covariates and the survival outcome of interest. In this paper, we present a brief review of the statistical methodologies established for the analysis of length-biased data under the Cox model, which is the most commonly adopted semiparametric model, and introduce an R package CoxPhLb that implements these methods. Specifically, the package includes features such as …


The R Package Nonprobest For Estimation In Non-Probability Surveys, M. Rueda, R. Ferri-García, L. Castro 2020 University of Granada

The R Package Nonprobest For Estimation In Non-Probability Surveys, M. Rueda, R. Ferri-García, L. Castro

The R Journal

Different inference procedures are proposed in the literature to correct selection bias that might be introduced with non-random sampling mechanisms. The R package NonProbEst enables the estimation of parameters using some of these techniques to correct selection bias in non-probability surveys. The mean and the total of the target variable are estimated using Propensity Score Adjustment, calibration, statistical matching, model-based, model-assisted and model-calibratated techniques. Confidence intervals can also obtained for each method. Machine learning algorithms can be used for estimating the propensities or for predicting the unknown values of the target variable for the non-sampled units. Variance of a given …


Variable Importance Plots: An Introduction To The Vip Package, Brandon M. Bartoszuk, Marek Gagolewski 2020 University of Cincinnati

Variable Importance Plots: An Introduction To The Vip Package, Brandon M. Bartoszuk, Marek Gagolewski

The R Journal

In the era of “big data”, it is becoming more of a challenge to not only build state-of-the-art predictive models, but also gain an understanding of what’s really going on in the data. For example, it is often of interest to know which, if any, of the predictors in a fitted model are relatively influential on the predicted outcome. Some modern algorithms—like random forests (RFs) and gradient boosted decision trees (GBMs)—have a natural way of quantifying the importance or relative influence of each feature. Other algorithms—like naive Bayes classifiers and support vector machines—are not capable of doing so and model-agnostic …


Tsmp: An R Package For Time Series With Matrix Profile, Francisco Bischoff, Pedro Pereira Rodriques 2020 Faculty of Medicine of the University of Porto

Tsmp: An R Package For Time Series With Matrix Profile, Francisco Bischoff, Pedro Pereira Rodriques

The R Journal

This article describes tsmp, an R package that implements the MP concept for TS. The tsmp package is a toolkit that allows all-pairs similarity joins, motif, discords and chains discovery, semantic segmentation, etc. Here we describe how the tsmp package may be used by showing some of the use-cases from the original articles and evaluate the algorithm speed in the R environment. This package can be downloaded at https://CRAN.R-project.org/package=tsmp.


R Foundation News, Torsten Hothorn 2020 Universität Zürich

R Foundation News, Torsten Hothorn

The R Journal

Membership fees and donations received between 2020-02-24 and 2020-09-08.


Projectmanagement: An R Package For Managing Projects, Juan Carlos Gonçalves-Dosantos, Ignacio García-Jurado, Julián Costa 2020 Universidade da Coruña

Projectmanagement: An R Package For Managing Projects, Juan Carlos Gonçalves-Dosantos, Ignacio García-Jurado, Julián Costa

The R Journal

Project management is an important body of knowledge and practices that comprises the planning, organisation and control of resources to achieve one or more pre-determined objectives. In this paper, we introduce ProjectManagement, a new R package that provides the necessary tools to manage projects in a broad sense, and illustrate its use by examples.


Npordtests: An R Package Of Nonparametric Tests For Equality Of Location Against Ordered Alternatives, Bulent Altunkaynak, Hamza Gamgam 2020 Gazi University Faculty of Science

Npordtests: An R Package Of Nonparametric Tests For Equality Of Location Against Ordered Alternatives, Bulent Altunkaynak, Hamza Gamgam

The R Journal

Ordered alternatives are an important statistical problem in many situation such as increased risk of congenital malformation caused by excessive alcohol consumption during pregnancy life test experiments, drug-screening studies, dose-finding studies, the dose-response studies, age-related response. There are numerous other examples of this nature. In this paper, we present the npordtests package to test the equality of locations for ordered alternatives. The package includes the Jonckheere Terpstra, Beier and Buning’s Adaptive, Modified Jonckheere-Terpstra, Terpstra-Magel, Ferdhiana Terpstra-Magel, KTP, S and Gaur’s Gc tests. A simulation study is conducted to determine which test is the most appropriate test for which scenario and …


Spinifex: An R Package For Creating A Manual Tour Of Low-Dimensional Projections Of Multivariate Data, Nicholas Spyrison, Dianne Cook 2020 Monash University

Spinifex: An R Package For Creating A Manual Tour Of Low-Dimensional Projections Of Multivariate Data, Nicholas Spyrison, Dianne Cook

The R Journal

Dynamic low-dimensional linear projections of multivariate data collectively known as tours provide an important tool for exploring multivariate data and models. The R package tourr provides functions for several types of tours: grand, guided, little, local and frozen. Each of these can be viewed dynamically, or saved into a data object for animation. This paper describes a new package, spinifex, which provides a manual tour of multivariate data where the projection coefficient of a single variable is controlled. The variable is rotated fully into the projection, or completely out of the projection. The resulting sequence of projections can be …


Mistr: A Computational Framework For Mixture And Composite Distributions, Lukas Sablica, Kurt Hornik 2020 Vienna University of Economics and Business

Mistr: A Computational Framework For Mixture And Composite Distributions, Lukas Sablica, Kurt Hornik

The R Journal

No abstract provided.


Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis 2020 WU Wirtschaftsuniversität Wien

Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis

The R Journal

In the past 8 months, 1554 new packages were added to the CRAN package repository. 96 packages were unarchived and 843 were archived. The following shows the growth of the number of active packages in the CRAN package repository:


Skew-T Expected Information Matrix Evaluation And Use For Standard Error Calculations, R. Douglas Martin, Chindhanai Uthaisaad, Daniel Z. Xia 2020 University of Washington

Skew-T Expected Information Matrix Evaluation And Use For Standard Error Calculations, R. Douglas Martin, Chindhanai Uthaisaad, Daniel Z. Xia

The R Journal

Skew-t distributions derived from skew-normal distributions, as developed by Azzalini and several co-workers, are popular because of their theoretical foundation and the availability of computational methods in the R package sn. One difficulty with this skew-t family is that the elements of the expected information matrix do not have closed form analytic formulas. Thus, we developed a numerical integration method of computing the expected information matrix in the R package skewtInfo. The accuracy of our expected information matrix calculation method was confirmed by comparing the result with that obtained using an observed information matrix for a very large sample …


Tools For Analyzing R Code The Tidy Way, Lucy D'Agostino McGowan, Sean Kross, Jeffrey Leek 2020 Wake Forest University

Tools For Analyzing R Code The Tidy Way, Lucy D'Agostino Mcgowan, Sean Kross, Jeffrey Leek

The R Journal

With the current emphasis on reproducibility and replicability, there is an increasing need to examine how data analyses are conducted. In order to analyze the between researcher variability in data analysis choices as well as the aspects within the data analysis pipeline that contribute to the variability in results, we have created two R packages: matahari and tidycode. These packages build on methods created for natural language processing; rather than allowing for the processing of natural language, we focus on R code as the substrate of interest. The matahari package facilitates the logging of everything that is typed in the …


Individual-Level Modelling Of Infectious Disease Data: Epiilm, Vineetha Warriyar, Waleed Almutiry, Rob Deardon 2020 University of Calgary

Individual-Level Modelling Of Infectious Disease Data: Epiilm, Vineetha Warriyar, Waleed Almutiry, Rob Deardon

The R Journal

In this article we introduce the R package EpiILM, which provides tools for simulation from, and inference for, discrete-time individual-level models of infectious disease transmission proposed by Deardon et al. (2010). The inference is set in a Bayesian framework and is carried out via Metropolis Hastings Markov chain Monte Carlo (MCMC). For its fast implementation, key functions are coded in Fortran. Both spatial and contact network models are implemented in the package and can be set in either susceptible-infected (SI) or susceptible-infected-removed (SIR) compartmental frameworks. Use of the package is demonstrated through examples involving both simulated and real data.


Conference Report: Why R? 2019, Michał Burdukiewicz, Filip Pietluch, Jarosław Chilimoniuk, Katarzyna Sidorczuk, Dominik Rafacz, Leon Eyrich Jessen, Stefan Rödiger, Marcin Kosiński, Piotr Wójcik 2020 Warsaw University of Technology

Conference Report: Why R? 2019, Michał Burdukiewicz, Filip Pietluch, Jarosław Chilimoniuk, Katarzyna Sidorczuk, Dominik Rafacz, Leon Eyrich Jessen, Stefan Rödiger, Marcin Kosiński, Piotr Wójcik

The R Journal

WhyR?conferences have been the hallmark of the Why R? Foundation (whyr.pl). Our goal has been to establish a series of international R-related events in Poland. After three years, weare happy to announce that our main event, the Why R? conference, has become one of the largest annual R conferences in Central Europe.


Difnlr: Generalized Logistic Regression Models For Dif And Ddf Detection, Adéla Hladká, Patrícia Martinková 2020 Institute of Computer Science of the Czech Academy of Sciences, Charles University

Difnlr: Generalized Logistic Regression Models For Dif And Ddf Detection, Adéla Hladká, Patrícia Martinková

The R Journal

Differential item functioning (DIF) and differential distractor functioning (DDF) are impor tant topics in psychometrics, pointing to potential unfairness in items with respect to minorities or different social groups. Various methods have been proposed to detect these issues. The difNLR R package extends DIF methods currently provided in other packages by offering approaches based on generalized logistic regression models that account for possible guessing or inattention, and by pro viding methods to detect DIF and DDF among ordinal and nominal data. In the current paper, we describe implementation of the main functions of the difNLR package, from data generation, through …


Survboost: An R Package For High-Dimensional Variable Selection In The Stratified Proportional Hazards Model Via Gradient Boosting, Emily Morris, Kevin He, Yanming Li, Yi Li, Jian Kang 2020 University of Michigan

Survboost: An R Package For High-Dimensional Variable Selection In The Stratified Proportional Hazards Model Via Gradient Boosting, Emily Morris, Kevin He, Yanming Li, Yi Li, Jian Kang

The R Journal

High-dimensional variable selection in the proportional hazards (PH) model has many successful applications in different areas. In practice, data may involve confounding variables that do not satisfy the PH assumption, in which case the stratified proportional hazards (SPH) model can be adopted to control the confounding effects by stratification without directly modeling the confounding effects. However, there is a lack of computationally efficient statistical software for high-dimensional variable selection in the SPH model. In this work an R package, SurvBoost, is developed to implement the gradient boosting algorithm for fitting the SPH model with high-dimensional covariate variables. Simulation studies …


Copulacenr: Copula Based Regression Models For Bivariate Censored Data In R, Tao Sun, Ying Ding 2020 Renmin University of China, University of Pittsburgh

Copulacenr: Copula Based Regression Models For Bivariate Censored Data In R, Tao Sun, Ying Ding

The R Journal

Bivariate time-to-event data frequently arise in research areas such as clinical trials and epidemiological studies, where the occurrence of two events are correlated. In many cases, the exact event times are unknown due to censoring. The copula model is a popular approach for modeling correlated bivariate censored data, in which the two marginal distributions and the between margin dependence are modeled separately. This article presents the R package CopulaCenR, which is designed for modeling and testing bivariate data under right or (general) interval censoring in a regression setting. It provides a variety of Archimedean copula functions including a flexible two-parameter …


Sortedeffects: Sorted Causal Effects In R, Schuowen Chen, Victor Chernozhukov, Iván Fernández-Val, Ye Luo 2020 Boston University

Sortedeffects: Sorted Causal Effects In R, Schuowen Chen, Victor Chernozhukov, Iván Fernández-Val, Ye Luo

The R Journal

Chernozhukov et al. (2018) proposed the sorted effect method for nonlinear regression models. This method consists of reporting percentiles of the partial effects, the sorted effects, in addition to the average effect commonly used to summarize the heterogeneity in the partial effects. They also propose to use the sorted effects to carry out classification analysis where the observational units are classified as most and least affected if their partial effect are above or below some tail sorted effects. The R package SortedEffects implements the estimation and inference methods therein and provides tools to visualize the results. This vignette serves as …


A Virtualization Based System Infrastructure For Dynamic Program Analysis, Jiaqi HONG 2020 Singapore Management University

A Virtualization Based System Infrastructure For Dynamic Program Analysis, Jiaqi Hong

Dissertations and Theses Collection (Open Access)

Dynamic malware analysis schemes either run the target program as is in an isolated environment assisted by additional hardware facilities or modify it with instrumentation code statically or dynamically. The hardware-assisted schemes usually trap the target during its execution to a more privileged environment based on the available hardware events. The more privileged environment is not accessible by the untrusted kernel, thus this approach is often applied for transparent and secure kernel analysis. Nevertheless, the isolated environment induces a virtual address gap between the analyzer and the target, which hinders effective and efficient memory introspection and undermines the correctness of …


Digital Commons powered by bepress