Towards A Complete Formal Semantics Of Rust,
2021
California Polytechnic State University, San Luis Obispo
Towards A Complete Formal Semantics Of Rust, Alexa White
Master's Theses
Rust is a relatively new programming language with a unique memory model designed to provide the ease of use of a high-level language as well as the power and control of a low-level language while preserving memory safety. In order to prove the safety and correctness of Rust and to provide analysis tools for its use cases, it is necessary to construct a formal semantics of the language. Existing efforts to construct such a semantic model are limited in their scope and none to date have successfully captured the complete functionality of the language. This thesis focuses on the K-Rust …
Qlens: Visual Analytics Of Multi-Step Problem-Solving Behaviors For Improving Question Design,
2021
Hong Kong University of Science and Technology
Qlens: Visual Analytics Of Multi-Step Problem-Solving Behaviors For Improving Question Design, Meng Xia, Reshika P. Velumani, Yong Wang, Huamin Qu, Xiaojuan Ma
Research Collection School Of Computing and Information Systems
With the rapid development of online education in recent years, there has been an increasing number of learning platforms that provide students with multi-step questions to cultivate their problem-solving skills. To guarantee the high quality of such learning materials, question designers need to inspect how students’ problem-solving processes unfold step by step to infer whether students’ problem-solving logic matches their design intent. They also need to compare the behaviors of different groups (e.g., students from different grades) to distribute questions to students with the right level of knowledge. The availability of fine-grained interaction data, such as mouse movement trajectories from …
Using Torchattacks To Improve The Robustness Of Models With Adversarial Training,
2021
Universidad Interamericana de Puerto Rico - Barranquitas
Using Torchattacks To Improve The Robustness Of Models With Adversarial Training, William S. Matos Díaz
Cybersecurity: Deep Learning Driven Cybersecurity Research in a Multidisciplinary Environment
Adversarial training has proven to be one of the most successful ways to defend models against adversarial examples. This process consists of training a model with an adversarial example to improve the robustness of the model. In this experiment, Torchattacks, a Pytorch library made for importing adversarial examples more easily, was used to determine which attack was the strongest. Later on, the strongest attack was used to train the model and make it more robust against adversarial examples. The datasets used to perform the experiments were MNIST and CIFAR-10. Both datasets were put to the test using PGD, FGSM, and …
Source Code Comment Classification Artificial Intelligence,
2021
The University of Akron
Source Code Comment Classification Artificial Intelligence, Cole Sutyak
Williams Honors College, Honors Research Projects
Source code comment classification is an important problem for future machine learning solutions. In particular, supervised machine learning solutions that have largely subjective data labels but are difficult to obtain the labels for. Machine learning problems are problems largely because of a lack of data. In machine learning solutions, it is better to have a large amount of mediocre data than it is to have a small amount of good data. While the mediocre data might not produce the best accuracy, it produces the best results because there is much more to learn from the problem.
In this project, data …
The R Journal (December 2020) 12(2): Complete Issue,
2020
University of Nebraska - Lincoln
The R Journal (December 2020) 12(2): Complete Issue, The R Foundation
The R Journal
Editorial, Michael J. Kane
Contributed Research Articles
The biglasso Package: A Memory- and Computation-Efficient Solver for Lasso Model Fitting with Big Data in R, Yaohui Zeng and Patrick Breheny
Comparing Multiple Survival Functions with Crossing Hazards in R, Hsin-wen Chang, Pei-Yuan Tsai, Jen-Tse Kao, and Guo-You Lan
A Unified Algorithm for the Non-Convex Penalized Estimation: The ncpen Package, Dongshin Kim, Sangin Lee, and Sunghoon Kwon
TULIP: A Toolbox for Linear Discriminant Analysis with Penalties, Yuqing Pan, Qing Mai, and Xin Zhang
fitzRoy: An R Package to Encourage Reproducible Sports Analysis, Robert Nguyen, James Day, David Warton, and Oscar Lane
Assembling …
On The Generation, Structure, And Semantics Of Grammar Patterns In Source Code Identifiers,
2020
Kent State University
On The Generation, Structure, And Semantics Of Grammar Patterns In Source Code Identifiers, Christian D. Newman,, Reem S. Alsuhaibani, Michael J. Decker, Anthony Peruma, Dishant Kaushik, Mohamed Wiem Mkaouer, Emily Hill
Articles
Identifier names are the atoms of program comprehension. Weak identifier names decrease developer productivity and degrade the performance of automated approaches that leverage identifier names in source code analysis; threatening many of the advantages which stand to be gained from advances in artificial intelligence and machine learning. Therefore, it is vital to support developers in naming and renaming identifiers. In this paper, we extend our prior work, which studies the primary method through which names evolve: rename refactorings. In our prior work, we contextualize rename changes by examining commit messages and other refactorings. In this extension, we further consider data …
Changes In R 3.6–4.0,
2020
Czech Technical University
Changes In R 3.6–4.0, Tomas Kalibera, Sebastian Meyer, Kurt Hornik
The R Journal
We give a selection of the most important changes in R 4.0.0 and in the R 3.6 release series. Some statistics on source code commits and bug tracking activities are also provided.
Analyzing Basket Trials Under Multisource Exchangeability Assumptions,
2020
Yale University
Analyzing Basket Trials Under Multisource Exchangeability Assumptions, Michael J. Kane, Nan Chen, Alexander M. Kaizer, Xun Jiang, H Amy Xia, Brian P. Hobbs
The R Journal
Basket designs are prospective clinical trials that are devised with the hypothesis that the presence of selected molecular features determine a patient’s subsequent response to a particular “targeted” treatment strategy. Basket trials are designed to enroll multiple clinical subpopulations to which it is assumed that the therapy in question offers beneficial efficacy in the presence of the targeted molecular profile. The treatment, however, may not offer acceptable efficacy to all subpopulations enrolled. Moreover, for rare disease settings, such as oncology wherein these trials have become popular, marginal measures of statistical evidence are difficult to interpret for sparsely enrolled subpopulations. Consequently, …
Motbfs: An R Package For Learning Hybrid Bayesian Networks Using Mixtures Of Truncated Basis Functions,
2020
University of Almería
Motbfs: An R Package For Learning Hybrid Bayesian Networks Using Mixtures Of Truncated Basis Functions, Inmaculada Pérez-Bernabé, Ana D. Maldonado, Antonio Salmerón, Thomas D. Nielsen
The R Journal
This paper introduces MoTBFs, an R package for manipulating mixtures of truncated basis functions. This class of functions allows the representation of joint probability distributions involving discrete and continuous variables simultaneously, and includes mixtures of truncated exponentials and mixtures of polynomials as special cases. The package implements functions for learning the parameters of univariate, multivariate, and conditional distributions, and provides support for parameter learning in Bayesian networks with both discrete and continuous variables. Probabilistic inference using forward sampling is also implemented. Part of the functionality of the MoTBFs package relies on the bnlearn package, which includes functions for learning the …
A Graphical Eda Tool With Ggplot2: Brinton,
2020
Servei Català de Trànsit, Universitat de Vic
A Graphical Eda Tool With Ggplot2: Brinton, Pere Millán-Martínez, Ramon Oller
The R Journal
We present brinton package, which we developed for graphical exploratory data analysis in R. Based on ggplot2, gridExtra and rmarkdown, brinton package introduces wideplot() graphics for exploring the structure of a dataset through a grid of variables and graphic types. It also introduces longplot() graphics, which present the entire catalog of available graphics for representing a particular variable using a grid of graphic types and variations on these types. Finally, it introduces the plotup() function, which complements the previous two functions in that it presents a particular graphic for a specific variable of a dataset. This set of functions is …
Nts: An R Package For Nonlinear Time Series Analysis,
2020
San Diego State University
Nts: An R Package For Nonlinear Time Series Analysis, Xialu Liu, Rong Chen, Ruey Tsay
The R Journal
Linear time series models are commonly used in analyzing dependent data and in forecasting. On the other hand, real phenomena often exhibit nonlinear behavior and the observed data show nonlinear dynamics. This paper introduces the R package NTS that offers various computational tools and nonlinear models for analyzing nonlinear dependent data. The package fills the gaps of several outstanding R packages for nonlinear time series analysis. Specifically, the NTS package covers the implementation of threshold autoregressive (TAR) models, autoregressive conditional mean models with exogenous variables (ACMx), functional autoregressive models, and state-space models. Users can also evaluate and compare the performance …
Aquadtree: An R Package For Quadtree Anonymization Of Point Data,
2020
University of Vic
Aquadtree: An R Package For Quadtree Anonymization Of Point Data, Raymond Lagonigro, Ramon Oller, Joan Carles Martori
The R Journal
The demand for precise data for analytical purposes grows rapidly among the research community and decision makers as more geographic information is being collected. Laws protecting data privacy are being enforced to prevent data disclosure. Statistical institutes and agencies need methods to preserve confidentiality while maintaining accuracy when disclosing geographic data. In this paper we present the AQuadtree package, a software intended to produce and deal with official spatial data making data privacy and accuracy compatible. The lack of specific methods in R to anonymize spatial data motivated the development of this package, providing an automatic aggregation tool to anonymize …
Kspm: A Package For Kernel Semi-Parametric Models,
2020
Montreal university
Kspm: A Package For Kernel Semi-Parametric Models, Catherine Schramm, Sébastien Jacquemont, Karim Oualkacha, Aurélie Labbe, Celia M. T. Greenwood
The R Journal
Kernel semi-parametric models and their equivalence with linear mixed models provide analysts with the flexibility of machine learning methods and a foundation for inference and tests of hypothesis. These models are not impacted by the number of predictor variables, since the kernel trick transforms them to a kernel matrix whose size only depends on the number of subjects. Hence, methods based on this model are appealing and numerous, however only a few R programs are available and none includes a complete set of features. Here, we present the KSPM package to fit the kernel semi-parametric model and its extensions in …
Ordinalclust: An R Package To Analyze Ordinal Data,
2020
Université de Lyon
Ordinalclust: An R Package To Analyze Ordinal Data, Margot Selosse, Julien Jacques, Christophe Biernacki
The R Journal
Ordinal data are used in many domains, especially when measurements are collected from people through observations, tests, or questionnaires. ordinalClust is an innovative R package dedicated to ordinal data that provides tools for modeling, clustering, co-clustering and classifying such data. Ordinal data are modeled using the BOS distribution, which is a model with two meaningful parameters referred to as "position" and "precision". The former indicates the mode of the distribution and the latter describes how scattered the data are around the mode: the user is able to easily interpret the distribution of their data when given these two parameters. The …
A Fast And Scalable Implementation Method For Competing Risks Data With The R Package Fastcmprsk,
2020
University of Southern California
A Fast And Scalable Implementation Method For Competing Risks Data With The R Package Fastcmprsk, Eric S. Kawaguchi, Jenny I. Shen, Gang Li, Marc A. Suchard
The R Journal
Advancements in medical informatics tools and high-throughput biological experimentation make large-scale biomedical data routinely accessible to researchers. Competing risks data are typical in biomedical studies where individuals are at risk to more than one cause (type of event) which can preclude the others from happening. The Fine and Gray (1999) proportional subdistribution hazards model is a popular and well-appreciated model for competing risks data and is currently implemented in a number of statistical software packages. However, current implementations are not computationally scalable for large-scale competing risks data. We have developed an R package, fastcmprsk, that uses a novel forward-backward scan …
Six Years Of Shiny In Research: Collaborative Development Of Web Tools In R,
2020
University of Adelaide
Six Years Of Shiny In Research: Collaborative Development Of Web Tools In R, Peter Kasprzak, Lachlan Mitchell, Olena Kravchuk, Andy Timmins
The R Journal
The use of Shiny in research publications is investigated over the six and a half years since the appearance of this popular web application framework for R, which has been utilised in many varied research areas. While it is demonstrated that the complexity of Shiny applications is limited by the background architecture, and real security concerns exist for novice app developers, the collaborative benefits are worth attention from the wider research community. Shiny simplifies the use of complex methodologies for people of different specialities, at the level of proficiency appropriate for the end user. This enables a diverse community of …
Assembling Pharmacometric Datasets In R: The Puzzle Package,
2020
Modeling Great Solution
Assembling Pharmacometric Datasets In R: The Puzzle Package, Mario González-Sales, Olivier Barrière, Pierre Olivier Tremblay, Guillaume Bonnefois, Julie Desrochers, Fahima Nekka
The R Journal
Pharmacometric analyses are integral components of the drug development process. The core of each pharmacometric analysis is a dataset. The time required to construct a pharmacometrics dataset can sometimes be higher than the effort required for the modeling per se. To simplify the process, the puzzle R package has been developed aimed at simplifying and facilitating the time consuming and error prone task of assembling pharmacometrics datasets.
Puzzle consist of a series of functions written in R. These functions create, from tabulated files, datasets that are compatible with the formatting requirements of the gold standard non-linear mixed effects modeling …
Fitzroy: An R Package To Encourage Reproducible Sports Analysis,
2020
University of New South Wales
Fitzroy: An R Package To Encourage Reproducible Sports Analysis, Robert Nguyen, James Day, David Warton, Oscar Lane
The R Journal
The importance of reproducibility, and the related issue of open access to data, has received a lot of recent attention. Momentum on these issues is gathering in the sports analytics community. While Australian Rules football (AFL) is the leading commercial sport in Australia, unlike popular international sports, there has been no mechanism for the public to access comprehensive statistics on players and teams. Expert commentary currently relies heavily on data that isn’t made readily accessible and this produces an unnecessary barrier for the development of an inclusive sports analytics community. We present the R package fitzRoy to provide easy access …
Comparing Multiple Survival Functions With Crossing Hazards In R,
2020
Academia Sinica
Comparing Multiple Survival Functions With Crossing Hazards In R, Hsin-Wen Chang, Pei-Yuan Tsai, Jen-Tse Kao, Guo-You Lan
The R Journal
It is frequently of interest in time-to-event analysis to compare multiple survival functions nonparametrically. However, when the hazard functions cross, tests in existing R packages do not perform well. To address the issue, we introduce the package survELtest, which provides tests for comparing multiple survival functions with possibly crossing hazards. Due to its powerful likelihood ratio formulation, this is the only R package to date that works when the hazard functions cross. We illustrate the use of the procedures in survELtest by applying them to data from randomized clinical trials and simulated datasets. We show that these methods lead …
The Biglasso Package: A Memory- And Computation-Efficient Solver For Lasso Model Fitting With Big Data In R,
2020
University of Iowa
The Biglasso Package: A Memory- And Computation-Efficient Solver For Lasso Model Fitting With Big Data In R, Yaohui Zeng, Patrick Breheny
The R Journal
Penalized regression models such as the lasso have been extensively applied to analyzing high-dimensional data sets. However, due to memory limitations, existing R packages like glmnet and ncvreg are not capable of fitting lasso-type models for ultrahigh-dimensional, multi-gigabyte data sets that are increasingly seen in many areas such as genetics, genomics, biomedical imaging, and high-frequency finance. In this research, we implement an R package called biglasso that tackles this challenge. biglasso utilizes memory-mapped files to store the massive data on the disk, only reading data into memory when necessary during model fitting, and is thus able to handle out-of-core computation …
