Open Access. Powered by Scholars. Published by Universities.®

2018

Discipline
Institution
Keyword
Publication
Publication Type

Articles 91 - 120 of 255

Full-Text Articles in Numerical Analysis and Scientific Computing

Offline Versus Online: A Meaningful Categorization Of Ties For Retweets, Felicia Natali, Feida Zhu Aug 2018

Offline Versus Online: A Meaningful Categorization Of Ties For Retweets, Felicia Natali, Feida Zhu

Research Collection School Of Computing and Information Systems

With the recent proliferation of news being shared through online social networks, it is crucial to determine how news is spread and what drives people to share certain stories. In this paper, we focus on the social networking site Twitter and analyse user’s retweets. We study retweeting patterns between offline and online friends, particularly, how tweet novelty and tweet topic differ between tweets retweeted by offline friends and those retweeted by online friends.


Deep Learning For Practical Image Recognition: Case Study On Kaggle Competitions, Xulei Yang, Zeng Zeng, Sin G. Teo, Li Wang, Vijay Chandrasekar, Steven C. H. Hoi Aug 2018

Deep Learning For Practical Image Recognition: Case Study On Kaggle Competitions, Xulei Yang, Zeng Zeng, Sin G. Teo, Li Wang, Vijay Chandrasekar, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

In past years, deep convolutional neural networks (DCNN) have achieved big successes in image classification and object detection, as demonstrated on ImageNet in academic field. However, There are some unique practical challenges remain for real-world image recognition applications, e.g., small size of the objects, imbalanced data distributions, limited labeled data samples, etc. In this work, we are making efforts to deal with these challenges through a computational framework by incorporating latest developments in deep learning. In terms of two-stage detection scheme, pseudo labeling, data augmentation, cross-validation and ensemble learning, the proposed framework aims to achieve better performances for practical image …


Mining Association Rules For Low-Frequency Itemsets, Jimmy Ming-Tai Wu, Justin Zhan, Sanket Chobe Jul 2018

Mining Association Rules For Low-Frequency Itemsets, Jimmy Ming-Tai Wu, Justin Zhan, Sanket Chobe

Computer Science Faculty Research

High utility itemset mining has become an important and critical operation in the Data Mining field. High utility itemset mining generates more profitable itemsets and the association among these itemsets, to make business decisions and strategies. Although, high utility is important, it is not the sole measure to decide efficient business strategies such as discount offers. It is very important to consider the pattern of itemsets based on the frequency as well as utility to predict more profitable itemsets. For example, in a supermarket or restaurant, beverages like champagne or wine might generate high utility (profit), but also sell less …


Predicting River Stage Using Recurrent Neural Networks, Eric Rohli Jul 2018

Predicting River Stage Using Recurrent Neural Networks, Eric Rohli

LSU Master's Theses

River stage prediction is an important problem in the water transportation industry. Accurate river stage predictions provide crucial information to barge and tow boat operators, port terminal captains, and lock management officials. Shallow river levels caused by prolonged drought impact the loading capacity of barges and tow boats. High river levels caused by excessive rainfall or snowmelt allow for greater tow capacities but make downstream transportation and lock management risky. Current academic river height prediction systems utilize either time series statistical analysis or machine learning algorithms to forecast future river heights, but systems that combine these two areas often limit …


Lp Algorithms For Portfolio Optimization: The Portfoliooptim Package, Andrzej Palczewski Jul 2018

Lp Algorithms For Portfolio Optimization: The Portfoliooptim Package, Andrzej Palczewski

The R Journal

The paper describes two algorithms for financial portfolio optimization with the following risk measures: CVaR, MAD, LSAD and dispersion CVaR. These algorithms can be applied to discrete distributions of asset returns since then the optimization problems can be reduced to linear programs. The first algorithm solves a simple recourse problem as described by Haneveld using Benders de composition method. The second algorithm finds an optimal portfolio with the smallest distance to a given benchmark portfolio and is an adaptation of the least norm solution (called also normal solution) of linear programs due to Zhao and Li. The algorithms are implemented …


Changes In R, R Core Team Jul 2018

Changes In R, R Core Team

The R Journal

CHANGES IN R 3.5.0 patched

CHANGES IN R 3.5.0


Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis Jul 2018

Changes On Cran, Kurt Hornik, Uwe Ligges, Achim Zeileis

The R Journal

In the past 7 months, 1178 new packages were added to the CRAN package repository. 18 packages were unarchived, 493 archived and none removed. The following shows the growth of the number of active packages in the CRAN package repository:


R Day Report, Fernando P. Mayer, Walmes M. Zeviani, Wagner H. Bonat, Elias T. Krainski, Paulo J. Ribeiro Jr. Jul 2018

R Day Report, Fernando P. Mayer, Walmes M. Zeviani, Wagner H. Bonat, Elias T. Krainski, Paulo J. Ribeiro Jr.

The R Journal

R Day1- National Meeting of R Users, took place on May, 22, 2018 at Federal University of Paraná (UFPR), Curitiba, Brazil. It was the first event in Brazil endorsed by The R Foundation.


Setmethods: An Add-On R Package For Advanced Qca, Ioana-Elena Oana, Carsten Q. Scheider Jul 2018

Setmethods: An Add-On R Package For Advanced Qca, Ioana-Elena Oana, Carsten Q. Scheider

The R Journal

This article presents the functionalities of the R package SetMethods, aimed at performing advanced set-theoretic analyses. This includes functions for performing set-theoretic multi-method research, set-theoretic theory evaluation, Enhanced Standard Analysis, diagnosing the impact of temporal, spatial, or substantive clusterings of the data on the results obtained via Qualitative Comparative Analysis (QCA), indirect calibration, and visualising QCA results via XY plots or radar charts. Each functionality is presented in turn, the conceptual idea and the logic behind the procedure being first summarized, and afterwards illustrated with data from Schneider et al. (2010).


Simple Features For R: Standardized Support For Spatial Vector Data, Edzer Pebesma Jul 2018

Simple Features For R: Standardized Support For Spatial Vector Data, Edzer Pebesma

The R Journal

Simple features are a standardized way of encoding spatial vector data (points, lines, polygons) in computers. The sf package implements simple features in R, and has roughly the same capacity for spatial vector data as packages sp, rgeos, and rgdal. We describe the need for this package, its place in the R package ecosystem, and its potential to connect R to other computer systems. We illustrate this with examples of its use.


Epistemic Game Theory: Putting Algorithms To Work, Bilge BaşEr, Nalan Cinemre Jul 2018

Epistemic Game Theory: Putting Algorithms To Work, Bilge BaşEr, Nalan Cinemre

The R Journal

The aim of this study is to construct an epistemic model in which each rational choice under common belief in rationality is supplemented by a type which expresses such a belief. In practice, the finding of type depends on manual solution approach with some mathematical operations in scope of the theory. This approach becomes less convenient with the growth of the size of the game. To solve this difficulty, a linear programming model is constructed for two-player, static and non-cooperative games to find the type that is supporting that player’s rational choice is optimal under common belief in rationality and …


Grpstring: An R Package For Analysis Of Groups Of Strings, Hui Tang, Elizabeth L. Day, Molly B. Atkinson, Norbert J. Pienta Jul 2018

Grpstring: An R Package For Analysis Of Groups Of Strings, Hui Tang, Elizabeth L. Day, Molly B. Atkinson, Norbert J. Pienta

The R Journal

The R package GrpString was developed as a comprehensive toolkit for quantitatively analyzing and comparing groups of strings. It offers functions for researchers and data analysts to prepare strings from event sequences, extract common patterns from strings, and compare patterns be tween string vectors. The package also finds transition matrices and complexity of strings, determines clusters in a string vector, and examines the statistical difference between two groups of strings.


Lba: An R Package For Latent Budget Analysis, Enio G. Jelihovschi, Ivan Bezerra Allaman Jul 2018

Lba: An R Package For Latent Budget Analysis, Enio G. Jelihovschi, Ivan Bezerra Allaman

The R Journal

The latent budget model is a mixture model for compositional data sets in which the entries, a contingency table, may be either realizations from a product multinomial distribution or distribution free. Based on this model, the latent budget analysis considers the interactions of two variables; the explanatory (row) and the response (column) variables. The package lba uses expectation-maximization and active constraints method (ACM) to carry out, respectively, the maximum likelihood and the least squares estimation of the model parameters. It contains three main functions, lba which performs the analysis, goodnessfit for model selection and goodness of fit and the plotting …


Icsoutlier: Unsupervised Outlier Detection For Low-Dimensional Contamination Authors: Structure, Aurore Archimbaud, Klaus Nordhausen, Anne Ruiz-Gazen Jul 2018

Icsoutlier: Unsupervised Outlier Detection For Low-Dimensional Contamination Authors: Structure, Aurore Archimbaud, Klaus Nordhausen, Anne Ruiz-Gazen

The R Journal

Detecting outliers in a multivariate and unsupervised context is an important and ongoing problem notably for quality control. Many statistical methods are already implemented in R and are briefly surveyed in the present paper. But only a few lead to the accurate identification of potential outliers in the case of a small level of contamination. In this particular context, the Invariant Coordinate Selection (ICS) method shows remarkable properties for identifying outliers that lie on a low-dimensional subspace in its first invariant components. It is implemented in the ICSOutlier package. The main function of the package, ics.outlier, offers the possibility of …


Onewaytests: An R Package For One-Way Tests In Independent Groups Designs, Osman Dag, Anil Dolgun, Naime Meric Konar Jul 2018

Onewaytests: An R Package For One-Way Tests In Independent Groups Designs, Osman Dag, Anil Dolgun, Naime Meric Konar

The R Journal

One-way tests in independent groups designs are the most commonly utilized statistical methods with applications on the experiments in medical sciences, pharmaceutical research, agriculture, biology, engineering, social sciences and so on. In this paper, we present the one-way tests package to investigate treatment effects on the dependent variable. The package offers the one-way tests in independent groups designs, which include ANOVA, Welch’s heteroscedastic F test, Welch’s heteroscedastic F test with trimmed means and Winsorized variances, Brown-Forsythe test, Alexander Govern test, James second order test and Kruskal-Wallis test. The package also provides pairwise comparisons, graphical approaches, and assesses variance homogeneity and …


Bayesian Testing, Variable Selection And Model Averaging In Linear Models Using R With Bayesvarsel, Gonzalo Garcia-Donato, Anabel Forte Jul 2018

Bayesian Testing, Variable Selection And Model Averaging In Linear Models Using R With Bayesvarsel, Gonzalo Garcia-Donato, Anabel Forte

The R Journal

In this paper, objective Bayesian methods for hypothesis testing and variable selection in linear models are considered. The focus is on BayesVarSel, an R package that computes posterior probabilities of hypotheses/models and provides a suite of tools to properly summarize the results. We introduce the usage of specific functions to compute several types of model averaging estimations and predictions weighted by posterior probabilities. BayesVarSel contains exact algorithms to perform fast computations in problems of small to moderate size and heuristic sampling methods to solve large problems. We illustrate the functionalities of the package with several data examples.


Tackling Uncertainties Of Species Distribution Model Projections With Package Mopa, M. Iturbide, J. Bedia, J.M. Gutiérrez Jul 2018

Tackling Uncertainties Of Species Distribution Model Projections With Package Mopa, M. Iturbide, J. Bedia, J.M. Gutiérrez

The R Journal

Species Distribution Models (SDMs) constitute an important tool to assist decision-making in environmental conservation and planning in the context of climate change. Nevertheless, SDM projections are affected by a wide range of uncertainty factors (related to training data, climate projections and SDM techniques), which limit their potential value and credibility. The new package mopa provides tools for designing comprehensive multi-factor SDM ensemble experiments, combining multiple sources of uncertainty (e.g. baseline climate, pseudo-absence realizations, SDM techniques, future projections) and allowing to assess their contribution to the overall spread of the ensemble projection. In addition, mopa is seamlessly integrated with the climate4R …


Panjen: An R Package For Ranking Transformations In A Linear Regression, Cathrine Ulla Jensen, Toke Emil Panduro Jul 2018

Panjen: An R Package For Ranking Transformations In A Linear Regression, Cathrine Ulla Jensen, Toke Emil Panduro

The R Journal

PanJen is an R-package for ranking transformations in linear regressions. It provides users with the ability to explore the relationship between a dependent variable and its independent variables. The package offers an easy and data-driven way to choose a functional form in multiple linear regression models by comparing a range of parametric transformations. The parametric functional forms are benchmarked against each other and a non-parametric transformation. The package allows users to generate plots that show the relation between a covariate and the dependent variable. Furthermore, PanJen will enable users to specify specific functional transformations, driven by a priori and theory-based …


Arco: An R Package To Estimate Artificial Counterfactuals, Yuri R. Fonseca, Ricardo P. Masini, Marcelo C. Medeiros, Gabriel F.R. Vasconcelos Jul 2018

Arco: An R Package To Estimate Artificial Counterfactuals, Yuri R. Fonseca, Ricardo P. Masini, Marcelo C. Medeiros, Gabriel F.R. Vasconcelos

The R Journal

In this paper we introduce the ArCo package for R which consists of a set of functions to implement the the Artificial Counterfactual (ArCo) methodology to estimate causal effects of an intervention (treatment) on aggregated data and when a control group is not necessarily available. The ArCo method is a two-step procedure, where in the first stage a counterfactual is estimated from a large panel of time series from a pool of untreated peers. In the second-stage, the average treatment effect over the post-intervention sample is computed. Standard inferential procedures are available. The package is illustrated with both simulated and …


Infotrad: An R Package For Estimating The Probability Of Informed Trading, Duygu Çelik, Murat Tiniç Jul 2018

Infotrad: An R Package For Estimating The Probability Of Informed Trading, Duygu Çelik, Murat Tiniç

The R Journal

The purpose of this paper is to introduce the R package InfoTrad for estimating the probability of informed trading (PIN) initially proposed by Easley et al. (1996). PIN is a popular information asymmetry measure that proxies the proportion of informed traders in the market. This study provides a short survey on alternative estimation techniques for the PIN. There are many problems documented in the existing literature in estimating PIN. InfoTrad package aims to address two problems. First, the sequential trading structure proposed by Easley et al. (1996) and later extended by Easley et al. (2002) is prone to sample selection …


Conference Report: Erum 2018, Gergely Daróczi Jul 2018

Conference Report: Erum 2018, Gergely Daróczi

The R Journal

The European R Users Meeting (eRum) is an international conference that aims at bringing together users of the R language living in Europe– in the years when the useR! conference is hosted outside of the continent.

The first eRum conference was held in 2016 in Poznan, Poland with around 250 attendees and 20 sessions spanning over 3 days, including more than 80 speakers. Around that time, we also held a smaller conference in Budapest: the first satRday event happened with 25 speakers and almost 200 attendees from 19 countries in 2016.

The eRum 2018 conference is heritage of these two …


Realvams: An R Package For Fitting A Multivariate Value-Added Model (Vam), Jennifer Broatch, Jennifer Green, Andrew Karl Jul 2018

Realvams: An R Package For Fitting A Multivariate Value-Added Model (Vam), Jennifer Broatch, Jennifer Green, Andrew Karl

The R Journal

We present RealVAMS, an R package for fitting a generalized linear mixed model to multimembership data with partially crossed and partially nested random effects. RealVAMS utilizes a multivariate generalized linear mixed model with pseudo-likelihood approximation for fitting normally distributed continuous response(s) jointly with a binary outcome. In an educational context, the model is referred to as a multidimensional value-added model, which extends previous theory to estimate the relationships between potential teacher contributions toward different student outcomes and to allow the consideration of a binary, real-world outcome such as graduation. The simultaneous joint modeling of continuous and binary outcomes was not …


Approximating The Sum Of Independent Non-Identical Binomial Random Variables, Boxiang Liu, Thomas Quertermous Jul 2018

Approximating The Sum Of Independent Non-Identical Binomial Random Variables, Boxiang Liu, Thomas Quertermous

The R Journal

The distribution of the sum of independent non-identical binomial random variables is frequently encountered in areas such as genomics, healthcare, and operations research. Analytical solutions for the density and distribution are usually cumbersome to find and difficult to compute. Several methods have been developed to approximate the distribution, among which is the saddlepoint approximation. However, implementation of the saddlepoint approximation is non-trivial. In this paper, we implement the saddlepoint approximation in the sinib package and provide two examples to illustrate its usage. One example uses simulated data while the other uses real-world healthcare data. The sinib package addresses the gap …


Editorial, John Verzani Jul 2018

Editorial, John Verzani

The R Journal

On behalf of the Editorial Board, I am pleased to present Volume 10, Issue 1 of the R Journal. This issue contains 36 contributed articles. The majority of which cover new or newly enhanced packages on CRAN.


Nonparametric Independence Tests And K-Sample Tests For Large Sample Sizes Using Package Hhg, Barak Brill, Yair Heller, Ruth Heller Jul 2018

Nonparametric Independence Tests And K-Sample Tests For Large Sample Sizes Using Package Hhg, Barak Brill, Yair Heller, Ruth Heller

The R Journal

Nonparametric tests of independence and k-sample tests are ubiquitous in modern applications, but they are typically computationally expensive. We present a family of nonparametric tests that are computationally efficient and powerful for detecting any type of dependence between a pair of univariate random variables. The computational complexity of the suggested tests is sub-quadratic in sample size, allowing calculation of test statistics for millions of observations. We survey both algorithms and the HHG package in which they are implemented, with usage examples showing the implementation of the proposed tests for both the independence case and the k-sample problem. The tests are …


Dimred And Coranking - Unifying Dimensionality Reduction In R, Guido Kraemer, Markus Reichstein, Miguel D. Mahecha Jul 2018

Dimred And Coranking - Unifying Dimensionality Reduction In R, Guido Kraemer, Markus Reichstein, Miguel D. Mahecha

The R Journal

“Dimensionality reduction” (DR) is a widely used approach to find low dimensional and interpretable representations of data that are natively embedded in high-dimensional spaces. DR ca nbe realized by a plethora of methods with different properties, objectives, and, hence, (dis)advantages. The resulting low-dimensional data embeddings are often difficult to compare with objective criteria. Here, we introduce the dimRed and coRanking packages for the R language. These open source software packages enable users to easily access multiple classical and advanced DR methods using a common interface. The packages also provide quality indicators for the embeddings and easy visualization of high dimensional …


Collections In R: Review And Proposal, Timothy Barry Jul 2018

Collections In R: Review And Proposal, Timothy Barry

The R Journal

R is a powerful tool for data processing, visualization, and modeling. However, R is slower than other languages used for similar purposes, such as Python. One reason for this is that R lacks base support for collections, abstract data types that store, manipulate, and return data (e.g., sets, maps, stacks). An exciting recent trend in the R extension ecosystem is the development of collection packages, packages that provide classes that implement common collections. At least 12 collection packages are available across the two major R extension repositories, the Comprehensive R Archive Network (CRAN) and Bioconductor. In this article, we compare …


Pstat: An R Package To Assess Population Differentiation In Phenotypic Traits, Stéphane Blondeau Da Silva, Anne Da Silva Jul 2018

Pstat: An R Package To Assess Population Differentiation In Phenotypic Traits, Stéphane Blondeau Da Silva, Anne Da Silva

The R Journal

The package Pstat calculates PST values to assess differentiation among populations from a set of quantitative traits and provides bootstrapped distributions and confidence intervals for PST. Variations of PST as a function of the parameter c/h2 are studied as well. The package implements different transformations of the measured phenotypic traits to eliminate variation resulting from allometric growth, including calculation of residuals from linear regression, Reist standardization, and the Aitchison transformation.


Residuals And Diagnostics For Binary And Ordinal Regression Models: An Introduction To The Sure Package, Brandon M. Greenwell, Andrew J. Mccarthy, Bradley C. Boehmke, Dungang Liu Jul 2018

Residuals And Diagnostics For Binary And Ordinal Regression Models: An Introduction To The Sure Package, Brandon M. Greenwell, Andrew J. Mccarthy, Bradley C. Boehmke, Dungang Liu

The R Journal

Residual diagnostics is an important topic in the classroom, but it is less often used in practice when the response is binary or ordinal. Part of the reason for this is that generalized models for discrete data, like cumulative link models and logistic regression, do not produce standard residuals that are easily interpreted as those in ordinary linear regression. In this paper, we introduce the R package sure, which implements a recently developed idea of SUrrogate REsiduals. We demonstrate the utility of the package in detection of cumulative link model misspecification with respect to mean structures, link functions, …


Hrm: An R Package For Analysing High-Dimensional Multi-Factor Repeated Measures Authors: Martin Happ, Solomon W. Harrar And Arne C. Bathke, Martin Happ, Solomon W. Harrar, Arne C. Bathke Jul 2018

Hrm: An R Package For Analysing High-Dimensional Multi-Factor Repeated Measures Authors: Martin Happ, Solomon W. Harrar And Arne C. Bathke, Martin Happ, Solomon W. Harrar, Arne C. Bathke

The R Journal

High-dimensional longitudinal data pose a serious challenge for statistical inference as many test statistics cannot be computed for high-dimensional data, or they do not maintain the nominal type-I error rate, or have very low power. Therefore, it is necessary to derive new inference methods capable of dealing with high dimensionality, and to make them available to statistics practitioners. One such method is implemented in the package HRM described in this article. This new method uses a similar approach as the Welch-Satterthwaite t-test approximation and works very well for high-dimensional data as long as the data distribution is not too skewed …