Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Programming Languages and Compilers

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 421 - 450 of 1844

Full-Text Articles in Computer Sciences

The R Journal (December 2021) 13(2): Complete Issue, The R Foundation Nov 2021

The R Journal (December 2021) 13(2): Complete Issue, The R Foundation

The R Journal

On behalf of the R Foundation and the Editorial board, I am pleased to present Volume 13 Issue 2 of the R Journal. This is the biggest issue ever!

First, some news from the Editorial board. A big thank you to Mike Kane, who has finished his term. As Editor-in-Chief in 2020, Mike expanded operations to include Associate Editors in the reviewing process. The R Journal now has a team of 20 Associate Editors. This has helped to manage the increasing number of submissions. We welcome new Associate Editors, Przemek Biecek, Chris Brunsdon, Mine Çetinkaya-Rundel, Kieran Healy, Adam Loy, Priyanga …


From Community Search To Community Understanding: A Multimodal Community Query Engine, Zhao Li, Pengcheng Zou, Xia Chen, Shichang Hu, Peng Zhang, Yumou Zhou, Bingsheng He, Yuchen Li, Xing Tang Nov 2021

From Community Search To Community Understanding: A Multimodal Community Query Engine, Zhao Li, Pengcheng Zou, Xia Chen, Shichang Hu, Peng Zhang, Yumou Zhou, Bingsheng He, Yuchen Li, Xing Tang

Research Collection School Of Computing and Information Systems

In this demo, we present an online multi-modal community query engine (MQE1 ) on Alibaba’s billion-scale heterogeneous network. MQE has two distinct features in comparison with existing community query engines. Firstly, MQE supports multimodal community search on heterogeneous graphs with keyword and image queries. Secondly, to facilitate community understanding in real business scenarios, MQE generates natural language descriptions for the retrieved community in combination with other useful demographic information. The distinct features of MQE benefit many downstream applications in Alibaba’s e-commerce platform like recommendation. Our experiments confirm the effectiveness and efficiency of MQE on graphs with billions of edges.


Wav-Bert: Cooperative Acoustic And Linguistic Representation Learning For Low-Resource Speech Recognition, Guolin Zheng, Yubei Xiao, Ke Gong, Pan Zhou, Xiaodan Liang, Liang Lin Nov 2021

Wav-Bert: Cooperative Acoustic And Linguistic Representation Learning For Low-Resource Speech Recognition, Guolin Zheng, Yubei Xiao, Ke Gong, Pan Zhou, Xiaodan Liang, Liang Lin

Research Collection School Of Computing and Information Systems

Unifying acoustic and linguistic representation learning has become increasingly crucial to transfer the knowledge learned on the abundance of high-resource language data for low-resource speech recognition. Existing approaches simply cascade pre-trained acoustic and language models to learn the transfer from speech to text. However, how to solve the representation discrepancy of speech and text is unexplored, which hinders the utilization of acoustic and linguistic information. Moreover, previous works simply replace the embedding layer of the pre-trained language model with the acoustic features, which may cause the catastrophic forgetting problem. In this work, we introduce Wav-BERT, a cooperative acoustic and linguistic …


Preparing School Leaders To Advance Equity In Computer Science Education, Julie Flapan, Jean J. Ryoo, Roxana Hadad, Joel Knudson Oct 2021

Preparing School Leaders To Advance Equity In Computer Science Education, Julie Flapan, Jean J. Ryoo, Roxana Hadad, Joel Knudson

Journal of Computer Science Integration

Background and Context: Most large-scale statewide initiatives of the Computer Science for All (CS for All) movement have focused on the classroom level. Critical questions remain about building school and district leadership capacity to support teachers while implementing equitable computer science education that is scalable and sustainable.

Objective: This statewide research-practice partnership, involving university researchers and school leaders from 14 local education agencies (LEA) from district and county offices, addresses the following research question: What do administrators identify as most helpful for understanding issues related to equitable computer science implementation when engaging with a guide and workshop we …


Building A More Sustainable And Accessible Internet: Lightweight Web Design With Html And Css, Chelsea Thompto Oct 2021

Building A More Sustainable And Accessible Internet: Lightweight Web Design With Html And Css, Chelsea Thompto

Assignment Prompts

While the internet has great potential to bring people together, if the internet was a country, it would be the 7th largest energy consumer on the planet. This is set to increase in years to come moving the internet even higher on this list to become the 4th largest energy consumer if it were to be a country. So, as artists and digital citizens it is imperative that we understand how to create and display the content we produce online in ways that are sustainable and accessible.
This assignment, while slated for Art 109, may be slotted into an earlier …


Design And Supervision Model Of Group Projects For Active Learning, Yi Meng Lau, Kyong Jin Shim, Swapna Gottipati Oct 2021

Design And Supervision Model Of Group Projects For Active Learning, Yi Meng Lau, Kyong Jin Shim, Swapna Gottipati

Research Collection School Of Computing and Information Systems

This research paper presents a group project framework for a second-year programming course, which was conducted during the COVID-19 pandemic. The framework offers well defined stages of the group project which allow students to work on their choice of a real-world problem, integrate their learnings from previous courses, and present a working solution. In the group project, students actively participate, reflect, and contribute to achieving the goals set in the learning objectives of the course. Our framework incorporates key features from Kolb’s Experiential Learning Theory (1984) and principles of active learning from Barnes (1989) to achieve active and experiential learning …


Teachers’ Engagement And Self-Efficacy In A Pk–12 Computer Science Teacher Virtual Community Of Practice, Robert Schwarzhaupt, Feng Liu, Joseph Wilson, Fanny Lee, Melissa Rasberry Sep 2021

Teachers’ Engagement And Self-Efficacy In A Pk–12 Computer Science Teacher Virtual Community Of Practice, Robert Schwarzhaupt, Feng Liu, Joseph Wilson, Fanny Lee, Melissa Rasberry

Journal of Computer Science Integration

Prekindergarten to 12th-grade teachers of computer science (CS) face many challenges, including isolation, limited CS professional development resources, and low levels of CS teaching self-efficacy that could be mitigated through communities of practice (CoPs). This study used survey data from 420 PK–12 CS teacher members of a virtual CoP, CS for All Teachers, to examine the needs of these teachers and how CS teaching self-efficacy, community engagement, and sharing behaviors vary by teachers’ instructional experiences and school levels taught. Results show that CS teachers primarily join the CoP to gain high-quality pedagogical, assessment, and instructional resources. The study also found …


The Development Of Teaching Case Studies To Explore Ethical Issues Associated With Computer Programming, Michael Collins, Damian Gordon, Dympna O'Sullivan Sep 2021

The Development Of Teaching Case Studies To Explore Ethical Issues Associated With Computer Programming, Michael Collins, Damian Gordon, Dympna O'Sullivan

Conference papers

In the past decade software products have become pervasive in many aspects of people’s lives around the world. Unfortunately, the quality of the experience an individual has interacting with that software is dependent on the quality of the software itself, and it is becoming more and more evident that many large software products contain a range of issues and errors, and these issues are not known to the developers of these systems, and they are unaware of the deleterious impacts of those issues on the individuals who use these systems. The authors of this paper are developing a new digital …


Does Bert Understand Idioms? A Probing-Based Empirical Study Of Bert Encodings Of Idioms, Minghuan Tan, Jing Jiang Sep 2021

Does Bert Understand Idioms? A Probing-Based Empirical Study Of Bert Encodings Of Idioms, Minghuan Tan, Jing Jiang

Research Collection School Of Computing and Information Systems

Understanding idioms is important in NLP. In this paper, we study to what extent pre-trained BERT model can encode the meaning of a potentially idiomatic expression (PIE) in a certain context. We make use of a few existing datasets and perform two probing tasks: PIE usage classification and idiom paraphrase identification. Our experiment results suggest that BERT indeed can separate the literal and idiomatic usages of a PIE with high accuracy. It is also able to encode the idiomatic meaning of a PIE to some extent.


Injecting Descriptive Meta-Information Into Pre-Trained Language Models With Hypernetworks, Wenying Duan, Xiaoxi He, Zimu Zhou, Hong Rao, Lothar Thiele Sep 2021

Injecting Descriptive Meta-Information Into Pre-Trained Language Models With Hypernetworks, Wenying Duan, Xiaoxi He, Zimu Zhou, Hong Rao, Lothar Thiele

Research Collection School Of Computing and Information Systems

Pre-trained language models have been widely adopted as backbones in various natural language processing tasks. However, existing pre-trained language models ignore the descriptive meta-information in the text such as the distinction between the title and the mainbody, leading to over-weighted attention to insignificant text. In this paper, we propose a hypernetwork-based architecture to model the descriptive meta-information and integrate it into pre-trained language models. Evaluations on three natural language processing tasks show that our method notably improves the performance of pre-trained language models and achieves the state-of-the-art results on keyphrase extraction.


Learning And Evaluating Chinese Idiom Embeddings, Minghuan Tan, Jing Jiang Sep 2021

Learning And Evaluating Chinese Idiom Embeddings, Minghuan Tan, Jing Jiang

Research Collection School Of Computing and Information Systems

We study the task of learning and evaluating Chinese idiom embeddings. We first construct a new evaluation dataset that contains idiom synonyms and antonyms. Observing that existing Chinese word embedding methods may not be suitable for learning idiom embeddings, we further present a BERT-based method that directly learns embedding vectors for individual idioms. We empirically compare representative existing methods and our method. We find that our method substantially outperforms existing methods on the evaluation dataset we have constructed.


Teaching Students How To Code Qualitative Data: An Experiential Activity Sequence For Training Novice Educational Researchers, Jennifer E. Lineback Aug 2021

Teaching Students How To Code Qualitative Data: An Experiential Activity Sequence For Training Novice Educational Researchers, Jennifer E. Lineback

University of South Florida (USF) M3 Publishing

Coursework on qualitative research methods is common in many collegiate departments, including psychology, nursing, sociology, and education. Instructors for these courses must identify meaningful activities to support their students’ learning of the domain. This paper presents the components of an experiential activity sequence centered on coding and coding scheme development. Each of the three component activities of this sequence is elaborated, as are the students’ experiences during their participation in the activities. Additionally, the issues concerning coding and coding scheme development that typically emerge from students’ participation in these activities are discussed. Results from implementations of both in-person (face-to-face) and …


Towards A Large-Scale Intelligent Mobile-Argumentation And Discovering Arguments, Controversial Topics And Topic-Oriented Focal Sets In Cyber-Argumentation, Najla Althuniyan Jul 2021

Towards A Large-Scale Intelligent Mobile-Argumentation And Discovering Arguments, Controversial Topics And Topic-Oriented Focal Sets In Cyber-Argumentation, Najla Althuniyan

Graduate Theses and Dissertations

User-generated content (UGC) platforms host different forms of information, such as audio, video, pictures, and text. They have many online applications, such as social media, blogs, photo and video sharing, customer reviews, debate, and deliberation platforms. Usually, the content of these platforms is provided and consumed by users. Most of these platforms, mainly social media and blogs, are often used for online discussion. These platforms offer tools for users to share and express opinions. Commonly, people from different backgrounds and origins discuss opinions about various issues over the Internet. Furthermore, discussions among users contain substantial information from which knowledge about …


An Automated Method To Enrich And Expand Consumer Health Vocabularies Using Glove Word Embeddings, Mohammed Ibrahim Jul 2021

An Automated Method To Enrich And Expand Consumer Health Vocabularies Using Glove Word Embeddings, Mohammed Ibrahim

Graduate Theses and Dissertations

Clear language makes communication easier between any two parties. However, a layman may have difficulty communicating with a professional due to not understanding the specialized terms common to the domain. In healthcare, it is rare to find a layman knowledgeable in medical jargon, which can lead to poor understanding of their condition and/or treatment. To bridge this gap, several professional vocabularies and ontologies have been created to map laymen medical terms to professional medical terms and vice versa. Many of the presented vocabularies are built manually or semi-automatically requiring large investments of time and human effort and consequently the slow …


Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao Jul 2021

Privacy-Preserving Cloud-Assisted Data Analytics, Wei Bao

Graduate Theses and Dissertations

Nowadays industries are collecting a massive and exponentially growing amount of data that can be utilized to extract useful insights for improving various aspects of our life. Data analytics (e.g., via the use of machine learning) has been extensively applied to make important decisions in various real world applications. However, it is challenging for resource-limited clients to analyze their data in an efficient way when its scale is large. Additionally, the data resources are increasingly distributed among different owners. Nonetheless, users' data may contain private information that needs to be protected.

Cloud computing has become more and more popular in …


Teaching Coding In A Virtual Environment: Overcoming Challenges, Marion S. Smith Jun 2021

Teaching Coding In A Virtual Environment: Overcoming Challenges, Marion S. Smith

Southwestern Business Administration Journal

Educational research suggests that teaching techniques are subject matter specific. Teaching techniques in introductory programming classes are centered around two approaches used by students in learning. One approach is where students develop a thorough understanding of what they are learning. This is referred to as “deep learning”. Other students use a “surface approach” where they perform the tasks required from them. The persona of the instructor and the choice of instructional materials used within a class determines which approach the student will adopt. Active teaching techniques fosters “deep learning”. With the need to adapt active teaching techniques to a virtual …


The R Journal (June 2021) 13(1): Complete Issue, The R Foundation Jun 2021

The R Journal (June 2021) 13(1): Complete Issue, The R Foundation

The R Journal

Editorial, Dianne Cook

Contributed Research Articles

SEEDCCA: An Integrated R-Package for Canonical Correlation Analysis and Partial Least Squares, Bo-Young Kim, Yunju Im, and Jae Keun Yoo

npcure: An R Package for Nonparametric Inference in Mixture Cure Models, Ana López-Cheda, M. Amalia Jácome, and Ignacio López-de-Ullibarri

A Method for Deriving Information from Running R Code, Mark P. J. van der Loo

JMcmprsk: An R Package for Joint Modelling of Longitudinal and Survival Data with Competing Risks, Hong Wang, Ning Li, Shanpeng Li, and Gang Li

Wide-to-tall Data Reshaping Using Regular Expressions and the nc Package, Toby Dylan Hocking

Linear Regression with …


Analyzing Dependence Between Point Processes In Time Using Indtestpp, Ana C. Cebrián, Jesús Asín Jun 2021

Analyzing Dependence Between Point Processes In Time Using Indtestpp, Ana C. Cebrián, Jesús Asín

The R Journal

The need to analyze the dependence between two or more point processes in time appears in many modeling problems related to the occurrence of events, such as the occurrence of climate events at different spatial locations or synchrony detection in spike train analysis. The package IndTestPP provides a general framework for all the steps in this type of analysis, and one of its main features is the implementation of three families of tests to study independence given the intensities of the processes, which are not only useful to assess independence but also to identify factors causing dependence. The package also …


Distr6: R6 Object-Oriented Probability Distributions Interface In R, Raphael Sonabend, Franz J. Király Jun 2021

Distr6: R6 Object-Oriented Probability Distributions Interface In R, Raphael Sonabend, Franz J. Király

The R Journal

distr6 is an object-oriented (OO) probability distributions interface leveraging the extensibility and scalability of R6 and the speed and efficiency of Rcpp. Over 50 probability distributions are currently implemented in the package with ‘core’ methods, including density, distribution, and generating functions, and more ‘exotic’ ones, including hazards and distribution function anti-derivatives. In addition to simple distributions, distr6 supports compositions such as truncation, mixtures, and product distributions. This paper presents the core functionality of the package and demonstrates examples for key use-cases. In addition, this paper provides a critical review of the object-oriented programming paradigms in R and describes some …


Krippendorffsalpha: An R Package For Measuring Agreement Using Krippendorff's Alpha Coefficient, John Hughes Jun 2021

Krippendorffsalpha: An R Package For Measuring Agreement Using Krippendorff's Alpha Coefficient, John Hughes

The R Journal

R package krippendorffsalpha provides tools for measuring agreement using Krippendorff’s α coefficient, a well-known nonparametric measure of agreement (also called inter-rater reliability and various other names). This article first develops Krippendorff’s α in a natural way and situates α among statistical procedures. Then, the usage of package krippendorffsalpha is illustrated via analyses of two datasets, the latter of which was collected during an imaging study of hip cartilage. The package permits users to apply the α methodology using built-in distance functions for the nominal, ordinal, interval, or ratio levels of measurement. User-defined distance functions are also supported. The fitting function …


The R Package Smicd: Statistical Methods For Interval-Censored Data, Paul Walter Jun 2021

The R Package Smicd: Statistical Methods For Interval-Censored Data, Paul Walter

The R Journal

The package allows the use of two new statistical methods for the analysis of intervalcensored data: 1) direct estimation/prediction of statistical indicators and 2) linear (mixed) regression analysis. Direct estimation of statistical indicators, for instance, poverty and inequality indicators, is facilitated by a non parametric kernel density algorithm. The algorithm is able to account for weights in the estimation of statistical indicators. The standard errors of the statistical indicators are estimated with a non parametric bootstrap. Furthermore, the package offers statistical methods for the estimation of linear and linear mixed regression models with an interval-censored dependent variable, particularly random slope …


Finding Optimal Normalizing Transformations Via Bestnormalize, Ryan A. Peterson Jun 2021

Finding Optimal Normalizing Transformations Via Bestnormalize, Ryan A. Peterson

The R Journal

The bestNormalize R package was designed to help users find a transformation that can effectively normalize a vector regardless of its actual distribution. Each of the many normalization techniques that have been developed has its own strengths and weaknesses, and deciding which to use until data are fully observed is difficult or impossible. This package facilitates choosing between a range of possible transformations and will automatically return the best one, i.e., the one that makes data look the most normal. To evaluate and compare the normalization efficacy across a suite of possible transformations, we developed a statistic based on a …


Robustness In Network (Robin): An R Package For Comparison And Validation Of Communities, Valeria Policastro, Dario Righelli, Annamaria Carissimo, Luisa Cutillo, Italia De Feis Jun 2021

Robustness In Network (Robin): An R Package For Comparison And Validation Of Communities, Valeria Policastro, Dario Righelli, Annamaria Carissimo, Luisa Cutillo, Italia De Feis

The R Journal

In network analysis, many community detection algorithms have been developed. However, their implementation leaves unaddressed the question of the statistical validation of the results. Here, we present robin (ROBustness In Network), an R package to assess the robustness of the community structure of a network found by one or more methods to give indications about their reliability. The procedure initially detects if the community structure found by a set of algorithms is statistically significant and then compares two selected detection algorithms on the same graph to choose the one that better fits the network of interest. We demonstrate the use …


Indexnumber: An R Package For Measuring The Evolution Of Magnitudes, Alejandro Saavedra-Nieves, Paula Saavedra-Nieves Jun 2021

Indexnumber: An R Package For Measuring The Evolution Of Magnitudes, Alejandro Saavedra-Nieves, Paula Saavedra-Nieves

The R Journal

Index numbers are descriptive statistical measures useful in economic settings for comparing simple and complex magnitudes registered, usually in two time periods. Although this theory has a large history, it still plays an important role in modern today’s societies where big amounts of economic data are available and need to be analyzed. After a detailed revision on classical index numbers in literature, this paper is focused on the description of the R package IndexNumber with strong capabilities for calculating them. Two of the four real data sets contained in this library are used for illustrating the determination of the index …


Pdynmc: A Package For Estimating Linear Dynamic Panel Data Models Based On Nonlinear Moment Conditions, Markus Fritsch, Andrew Adrian Yu Pua, Joachim Schnurbus Jun 2021

Pdynmc: A Package For Estimating Linear Dynamic Panel Data Models Based On Nonlinear Moment Conditions, Markus Fritsch, Andrew Adrian Yu Pua, Joachim Schnurbus

The R Journal

This paper introduces pdynmc, an R package that provides users sufficient flexibility and precise control over the estimation and inference in linear dynamic panel data models. The package primarily allows for the inclusion of nonlinear moment conditions and the use of iterated GMM; additionally, visualizations for data structure and estimation results are provided. The current implementation reflects recent developments in literature, uses sensible argument defaults, and aligns commercial and noncommercial estimation commands. Since the understanding of the model assumptions is vital for setting up plausible estimation routines, we provide a broad introduction of linear dynamic panel data models directed towards …


Benchmarking R Packages For Calculation Of Persistent Homology, Eashwar V. Somasundaram, Shael E. Brown, Adam Litzler, Jacob G. Scott, Raoul R. Wadhwa Jun 2021

Benchmarking R Packages For Calculation Of Persistent Homology, Eashwar V. Somasundaram, Shael E. Brown, Adam Litzler, Jacob G. Scott, Raoul R. Wadhwa

The R Journal

Several persistent homology software libraries have been implemented in R. Specifically, the Dionysus, GUDHI, and Ripser libraries have been wrapped by the TDA and TDAstats CRAN packages. These software represent powerful analysis tools that are computationally expensive and, to our knowledge, have not been formally benchmarked. Here, we analyze runtime and memory growth for the 2 R packages and the 3 underlying libraries. We find that datasets with less than 3 dimensions can be evaluated with persistent homology fastest by the GUDHI library in the TDA package. For higher-dimensional datasets, the Ripser library in the TDAstats package is the fastest. …


Unidimensional And Multidimensional Methods For Recurrence Quantification Analysis With Crqa, Moreno I. Coco, Dan Mønster, Giuseppe Leonardi, Rick Dale, Sebastian Wallot Jun 2021

Unidimensional And Multidimensional Methods For Recurrence Quantification Analysis With Crqa, Moreno I. Coco, Dan Mønster, Giuseppe Leonardi, Rick Dale, Sebastian Wallot

The R Journal

Recurrence quantification analysis is a widely used method for characterizing patterns in time series. This article presents a comprehensive survey for conducting a wide range of recurrence-based analyses to quantify the dynamical structure of single and multivariate time series and capture coupling properties underlying leader-follower relationships. The basics of recurrence quantification analysis (RQA) and all its variants are formally introduced step-by-step from the simplest auto-recurrence to the most advanced multivariate case. Importantly, we show how such RQA methods can be deployed under a single computational framework in R using a substantially renewed version of our crqa 2.0 package. This package …


The Bdpar Package: Big Data Pipelining Architecture For R, Miguel Ferreiro-Díaz, Tomás R. Cotos-Yáñez, José R. Méndez, David Ruano-Ordás Jun 2021

The Bdpar Package: Big Data Pipelining Architecture For R, Miguel Ferreiro-Díaz, Tomás R. Cotos-Yáñez, José R. Méndez, David Ruano-Ordás

The R Journal

In the last years, big data has become a useful paradigm for taking advantage of multiple sources to find relevant knowledge in real domains (such as the design of personalized marketing campaigns or helping to palliate the effects of several fatal diseases). Big data programming tools and methods have evolved over time from a MapReduce to a pipeline-based archetype. Concretely the use of pipelining schemes has become the most reliable way of processing and analyzing large amounts of data. To this end, this work introduces bdpar, a new highly customizable pipeline-based framework (using the OOP paradigm provided by R6 …


Exprior: An R Package For The Formulation Of Ex-Situ Priors, Falk Heße, Karina Cucchi, Nura Kawa, Yoram Rubin Jun 2021

Exprior: An R Package For The Formulation Of Ex-Situ Priors, Falk Heße, Karina Cucchi, Nura Kawa, Yoram Rubin

The R Journal

The exPrior package implements a procedure for formulating informative priors of geostatistical properties for a target field site, called ex-situ priors and introduced in Cucchi et al. (2019). The procedure uses a Bayesian hierarchical model to assimilate multiple types of data coming from multiple sites considered as similar to the target site. This prior summarizes the information contained in the data in the form of a probability density function that can be used to better inform further geostatistical investigations at the site. The formulation of the prior uses ex-situ data, where the data set can either be gathered by the …


Linear Regression With Stationary Errors: The R Package Slm, Emmanuel Caron, Jérôme Dedecker, Bertrand Michel Jun 2021

Linear Regression With Stationary Errors: The R Package Slm, Emmanuel Caron, Jérôme Dedecker, Bertrand Michel

The R Journal

This paper introduces the R package slm, which stands for Stationary Linear Models. The package contains a set of statistical procedures for linear regression in the general context where the error process is strictly stationary with a short memory. We work in the setting of Hannan (1973), who proved the asymptotic normality of the (normalized) least squares estimators (LSE) under very mild conditions on the error process. We propose different ways to estimate the asymptotic covariance matrix of the LSE and then to correct the type I error rates of the usual tests on the parameters (as well as confidence …