Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year
File Type

Articles 2251 - 2280 of 8479

Full-Text Articles in Computer Sciences

Watch Your Flavors: Augmenting People's Flavor Perceptions And Associated Emotions Based On Videos Watched While Eating, Meetha Nesam James, Nimesha Ranasinghe, Anthony Tang, Lora Oehlberg May 2022

Watch Your Flavors: Augmenting People's Flavor Perceptions And Associated Emotions Based On Videos Watched While Eating, Meetha Nesam James, Nimesha Ranasinghe, Anthony Tang, Lora Oehlberg

Research Collection School Of Computing and Information Systems

People engage in different activities while eating alone, such as watching television or scrolling through social media on their phones. However, the impacts of these visual contents on human cognitive processes, particularly related to flavor perception and its attributes, are still not thoroughly explored. This paper presents a user study to evaluate the influence of six different types of video content (including nature, cooking, and a new food video genre known as mukbang) on people’s flavor perceptions in terms of taste sensations, liking, and emotions while eating plain white rice. Our findings revealed that the participants’ flavor perceptions are augmented …


Adaptive Task Planning For Large-Scale Robotized Warehouses, Dingyuan Shi, Yongxin Tong, Zimu Zhou, Ke Xu, Wenzhe Tan, Hongbo Li May 2022

Adaptive Task Planning For Large-Scale Robotized Warehouses, Dingyuan Shi, Yongxin Tong, Zimu Zhou, Ke Xu, Wenzhe Tan, Hongbo Li

Research Collection School Of Computing and Information Systems

Robotized warehouses are deployed to automatically distribute millions of items brought by the massive logistic orders from e-commerce. A key to automated item distribution is to plan paths for robots, also known as task planning, where each task is to deliver racks with items to pickers for processing and then return the rack back. Prior solutions are unfit for large-scale robotized warehouses due to the inflexibility to time-varying item arrivals and the low efficiency for high throughput. In this paper, we propose a new task planning problem called TPRW, which aims to minimize the end-to-end makespan that incorporates the entire …


Deep Depression Prediction On Longitudinal Data Via Joint Anomaly Ranking And Classification, Guansong Pang, Ngoc Thien Anh Pham, Emma Baker, Rebecca Bentley, Anton Van Den Hengel May 2022

Deep Depression Prediction On Longitudinal Data Via Joint Anomaly Ranking And Classification, Guansong Pang, Ngoc Thien Anh Pham, Emma Baker, Rebecca Bentley, Anton Van Den Hengel

Research Collection School Of Computing and Information Systems

A wide variety of methods have been developed for identifying depression, but they focus primarily on measuring the degree to which individuals are suffering from depression currently. In this work we explore the possibility of predicting future depression using machine learning applied to longitudinal socio-demographic data. In doing so we show that data such as housing status, and the details of the family environment, can provide cues for predicting future psychiatric disorders. To this end, we introduce a novel deep multi-task recurrent neural network to learn time-dependent depression cues. The depression prediction task is jointly optimized with two auxiliary anomaly …


Xai4fl: Enhancing Spectrum-Based Fault Localization With Explainable Artificial Intelligence, Ratnadira Widyasari, Gede Artha Azriadi Prana, Stefanus Agus Haryono, Yuan Tian, Hafil Noer Zachiary, David Lo May 2022

Xai4fl: Enhancing Spectrum-Based Fault Localization With Explainable Artificial Intelligence, Ratnadira Widyasari, Gede Artha Azriadi Prana, Stefanus Agus Haryono, Yuan Tian, Hafil Noer Zachiary, David Lo

Research Collection School Of Computing and Information Systems

Manually finding the program unit (e.g., class, method, or statement) responsible for a fault is tedious and time-consuming. To mitigate this problem, many fault localization techniques have been proposed. A popular family of such techniques is spectrum-based fault localization (SBFL), which takes program execution traces (spectra) of failed and passed test cases as input and applies a ranking formula to compute a suspiciousness score for each program unit. However, most existing SBFL techniques fail to consider two facts: 1) not all failed test cases contribute equally to a considered fault(s), and 2) program units collaboratively contribute to the failure/pass of …


Competition And Third-Party Platform-Integration In Ride-Sourcing Markets, Yaqian Zhou, Hai Yang, Jintao Ke, Hai Wang, Xinwei Li May 2022

Competition And Third-Party Platform-Integration In Ride-Sourcing Markets, Yaqian Zhou, Hai Yang, Jintao Ke, Hai Wang, Xinwei Li

Research Collection School Of Computing and Information Systems

Recently, some third-party integrators attempt to integrate the ride services offered by multiple independent ride-sourcing platforms. Accordingly, passengers can request ride through the integrators and receive ride service from any one of the ride-sourcing platforms. This novel business model, termed as third-party platform-integration in this work, has potentials to alleviate market fragmentation cost resulting from demand splitting among multiple platforms. Although most existing studies focus on operation strategies for one single monopolist platform, much less is known about the competition and platform-integration and their implications on operation strategy and system efficiency. In this work, we propose mathematical models to describe …


Gdefects4dl: A Dataset Of General Real-World Deep Learning Program Defects, Yunkai Liang, Yun Lin, Xuezhi Song, Jun Sun, Zhiyong Feng, Jin Song Dong May 2022

Gdefects4dl: A Dataset Of General Real-World Deep Learning Program Defects, Yunkai Liang, Yun Lin, Xuezhi Song, Jun Sun, Zhiyong Feng, Jin Song Dong

Research Collection School Of Computing and Information Systems

The development of deep learning programs, as a new programming paradigm, is observed to suffer from various defects. Emerging research works have been proposed to detect, debug, and repair deep learning bugs, which drive the need to construct the bug benchmarks. In this work, we present gDefects4DL, a dataset for general bugs of deep learning programs. Comparing to existing datasets, gDefects4DL collects bugs where the root causes and fix solutions can be well generalized to other projects. Our general bugs include deep learning program bugs such as (1) violation of deep learning API usage pattern (e.g., the standard to implement …


Cost-Effective And Collaborative Methods To Author Video's Scene Description For Blind People, Rosiana Natalie May 2022

Cost-Effective And Collaborative Methods To Author Video's Scene Description For Blind People, Rosiana Natalie

Research Collection School Of Computing and Information Systems

The majority of online video content remains inaccessible for blind people due to the lack of audio descriptions. Content creators have traditionally relied on professionals to author audio descriptions, but their service is costly and not readily available. In this research, I introduce four threads of research that I will conduct for my Ph.D. dissertation, aimed to create methods and tools that are both time- and cost-effective in providing good quality audio descriptions. They are: (i) The development and evaluation of mixed-ability collaboration authoring tool, (ii) The formative study to uncover the feedback pattern from the reviewer, (iii) the evaluation …


Detecting False Alarms From Automatic Static Analysis Tools: How Far Are We?, Hong Jin Kang, Khai Loong Aw, David Lo May 2022

Detecting False Alarms From Automatic Static Analysis Tools: How Far Are We?, Hong Jin Kang, Khai Loong Aw, David Lo

Research Collection School Of Computing and Information Systems

Automatic static analysis tools (ASATs), such as Findbugs, have a high false alarm rate. The large number of false alarms produced poses a barrier to adoption. Researchers have proposed the use of machine learning to prune false alarms and present only actionable warnings to developers. The state-of-the-art study has identified a set of “Golden Features” based on metrics computed over the characteristics and history of the file, code, and warning. Recent studies show that machine learning using these features is extremely effective and that they achieve almost perfect performance. We perform a detailed analysis to better understand the strong performance …


Practitioners' Expectations On Automated Code Comment Generation, Xing Hu, Xin Xia, David Lo, Zhiyuan Wan, Qiuyuan Chen, Thomas Zimmermann May 2022

Practitioners' Expectations On Automated Code Comment Generation, Xing Hu, Xin Xia, David Lo, Zhiyuan Wan, Qiuyuan Chen, Thomas Zimmermann

Research Collection School Of Computing and Information Systems

Good comments are invaluable assets to software projects, as they help developers understand and maintain projects. However, due to some poor commenting practices, comments are often missing or inconsistent with the source code. Software engineering practitioners often spend a significant amount of time and effort reading and understanding programs without or with poor comments. To counter this, researchers have proposed various techniques to automatically generate code comments in recent years, which can not only save developers time writing comments but also help them better understand existing software projects. However, it is unclear whether these techniques can alleviate comment issues and …


On The Effectiveness Of Pretrained Models For Api Learning, Mohammad Abdul Hadi, Imam Nur Bani Yusuf, Thung Ferdian, Gia Kien Luong, Lingxiao Jiang, Fatemeh H. Fard, David Lo May 2022

On The Effectiveness Of Pretrained Models For Api Learning, Mohammad Abdul Hadi, Imam Nur Bani Yusuf, Thung Ferdian, Gia Kien Luong, Lingxiao Jiang, Fatemeh H. Fard, David Lo

Research Collection School Of Computing and Information Systems

Developers frequently use APIs to implement certain functionalities, such as parsing Excel Files, reading and writing text files line by line, etc. Developers can greatly benefit from automatic API usage sequence generation based on natural language queries for building applications in a faster and cleaner manner. Existing approaches utilize information retrieval models to search for matching API sequences given a query or use RNN-based encoder-decoder to generate API sequences. As it stands, the first approach treats queries and API names as bags of words. It lacks deep comprehension of the semantics of the queries. The latter approach adapts a neural …


Is Surprisal In Issue Trackers Actionable?, James Caddy, Markus Wagner, Christoph Treude, Earl T. Barr, Miltiadis Allamanis May 2022

Is Surprisal In Issue Trackers Actionable?, James Caddy, Markus Wagner, Christoph Treude, Earl T. Barr, Miltiadis Allamanis

Research Collection School Of Computing and Information Systems

Background. From information theory, surprisal is a measurement of how unexpected an event is. Statistical language models provide a probabilistic approximation of natural languages, and because surprisal is constructed with the probability of an event occuring, it is therefore possible to determine the surprisal associated with English sentences. The issues and pull requests of software repository issue trackers give insight into the development process and likely contain the surprising events of this process. Objective. Prior works have identified that unusual events in software repositories are of interest to developers, and use simple code metrics-based methods for detecting them. In this …


Linkbreaker: Breaking The Backdoor-Trigger Link In Dnns Via Neurons Consistency Check, Zhenzhu Chen, Shang Wang, Anmin Fu, Yansong Gao, Shui Yu, Robert H. Deng May 2022

Linkbreaker: Breaking The Backdoor-Trigger Link In Dnns Via Neurons Consistency Check, Zhenzhu Chen, Shang Wang, Anmin Fu, Yansong Gao, Shui Yu, Robert H. Deng

Research Collection School Of Computing and Information Systems

Backdoor attacks cause model misbehaving by first implanting backdoors in deep neural networks (DNNs) during training and then activating the backdoor via samples with triggers during inference. The compromised models could pose serious security risks to artificial intelligence systems, such as misidentifying 'stop' traffic sign into '80km/h'. In this paper, we investigate the connection characteristic between the backdoor and the trigger in DNNs and observe the fact that the backdoor is implanted via establishing a link between a cluster of neurons, representing the backdoor, and the triggers. Based on this observation, we design LinkBreaker, a new generic scheme for defending …


Topic-Guided Conversational Recommender In Multiple Domains, Lizi Liao, Ryuichi Takanobu, Yunshan Ma, Xun Yang, Minlie Huang, Tat-Seng Chua May 2022

Topic-Guided Conversational Recommender In Multiple Domains, Lizi Liao, Ryuichi Takanobu, Yunshan Ma, Xun Yang, Minlie Huang, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Conversational systems have recently attracted significant attention. Both the research community and industry believe that it will exert huge impact on human-computer interaction, and specifically, the IR/RecSys community has begun to explore Conversational Recommendation. In real-life scenarios, such systems are often urgently needed in helping users accomplishing different tasks under various situations. However, existing works still face several shortcomings: (1) Most efforts are largely confined in single task setting. They fall short of hands in handling tasks across domains. (2) Aside from soliciting user preference from dialogue history, a conversational recommender naturally has access to the back-end data structure which …


Quid Pro Quo: An Exploration Of Reciprocity In Code Review, Carlos Gavidia-Calderon, Donggyun Han, Amel Bennaceur May 2022

Quid Pro Quo: An Exploration Of Reciprocity In Code Review, Carlos Gavidia-Calderon, Donggyun Han, Amel Bennaceur

Research Collection School Of Computing and Information Systems

We explore the role of reciprocity in code review processes. Reciprocity manifests itself in two ways: 1) reviewing code for others translates to accepted code contributions, and 2) having contributions accepted increases the reviews made for others. We use vector autoregressive (VAR) models to explore the causal relation between reviews performed and accepted contributions. After fitting VAR models for 24 active open-source developers, we found evidence of reciprocity in 6 of them. These results suggest reciprocity does play a role in code review, that can potentially be exploited to increase reviewer participation.


Software Engineering User Study Recruitment On Prolific: An Experience Report, Brittany Reid, Markus Wagner, Marcelo D’Amorim, Christoph Treude May 2022

Software Engineering User Study Recruitment On Prolific: An Experience Report, Brittany Reid, Markus Wagner, Marcelo D’Amorim, Christoph Treude

Research Collection School Of Computing and Information Systems

Online participant recruitment platforms such as Prolific have been gaining popularity in research, as they enable researchers to easily access large pools of participants. However, participant quality can be an issue; participants may give incorrect information to gain access to more studies, adding unwanted noise to results. This paper details our experience recruiting participants from Prolific for a user study requiring programming skills in Node.js, with the aim of helping other researchers conduct similar studies. We explore a method of recruiting programmer participants using prescreening validation, attention checks and a series of programming knowledge questions. We received 680 responses, and …


On Recruiting Experienced Github Contributors For Interviews And Surveys On Prolific, Felipe Ebert, Alexander Serebrenik, Christoph Treude, Nicole Novielli, Fernando Castor May 2022

On Recruiting Experienced Github Contributors For Interviews And Surveys On Prolific, Felipe Ebert, Alexander Serebrenik, Christoph Treude, Nicole Novielli, Fernando Castor

Research Collection School Of Computing and Information Systems

Software engineering researchers have been using general purpose online tools for crowd-sourcing for quite some time. Those tools can be useful to recruit participants for research studies as they are paid for their time. However, those tools should be used carefully. In this paper, we have described the issues we faced when recruiting participants on Prolific. We used Prolific to recruit open-source developers with experience in submitting and reviewing pull requests (PRs). However, we did not succeed in obtaining valid participants for either the interview or the survey, which led us to change the approach of our study and not …


Diffusion Of Ai Governance, Langtao Chen, Brenda Eschenbrenner, Fiona Fui-Hoon Nah, Keng Siau, Yuzhou Qian May 2022

Diffusion Of Ai Governance, Langtao Chen, Brenda Eschenbrenner, Fiona Fui-Hoon Nah, Keng Siau, Yuzhou Qian

Research Collection School Of Computing and Information Systems

Artificial intelligence (AI) has the potential to address social, economic, and environmental challenges. However, effective use of AI in organizations relies on the establishment of an AI governance framework. Although existing studies have discussed a variety of issues raised by AI-based systems and proposed AI governance frameworks to overcome those issues, organizations face challenges in adopting AI governance. Informed by innovation diffusion theory, this research evaluates the impact of internal and external influences on AI governance adoption between highly regulated and less regulated industries. We also assess the effect of adopting AI governance on organizational performance. Findings from this study …


Context Modeling With Evidence Filter For Multiple Choice Question Answering, Sicheng Yu, Hao Zhang, Wei Jing, Jing Jiang May 2022

Context Modeling With Evidence Filter For Multiple Choice Question Answering, Sicheng Yu, Hao Zhang, Wei Jing, Jing Jiang

Research Collection School Of Computing and Information Systems

Multiple-Choice Question Answering (MCQA) is one of the challenging tasks in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that supports the correct answer. In OpenbookQA dataset [1], the requirement of extracting "evidence" is particularly important due to the mutual independence of sentences in the context. Existing work tackles this problem by annotated evidence or distant supervision with rules which overly rely on human efforts. To address the challenge, we propose a simple yet effective approach termed evidence filtering to model the relationships between the encoded contexts with respect to different options …


Smile: Secure Memory Introspection For Live Enclave, Lei Zhou, Xuhua Ding, Zhang Fengwei May 2022

Smile: Secure Memory Introspection For Live Enclave, Lei Zhou, Xuhua Ding, Zhang Fengwei

Research Collection School Of Computing and Information Systems

SGX enclaves prevent external software from accessing their memory. This feature conflicts with legitimate needs for enclave memory introspection, e.g., runtime stack collection on an enclave under a return-oriented-programming attack. We propose SMILE for enclave owners to acquire live enclave contents with the assistance of a semi-trusted agent installed by the host platform’s vendor as a plug-in of the System Management Interrupt handler. SMILE authenticates the enclave under introspection without trusting the kernel nor depending on the SGX attestation facility. SMILE is enclave security preserving as breaking of SMILE does not undermine enclave security. It allows a cloud server to …


Benchmarking Library Recognition In Tweets, Ting Zhang, Divya Prabha Chandrasekaran, Ferdian Thung, David Lo May 2022

Benchmarking Library Recognition In Tweets, Ting Zhang, Divya Prabha Chandrasekaran, Ferdian Thung, David Lo

Research Collection School Of Computing and Information Systems

Software developers often use social media (such as Twitter) to shareprogramming knowledge such as new tools, sample code snippets,and tips on programming. One of the topics they talk about is thesoftware library. The tweets may contain useful information abouta library. A good understanding of this information, e.g., on thedeveloper’s views regarding a library can be beneficial to weigh thepros and cons of using the library as well as the general sentimentstowards the library. However, it is not trivial to recognize whethera word actually refers to a library or other meanings. For example,a tweet mentioning the word “pandas" may refer to …


Structure-Aware Visualization Retrieval, Haotian Li, Yong Wang, Aoyu Wu, Huan Wei, Huamin. Qu May 2022

Structure-Aware Visualization Retrieval, Haotian Li, Yong Wang, Aoyu Wu, Huan Wei, Huamin. Qu

Research Collection School Of Computing and Information Systems

With the wide usage of data visualizations, a huge number of Scalable Vector Graphic (SVG)-based visualizations have been created and shared online. Accordingly, there has been an increasing interest in exploring how to retrieve perceptually similar visualizations from a large corpus, since it can benefit various downstream applications such as visualization recommendation. Existing methods mainly focus on the visual appearance of visualizations by regarding them as bitmap images. However, the structural information intrinsically existing in SVG-based visualizations is ignored. Such structural information can delineate the spatial and hierarchical relationship among visual elements, and characterize visualizations thoroughly from a new perspective. …


Who Will Support My Project? Interactive Search Of Potential Crowdfunding Investors Through Insearch., Songheng Zhang, Yong Wang, Haotian Li, Wanyu Zhang May 2022

Who Will Support My Project? Interactive Search Of Potential Crowdfunding Investors Through Insearch., Songheng Zhang, Yong Wang, Haotian Li, Wanyu Zhang

Research Collection School Of Computing and Information Systems

Crowdfunding provides project founders with a convenient way to reach online investors. However, it is challenging for founders to find the most potential investors and successfully raise money for their projects on crowdfunding platforms. A few machine learning based methods have been proposed to recommend investors’ interest in a specific crowdfunding project, but they fail to provide project founders with explanations in detail for these recommendations, thereby leading to an erosion of trust in predicted investors. To help crowdfunding founders find truly interested investors, we conducted semi-structured interviews with four crowdfunding experts and presentsinSearch, a visual analytic system. inSearch allows …


Rumorlens: Interactive Analysis And Validation Of Suspected Rumors On Social Media, Ran Wang, Kehan Du, Qianhe Chen, Yifei Zhao, Mojie Tang, Hongxi Tao, Shipan Wang, Yiyao Li, Yong Wang May 2022

Rumorlens: Interactive Analysis And Validation Of Suspected Rumors On Social Media, Ran Wang, Kehan Du, Qianhe Chen, Yifei Zhao, Mojie Tang, Hongxi Tao, Shipan Wang, Yiyao Li, Yong Wang

Research Collection School Of Computing and Information Systems

With the development of social media, various rumors can be easily spread on the Internet and such rumors can have serious negative effects on society. Thus, it has become a critical task for social media platforms to deal with suspected rumors. However, due to the lack of effective tools, it is often difficult for platform administrators to analyze and validate rumors from a large volume of information on a social media platform efficiently. We have worked closely with social media platform administrators for four months to summarize their requirements of identifying and analyzing rumors, and further proposed an interactive visual …


Static Inference Meets Deep Learning: A Hybrid Type Inference Approach For Python, Yun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao, David Lo, Qirun Zhang, Michael R. Lyu May 2022

Static Inference Meets Deep Learning: A Hybrid Type Inference Approach For Python, Yun Peng, Cuiyun Gao, Zongjie Li, Bowei Gao, David Lo, Qirun Zhang, Michael R. Lyu

Research Collection School Of Computing and Information Systems

Type inference for dynamic programming languages such as Python is an important yet challenging task. Static type inference techniques can precisely infer variables with enough static constraints but are unable to handle variables with dynamic features. Deep learning (DL) based approaches are feature-agnostic, but they cannot guarantee the correctness of the predicted types. Their performance significantly depends on the quality of the training data (i.e., DL models perform poorly on some common types that rarely appear in the training dataset). It is interesting to note that the static and DL-based approaches offer complementary benefits. Unfortunately, to our knowledge, precise type …


Ptm4tag: Sharpening Tag Recommendation Of Stack Overflow Posts With Pre-Trained Models, Junda He, Bowen Xu, Zhou Yang, Donggyun Han, Chengran Yang, David Lo May 2022

Ptm4tag: Sharpening Tag Recommendation Of Stack Overflow Posts With Pre-Trained Models, Junda He, Bowen Xu, Zhou Yang, Donggyun Han, Chengran Yang, David Lo

Research Collection School Of Computing and Information Systems

Stack Overflow is often viewed as one of the most influential Software Question & Answer (SQA) websites, containing millions of programming-related questions and answers. Tags play a critical role in efficiently structuring the contents in Stack Overflow and are vital to support a range of site operations, e.g., querying relevant contents. Poorly selected tags often introduce extra noise and redundancy, which raises problems like tag synonym and tag explosion. Thus, an automated tag recommendation technique that can accurately recommend high-quality tags is desired to alleviate the problems mentioned above.


Automated Identification Of Libraries From Vulnerability Data: Can We Do Better?, Stefanus A. Haryono, Hong Jin Kang, Abhishek Sharma, Asankhaya Sharma, Andrew E. Santosa, Ming Yi Ang, David Lo May 2022

Automated Identification Of Libraries From Vulnerability Data: Can We Do Better?, Stefanus A. Haryono, Hong Jin Kang, Abhishek Sharma, Asankhaya Sharma, Andrew E. Santosa, Ming Yi Ang, David Lo

Research Collection School Of Computing and Information Systems

Software engineers depend heavily on software libraries and have to update their dependencies once vulnerabilities are found in them. Software Composition Analysis (SCA) helps developers identify vulnerable libraries used by an application. A key challenge is the identification of libraries related to a given reported vulnerability in the National Vulnerability Database (NVD), which may not explicitly indicate the affected libraries. Recently, researchers have tried to address the problem of identifying the libraries from an NVD report by treating it as an extreme multi-label learning (XML) problem, characterized by its large number of possible labels and severe data sparsity. As input, …


Simple Or Complex? Together For A More Accurate Just-In-Time Defect Predictor, Xin Zhou, Donggyun Han, David Lo May 2022

Simple Or Complex? Together For A More Accurate Just-In-Time Defect Predictor, Xin Zhou, Donggyun Han, David Lo

Research Collection School Of Computing and Information Systems

Just-In-Time (JIT) defect prediction aims to automatically predict whether a commit is defective or not, and has been widely studied in recent years. In general, most studies can be classified into two categories: 1) simple models using traditional machine learning classifiers with hand-crafted features, and 2) complex models using deep learning techniques to automatically extract features. Hand-crafted features used by simple models are based on expert knowledge but may not fully represent the semantic meaning of the commits. On the other hand, deep learning-based features used by complex models represent the semantic meaning of commits but may not reflect useful …


On The Transferability Of Pre-Trained Language Models For Low-Resource Programming Languages, Fuxiang Chen, Fatemeh H. Fard, David Lo, Timofey Bryksin May 2022

On The Transferability Of Pre-Trained Language Models For Low-Resource Programming Languages, Fuxiang Chen, Fatemeh H. Fard, David Lo, Timofey Bryksin

Research Collection School Of Computing and Information Systems

A recent study by Ahmed and Devanbu reported that using a corpus of code written in multilingual datasets to fine-tune multilingual Pre-trained Language Models (PLMs) achieves higher performance as opposed to using a corpus of code written in just one programming language. However, no analysis was made with respect to fine-tuning monolingual PLMs. Furthermore, some programming languages are inherently different and code written in one language usually cannot be interchanged with the others, i.e., Ruby and Java code possess very different structure. To better understand how monolingual and multilingual PLMs affect different programming languages, we investigate 1) the performance of …


Arsearch: Searching For Api Related Resources From Stack Overflow And Github, Kien Luong, Ferdian Thung, David Lo May 2022

Arsearch: Searching For Api Related Resources From Stack Overflow And Github, Kien Luong, Ferdian Thung, David Lo

Research Collection School Of Computing and Information Systems

Stack Overflow and GitHub are two popular platforms containing API-related resources for developers to learn how to use APIs. The platforms are good sources for information about API such as code examples, usages, sentiment, bug reports, etc. However, it is difficult to collect the correct resources regarding a particular API due to the ambiguity of an API method name. An API method name mentioned in the text would only refer to one API, but the method name could match with different APIs. To help people in finding the correct resources for a particular API, we introduce ARSearch. ARSearch finds Stack …


Structure-Aware Visualization Retrieval, Haotian Li, Yong Wang, Wu Aoyu, Huan Wei, Huamin Qu May 2022

Structure-Aware Visualization Retrieval, Haotian Li, Yong Wang, Wu Aoyu, Huan Wei, Huamin Qu

Research Collection School Of Computing and Information Systems

With the wide usage of data visualizations, a huge number of Scalable Vector Graphic (SVG)-based visualizations have been created and shared online. Accordingly, there has been an increasing interest in exploring how to retrieve perceptually similar visualizations from a large corpus, since it can beneft various downstream applications such as visualization recommendation. Existing methods mainly focus on the visual appearance of visualizations by regarding them as bitmap images. However, the structural information intrinsically existing in SVG-based visualizations is ignored. Such structural information can delineate the spatial and hierarchical relationship among visual elements, and characterize visualizations thoroughly from a new perspective. …