Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (239)
- Programming Languages and Compilers (198)
- Artificial Intelligence and Robotics (142)
- Engineering (142)
- Computer Engineering (126)
-
- Information Security (109)
- OS and Networks (80)
- Graphics and Human Computer Interfaces (78)
- Social and Behavioral Sciences (76)
- Numerical Analysis and Scientific Computing (59)
- Theory and Algorithms (46)
- Business (43)
- Computer and Systems Architecture (39)
- Digital Communications and Networking (39)
- Medicine and Health Sciences (33)
- Communication (30)
- Education (28)
- Systems Architecture (21)
- Health Information Technology (20)
- Sociology (18)
- Finance and Financial Management (15)
- Gerontology (15)
- Public Affairs, Public Policy and Public Administration (15)
- Higher Education (13)
- Social Media (13)
- Transportation (12)
- Technology and Innovation (11)
- Keyword
-
- Deep learning (54)
- Software engineering (49)
- Empirical study (43)
- Machine learning (35)
- Software (30)
-
- Android (29)
- Model Check (29)
- Collaboration (26)
- Deep Learning (24)
- GitHub (22)
- Stack Overflow (22)
- Fuzzing (21)
- Security (21)
- Testing (20)
- Data mining (19)
- Large language models (18)
- Programming (18)
- Software Engineering (18)
- Code search (17)
- Computer bugs (17)
- Codes (16)
- Empirical Study (16)
- Information retrieval (16)
- Large Language Models (16)
- Large language model (16)
- Software testing (15)
- Large Language Model (14)
- Linear Temporal Logic (14)
- Software maintenance (14)
- Vulnerability detection (14)
- Publication Year
- Publication
- Publication Type
Articles 841 - 870 of 2211
Full-Text Articles in Software Engineering
Novel Deep Learning Methods Combined With Static Analysis For Source Code Processing, Duy Quoc Nghi Bui
Novel Deep Learning Methods Combined With Static Analysis For Source Code Processing, Duy Quoc Nghi Bui
Dissertations and Theses Collection (Open Access)
It is desirable to combine machine learning and program analysis so that one can leverage the best of both to increase the performance of software analytics. On one side, machine learning can analyze the source code of thousands of well-written software projects that can uncover patterns that partially characterize software that is reliable, easy to read, and easy to maintain. On the other side, the program analysis can be used to define rigorous and unique rules that are only available in programming languages, which enrich the representation of source code and help the machine learning to capture the patterns better. …
Spark: Spatial-Aware Online Incremental Attack Against Visual Tracking, Qing Guo, Xiaofei Xie, Felix Juefei-Xu, Lei Ma, Zhongguo Li, Wanli Xue, Wei Feng, Yang Liu
Spark: Spatial-Aware Online Incremental Attack Against Visual Tracking, Qing Guo, Xiaofei Xie, Felix Juefei-Xu, Lei Ma, Zhongguo Li, Wanli Xue, Wei Feng, Yang Liu
Research Collection School Of Computing and Information Systems
Adversarial attacks of deep neural networks have been intensively studied on image, audio, and natural language classification tasks. Nevertheless, as a typical while important real-world application, the adversarial attacks of online video tracking that traces an object’s moving trajectory instead of its category are rarely explored. In this paper, we identify a new task for the adversarial attack to visual tracking: online generating imperceptible perturbations that mislead trackers along with an incorrect (Untargeted Attack, UA) or specified trajectory (Targeted Attack, TA). To this end, we first propose a spatial-aware basic attack by adapting existing attack methods, i.e., FGSM, BIM, and …
Commanding And Re-Dictation: Developing Eyes-Free Voice-Based Interaction For Editing Dictated Text, Debjyoti Ghosh, Can Liu, Shengdong Zhao, Kotaro Hara
Commanding And Re-Dictation: Developing Eyes-Free Voice-Based Interaction For Editing Dictated Text, Debjyoti Ghosh, Can Liu, Shengdong Zhao, Kotaro Hara
Research Collection School Of Computing and Information Systems
Existing voice-based interfaces have limited support for text editing, especially when seeing the text is difficult, e.g., while walking or cooking. This research develops voice interaction techniques for eyes-free text editing. First, with a Wizard-of-Oz study, we identified two primary user strategies: using commands, e.g., “replace go with goes” and re-dictating over an erroneous portion, e.g., correcting “he go there” by saying “he goes there.” To support these user strategies with an actual system implementation, we developed two eyes-free voice interaction techniques, Commanding and Re-dictation, and evaluated them with a controlled experiment. Results showed that while Re-dictation performs significantly better …
Prevalence, Contents And Automatic Detection Of Kl-Satd, Leevi Rantala, Mika Mantyla, David Lo
Prevalence, Contents And Automatic Detection Of Kl-Satd, Leevi Rantala, Mika Mantyla, David Lo
Research Collection School Of Computing and Information Systems
When developers use different keywords such as TODO and FIXME in source code comments to describe self-admitted technical debt (SATD), we refer it as Keyword-Labeled SATD (KL-SATD). We study KL-SATD from 33 software repositories with 13,588 KL-SATD comments. We find that the median percentage of KL-SATD comments among all comments is only 1,52%. We find that KL-SATD comment contents include words expressing code changes and uncertainty, such as remove, fix, maybe and probably. This makes them different compared to other comments. KL-SATD comment contents are similar to manually labeled SATD comments of prior work. Our machine learning classifier using logistic …
A Longitudinal Study Of A Capstone Course, Benjamin Gan, Eng Lieh Ouh, Yin Yin Fiona Lee
A Longitudinal Study Of A Capstone Course, Benjamin Gan, Eng Lieh Ouh, Yin Yin Fiona Lee
Research Collection School Of Computing and Information Systems
This is a 7 years study on a capstone course completed by 1700+ students for 200+ organizations involving 300+ projects. Student teams deliver a system to solve real-world problems proposed by industry partners. We want to understand what independent variables influence student performance. We analyzed the deployment status of systems delivered, the type of organization/industry, the number of meetings and the technology used. Our results show some organization value proof of concept over fully deployed systems, student strengths are in Infocomm and Finance projects, the number of meetings is a weak correlation to performance and best performing projects are fully …
Rethinking Pruning For Accelerating Deep Inference At The Edge, Dawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong, Ke Xu, Lothar Thiele
Rethinking Pruning For Accelerating Deep Inference At The Edge, Dawei Gao, Xiaoxi He, Zimu Zhou, Yongxin Tong, Ke Xu, Lothar Thiele
Research Collection School Of Computing and Information Systems
There is a growing trend to deploy deep neural networks at the edge for high-accuracy, real-time data mining and user interaction. Applications such as speech recognition and language understanding often apply a deep neural network to encode an input sequence and then use a decoder to generate the output sequence. A promising technique to accelerate these applications on resource-constrained devices is network pruning, which compresses the size of the deep neural network without severe drop in inference accuracy. However, we observe that although existing network pruning algorithms prove effective to speed up the prior deep neural network, they lead to …
How Practitioners Perceive Automated Bug Report Management Techniques, Weiqin Zou, David Lo, Zhenyu Chen, Xin Xia, Yang Feng, Baowen Xu
How Practitioners Perceive Automated Bug Report Management Techniques, Weiqin Zou, David Lo, Zhenyu Chen, Xin Xia, Yang Feng, Baowen Xu
Research Collection School Of Computing and Information Systems
Bug reports play an important role in the process of debugging and fixing bugs. To reduce the burden of bug report managers and facilitate the process of bug fixing, a great amount of software engineering research has been invested into automated bug report management techniques. However, the verdict is still open whether such techniques are actually required and applicable outside of the theoretical research domain. To fill this gap, in this paper, we conducted a survey among 327 practitioners to gain their insights into various categories of automated bug report management techniques. Specifically, in the survey, we asked them to …
Refactoring From 9 To 5? What And When Employees And Volunteers Contribute To Oss, Luiz Felipe Dias, Caio Barbosa, Gustavo Pinto, Igor Steinmacher, Baldoino Fonseca, Márcio Ribeiro, Christoph Treude, Daniel Alencar Da Costa
Refactoring From 9 To 5? What And When Employees And Volunteers Contribute To Oss, Luiz Felipe Dias, Caio Barbosa, Gustavo Pinto, Igor Steinmacher, Baldoino Fonseca, Márcio Ribeiro, Christoph Treude, Daniel Alencar Da Costa
Research Collection School Of Computing and Information Systems
In this paper we characterize the contributions made by employees (developers that work for GitHub, the company) and volunteers (developers that use GitHub, the platform) to OSS projects maintained by GitHub (the company) on GitHub (the platform). By mining activities performed in five well-known company-owned OSS projects, we investigate what they do and when they do it. We found that the majority of the volunteers' contributions are related to reengineering (e.g., refactoring), while employees focus more on management (e.g., documentation). When it comes to the working hours, we found that contributions are made mostly from 9am-5pm, even for the volunteers.
Keen2act: Activity Recommendation In Online Social Collaborative Platforms, Roy Ka-Wei Lee, Thong Hoang, Richard J. Oentaryo, David Lo
Keen2act: Activity Recommendation In Online Social Collaborative Platforms, Roy Ka-Wei Lee, Thong Hoang, Richard J. Oentaryo, David Lo
Research Collection School Of Computing and Information Systems
Social collaborative platforms such as GitHub and Stack Overflow have been increasingly used to improve work productivity via collaborative efforts. To improve user experiences in these platforms, it is desirable to have a recommender system that can suggest not only items (e.g., a GitHub repository) to a user, but also activities to be performed on the suggested items (e.g., forking a repository). To this end, we propose a new approach dubbed Keen2Act, which decomposes the recommendation problem into two stages: the Keen and Act steps. The Keen step identifies, for a given user, a (sub)set of items in which he/she …
Sentiment Analysis Over Collaborative Relationships In Open Source Software Projects, Lingjia Li, Jian Cao, David Lo
Sentiment Analysis Over Collaborative Relationships In Open Source Software Projects, Lingjia Li, Jian Cao, David Lo
Research Collection School Of Computing and Information Systems
Sentiments and collaboration efficiency are key factors in the success of the open source software (OSS) development process. However, in the software engineering domain, no studies have been conducted to analyze the effect between collaborators' sentiments, and the role of sentiment in collaborative relationships during the development process. In this study, we apply sentiment analysis and statistical analysis on collaboration artifacts over five projects on GitHub. We use sentiment consistency to quantify the relation between sentiments in collaborative relationships. It is found that sentiment consistency is positively correlated with the closeness of collaborative relationships and collaborators' overall sentiment states. We …
Automatic Android Deprecated-Api Usage Update By Learning From Single Updated Example, Stefanus A. Haryono, Ferdian Thung, Hong Jin Kang, Lucas Serrano, Gilles Muller, Julia Lawall, David Lo, Lingxiao Jiang
Automatic Android Deprecated-Api Usage Update By Learning From Single Updated Example, Stefanus A. Haryono, Ferdian Thung, Hong Jin Kang, Lucas Serrano, Gilles Muller, Julia Lawall, David Lo, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Due to the deprecation of APIs in the Android operating system, developers have to update usages of the APIs to ensure that their applications work for both the past and current versions of Android. Such updates may be widespread, non-trivial, and time-consuming. Therefore, automation of such updates will be of great benefit to developers. AppEvolve, which is the state-of-the-art tool for automating such updates, relies on having before- and after-update examples to learn from. In this work, we propose an approach named CocciEvolve that performs such updates using only a single after-update example. CocciEvolve learns edits by extracting the relevant …
Psc2code: Denoising Code Extraction From Programming Screencasts, Lingfeng Bao, Zhenchang Xing, Xin Xia, David Lo, Minghui Wu, Xiaohu Yang
Psc2code: Denoising Code Extraction From Programming Screencasts, Lingfeng Bao, Zhenchang Xing, Xin Xia, David Lo, Minghui Wu, Xiaohu Yang
Research Collection School Of Computing and Information Systems
Programming screencasts have become a pervasive resource on the Internet, which help developers learn new programming technologies or skills. The source code in programming screencasts is an important and valuable information for developers. But the streaming nature of programming screencasts (i.e., a sequence of screen-captured images) limits the ways that developers can interact with the source code in the screencasts. Many studies use the Optical Character Recognition (OCR) technique to convert screen images (also referred to as video frames) into textual content, which can then be indexed and searched easily. However, noisy screen images significantly affect the quality of source …
Automated Synthesis Of Local Time Requirement For Service Composition, Étienne André, Tian Huat Tan, Manman Chen, Shuang Liu, Jun Sun, Yang Liu, Jin Song Dong
Automated Synthesis Of Local Time Requirement For Service Composition, Étienne André, Tian Huat Tan, Manman Chen, Shuang Liu, Jun Sun, Yang Liu, Jin Song Dong
Research Collection School Of Computing and Information Systems
Service composition aims at achieving a business goal by composing existing service-based applications or components. The response time of a service is crucial, especially in time-critical business environments, which is often stated as a clause in service-level agreements between service providers and service users. To meet the guaranteed response time requirement of a composite service, it is important to select a feasible set of component services such that their response time will collectively satisfy the response time requirement of the composite service. In this work, we use the BPEL modeling language that aims at specifying Web services. We extend it …
Objsim: Efficient Testing Of Cyber-Physical Systems, Jun Sun, Zijiang Yang
Objsim: Efficient Testing Of Cyber-Physical Systems, Jun Sun, Zijiang Yang
Research Collection School Of Computing and Information Systems
Cyber-physical systems (CPSs) play a critical role in automating public infrastructure and thus attract wide range of attacks. Assessing the effectiveness of defense mechanisms is challenging as realistic sets of attacks to test them against are not always available. In this short paper, we briefly describe smart fuzzing, an automated, machine learning guided technique for systematically producing test suites of CPS network attacks. Our approach uses predictive ma- chine learning models and meta-heuristic search algorithms to guide the fuzzing of actuators so as to drive the CPS into different unsafe physical states. The approach has been proven effective on two …
Recovering Fitness Gradients For Interprocedural Boolean Flags In Search-Based Testing, Yun Lin, Jun Sun, Gordon Fraser, Ziheng Xiu, Ting Liu, Jin Song Dong
Recovering Fitness Gradients For Interprocedural Boolean Flags In Search-Based Testing, Yun Lin, Jun Sun, Gordon Fraser, Ziheng Xiu, Ting Liu, Jin Song Dong
Research Collection School Of Computing and Information Systems
In Search-based Software Testing (SBST), test generation is guided by fitness functions that estimate how close a test case is to reach an uncovered test goal (e.g., branch). A popular fitness function estimates how close conditional statements are to evaluating to true or false, i.e., the branch distance. However, when conditions read Boolean variables (e.g., if(x && y)), the branch distance provides no gradient for the search, since a Boolean can either be true or false. This flag problem can be addressed by transforming individual procedures such that Boolean flags are replaced with numeric comparisons that provide better guidance for …
Global Pac Bounds For Learning Discrete Time Markov Chains, Hugo Bazille, Blaise Genest, Cyrille Jegourel, Jun Sun
Global Pac Bounds For Learning Discrete Time Markov Chains, Hugo Bazille, Blaise Genest, Cyrille Jegourel, Jun Sun
Research Collection School Of Computing and Information Systems
Learning models from observations of a system is a powerful tool with many applications. In this paper, we consider learning Discrete Time Markov Chains (DTMC), with different methods such as frequency estimation or Laplace smoothing. While models learnt with such methods converge asymptotically towards the exact system, a more practical question in the realm of trusted machine learning is how accurate a model learnt with a limited time budget is. Existing approaches provide bounds on how close the model is to the original system, in terms of bounds on local (transition) probabilities, which has unclear implication on the global behavior. …
Active Fuzzing For Testing And Securing Cyber-Physical Systems, Yuqi Chen, Bohan Xuan, Christopher M. Poskitt, Jun Sun, Fan Zhang
Active Fuzzing For Testing And Securing Cyber-Physical Systems, Yuqi Chen, Bohan Xuan, Christopher M. Poskitt, Jun Sun, Fan Zhang
Research Collection School Of Computing and Information Systems
Cyber-physical systems (CPSs) in critical infrastructure face a pervasive threat from attackers, motivating research into a variety of countermeasures for securing them. Assessing the effectiveness of these countermeasures is challenging, however, as realistic benchmarks of attacks are difficult to manually construct, blindly testing is ineffective due to the enormous search spaces and resource requirements, and intelligent fuzzing approaches require impractical amounts of data and network access. In this work, we propose active fuzzing, an automatic approach for finding test suites of packet-level CPS network attacks, targeting scenarios in which attackers can observe sensors and manipulate packets, but have no existing …
Spinfer: Inferring Semantic Patches For The Linux Kernel, Lucas Serrano, Van-Anh Nguyen, Ferdian Thung, Lingxiao Jiang, David Lo, Julia Lawall, Gilles Muller
Spinfer: Inferring Semantic Patches For The Linux Kernel, Lucas Serrano, Van-Anh Nguyen, Ferdian Thung, Lingxiao Jiang, David Lo, Julia Lawall, Gilles Muller
Research Collection School Of Computing and Information Systems
In a large software system such as the Linux kernel, there is a continual need for large-scale changes across many source files, triggered by new needs or refined design decisions. In this paper, we propose to ease such changes by suggesting transformation rules to developers, inferred automatically from a collection of examples. Our approach can help automate large-scale changes as well as help understand existing large-scale changes, by highlighting the various cases that the developer who performed the changes has taken into account. We have implemented our approach as a tool, Spinfer. We evaluate Spinfer on a range of challenging …
Optimising The Fit Of Stack Overflow Code Snippets Into Existing Code, Brittany Reid, Christoph Treude, Markus Wagner
Optimising The Fit Of Stack Overflow Code Snippets Into Existing Code, Brittany Reid, Christoph Treude, Markus Wagner
Research Collection School Of Computing and Information Systems
Software developers often reuse code from online sources such as Stack Overflow within their projects. However, the process of searching for code snippets and integrating them within existing source code can be tedious. In order to improve efficiency and reduce time spent on code reuse, we present an automated code reuse tool for the Eclipse IDE (Integrated Developer Environment), NLP2TestableCode. NLP2TestableCode can not only search for Java code snippets using natural language tasks, but also evaluate code snippets based on a user’s existing code, modify snippets to improve fit and correct errors, before presenting the user with the best snippet, …
Mining And Predicting Micro-Process Patterns Of Issue Resolution For Open Source Software Projects, Yiran Wang, Jian Cao, David Lo
Mining And Predicting Micro-Process Patterns Of Issue Resolution For Open Source Software Projects, Yiran Wang, Jian Cao, David Lo
Research Collection School Of Computing and Information Systems
Addressing issue reports is an integral part of open source software (OSS) projects. Although several studies have attempted to discover the factors that affect issue resolution, few pay attention to the underlying micro-process patterns of resolution processes. Discovering these micro-patterns will help us understand the dynamics of issue resolution processes so that we can manage and improve them in better ways. Of the various types of issues, those relating to corrective maintenance account for nearly half hence resolving these issues efficiently is critical for the success of OSS projects. Therefore, we apply process mining techniques to discover the micro-patterns of …
How Are Deep Learning Models Similar? An Empirical Study On Clone Analysis Of Deep Learning Software, Xiongfei Wu, Liangyu Qin, Bing Yu, Xiaofei Xie, Lei Ma, Yinxing Xue, Yang Liu, Jianjun Zhao
How Are Deep Learning Models Similar? An Empirical Study On Clone Analysis Of Deep Learning Software, Xiongfei Wu, Liangyu Qin, Bing Yu, Xiaofei Xie, Lei Ma, Yinxing Xue, Yang Liu, Jianjun Zhao
Research Collection School Of Computing and Information Systems
Deep learning (DL) has been successfully applied to many cutting-edge applications, e.g., image processing, speech recognition, and natural language processing. As more and more DL software is made open-sourced, publicly available, and organized in model repositories and stores (Model Zoo, ModelDepot), there comes a need to understand the relationships of these DL models regarding their maintenance and evolution tasks. Although clone analysis has been extensively studied for traditional software, up to the present, clone analysis has not been investigated for DL software. Since DL software adopts the data-driven development paradigm, it is still not clear whether and to what extent …
Privacy-Enhanced Remote Data Integrity Checking With Updatable Timestamp, Tong Wu, Guomin Yang, Yi Mu, Rongmao Chen, Shengmin Xu
Privacy-Enhanced Remote Data Integrity Checking With Updatable Timestamp, Tong Wu, Guomin Yang, Yi Mu, Rongmao Chen, Shengmin Xu
Research Collection School Of Computing and Information Systems
Remote data integrity checking (RDIC) enables clients to verify whether the outsourced data is intact without keeping a copy locally or downloading it. Nevertheless, the existing RDIC schemes do not support the pay-as-you-go (PAYG) payment model, where the payment is decided by the volume and duration of the outsourced data. Specifically, none of the existing works have considered the client’s control over changes in storage duration. In this paper, we propose an RDIC scheme to simultaneously check the data content and storage duration represented by an updatable timestamp via the third-party auditor (TPA). Also, our proposed scheme achieves indistinguishable privacy …
Is Using Deep Learning Frameworks Free?: Characterizing Technical Debt In Deep Learning Frameworks, Jiakun Liu, Qiao Huang, Xin Xia, Emad Shihab, David Lo, Shanping Li
Is Using Deep Learning Frameworks Free?: Characterizing Technical Debt In Deep Learning Frameworks, Jiakun Liu, Qiao Huang, Xin Xia, Emad Shihab, David Lo, Shanping Li
Research Collection School Of Computing and Information Systems
Developers of deep learning applications (shortened as application developers) commonly use deep learning frameworks in their projects. However, due to time pressure, market competition, and cost reduction, developers of deep learning frameworks (shortened as framework developers) often have to sacrifice software quality to satisfy a shorter completion time. This practice leads to technical debt in deep learning frameworks, which results in the increasing burden to both the application developers and the framework developers in future development.In this paper, we analyze the comments indicating technical debt (self-admitted technical debt) in 7 of the most popular open-source deep learning frameworks. Although framework …
Revisiting Supervised And Unsupervised Methods For Effort-Aware Cross-Project Defect Prediction, Chao Ni, Xin Xia, David Lo, Xiang Chen, Qing Gu
Revisiting Supervised And Unsupervised Methods For Effort-Aware Cross-Project Defect Prediction, Chao Ni, Xin Xia, David Lo, Xiang Chen, Qing Gu
Research Collection School Of Computing and Information Systems
Cross-project defect prediction (CPDP), aiming to apply defect prediction models built on source projects to a target project, has been an active research topic. A variety of supervised CPDP methods and some simple unsupervised CPDP methods have been proposed. In a recent study, Zhou et al. found that simple unsupervised CPDP methods (i.e., ManualDown and ManualUp) have a prediction performance comparable or even superior to complex supervised CPDP methods. Therefore, they suggested that the ManualDown should be treated as the baseline when considering non-effort-aware performance measures (NPMs) and the ManualUp should be treated as the baseline when considering effort-aware performance …
A Virtualization Based System Infrastructure For Dynamic Program Analysis, Jiaqi Hong
A Virtualization Based System Infrastructure For Dynamic Program Analysis, Jiaqi Hong
Dissertations and Theses Collection (Open Access)
Dynamic malware analysis schemes either run the target program as is in an isolated environment assisted by additional hardware facilities or modify it with instrumentation code statically or dynamically. The hardware-assisted schemes usually trap the target during its execution to a more privileged environment based on the available hardware events. The more privileged environment is not accessible by the untrusted kernel, thus this approach is often applied for transparent and secure kernel analysis. Nevertheless, the isolated environment induces a virtual address gap between the analyzer and the target, which hinders effective and efficient memory introspection and undermines the correctness of …
Cc2vec: Distributed Representations Of Code Changes, Thong Hoang, Hong Jin Kang, Julia Lawall, David Lo
Cc2vec: Distributed Representations Of Code Changes, Thong Hoang, Hong Jin Kang, Julia Lawall, David Lo
Research Collection School Of Computing and Information Systems
Existing work on software patches often use features specific to a single task. These works often rely on manually identified features, and human effort is required to identify these features for each task. In this work, we propose CC2Vec, a neural network model that learns a representation of code changes guided by their accompanying log messages, which represent the semantic intent of the code changes. CC2Vec models the hierarchical structure of a code change with the help of the attention mechanism and usesmultiple comparison functions to identify the differences between the removed and added code. To evaluate if CC2Vec can …
Mutation Testing Of Smart Contracts At Scale, Pieter Hartel, Richard Schumi
Mutation Testing Of Smart Contracts At Scale, Pieter Hartel, Richard Schumi
Research Collection School Of Computing and Information Systems
It is crucial that smart contracts are tested thoroughly due to their immutable nature. Even small bugs in smart contracts can lead to huge monetary losses. However, testing is not enough; it is also important to ensure the quality and completeness of the tests. There are already several approaches that tackle this challenge with mutation testing, but their effectiveness is questionable since they only considered small contract samples. Hence, we evaluate the quality of smart contract mutation testing at scale. We choose the most promising of the existing (smart contract specific) mutation operators, analyse their effectiveness in terms of killability …
A Machine Learning Approach For Vulnerability Curation, Yang Chen, Andrew E. Santosa, Ming Yi Ang, Abhishek Sharma, Asankhaya Sharma, David Lo
A Machine Learning Approach For Vulnerability Curation, Yang Chen, Andrew E. Santosa, Ming Yi Ang, Abhishek Sharma, Asankhaya Sharma, David Lo
Research Collection School Of Computing and Information Systems
Software composition analysis depends on database of open-source library vulerabilities, curated by security researchers using various sources, such as bug tracking systems, commits, and mailing lists. We report the design and implementation of a machine learning system to help the curation by by automatically predicting the vulnerability-relatedness of each data item. It supports a complete pipeline from data collection, model training and prediction, to the validation of new models before deployment. It is executed iteratively to generate better models as new input data become available. We use self-training to significantly and automatically increase the size of the training dataset, opportunistically …
Semantic Understanding Of Smart Contracts: Executable Operational Semantics Of Solidity, Jiao Jiao, Shuanglong Kan, Shang Wei Lin, David Sanan, Yang Liu, Jun Sun
Semantic Understanding Of Smart Contracts: Executable Operational Semantics Of Solidity, Jiao Jiao, Shuanglong Kan, Shang Wei Lin, David Sanan, Yang Liu, Jun Sun
Research Collection School Of Computing and Information Systems
Bitcoin has been a popular research topic recently. Ethereum (ETH), a second generation of cryptocurrency, extends Bitcoin's design by offering a Turing-complete programming language called Solidity to develop smart contracts. Smart contracts allow creditable execution of contracts on EVM (Ethereum Virtual Machine) without third parties. Developing correct and secure smart contracts is challenging due to the decentralized computation nature of the blockchain. Buggy smart contracts may lead to huge financial loss. Furthermore, smart contracts are very hard, if not impossible, to patch once they are deployed. Thus, there is a recent surge of interest in analyzing and verifying smart contracts. …
Chaff From The Wheat: Characterizing And Determining Valid Bug Reports, Yuanrui Fan, Xin Xia, David Lo, Ahmed E. Hassan
Chaff From The Wheat: Characterizing And Determining Valid Bug Reports, Yuanrui Fan, Xin Xia, David Lo, Ahmed E. Hassan
Research Collection School Of Computing and Information Systems
Developers use bug reports to triage and fix bugs. When triaging a bug report, developers must decide whether the bug report is valid (i.e., a real bug). A large amount of bug reports are submitted every day, with many of them end up being invalid reports. Manually determining valid bug report is a difficult and tedious task. Thus, an approach that can automatically analyze the validity of a bug report and determine whether a report is valid can help developers prioritize their triaging tasks and avoid wasting time and effort on invalid bug reports. In this study, motivated by the …