Combining Query Reduction And Expansion For Text-Retrieval-Based Bug Localization,
2021
Singapore Management University
Combining Query Reduction And Expansion For Text-Retrieval-Based Bug Localization, Juan Manuel Florez, Oscar Chaparro, Christoph Treude, Andrian Marcus
Research Collection School Of Computing and Information Systems
Automated text-retrieval-based bug localization (TRBL) techniques normally use the full text of a bug report to formulate a query and retrieve parts of the code that are buggy. Previous research has shown that reducing the size of the query increases the effectiveness of TRBL. On the other hand, researchers also found improvements when expanding the query (i.e., adding more terms). In this paper, we bring these two views together to reformulate queries for TRBL. Specifically, we improve discourse-based query reduction strategies, by adopting a combinatorial approach and using task phrases from bug reports, and combine them with a state-of-the-art query …
Csci 40500/77100: Software Engineering,
2021
CUNY Hunter College
Csci 40500/77100: Software Engineering, Raffi Khatchadourian
Open Educational Resources
This course is intended to be an introductory survey on the fundamental concepts and principles that underlie current and emerging methods, tools, and techniques for the efficient engineering of high-quality software systems. This may include understanding and appreciating problems in large-scale software development such as functional analysis of information processing systems, system design concepts, timing estimates, documentation, and system testing.
Csci 40500/77100: Software Engineering,
2021
CUNY Hunter College
Csci 40500/77100: Software Engineering, Raffi Khatchadourian
Open Educational Resources
This course is intended to be an introductory survey on the fundamental concepts and principles that underlie current and emerging methods, tools, and techniques for the efficient engineering of high-quality software systems. This may include understanding and appreciating problems in large-scale software development such as functional analysis of information processing systems, system design concepts, timing estimates, documentation, and system testing.
Treecaps: Tree-Based Capsule Networks For Source Code Processing,
2021
Singapore Management University
Treecaps: Tree-Based Capsule Networks For Source Code Processing, Duy Quoc Nghi Bui, Yijun Yu, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Recently program learning techniques have been proposed to process source code based on syntactical structures (e.g., Abstract Syntax Trees) and/or semantic information (e.g., Dependency Graphs). While graphs may be better at capturing various viewpoints of code semantics than trees, constructing graph inputs from code need static code semantic analysis that may not be accurate and introduces noise during learning. On the other hand, syntax trees are precisely defined according to the language grammar and easier to construct and process than graphs. We propose a new tree-based learning technique, named TreeCaps, by fusing capsule networks with tree-based convolutional neural networks, to …
Revman: Revenue-Aware Multi-Task Online Insurance Recommendation,
2021
Jilin University
Revman: Revenue-Aware Multi-Task Online Insurance Recommendation, Yu Li, Yi Zhang, Lu Gan, Gengwei Hong, Zimu Zhou, Qiang Li
Research Collection School Of Computing and Information Systems
Online insurance is a new type of e-commerce with exponential growth. An effective recommendation model that maximizes the total revenue of insurance products listed in multiple customized sales scenarios is crucial for the success of online insurance business. Prior recommendation models are ineffective because they fail to characterize the complex relatedness of insurance products in multiple sales scenarios and maximize the overall conversion rate rather than the total revenue. Even worse, it is impractical to collect training data online for total revenue maximization due to the business logic of online insurance. We propose RevMan, a Revenue-aware Multi-task Network for online …
An Exploratory Study On The Introduction And Removal Of Different Types Of Technical Debt In Deep Learning Frameworks,
2021
Singapore Management University
An Exploratory Study On The Introduction And Removal Of Different Types Of Technical Debt In Deep Learning Frameworks, Jiakun Liu, Qiao Huang, Xin Xia, Emad Shihab, David Lo, Shanping Li
Research Collection School Of Computing and Information Systems
To complete tasks faster, developers often have to sacrifice the quality of the software. Such compromised practice results in the increasing burden to developers in future development. The metaphor, technical debt, describes such practice. Prior research has illustrated the negative impact of technical debt, and many researchers investigated how developers deal with a certain type of technical debt. However, few studies focused on the removal of different types of technical debt in practice. To fill this gap, we use the introduction and removal of different types of self-admitted technical debt (i.e., SATD) in 7 deep learning frameworks as an example. …
Fault Analysis And Debugging Of Microservice Systems: Industrial Survey, Benchmark System, And Empirical Study,
2021
Fudan University
Fault Analysis And Debugging Of Microservice Systems: Industrial Survey, Benchmark System, And Empirical Study, Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Wenhai Li, Dan Ding
Research Collection School Of Computing and Information Systems
The complexity and dynamism of microservice systems pose unique challenges to a variety of software engineering tasks such as fault analysis and debugging. In spite of the prevalence and importance of microservices in industry, there is limited research on the fault analysis and debugging of microservice systems. To fill this gap, we conduct an industrial survey to learn typical faults of microservice systems, current practice of debugging, and the challenges faced by developers in practice. We then develop a medium-size benchmark microservice system (being the largest and most complex open source microservice system within our knowledge) and replicate 22 industrial …
Understanding Adversarial Robustness Via Critical Attacking Route,
2021
Beijing University of Aeronautics and Astronautics (Beihang University)
Understanding Adversarial Robustness Via Critical Attacking Route, Tianlin Li, Aishan Liu, Xianglong Liu, Yitao Xu, Chongzhi Zhang, Xiaofei Xie
Research Collection School Of Computing and Information Systems
Deep neural networks (DNNs) are vulnerable to adversarial examples which are generated by inputs with imperceptible perturbations. Understanding adversarial robustness of DNNs has become an important issue, which would for certain result in better practical deep learning applications. To address this issue, we try to explain adversarial robustness for deep models from a new perspective of critical attacking route, which is computed by a gradient-based influence propagation strategy. Similar to rumor spreading in social net-works, we believe that adversarial noises are amplified and propagated through the critical attacking route. By exploiting neurons' influences layer by layer, we compose the critical …
Decision-Guided Weighted Automata Extraction From Recurrent Neural Networks,
2021
Singapore Management University
Decision-Guided Weighted Automata Extraction From Recurrent Neural Networks, Xiyue Zhang, Xiaoning Du, Xiaofei Xie, Lei Ma, Yang Liu, Meng Sun
Research Collection School Of Computing and Information Systems
Recurrent Neural Networks (RNNs) have demonstrated their effectiveness in learning and processing sequential data (e.g., speech and natural language). However, due to the black-box nature of neural networks, understanding the decision logic of RNNs is quite challenging. Some recent progress has been made to approximate the behavior of an RNN by weighted automata. They provide better interpretability, but still suffer from poor scalability. In this paper, we propose a novel approach to extracting weighted automata with the guidance of a target RNN’s decision and context information. In particular, we identify the patterns of RNN’s step-wise predictive decisions to instruct the …
Visual Analysis Of Discrimination In Machine Learning,
2021
Hong Kong University of Science and Technology
Visual Analysis Of Discrimination In Machine Learning, Qianwen Wang, Zhenghua Xu, Zhutian Chen, Yong Wang, Shixia Liu, Huamin Qu
Research Collection School Of Computing and Information Systems
The growing use of automated decision-making in critical applications, such as crime prediction and college admission, has raised questions about fairness in machine learning. How can we decide whether different treatments are reasonable or discriminatory? In this paper, we investigate discrimination in machine learning from a visual analytics perspective and propose an interactive visualization tool, DiscriLens, to support a more comprehensive analysis. To reveal detailed information on algorithmic discrimination, DiscriLens identifies a collection of potentially discriminatory itemsets based on causal modeling and classification rules mining. By combining an extended Euler diagram with a matrix-based visualization, we develop a novel set …
Efficientderain: Learning Pixel-Wise Dilation Filtering For High-Efficiency Single-Image Deraining,
2021
Nanyang Technological University
Efficientderain: Learning Pixel-Wise Dilation Filtering For High-Efficiency Single-Image Deraining, Qing Guo, Jingyang Sun, Felix Juefei-Xu, Lei Ma, Xiaofei Xie, Wei Feng, Yang Liu, Jianjun Zhao
Research Collection School Of Computing and Information Systems
Single-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however, significantly affects these methods’ efficiency and effectiveness for many efficiency-critical applications. To fill this gap, in this paper, we regard the single-image deraining as a general image-enhancing problem and originally propose a model-free deraining method, i.e., EfficientDeRain, which is able to process a rainy image within 10 ms (i.e., around 6 ms on average), over 80 times …
Terrace-Based Food Counting And Segmentation,
2021
Singapore Management University
Terrace-Based Food Counting And Segmentation, Huu-Thanh Nguyen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
This paper represents object instance as a terrace, where the height of terrace corresponds to object attention while the evolution of layers from peak to sea level represents the complexity in drawing the finer boundary of an object. A multitask neural network is presented to learn the terrace representation. The attention of terrace is leveraged for instance counting, and the layers provide prior for easy-to-hard pathway of progressive instance segmentation. We study the model for counting and segmentation for a variety of food instances, ranging from Chinese, Japanese to Western food. This paper presents how the terrace model deals with …
Qlens: Visual Analytics Of Multi-Step Problem-Solving Behaviors For Improving Question Design,
2021
Hong Kong University of Science and Technology
Qlens: Visual Analytics Of Multi-Step Problem-Solving Behaviors For Improving Question Design, Meng Xia, Reshika P. Velumani, Yong Wang, Huamin Qu, Xiaojuan Ma
Research Collection School Of Computing and Information Systems
With the rapid development of online education in recent years, there has been an increasing number of learning platforms that provide students with multi-step questions to cultivate their problem-solving skills. To guarantee the high quality of such learning materials, question designers need to inspect how students’ problem-solving processes unfold step by step to infer whether students’ problem-solving logic matches their design intent. They also need to compare the behaviors of different groups (e.g., students from different grades) to distribute questions to students with the right level of knowledge. The availability of fine-grained interaction data, such as mouse movement trajectories from …
Change Request Prediction And Effort Estimation In An Evolving Software System,
2021
University of Denver
Change Request Prediction And Effort Estimation In An Evolving Software System, Lamees Abdullah Alhazzaa
Electronic Theses and Dissertations
Prediction of software defects has been the focus of many researchers in empirical software engineering and software maintenance because of its significance in providing quality estimates from the project management perspective for an evolving legacy system. Software Reliability Growth Models (SRGM) have been used to predict future defects in a software release. Modern software engineering databases contain Change Requests (CR), which include both defects and other maintenance requests. Our goal is to use defect prediction methods to help predict CRs in an evolving legacy system.
Limited research has been done in defect prediction using curve-fitting methods evolving software systems, with …
Text Classification Using Novel Term Weighting Scheme-Based Improved Tf-Idf For Internet Media Reports,
2021
Beijing University of Chemical Technology
Text Classification Using Novel Term Weighting Scheme-Based Improved Tf-Idf For Internet Media Reports, Zhiying Jiang Phd, Bo Gao, Yanlin He, Yongming Han, Paul Doyle, Qunxiong Zhu
Other
With the rapid development of the internet technology, a large amount of internet text data can be obtained. The text classification (TC) technology plays a very important role in processing massive text data, but the accuracy of classification is directly affected by the performance of term weighting in TC. Due to the original design of information retrieval (IR), term frequency-inverse document frequency (TF-IDF) is not effective enough for TC, especially for processing text data with unbalanced distributions in internet media reports. Therefore, the variance between the DF value of a particular term and the average of all DFs , namely, …
Tagtools: Software Supporting Biologging Research,
2021
Calvin University
Tagtools: Software Supporting Biologging Research, Sam Fynewever, Racheal Tejevbo, Stacy L. De Ruiter
Summer Research
Observing animals in obscure habitats has long posed a challenge in biology. Biologging addresses this broad problem, obtaining data about animals like their acceleration, GPS position, etc. using electronic tags. Data from tags can lead to important knowledge of animal behavior, as DeRuiter et al. (2013) have shown1 . However, data is often shared raw, causing inconsistencies. This project maintains TagTools, software addressing this specific problem. TagToolsfunctions (in Matlab, Octave & R) calibrate data and help analyze behavior. TagTools software is traditionally taught at in-person workshops. Our old documents (practicals) were hosted online, but had bugs & were very long, …
An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags,
2021
Bowling Green State University
An Ensemble Approach For Annotating Source Code Identifiers With Part-Of-Speech Tags, Christian D. Newman,, Michael J. Decker, Reem S. Alsuhaibani, Anthony Peruma, Mohamed Wiem Mkaouer, Satyajit Mohapatra, Tejal Vishnoi, Marcos Zampieri, Timothy Sheldon, Emily Hill
Articles
This paper presents an ensemble part-of-speech tagging approach for source code identifiers. Ensemble tagging is a technique that uses machine-learning and the output from multiple part-of-speech taggers to annotate natural language text at a higher quality than the part-of-speech taggers are able to obtain independently. Our ensemble uses three state-of-the-art part-of-speech taggers: SWUM, POSSE, and Stanford. We study the quality of the ensemble's annotations on five different types of identifier names: function, class, attribute, parameter, and declaration statement at the level of both individual words and full identifier names. We also study and discuss the weaknesses of our tagger to …
Applications Of Machine Learning To Facilitate Software Engineering And Scientific Computing,
2021
Chapman University
Applications Of Machine Learning To Facilitate Software Engineering And Scientific Computing, Natalie Best
Computational and Data Sciences (PhD) Dissertations
The use of machine learning has risen in recent years, though many areas remain unexplored due to lack of data or lack of computational tools. This dissertation explores machine learning approaches in case studies involving image classification and natural language processing. In addition, a software library in the form of two-way bridge connecting deep learning models in Keras with ones available in the Fortran programming language is also presented.
In Chapter 2, we explore the applicability of transfer learning utilizing models pre-trained on non-software engineering data applied to the problem of classifying software unified modeling language diagrams where data is …
Soda: An Open-Source Library For Visualizing Biological Sequence Annotation,
2021
The University Of Montana
Soda: An Open-Source Library For Visualizing Biological Sequence Annotation, Jack W. Roddy, Travis J. Wheeler
Graduate Student Theses, Dissertations, & Professional Papers
Genome annotation is the process of identifying and labeling known genetic sequences or features within a genome. Across the various subfields within modern molecular biology, there is a common need for the visualization of such annotations. Genomic data is often visualized on web browser platforms, providing users with easy access to visualization tools without the need for installing any software or, in many cases, underlying datasets. While there exists a broad range of web-based visualization tools, there is, to my knowledge, no lightweight, modern library tailored towards the visualization of genomic data. Instead, developers charged with the task of producing …
Privattnet: Predicting Privacy Risks In Images Using Visual Attention,
2021
Institute for High Performance Computing
Privattnet: Predicting Privacy Risks In Images Using Visual Attention, Zhang Chen, Thivya Kandappu, Vigneshwaran Subbaraju
Research Collection School Of Computing and Information Systems
Visual privacy concerns associated with image sharing is a critical issue that need to be addressed to enable safe and lawful use of online social platforms. Users of social media platforms often suffer from no guidance in sharing sensitive images in public, and often face with social and legal consequences. Given the recent success of visual attention based deep learning methods in measuring abstract phenomena like image memorability, we are motivated to investigate whether visual attention based methods could be useful in measuring psychophysical phenomena like “privacy sensitivity”. In this paper we propose PrivAttNet – a visual attention based approach, …
