Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 4171 - 4200 of 9025

Full-Text Articles in Computer Sciences

Practitioners' Views On Good Software Testing Practices, Pavneet S. Kochhar, Xin Xia, David Lo May 2019

Practitioners' Views On Good Software Testing Practices, Pavneet S. Kochhar, Xin Xia, David Lo

Research Collection School Of Computing and Information Systems

Software testing is an integral part of software development process. Unfortunately, for many projects, bugs are prevalent despite testing effort, and testing continues to cost significant amount of time and resources. This brings forward the issue of test case quality and prompts us to investigate what make good test cases. To answer this important question, we interview 21 and survey 261 practitioners, who come from many small to large companies and open source projects distributed in 27 countries, to create and validate 29 hypotheses that describe characteristics of good test cases and testing practices. These characteristics span multiple dimensions including …


Patchnet: A Tool For Deep Patch Classification, Thong Hoang, Julia Lawall, Richard J. Oentaryo, Yuan Tian, David Lo May 2019

Patchnet: A Tool For Deep Patch Classification, Thong Hoang, Julia Lawall, Richard J. Oentaryo, Yuan Tian, David Lo

Research Collection School Of Computing and Information Systems

This work proposes PatchNet, an automated tool based on hierarchical deep learning for classifying patches by extracting features from commit messages and code changes. PatchNet contains a deep hierarchical structure that mirrors the hierarchical and sequential structure of a code change, differentiating it from the existing deep learning models on source code. PatchNet provides several options allowing users to selectparameters for the training process. The tool has been validated in the context of automatic identification of stable-relevant patches in the Linux kernel and is potentially applicable to automate other software engineering tasks that can be formulated as patch classification problems. …


How Practitioners Perceive Coding Proficiency, Xin Xia, Zhiyuan Wan, Pavneet S. Kochhar, David Lo May 2019

How Practitioners Perceive Coding Proficiency, Xin Xia, Zhiyuan Wan, Pavneet S. Kochhar, David Lo

Research Collection School Of Computing and Information Systems

Coding proficiency is essential to software practitioners. Unfortunately, our understanding on coding proficiency often translates to vague stereotypes, e.g., “able to write good code”. The lack of specificity hinders employers from measuring a software engineer’s coding proficiency, and software engineers from improving their coding proficiency skills. This raises an important question: what skills matter to improve one’s coding proficiency. To answer this question, we perform an empirical study by surveying 340 software practitioners from 33 countries across 5 continents. We first identify 38 coding proficiency skills grouped into nine categories by interviewing 15 developers from three companies. We then ask …


How To Derive Causal Insights For Digital Commerce In China? A Research Commentary On Computational Social Science Methods, David C.W. Phang, Kanliang Wang, Qiu-Hong Wang, Robert John Kauffman, Maurizio Naldi May 2019

How To Derive Causal Insights For Digital Commerce In China? A Research Commentary On Computational Social Science Methods, David C.W. Phang, Kanliang Wang, Qiu-Hong Wang, Robert John Kauffman, Maurizio Naldi

Research Collection School Of Computing and Information Systems

The transformation of empirical research due to the arrival of big data analytics and data science, as well as the new availability of methods that emphasize causal inference, are moving forward at full speed. In this Research Commentary, we examine the extent to which this has the potential to influence how e-commerce research is conducted. China offers the ultimate in data-at-scale settings, and the construction of real-world natural experiments. Chinese e-commerce includes some of the largest firms involved in e-commerce, mobile commerce, social media and social networks. This article was written to encourage young faculty and doctoral students to engage …


Adaptive Resonance Theory (Art) For Social Media Analytics, Lei Meng, Ah-Hwee Tan, Donald C. Ii Wunsch May 2019

Adaptive Resonance Theory (Art) For Social Media Analytics, Lei Meng, Ah-Hwee Tan, Donald C. Ii Wunsch

Research Collection School Of Computing and Information Systems

The last decade has witnessed how social media in the era of Web 2.0 reshapes the way people communicate, interact, and entertain in daily life and incubates the prosperity of various user-centric platforms, such as social networking, question answering, massive open online courses (MOOC), and e-commerce platforms. The available rich user-generated multimedia data on the web has evolved traditional ways of understanding multimedia research and has led to numerous emerging topics on human-centric analytics and services, such as user profiling, social network mining, crowd behavior analysis, and personalized recommendation. Clustering, as an important tool for mining information groups and in-group …


Deepjit: An End-To-End Deep Learning Framework For Just-In-Time Defect Prediction, Thong Hoang, Hoa Khanh Dam, Yasutaka Kamei, David Lo, Naoyasu Ubayashi May 2019

Deepjit: An End-To-End Deep Learning Framework For Just-In-Time Defect Prediction, Thong Hoang, Hoa Khanh Dam, Yasutaka Kamei, David Lo, Naoyasu Ubayashi

Research Collection School Of Computing and Information Systems

Software quality assurance efforts often focus on identifying defective code. To find likely defective code early, change-level defect prediction – aka. Just-In-Time (JIT) defect prediction – has been proposed. JIT defect prediction models identify likely defective changes and they are trained using machine learning techniques with the assumption that historical changes are similar to future ones. Most existing JIT defect prediction approaches make use of manually engineered features. Unlike those approaches, in this paper, we propose an end-to-end deep learning framework, named DeepJIT, that automatically extracts features from commit messages and code changes and use them to identify defects. Experiments …


On Reliability Of Patch Correctness Assessment, Xuan-Bach D. Le, Lingfeng Bao, David Lo, Xin Xia, Shanping Li, Corina S. Pasareanu May 2019

On Reliability Of Patch Correctness Assessment, Xuan-Bach D. Le, Lingfeng Bao, David Lo, Xin Xia, Shanping Li, Corina S. Pasareanu

Research Collection School Of Computing and Information Systems

Current state-of-the-art automatic software repair (ASR) techniques rely heavily on incomplete specifications, or test suites, to generate repairs. This, however, may cause ASR tools to generate repairs that are incorrect and hard to generalize. To assess patch correctness, researchers have been following two methods separately: (1) Automated annotation, wherein patches are automatically labeled by an independent test suite (ITS) – a patch passing the ITS is regarded as correct or generalizable, and incorrect otherwise, (2) Author annotation, wherein authors of ASR techniques manually annotate the correctness labels of patches generated by their and competing tools. While automated annotation cannot ascertain …


Emerging App Issue Identification From User Feedback: Experience On Wechat, Cuiyun Gao, Wujie Zheng, Yuetang Deng, David Lo, Jichuan Zeng, Michael R. Lyu, Irwin King May 2019

Emerging App Issue Identification From User Feedback: Experience On Wechat, Cuiyun Gao, Wujie Zheng, Yuetang Deng, David Lo, Jichuan Zeng, Michael R. Lyu, Irwin King

Research Collection School Of Computing and Information Systems

It is vital for popular mobile apps with large numbers of users to release updates with rich features while keeping stable user experience. Timely and accurately locating emerging app issues can greatly help developers to maintain and update apps. User feedback (i.e., user reviews) is a crucial channel between app developers and users, delivering a stream of information about bugs and features that concern users. Methods to identify emerging issues based on user feedback have been proposed in the literature, however, their applicability in industry has not been explored. We apply the recent method IDEA to WeChat, a popular messenger …


Patchnet: A Tool For Deep Patch Classification, Thong Hoang, Julia Lawall, Richard J. Oentaryo, Yuan Tian, David Lo May 2019

Patchnet: A Tool For Deep Patch Classification, Thong Hoang, Julia Lawall, Richard J. Oentaryo, Yuan Tian, David Lo

Research Collection School Of Computing and Information Systems

This work proposes PatchNet, an automated tool based on hierarchical deep learning for classifying patches by extracting features from commit messages and code changes. PatchNet contains a deep hierarchical structure that mirrors the hierarchical and sequential structure of a code change, differentiating it from the existing deep learning models on source code. PatchNet provides several options allowing users to select parameters for the training process. The tool has been validated in the context of automatic identification of stable-relevant patches in the Linux kernel and is potentially applicable to automate other software engineering tasks that can be formulated as patch classification …


Detect Rumors On Twitter By Promoting Information Campaigns With Generative Adversarial Learning, Jing Ma, Wei Gao, Kam-Fai Wong May 2019

Detect Rumors On Twitter By Promoting Information Campaigns With Generative Adversarial Learning, Jing Ma, Wei Gao, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

Rumors can cause devastating consequences to individual and/or society. Analysis shows that widespread of rumors typically results from deliberately promoted information campaigns which aim to shape collective opinions on the concerned news events. In this paper, we attempt to fight such chaos with itself to make automatic rumor detection more robust and effective. Our idea is inspired by adversarial learning method originated from Generative Adversarial Networks (GAN). We propose a GAN-style approach, where a generator is designed to produce uncertain or conflicting voices, complicating the original conversational threads in order to pressurize the discriminator to learn stronger rumor indicative representations …


Personalized Web Image Organization, Lei Meng, Ah-Hwee Tan, Donald C. Wunsch May 2019

Personalized Web Image Organization, Lei Meng, Ah-Hwee Tan, Donald C. Wunsch

Research Collection School Of Computing and Information Systems

Due to the problem of semantic gap, i.e. the visual content of an image may not represent its semantics well, existing efforts on web image organization usually transform this task to clustering the surrounding text. However, because the surrounding text is usually short and the words therein usually appear only once, existing text clustering algorithms can hardly use the statistical information for image representation and may achieve downgraded performance with higher computational cost caused by learning from noisy tags. This chapter presents using the Probabilistic ART with user preference architecture, as introduced in Sects. 3.5 and 3.4, for personalized web …


Online Multimodal Co-Indexing And Retrieval Of Social Media Data, Lei Meng, Ah-Hwee Tan, Donald C. Wunsch May 2019

Online Multimodal Co-Indexing And Retrieval Of Social Media Data, Lei Meng, Ah-Hwee Tan, Donald C. Wunsch

Research Collection School Of Computing and Information Systems

Effective indexing of social media data is key to searching for information on the social Web. However, the characteristics of social media data make it a challenging task. The large-scale and streaming nature is the first challenge, which requires the indexing algorithm to be able to efficiently update the indexing structure when receiving data streams. The second challenge is utilizing the rich meta-information of social media data for a better evaluation of the similarity between data objects and for a more semantically meaningful indexing of the data, which may allow the users to search for them using the different types …


Towards Zero Knowledge Learning For Cross Language Api Mappings, Duy Quoc Nghi Bui May 2019

Towards Zero Knowledge Learning For Cross Language Api Mappings, Duy Quoc Nghi Bui

Research Collection School Of Computing and Information Systems

Programmers often need to migrate programs from one language or platform to another in order to implement functionality, instead of rewriting the code from scratch. However, most techniques proposed to identify API mappings across languages and facilitate automated program translation require manually curated parallel corpora that contain already mapped API seeds or functionally-equivalent code using the APIs in two different languages so that the techniques can have an anchor to map APIs. To alleviate the need of curating parallel data and to generalize the applicability of program translation techniques, we develop a new automated approach for identifying API mappings across …


Faster: Fusion Analytics For Public Transport Event Response Industrial Applications Track, Sebastien Blandin, Laura Wynter, Hasan Poonawala, Sean Laguna, Basile Dura May 2019

Faster: Fusion Analytics For Public Transport Event Response Industrial Applications Track, Sebastien Blandin, Laura Wynter, Hasan Poonawala, Sean Laguna, Basile Dura

Research Collection School Of Computing and Information Systems

The Autonomous Agents and Multiagent Systems (AAMAS) conference series gathers researchers from around the world to share the latest advances in the field. It is the premier forum for research in the theory and practice of autonomous agents and multiagent systems. AAMAS 2002, the first of the series, was held in Bologna, followed by Melbourne (2003), New York (2004), Utrecht (2005), Hakodate (2006), Honolulu (2007), Estoril (2008), Budapest (2009), Toronto (2010), Taipei (2011), Valencia (2012), Saint Paul (2013), Paris (2014), Istanbul (2015), Singapore (2016), São Paulo (2017) and Stockholm (2018). This volume is the proceedings of AAMAS 2019, the 18th …


Organizing For Artificial Intelligence (Ai) Technologies, Sukti Ghosh May 2019

Organizing For Artificial Intelligence (Ai) Technologies, Sukti Ghosh

Research Collection Lee Kong Chian School Of Business

This study focuses on organisation design choices as tools for addressing the management challenges of commercialising AI technologies for competitive advantage. It explores how design choices address fundamental problems of organising in such context, illustrating notable design features observed. Additionally, it examines external alignment and internal coherence reiterating interdependencies in organisation’s design choices, when adapting to exogenous changes due to emerging AI technologies.


Learning Two-Layer Neural Networks With Symmetric Inputs, Rong Ge, Rohith Kuditipudi, Zhize Li, Xiang Wang May 2019

Learning Two-Layer Neural Networks With Symmetric Inputs, Rong Ge, Rohith Kuditipudi, Zhize Li, Xiang Wang

Research Collection School Of Computing and Information Systems

We give a new algorithm for learning a two-layer neural network under a very general class of input distributions. Assuming there is a ground-truth two-layer network $y = A \sigma(Wx) + \xi$, where A, W are weight matrices, $\xi$ represents noise, and the number of neurons in the hidden layer is no larger than the input or output, our algorithm is guaranteed to recover the parameters A, W of the ground-truth network. The only requirement on the input x is that it is symmetric, which still allows highly complicated and structured input. Our algorithm is based on the method-of-moments framework …


Witt: Querying Technology Terms Based On Automated Classification, Mathieu Nassif, Christoph Treude, Martin P. Robillard May 2019

Witt: Querying Technology Terms Based On Automated Classification, Mathieu Nassif, Christoph Treude, Martin P. Robillard

Research Collection School Of Computing and Information Systems

Witt is a tool that systematically and automatically categorizes software technologies using original information extraction algorithms applied to Stack Overflow and Wikipedia. Witt takes as input a term, such as "django", and returns one or more categories that describe it (e.g., "framework"), along with attributes that further qualify it (e.g., "web-application"). Our comparative evaluation of Witt against six independent taxonomy tools showed that, when applied to software terms, Witt has better coverage than alternative solutions, without a corresponding degradation in the number of spurious results. The information extracted by Witt is available through the Witt Web Application, which allows users …


9.6 Million Links In Source Code Comments: Purpose, Evolution, And Decay, Hideaki Hata, Christoph Treude, Raula Gaikovina Kula, Takashi Ishio May 2019

9.6 Million Links In Source Code Comments: Purpose, Evolution, And Decay, Hideaki Hata, Christoph Treude, Raula Gaikovina Kula, Takashi Ishio

Research Collection School Of Computing and Information Systems

Links are an essential feature of the World Wide Web, and source code repositories are no exception. However, despite their many undisputed benefits, links can suffer from decay, insufficient versioning, and lack of bidirectional traceability. In this paper, we investigate the role of links contained in source code comments from these perspectives. We conducted a large-scale study of around 9.6 million links to establish their prevalence, and we used a mixed-methods approach to identify the links' targets, purposes, decay, and evolutionary aspects. We found that links are prevalent in source code repositories, that licenses, software homepages, and specifications are common …


Automatically Generating Documentation For Lambda Expressions In Java, Anwar Alqaimi, Patanamon Thongtanunam, Christoph Treude May 2019

Automatically Generating Documentation For Lambda Expressions In Java, Anwar Alqaimi, Patanamon Thongtanunam, Christoph Treude

Research Collection School Of Computing and Information Systems

When lambda expressions were introduced to the Java programming language as part of the release of Java 8 in 2014, they were the language’s first step into functional programming. Since lambda expressions are still relatively new, not all developers use or understand them. In this paper, we first present the results of an empirical study to determine how frequently developers of GitHub repositories make use of lambda expressions and how they are documented. We find that 11% of Java GitHub repositories use lambda expressions, and that only 6% of the lambda expressions are accompanied by source code comments. We then …


Predicting Good Configurations For Github And Stack Overflow Topic Models, Christoph Treude, Markus Wagner May 2019

Predicting Good Configurations For Github And Stack Overflow Topic Models, Christoph Treude, Markus Wagner

Research Collection School Of Computing and Information Systems

Software repositories contain large amounts of textual data, ranging from source code comments and issue descriptions to questions, answers, and comments on Stack Overflow. To make sense of this textual data, topic modelling is frequently used as a text-mining tool for the discovery of hidden semantic structures in text bodies. Latent Dirichlet allocation (LDA) is a commonly used topic model that aims to explain the structure of a corpus by grouping texts. LDA requires multiple parameters to work well, and there are only rough and sometimes conflicting guidelines available on how these parameters should be set. In this paper, we …


Sotorrent: Studying The Origin, Evolution, And Usage Of Stack Overflow Code Snippets, Sebastian Baltes, Christoph Treude, Stephan Diehl May 2019

Sotorrent: Studying The Origin, Evolution, And Usage Of Stack Overflow Code Snippets, Sebastian Baltes, Christoph Treude, Stephan Diehl

Research Collection School Of Computing and Information Systems

Stack Overflow (SO) is the most popular questionand-answer website for software developers, providing a large amount of copyable code snippets. Like other software artifacts, code on SO evolves over time, for example when bugs are fixed or APIs are updated to the most recent version. To be able to analyze how code and the surrounding text on SO evolves, we built SOTorrent, an open dataset based on the official SO data dump. SOTorrent provides access to the version history of SO content at the level of whole posts and individual text and code blocks. It connects code snippets from SO …


Neural Multimodal Belief Tracker With Adaptive Attention For Dialogue Systems, Zheng Zhang, Lizi Liao, Minlie Huang, Xiaoyan Zhu, Tat-Seng Chua May 2019

Neural Multimodal Belief Tracker With Adaptive Attention For Dialogue Systems, Zheng Zhang, Lizi Liao, Minlie Huang, Xiaoyan Zhu, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Multimodal dialogue systems are attracting increasing attention with a more natural and informative way for human-computer interaction. As one of its core components, the belief tracker estimates the user's goal at each step of the dialogue and provides a direct way to validate the ability of dialogue understanding. However, existing studies on belief trackers are largely limited to textual modality, which cannot be easily extended to capture the rich semantics in multimodal systems such as those with product images. For example, in fashion domain, the visual appearance of clothes play a crucial role in understanding the user's intention. In this …


Cure: Flexible Categorical Data Representation By Hierarchical Coupling Learning, Songlei Jian, Guansong Pang, Longbing Cao, Kai Lu, Hang Gao May 2019

Cure: Flexible Categorical Data Representation By Hierarchical Coupling Learning, Songlei Jian, Guansong Pang, Longbing Cao, Kai Lu, Hang Gao

Research Collection School Of Computing and Information Systems

The representation of categorical data with hierarchical value coupling relationships (i.e., various value-to-value cluster interactions) is very critical yet challenging for capturing complex data characteristics in learning tasks. This paper proposes a novel and flexible coupled unsupervised categorical data representation (CURE) framework, which not only captures the hierarchical couplings but is also flexible enough to be instantiated for contrastive learning tasks. CURE first learns the value clusters of different granularities based on multiple value coupling functions and then learns the value representation from the couplings between the obtained value clusters. With two complementary value coupling functions, CURE is instantiated into …


Peerlens: Peer-Inspired Interactive Learning Path Planning In Online Question Pool, Meng Xia, Mingfei Sun, Huan Wei, Qing Chen, Yong Wang, Lei Shi, Huamin Qu, Xiaojuan Ma May 2019

Peerlens: Peer-Inspired Interactive Learning Path Planning In Online Question Pool, Meng Xia, Mingfei Sun, Huan Wei, Qing Chen, Yong Wang, Lei Shi, Huamin Qu, Xiaojuan Ma

Research Collection School Of Computing and Information Systems

Online question pools like LeetCode provide hands-on exercises of skills and knowledge. However, due to the large volume of questions and the intent of hiding the tested knowledge behind them, many users find it hard to decide where to start or how to proceed based on their goals and performance. To overcome these limitations, we present PeerLens, an interactive visual analysis system that enables peer-inspired learning path planning. PeerLens can recommend a customized, adaptable sequence of practice questions to individual learners, based on the exercise history of other users in a similar learning scenario. We propose a new way to …


Beyond Autonomy: The Self And Life Of Social Agents, Budhitama Subagdja, Ah-Hwee Tan May 2019

Beyond Autonomy: The Self And Life Of Social Agents, Budhitama Subagdja, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Agents have gained popularity nowadays as virtual assistants and companions of their human users supporting daily activities in many aspects of personal life. Designed to be sociable, an agent engages its user(s) to communicate and even develop friendships. Rather than just as a lifeless toy, it is supposed to be perceived as an individual with its own personality, experiences, and social life. In this paper, we seek to highlight self-hood as another dimension that characterizes an agent. Besides levels of autonomy and reasoning, an agent can be defined based on its capacity to process and reflect on its own self …


The Challenges Of Creating Engaging Content: Results From A Focus Group Study Of A Popular News Media Organization, Kholoud Khalil Aldous, Jisun An, Bernard J. Jansen May 2019

The Challenges Of Creating Engaging Content: Results From A Focus Group Study Of A Popular News Media Organization, Kholoud Khalil Aldous, Jisun An, Bernard J. Jansen

Research Collection School Of Computing and Information Systems

The process of content creation for distribution via social media platforms is not a trivial one for social media editors as the goal of creating both serious and engaging content is challenging, with no clear or differing guidelines or rules across and between platforms. For creators of serious content, such as news organizations, advertisers, or educational institutions, engagement has a deeper meaning beyond likes, shares, etc. that is aimed at the audience actually processing the underlying content associated with a social media post. In this research, we report findings from a group study that aimed to understand the process and …


Interaction-Aware Arrangement For Event-Based Social Networks, Feifei Kou, Zimu Zhou, Hao Cheg, Junping Du, Yexuan Shi, Pan Xu Apr 2019

Interaction-Aware Arrangement For Event-Based Social Networks, Feifei Kou, Zimu Zhou, Hao Cheg, Junping Du, Yexuan Shi, Pan Xu

Research Collection School Of Computing and Information Systems

No abstract provided.


Artificial Intelligence, Real Concerns…And Cash, Singapore Management University Apr 2019

Artificial Intelligence, Real Concerns…And Cash, Singapore Management University

Perspectives@SMU

Regulating development of self-aware robots is crucial. Data privacy is key to user-app power dynamic


Big Data And The Consumer, Seema Chokshi Apr 2019

Big Data And The Consumer, Seema Chokshi

MITB Thought Leadership Series

What is big data? The intuitive meaning of the phrase ‘big data’ might be “data that is huge in quantity”. But is that interpretation enough? Data of this type has existed for as long as humans have made records of their work. Some of the earliest writings, such as cuneiform, contain vast amounts of data covering areas as diverse as law, mapping and mathematical equations.


Question Answering With Textual Sequence Matching, Shuohang Wang Apr 2019

Question Answering With Textual Sequence Matching, Shuohang Wang

Dissertations and Theses Collection (Open Access)

Question answering (QA) is one of the most important applications in natural language processing. With the explosive text data from the Internet, intelligently getting answers of questions will help humans more efficiently collect useful information. My research in this thesis mainly focuses on solving question answering problem with textual sequence matching model which is to build vectorized representations for pairs of text sequences to enable better reasoning. And our thesis consists of three major parts.

In Part I, we propose two general models for building vectorized representations over a pair of sentences, which can be directly used to solve the …