An Empirical Study Of The Discreteness Prior In Low-Rank Matrix Completion,
2021
Singapore Management University
An Empirical Study Of The Discreteness Prior In Low-Rank Matrix Completion, Rodrigo Alves, Antoine Ledent, Renato Assunção, Marius And Kloft
Research Collection School Of Computing and Information Systems
A reasonable assumption in recommender systems is that the rows (users) and columns (items) of the rating matrix can be split into groups (communities) with the following property: each entry of the matrix is the sum of components corresponding to community behavior and a purely low-rank component corresponding to individual behavior. We investigate (1) whether such a structure is present in real-world datasets, (2) whether the knowledge of the existence of such structure alone can improve performance, without explicit information about the community memberships. To these ends, we formulate a joint optimization problem over all (completed matrix, set of communities) …
A Survey On Ml4vis: Applying Machine Learning Advances To Data Visualization,
2021
Singapore Management University
A Survey On Ml4vis: Applying Machine Learning Advances To Data Visualization, Qianwen Wang, Zhutian Chen, Yong Wang, Huamin Qu
Research Collection School Of Computing and Information Systems
Inspired by the great success of machine learning (ML), researchers have applied ML techniques to visualizations to achieve a better design, development, and evaluation of visualizations. This branch of studies, known as ML4VIS, is gaining increasing research attention in recent years. To successfully adapt ML techniques for visualizations, a structured understanding of the integration of ML4VIS is needed. In this article, we systematically survey 88 ML4VIS studies, aiming to answer two motivating questions: “what visualization processes can be assisted by ML?” and “how ML techniques can be used to solve visualization problems? ” This survey reveals seven main processes where …
Vehicle Routing: Review Of Benchmark Datasets,
2021
Singapore Management University
Vehicle Routing: Review Of Benchmark Datasets, Aldy Gunawan, Graham Kendall, Barry Mccollum, Hsin-Vonn Seow, Lai Soon Lee
Research Collection School Of Computing and Information Systems
The Vehicle Routing Problem (VRP) was formally presented to the scientific literature in 1959 by Dantzig and Ramser (DOI:10.1287/mnsc.6.1.80). Sixty years on, the problem is still heavily researched, with hundreds of papers having been published addressing this problem and the many variants that now exist. Many datasets have been proposed to enable researchers to compare their algorithms using the same problem instances where either the best known solution is known or, in some cases, the optimal solution is known. In this survey paper, we provide a list of Vehicle Routing Problem datasets, categorized to enable researchers to have easy access …
Effective Digital Learning Practices For Is Design Courses During Covid-19,
2021
Singapore Management University
Effective Digital Learning Practices For Is Design Courses During Covid-19, Eng Lieh Ouh, Benjamin Gan
Research Collection School Of Computing and Information Systems
The COVID-19 pandemic has pushed educational institutions to adopt digital learning for an extended period. This research studies the effectiveness of digital learning practices based on student feedback data collected for two Information Systems design courses: human interaction design and solution architecture design. This paper leverages the data to analyze the effectiveness of a set of digital learning practices: ZOOM lectures, polling or Kahoot questions, self-reflection, virtual exercises and virtual mentorship. Our research questions are on the effectiveness of these learning practices to keep the student’s interest and learn the course materials. The research compares each learning practice and the …
Modeling Transitions Of Focal Entities For Conversational Knowledge Base Question Answering,
2021
Singapore Management University
Modeling Transitions Of Focal Entities For Conversational Knowledge Base Question Answering, Yunshi Lan, Jing Jiang
Research Collection School Of Computing and Information Systems
Conversational KBQA is about answering a sequence of questions related to a KB. Follow-up questions in conversational KBQA often have missing information referring to entities from the conversation history. In this paper, we propose to model these implied entities, which we refer to as the focal entities of the conversation. We propose a novel graph-based model to capture the transitions of focal entities and apply a graph neural network to derive a probability distribution of focal entities for each question, which is then combined with a standard KBQA module to perform answer ranking. Our experiments on two datasets demonstrate the …
Thunderrw: An In-Memory Graph Random Walk Engine,
2021
Singapore Management University
Thunderrw: An In-Memory Graph Random Walk Engine, Shixuan Sun, Yuhang Chen, Shengliang Lu, Bingsheng He, Yuchen Li
Research Collection School Of Computing and Information Systems
As random walk is a powerful tool in many graph processing, mining and learning applications, this paper proposes an efficient inmemory random walk engine named ThunderRW. Compared with existing parallel systems on improving the performance of a single graph operation, ThunderRW supports massive parallel random walks. The core design of ThunderRW is motivated by our profiling results: common RW algorithms have as high as 73.1% CPU pipeline slots stalled due to irregular memory access, which suffers significantly more memory stalls than the conventional graph workloads such as BFS and SSSP. To improve the memory efficiency, we first design a generic …
Context-Aware Outstanding Fact Mining From Knowledge Graphs,
2021
Singapore Management University
Context-Aware Outstanding Fact Mining From Knowledge Graphs, Yueji Yang, Yuchen Li, Panagiotis Karras, Anthony Tung
Research Collection School Of Computing and Information Systems
An Outstanding Fact (OF) is an attribute that makes a target entity stand out from its peers. The mining of OFs has important applications, especially in Computational Journalism, such as news promotion, fact-checking, and news story finding. However, existing approaches to OF mining: (i) disregard the context in which the target entity appears, hence may report facts irrelevant to that context; and (ii) require relational data, which are often unavailable or incomplete in many application domains. In this paper, we introduce the novel problem of mining Contextaware Outstanding Facts (COFs) for a target entity under a given context specified by …
Are Missing Links Predictable? An Inferential Benchmark For Knowledge Graph Completion,
2021
Singapore Management University
Are Missing Links Predictable? An Inferential Benchmark For Knowledge Graph Completion, Yixin Cao, Xiang Ji, Xin Lv, Juanzi Li, Yonggang Wen, Hanwang Zhang
Research Collection School Of Computing and Information Systems
We present InferWiki, a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns. First, each testing sample is predictable with supportive data in the training set. To ensure it, we propose to utilize rule-guided train/test generation, instead of conventional random split. Second, InferWiki initiates the evaluation following the open-world assumption and improves the inferential difficulty of the closed-world assumption, by providing manually annotated negative and unknown triples. Third, we include various inference patterns (e.g., reasoning path length and types) for comprehensive evaluation. In experiments, we curate two settings of InferWiki varying in sizes …
The 4th Workshop On Heterogeneous Information Network Analysis And Applications (Hena 2021),
2021
Singapore Management University
The 4th Workshop On Heterogeneous Information Network Analysis And Applications (Hena 2021), Chuan Shi, Yuan Fang, Yanfang Ye, Jiawei Zhang
Research Collection School Of Computing and Information Systems
The 4th Workshop on Heterogeneous Information Network Analysis and Applications (HENA 2021) is co-located with the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. The goal of this workshop is to bring together researchers and practitioners in the field and provide a forum for sharing new techniques and applications in heterogeneous information network analysis. This workshop has an exciting program that spans a number of subtopics, such as heterogeneous network embedding and graph neural networks, data mining techniques on heterogeneous information networks, and applications of heterogeneous information network analysis. The workshop program includes several invited speakers, lively discussion …
Automating The Removal Of Obsolete Todo Comments,
2021
Singapore Management University
Automating The Removal Of Obsolete Todo Comments, Zhipeng Gao, Xin Xia, David Lo, John C. Grundy, Thomas Zimmermann
Research Collection School Of Computing and Information Systems
TODO comments are very widely used by software developers to describe their pending tasks during software development. However, after performing the task developers sometimes neglect or simply forget to remove the TODO comment, resulting in obsolete TODO comments. These obsolete TODO comments can confuse development teams and may cause the introduction of bugs in the future, decreasing the software’s quality and maintainability. Manually identifying obsolete TODO comments is time-consuming and expensive. It is thus necessary to detect obsolete TODO comments and remove them automatically before they cause any unwanted side effects. In this work, we propose a novel model, named …
A Survey On Complex Knowledge Base Question Answering: Methods, Challenges And Solutions,
2021
Singapore Management University
A Survey On Complex Knowledge Base Question Answering: Methods, Challenges And Solutions, Yunshi Lan, Gaole He, Jinhao Jiang, Jing Jiang, Wayne Xin Zhao, Ji-Rong Wen
Research Collection School Of Computing and Information Systems
Knowledge base question answering (KBQA) aims to answer a question over a knowledge base (KB). Recently, a large number of studies focus on semantically or syntactically complicated questions. In this paper, we elaborately summarize the typical challenges and solutions for complex KBQA. We begin with introducing the background about the KBQA task. Next, we present the two mainstream categories of methods for complex KBQA, namely semantic parsing-based (SP-based) methods and information retrieval-based (IR-based) methods. We then review the advanced methods comprehensively from the perspective of the two categories. Specifically, we explicate their solutions to the typical challenges. Finally, we conclude …
Crossasr++: A Modular Differential Testing Framework For Automatic Speech Recognition,
2021
Singapore Management University
Crossasr++: A Modular Differential Testing Framework For Automatic Speech Recognition, Muhammad Hilmi Asyrofi, Zhou Yang, David Lo
Research Collection School Of Computing and Information Systems
Developers need to perform adequate testing to ensure the quality of Automatic Speech Recognition (ASR) systems. However, manually collecting required test cases is tedious and time-consuming. Our recent work proposes CrossASR, a differential testing method for ASR systems. This method first utilizes Text-to-Speech (TTS) to generate audios from texts automatically and then feed these audios into different ASR systems for cross-referencing to uncover failed test cases. It also leverages a failure estimator to find failing test cases more efficiently. Such a method is inherently self-improvable: the performance can increase by leveraging more advanced TTS and ASR systems. So, in this …
Pre-Training On Large-Scale Heterogeneous Graph,
2021
Singapore Management University
Pre-Training On Large-Scale Heterogeneous Graph, Xunqiang Jiang, Tianrui Jia, Yuan Fang, Chuan Shi, Zhe Lin, Hui Wang
Research Collection School Of Computing and Information Systems
Graph neural networks (GNNs) emerge as the state-of-the-art representation learning methods on graphs and often rely on a large amount of labeled data to achieve satisfactory performance. Recently, in order to relieve the label scarcity issues, some works propose to pre-train GNNs in a self-supervised manner by distilling transferable knowledge from the unlabeled graph structures. Unfortunately, these pre-training frameworks mainly target at homogeneous graphs, while real interaction systems usually constitute large-scale heterogeneous graphs, containing different types of nodes and edges, which leads to new challenges on structure heterogeneity and scalability for graph pre-training. In this paper, we first study the …
Characterizing Search Activities On Stack Overflow,
2021
Singapore Management University
Characterizing Search Activities On Stack Overflow, Jiakun Liu, Sebastian Baltes, Christoph Treude, David Lo, Yun Zhang, Xin Xia
Research Collection School Of Computing and Information Systems
To solve programming issues, developers commonly search on Stack Overflow to seek potential solutions. However, there is a gap between the knowledge developers are interested in and the knowledge they are able to retrieve using search engines. To help developers efficiently retrieve relevant knowledge on Stack Overflow, prior studies proposed several techniques to reformulate queries and generate summarized answers. However, few studies performed a large-scale analysis using real-world search logs. In this paper, we characterize how developers search on Stack Overflow using such logs. By doing so, we identify the challenges developers face when searching on Stack Overflow and seek …
Integrating Knowledge Compilation With Reinforcement Learning For Routes,
2021
Singapore Management University
Integrating Knowledge Compilation With Reinforcement Learning For Routes, Jiajing Ling, Kushagra Chandak, Akshat Kumar
Research Collection School Of Computing and Information Systems
Sequential multiagent decision-making under partial observability and uncertainty poses several challenges. Although multiagent reinforcement learning (MARL) approaches have increased the scalability, addressing combinatorial domains is still challenging as random exploration by agents is unlikely to generate useful reward signals. We address cooperative multiagent pathfinding under uncertainty and partial observability where agents move from their respective sources to destinations while also satisfying constraints (e.g., visiting landmarks). Our main contributions include: (1) compiling domain knowledge such as underlying graph connectivity and domain constraints into propositional logic based decision diagrams, (2) developing modular techniques to integrate such knowledge with deep MARL algorithms, and …
Toward Explainable Deep Anomaly Detection,
2021
Singapore Management University
Toward Explainable Deep Anomaly Detection, Guansong Pang, Charu Aggarwal
Research Collection School Of Computing and Information Systems
Anomaly explanation, also known as anomaly localization, is as important as, if not more than, anomaly detection in many realworld applications. However, it is challenging to build explainable detection models due to the lack of anomaly-supervisory information and the unbounded nature of anomaly; most existing studies exclusively focus on the detection task only, including the recently emerging deep learning-based anomaly detection that leverages neural networks to learn expressive low-dimensional representations or anomaly scores for the detection task. Deep learning models, including deep anomaly detection models, are often constructed as black boxes, which have been criticized for the lack of explainability …
Metaxmorph: Hierarchical Transformation Of Data With Metadata,
2021
Utah State University
Metaxmorph: Hierarchical Transformation Of Data With Metadata, Shubham Airan
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
This research is about transforming data. Data comes in different shapes; it can be structured as a graph, a tree, a collection of tables, or some other shape. In this thesis, we focus on data structured as a tree, which is known as hierarchical data. The same data could be structured in many different tree shapes. Previously it was shown how to transform data from one tree shape, one hierarchy to another without losing any information. But sometimes the pieces of the hierarchy are annotated or associated with metadata, that is, with data about the data itself. The metadata can …
Deep Learning Data And Indexes In A Database,
2021
Utah State University
Deep Learning Data And Indexes In A Database, Vishal Sharma
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
A database is used to store and retrieve data, which is a critical component for any software application. Databases requires configuration for efficiency, however, there are tens of configuration parameters. It is a challenging task to manually configure a database. Furthermore, a database must be reconfigured on a regular basis to keep up with newer data and workload. The goal of this thesis is to use the query workload history to autonomously configure the database and improve its performance. We achieve proposed work in four stages: (i) we develop an index recommender using deep reinforcement learning for a standalone database. …
Socio-Technical Perspective For Electronic Tax Information System In Tanzania,
2021
Department of Computer Science and Engineering, University of Dar es Salaam
Socio-Technical Perspective For Electronic Tax Information System In Tanzania, Lucas Ngowi, Ellen Kalinga
Tanzania Journal of Engineering and Technology (TJET)
Socio-technical systems theory has rarely been used by system architects in setting up computing systems. However, the role of socio-technical concepts in computing, which is becoming social in nature, has made the concepts more relevant and commercial. Tax information systems are examples of such systems because they are influenced by external variables such as the political environment, technological trends, and social environment, introducing complexity in their deployment and determining the type of e-services and their delivery to a diverse group of people. It was observed that in Tanzania there is resistance, reluctance and minimal use of electronic tax system because …
The Is Social Continuance Model: Using Conversational Agents To Support Co-Creation,
2021
University of South Florida
The Is Social Continuance Model: Using Conversational Agents To Support Co-Creation, Naif Alawi
USF Tampa Graduate Theses and Dissertations
With the rise of Agentic IS Artifact and the increasing integration of this technology within organizations, our understanding of the impact of this technology on individuals remains limited. Although IS use literature provides important guidance for organization to increase employees’ willingness to work with new technology implementations, the utilitarian view of prior IS use limits its application in light of the new evolving social interaction between humans and Agentic IS Artifacts. To that end, we contribute to the IS use literature by implementing a social view to understand the impact of Agentic IS Artifacts on an individual’s perception and behavior. …
