Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

7,251 Full-Text Articles 10,409 Authors 4,901,411 Downloads 214 Institutions

All Articles in Databases and Information Systems

Faceted Search

7,251 full-text articles. Page 83 of 268.

An Empirical Study Of The Discreteness Prior In Low-Rank Matrix Completion, Rodrigo ALVES, Antoine LEDENT, Renato ASSUNÇÃO, Marius and KLOFT 2021 Singapore Management University

An Empirical Study Of The Discreteness Prior In Low-Rank Matrix Completion, Rodrigo Alves, Antoine Ledent, Renato Assunção, Marius And Kloft

Research Collection School Of Computing and Information Systems

A reasonable assumption in recommender systems is that the rows (users) and columns (items) of the rating matrix can be split into groups (communities) with the following property: each entry of the matrix is the sum of components corresponding to community behavior and a purely low-rank component corresponding to individual behavior. We investigate (1) whether such a structure is present in real-world datasets, (2) whether the knowledge of the existence of such structure alone can improve performance, without explicit information about the community memberships. To these ends, we formulate a joint optimization problem over all (completed matrix, set of communities) …


A Survey On Ml4vis: Applying Machine Learning Advances To Data Visualization, Qianwen WANG, Zhutian CHEN, Yong WANG, Huamin QU 2021 Singapore Management University

A Survey On Ml4vis: Applying Machine Learning Advances To Data Visualization, Qianwen Wang, Zhutian Chen, Yong Wang, Huamin Qu

Research Collection School Of Computing and Information Systems

Inspired by the great success of machine learning (ML), researchers have applied ML techniques to visualizations to achieve a better design, development, and evaluation of visualizations. This branch of studies, known as ML4VIS, is gaining increasing research attention in recent years. To successfully adapt ML techniques for visualizations, a structured understanding of the integration of ML4VIS is needed. In this article, we systematically survey 88 ML4VIS studies, aiming to answer two motivating questions: “what visualization processes can be assisted by ML?” and “how ML techniques can be used to solve visualization problems? ” This survey reveals seven main processes where …


Vehicle Routing: Review Of Benchmark Datasets, Aldy GUNAWAN, Graham KENDALL, Barry McCollum, Hsin-Vonn SEOW, Lai Soon LEE 2021 Singapore Management University

Vehicle Routing: Review Of Benchmark Datasets, Aldy Gunawan, Graham Kendall, Barry Mccollum, Hsin-Vonn Seow, Lai Soon Lee

Research Collection School Of Computing and Information Systems

The Vehicle Routing Problem (VRP) was formally presented to the scientific literature in 1959 by Dantzig and Ramser (DOI:10.1287/mnsc.6.1.80). Sixty years on, the problem is still heavily researched, with hundreds of papers having been published addressing this problem and the many variants that now exist. Many datasets have been proposed to enable researchers to compare their algorithms using the same problem instances where either the best known solution is known or, in some cases, the optimal solution is known. In this survey paper, we provide a list of Vehicle Routing Problem datasets, categorized to enable researchers to have easy access …


Effective Digital Learning Practices For Is Design Courses During Covid-19, Eng Lieh OUH, Benjamin GAN 2021 Singapore Management University

Effective Digital Learning Practices For Is Design Courses During Covid-19, Eng Lieh Ouh, Benjamin Gan

Research Collection School Of Computing and Information Systems

The COVID-19 pandemic has pushed educational institutions to adopt digital learning for an extended period. This research studies the effectiveness of digital learning practices based on student feedback data collected for two Information Systems design courses: human interaction design and solution architecture design. This paper leverages the data to analyze the effectiveness of a set of digital learning practices: ZOOM lectures, polling or Kahoot questions, self-reflection, virtual exercises and virtual mentorship. Our research questions are on the effectiveness of these learning practices to keep the student’s interest and learn the course materials. The research compares each learning practice and the …


Modeling Transitions Of Focal Entities For Conversational Knowledge Base Question Answering, Yunshi LAN, Jing JIANG 2021 Singapore Management University

Modeling Transitions Of Focal Entities For Conversational Knowledge Base Question Answering, Yunshi Lan, Jing Jiang

Research Collection School Of Computing and Information Systems

Conversational KBQA is about answering a sequence of questions related to a KB. Follow-up questions in conversational KBQA often have missing information referring to entities from the conversation history. In this paper, we propose to model these implied entities, which we refer to as the focal entities of the conversation. We propose a novel graph-based model to capture the transitions of focal entities and apply a graph neural network to derive a probability distribution of focal entities for each question, which is then combined with a standard KBQA module to perform answer ranking. Our experiments on two datasets demonstrate the …


Thunderrw: An In-Memory Graph Random Walk Engine, Shixuan SUN, Yuhang CHEN, Shengliang LU, Bingsheng HE, Yuchen LI 2021 Singapore Management University

Thunderrw: An In-Memory Graph Random Walk Engine, Shixuan Sun, Yuhang Chen, Shengliang Lu, Bingsheng He, Yuchen Li

Research Collection School Of Computing and Information Systems

As random walk is a powerful tool in many graph processing, mining and learning applications, this paper proposes an efficient inmemory random walk engine named ThunderRW. Compared with existing parallel systems on improving the performance of a single graph operation, ThunderRW supports massive parallel random walks. The core design of ThunderRW is motivated by our profiling results: common RW algorithms have as high as 73.1% CPU pipeline slots stalled due to irregular memory access, which suffers significantly more memory stalls than the conventional graph workloads such as BFS and SSSP. To improve the memory efficiency, we first design a generic …


Context-Aware Outstanding Fact Mining From Knowledge Graphs, Yueji YANG, Yuchen LI, Panagiotis KARRAS, Anthony TUNG 2021 Singapore Management University

Context-Aware Outstanding Fact Mining From Knowledge Graphs, Yueji Yang, Yuchen Li, Panagiotis Karras, Anthony Tung

Research Collection School Of Computing and Information Systems

An Outstanding Fact (OF) is an attribute that makes a target entity stand out from its peers. The mining of OFs has important applications, especially in Computational Journalism, such as news promotion, fact-checking, and news story finding. However, existing approaches to OF mining: (i) disregard the context in which the target entity appears, hence may report facts irrelevant to that context; and (ii) require relational data, which are often unavailable or incomplete in many application domains. In this paper, we introduce the novel problem of mining Contextaware Outstanding Facts (COFs) for a target entity under a given context specified by …


Are Missing Links Predictable? An Inferential Benchmark For Knowledge Graph Completion, Yixin CAO, Xiang JI, Xin LV, Juanzi LI, Yonggang WEN, Hanwang ZHANG 2021 Singapore Management University

Are Missing Links Predictable? An Inferential Benchmark For Knowledge Graph Completion, Yixin Cao, Xiang Ji, Xin Lv, Juanzi Li, Yonggang Wen, Hanwang Zhang

Research Collection School Of Computing and Information Systems

We present InferWiki, a Knowledge Graph Completion (KGC) dataset that improves upon existing benchmarks in inferential ability, assumptions, and patterns. First, each testing sample is predictable with supportive data in the training set. To ensure it, we propose to utilize rule-guided train/test generation, instead of conventional random split. Second, InferWiki initiates the evaluation following the open-world assumption and improves the inferential difficulty of the closed-world assumption, by providing manually annotated negative and unknown triples. Third, we include various inference patterns (e.g., reasoning path length and types) for comprehensive evaluation. In experiments, we curate two settings of InferWiki varying in sizes …


The 4th Workshop On Heterogeneous Information Network Analysis And Applications (Hena 2021), Chuan SHI, Yuan FANG, Yanfang YE, Jiawei ZHANG 2021 Singapore Management University

The 4th Workshop On Heterogeneous Information Network Analysis And Applications (Hena 2021), Chuan Shi, Yuan Fang, Yanfang Ye, Jiawei Zhang

Research Collection School Of Computing and Information Systems

The 4th Workshop on Heterogeneous Information Network Analysis and Applications (HENA 2021) is co-located with the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. The goal of this workshop is to bring together researchers and practitioners in the field and provide a forum for sharing new techniques and applications in heterogeneous information network analysis. This workshop has an exciting program that spans a number of subtopics, such as heterogeneous network embedding and graph neural networks, data mining techniques on heterogeneous information networks, and applications of heterogeneous information network analysis. The workshop program includes several invited speakers, lively discussion …


Automating The Removal Of Obsolete Todo Comments, Zhipeng GAO, Xin XIA, David LO, John C. GRUNDY, Thomas ZIMMERMANN 2021 Singapore Management University

Automating The Removal Of Obsolete Todo Comments, Zhipeng Gao, Xin Xia, David Lo, John C. Grundy, Thomas Zimmermann

Research Collection School Of Computing and Information Systems

TODO comments are very widely used by software developers to describe their pending tasks during software development. However, after performing the task developers sometimes neglect or simply forget to remove the TODO comment, resulting in obsolete TODO comments. These obsolete TODO comments can confuse development teams and may cause the introduction of bugs in the future, decreasing the software’s quality and maintainability. Manually identifying obsolete TODO comments is time-consuming and expensive. It is thus necessary to detect obsolete TODO comments and remove them automatically before they cause any unwanted side effects. In this work, we propose a novel model, named …


A Survey On Complex Knowledge Base Question Answering: Methods, Challenges And Solutions, Yunshi LAN, Gaole HE, Jinhao JIANG, Jing JIANG, Wayne Xin ZHAO, Ji-Rong WEN 2021 Singapore Management University

A Survey On Complex Knowledge Base Question Answering: Methods, Challenges And Solutions, Yunshi Lan, Gaole He, Jinhao Jiang, Jing Jiang, Wayne Xin Zhao, Ji-Rong Wen

Research Collection School Of Computing and Information Systems

Knowledge base question answering (KBQA) aims to answer a question over a knowledge base (KB). Recently, a large number of studies focus on semantically or syntactically complicated questions. In this paper, we elaborately summarize the typical challenges and solutions for complex KBQA. We begin with introducing the background about the KBQA task. Next, we present the two mainstream categories of methods for complex KBQA, namely semantic parsing-based (SP-based) methods and information retrieval-based (IR-based) methods. We then review the advanced methods comprehensively from the perspective of the two categories. Specifically, we explicate their solutions to the typical challenges. Finally, we conclude …


Crossasr++: A Modular Differential Testing Framework For Automatic Speech Recognition, Muhammad Hilmi ASYROFI, Zhou YANG, David LO 2021 Singapore Management University

Crossasr++: A Modular Differential Testing Framework For Automatic Speech Recognition, Muhammad Hilmi Asyrofi, Zhou Yang, David Lo

Research Collection School Of Computing and Information Systems

Developers need to perform adequate testing to ensure the quality of Automatic Speech Recognition (ASR) systems. However, manually collecting required test cases is tedious and time-consuming. Our recent work proposes CrossASR, a differential testing method for ASR systems. This method first utilizes Text-to-Speech (TTS) to generate audios from texts automatically and then feed these audios into different ASR systems for cross-referencing to uncover failed test cases. It also leverages a failure estimator to find failing test cases more efficiently. Such a method is inherently self-improvable: the performance can increase by leveraging more advanced TTS and ASR systems. So, in this …


Pre-Training On Large-Scale Heterogeneous Graph, Xunqiang JIANG, Tianrui JIA, Yuan FANG, Chuan SHI, Zhe LIN, Hui WANG 2021 Singapore Management University

Pre-Training On Large-Scale Heterogeneous Graph, Xunqiang Jiang, Tianrui Jia, Yuan Fang, Chuan Shi, Zhe Lin, Hui Wang

Research Collection School Of Computing and Information Systems

Graph neural networks (GNNs) emerge as the state-of-the-art representation learning methods on graphs and often rely on a large amount of labeled data to achieve satisfactory performance. Recently, in order to relieve the label scarcity issues, some works propose to pre-train GNNs in a self-supervised manner by distilling transferable knowledge from the unlabeled graph structures. Unfortunately, these pre-training frameworks mainly target at homogeneous graphs, while real interaction systems usually constitute large-scale heterogeneous graphs, containing different types of nodes and edges, which leads to new challenges on structure heterogeneity and scalability for graph pre-training. In this paper, we first study the …


Characterizing Search Activities On Stack Overflow, Jiakun LIU, Sebastian BALTES, Christoph TREUDE, David LO, Yun ZHANG, Xin XIA 2021 Singapore Management University

Characterizing Search Activities On Stack Overflow, Jiakun Liu, Sebastian Baltes, Christoph Treude, David Lo, Yun Zhang, Xin Xia

Research Collection School Of Computing and Information Systems

To solve programming issues, developers commonly search on Stack Overflow to seek potential solutions. However, there is a gap between the knowledge developers are interested in and the knowledge they are able to retrieve using search engines. To help developers efficiently retrieve relevant knowledge on Stack Overflow, prior studies proposed several techniques to reformulate queries and generate summarized answers. However, few studies performed a large-scale analysis using real-world search logs. In this paper, we characterize how developers search on Stack Overflow using such logs. By doing so, we identify the challenges developers face when searching on Stack Overflow and seek …


Integrating Knowledge Compilation With Reinforcement Learning For Routes, Jiajing LING, Kushagra CHANDAK, Akshat KUMAR 2021 Singapore Management University

Integrating Knowledge Compilation With Reinforcement Learning For Routes, Jiajing Ling, Kushagra Chandak, Akshat Kumar

Research Collection School Of Computing and Information Systems

Sequential multiagent decision-making under partial observability and uncertainty poses several challenges. Although multiagent reinforcement learning (MARL) approaches have increased the scalability, addressing combinatorial domains is still challenging as random exploration by agents is unlikely to generate useful reward signals. We address cooperative multiagent pathfinding under uncertainty and partial observability where agents move from their respective sources to destinations while also satisfying constraints (e.g., visiting landmarks). Our main contributions include: (1) compiling domain knowledge such as underlying graph connectivity and domain constraints into propositional logic based decision diagrams, (2) developing modular techniques to integrate such knowledge with deep MARL algorithms, and …


Toward Explainable Deep Anomaly Detection, Guansong PANG, Charu AGGARWAL 2021 Singapore Management University

Toward Explainable Deep Anomaly Detection, Guansong Pang, Charu Aggarwal

Research Collection School Of Computing and Information Systems

Anomaly explanation, also known as anomaly localization, is as important as, if not more than, anomaly detection in many realworld applications. However, it is challenging to build explainable detection models due to the lack of anomaly-supervisory information and the unbounded nature of anomaly; most existing studies exclusively focus on the detection task only, including the recently emerging deep learning-based anomaly detection that leverages neural networks to learn expressive low-dimensional representations or anomaly scores for the detection task. Deep learning models, including deep anomaly detection models, are often constructed as black boxes, which have been criticized for the lack of explainability …


Metaxmorph: Hierarchical Transformation Of Data With Metadata, Shubham Airan 2021 Utah State University

Metaxmorph: Hierarchical Transformation Of Data With Metadata, Shubham Airan

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

This research is about transforming data. Data comes in different shapes; it can be structured as a graph, a tree, a collection of tables, or some other shape. In this thesis, we focus on data structured as a tree, which is known as hierarchical data. The same data could be structured in many different tree shapes. Previously it was shown how to transform data from one tree shape, one hierarchy to another without losing any information. But sometimes the pieces of the hierarchy are annotated or associated with metadata, that is, with data about the data itself. The metadata can …


Deep Learning Data And Indexes In A Database, Vishal Sharma 2021 Utah State University

Deep Learning Data And Indexes In A Database, Vishal Sharma

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

A database is used to store and retrieve data, which is a critical component for any software application. Databases requires configuration for efficiency, however, there are tens of configuration parameters. It is a challenging task to manually configure a database. Furthermore, a database must be reconfigured on a regular basis to keep up with newer data and workload. The goal of this thesis is to use the query workload history to autonomously configure the database and improve its performance. We achieve proposed work in four stages: (i) we develop an index recommender using deep reinforcement learning for a standalone database. …


Socio-Technical Perspective For Electronic Tax Information System In Tanzania, Lucas Ngowi, Ellen Kalinga 2021 Department of Computer Science and Engineering, University of Dar es Salaam

Socio-Technical Perspective For Electronic Tax Information System In Tanzania, Lucas Ngowi, Ellen Kalinga

Tanzania Journal of Engineering and Technology (TJET)

Socio-technical systems theory has rarely been used by system architects in setting up computing systems. However, the role of socio-technical concepts in computing, which is becoming social in nature, has made the concepts more relevant and commercial. Tax information systems are examples of such systems because they are influenced by external variables such as the political environment, technological trends, and social environment, introducing complexity in their deployment and determining the type of e-services and their delivery to a diverse group of people. It was observed that in Tanzania there is resistance, reluctance and minimal use of electronic tax system because …


The Is Social Continuance Model: Using Conversational Agents To Support Co-Creation, Naif Alawi 2021 University of South Florida

The Is Social Continuance Model: Using Conversational Agents To Support Co-Creation, Naif Alawi

USF Tampa Graduate Theses and Dissertations

With the rise of Agentic IS Artifact and the increasing integration of this technology within organizations, our understanding of the impact of this technology on individuals remains limited. Although IS use literature provides important guidance for organization to increase employees’ willingness to work with new technology implementations, the utilitarian view of prior IS use limits its application in light of the new evolving social interaction between humans and Agentic IS Artifacts. To that end, we contribute to the IS use literature by implementing a social view to understand the impact of Agentic IS Artifacts on an individual’s perception and behavior. …


Digital Commons powered by bepress