Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

7,251 Full-Text Articles 10,409 Authors 4,901,411 Downloads 214 Institutions

All Articles in Databases and Information Systems

Faceted Search

7,251 full-text articles. Page 66 of 268.

Learning Transferable Perturbations For Image Captioning, Hanjie WU, Yongtuo LIU, Hongmin CAI, Shengfeng HE 2022 Singapore Management University

Learning Transferable Perturbations For Image Captioning, Hanjie Wu, Yongtuo Liu, Hongmin Cai, Shengfeng He

Research Collection School Of Computing and Information Systems

Present studies have discovered that state-of-the-art deep learning models can be attacked by small but well-designed perturbations. Existing attack algorithms for the image captioning task is time-consuming, and their generated adversarial examples cannot transfer well to other models. To generate adversarial examples faster and stronger, we propose to learn the perturbations by a generative model that is governed by three novel loss functions. Image feature distortion loss is designed to maximize the encoded image feature distance between original images and the corresponding adversarial examples at the image domain, and local-global mismatching loss is introduced to separate the mapping encoding representation …


Unified Route Planning For Shared Mobility: An Insertion-Based Framework, Yongxin TONG, Yuxiang ZENG, Zimu ZHOU, Lei CHEN, Ke. XU 2022 Singapore Management University

Unified Route Planning For Shared Mobility: An Insertion-Based Framework, Yongxin Tong, Yuxiang Zeng, Zimu Zhou, Lei Chen, Ke. Xu

Research Collection School Of Computing and Information Systems

There has been a dramatic growth of shared mobility applications such as ride-sharing, food delivery, and crowdsourced parcel delivery. Shared mobility refers to transportation services that are shared among users, where a central issue is route planning. Given a set of workers and requests, route planning finds for each worker a route, i.e., a sequence of locations to pick up and drop off passengers/parcels that arrive from time to time, with different optimization objectives. Previous studies lack practicability due to their conflicted objectives and inefficiency in inserting a new request into a route, a basic operation called insertion. In addition, …


Learning Semantically Rich Network-Based Multi-Modal Mobile User Interface Embeddings, Meng Kiat Gary ANG, Ee-peng LIM 2022 Singapore Management University

Learning Semantically Rich Network-Based Multi-Modal Mobile User Interface Embeddings, Meng Kiat Gary Ang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Semantically rich information from multiple modalities - text, code, images, categorical and numerical data - co-exist in the user interface (UI) design of mobile applications. Moreover, each UI design is composed of inter-linked UI entities which support different functions of an application, e.g., a UI screen comprising a UI taskbar, a menu and multiple button elements. Existing UI representation learning methods unfortunately are not designed to capture multi-modal and linkage structure between UI entities. To support effective search and recommendation applications over mobile UIs, we need UI representations that integrate latent semantics present in both multi-modal information and linkages between …


Mmekg: Multi-Modal Event Knowledge Graph Towards Universal Representation Across Modalities, Yubo MA, Zehao WANG, Mukai LI, Yixin CAO, Meiqi CHEN, Xinze LI, Wenqi SUN, Kunquan DENG, Kun WANG, Aixin SUN, Jing SHAO 2022 Singapore Management University

Mmekg: Multi-Modal Event Knowledge Graph Towards Universal Representation Across Modalities, Yubo Ma, Zehao Wang, Mukai Li, Yixin Cao, Meiqi Chen, Xinze Li, Wenqi Sun, Kunquan Deng, Kun Wang, Aixin Sun, Jing Shao

Research Collection School Of Computing and Information Systems

Events are fundamental building blocks of realworld happenings. In this paper, we present a large-scale, multi-modal event knowledge graph named MMEKG. MMEKG unifies different modalities of knowledge via events, which complement and disambiguate each other. Specifically, MMEKG incorporates (i) over 990 thousand concept events with 644 relation types to cover most types of happenings, and (ii) over 863 million instance events connected through 934 million relations, which provide rich contextual information in texts and/or images. To collect billion-scale instance events and relations among them, we additionally develop an efficient yet effective pipeline for textual/visual knowledge extraction system. We also develop …


Prompt For Extraction? Paie: Prompting Argument Interaction For Event Argument Extraction, Yubo MA, Zehao WANG, Yixin CAO, Mukai LI, Meiqi CHEN, Kun WANG, Jing SHAO 2022 Singapore Management University

Prompt For Extraction? Paie: Prompting Argument Interaction For Event Argument Extraction, Yubo Ma, Zehao Wang, Yixin Cao, Mukai Li, Meiqi Chen, Kun Wang, Jing Shao

Research Collection School Of Computing and Information Systems

In this paper, we propose an effective yet efficient model PAIE for both sentence-level and document-level Event Argument Extraction (EAE), which also generalizes well when there is a lack of training data. On the one hand, PAIE utilizes prompt tuning for extractive objectives to take the best advantages of Pre-trained Language Models (PLMs). It introduces two span selectors based on the prompt to select start/end tokens among input texts for each role. On the other hand, it captures argument interactions via multi-role prompts and conducts joint optimization with optimal span assignments via a bipartite matching loss. Also, with a flexible …


Translate-Train Embracing Translationese Artifacts, Sicheng YU, Qianru SUN, Hao ZHANG, Jing JIANG 2022 Singapore Management University

Translate-Train Embracing Translationese Artifacts, Sicheng Yu, Qianru Sun, Hao Zhang, Jing Jiang

Research Collection School Of Computing and Information Systems

Translate-train is a general training approach to multilingual tasks. The key idea is to use the translator of the target language to generate training data to mitigate the gap between the source and target languages. However, its performance is often hampered by the artifacts in the translated texts (translationese). We discover that such artifacts have common patterns in different languages and can be modeled by deep learning, and subsequently propose an approach to conduct translate-train using Translationese Embracing the effect of Artifacts (TEA). TEA learns to mitigate such effect on the training data of a source language (whose original and …


Storm The Capitol: Linking Offline Political Speech And Online Twitter Extra-Representational Participation On Qanon And The January 6 Insurrection, Claire Seungeun LEE, Juan MERIZALDE, John D. COLAUTTI, Jisun AN, Haewoon KWAK 2022 Singapore Management University

Storm The Capitol: Linking Offline Political Speech And Online Twitter Extra-Representational Participation On Qanon And The January 6 Insurrection, Claire Seungeun Lee, Juan Merizalde, John D. Colautti, Jisun An, Haewoon Kwak

Research Collection School Of Computing and Information Systems

The transfer of power stemming from the 2020 presidential election occurred during an unprecedented period in United States history. Uncertainty from the COVID-19 pandemic, ongoing societal tensions, and a fragile economy increased societal polarization, exacerbated by the outgoing president's offline rhetoric. As a result, online groups such as QAnon engaged in extra political participation beyond the traditional platforms. This research explores the link between offline political speech and online extra-representational participation by examining Twitter within the context of the January 6 insurrection. Using a mixed-methods approach of quantitative and qualitative thematic analyses, the study combines offline speech information with Twitter …


Unified And Incremental Simrank: Index-Free Approximation With Scheduled Principle (Extended Abstract), Fanwei ZHU, Yuan FANG, Kai ZHANG, Kevin Chen-Chuan CHANG, Hongtai CAO, Zhen JIANG, Minghui WU 2022 Singapore Management University

Unified And Incremental Simrank: Index-Free Approximation With Scheduled Principle (Extended Abstract), Fanwei Zhu, Yuan Fang, Kai Zhang, Kevin Chen-Chuan Chang, Hongtai Cao, Zhen Jiang, Minghui Wu

Research Collection School Of Computing and Information Systems

SimRank is a popular link-based similarity measure on graphs. It enables a variety of applications with different modes of querying. In this paper, we propose UISim, a unified and incremental framework for all SimRank modes based on a scheduled approximation principle. UISim processes queries with incremental and prioritized exploration of the entire computation space, and thus allows flexible tradeoff of time and accuracy. On the other hand, it creates and shares common “building blocks” for online computation without relying on indexes, and thus is efficient to handle both static and dynamic graphs. Our experiments on various real-world graphs show that …


Context Modeling With Evidence Filter For Multiple Choice Question Answering, Sicheng YU, Hao ZHANG, Wei JING, Jing JIANG 2022 Singapore Management University

Context Modeling With Evidence Filter For Multiple Choice Question Answering, Sicheng Yu, Hao Zhang, Wei Jing, Jing Jiang

Research Collection School Of Computing and Information Systems

Multiple-Choice Question Answering (MCQA) is one of the challenging tasks in machine reading comprehension. The main challenge in MCQA is to extract "evidence" from the given context that supports the correct answer. In OpenbookQA dataset [1], the requirement of extracting "evidence" is particularly important due to the mutual independence of sentences in the context. Existing work tackles this problem by annotated evidence or distant supervision with rules which overly rely on human efforts. To address the challenge, we propose a simple yet effective approach termed evidence filtering to model the relationships between the encoded contexts with respect to different options …


Structure-Aware Visualization Retrieval, Haotian LI, Yong WANG, Aoyu WU, Huan WEI, Huamin. QU 2022 Singapore Management University

Structure-Aware Visualization Retrieval, Haotian Li, Yong Wang, Aoyu Wu, Huan Wei, Huamin. Qu

Research Collection School Of Computing and Information Systems

With the wide usage of data visualizations, a huge number of Scalable Vector Graphic (SVG)-based visualizations have been created and shared online. Accordingly, there has been an increasing interest in exploring how to retrieve perceptually similar visualizations from a large corpus, since it can benefit various downstream applications such as visualization recommendation. Existing methods mainly focus on the visual appearance of visualizations by regarding them as bitmap images. However, the structural information intrinsically existing in SVG-based visualizations is ignored. Such structural information can delineate the spatial and hierarchical relationship among visual elements, and characterize visualizations thoroughly from a new perspective. …


Arseek: Identifying Api Resource Using Code And Discussion On Stack Overflow, Gia Kien LUONG, Mohammad HADI, Thung Ferdian, Fatemeh H. FARD, David LO 2022 Singapore Management University

Arseek: Identifying Api Resource Using Code And Discussion On Stack Overflow, Gia Kien Luong, Mohammad Hadi, Thung Ferdian, Fatemeh H. Fard, David Lo

Research Collection School Of Computing and Information Systems

It is not a trivial problem to collect API-relevant examples, usages, and mentions on venues such as Stack Overflow. It requires efforts to correctly recognize whether the discussion refers to the API method that developers/tools are searching for. The content of the Stack Overflow thread, which consists of both text paragraphs describing the involvement of the API method in the discussion and the code snippets containing the API invocation, may refer to the given API method. Leveraging this observation, we develop ARSeek, a context-specific algorithm to capture the semantic and syntactic information of the paragraphs and code snippets in a …


An Exploratory Study On Code Attention In Bert, Rishab SHARMA, Fuxiang CHEN, Fatemeh H. FARD, David LO 2022 Singapore Management University

An Exploratory Study On Code Attention In Bert, Rishab Sharma, Fuxiang Chen, Fatemeh H. Fard, David Lo

Research Collection School Of Computing and Information Systems

Many recent models in software engineering introduced deep neural models based on the Transformer architecture or use transformerbased Pre-trained Language Models (PLM) trained on code. Although these models achieve the state of the arts results in many downstream tasks such as code summarization and bug detection, they are based on Transformer and PLM, which are mainly studied in the Natural Language Processing (NLP) field. The current studies rely on the reasoning and practices from NLP for these models in code, despite the differences between natural languages and programming languages. There is also limited literature on explaining how code is modeled. …


Arsearch: Searching For Api Related Resources From Stack Overflow And Github, Kien LUONG, Ferdian THUNG, David LO 2022 Singapore Management University

Arsearch: Searching For Api Related Resources From Stack Overflow And Github, Kien Luong, Ferdian Thung, David Lo

Research Collection School Of Computing and Information Systems

Stack Overflow and GitHub are two popular platforms containing API-related resources for developers to learn how to use APIs. The platforms are good sources for information about API such as code examples, usages, sentiment, bug reports, etc. However, it is difficult to collect the correct resources regarding a particular API due to the ambiguity of an API method name. An API method name mentioned in the text would only refer to one API, but the method name could match with different APIs. To help people in finding the correct resources for a particular API, we introduce ARSearch. ARSearch finds Stack …


Data Pricing In Machine Learning Pipelines, Zicun CONG, Xuan LUO, Jian PEI, Feida ZHU, Yong ZHANG 2022 Singapore Management University

Data Pricing In Machine Learning Pipelines, Zicun Cong, Xuan Luo, Jian Pei, Feida Zhu, Yong Zhang

Research Collection School Of Computing and Information Systems

Machine learning is disruptive. At the same time, machine learning can only succeed by collaboration among many parties in multiple steps naturally as pipelines in an eco-system, such as collecting data for possible machine learning applications, collaboratively training models by multiple parties and delivering machine learning services to end users. Data are critical and penetrating in the whole machine learning pipelines. As machine learning pipelines involve many parties and, in order to be successful, have to form a constructive and dynamic eco-system, marketplaces and data pricing are fundamental in connecting and facilitating those many parties. In this article, we survey …


Optimal In‐Place Suffix Sorting, Zhize LI, Jian LI, Hongwei HUO 2022 Singapore Management University

Optimal In‐Place Suffix Sorting, Zhize Li, Jian Li, Hongwei Huo

Research Collection School Of Computing and Information Systems

The suffix array is a fundamental data structure for many applications that involve string searching and data compression. Designing time/space-efficient suffix array construction algorithms has attracted significant attention and considerable advances have been made for the past 20 years. We obtain the \emph{first} in-place suffix array construction algorithms that are optimal both in time and space for (read-only) integer alphabets. Concretely, we make the following contributions: 1. For integer alphabets, we obtain the first suffix sorting algorithm which takes linear time and uses only $O(1)$ workspace (the workspace is the total space needed beyond the input string and the output …


Cancel Culture: Who Or What Will Be Next?, Christine Trumper 2022 Bryant University

Cancel Culture: Who Or What Will Be Next?, Christine Trumper

Honors Projects in Data Science

This paper utilizes Data Science and Applied Statistic techniques, to perform an analytical dive into Cancel Culture as it is referenced and used on Twitter. The research focuses on analyzing how Cancel Culture has affected the sentiment of Twitter, specifically how it impacts prominent topics in the media that have occurred between February 2021 to September 2021. The development of a topic and sentiment analysis will be based on 1,302,844 Tweets collected using Twitter’s API. Cancel Culture became popularized on social media in the past few years and there is little concrete information regarding its process and the demographics it …


Performance Comparison Of The Filesystem And Embedded Key-Value Databases, Jesse Hines, Nicholas Cunningham 2022 Southern Adventist University

Performance Comparison Of The Filesystem And Embedded Key-Value Databases, Jesse Hines, Nicholas Cunningham

Campus Research Month

A common scenario when developing local applications is storing many records and then retrieving them by ID. A developer can simply save the records as files or use an embedded database. Large numbers of files can slow down filesystems, but developers may want to avoid a dependency on an embedded database if it offers little benefit for their use case. We will compare the performance for the insert, update, get and delete operations and the space efficiency of storing records as files vs. using key-value embedded databases including RocksDB, LevelDB, Berkley DB, and SQLite.


Database Query Execution Through Virtual Reality, Logan Bateman, Marc Butler 2022 Southern Adventist University

Database Query Execution Through Virtual Reality, Logan Bateman, Marc Butler

Campus Research Month

Building database queries often requires technical knowledge of a query language. However, company employees, such as executives, managers, and others (outside of software research and development, generally) may not have the pre-required knowledge to accurately construct and execute database queries. This paper proposes an approach to constructing database queries using virtual reality. This approach utilizes natural hand or controller gestures which map to various components of building and visualizing database queries.


Realtime Visualization Of Kafka Architectures, Matthew Jensen, Miro Manestar 2022 Southern Adventist University

Realtime Visualization Of Kafka Architectures, Matthew Jensen, Miro Manestar

Campus Research Month

Apache Kafka specializes in the transfer of incredibly large amounts of data in real-time between devices. However, it can be difficult to comprehend the inner workings of Kafka. Often, to get real-time data, a user must run complicated commands from within the Kafka CLI. Our contribution is a tool that monitors Kafka consumers, producers, and topics, and displays the flow of events between them in a web-based dashboard. This dashboard serves to reduce the complexity of Kafka and enables users unfamiliar with the platform and protocol to better understand how their architecture is configured.


Novel 360-Degree Camera, Ian Gauger, Andrew Kurtz, Zakariya Niazi 2022 Circle Optics

Novel 360-Degree Camera, Ian Gauger, Andrew Kurtz, Zakariya Niazi

Frameless

Circle Optics is developing novel technology for low-parallax, real time, panoramic image capture using an integrated array of multiple adjacent polygonal-edged cameras. This technology can be optimized and deployed for a variety of markets, including cinematic VR. Circle Optics’ existing prototype, Hydra Alpha, will be demonstrated.


Digital Commons powered by bepress