Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

7,251 Full-Text Articles 10,409 Authors 4,901,411 Downloads 214 Institutions

All Articles in Databases and Information Systems

Faceted Search

7,251 full-text articles. Page 155 of 268.

Guest Editor's Introduction To The Special Issue On Source Code Analysis And Manipulation (Scam 2015), Foutse KHOMH, David LO, Michael W. GODFREY 2017 Singapore Management University

Guest Editor's Introduction To The Special Issue On Source Code Analysis And Manipulation (Scam 2015), Foutse Khomh, David Lo, Michael W. Godfrey

Research Collection School Of Computing and Information Systems

We are happy to introduce you to this special issue that presents selected papers from the 15th IEEE International Working Conference on Source Code Analysis and Manipulation (SCAM 2015). SCAM is a leading conference that brings together researchers and practitioners working on theory, techniques, and applications that concern analysis and/or manipulation of the source code of computer systems. While much attention in the wider software engineering community is properly directed towards other aspects of systems development and evolution, such as specification, design, and requirements engineering, it is the source code that contains the only precise description of the behavior of …


An Integrated Framework For Modeling And Predicting Spatiotemporal Phenomena In Urban Environments, Tuc Viet LE 2017 Singapore Management University

An Integrated Framework For Modeling And Predicting Spatiotemporal Phenomena In Urban Environments, Tuc Viet Le

Dissertations and Theses Collection (Open Access)

This thesis proposes a general solution framework that integrates methods in machine learning in creative ways to solve a diverse set of problems arising in urban environments. It particularly focuses on modeling spatiotemporal data for the purpose of predicting urban phenomena. Concretely, the framework is applied to solve three specific real-world problems: human mobility prediction, trac speed prediction and incident prediction. For human mobility prediction, I use visitor trajectories collected a large theme park in Singapore as a simplified microcosm of an urban area. A trajectory is an ordered sequence of attraction visits and corresponding timestamps produced by a visitor. …


Scalable Online Kernel Learning, Jing LU 2017 Singapore Management University

Scalable Online Kernel Learning, Jing Lu

Dissertations and Theses Collection (Open Access)

One critical deficiency of traditional online kernel learning methods is their increasing and unbounded number of support vectors (SV’s), making them inefficient and non-scalable for large-scale applications. Recent studies on budget online learning have attempted to overcome this shortcoming by bounding the number of SV’s. Despite being extensively studied, budget algorithms usually suffer from several drawbacks.
First of all, although existing algorithms attempt to bound the number of SV’s at each iteration, most of them fail to bound the number of SV’s for the final averaged classifier, which is commonly used for online-to-batch conversion. To solve this problem, we propose …


Vireo @ Trecvid 2017: Video-To-Text, Ad-Hoc Video Search And Video Hyperlinking, Phuong Anh NGUYEN, Qing LI, Zhi-Qi CHENG, Yi-Jie LU, Hao ZHANG, Xiao WU, Chong-wah NGO 2017 Singapore Management University

Vireo @ Trecvid 2017: Video-To-Text, Ad-Hoc Video Search And Video Hyperlinking, Phuong Anh Nguyen, Qing Li, Zhi-Qi Cheng, Yi-Jie Lu, Hao Zhang, Xiao Wu, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

In this paper, we describe the systems developed for Video-to-Text (VTT), Ad-hoc Video Search (AVS) and Video Hyper-linking (LNK) tasks at TRECVID 2017 [1] and the achieved results.


Color-Sketch Simulator: A Guide For Color-Based Visual Known-Item Search, Jakub LOKOČ, Anh Nguyen PHUONG, Marta VOMLELOVÁ, Chong-wah NGO 2017 Charles University

Color-Sketch Simulator: A Guide For Color-Based Visual Known-Item Search, Jakub Lokoč, Anh Nguyen Phuong, Marta Vomlelová, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

In order to evaluate the effectiveness of a color-sketch retrieval system for a given multimedia database, tedious evaluations involving real users are required as users are in the center of query sketch formulation. However, without any prior knowledge about the bottlenecks of the underlying sketch-based retrieval model, the evaluations may focus on wrong settings and thus miss the desired effect. Furthermore, users have usually no clues or recommendations to draw color-sketches effectively. In this paper, we aim at a preliminary analysis to identify potential bottlenecks of a flexible color-sketch retrieval model. We present a formal framework based on position-color feature …


Modeling Check-In Behavior With Geographical Neighborhood Influence Of Venues, Thanh Nam DOAN, Ee Peng LIM 2017 Singapore Management University

Modeling Check-In Behavior With Geographical Neighborhood Influence Of Venues, Thanh Nam Doan, Ee Peng Lim

Research Collection School Of Computing and Information Systems

With many users adopting location-based social networks (LBSNs) to share their daily activities, LBSNs become a gold mine for researchers to study human check-in behavior. Modeling such behavior can benefit many useful applications such as urban planning and location-aware recommender systems. Unlike previous studies [4,6,12,17] that focus on the effect of distance on users checking in venues, we consider two venue-specific effects of geographical neighborhood influence, namely, spatial homophily and neighborhood competition. The former refers to the fact that venues share more common features with their spatial neighbors, while the latter captures the rivalry of a venue and its nearby …


Predicting Indoor Crowd Density Using Column-Structured Deep Neural Network, Akihito SUDO, Teck Hou (DENG Dehao) TENG, Hoong Chuin LAU, Yoshihide SEKIMOTO 2017 Shizuoko University

Predicting Indoor Crowd Density Using Column-Structured Deep Neural Network, Akihito Sudo, Teck Hou (Deng Dehao) Teng, Hoong Chuin Lau, Yoshihide Sekimoto

Research Collection School Of Computing and Information Systems

This work proposes a deep neural network approach known as the column-structured deep neural network (COL-DNN-R) for predicting crowd density in an indoor environment using historical Wi-Fi traces of individual visitors. With a structure designed to minimize feature engineering, COL-DNN accepts raw features such as crowd density, opening and closing hours and peak visitor counts for extracting features. The extracted features are used by a regression model R for predicting the crowd densities. Standard regression models such as MLP, RF and SVM can be used as R. Experiments are performed to investigate the effect of feature representation and model structure …


Leveraging Social Analytics Data For Identifying Customer Segments For Online News Media, JANSEN, Bernard J, Soon-Gyo JUNG, Jisun AN, Haewoon KWAK, HAEWOON KWAK 2017 Singapore Management University

Leveraging Social Analytics Data For Identifying Customer Segments For Online News Media, Jansen, Bernard J, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Haewoon Kwak

Research Collection School Of Computing and Information Systems

In this work, we describe a methodology for leveraging large amounts of customer interaction data with online content from major social media platforms in order to isolate meaningful customer segments. The methodology is robust in that it can rapidly identify diverse customer segments using solely online behaviors and then associate these behavioral customer segments with the related distinct demographic segments, presenting a holistic picture of the customer base of an organization. We validate our methodology via the implementation of a working system that rapidly and in near real-time processes tens of millions of online customer interactions with content posted on …


Eeg-Based Emotion Recognition Via Fast And Robust Feature Smoothing, Cheng TANG, Di WANG, Ah-hwee TAN, Chunyan MIAO 2017 Nanyang Technological University

Eeg-Based Emotion Recognition Via Fast And Robust Feature Smoothing, Cheng Tang, Di Wang, Ah-Hwee Tan, Chunyan Miao

Research Collection School Of Computing and Information Systems

Electroencephalograph (EEG) signals reveal much of our brain states and have been widely used in emotion recognition. However, the recognition accuracy is hardly ideal mainly due to the following reasons: (i) the features extracted from EEG signals may not solely reflect one’s emotional patterns and their quality is easily affected by noise; and (ii) increasing feature dimension may enhance the recognition accuracy, but it often requires extra computation time. In this paper, we propose a feature smoothing method to alleviate the aforementioned problems. Specifically, we extract six statistical features from raw EEG signals and apply a simple yet cost-effective feature …


Second-Order Online Active Learning And Its Applications, Shuji HAO, Jing LU, Peilin ZHAO, Chi ZHANG, Steven C. H. HOI, Chunyan MIAO 2017 Institute of High Performance Computing

Second-Order Online Active Learning And Its Applications, Shuji Hao, Jing Lu, Peilin Zhao, Chi Zhang, Steven C. H. Hoi, Chunyan Miao

Research Collection School Of Computing and Information Systems

The goal of online active learning is to learn predictive models from a sequence of unlabeled data given limited label querybudget. Unlike conventional online learning tasks, online active learning is considerably more challenging because of two reasons.Firstly, it is difficult to design an effective query strategy to decide when is appropriate to query the label of an incoming instance givenlimited query budget. Secondly, it is also challenging to decide how to update the predictive models effectively whenever the true labelof an instance is queried. Most existing approaches for online active learning are often based on a family of first-order online …


Highly Efficient Mining Of Overlapping Clusters In Signed Weighted Networks, Tuan-Anh HOANG, Ee-peng LIM 2017 Leibniz University of Hanover

Highly Efficient Mining Of Overlapping Clusters In Signed Weighted Networks, Tuan-Anh Hoang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

In many practical contexts, networks are weighted as their links are assigned numerical weights representing relationship strengths or intensities of inter-node interaction. Moreover, the links' weight can be positive or negative, depending on the relationship or interaction between the connected nodes. The existing methods for network clustering however are not ideal for handling very large signed weighted networks. In this paper, we present a novel method called LPOCSIN (short for "Linear Programming based Overlapping Clustering on Signed Weighted Networks") for efficient mining of overlapping clusters in signed weighted networks. Different from existing methods that rely on computationally expensive cluster cohesiveness …


On Analyzing Job Hop Behavior And Talent Flow Networks, Richard J. OENTARYO, Xavier Jayaraj Siddarth ASHOK, Ee-peng LIM, Philips Kokoh PRASETYO 2017 McLaren Applied Technologies

On Analyzing Job Hop Behavior And Talent Flow Networks, Richard J. Oentaryo, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim, Philips Kokoh Prasetyo

Research Collection School Of Computing and Information Systems

Analyzing job hopping behavior is important for theunderstanding of job preference and career progression of working individuals.When analyzed at the workforce population level, job hop analysis helps to gaininsights of talent flow and organization competition. Traditionally, surveysare conducted on job seekers and employers to study job behavior. While surveysare good at getting direct user input to specially designed questions, they areoften not scalable and timely enough to cope with fast-changing job landscape.In this paper, we present a data science approach to analyze job hops performedby about 490,000 working professionals located in a city using their publiclyshared profiles. We develop several …


Collaborative Topic Regression With Denoising Autoencoder For Content And Community Co-Representation, Trong T. NGUYEN, Hady W. LAUW 2017 Singapore Management University

Collaborative Topic Regression With Denoising Autoencoder For Content And Community Co-Representation, Trong T. Nguyen, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Personalized recommendation of items frequently faces scenarios where we have sparse observations on users' adoption of items. In the literature, there are two promising directions. One is to connect sparse items through similarity in content. The other is to connect sparse users through similarity in social relations. We seek to integrate both types of information, in addition to the adoption information, within a single integrated model. Our proposed method models item content via a topic model, and user communities via an autoencoder model, while bridging a user's community-based preference to her topic-based preference. Experiments on public real-life data showcase the …


Indexable Bayesian Personalized Ranking For Efficient Top-K Recommendation, Dung D. LE, Hady W. LAUW 2017 Singapore Management University

Indexable Bayesian Personalized Ranking For Efficient Top-K Recommendation, Dung D. Le, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Top-k recommendation seeks to deliver a personalized recommendation list of k items to a user. The dual objectives are (1) accuracy in identifying the items a user is likely to prefer, and (2) efficiency in constructing the recommendation list in real time. One direction towards retrieval efficiency is to formulate retrieval as approximate k nearest neighbor (kNN) search aided by indexing schemes, such as locality-sensitive hashing, spatial trees, and inverted index. These schemes, applied on the output representations of recommendation algorithms, speed up the retrieval process by automatically discarding a large number of potentially irrelevant items when given a user …


Answerbot: Automated Generation Of Answer Summary To Developers’ Technical Questions, Bowen XU, Zhenchang XING, Xin XIA, David LO 2017 Singapore Management University

Answerbot: Automated Generation Of Answer Summary To Developers’ Technical Questions, Bowen Xu, Zhenchang Xing, Xin Xia, David Lo

Research Collection School Of Computing and Information Systems

The prevalence of questions and answers on domain-specific Q&A sites like Stack Overflow constitutes a core knowledge asset for software engineering domain. Although search engines can return a list of questions relevant to a user query of some technical question, the abundance of relevant posts and the sheer amount of information in them makes it difficult for developers to digest them and find the most needed answers to their questions. In this work, we aim to help developers who want to quickly capture the key points of several answer posts relevant to a technical question before they read the details …


Sourcevote: Fusing Multi-Valued Data Via Inter-Source Agreements, Xiu Susie FANG, Quan Z. SHENG, Xianzhi WANG, Mahmoud BARHAMGI, Lina YAO, Anne H.H. NGU 2017 Macquarie University

Sourcevote: Fusing Multi-Valued Data Via Inter-Source Agreements, Xiu Susie Fang, Quan Z. Sheng, Xianzhi Wang, Mahmoud Barhamgi, Lina Yao, Anne H.H. Ngu

Research Collection School Of Computing and Information Systems

Data fusion is a fundamental research problem of identifying true values of data items of interest from conflicting multi-sourced data. Although considerable research efforts have been conducted on this topic, existing approaches generally assume every data item has exactly one true value, which fails to reflect the real world where data items with multiple true values widely exist. In this paper, we propose a novel approach,SourceVote, to estimate value veracity for multi-valued data items. SourceVote models the endorsement relations among sources by quantifying their two-sided inter-source agreements. In particular, two graphs are constructed to model inter-source relations. Then two aspects …


Tweet Geolocation: Leveraging Location, User And Peer Signals, Wen-Haw CHONG, Ee Peng LIM 2017 Singapore Management University

Tweet Geolocation: Leveraging Location, User And Peer Signals, Wen-Haw Chong, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Which venue is a tweet posted from? We referred this as fine-grained geolocation. To solve this problem effectively, we develop novel techniques to exploit each posting user's content history. This is motivated by our finding that most users do not share their visitation history, but have ample content history from tweet posts. We formulate fine-grained geolocation as a ranking problem whereby given a test tweet, we rank candidate venues. We propose several models that leverage on three types of signals from locations, users and peers. Firstly, the location signals are words that are indicative of venues. We propose a location-indicative …


Interactive Social Recommendation, Xin WANG, Steven C. H. HOI, Chenghao LIU, Martin ESTER 2017 Tsinghua University

Interactive Social Recommendation, Xin Wang, Steven C. H. Hoi, Chenghao Liu, Martin Ester

Research Collection School Of Computing and Information Systems

Social recommendation has been an active research topic over the last decade, based on the assumption that social information from friendship networks is beneficial for improving recommendation accuracy, especially when dealing with cold-start users who lack sufficient past behavior information for accurate recommendation. However, it is nontrivial to use such information, since some of a person's friends may share similar preferences in certain aspects, but others may be totally irrelevant for recommendations. Thus one challenge is to explore and exploit the extend to which a user trusts his/her friends when utilizing social information to improve recommendations. On the other hand, …


Unsupervised Topic Hypergraph Hashing For Efficient Mobile Image Retrieval, Lei ZHU, Jialie SHEN, Liang XIE, Zhiyong CHENG 2017 Singapore Management University

Unsupervised Topic Hypergraph Hashing For Efficient Mobile Image Retrieval, Lei Zhu, Jialie Shen, Liang Xie, Zhiyong Cheng

Research Collection School Of Computing and Information Systems

Hashing compresses high-dimensional features into compact binary codes. It is one of the promising techniques to support efficient mobile image retrieval, due to its low data transmission cost and fast retrieval response. However, most of existing hashing strategies simply rely on low-level features. Thus, they may generate hashing codes with limited discriminative capability. Moreover, many of them fail to exploit complex and high-order semantic correlations that inherently exist among images. Motivated by these observations, we propose a novel unsupervised hashing scheme, called topic hypergraph hashing (THH), to address the limitations. THH effectively mitigates the semantic shortage of hashing codes by …


Selective Value Coupling Learning For Detecting Outliers In High-Dimensional Categorical Data, Guansong PANG, Hongzuo XU, CAO Longbing, Wentao ZHAO 2017 Singapore Management University

Selective Value Coupling Learning For Detecting Outliers In High-Dimensional Categorical Data, Guansong Pang, Hongzuo Xu, Cao Longbing, Wentao Zhao

Research Collection School Of Computing and Information Systems

This paper introduces a novel framework, namely SelectVC and its instance POP, for learning selective value couplings (i.e., interactions between the full value set and a set of outlying values) to identify outliers in high-dimensional categorical data. Existing outlier detection methods work on a full data space or feature subspaces that are identified independently from subsequent outlier scoring. As a result, they are significantly challenged by overwhelming irrelevant features in high-dimensional data due to the noise brought by the irrelevant features and its huge search space. In contrast, SelectVC works on a clean and condensed data space spanned by selective …


Digital Commons powered by bepress