Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 2071 - 2100 of 3560

Full-Text Articles in Databases and Information Systems

Cnl: Collective Network Linkage Across Heterogeneous Social Platforms, Ming Gao, Ee-Peng Lim, David Lo, Feida Zhu, Philips Kokoh Prasetyo, Aoying Zhou Nov 2015

Cnl: Collective Network Linkage Across Heterogeneous Social Platforms, Ming Gao, Ee-Peng Lim, David Lo, Feida Zhu, Philips Kokoh Prasetyo, Aoying Zhou

Research Collection School Of Computing and Information Systems

The popularity of social media has led many users to create accounts with different online social networks. Identifying these multiple accounts belonging to same user is of critical importance to user profiling, community detection, user behavior understanding and product recommendation. Nevertheless, linking users across heterogeneous social networks is challenging due to large network sizes, heterogeneous user attributes and behaviors in different networks, and noises in user generated data. In this paper, we propose an unsupervised method, Collective Network Linkage (CNL), to link users across heterogeneous social networks. CNL incorporates heterogeneous attributes and social features unique to social network users, handles …


Not All Trips Are Equal: Analyzing Foursquare Check-Ins Of Trips And City Visitors, Wen Haw Chong, Bingtian Dai, Ee Peng Lim Nov 2015

Not All Trips Are Equal: Analyzing Foursquare Check-Ins Of Trips And City Visitors, Wen Haw Chong, Bingtian Dai, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Location-Based Social Networks (LBSN) such as Foursquare allow users to indicate venue visits via check-ins. This results in much fine grained context-rich data, useful for studying user mobility. In this work, we use check-ins to characterize trips and visitors to two cities, where visitors are defined as having their home cities elsewhere. First, we divide trips into two duration types: long and short. We then show that trip types differ in check-in distributions over venue categories, time slots, as well as check-in intensity. Based on the trip types, we then divide visitors into long-term and short-term visitors. We compare visitor …


Modelling Cascades Over Time In Microblogs, Xie Wei, Feida Zhu, Siyuan Liu, Ke Wang Nov 2015

Modelling Cascades Over Time In Microblogs, Xie Wei, Feida Zhu, Siyuan Liu, Ke Wang

Research Collection School Of Computing and Information Systems

One of the most important features of microblogging services such as Twitter is how easy it is to re-share a piece of information across the network through various user connections, forming what we call a "cascade". Business applications such as viral marketing have driven a tremendous amount of research effort predicting whether a certain cascade will go viral. Yet the rarity of viral cascades in real data poses a challenge to all existing prediction methods. One solution is to simulate cascades that well fit the real viral ones, which requires our ability to tell how a certain cascade grows over …


Analysis Of Aspects And Star Ratings In Consumer Reviews, Maruthi Prithivirajan, Vivian Lai, Kyong Jin Shim Nov 2015

Analysis Of Aspects And Star Ratings In Consumer Reviews, Maruthi Prithivirajan, Vivian Lai, Kyong Jin Shim

Research Collection School Of Computing and Information Systems

This paper presents an analysis of star ratings in consumer reviews in Yelp, an online social platform for sharing consumer reviews about local businesses. In particular, we analyze consumer reviews about food businesses. We analyze how well or poorly the star ratings (on a scale of one star to five stars) associated with these reviews tally with the sentiment derived from the textual portion of the consumer review.


Dictionary Pair Learning On Grassmann Manifolds For Image Denoising, Xianhua Zeng, Wei Bian, Wei Liu, Jialie Shen, Dacheng Tao Nov 2015

Dictionary Pair Learning On Grassmann Manifolds For Image Denoising, Xianhua Zeng, Wei Bian, Wei Liu, Jialie Shen, Dacheng Tao

Research Collection School Of Computing and Information Systems

Image denoising is a fundamental problem in computer vision and image processing that holds considerable practical importance for real-world applications. The traditional patch-based and sparse coding-driven image denoising methods convert 2D image patches into 1D vectors for further processing. Thus, these methods inevitably break down the inherent 2D geometric structure of natural images. To overcome this limitation pertaining to the previous image denoising methods, we propose a 2D image denoising model, namely, the dictionary pair learning (DPL) model, and we design a corresponding algorithm called the DPL on the Grassmann-manifold (DPLG) algorithm. The DPLG algorithm first learns an initial dictionary …


Where Are The Passengers? A Grid-Based Gaussian Mixture Model For Taxi Bookings, Meng-Fen Chiang, Tuan Anh Hoang, Ee-Peng Lim Nov 2015

Where Are The Passengers? A Grid-Based Gaussian Mixture Model For Taxi Bookings, Meng-Fen Chiang, Tuan Anh Hoang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Taxi bookings are events where requests for taxis are made by passengers either over voice calls or mobile apps. As the demand for taxis changes with space and time, it is important to model both the space and temporal dimensions in dynamic booking data. Several applications can benefit from a good taxi booking model. These include the prediction of number of bookings at certain location and time of the day, and the detection of anomalous booking events. In this paper, we propose a Grid-based Gaussian Mixture Model (GGMM) with spatio-temporal dimensions that groups booking data into a number of spatio-temporal …


Using Digital Genomics To Create An Intelligent Enterprise, Mario Domingo Nov 2015

Using Digital Genomics To Create An Intelligent Enterprise, Mario Domingo

Asian Management Insights

Every business knows that it needs to leverage customer data, but few know the potential it has to transform business processes, decisions and performance.


A Method And System For Sentiment Classification And Emotion Classification [Us Patent 20170308523a1], Zhaoxia Wang, Rick Siow Mong Goh, Yinping Yang Nov 2015

A Method And System For Sentiment Classification And Emotion Classification [Us Patent 20170308523a1], Zhaoxia Wang, Rick Siow Mong Goh, Yinping Yang

Research Collection School Of Computing and Information Systems

A system and a method for classifying text messages, such as social media messages into sentiment valence categories are provided. The system comprising a module for decomposing text messages, a module for cleaning text messages, a module for producing feature data of text messages, and a module for classifying text messages into sentiment valence categories. The module for decomposing text messages is configured to: receive a text message, parse the text message into separate portions in response to parsing criteria based on sentence delimiters, wherein the separate portions are sentences, phrases and words, and rejoin at least some of the …


Face Recognition On Large-Scale Video In The Wild With Hybrid Euclidean-And-Riemannian Metric Learning, Zhiwu Huang, R. Wang, S. Shan, X Chen Oct 2015

Face Recognition On Large-Scale Video In The Wild With Hybrid Euclidean-And-Riemannian Metric Learning, Zhiwu Huang, R. Wang, S. Shan, X Chen

Research Collection School Of Computing and Information Systems

Face recognition on large-scale video in the wild is becoming increasingly important due to the ubiquity of video data captured by surveillance cameras, handheld devices, Internet uploads, and other sources. By treating each video as one image set, set-based methods recently have made great success in the field of video-based face recognition. In the wild world, videos often contain extremely complex data variations and thus pose a big challenge of set modeling for set-based methods. In this paper, we propose a novel Hybrid Euclidean-and-Riemannian Metric Learning (HERML) method to fuse multiple statistics of image set. Specifically, we represent each image …


Assessing Developer Contribution With Repository Mining-Based Metrics, Jalerson Lima, Christoph Treude, Fernando Figueira Filho, Uirá Kulesza Oct 2015

Assessing Developer Contribution With Repository Mining-Based Metrics, Jalerson Lima, Christoph Treude, Fernando Figueira Filho, Uirá Kulesza

Research Collection School Of Computing and Information Systems

Productivity as a result of individual developers' contributions is an important aspect for software companies to maintain their competitiveness in the market. However, there is no consensus in the literature on how to measure productivity or developer contribution. While some repository mining-based metrics have been proposed, they lack validation in terms of their applicability and usefulness from the individuals who will use them to assess developer contribution: team and project leaders. In this paper, we propose the design of a suite of metrics for the assessment of developer contribution, based on empirical evidence obtained from project and team leaders. In …


Scheduled Approximation For Personalized Pagerank With Utility-Based Hub Selection, Fanwei Zhu, Yuan Fang, Kevin Chen-Chuan Chang, Jing Ying Oct 2015

Scheduled Approximation For Personalized Pagerank With Utility-Based Hub Selection, Fanwei Zhu, Yuan Fang, Kevin Chen-Chuan Chang, Jing Ying

Research Collection School Of Computing and Information Systems

As Personalized PageRank has been widely leveraged for ranking on a graph, the efficient computation of Personalized PageRank Vector (PPV) becomes a prominent issue. In this paper, we propose FastPPV, an approximate PPV computation algorithm that is incremental and accuracy-aware. Our approach hinges on a novel paradigm of scheduled approximation: the computation is partitioned and scheduled for processing in an “organized” way, such that we can gradually improve our PPV estimation in an incremental manner and quantify the accuracy of our approximation at query time. Guided by this principle, we develop an efficient hub-based realization, where we adopt the metric …


Detect Rumors Using Time Series Of Social Context Information On Microblogging Websites, Jing Ma, Wei Gao, Zhongyu Wei, Yueming Lu, Kam-Fai Wong Oct 2015

Detect Rumors Using Time Series Of Social Context Information On Microblogging Websites, Jing Ma, Wei Gao, Zhongyu Wei, Yueming Lu, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

Automatically identifying rumors from online social media especially microblogging websites is an important research issue. Most of existing work for rumor detection focuses on modeling features related to microblog contents, users and propagation patterns, but ignore the importance of the variation of these social context features during the message propagation over time. In this study, we propose a novel approach to capture the temporal characteristics of these features based on the time series of rumor's lifecycle, for which time series modeling technique is applied to incorporate various social context information. Our experiments using the events in two microblog datasets confirm …


Choosing Your Weapons: On Sentiment Analysis Tools For Software Engineering Research, Robbert Jongeling, Subhajit Datta, Alexander Serebrenik Oct 2015

Choosing Your Weapons: On Sentiment Analysis Tools For Software Engineering Research, Robbert Jongeling, Subhajit Datta, Alexander Serebrenik

Research Collection School Of Computing and Information Systems

Recent years have seen an increasing attention to social aspects of software engineering, including studies of emotions and sentiments experienced and expressed by the software developers. Most of these studies reuse existing sentiment analysis tools such as SentiStrength and NLTK. However, these tools have been trained on product reviews and movie reviews and, therefore, their results might not be applicable in the software engineering domain. In this paper we study whether the sentiment analysis tools agree with the sentiment recognized by human evaluators (as reported in an earlier study) as well as with each other. Furthermore, we evaluate the impact …


The Importance Of Being Isolated: An Empirical Study On Chromium Reviews, Subhajit Datta, Devarshi Bhatt, Manish Jain, Proshanta Sarkar, Santonu Sarkar Oct 2015

The Importance Of Being Isolated: An Empirical Study On Chromium Reviews, Subhajit Datta, Devarshi Bhatt, Manish Jain, Proshanta Sarkar, Santonu Sarkar

Research Collection School Of Computing and Information Systems

As large scale software development has become more collaborative, and software teams more globally distributed, several studies have explored how developer interaction influences software development outcomes. The emphasis so far has been largely on outcomes like defect count, the time to close modification requests etc. In the paper, we examine data from the Chromium project to understand how different aspects of developer discussion relate to the closure time of reviews. On the basis of analyzing reviews discussed by 2000+ developers, our results indicate that quicker closure of reviews owned by a developer relates to higher reception of information and insights …


Social Tag Relevance Estimation Via Ranking-Oriented Neighbour Voting, Chaoran Cui, Jialie Shen, Jun Ma, Tao Lian Oct 2015

Social Tag Relevance Estimation Via Ranking-Oriented Neighbour Voting, Chaoran Cui, Jialie Shen, Jun Ma, Tao Lian

Research Collection School Of Computing and Information Systems

User-generated tags associated with social images are frequently imprecise and incomplete. Therefore, a fundamental challenge in tag-based applications is the problem of tag relevance estimation, which concerns how to interpret and quantify the relevance of a tag with respect to the contents of an image. In this paper, we address the key problem from a new perspective of learning to rank, and develop a novel approach to facilitate tag relevance estimation to directly optimize the ranking performance of tag-based image search. A supervision step is introduced into the neighbour voting scheme, in which tag relevance is estimated by accumulating votes …


Structural Constraints For Multipartite Entity Resolution With Markov Logic Network, Tengyuan Ye, Hady W. Lauw Oct 2015

Structural Constraints For Multipartite Entity Resolution With Markov Logic Network, Tengyuan Ye, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Multipartite entity resolution seeks to match entity mentions across several collections. An entity mention is presumed unique within a collection, and thus could match at most one entity mention in each of the other collections. In addition to domain-specific features considered in entity resolution, there are a number of domain-invariant structural contraints that apply in this scenario, including one-to-one assignment as well as cross-collection transitivity. We propose a principled solution to the multipartite entity resolution problem, building on the foundation of Markov Logic Network (MLN) that combines probabilistic graphical model and first-order logic. We describe how the domain-invariant structural constraints …


Learning Relative Similarity From Data Streams: Active Online Learning Approaches, Shuji Hao, Peilin Zhao, Steven C. H. Hoi, Chunyan Miao Oct 2015

Learning Relative Similarity From Data Streams: Active Online Learning Approaches, Shuji Hao, Peilin Zhao, Steven C. H. Hoi, Chunyan Miao

Research Collection School Of Computing and Information Systems

Relative similarity learning, as an important learning scheme for information retrieval, aims to learn a bi-linear similarity function from a collection of labeled instance-pairs, and the learned function would assign a high similarity value for a similar instance-pair and a low value for a dissimilar pair. Existing algorithms usually assume the labels of all the pairs in data streams are always made available for learning. However, this is not always realistic in practice since the number of possible pairs is quadratic to the number of instances in the database, and manually labeling the pairs could be very costly and time …


On Robust Image Spam Filtering Via Comprehensive Visual Modeling, Jialie Shen, Deng, Robert H., Zhiyong Cheng, Liqiang Nie, Shuicheng Yan Oct 2015

On Robust Image Spam Filtering Via Comprehensive Visual Modeling, Jialie Shen, Deng, Robert H., Zhiyong Cheng, Liqiang Nie, Shuicheng Yan

Research Collection School Of Computing and Information Systems

The Internet has brought about fundamental changes in the way peoples generate and exchange media information. Over the last decade, unsolicited message images (image spams) have become one of the most serious problems for Internet service providers (ISPs), business firms and general end users. In this paper, we report a novel system called RoBoTs (Robust BoosTrap based spam detector) to support accurate and robust image spam filtering. The system is developed based on multiple visual properties extracted from different levels of granularity, aiming to capture more discriminative contents for effective spam image identification. In addition, a resampling based learning framework …


Two Formulas For Success In Social Media: Learning And Network Effects, Liangfei Qiu, Qian Tang, Andrew B. Whinston Oct 2015

Two Formulas For Success In Social Media: Learning And Network Effects, Liangfei Qiu, Qian Tang, Andrew B. Whinston

Research Collection School Of Computing and Information Systems

Recent years have witnessed an unprecedented explosion in information technology that enables dynamic diffusion of user-generated content in social networks. Online videos, in particular, have changed the landscape of marketing and entertainment, competing with premium content and spurring business innovations. In the present study, we examine how learning and network effects drive the diffusion of online videos. While learning happens through informational externalities, network effects are direct payoff externalities. Using a unique data set from YouTube, we empirically identify learning and network effects separately, and find that both mechanisms have statistically and economically significant effects on video views; furthermore, the …


Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong Sep 2015

Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

A microblog repost tree provides strong clues on how an event described therein develops. To help social media users capture the main clues of events on microblogging sites, we propose a novel repost tree summarization framework by effectively differentiating two kinds of messages on repost trees called leaders and followers, which are derived from contentlevel structure information, i.e., contents of messages and the reposting relations. To this end, Conditional Random Fields (CRF) model is used to detect leaders across repost tree paths. We then present a variant of random-walk-based summarization model to rank and select salient messages based on the …


Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong Sep 2015

Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

A microblog repost tree provides strong clues on how an event described therein develops. To help social media users capture the main clues of events on microblogging sites, we propose a novel repost tree summarization framework by effectively differentiating two kinds of messages on repost trees called leaders and followers, which are derived from contentlevel structure information, i.e., contents of messages and the reposting relations. To this end, Conditional Random Fields (CRF) model is used to detect leaders across repost tree paths. We then present a variant of random-walk-based summarization model to rank and select salient messages based on the …


Name List Only? Target Entity Disambiguation In Short Texts, Yixin Cao, Juanzi Li, Xiaofei Guo, Shuanhu Bai, Heng Ji, Jie Tang Sep 2015

Name List Only? Target Entity Disambiguation In Short Texts, Yixin Cao, Juanzi Li, Xiaofei Guo, Shuanhu Bai, Heng Ji, Jie Tang

Research Collection School Of Computing and Information Systems

Target entity disambiguation (TED), the task of identifying target entities of the same domain, has been recognized as a critical step in various important applications. In this paper, we propose a graphbased model called TremenRank to collectively identify target entities in short texts given a name list only. TremenRank propagates trust within the graph, allowing for an arbitrary number of target entities and texts using inverted index technology. Furthermore, we design a multi-layer directed graph to assign different trust levels to short texts for better performance. The experimental results demonstrate that our model outperforms state-of-the-art methods with an average gain …


Real-Time Targeted Influence Maximization For Online Advertisements, Yuchen Li, Dongxiang Zhang, Kian-Lee Tan Sep 2015

Real-Time Targeted Influence Maximization For Online Advertisements, Yuchen Li, Dongxiang Zhang, Kian-Lee Tan

Research Collection School Of Computing and Information Systems

Advertising in social network has become a multi-billion dollar industry. A main challenge is to identify key influencers who can effectively contribute to the dissemination of information. Although the influence maximization problem, which finds a seed set of k most influential users based on certain propagation models, has been well studied, it is not target-aware and cannot be directly applied to online advertising. In this paper, we propose a new problem, named Keyword-Based Targeted Influence Maximization (KB-TIM), to find a seed set that maximizes the expected influence over users who are relevant to a given advertisement. To solve the problem, …


Mining Revenue-Maximizing Bundling Configuration, Loc Do, Hady Wirawan Lauw, Ke Wang Sep 2015

Mining Revenue-Maximizing Bundling Configuration, Loc Do, Hady Wirawan Lauw, Ke Wang

Research Collection School Of Computing and Information Systems

With greater prevalence of social media, there is an increasing amount of user-generated data revealing consumer preferences for various products and services. Businesses seek to harness this wealth of data to improve their marketing strategies. Bundling, or selling two or more items for one price is a highly-practiced marketing strategy. In this paper, we address the bundle configuration problem from the data-driven perspective. Given a set of items in a seller’s inventory, we seek to determine which items should belong to which bundle so as to maximize the total revenue, by mining consumer preferences data. We show that this problem …


Cobweb: A Robust Map Update System Using Gps Trajectories, Zhangqing Shan, Hao Wu, Weiwei Sun, Baihua Zheng Sep 2015

Cobweb: A Robust Map Update System Using Gps Trajectories, Zhangqing Shan, Hao Wu, Weiwei Sun, Baihua Zheng

Research Collection School Of Computing and Information Systems

The accuracy and completeness of a digital map plays a critical role in determining the quality of most location-based services. Unfortunately, road networks change frequently. Consequently, we study the issue of automatic map update in this paper. We propose a system called COBWEB which takes all the unmatched trajectories as input and generates the missing road segments with both the geometry properties and topology features well preserved. We conduct a comprehensive experimental study via real trajectory data generated by roughly 15,000 taxis in Singapore within a 5-month period. Compared with existing work, COBWEB demonstrates a better and more stable performance …


Maximum Rank Query, Kyriakos Mouratidis, Jilian Zhang, Hwee Hwa Pang Sep 2015

Maximum Rank Query, Kyriakos Mouratidis, Jilian Zhang, Hwee Hwa Pang

Research Collection School Of Computing and Information Systems

The top-k query is a common means to shortlist a number of options from a set of alternatives, based on the user's preferences. Typically, these preferences are expressed as a vector of query weights, defined over the options' attributes. The query vector implicitly associates each alternative with a numeric score, and thus imposes a ranking among them. The top-k result includes the k options with the highest scores. In this context, we define the maximum rank query (MaxRank). Given a focal option in a set of alternatives, the MaxRank problem is to compute the highest rank this option may achieve …


Tagcombine: Recommending Tags To Contents In Software Information Sites, Xin Yu Wang, Xin Xia, David Lo Sep 2015

Tagcombine: Recommending Tags To Contents In Software Information Sites, Xin Yu Wang, Xin Xia, David Lo

Research Collection School Of Computing and Information Systems

Nowadays, software engineers use a variety of online media to search and become informed of new and interesting technologies, and to learn from and help one another. We refer to these kinds of online media which help software engineers improve their performance in software development, maintenance, and test processes as software information sites. In this paper, we propose TagCombine, an automatic tag recommendation method which analyzes objects in software information sites. TagCombine has three different components: 1) multi-label ranking component which considers tag recommendation as a multi-label learning problem; 2) similarity-based ranking component which recommends tags from similar objects; 3) …


Answering Why-Not Questions On Reverse Top-K Queries, Yunjun Gao, Qing Liu, Gang Chen, Baihua Zheng, Linlin Zhou Sep 2015

Answering Why-Not Questions On Reverse Top-K Queries, Yunjun Gao, Qing Liu, Gang Chen, Baihua Zheng, Linlin Zhou

Research Collection School Of Computing and Information Systems

Why-not questions, which aim to seek clarifications on the missing tuples for query results, have recently received considerable attention from the database community. In this paper, we systematically explore why-not questions on reverse top-k queries, owing to its importance in multi-criteria decision making. Given an initial reverse top-k query and a missing/why-not weighting vector set Wm that is absent from the query result, why-not questions on reverse top-k queries explain why Wm does not appear in the query result and provide suggestions on how to refine the initial query with minimum penalty to include Wm in the refined query result. …


Towards Opinion Summarization From Online Forums, Ding Ying, Jing Jiang Sep 2015

Towards Opinion Summarization From Online Forums, Ding Ying, Jing Jiang

Research Collection School Of Computing and Information Systems

Summarizing opinions expressed in online forums can potentially benefit many people. However, special characteristics of this problem may require changes to standard text summarization techniques. In this work, we present our initial attempt at extractive summarization of opinionated online forum threads. Given the nature of user generated content in online discussion forums, we hypothesize that besides relevance, text quality and subjectivity also play important roles in deciding which sentences are good summary sentences. We therefore construct an annotated corpus to facilitate our study of extractive summarization of online discussion forums. We define a set of features to capture relevance, text …


A Joint Model Of Product Properties, Aspects And Ratings For Online Reviews, Ding Ying, Jing Jiang Sep 2015

A Joint Model Of Product Properties, Aspects And Ratings For Online Reviews, Ding Ying, Jing Jiang

Research Collection School Of Computing and Information Systems

Product review mining is an important task that can benefit both businesses and consumers. Lately a number of models combining collaborative filtering and content analysis to model reviews have been proposed, among which the Hidden Factors as Topics (HFT) model is a notable one. In this work, we propose a new model on top of HFT to separate product properties and aspects. Product properties are intrinsic to certain products (e.g. types of cuisines of restaurants) whereas aspects are dimensions along which products in the same category can be compared (e.g. service quality of restaurants). Our proposed model explicitly separates the …