Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year

Articles 421 - 450 of 1024

Full-Text Articles in Numerical Analysis and Scientific Computing

Collaborative Topic Regression With Denoising Autoencoder For Content And Community Co-Representation, Trong T. Nguyen, Hady W. Lauw Nov 2017

Collaborative Topic Regression With Denoising Autoencoder For Content And Community Co-Representation, Trong T. Nguyen, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Personalized recommendation of items frequently faces scenarios where we have sparse observations on users' adoption of items. In the literature, there are two promising directions. One is to connect sparse items through similarity in content. The other is to connect sparse users through similarity in social relations. We seek to integrate both types of information, in addition to the adoption information, within a single integrated model. Our proposed method models item content via a topic model, and user communities via an autoencoder model, while bridging a user's community-based preference to her topic-based preference. Experiments on public real-life data showcase the …


Indexable Bayesian Personalized Ranking For Efficient Top-K Recommendation, Dung D. Le, Hady W. Lauw Nov 2017

Indexable Bayesian Personalized Ranking For Efficient Top-K Recommendation, Dung D. Le, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Top-k recommendation seeks to deliver a personalized recommendation list of k items to a user. The dual objectives are (1) accuracy in identifying the items a user is likely to prefer, and (2) efficiency in constructing the recommendation list in real time. One direction towards retrieval efficiency is to formulate retrieval as approximate k nearest neighbor (kNN) search aided by indexing schemes, such as locality-sensitive hashing, spatial trees, and inverted index. These schemes, applied on the output representations of recommendation algorithms, speed up the retrieval process by automatically discarding a large number of potentially irrelevant items when given a user …


Answerbot: Automated Generation Of Answer Summary To Developers’ Technical Questions, Bowen Xu, Zhenchang Xing, Xin Xia, David Lo Nov 2017

Answerbot: Automated Generation Of Answer Summary To Developers’ Technical Questions, Bowen Xu, Zhenchang Xing, Xin Xia, David Lo

Research Collection School Of Computing and Information Systems

The prevalence of questions and answers on domain-specific Q&A sites like Stack Overflow constitutes a core knowledge asset for software engineering domain. Although search engines can return a list of questions relevant to a user query of some technical question, the abundance of relevant posts and the sheer amount of information in them makes it difficult for developers to digest them and find the most needed answers to their questions. In this work, we aim to help developers who want to quickly capture the key points of several answer posts relevant to a technical question before they read the details …


A Fast Trajectory Outlier Detection Approach Via Driving Behavior Modeling, Hao Wu, Weiwei Sun, Baihua Zheng Nov 2017

A Fast Trajectory Outlier Detection Approach Via Driving Behavior Modeling, Hao Wu, Weiwei Sun, Baihua Zheng

Research Collection School Of Computing and Information Systems

Trajectory outlier detection is a fundamental building block for many location-based service (LBS) applications, with a large application base. We dedicate this paper on detecting the outliers from vehicle trajectories efficiently and effectively. In addition, we want our solution to be able to issue an alarm early when an outlier trajectory is only partially observed (i.e., the trajectory has not yet reached the destination). Most existing works study the problem on general Euclidean trajectories and require accesses to the historical trajectory database or computations on the distance metric that are very expensive. Furthermore, few of existing works consider some specific …


Visual Sentiment Analysis For Review Images With Item-Oriented And User-Oriented Cnn, Quoc Tuan Truong, Hady W. Lauw Oct 2017

Visual Sentiment Analysis For Review Images With Item-Oriented And User-Oriented Cnn, Quoc Tuan Truong, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Online reviews are prevalent. When recounting their experience with a product, service, or venue, in addition to textual narration, a reviewer frequently includes images as photographic record. While textual sentiment analysis has been widely studied, in this paper we are interested in visual sentiment analysis to infer whether a given image included as part of a review expresses the overall positive or negative sentiment of that review. Visual sentiment analysis can be formulated as image classification using deep learning methods such as Convolutional Neural Networks or CNN. However, we observe that the sentiment captured within an image may be affected …


On Negative Results When Using Sentiment Analysis Tools For Software Engineering Research, Robbert Jongeling, Proshanta Sarkar, Subhajit Datta, Alexander Serebrenik Oct 2017

On Negative Results When Using Sentiment Analysis Tools For Software Engineering Research, Robbert Jongeling, Proshanta Sarkar, Subhajit Datta, Alexander Serebrenik

Research Collection School Of Computing and Information Systems

Recent years have seen an increasing attention to social aspects of software engineering, including studies of emotions and sentiments experienced and expressed by the software developers. Most of these studies reuse existing sentiment analysis tools such as SentiStrength and NLTK. However, these tools have been trained on product reviews and movie reviews and, therefore, their results might not be applicable in the software engineering domain. In this paper we study whether the sentiment analysis tools agree with the sentiment recognized by human evaluators (as reported in an earlier study) as well as with each other. Furthermore, we evaluate the impact …


A Conceptual Framework For Analyzing Students' Feedback, Venky Shankararaman, Swapna Gottipati, Sandy Gan Oct 2017

A Conceptual Framework For Analyzing Students' Feedback, Venky Shankararaman, Swapna Gottipati, Sandy Gan

Research Collection School Of Computing and Information Systems

In academic institutions it is normal practice that at the end of each term,students are required to complete a questionnaire that is designed to gather students’perceptions of the instructor and their learning experience in the course. This questionnaire comprises of Likert-scale questions and qualitative questions.One of the important goals of this exercise is to enable the instructor and the senior management to examine the feedback and then enhance students’ learning experience. In most universities, including our own, a lot of attention is paid to the quantitative feedback, which is summarized and statistical comparisons are computed, analysed and presented. However, the …


Graphh: High Performance Big Graph Analytics In Small Clusters, Peng Sun, Yonggang Wen, Nguyen Binh Duong Ta, Xiaokui Xiao Sep 2017

Graphh: High Performance Big Graph Analytics In Small Clusters, Peng Sun, Yonggang Wen, Nguyen Binh Duong Ta, Xiaokui Xiao

Research Collection School Of Computing and Information Systems

It is common for real-world applications to analyze big graphs using distributed graph processing systems. Popular in-memory systems require an enormous amount of resources to handle big graphs. While several out-of-core approaches have been proposed for processing big graphs on disk, the high disk I/O overhead could significantly reduce performance. In this paper, we propose GraphH to enable highperformance big graph analytics in small clusters. Specifically, we design a two-stage graph partition scheme to evenly divide the input graph into partitions, and propose a GAB (GatherApply-Broadcast) computation model to make each worker process a partition in memory at a time. …


Modeling Trajectories With Recurrent Neural Networks, Hao Wu, Ziyang Chen, Weiwei Sun, Baihua Zheng, Wei Wang Aug 2017

Modeling Trajectories With Recurrent Neural Networks, Hao Wu, Ziyang Chen, Weiwei Sun, Baihua Zheng, Wei Wang

Research Collection School Of Computing and Information Systems

Modeling trajectory data is a building block for many smart-mobility initiatives. Existing approaches apply shallow models such as Markov chain and inverse reinforcement learning to model trajectories, which cannot capture the long-term dependencies. On the other hand, deep models such as Recurrent Neura lNetwork (RNN) have demonstrated their strength of modeling variable length sequences. However, directly adopting RNN to model trajectories is not appropriate because of the unique topological constraints faced by trajectories. Motivated by these findings, we design two RNN-based models which can make full advantage of the strength of RNN to capture variable length sequence and meanwhile to …


Deepfacade: A Deep Learning Approach To Facade Parsing, Hantang Liu, Jialiang Zhang, Jianke Zhu, Steven C. H. Hoi Aug 2017

Deepfacade: A Deep Learning Approach To Facade Parsing, Hantang Liu, Jialiang Zhang, Jianke Zhu, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

The parsing of building facades is a key component to the problem of 3D street scenes reconstruction, which is long desired in computer vision. In this paper, we propose a deep learning based method for segmenting a facade into semantic categories. Man-made structures often present the characteristic of symmetry. Based on this observation, we propose a symmetric regularizer for training the neural network. Our proposed method can make use of both the power of deep neural networks and the structure of man-made architectures. We also propose a method to refine the segmentation results using bounding boxes generated by the Region …


Basket-Sensitive Personalized Item Recommendation, Duc Trong Le, Hady W. Lauw, Yuan Fang Aug 2017

Basket-Sensitive Personalized Item Recommendation, Duc Trong Le, Hady W. Lauw, Yuan Fang

Research Collection School Of Computing and Information Systems

Personalized item recommendation is useful in narrowing down the list of options provided to a user. In this paper, we address the problem scenario where the user is currently holding a basket of items, and the task is to recommend an item to be added to the basket. Here, we assume that items currently in a basket share some association based on an underlying latent need, e.g., ingredients to prepare some dish, spare parts of some device. Thus, it is important that a recommended item is relevant not only to the user, but also to the existing items in the …


Generating Cultural Personas From Social Data: A Perspective Of Middle Eastern Users, Salminen Joni, Sercan Sengün, Haewoon Kwak, Bernard Jansen, Jisun An, Soon-Gyo Jung, Sarah Vieweg, D. Fox Harrell Aug 2017

Generating Cultural Personas From Social Data: A Perspective Of Middle Eastern Users, Salminen Joni, Sercan Sengün, Haewoon Kwak, Bernard Jansen, Jisun An, Soon-Gyo Jung, Sarah Vieweg, D. Fox Harrell

Research Collection School Of Computing and Information Systems

We conduct a mixed-method study to better understand the content consumption patterns of Middle Eastern social media users and to explore new ways to present online data by using automatic persona generation. First, we analyze millions of content interactions on YouTube to dynamically generate personas describing behavioral patterns of different demographic groups. Second, we analyze interview data on social media users in the Middle Eastern region to generate additional insights into the dynamically generated personas. Our findings provide insights into social media users in the Middle East, as well as present a novel methodology of using computational analysis and qualitative …


Multiplex Media Attention And Disregard Network Among 129 Countries, Haewoon Kwak, Jisun An Aug 2017

Multiplex Media Attention And Disregard Network Among 129 Countries, Haewoon Kwak, Jisun An

Research Collection School Of Computing and Information Systems

We built a multiplex media attention and disregard network (MADN) among 129 countries over 212 days. By characterizing the MADN from multiple levels, we found that it is formed primarily by skewed, hierarchical, and asymmetric relationships. Also, we found strong evidence that our news world is becoming a "global village." However, at the same time, unique attention blocks of the Middle East and North Africa (MENA) region, as well as Russia and its neighbors, still exist.


Personas For Content Creators Via Decomposed Aggregate Audience Statistics, Jisun An, Haewoon Kwak, Bernard J. Jansen Aug 2017

Personas For Content Creators Via Decomposed Aggregate Audience Statistics, Jisun An, Haewoon Kwak, Bernard J. Jansen

Research Collection School Of Computing and Information Systems

We propose a novel method for generating personas based on online user data for the increasingly common situation of content creators distributing products via online platforms. We use non-negative matrix factorization to identify user segments and develop personas by adding personality such as names and photos. Our approach can develop accurate personas representing real groups of people using online user data, versus relying on manually gathered data.


A Domain Based Approach To Social Relation Recognition, Qianru Sun, Bernt Schiele, Mario Fritz Jul 2017

A Domain Based Approach To Social Relation Recognition, Qianru Sun, Bernt Schiele, Mario Fritz

Research Collection School Of Computing and Information Systems

Social relations are the foundation of human daily life. Developing techniques to analyze such relations from visual data bears great potential to build machines that better understand us and are capable of interacting with us at a social level. Previous investigations have remained partial due to the overwhelming diversity and complexity of the topic and consequently have only focused on a handful of social relations. In this paper, we argue that the domain-based theory from social psychology is a great starting point to systematically approach this problem. The theory provides coverage of all aspects of social relations and equally is …


Demographics Of News Sharing In The U.S. Twittersphere, Julio C.S. Reis, Haewoon Kwak, Jisun An, Johnnatan Messias, Benevenuto Fabrıcio. Jul 2017

Demographics Of News Sharing In The U.S. Twittersphere, Julio C.S. Reis, Haewoon Kwak, Jisun An, Johnnatan Messias, Benevenuto Fabrıcio.

Research Collection School Of Computing and Information Systems

The widespread adoption and dissemination of online news through social media systems have been revolutionizing many segments of our society and ultimately our daily lives. In these systems, users can play a central role as they share content to their friends. Despite that, little is known about news spreaders in social media. In this paper, we provide the first of its kind in-depth characterization of news spreaders in social media. In particular, we investigate their demographics, what kind of content they share, and the audience they reach. Among our main findings, we show that males and white users tend to …


A Nash Equilibrium Formulation Of A Tradable Credits Scheme For Incentivizing Transport Choices: From Next-Generation Public Transport Mode Choice To Hot Lanes, Salem Lahlou, Laura Wynter Jul 2017

A Nash Equilibrium Formulation Of A Tradable Credits Scheme For Incentivizing Transport Choices: From Next-Generation Public Transport Mode Choice To Hot Lanes, Salem Lahlou, Laura Wynter

Research Collection School Of Computing and Information Systems

We consider a tradable credits scheme for binary transport games where one option is faster (or more comfortable) than the other, but its quality of service suffers when usage is high. Applications can be found in mode choice (public transit versus road transport), premium (i.e., express bus) versus ordinary public transit, and fast (e.g., high-occupancy toll, or HOT) versus regular lanes on expressways. We are motivated in particular by the choice between public transport and use of the road network as a privilege to be discouraged. In a future where GPS-based time-distance-place road charging exists, such next-generation transport management strategies …


A Weighted Maximum Matching Algorithm For Influence Maximization And Structural Controllability, Giorgio Sartor, Yeow Khiang Chia, Laura Wynter, Justin Ruths Jul 2017

A Weighted Maximum Matching Algorithm For Influence Maximization And Structural Controllability, Giorgio Sartor, Yeow Khiang Chia, Laura Wynter, Justin Ruths

Research Collection School Of Computing and Information Systems

Structural control and influence maximization on networks both admit the problem of selecting a particular subset of nodes. In structural control, the subset of nodes should guarantee the controllability of the network (in the usual sense) for almost any combination of weights. In influence maximization, given a diffusion process over the network, the chosen subset of nodes (of a given cardinality) should produce the greatest diffusive influence over the rest of the network. While structural control exploits only the structure of the network, influence maximization depends both on the structure and the weights of the edges. We modify an algorithm …


Compress: A Comprehensive Framework Of Trajectory Compression In Road Networks, Yunheng Han, Weiwei Sun, Baihua Zheng Jun 2017

Compress: A Comprehensive Framework Of Trajectory Compression In Road Networks, Yunheng Han, Weiwei Sun, Baihua Zheng

Research Collection School Of Computing and Information Systems

More and more advanced technologies have become available to collect and integrate an unprecedented amount of data from multiple sources, including GPS trajectories about the traces of moving objects. Given the fact that GPS trajectories are vast in size while the information carried by the trajectories could be redundant, we focus on trajectory compression in this article. As a systematic solution, we propose a comprehensive framework, namely, COMPRESS (Comprehensive Paralleled Road-Network-Based Trajectory Compression), to compress GPS trajectory data in an urban road network. In the preprocessing step, COMPRESS decomposes trajectories into spatial paths and temporal sequences, with a thorough justification …


On Self-Selection Biases In Online Product Reviews, Nan Hu, Paul A. Pavlou, Jie Zhang Jun 2017

On Self-Selection Biases In Online Product Reviews, Nan Hu, Paul A. Pavlou, Jie Zhang

Research Collection School Of Computing and Information Systems

Online product reviews help consumers infer product quality, and the mean (average) rating is often used as a proxy for product quality. However, two self-selection biases, acquisition bias (mostly consumers with a favorable predisposition acquire a product and hence write a product review) and underreporting bias (consumers with extreme, either positive or negative, ratings are more likely to write reviews than consumers with moderate product ratings), render the mean rating a biased estimator of product quality, and they result in the well-known J-shaped (positively skewed, asymmetric, bimodal) distribution of online product reviews. To better understand the nature and consequences of …


Continuous Top-K Monitoring On Document Streams, Leong Hou U, Junjie Zhang, Kyriakos Mouratidis, Ye Li May 2017

Continuous Top-K Monitoring On Document Streams, Leong Hou U, Junjie Zhang, Kyriakos Mouratidis, Ye Li

Research Collection School Of Computing and Information Systems

The efficient processing of document streams plays an important role in many information filtering systems. Emerging applications, such as news update filtering and social network notifications, demand presenting end-users with the most relevant content to their preferences. In this work, user preferences are indicated by a set of keywords. A central server monitors the document stream and continuously reports to each user the top-k documents that are most relevant to her keywords. Our objective is to support large numbers of users and high stream rates, while refreshing the top-k results almost instantaneously. Our solution abandons the traditional frequency-ordered indexing approach. …


Who Will Leave The Company?: A Large-Scale Industry Study Of Developer Turnover By Mining Monthly Work Report, Lingfeng Bao, Zhenchang Xing, Xin Xia, David Lo, Shanping Li May 2017

Who Will Leave The Company?: A Large-Scale Industry Study Of Developer Turnover By Mining Monthly Work Report, Lingfeng Bao, Zhenchang Xing, Xin Xia, David Lo, Shanping Li

Research Collection School Of Computing and Information Systems

Software developer turnover has become a big challenge for information technology (IT) companies. The departure of key software developers might cause big loss to an IT company since they also depart with important business knowledge and critical technical skills. Understanding developer turnover is very important for IT companies to retain talented developers and reduce the loss due to developers' departure. Previous studies mainly perform qualitative observations or simple statistical analysis of developers' activity data to understand developer turnover. In this paper, we investigate whether we can predict the turnover of software developers in non-open source companies by automatically analyzing monthly …


A Neural Network Model For Semi-Supervised Review Aspect Identification, Ying Ding, Changlong Yu, Jing Jiang May 2017

A Neural Network Model For Semi-Supervised Review Aspect Identification, Ying Ding, Changlong Yu, Jing Jiang

Research Collection School Of Computing and Information Systems

Aspect identification is an important problem in opinion mining. It is usually solved in an unsupervised manner, and topic models have been widely used for the task. In this work, we propose a neural network model to identify aspects from reviews by learning their distributional vectors. A key difference of our neural network model from topic models is that we do not use multinomial word distributions but instead embedding vectors to generate words. Furthermore, to leverage review sentences labeled with aspect words, a sequence labeler based on Recurrent Neural Networks (RNNs) is incorporated into our neural network. The resulting model …


Persona Generation From Aggregated Social Media Data, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Moeed Ahmad, Lene Nielsen, Bernard J. Jansen May 2017

Persona Generation From Aggregated Social Media Data, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Moeed Ahmad, Lene Nielsen, Bernard J. Jansen

Research Collection School Of Computing and Information Systems

We develop a methodology for persona generation using real time social media data for the distribution of products via online platforms. From a large social media account containing more than 30 million interactions from users from 181 countries engaging with more than 4,200 digital products produced by a global media corporation, we demonstrate that our methodology can first identify both distinct and impactful user segments and then create persona descriptions by automatically adding pertinent features, such as names, photos, and personal attributes. We validate our approach by implementing the methodology into an actual working system that leverages large scale online …


Predicting The Impact Of Software Engineering Topics: An Empirical Study, Santonu Sarkar, Rumana Lakdawala, Subhajit Datta Apr 2017

Predicting The Impact Of Software Engineering Topics: An Empirical Study, Santonu Sarkar, Rumana Lakdawala, Subhajit Datta

Research Collection School Of Computing and Information Systems

Predicting the future is hard, more so in active research areas. In this paper, we customize an established model for citation prediction of research papers and apply it on research topics. We argue that research topics, rather than individual publications, have wider relevance in the research ecosystem, for individuals as well as organizations. In this study, topics are extracted from a corpus of software engineering publications covering 55,000+ papers written by more than 70,000 authors across 56 publication venues, over a span of 38 years, using natural language processing techniques. We demonstrate how critical aspects of the original paper-based prediction …


Modeling Topics And Behavior Of Microbloggers: An Integrated Approach, Tuan Anh Hoang, Ee-Peng Lim Apr 2017

Modeling Topics And Behavior Of Microbloggers: An Integrated Approach, Tuan Anh Hoang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Microblogging encompasses both user-generated content and behavior. When modeling microblogging data, one has to consider personal and background topics, as well as how these topics generate the observed content and behavior. In this article, we propose the Generalized Behavior-Topic (GBT) model for simultaneously modeling background topics and users' topical interest in microblogging data. GBT considers multiple topical communities (or realms) with different background topical interests while learning the personal topics of each user and the user's dependence on realms to generate both content and behavior. This differentiates GBT from other previous works that consider either one realm only or content …


Assessing The Language Of Chat For Teamwork Dialogue, Antonette Shibani, Elizabeth Koh, Vivian Lai, Kyong Jin Shim Apr 2017

Assessing The Language Of Chat For Teamwork Dialogue, Antonette Shibani, Elizabeth Koh, Vivian Lai, Kyong Jin Shim

Research Collection School Of Computing and Information Systems

In technology enhanced language learning, many pedagogical activities involve students in online discussion such as synchronous chat, in order to help them practice their language skills. Besides developing the language competency of students, it is also crucial to nurture their teamwork competencies for today's global and complex environment. Language communication is an important glue of teamwork. In order to assess the language of chat for teamwork dimensions, several text mining methods are pos sible. However, difficulties arise such as pre-processing being a black box and classification approaches and algorithms being dependent on the context. To address these issues, the study …


Characterizing Malicious Android Apps By Mining Topic-Specific Data Flow Signatures, Xinli Yang, David Lo, Li Li, Xin Xia, Tegawendé F. Bissyande, Jacques Klein Apr 2017

Characterizing Malicious Android Apps By Mining Topic-Specific Data Flow Signatures, Xinli Yang, David Lo, Li Li, Xin Xia, Tegawendé F. Bissyande, Jacques Klein

Research Collection School Of Computing and Information Systems

Context: State-of-the-art works on automated detection of Android malware have leveraged app descriptions to spot anomalies w.r.t the functionality implemented, or have used data flow information as a feature to discriminate malicious from benign apps. Although these works have yielded promising performance,we hypothesize that these performances can be improved by a better understanding of malicious behavior. Objective: To characterize malicious apps, we take into account both information on app descriptions,which are indicative of apps’ topics, and information on sensitive data flow, which can be relevant todiscriminate malware from benign apps. Method: In this paper, we propose a topic-specific approach to …


Comparative Relation Generative Model, Maksim Tkachenko, Hady W. Lauw Apr 2017

Comparative Relation Generative Model, Maksim Tkachenko, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Online reviews are important decision aids to consumers. Other than helping users to evaluate individual products, reviews also support comparison shopping by comparing two (or more) products based on a specific aspect. However, making a comparison across two different reviews, written by different authors, is not always equitable due to the different standards and preferences of authors. Therefore, we focus on comparative sentences, whereby two products are compared directly by a review author within a sentence. We study the problem of comparative relation mining. Given a set of comparative sentences, each relating a pair of entities, our objective is three-fold: …


Inferring User Consumption Preferences From Social Media, Yang Li, Jing Jiang, Ting Liu Mar 2017

Inferring User Consumption Preferences From Social Media, Yang Li, Jing Jiang, Ting Liu

Research Collection School Of Computing and Information Systems

Social Media has already become a new arena of our lives and involved different aspects of our social presence. Users' personal information and activities on social media presumably reveal their personal interests, which offer great opportunities for many e-commerce applications. In this paper, we propose a principled latent variable model to infer user consumption preferences at the category level (e.g. inferring what categories of products a user would like to buy). Our model naturally links users' published content and following relations on microblogs with their consumption behaviors on e-commerce websites. Experimental results show our model outperforms the state-of-the-art methods significantly …