Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 421 - 450 of 808

Full-Text Articles in Numerical Analysis and Scientific Computing

Real Time Event Detection In Twitter, Xun Wang, Feida Zhu, Jing Jiang, Sujian Li Jun 2013

Real Time Event Detection In Twitter, Xun Wang, Feida Zhu, Jing Jiang, Sujian Li

Research Collection School Of Computing and Information Systems

Event detection has been an important task for a long time. When it comes to Twitter, new problems are presented. Twitter data is a huge temporal data flow with much noise and various kinds of topics. Traditional sophisticated methods with a high computational complexity aren’t designed to handle such data flow efficiently. In this paper, we propose a mixture Gaussian model for bursty word extraction in Twitter and then employ a novel time-dependent HDP model for new topic detection. Our model can grasp new events, the location and the time an event becomes bursty promptly and accurately. Experiments show the …


A Latent Variable Model For Viewpoint Discovery From Threaded Forum Posts, Minghui Qiu, Jing Jiang Jun 2013

A Latent Variable Model For Viewpoint Discovery From Threaded Forum Posts, Minghui Qiu, Jing Jiang

Research Collection School Of Computing and Information Systems

Threaded discussion forums provide an important social media platform. Its rich user generated content has served as an important source of public feedback. To automatically discover the viewpoints or stances on hot issues from forum threads is an important and useful task. In this paper, we propose a novel latent variable model for viewpoint discovery from threaded forum posts. Our model is a principled generative latent variable model which captures three important factors: viewpoint specific topic preference, user identity and user interactions. Evaluation results show that our model clearly outperforms a number of baseline models in terms of both clustering …


Mining User Relations From Online Discussions Using Sentiment Analysis And Probabilistic Matrix Factorization, Minghui Qiu, Liu Yang, Jing Jiang Jun 2013

Mining User Relations From Online Discussions Using Sentiment Analysis And Probabilistic Matrix Factorization, Minghui Qiu, Liu Yang, Jing Jiang

Research Collection School Of Computing and Information Systems

Advances in sentiment analysis have enabled extraction of user relations implied in online textual exchanges such as forum posts. However, recent studies in this direction only consider direct relation extraction from text. As user interactions can be sparse in online discussions, we propose to apply collaborative filtering through probabilistic matrix factorization to generalize and improve the opinion matrices extracted from forum posts. Experiments with two tasks show that the learned latent factor representation can give good performance on a relation polarity prediction task and improve the performance of a subgroup detection task.


R-Energy For Evaluating Robustness Of Dynamic Networks, Ming Gao, Ee Peng Lim, David Lo May 2013

R-Energy For Evaluating Robustness Of Dynamic Networks, Ming Gao, Ee Peng Lim, David Lo

Research Collection School Of Computing and Information Systems

The robustness of a network is determined by how well its vertices are connected to one another so as to keep the network strong and sustainable. As the network evolves its robustness changes and may reveal events as well as periodic trend patterns that affect the interactions among users in the network. In this paper, we develop R-energy as a new measure of network robustness based on the spectral analysis of normalized Laplacian matrix. R-energy can cope with disconnected networks, and is efficient to compute with a time complexity of O (jV j + jEj) where V and E are …


Fans: Face Annotation By Searching Large-Scale Web Facial Images, Steven Hoi, Dayong Wang, I Yeu Cheng, Elmer Lin, Jianke Zhu, Ying He, Chunyan Miao May 2013

Fans: Face Annotation By Searching Large-Scale Web Facial Images, Steven Hoi, Dayong Wang, I Yeu Cheng, Elmer Lin, Jianke Zhu, Ying He, Chunyan Miao

Research Collection School Of Computing and Information Systems

Auto face annotation is an important technique for many real-world applications, such as online photo album management, new video summarization, and so on. It aims to automatically detect human faces from a photo image and further name the faces with the corresponding human names. Recently, mining web facial images on the internet has emerged as a promising paradigm towards auto face annotation. In this paper, we present a demonstration system of search-based face annotation: FANS - Face ANnotation by Searching large-scale web facial images. Given a query facial image for annotation, we first retrieve a short list of the most …


Your Love Is Public Now: Questioning The Use Of Personal Information In Authentication, Payas Gupta, Swapna Gottipati, Jing Jiang, Debin Gao May 2013

Your Love Is Public Now: Questioning The Use Of Personal Information In Authentication, Payas Gupta, Swapna Gottipati, Jing Jiang, Debin Gao

Research Collection School Of Computing and Information Systems

Most social networking platforms protect user's private information by limiting access to it to a small group of members, typically friends of the user, while allowing (virtually) everyone's access to the user's public data. In this paper, we exploit public data available on Facebook to infer users' undisclosed interests on their profile pages. In particular, we infer their undisclosed interests from the public data fetched using Graph APIs provided by Facebook. We demonstrate that simply liking a Facebook page does not corroborate that the user is interested in the page. Instead, we perform sentiment-oriented mining on various attributes of a …


It Is Not Just What We Say, But How We Say Them: Lda-Based Behavior-Topic Model, Minghui Qiu, Feida Zhu, Jing Jiang May 2013

It Is Not Just What We Say, But How We Say Them: Lda-Based Behavior-Topic Model, Minghui Qiu, Feida Zhu, Jing Jiang

Research Collection School Of Computing and Information Systems

Textual information exchanged among users on online social network platforms provides deep understanding into users' interest and behavioral patterns. However, unlike traditional text-dominant settings such as o ine publishing, one distinct feature for online social network is users' rich interactions with the textual content, which, unfortunately, has not yet been well incorporated in the existing topic modeling frameworks. In this paper, we propose an LDA-based behavior-topic model (B-LDA) which jointly models user topic interests and behavioral patterns. We focus the study of the model on online social network settings such as microblogs like Twitter where the textual content is relatively …


Retweeting: An Act Of Viral Users, Susceptible Users, Or Viral Topics?, Tuan-Anh Hoang, Ee Peng Lim May 2013

Retweeting: An Act Of Viral Users, Susceptible Users, Or Viral Topics?, Tuan-Anh Hoang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

When a user retweets, there are three behavioral factors that cause the actions. They are the topic virality, user virality and user susceptibility. Topic virality captures the degree to which a topic attracts retweets by users. For each topic, user virality and susceptibility refer to the likelihood that a user attracts retweets and performs retweeting respectively. To model a set of observed retweet data as a result of these three topic specific factors, we first represent the retweets as a three-dimensional tensor of the tweet authors, their followers, and the tweets themselves. We then propose the V 2S model, a …


Vistruclizer: A Structural Visualizer For Multi-Dimensional Social Networks, Bingtian Dai, Agus Trisnajaya Kwee, Ee Peng Lim Apr 2013

Vistruclizer: A Structural Visualizer For Multi-Dimensional Social Networks, Bingtian Dai, Agus Trisnajaya Kwee, Ee Peng Lim

Research Collection School Of Computing and Information Systems

With the popularity of Web 2.0 sites, social networks today increasingly involve different kinds of relationships among different types of users in a single network. Such social networks are said to be multi-dimensional. Analyzing multi-dimensional networks is a challenging research task that requires intelligent visualization techniques. In this paper, we therefore propose a visual analytics tool called ViStruclizer to analyze structures embedded in a multi-dimensional social network. ViStruclizer incorporates structure analyzers that summarize social networks into both node clusters each representing a set of users, and edge clusters representing relationships between users in the node clusters. ViStruclizer supports user interactions …


Roundtriprank: Graph-Based Proximity With Importance And Specificity, Yuan Fang, Kevin Chen-Chuan Chang, Hady W. Lauw Apr 2013

Roundtriprank: Graph-Based Proximity With Importance And Specificity, Yuan Fang, Kevin Chen-Chuan Chang, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Graph-based proximity has many applications with different ranking needs. However, most previous works only stress the sense of importance by finding "popular” results for a query. Often times important results are overly general without being well-tailored to the query, lacking a sense of specificity— which only emerges recently. Even then, the two senses are treated independently, and only combined empirically. In this paper, we generalize the well-studied importance-based random walk into a round trip and develop RoundTripRank, seamlessly integrating specificity and importance in one coherent process. We also recognize the need for a flexible trade-off between the two senses, and …


Dynamic Label Propagation In Social Networks, Juan Du, Feida Zhu, Ee Peng Lim Apr 2013

Dynamic Label Propagation In Social Networks, Juan Du, Feida Zhu, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Label propagation has been studied for many years, starting from a set of nodes with labels and then propagating to those without labels. In social networks, building complete user profiles like interests and affiliations contributes to the systems like link prediction, personalized feeding, etc. Since the labels for each user are mostly not filled, we often employ some people to label these users. And therefore, the cost of human labeling is high if the data set is large. To reduce the expense, we need to select the optimal data set for labeling, which produces the best propagation result. In this …


Finding The Optimal Social Trust Path For The Selection Of Trustworthy Service Providers In Complex Social Networks, Guanfeng Liu, Yan Wang, Mehmet A. Orgun, Ee Peng Lim Apr 2013

Finding The Optimal Social Trust Path For The Selection Of Trustworthy Service Providers In Complex Social Networks, Guanfeng Liu, Yan Wang, Mehmet A. Orgun, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Online social networks have provided the infrastructure for a number of emerging applications in recent years, e.g., for the recommendation of service providers or the recommendation of files as services. In these applications, trust is one of the most important factors in decision making by a service consumer, requiring the evaluation of the trustworthiness of a service provider along the social trust paths from a service consumer to the service provider. However, there are usually many social trust paths between two participants who are unknown to one another. In addition, some social information, such as social relationships between participants and …


Image Collection Summarization Via Dictionary Learning For Sparse Representation, Chunlei Yang, Jialie Shen, Jinye Peng, Jianping Fan Mar 2013

Image Collection Summarization Via Dictionary Learning For Sparse Representation, Chunlei Yang, Jialie Shen, Jinye Peng, Jianping Fan

Research Collection School Of Computing and Information Systems

In this paper, a novel approach is developed to achieve automatic image collection summarization. The effectiveness of the summary is reflected by its ability to reconstruct the original set or each individual image in the set. We have leveraged the dictionary learning for sparse representation model to construct the summary and to represent the image. Specifically we reformulate the summarization problem into a dictionary learning problem by selecting bases which can be sparsely combined to represent the original image and achieve a minimum global reconstruction error, such as MSE (Mean Square Error). The resulting “Sparse Least Square” problem is NP-hard, …


Online Multi-Modal Distance Learning For Scalable Multimedia Retrieval, Hao Xia, Pengcheng Wu, Steven C. H. Hoi Feb 2013

Online Multi-Modal Distance Learning For Scalable Multimedia Retrieval, Hao Xia, Pengcheng Wu, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

In many real-word scenarios, e.g., multimedia applications, data often originates from multiple heterogeneous sources or are represented by diverse types of representation, which is often referred to as "multi-modal data". The definition of distance between any two objects/items on multi-modal data is a key challenge encountered by many real-world applications, including multimedia retrieval. In this paper, we present a novel online learning framework for learning distance functions on multi-modal data through the combination of multiple kernels. In order to attack large-scale multimedia applications, we propose Online Multi-modal Distance Learning (OMDL) algorithms, which are significantly more efficient and scalable than the …


Hypergraph Index: An Index For Context-Aware Nearest Neighbor Query On Social Networks, Yazhe Wang, Baihua Zheng Jan 2013

Hypergraph Index: An Index For Context-Aware Nearest Neighbor Query On Social Networks, Yazhe Wang, Baihua Zheng

Research Collection School Of Computing and Information Systems

Social network has been touted as the No. 2 innovation in a recent IEEE Spectrum Special Report on “Top 11 Technologies of the Decade”, and it has cemented its status as a bona fide Internet phenomenon. With more and more people starting using social networks to share ideas, activities, events, and interests with other members within the network, social networks contain a huge amount of content. However, it might not be easy to navigate social networks to find specific information. In this paper, we define a new type of queries, namely context-aware nearest neighbor (CANN) search over social network to …


Towards Next-Generation Multimedia Recommendation Systems, Jialie Shen, Shuicheng Yan, Xian-Sheng Hua Jan 2013

Towards Next-Generation Multimedia Recommendation Systems, Jialie Shen, Shuicheng Yan, Xian-Sheng Hua

Research Collection School Of Computing and Information Systems

Empowered by advances in information technology, such as social media network, digital library and mobile computing, there emerges an ever-increasing amounts of multimedia data. As the key technology to address the problem of information overload, multimedia recommendation system has been received a lot of attentions from both industry and academia. This course aims to 1) provide a series of detailed review of state-of-the-art in multimedia recommendation; 2) analyze key technical challenges in developing and evaluating next generation multimedia recommendation systems from different perspectives and 3) give some predictions about the road lies ahead of us.


Cqarank: Jointly Model Topics And Expertise In Community Question Answering, Liu Yang, Minghui Qiu, Swapna Gottopati, Feida Zhu, Jing Jiang, Huiping Sun, Zhong Chen Jan 2013

Cqarank: Jointly Model Topics And Expertise In Community Question Answering, Liu Yang, Minghui Qiu, Swapna Gottopati, Feida Zhu, Jing Jiang, Huiping Sun, Zhong Chen

Research Collection School Of Computing and Information Systems

Community Question Answering (CQA) websites, where people share expertise on open platforms, have become large repositories of valuable knowledge. To bring the best value out of these knowledge repositories, it is critically important for CQA services to know how to find the right experts, retrieve archived similar questions and recommend best answers to new questions. To tackle this cluster of closely related problems in a principled approach, we proposed Topic Expertise Model (TEM), a novel probabilistic generative model with GMM hybrid, to jointly model topics and expertise by integrating textual content model and link structure analysis. Based on TEM results, …


Multimedia Recommendation: Technology And Techniques, Jialie Shen, Meng Wang, Shuicheng Yan, Peng Cui Jan 2013

Multimedia Recommendation: Technology And Techniques, Jialie Shen, Meng Wang, Shuicheng Yan, Peng Cui

Research Collection School Of Computing and Information Systems

In recent years, we have witnessed a rapid growth in the availability of digital multimedia on various application platforms and domains. Consequently, the problem of information overload has become more and more serious. In order to tackle the challenge, various multimedia recommendation technologies have been developed by different research communities (e.g., multimedia systems, information retrieval, machine learning and computer version). Meanwhile, many commercial web systems (e.g., Flick, YouTube, and Last.fm) have successfully applied recommendation techniques to provide users personalized content and services in a convenient and flexible way. When looking back, the information retrieval (IR) community has a long history …


Business Intelligence And Analytics: Research Directions, Ee Peng Lim, Hsinchun Chen, Guoqing Chen Jan 2013

Business Intelligence And Analytics: Research Directions, Ee Peng Lim, Hsinchun Chen, Guoqing Chen

Research Collection School Of Computing and Information Systems

Business intelligence and analytics (BIA) is about the development of technologies, systems, practices, and applications to analyze critical business data so as to gain new insights about business and markets. The new insights can be used for improving products and services, achieving better operational efficiency, and fostering customer relationships. In this article, we will categorize BIA research activities into three broad research directions: (a) big data analytics, (b) text analytics, and (c) network analytics. The article aims to review the state-of-the-art techniques and models and to summarize their use in BIA applications. For each research direction, we will also determine …


Cost-Sensitive Online Classification, Jialei Wang, Peilin Zhao, Steven C. H. Hoi Dec 2012

Cost-Sensitive Online Classification, Jialei Wang, Peilin Zhao, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Both cost-sensitive classification and online learning have been extensively studied in data mining and machine learning communities, respectively. However, very limited study addresses an important intersecting problem, that is, “Cost-Sensitive Online Classification". In this paper, we formally study this problem, and propose a new framework for Cost-Sensitive Online Classification by directly optimizing cost-sensitive measures using online gradient descent techniques. Specifically, we propose two novel cost-sensitive online classification algorithms, which are designed to directly optimize two well-known cost-sensitive measures: (i) maximization of weighted sum of sensitivity and specificity, and (ii) minimization of weighted misclassification cost. We analyze the theoretical bounds of …


A Survey Of Recommender Systems In Twitter, Su Mon Kywe, Ee Peng Lim, Feida Zhu Dec 2012

A Survey Of Recommender Systems In Twitter, Su Mon Kywe, Ee Peng Lim, Feida Zhu

Research Collection School Of Computing and Information Systems

Twitter is a social information network where short messages or tweets are shared among a large number of users through a very simple messaging mechanism. With a population of more than 100M users generating more than 300M tweets each day, Twitter users can be easily overwhelmed by the massive amount of information available and the huge number of people they can interact with. To overcome the above information overload problem, recommender systems can be introduced to help users make the appropriate selection. Researchers have began to study recommendation problems in Twitter but their works usually address individual recommendation tasks. There …


Detecting Anomalies In Bipartite Graphs With Mutual Dependency Principles, Hanbo Dai, Feida Zhu, Ee Peng Lim, Hwee Hwa Pang Dec 2012

Detecting Anomalies In Bipartite Graphs With Mutual Dependency Principles, Hanbo Dai, Feida Zhu, Ee Peng Lim, Hwee Hwa Pang

Research Collection School Of Computing and Information Systems

Bipartite graphs can model many real life applications including users-rating-products in online marketplaces, users-clicking-webpages on the World Wide Web and users referring users in social networks. In these graphs, the anomalousness of nodes in one partite often depends on that of their connected nodes in the other partite. Previous studies have shown that this dependency can be positive (the anomalousness of a node in one partite increases or decreases along with that of its connected nodes in the other partite) or negative (the anomalousness of a node in one partite rises or falls in opposite direction to that of its …


Impact Of Multimedia In Sina Weibo: Popularity And Life Span, Xun Zhao, Feida Zhu, Weining Qian, Aoying Zhou Nov 2012

Impact Of Multimedia In Sina Weibo: Popularity And Life Span, Xun Zhao, Feida Zhu, Weining Qian, Aoying Zhou

Research Collection School Of Computing and Information Systems

Multimedia contents such as images and videos are widely used in social network sites nowadays. Sina Weibo, a Chinese microblogging service, is one of the first microblog platforms to incorporate multimedia content sharing features. This work provides statistical analysis on how multimedia contents are produced, consumed, and propagated in Sina Weibo. Based on 230 million tweets and 1.8 million user profiles in Sina Weibo, we study the impact of multimedia contents on the popularity of both users and tweets as well as tweet life span. Our preliminary study shows that multimedia tweets dominant pure text ones in SinaWeibo. Multimedia contents …


Divad: A Dynamic And Interactive Visual Analytical Dashboard For Exploring And Analyzing Transport Data, Tin Seong Kam, Ketan Barshikar, Shaun Jun Hua Tan Nov 2012

Divad: A Dynamic And Interactive Visual Analytical Dashboard For Exploring And Analyzing Transport Data, Tin Seong Kam, Ketan Barshikar, Shaun Jun Hua Tan

Research Collection School Of Computing and Information Systems

The advances in location-based data collection technologies such as GPS, RFID etc. and the rapid reduction of their costs provide us with a huge and continuously increasing amount of data about movement of vehicles, people and goods in an urban area. This explosive growth of geospatially-referenced data has far outpaced the planner’s ability to utilize and transform the data into insightful information thus creating an adverse impact on the return on the investment made to collect and manage this data. Addressing this pressing need, we designed and developed DIVAD, a dynamic and interactive visual analytics dashboard to allow city planners …


A Unified Learning Framework For Auto Face Annotation By Mining Web Facial Images, Dayong Wang, Steven C. H. Hoi, Ying He Nov 2012

A Unified Learning Framework For Auto Face Annotation By Mining Web Facial Images, Dayong Wang, Steven C. H. Hoi, Ying He

Research Collection School Of Computing and Information Systems

Auto face annotation plays an important role in many real-world multimedia information and knowledge management systems. Recently there is a surge of research interests in mining weakly-labeled facial images on the internet to tackle this long-standing research challenge in computer vision and image understanding. In this paper, we present a novel unified learning framework for face annotation by mining weakly labeled web facial images through interdisciplinary efforts of combining sparse feature representation, content-based image retrieval, transductive learning and inductive learning techniques. In particular, we first introduce a new search-based face annotation paradigm using transductive learning, and then propose an effective …


Fast And Accurate Psd Matrix Estimation By Row Reduction, Hiroshi Kuwajima, Takashi Washio, Ee Peng Lim Nov 2012

Fast And Accurate Psd Matrix Estimation By Row Reduction, Hiroshi Kuwajima, Takashi Washio, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Fast and accurate estimation of missing relations, e.g., similarity, distance and kernel, among objects is now one of the most important techniques required by major data mining tasks, because the missing information of the relations is needed in many applications such as economics, psychology, and social network communities. Though some approaches have been proposed in the last several years, the practical balance between their required computation amount and obtained accuracy are insufficient for some class of the relation estimation. The objective of this paper is to formalize a problem to quickly and efficiently estimate missing relations among objects from the …


Entity Synonyms For Structured Web Search, Tao Cheng, Hady W. Lauw, Stelios Paparizos Oct 2012

Entity Synonyms For Structured Web Search, Tao Cheng, Hady W. Lauw, Stelios Paparizos

Research Collection School Of Computing and Information Systems

Nowadays, there are many queries issued to search engines targeting at finding values from structured data (e.g., movie showtime of a specific location). In such scenarios, there is often a mismatch between the values of structured data (how content creators describe entities) and the web queries (how different users try to retrieve them). Therefore, recognizing the alternative ways people use to reference an entity, is crucial for structured web search. In this paper, we study the problem of automatic generation of entity synonyms over structured data toward closing the gap between users and structured data. We propose an offline, data-driven …


Influentials, Novelty, And Social Contagion: The Viral Power Of Average Friends, Close Communities, And Old News, Nicholas Harrigan, Palakorn Achananuparp, Ee Peng Lim Oct 2012

Influentials, Novelty, And Social Contagion: The Viral Power Of Average Friends, Close Communities, And Old News, Nicholas Harrigan, Palakorn Achananuparp, Ee Peng Lim

Research Collection School Of Computing and Information Systems

What is the effect of (1) popular individuals, and (2) community structures on the retransmission of socially contagious behavior? We examine a community of Twitter users over a five month period, operationalizing social contagion as ‘retweeting’, and social structure as the count of subgraphs (small patterns of ties and nodes) between users in the follower/following network. We find that popular individuals act as ‘inefficient hubs’ for social contagion: they have limited attention, are overloaded with inputs, and therefore display limited responsiveness to viral messages. We argue this contradicts the ‘law of the few’ and ‘influentials hypothesis’. We find that community …


Boosting Multi-Kernel Locality-Sensitive Hashing For Scalable Image Retrieval, Hao Xia, Steven C. H. Hoi, Pengcheng Wu, Rong Jin Aug 2012

Boosting Multi-Kernel Locality-Sensitive Hashing For Scalable Image Retrieval, Hao Xia, Steven C. H. Hoi, Pengcheng Wu, Rong Jin

Research Collection School Of Computing and Information Systems

Similarity search is a key challenge for multimedia retrieval applications where data are usually represented in high-dimensional space. Among various algorithms proposed for similarity search in high-dimensional space, Locality-Sensitive Hashing (LSH) is the most popular one, which recently has been extended to Kernelized Locality-Sensitive Hashing (KLSH) by exploiting kernel similarity for better retrieval efficacy. Typically, KLSH works only with a single kernel, which is often limited in real-world multimedia applications, where data may originate from multiple resources or can be represented in several different forms. For example, in content-based multimedia retrieval, a variety of features can be extracted to represent …


Modeling Concept Dynamics For Large Scale Music Search, Jialie Shen, Hwee Hwa Pang, Meng Wang, Shuicheng Yan Aug 2012

Modeling Concept Dynamics For Large Scale Music Search, Jialie Shen, Hwee Hwa Pang, Meng Wang, Shuicheng Yan

Research Collection School Of Computing and Information Systems

Continuing advances in data storage and communication technologies have led to an explosive growth in digital music collections. To cope with their increasing scale, we need effective Music Information Retrieval (MIR) capabilities like tagging, concept search and clustering. Integral to MIR is a framework for modelling music documents and generating discriminative signatures for them. In this paper, we introduce a multimodal, layered learning framework called DMCM. Distinguished from the existing approaches that encode music as an ensemble of order-less feature vectors, our framework extracts from each music document a variety of acoustic features, and translates them into low-level encodings over …