Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Numerical Analysis and Scientific Computing (671)
- Social and Behavioral Sciences (374)
- Artificial Intelligence and Robotics (360)
- Graphics and Human Computer Interfaces (313)
- Business (253)
-
- Software Engineering (239)
- Communication (234)
- Social Media (202)
- Engineering (199)
- Theory and Algorithms (178)
- Computer Engineering (171)
- Information Security (148)
- OS and Networks (116)
- Programming Languages and Compilers (98)
- E-Commerce (84)
- Data Storage Systems (75)
- Medicine and Health Sciences (70)
- Public Affairs, Public Policy and Public Administration (60)
- Education (55)
- Management Information Systems (53)
- International and Area Studies (51)
- Asian Studies (50)
- Health Information Technology (48)
- Transportation (47)
- Finance and Financial Management (43)
- Digital Communications and Networking (32)
- Technology and Innovation (32)
- Keyword
-
- Social media (59)
- Machine learning (56)
- Online learning (46)
- Deep learning (43)
- Data mining (42)
-
- Artificial intelligence (36)
- Twitter (30)
- Query processing (29)
- Classification (26)
- Neural networks (25)
- Reinforcement learning (25)
- Deep Learning (24)
- Algorithms (23)
- Clustering (21)
- Social network (21)
- Algorithm (20)
- Graph neural networks (20)
- Machine Learning (20)
- Natural language processing (20)
- Recommender systems (20)
- Semantics (20)
- Task analysis (20)
- Anomaly detection (19)
- Cloud computing (19)
- Visualization (19)
- Image retrieval (18)
- Performance (18)
- Sentiment analysis (18)
- Singapore (18)
- Social networks (17)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (3441)
- Dissertations and Theses Collection (Open Access) (58)
- Research Collection Lee Kong Chian School Of Business (11)
- Asian Management Insights (8)
- Research Collection School Of Accountancy (7)
-
- Dissertations and Theses Collection (5)
- PhD Student’s Publications Collection (5)
- Research Collection College of Integrative Studies (5)
- Research Collection Yong Pung How School Of Law (5)
- MITB Thought Leadership Series (3)
- LARC Research Publications (2)
- Perspectives@SMU (2)
- Research Collection School of Computing and Information Systems (2)
- 2024 AI for Research Week (1)
- CCX Research (1)
- Research Collection School Of Economics (1)
- Research Collection School of Accountancy (1)
- Research Collection School of Social Sciences (1)
- Research@SMU Infographics (1)
- Publication Type
Articles 2071 - 2100 of 3560
Full-Text Articles in Databases and Information Systems
Cnl: Collective Network Linkage Across Heterogeneous Social Platforms, Ming Gao, Ee-Peng Lim, David Lo, Feida Zhu, Philips Kokoh Prasetyo, Aoying Zhou
Cnl: Collective Network Linkage Across Heterogeneous Social Platforms, Ming Gao, Ee-Peng Lim, David Lo, Feida Zhu, Philips Kokoh Prasetyo, Aoying Zhou
Research Collection School Of Computing and Information Systems
The popularity of social media has led many users to create accounts with different online social networks. Identifying these multiple accounts belonging to same user is of critical importance to user profiling, community detection, user behavior understanding and product recommendation. Nevertheless, linking users across heterogeneous social networks is challenging due to large network sizes, heterogeneous user attributes and behaviors in different networks, and noises in user generated data. In this paper, we propose an unsupervised method, Collective Network Linkage (CNL), to link users across heterogeneous social networks. CNL incorporates heterogeneous attributes and social features unique to social network users, handles …
Not All Trips Are Equal: Analyzing Foursquare Check-Ins Of Trips And City Visitors, Wen Haw Chong, Bingtian Dai, Ee Peng Lim
Not All Trips Are Equal: Analyzing Foursquare Check-Ins Of Trips And City Visitors, Wen Haw Chong, Bingtian Dai, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Location-Based Social Networks (LBSN) such as Foursquare allow users to indicate venue visits via check-ins. This results in much fine grained context-rich data, useful for studying user mobility. In this work, we use check-ins to characterize trips and visitors to two cities, where visitors are defined as having their home cities elsewhere. First, we divide trips into two duration types: long and short. We then show that trip types differ in check-in distributions over venue categories, time slots, as well as check-in intensity. Based on the trip types, we then divide visitors into long-term and short-term visitors. We compare visitor …
Modelling Cascades Over Time In Microblogs, Xie Wei, Feida Zhu, Siyuan Liu, Ke Wang
Modelling Cascades Over Time In Microblogs, Xie Wei, Feida Zhu, Siyuan Liu, Ke Wang
Research Collection School Of Computing and Information Systems
One of the most important features of microblogging services such as Twitter is how easy it is to re-share a piece of information across the network through various user connections, forming what we call a "cascade". Business applications such as viral marketing have driven a tremendous amount of research effort predicting whether a certain cascade will go viral. Yet the rarity of viral cascades in real data poses a challenge to all existing prediction methods. One solution is to simulate cascades that well fit the real viral ones, which requires our ability to tell how a certain cascade grows over …
Analysis Of Aspects And Star Ratings In Consumer Reviews, Maruthi Prithivirajan, Vivian Lai, Kyong Jin Shim
Analysis Of Aspects And Star Ratings In Consumer Reviews, Maruthi Prithivirajan, Vivian Lai, Kyong Jin Shim
Research Collection School Of Computing and Information Systems
This paper presents an analysis of star ratings in consumer reviews in Yelp, an online social platform for sharing consumer reviews about local businesses. In particular, we analyze consumer reviews about food businesses. We analyze how well or poorly the star ratings (on a scale of one star to five stars) associated with these reviews tally with the sentiment derived from the textual portion of the consumer review.
Dictionary Pair Learning On Grassmann Manifolds For Image Denoising, Xianhua Zeng, Wei Bian, Wei Liu, Jialie Shen, Dacheng Tao
Dictionary Pair Learning On Grassmann Manifolds For Image Denoising, Xianhua Zeng, Wei Bian, Wei Liu, Jialie Shen, Dacheng Tao
Research Collection School Of Computing and Information Systems
Image denoising is a fundamental problem in computer vision and image processing that holds considerable practical importance for real-world applications. The traditional patch-based and sparse coding-driven image denoising methods convert 2D image patches into 1D vectors for further processing. Thus, these methods inevitably break down the inherent 2D geometric structure of natural images. To overcome this limitation pertaining to the previous image denoising methods, we propose a 2D image denoising model, namely, the dictionary pair learning (DPL) model, and we design a corresponding algorithm called the DPL on the Grassmann-manifold (DPLG) algorithm. The DPLG algorithm first learns an initial dictionary …
Where Are The Passengers? A Grid-Based Gaussian Mixture Model For Taxi Bookings, Meng-Fen Chiang, Tuan Anh Hoang, Ee-Peng Lim
Where Are The Passengers? A Grid-Based Gaussian Mixture Model For Taxi Bookings, Meng-Fen Chiang, Tuan Anh Hoang, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Taxi bookings are events where requests for taxis are made by passengers either over voice calls or mobile apps. As the demand for taxis changes with space and time, it is important to model both the space and temporal dimensions in dynamic booking data. Several applications can benefit from a good taxi booking model. These include the prediction of number of bookings at certain location and time of the day, and the detection of anomalous booking events. In this paper, we propose a Grid-based Gaussian Mixture Model (GGMM) with spatio-temporal dimensions that groups booking data into a number of spatio-temporal …
Using Digital Genomics To Create An Intelligent Enterprise, Mario Domingo
Using Digital Genomics To Create An Intelligent Enterprise, Mario Domingo
Asian Management Insights
Every business knows that it needs to leverage customer data, but few know the potential it has to transform business processes, decisions and performance.
A Method And System For Sentiment Classification And Emotion Classification [Us Patent 20170308523a1], Zhaoxia Wang, Rick Siow Mong Goh, Yinping Yang
A Method And System For Sentiment Classification And Emotion Classification [Us Patent 20170308523a1], Zhaoxia Wang, Rick Siow Mong Goh, Yinping Yang
Research Collection School Of Computing and Information Systems
A system and a method for classifying text messages, such as social media messages into sentiment valence categories are provided. The system comprising a module for decomposing text messages, a module for cleaning text messages, a module for producing feature data of text messages, and a module for classifying text messages into sentiment valence categories. The module for decomposing text messages is configured to: receive a text message, parse the text message into separate portions in response to parsing criteria based on sentence delimiters, wherein the separate portions are sentences, phrases and words, and rejoin at least some of the …
Face Recognition On Large-Scale Video In The Wild With Hybrid Euclidean-And-Riemannian Metric Learning, Zhiwu Huang, R. Wang, S. Shan, X Chen
Face Recognition On Large-Scale Video In The Wild With Hybrid Euclidean-And-Riemannian Metric Learning, Zhiwu Huang, R. Wang, S. Shan, X Chen
Research Collection School Of Computing and Information Systems
Face recognition on large-scale video in the wild is becoming increasingly important due to the ubiquity of video data captured by surveillance cameras, handheld devices, Internet uploads, and other sources. By treating each video as one image set, set-based methods recently have made great success in the field of video-based face recognition. In the wild world, videos often contain extremely complex data variations and thus pose a big challenge of set modeling for set-based methods. In this paper, we propose a novel Hybrid Euclidean-and-Riemannian Metric Learning (HERML) method to fuse multiple statistics of image set. Specifically, we represent each image …
Assessing Developer Contribution With Repository Mining-Based Metrics, Jalerson Lima, Christoph Treude, Fernando Figueira Filho, Uirá Kulesza
Assessing Developer Contribution With Repository Mining-Based Metrics, Jalerson Lima, Christoph Treude, Fernando Figueira Filho, Uirá Kulesza
Research Collection School Of Computing and Information Systems
Productivity as a result of individual developers' contributions is an important aspect for software companies to maintain their competitiveness in the market. However, there is no consensus in the literature on how to measure productivity or developer contribution. While some repository mining-based metrics have been proposed, they lack validation in terms of their applicability and usefulness from the individuals who will use them to assess developer contribution: team and project leaders. In this paper, we propose the design of a suite of metrics for the assessment of developer contribution, based on empirical evidence obtained from project and team leaders. In …
Scheduled Approximation For Personalized Pagerank With Utility-Based Hub Selection, Fanwei Zhu, Yuan Fang, Kevin Chen-Chuan Chang, Jing Ying
Scheduled Approximation For Personalized Pagerank With Utility-Based Hub Selection, Fanwei Zhu, Yuan Fang, Kevin Chen-Chuan Chang, Jing Ying
Research Collection School Of Computing and Information Systems
As Personalized PageRank has been widely leveraged for ranking on a graph, the efficient computation of Personalized PageRank Vector (PPV) becomes a prominent issue. In this paper, we propose FastPPV, an approximate PPV computation algorithm that is incremental and accuracy-aware. Our approach hinges on a novel paradigm of scheduled approximation: the computation is partitioned and scheduled for processing in an “organized” way, such that we can gradually improve our PPV estimation in an incremental manner and quantify the accuracy of our approximation at query time. Guided by this principle, we develop an efficient hub-based realization, where we adopt the metric …
Detect Rumors Using Time Series Of Social Context Information On Microblogging Websites, Jing Ma, Wei Gao, Zhongyu Wei, Yueming Lu, Kam-Fai Wong
Detect Rumors Using Time Series Of Social Context Information On Microblogging Websites, Jing Ma, Wei Gao, Zhongyu Wei, Yueming Lu, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
Automatically identifying rumors from online social media especially microblogging websites is an important research issue. Most of existing work for rumor detection focuses on modeling features related to microblog contents, users and propagation patterns, but ignore the importance of the variation of these social context features during the message propagation over time. In this study, we propose a novel approach to capture the temporal characteristics of these features based on the time series of rumor's lifecycle, for which time series modeling technique is applied to incorporate various social context information. Our experiments using the events in two microblog datasets confirm …
Choosing Your Weapons: On Sentiment Analysis Tools For Software Engineering Research, Robbert Jongeling, Subhajit Datta, Alexander Serebrenik
Choosing Your Weapons: On Sentiment Analysis Tools For Software Engineering Research, Robbert Jongeling, Subhajit Datta, Alexander Serebrenik
Research Collection School Of Computing and Information Systems
Recent years have seen an increasing attention to social aspects of software engineering, including studies of emotions and sentiments experienced and expressed by the software developers. Most of these studies reuse existing sentiment analysis tools such as SentiStrength and NLTK. However, these tools have been trained on product reviews and movie reviews and, therefore, their results might not be applicable in the software engineering domain. In this paper we study whether the sentiment analysis tools agree with the sentiment recognized by human evaluators (as reported in an earlier study) as well as with each other. Furthermore, we evaluate the impact …
The Importance Of Being Isolated: An Empirical Study On Chromium Reviews, Subhajit Datta, Devarshi Bhatt, Manish Jain, Proshanta Sarkar, Santonu Sarkar
The Importance Of Being Isolated: An Empirical Study On Chromium Reviews, Subhajit Datta, Devarshi Bhatt, Manish Jain, Proshanta Sarkar, Santonu Sarkar
Research Collection School Of Computing and Information Systems
As large scale software development has become more collaborative, and software teams more globally distributed, several studies have explored how developer interaction influences software development outcomes. The emphasis so far has been largely on outcomes like defect count, the time to close modification requests etc. In the paper, we examine data from the Chromium project to understand how different aspects of developer discussion relate to the closure time of reviews. On the basis of analyzing reviews discussed by 2000+ developers, our results indicate that quicker closure of reviews owned by a developer relates to higher reception of information and insights …
Social Tag Relevance Estimation Via Ranking-Oriented Neighbour Voting, Chaoran Cui, Jialie Shen, Jun Ma, Tao Lian
Social Tag Relevance Estimation Via Ranking-Oriented Neighbour Voting, Chaoran Cui, Jialie Shen, Jun Ma, Tao Lian
Research Collection School Of Computing and Information Systems
User-generated tags associated with social images are frequently imprecise and incomplete. Therefore, a fundamental challenge in tag-based applications is the problem of tag relevance estimation, which concerns how to interpret and quantify the relevance of a tag with respect to the contents of an image. In this paper, we address the key problem from a new perspective of learning to rank, and develop a novel approach to facilitate tag relevance estimation to directly optimize the ranking performance of tag-based image search. A supervision step is introduced into the neighbour voting scheme, in which tag relevance is estimated by accumulating votes …
Structural Constraints For Multipartite Entity Resolution With Markov Logic Network, Tengyuan Ye, Hady W. Lauw
Structural Constraints For Multipartite Entity Resolution With Markov Logic Network, Tengyuan Ye, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Multipartite entity resolution seeks to match entity mentions across several collections. An entity mention is presumed unique within a collection, and thus could match at most one entity mention in each of the other collections. In addition to domain-specific features considered in entity resolution, there are a number of domain-invariant structural contraints that apply in this scenario, including one-to-one assignment as well as cross-collection transitivity. We propose a principled solution to the multipartite entity resolution problem, building on the foundation of Markov Logic Network (MLN) that combines probabilistic graphical model and first-order logic. We describe how the domain-invariant structural constraints …
Learning Relative Similarity From Data Streams: Active Online Learning Approaches, Shuji Hao, Peilin Zhao, Steven C. H. Hoi, Chunyan Miao
Learning Relative Similarity From Data Streams: Active Online Learning Approaches, Shuji Hao, Peilin Zhao, Steven C. H. Hoi, Chunyan Miao
Research Collection School Of Computing and Information Systems
Relative similarity learning, as an important learning scheme for information retrieval, aims to learn a bi-linear similarity function from a collection of labeled instance-pairs, and the learned function would assign a high similarity value for a similar instance-pair and a low value for a dissimilar pair. Existing algorithms usually assume the labels of all the pairs in data streams are always made available for learning. However, this is not always realistic in practice since the number of possible pairs is quadratic to the number of instances in the database, and manually labeling the pairs could be very costly and time …
On Robust Image Spam Filtering Via Comprehensive Visual Modeling, Jialie Shen, Deng, Robert H., Zhiyong Cheng, Liqiang Nie, Shuicheng Yan
On Robust Image Spam Filtering Via Comprehensive Visual Modeling, Jialie Shen, Deng, Robert H., Zhiyong Cheng, Liqiang Nie, Shuicheng Yan
Research Collection School Of Computing and Information Systems
The Internet has brought about fundamental changes in the way peoples generate and exchange media information. Over the last decade, unsolicited message images (image spams) have become one of the most serious problems for Internet service providers (ISPs), business firms and general end users. In this paper, we report a novel system called RoBoTs (Robust BoosTrap based spam detector) to support accurate and robust image spam filtering. The system is developed based on multiple visual properties extracted from different levels of granularity, aiming to capture more discriminative contents for effective spam image identification. In addition, a resampling based learning framework …
Two Formulas For Success In Social Media: Learning And Network Effects, Liangfei Qiu, Qian Tang, Andrew B. Whinston
Two Formulas For Success In Social Media: Learning And Network Effects, Liangfei Qiu, Qian Tang, Andrew B. Whinston
Research Collection School Of Computing and Information Systems
Recent years have witnessed an unprecedented explosion in information technology that enables dynamic diffusion of user-generated content in social networks. Online videos, in particular, have changed the landscape of marketing and entertainment, competing with premium content and spurring business innovations. In the present study, we examine how learning and network effects drive the diffusion of online videos. While learning happens through informational externalities, network effects are direct payoff externalities. Using a unique data set from YouTube, we empirically identify learning and network effects separately, and find that both mechanisms have statistically and economically significant effects on video views; furthermore, the …
Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong
Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
A microblog repost tree provides strong clues on how an event described therein develops. To help social media users capture the main clues of events on microblogging sites, we propose a novel repost tree summarization framework by effectively differentiating two kinds of messages on repost trees called leaders and followers, which are derived from contentlevel structure information, i.e., contents of messages and the reposting relations. To this end, Conditional Random Fields (CRF) model is used to detect leaders across repost tree paths. We then present a variant of random-walk-based summarization model to rank and select salient messages based on the …
Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong
Using Content-Level Structures For Summarizing Microblog Repost Trees, Jing Li, Wei Gao, Zhongyu Wei, Baolin Peng, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
A microblog repost tree provides strong clues on how an event described therein develops. To help social media users capture the main clues of events on microblogging sites, we propose a novel repost tree summarization framework by effectively differentiating two kinds of messages on repost trees called leaders and followers, which are derived from contentlevel structure information, i.e., contents of messages and the reposting relations. To this end, Conditional Random Fields (CRF) model is used to detect leaders across repost tree paths. We then present a variant of random-walk-based summarization model to rank and select salient messages based on the …
Name List Only? Target Entity Disambiguation In Short Texts, Yixin Cao, Juanzi Li, Xiaofei Guo, Shuanhu Bai, Heng Ji, Jie Tang
Name List Only? Target Entity Disambiguation In Short Texts, Yixin Cao, Juanzi Li, Xiaofei Guo, Shuanhu Bai, Heng Ji, Jie Tang
Research Collection School Of Computing and Information Systems
Target entity disambiguation (TED), the task of identifying target entities of the same domain, has been recognized as a critical step in various important applications. In this paper, we propose a graphbased model called TremenRank to collectively identify target entities in short texts given a name list only. TremenRank propagates trust within the graph, allowing for an arbitrary number of target entities and texts using inverted index technology. Furthermore, we design a multi-layer directed graph to assign different trust levels to short texts for better performance. The experimental results demonstrate that our model outperforms state-of-the-art methods with an average gain …
Real-Time Targeted Influence Maximization For Online Advertisements, Yuchen Li, Dongxiang Zhang, Kian-Lee Tan
Real-Time Targeted Influence Maximization For Online Advertisements, Yuchen Li, Dongxiang Zhang, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
Advertising in social network has become a multi-billion dollar industry. A main challenge is to identify key influencers who can effectively contribute to the dissemination of information. Although the influence maximization problem, which finds a seed set of k most influential users based on certain propagation models, has been well studied, it is not target-aware and cannot be directly applied to online advertising. In this paper, we propose a new problem, named Keyword-Based Targeted Influence Maximization (KB-TIM), to find a seed set that maximizes the expected influence over users who are relevant to a given advertisement. To solve the problem, …
Mining Revenue-Maximizing Bundling Configuration, Loc Do, Hady Wirawan Lauw, Ke Wang
Mining Revenue-Maximizing Bundling Configuration, Loc Do, Hady Wirawan Lauw, Ke Wang
Research Collection School Of Computing and Information Systems
With greater prevalence of social media, there is an increasing amount of user-generated data revealing consumer preferences for various products and services. Businesses seek to harness this wealth of data to improve their marketing strategies. Bundling, or selling two or more items for one price is a highly-practiced marketing strategy. In this paper, we address the bundle configuration problem from the data-driven perspective. Given a set of items in a seller’s inventory, we seek to determine which items should belong to which bundle so as to maximize the total revenue, by mining consumer preferences data. We show that this problem …
Cobweb: A Robust Map Update System Using Gps Trajectories, Zhangqing Shan, Hao Wu, Weiwei Sun, Baihua Zheng
Cobweb: A Robust Map Update System Using Gps Trajectories, Zhangqing Shan, Hao Wu, Weiwei Sun, Baihua Zheng
Research Collection School Of Computing and Information Systems
The accuracy and completeness of a digital map plays a critical role in determining the quality of most location-based services. Unfortunately, road networks change frequently. Consequently, we study the issue of automatic map update in this paper. We propose a system called COBWEB which takes all the unmatched trajectories as input and generates the missing road segments with both the geometry properties and topology features well preserved. We conduct a comprehensive experimental study via real trajectory data generated by roughly 15,000 taxis in Singapore within a 5-month period. Compared with existing work, COBWEB demonstrates a better and more stable performance …
Maximum Rank Query, Kyriakos Mouratidis, Jilian Zhang, Hwee Hwa Pang
Maximum Rank Query, Kyriakos Mouratidis, Jilian Zhang, Hwee Hwa Pang
Research Collection School Of Computing and Information Systems
The top-k query is a common means to shortlist a number of options from a set of alternatives, based on the user's preferences. Typically, these preferences are expressed as a vector of query weights, defined over the options' attributes. The query vector implicitly associates each alternative with a numeric score, and thus imposes a ranking among them. The top-k result includes the k options with the highest scores. In this context, we define the maximum rank query (MaxRank). Given a focal option in a set of alternatives, the MaxRank problem is to compute the highest rank this option may achieve …
Tagcombine: Recommending Tags To Contents In Software Information Sites, Xin Yu Wang, Xin Xia, David Lo
Tagcombine: Recommending Tags To Contents In Software Information Sites, Xin Yu Wang, Xin Xia, David Lo
Research Collection School Of Computing and Information Systems
Nowadays, software engineers use a variety of online media to search and become informed of new and interesting technologies, and to learn from and help one another. We refer to these kinds of online media which help software engineers improve their performance in software development, maintenance, and test processes as software information sites. In this paper, we propose TagCombine, an automatic tag recommendation method which analyzes objects in software information sites. TagCombine has three different components: 1) multi-label ranking component which considers tag recommendation as a multi-label learning problem; 2) similarity-based ranking component which recommends tags from similar objects; 3) …
Answering Why-Not Questions On Reverse Top-K Queries, Yunjun Gao, Qing Liu, Gang Chen, Baihua Zheng, Linlin Zhou
Answering Why-Not Questions On Reverse Top-K Queries, Yunjun Gao, Qing Liu, Gang Chen, Baihua Zheng, Linlin Zhou
Research Collection School Of Computing and Information Systems
Why-not questions, which aim to seek clarifications on the missing tuples for query results, have recently received considerable attention from the database community. In this paper, we systematically explore why-not questions on reverse top-k queries, owing to its importance in multi-criteria decision making. Given an initial reverse top-k query and a missing/why-not weighting vector set Wm that is absent from the query result, why-not questions on reverse top-k queries explain why Wm does not appear in the query result and provide suggestions on how to refine the initial query with minimum penalty to include Wm in the refined query result. …
Towards Opinion Summarization From Online Forums, Ding Ying, Jing Jiang
Towards Opinion Summarization From Online Forums, Ding Ying, Jing Jiang
Research Collection School Of Computing and Information Systems
Summarizing opinions expressed in online forums can potentially benefit many people. However, special characteristics of this problem may require changes to standard text summarization techniques. In this work, we present our initial attempt at extractive summarization of opinionated online forum threads. Given the nature of user generated content in online discussion forums, we hypothesize that besides relevance, text quality and subjectivity also play important roles in deciding which sentences are good summary sentences. We therefore construct an annotated corpus to facilitate our study of extractive summarization of online discussion forums. We define a set of features to capture relevance, text …
A Joint Model Of Product Properties, Aspects And Ratings For Online Reviews, Ding Ying, Jing Jiang
A Joint Model Of Product Properties, Aspects And Ratings For Online Reviews, Ding Ying, Jing Jiang
Research Collection School Of Computing and Information Systems
Product review mining is an important task that can benefit both businesses and consumers. Lately a number of models combining collaborative filtering and content analysis to model reviews have been proposed, among which the Hidden Factors as Topics (HFT) model is a notable one. In this work, we propose a new model on top of HFT to separate product properties and aspects. Product properties are intrinsic to certain products (e.g. types of cuisines of restaurants) whereas aspects are dimensions along which products in the same category can be compared (e.g. service quality of restaurants). Our proposed model explicitly separates the …