Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year

Articles 931 - 960 of 3441

Full-Text Articles in Databases and Information Systems

Characterizing Search Activities On Stack Overflow, Jiakun Liu, Sebastian Baltes, Christoph Treude, David Lo, Yun Zhang, Xin Xia Aug 2021

Characterizing Search Activities On Stack Overflow, Jiakun Liu, Sebastian Baltes, Christoph Treude, David Lo, Yun Zhang, Xin Xia

Research Collection School Of Computing and Information Systems

To solve programming issues, developers commonly search on Stack Overflow to seek potential solutions. However, there is a gap between the knowledge developers are interested in and the knowledge they are able to retrieve using search engines. To help developers efficiently retrieve relevant knowledge on Stack Overflow, prior studies proposed several techniques to reformulate queries and generate summarized answers. However, few studies performed a large-scale analysis using real-world search logs. In this paper, we characterize how developers search on Stack Overflow using such logs. By doing so, we identify the challenges developers face when searching on Stack Overflow and seek …


Learning And Exploiting Shaped Reward Models For Large Scale Multiagent Rl, Arambam James Singh, Akshat Kumar, Hoong Chuin Lau Aug 2021

Learning And Exploiting Shaped Reward Models For Large Scale Multiagent Rl, Arambam James Singh, Akshat Kumar, Hoong Chuin Lau

Research Collection School Of Computing and Information Systems

Many real world systems involve interaction among large number of agents to achieve a common goal, for example, air traffic control. Several model-free RL algorithms have been proposed for such settings. A key limitation is that the empirical reward signal in model-free case is not very effective in addressing the multiagent credit assignment problem, which determines an agent's contribution to the team's success. This results in lower solution quality and high sample complexity. To address this, we contribute (a) an approach to learn a differentiable reward model for both continuous and discrete action setting by exploiting the collective nature of …


Data Pricing And Data Asset Governance In The Ai Era, Jian Pei, Feida Zhu, Zicun Cong, Luo Xuan, Liu Huiwen, Xin Mu Aug 2021

Data Pricing And Data Asset Governance In The Ai Era, Jian Pei, Feida Zhu, Zicun Cong, Luo Xuan, Liu Huiwen, Xin Mu

Research Collection School Of Computing and Information Systems

Data is one of the most critical resources in the AI Era. While substantial research has been dedicated to training machine learning models using various types of data, much less efforts have been invested in the exploration of assessing and governing data assets in end-to-end processes of machine learning and data science, that is, the pipeline where data is collected and processed, and then machine learning models are produced, requested, deployed, shared and evolved. To provide a state-of-the-art overall picture of this important and novel area and advocate the related research and development, we present a tutorial addressing two essential …


Toward Explainable Deep Anomaly Detection, Guansong Pang, Charu Aggarwal Aug 2021

Toward Explainable Deep Anomaly Detection, Guansong Pang, Charu Aggarwal

Research Collection School Of Computing and Information Systems

Anomaly explanation, also known as anomaly localization, is as important as, if not more than, anomaly detection in many realworld applications. However, it is challenging to build explainable detection models due to the lack of anomaly-supervisory information and the unbounded nature of anomaly; most existing studies exclusively focus on the detection task only, including the recently emerging deep learning-based anomaly detection that leverages neural networks to learn expressive low-dimensional representations or anomaly scores for the detection task. Deep learning models, including deep anomaly detection models, are often constructed as black boxes, which have been criticized for the lack of explainability …


A Mean-Field Markov Decision Process Model For Spatial-Temporal Subsidies In Ride-Sourcing Markets, Zheng Zhu, Jintao Ke, Hai Wang Jul 2021

A Mean-Field Markov Decision Process Model For Spatial-Temporal Subsidies In Ride-Sourcing Markets, Zheng Zhu, Jintao Ke, Hai Wang

Research Collection School Of Computing and Information Systems

Ride-sourcing services are increasingly popular because of their ability to accommodate on-demand travel needs. A critical issue faced by ride-sourcing platforms is the supply-demand imbalance, as a result of which drivers may spend substantial time on idle cruising and picking up remote passengers. Some platforms attempt to mitigate the imbalance by providing relocation guidance for idle drivers who may have their own self-relocation strategies and decline to follow the suggestions. Platforms then seek to induce drivers to system-desirable locations by offering them subsidies. This paper proposes a mean-field Markov decision process (MF-MDP) model to depict the dynamics in ride-sourcing markets …


Meta-Inductive Node Classification Across Graphs, Zhihao Wen, Yuan Fang, Zemin Liu Jul 2021

Meta-Inductive Node Classification Across Graphs, Zhihao Wen, Yuan Fang, Zemin Liu

Research Collection School Of Computing and Information Systems

Semi-supervised node classification on graphs is an important research problem, with many real-world applications in information retrieval such as content classification on a social network and query intent classification on an e-commerce query graph. While traditional approaches are largely transductive, recent graph neural networks (GNNs) integrate node features with network structures, thus enabling inductive node classification models that can be applied to new nodes or even new graphs in the same feature space. However, inter-graph differences still exist across graphs within the same domain. Thus, training just one global model (e.g., a state-of-the-art GNN) to handle all new graphs, whilst …


Marina: Faster Non-Convex Distributed Learning With Compression, Eduard Gorbunov, Konstantin Burlachenko, Zhize Li, Peter Richtarik Jul 2021

Marina: Faster Non-Convex Distributed Learning With Compression, Eduard Gorbunov, Konstantin Burlachenko, Zhize Li, Peter Richtarik

Research Collection School Of Computing and Information Systems

We develop and analyze MARINA: a new communication efficient method for non-convex distributed learning over heterogeneous datasets. MARINA employs a novel communication compression strategy based on the compression of gradient differences that is reminiscent of but different from the strategy employed in the DIANA method of Mishchenko et al. (2019). Unlike virtually all competing distributed first-order methods, including DIANA, ours is based on a carefully designed biased gradient estimator, which is the key to its superior theoretical and practical performance. The communication complexity bounds we prove for MARINA are evidently better than those of all previous first-order methods. Further, we …


Unified Conversational Recommendation Policy Learning Via Graph-Based Reinforcement Learning, Yang Deng, Yaliang Li, Fei Sun, Bolin Ding, Wai Lam Jul 2021

Unified Conversational Recommendation Policy Learning Via Graph-Based Reinforcement Learning, Yang Deng, Yaliang Li, Fei Sun, Bolin Ding, Wai Lam

Research Collection School Of Computing and Information Systems

Conversational recommender systems (CRS) enable the traditional recommender systems to explicitly acquire user preferences towards items and attributes through interactive conversations. Reinforcement learning (RL) is widely adopted to learn conversational recommendation policies to decide what attributes to ask, which items to recommend, and when to ask or recommend, at each conversation turn. However, existing methods mainly target at solving one or two of these three decision-making problems in CRS with separated conversation and recommendation components, which restrict the scalability and generality of CRS and fall short of preserving a stable training procedure. In the light of these challenges, we propose …


Users’ Reception Of Product Recommendations: Analyses Based On Eye Tracking Data, Feiyan Jia, Yani Shi, Choon Ling Sia, Chuan-Hoo Tan, Fiona Fui-Hoon Nah, Keng Siau Jul 2021

Users’ Reception Of Product Recommendations: Analyses Based On Eye Tracking Data, Feiyan Jia, Yani Shi, Choon Ling Sia, Chuan-Hoo Tan, Fiona Fui-Hoon Nah, Keng Siau

Research Collection School Of Computing and Information Systems

Based on eye tracking technology, we study consumers’ overall attention to recommendations appearing at different time settings (i.e., early, mid, and late) and their attention to different information contained in each recommendation, such as recommendation signs, product descriptions, and reviews. By investigating consumers’ eye movement patterns and attention distributions on recommendations, we open the “black box” of why consumers’ reception to recommendations appearing at different time settings varies. The product preference construction literature and mindset theory help to explain why the early recommendations receive the most attention. The need for justification helps to explain why the late recommendations should receive …


Oesense: Employing Occlusion Effect For In-Ear Human Sensing, Dong Ma, Andrea Ferlini, Cecilia Mascolo Jul 2021

Oesense: Employing Occlusion Effect For In-Ear Human Sensing, Dong Ma, Andrea Ferlini, Cecilia Mascolo

Research Collection School Of Computing and Information Systems

Smart earbuds are recognized as a new wearable platform for personal-scale human motion sensing. However, due to the interference from head movement or background noise, commonly-used modalities (e.g. accelerometer and microphone) fail to reliably detect both intense and light motions. To obviate this, we propose OESense, an acoustic-based in-ear system for general human motion sensing. The core idea behind OESense is the joint use of the occlusion effect (i.e., the enhancement of low-frequency components of bone-conducted sounds in an occluded ear canal) and inward-facing microphone, which naturally boosts the sensing signal and suppresses external interference. We prototype OESense as an …


Frameaxis: Characterizing Microframe Bias And Intensity With Word Embedding, Haewoon Kwak, Jisun An, Elise Jing Jing, Yong-Yeol Ahn Jul 2021

Frameaxis: Characterizing Microframe Bias And Intensity With Word Embedding, Haewoon Kwak, Jisun An, Elise Jing Jing, Yong-Yeol Ahn

Research Collection School Of Computing and Information Systems

Framing is a process of emphasizing a certain aspect of an issue over the others, nudging readers or listeners towards different positions on the issue even without making a biased argument. Here, we propose FrameAxis, a method for characterizing documents by identifying the most relevant semantic axes (“microframes”) that are overrepresented in the text using word embedding. Our unsupervised approach can be readily applied to large datasets because it does not require manual annotations. It can also provide nuanced insights by considering a rich set of semantic axes. FrameAxis is designed to quantitatively tease out two important dimensions of how …


Addressing The ‘Unseens’: Digital Wellbeing In The Remote Workplace, Holtjona Galanxhi, Fiona Fui-Hoon Nah Jul 2021

Addressing The ‘Unseens’: Digital Wellbeing In The Remote Workplace, Holtjona Galanxhi, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

The ubiquity of sophisticated devices, along with uninterrupted access to the Internet and organizational computerized systems, allows for the “anyplace” workplace to be established. Technology has the potential to deliberately or inadvertently impact psychological wellbeing. Specific psychological demands are inadvertently imposed on remote employees whose permanent online presence is required. Hence, it is important to understand factors affecting digital wellbeing and steps that can be taken to maximize the wellbeing of remote employees. This paper provides suggestions for future research on studying the digital wellbeing of (fully or partially) remote employees. A research framework is proposed to demonstrate the different …


Make It Easy: An Effective End-To-End Entity Alignment Framework, Congcong Ge, Xiaoze Liu, Lu Chen Chen, Baihua Zheng, Yunjun Gao Jul 2021

Make It Easy: An Effective End-To-End Entity Alignment Framework, Congcong Ge, Xiaoze Liu, Lu Chen Chen, Baihua Zheng, Yunjun Gao

Research Collection School Of Computing and Information Systems

Entity alignment (EA) is a prerequisite for enlarging the coverage of a unified knowledge graph. Previous EA approaches either restrain the performance due to inadequate information utilization or need labor-intensive pre-processing to get external or reliable information to perform the EA task. This paper proposes EASY, an effective end-to-end EA framework, which is able to (i) remove the labor-intensive pre-processing by fully discovering the name information provided by the entities themselves; and (ii) jointly fuse the features captured by the names of entities and the structural information of the graph to improve the EA results. Specifically, EASY first introduces NEAP, …


Variational Learning From Implicit Bandit Feedback, Quoc Tuan Truong, Hady W. Lauw Jul 2021

Variational Learning From Implicit Bandit Feedback, Quoc Tuan Truong, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Recommendations are prevalent in Web applications (e.g., search ranking, item recommendation, advertisement placement). Learning from bandit feedback is challenging due to the sparsity of feedback limited to system-provided actions. In this work, we focus on batch learning from logs of recommender systems involving both bandit and organic feedbacks. We develop a probabilistic framework with a likelihood function for estimating not only explicit positive observations but also implicit negative observations inferred from the data. Moreover, we introduce a latent variable model for organic-bandit feedbacks to robustly capture user preference distributions. Next, we analyze the behavior of the new likelihood under two …


A Differentially Private Task Planning Framework For Spatial Crowdsourcing, Qian Tao, Yongxin Tong, Shuyuan Li, Yuxiang Zeng, Zimu Zhou, Ke Xu Jul 2021

A Differentially Private Task Planning Framework For Spatial Crowdsourcing, Qian Tao, Yongxin Tong, Shuyuan Li, Yuxiang Zeng, Zimu Zhou, Ke Xu

Research Collection School Of Computing and Information Systems

Spatial crowdsourcing has stimulated various new applications such as taxi calling and food delivery. A key enabler for these spatial crowdsourcing based applications is to plan routes for crowd workers to execute tasks given diverse requirements of workers and the spatial crowdsourcing platform. Despite extensive studies on task planning in spatial crowdsourcing, few have accounted for the location privacy of tasks, which may be misused by an untrustworthy platform. In this paper, we explore efficient task planning for workers while protecting the locations of tasks. Specifically, we define the Privacy-Preserving Task Planning (PPTP) problem, which aims at both total revenue …


A Coprocessor-Based Introspection Framework Via Intel Management Engine, Lei Zhou, Fengwei Zhang, Jidong Xiao, Kevin Leach, Westley Weimer, Xuhua Ding, Guojun Wang Jul 2021

A Coprocessor-Based Introspection Framework Via Intel Management Engine, Lei Zhou, Fengwei Zhang, Jidong Xiao, Kevin Leach, Westley Weimer, Xuhua Ding, Guojun Wang

Research Collection School Of Computing and Information Systems

During the past decade, virtualization-based (e.g., virtual machine introspection) and hardware-assisted approaches (e.g., x86 SMM and ARM TrustZone) have been used to defend against low-level malware such as rootkits. However, these approaches either require a large Trusted Computing Base (TCB) or they must share CPU time with the operating system, disrupting normal execution. In this article, we propose an introspection framework called NIGHTHAWK that transparently checks system integrity and monitor the runtime state of target system. NIGHTHAWK leverages the Intel Management Engine (IME), a co-processor that runs in isolation from the main CPU. By using the IME, our approach has …


Dehumor: Visual Analytics For Decomposing Humor, Xingbo Wang, Yao Ming, Tongshuang Wu, Haipeng Zeng, Yong Wang, Huamin Qu Jul 2021

Dehumor: Visual Analytics For Decomposing Humor, Xingbo Wang, Yao Ming, Tongshuang Wu, Haipeng Zeng, Yong Wang, Huamin Qu

Research Collection School Of Computing and Information Systems

Despite being a critical communication skill, grasping humor is challenginga successful use of humor requires a mixture of both engaging content build-up and an appropriate vocal delivery (e.g., pause). Prior studies on computational humor emphasize the textual and audio features immediately next to the punchline, yet overlooking longer-term context setup. Moreover, the theories are usually too abstract for understanding each concrete humor snippet. To fill in the gap, we develop DeHumor, a visual analytical system for analyzing humorous behaviors in public speaking. To intuitively reveal the building blocks of each concrete example, DeHumor decomposes each humorous video into multimodal features …


Integrated Framework For Developing Instructional Videos For Foundational Computing Courses, Kyong Jin Shim, Gottipati Swapna, Yi Meng Lau Jul 2021

Integrated Framework For Developing Instructional Videos For Foundational Computing Courses, Kyong Jin Shim, Gottipati Swapna, Yi Meng Lau

Research Collection School Of Computing and Information Systems

Instructional videos are widely used in higher education due to their effectiveness and flexibility of personalized learning features. Computing courses usually focuses on programming, user interface design, server connectivity, data storage, and architecture, among others. The design of instructional videos varies in not only the course content but also the style of content creation. We propose an integrated framework, Computing Videos Design Framework (CVDF), for designing and developing instructional videos for computing courses. CVDF combines the cognitive skills from Bloom’s taxonomy, video design principles, and course learning outcomes for designing different types of instructional videos. We apply the framework to …


Page: A Simple And Optimal Probabilistic Gradient Estimator For Nonconvex Optimization, Zhize Li, Hongyan Bao, Xiangliang Zhang, Peter Richtarik Jul 2021

Page: A Simple And Optimal Probabilistic Gradient Estimator For Nonconvex Optimization, Zhize Li, Hongyan Bao, Xiangliang Zhang, Peter Richtarik

Research Collection School Of Computing and Information Systems

In this paper, we propose a novel stochastic gradient estimator---ProbAbilistic Gradient Estimator (PAGE)---for nonconvex optimization. PAGE is easy to implement as it is designed via a small adjustment to vanilla SGD: in each iteration, PAGE uses the vanilla minibatch SGD update with probability $p_t$ or reuses the previous gradient with a small adjustment, at a much lower computational cost, with probability $1-p_t$. We give a simple formula for the optimal choice of $p_t$. Moreover, we prove the first tight lower bound $\Omega(n+\frac{\sqrt{n}}{\epsilon^2})$ for nonconvex finite-sum problems, which also leads to a tight lower bound $\Omega(b+\frac{\sqrt{b}}{\epsilon^2})$ for nonconvex online problems, where …


Paying Attention To Video Object Pattern Understanding, Wenguan Wang, Jianbing Shen, Xiankai Lu, Steven C. H. Hoi, Haibin Ling Jul 2021

Paying Attention To Video Object Pattern Understanding, Wenguan Wang, Jianbing Shen, Xiankai Lu, Steven C. H. Hoi, Haibin Ling

Research Collection School Of Computing and Information Systems

This paper conducts a systematic study on the role of visual attention in video object pattern understanding. By elaborately annotating three popular video segmentation datasets (DAVIS) with dynamic eye-tracking data in the unsupervised video object segmentation (UVOS) setting. For the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgments during dynamic, task-driven viewing. Such novel observations provide an in-depth insight of the underlying rationale behind video object pattens. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention …


Exploring Cross-Modality Utilization In Recommender Systems, Quoc Tuan Truong, Aghiles Salah, Thanh-Binh Tran, Jingyao Guo, Hady W. Lauw Jul 2021

Exploring Cross-Modality Utilization In Recommender Systems, Quoc Tuan Truong, Aghiles Salah, Thanh-Binh Tran, Jingyao Guo, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Multimodal recommender systems alleviate the sparsity of historical user-item interactions. They are commonly catalogued based on the type of auxiliary data (modality) they leverage, such as preference data plus user-network (social), user/item texts (textual), or item images (visual) respectively. One consequence of this categorization is the tendency for virtual walls to arise between modalities. For instance, a study involving images would compare to only baselines ostensibly designed for images. However, a closer look at existing models' statistical assumptions about any one modality would reveal that many could work just as well with other modalities. Therefore, we pursue a systematic investigation …


Boosting Video Representation Learning With Multi-Faceted Integration, Zhaofan Qiu, Yao Ting, Chong-Wah Ngo, Xiao-Ping Zhang, Dong Wu, Tao Mei Jun 2021

Boosting Video Representation Learning With Multi-Faceted Integration, Zhaofan Qiu, Yao Ting, Chong-Wah Ngo, Xiao-Ping Zhang, Dong Wu, Tao Mei

Research Collection School Of Computing and Information Systems

Video content is multifaceted, consisting of objects, scenes, interactions or actions. The existing datasets mostly label only one of the facets for model training, resulting in the video representation that biases to only one facet depending on the training dataset. There is no study yet on how to learn a video representation from multifaceted labels, and whether multifaceted information is helpful for video representation learning. In this paper, we propose a new learning framework, MUlti-Faceted Integration (MUFI), to aggregate facets from different datasets for learning a representation that could reflect the full spectrum of video content. Technically, MUFI formulates the …


How-To Present News On Social Media: A Causal Analysis Of Editing News Headlines For Boosting User Engagement, Kunwoo Park, Haewoon Kwak, Jisun An, Sanjay Chawla Jun 2021

How-To Present News On Social Media: A Causal Analysis Of Editing News Headlines For Boosting User Engagement, Kunwoo Park, Haewoon Kwak, Jisun An, Sanjay Chawla

Research Collection School Of Computing and Information Systems

To reach a broader audience and optimize traffic toward news articles, media outlets commonly run social media accounts and share their content with a short text summary. Despite its importance of writing a compelling message in sharing articles, the research community does not own a sufficient understanding of what kinds of editing strategies effectively promote audience engagement. In this study, we aim to fill the gap by analyzing media outlets' current practices using a data-driven approach. We first build a parallel corpus of original news articles and their corresponding tweets that eight media outlets shared. Then, we explore how those …


On M-Impact Regions And Standing Top-K Influence Problems, Bo Tang, Kyriakos Mouratidis, Mingji Han Jun 2021

On M-Impact Regions And Standing Top-K Influence Problems, Bo Tang, Kyriakos Mouratidis, Mingji Han

Research Collection School Of Computing and Information Systems

In this paper, we study the ��-impact region problem (mIR). In a context where users look for available products with top-�� queries, mIR identifies the part of the product space that attracts the most user attention. Specifically, mIR determines the kind of attribute values that lead a (new or existing) product to the top-�� result for at least a fraction of the user population. mIR has several applications, ranging from effective marketing to product improvement. Importantly, it also leads to (exact and efficient) solutions for standing top-�� impact problems, which were previously solved heuristically only, or whose current solutions face …


Riding Through The Silver Tsunami: A Data Driven Approach To Improve Senior Citizens’ Engagement With Community Senior Activity Centres, Joshua Jie Feng Lam, Hwee-Pink Tan Jun 2021

Riding Through The Silver Tsunami: A Data Driven Approach To Improve Senior Citizens’ Engagement With Community Senior Activity Centres, Joshua Jie Feng Lam, Hwee-Pink Tan

Research Collection School Of Computing and Information Systems

In Singapore, 1 in 4 persons will be elderly by 2030 In preparation for the Silver Tsunami, the Singapore government and community care providers have collaborations to promote active, independent living amongst elders Current implementation of data driven population health is focused on well being indices using data collected from the general population There is no literature on the use of data analytics in assessing elder


Hierarchical Reinforcement Learning: A Comprehensive Survey, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan, Chai Quek Jun 2021

Hierarchical Reinforcement Learning: A Comprehensive Survey, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan, Chai Quek

Research Collection School Of Computing and Information Systems

Hierarchical Reinforcement Learning (HRL) enables autonomous decomposition of challenging long-horizon decision-making tasks into simpler subtasks. During the past years, the landscape of HRL research has grown profoundly, resulting in copious approaches. A comprehensive overview of this vast landscape is necessary to study HRL in an organized manner. We provide a survey of the diverse HRL approaches concerning the challenges of learning hierarchical policies, subtask discovery, transfer learning, and multi-agent learning using HRL. The survey is presented according to a novel taxonomy of the approaches. Based on the survey, a set of important open problems is proposed to motivate the future …


On Predicting Personal Values Of Social Media Users Using Community-Specific Language Features And Personal Value Correlation, Amila Silva, Pei Chi Lo, Ee-Peng Lim Jun 2021

On Predicting Personal Values Of Social Media Users Using Community-Specific Language Features And Personal Value Correlation, Amila Silva, Pei Chi Lo, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Personal values have significant influence on individuals’ behaviors, preferences, and decision making. It is therefore not a surprise that personal values of a person could influence his or her social media content and activities. Instead of getting users to complete personal value questionnaire, researchers have looked into a non-intrusive and highly scalable approach to predict personal values using user-generated social media data. Nevertheless, geographical differences in word usage and profile information are issues to be addressed when designing such prediction models. In this work, we focus on analyzing Singapore users’ personal values, and developing effective models to predict their personal …


Gpu-Accelerated Graph Label Propagation For Real-Time Fraud Detection, Chang Ye, Yuchen Li, Bingsheng He, Zhao Li, Jianling Sun Jun 2021

Gpu-Accelerated Graph Label Propagation For Real-Time Fraud Detection, Chang Ye, Yuchen Li, Bingsheng He, Zhao Li, Jianling Sun

Research Collection School Of Computing and Information Systems

Fraud detection is a pressing challenge for most financial and commercial platforms. In this paper, we study the processing pipeline of fraud detection in a large e-commerce platform of TaoBao. Graph label propagation (LP) is a core component in this pipeline to detect suspicious clusters from the user-interaction graph. Furthermore, the run-time of the LP component occupies 75% overhead of TaoBao’s automated detection pipeline. To enable real-time fraud detection, we propose a GPU-based framework, called GLP, to support large-scale LP workloads in enterprises. We have identified two key challenges when integrating GPU acceleration into TaoBao’s data processing pipeline: (1) programmability …


Cache-Efficient Fork-Processing Patterns On Large Graphs, Shengliang Lu, Shixuan Sun, Johns Paul, Yuchen Li, Bingsheng He Jun 2021

Cache-Efficient Fork-Processing Patterns On Large Graphs, Shengliang Lu, Shixuan Sun, Johns Paul, Yuchen Li, Bingsheng He

Research Collection School Of Computing and Information Systems

As large graph processing emerges, we observe a costly fork-processing pattern (FPP) that is common in many graph algorithms. The unique feature of the FPP is that it launches many independent queries from different source vertices on the same graph. For example, an algorithm in analyzing the network community profile can execute Personalized PageRanks that start from tens of thousands of source vertices at the same time. We study the efficiency of handling FPPs in state-of-the-art graph processing systems on multi-core architectures, including Ligra, Gemini, and GraphIt. We find that those systems suffer from severe cache miss penalty because of …


Minimum Coresets For Maxima Representation Of Multidimensional Data, Yanhao Wang, Michael Mathioudakis, Yuchen Li, Kian-Lee Tan Jun 2021

Minimum Coresets For Maxima Representation Of Multidimensional Data, Yanhao Wang, Michael Mathioudakis, Yuchen Li, Kian-Lee Tan

Research Collection School Of Computing and Information Systems

Coresets are succinct summaries of large datasets such that, for a given problem, the solution obtained from a coreset is provably competitive with the solution obtained from the full dataset. As such, coreset-based data summarization techniques have been successfully applied to various problems, e.g., geometric optimization, clustering, and approximate query processing, for scaling them up to massive data. In this paper, we study coresets for the maxima representation of multidimensional data: Given a set �� of points in R �� , where �� is a small constant, and an error parameter �� ∈ (0, 1), a subset �� ⊆ �� …