Upsizer: Synthetically Scaling An Empirical Relational Database,
2013
National University of Singapore
Upsizer: Synthetically Scaling An Empirical Relational Database, Y. C. Tay, Bing Tian Dai, Daniel T. Wang, Eldora Y. Sun, Yong Lin, Yuting Lin
Research Collection School Of Computing and Information Systems
The TPC benchmarks have helped users evaluate database system performance at different scales. Although each benchmark is domain-specific, it is not equally relevant to different applications in the same domain. The present proliferation of applications also leaves many of them uncovered by the very limited number of current TPC benchmarks. There is therefore a need to develop tools for application-specific database benchmarking. This paper presents UpSizeR, a software that addresses the Dataset Scaling Problem: Given an empirical set of relational tables D and a scale factor s, generate a database state e D that is similar to D but s …
A Link-Bridged Topic Model For Cross-Domain Document Classification,
2013
Singapore Management University
A Link-Bridged Topic Model For Cross-Domain Document Classification, Pei Yang, Wei Gao, Qi Tan, Kam-Fai Wong
Research Collection School Of Computing and Information Systems
Transfer learning utilizes labeled data available from some related domain (source domain) for achieving effective knowledge transformation to the target domain. However, most state-of-the-art cross-domain classification methods treat documents as plain text and ignore the hyperlink (or citation) relationship existing among the documents. In this paper, we propose a novel cross-domain document classification approach called Link-Bridged Topic model (LBT). LBT consists of two key steps. Firstly, LBT utilizes an auxiliary link network to discover the direct or indirect co-citation relationship among documents by embedding the background knowledge into a graph kernel. The mined co-citation relationship is leveraged to bridge the …
Social Listening For Customer Acquisition,
2013
Singapore Management University
Social Listening For Customer Acquisition, Juan Du, Biying Tan, Feida Zhu, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Social network analysis has received much attention from corporations recently. Corporations are trying to utilize social media platforms such as Twitter, Facebook and Sina Weibo to expand their own markets. Our system is an online tool to assist these corporations to 1) find potential customers, and 2) track a list of users by specific events from social networks. We employ both textual and network information, and thus produce a keyword-based relevance score for each user in pre-defined dimensions, which indicates the probability of the adoption of a product. Based on the score and its trend, out tool is able to …
Information Vs Interaction: An Alternative User Ranking Model For Social Networks,
2013
Singapore Management University
Information Vs Interaction: An Alternative User Ranking Model For Social Networks, Wei Xie, Ai Phuong Hoang, Feida Zhu, Ee Peng Lim
Research Collection School Of Computing and Information Systems
The recent years have seen an unprecedented boom of social network services, such as Twitter, which boasts over 200 million users. In such big social platforms, the influential users are ideal targets for viral marketing to potentially reach an audience of maximal size. Most proposed algorithms rely on the linkage structure of the respective underlying network to determine the information flow and hence indicate a users influence. From social interaction perspective, we built a model based on the dynamic user interactions constantly taking place on top of these linkage structures. In particular, in the Twitter setting we supposed a principle …
Mining Fraudulent Patterns In Online Advertising,
2013
Singapore Management University
Mining Fraudulent Patterns In Online Advertising, Richard J. Oentaryo, Ee-Peng Lim
Research Collection School Of Computing and Information Systems
Advances in web technologies have rendered onlineadvertising as an effective means for small and large businesses to target different market segments on the fly. Online advertising is a huge industry. According to Gartner Inc., worldwide online advertising revenue is projected tohit $11.4 billion in 2013, up from $9.6 billion in 2012. Global revenue will also reach $24.5 billion in 2016, with online advertising creating opportunities for app developers, advertising networks, and service providersin various regions. An online advertising ecosystem is typically coordinated by an advertising commissioner, acting as a broker between advertisers and content publishers. An advertiser plans a budget, …
Electroweak Measurements In Electron-Positron Collisions At W-Boson-Pair Energies At Lep,
2013
Singapore Management University
Electroweak Measurements In Electron-Positron Collisions At W-Boson-Pair Energies At Lep, S. Schael, Manoj Thulasidas
Research Collection School Of Computing and Information Systems
Electroweak measurements performed with data taken at the electron–positron collider LEP at CERN from 1995 to 2000 are reported. The combined data set considered in this report corresponds to a total luminosity of about 3 fb −1 collected by the four LEP experiments ALEPH, DELPHI, L3 and OPAL, at centre-of-mass energies ranging from 130 GeV to 209 GeV. Combining the published results of the four LEP experiments, the measurements include total and differential cross-sections in photon-pair, fermion-pair and four-fermion production, the latter resulting from both double-resonant WW and ZZ production as well as singly resonant production. Total and differential cross-sections …
Second Order Online Collaborative Filtering,
2013
Singapore Management University
Second Order Online Collaborative Filtering, Jing Lu, Steven C. H. Hoi, Jialei Wang, Peilin Zhao
Research Collection School Of Computing and Information Systems
Collaborative Filtering (CF) is one of the most successful learning techniques in building real-world recommender systems. Traditional CF algorithms are often based on batch machine learning methods which suffer from several critical drawbacks, e.g., extremely expensive model retraining cost whenever new samples arrive, unable to capture the latest change of user preferences over time, and high cost and slow reaction to new users or products extension. Such limitations make batch learning based CF methods unsuitable for real-world online applications where data often arrives sequentially and user preferences may change dynamically and rapidly. To address these limitations, we investigate online collaborative …
Social Informatics,
2013
Singapore Management University
Social Informatics, Adam Jatowt, Ee-Peng Lim, Ying Ding, Asako Miura, Taro Tezuka, Gael Dias, Katsumi Tanaka, Andrew J. Flanagin, Bing Tian Dai
Research Collection School Of Computing and Information Systems
This book constitutes the proceedings of the 5th International Conference on Social Informatics, SocInfo 2013, held in Kyoto, Japan, in November 2013. The 23 full papers, 15 short papers, and three poster papers included in this volume were carefully reviewed and selected from 103 submissions. The papers present original research work on studying the interplay between socially-centric platforms and social phenomena.
Using Micro-Reviews To Select An Efficient Set Of Reviews,
2013
Singapore Management University
Using Micro-Reviews To Select An Efficient Set Of Reviews, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas
Research Collection School Of Computing and Information Systems
Online reviews are an invaluable resource for web users trying to make decisions regarding products or services. However, the abundance of review content, as well as the unstructured, lengthy, and verbose nature of reviews make it hard for users to locate the appropriate reviews, and distill the useful information. With the recent growth of social networking and micro-blogging services, we observe the emergence of a new type of online review content, consisting of bite-sized, 140 character-long reviews often posted reactively on the spot via mobile devices. These micro-reviews are short, concise, and focused, nicely complementing the lengthy, elaborate, and verbose …
Efficient Index-Based Approaches For Skyline Queries In Location-Based Applications,
2013
Amazon.com
Efficient Index-Based Approaches For Skyline Queries In Location-Based Applications, Ken C. K. Lee, Baihua Zheng, Cindy Chen, Chi-Yin Chow
Research Collection School Of Computing and Information Systems
Enriching many location-based applications, various new skyline queries are proposed and formulated based on the notion of locational dominance, which extends conventional one by taking objects' nearness to query positions into account additional to objects' nonspatial attributes. To answer a representative class of skyline queries for location-based applications efficiently, this paper presents two index-based approaches, namely, augmented R-tree and dominance diagram. Augmented R-tree extends R-tree by including aggregated nonspatial attributes in index nodes to enable dominance checks during index traversal. Dominance diagram is a solution-based approach, by which each object is associated with a precomputed nondominance scope wherein query points …
Classification In P2p Networks With Cascade Support Vendor Machines,
2013
Nanyang Technological University
Classification In P2p Networks With Cascade Support Vendor Machines, Hock Hee Ang, Vivekanand Gopalkrishnan, Steven C. H. Hoi, Wee-Keong Ng
Research Collection School Of Computing and Information Systems
Classification in Peer-to-Peer (P2P) networks is important to many real applications, such as distributed intrusion detection, distributed recommendation systems, and distributed antispam detection. However, it is very challenging to perform classification in P2P networks due to many practical issues, such as scalability, peer dynamism, and asynchronism. This article investigates the practical techniques of constructing Support Vector Machine (SVM) classifiers in the P2P networks. In particular, we demonstrate how to efficiently cascade SVM in a P2P network with the use of reduced SVM. In addition, we propose to fuse the concept of cascade SVM with bootstrap aggregation to effectively balance the …
Multimedia Modeling,
2013
Singapore Management University
Multimedia Modeling, Chong-Wah Ngo, Klaus Schoeffmann, Yiannis Andreopoulos, Christian Breiteneder
Research Collection School Of Computing and Information Systems
Multimedia modeling aims to study computational models for addressing real-world multimedia problems from various perspectives, including information fusion, perceptual understanding, performance evaluation and social media. The topic becomes increasingly important with the massive amount of data available over the Internet, representing different pieces of information in heterogeneous forms that need to be consolidated before being used for multimedia problems. On the other hand, the advancement in technologies such as mobile and sensing devices drive the needs for revisiting the existing models for not only dealing with audio-visual cues but also incorporating various sensory modalities that have potential in providing cheaper …
Why Do I Retweet It? An Information Propagation Model For Microblogs,
2013
Singapore Management University
Why Do I Retweet It? An Information Propagation Model For Microblogs, Fabio Pezzoni, Jisun An, Andrea Passarella, Jon Crowcroft, Marco Conti
Research Collection School Of Computing and Information Systems
Microblogging platforms are Web 2.0 services that represent a suitable environment for studying how information is propagated in social networks and how users can become influential. In this work we analyse the impact of the network features and of the users' behaviour on the information diffusion. Our analysis highlights a strong relation between the level of visibility of a message in the flow of information seen by a user and the probability that the user further disseminates the message. In addition, we also highlight the existence of other latent factors that impact on the dissemination probability, correlated with the properties …
City Notifications As A Data Source For Traffic Management,
2013
Wright State University - Main Campus
City Notifications As A Data Source For Traffic Management, Pramod Anantharam, Biplav Srivastava
Kno.e.sis Publications
A common problem for cities of developing countries like India in managing traffic is the lack of basic automated instrumentation to track road conditions or vehicle locations. Still, to help their citizens make informed travel decisions based on changing city dynamics; many cities have an authorized, city-initiated, notification service in place to alert subscribing commuters about road conditions. Here, alternative means may be used to create informal textual notifications e.g., inputs from field personnel, citizen updates, and pre-authorized events from city calendar. In this paper, we show that collections of such notifications, when processed with information extraction techniques, can turn …
Suddenly...I'M Consulting On Data Management Plans! Data Management Plan Consultant Checklist,
2013
University of Nebraska - Lincoln
Suddenly...I'M Consulting On Data Management Plans! Data Management Plan Consultant Checklist, Kiyomi D. Deards
University of Nebraska-Lincoln Libraries: Presentations
This webinar will outline the most important questions to ask, and the best resources available, for those who "suddenly" will be consulting on data management plans.
Improving Reuse In Software Development For The Life Sciences,
2013
Purdue University
Improving Reuse In Software Development For The Life Sciences, Nicholas Vincent Iannotti
Open Access Dissertations
The last several years have seen unprecedented advancements in the application of technology to the life sciences, particularly in the area of data generation. Novel scientific insights are now often driven primarily by software development supporting new multidisciplinary and increasingly multifaceted data analysis. However, despite the availability of tools such as best practice frameworks, the current rate of software development is not able to keep up with the needs of scientists. This bottleneck in software development is largely due to code reuse generally not being applied in practice.
This dissertation presents Legwork, a class library of reuse-optimized design pattern implementations …
Search Tool That Utilizes Scientific Metadata Matched Against User-Entered Parameters,
2013
Portland State University
Search Tool That Utilizes Scientific Metadata Matched Against User-Entered Parameters, Veronika Margaret Megler, David Maier
Computer Science Faculty Publications and Presentations
A method for providing proximate dataset recommendations can begin with the creation of metadata records corresponding to datasets that represent scientific data by a scientific dataset search tool. The metadata records can conform to a standardized structural definition, and may be hierarchical. Values for the data elements of the metadata records can be contained within the datasets. Metadata records with a value that is proximate to a user-entered search parameter can be identified. A proximity score can be calculated for each identified metadata record. The proximity score can express a relevance of the corresponding dataset to the user-entered search parameters. …
Designing Mobile Educational Games On Voter‟S Education: A Tale Of Three Engines,
2013
Ateneo de Manila University
Designing Mobile Educational Games On Voter‟S Education: A Tale Of Three Engines, Ma. Regina Justina E. Estuar, Nadia Rowena C. Leetian, Michael B. Syson
Department of Information Systems & Computer Science Faculty Publications
The rapid growth of mobile learning is influenced by the ability to access learning content anytime and anywhere. The on demand capability is available because mobile devices allow for convergence of internet and communications technologies. At the same time, the availability of engines makes development of mobile applications faster and seamless. However, not all mobile development engines are alike. This paper discusses on the development of mobile learning applications using mobile development engines in teaching Filipinos on responsible voting. Specifically, this paper discusses how AndEngine, Ren’Py, and homegrown Usbong were used to develop a mobile board game and a mobile …
Modeling Interaction Features For Debate Side Clustering,
2013
Singapore Management University
Modeling Interaction Features For Debate Side Clustering, Minghui Qiu, Liu Yang, Jing Jiang
Research Collection School Of Computing and Information Systems
Online discussion forums are popular social media platforms for users to express their opinions and discuss controversial issues with each other. To automatically identify the sides/stances of posts or users from textual content in forums is an important task to help mine online opinions. To tackle the task, it is important to exploit user posts that implicitly contain support and dispute (interaction) information. The challenge we face is how to mine such interaction information from the content of posts and how to use them to help identify stances. This paper proposes a two-stage solution based on latent variable models: an …
On Effects Of Visual Query Complexity,
2013
Singapore Management University
On Effects Of Visual Query Complexity, Jialie Shen, Cheng Zhiyong
Research Collection School Of Computing and Information Systems
As an effective technique to manage large scale image collections, content-based image retrieval (CBIR) has been received great attentions and became a very active research domain in recent years. While assessing system performance is one of the key factors for the related technological advancement, relatively little attention has been paid to model and analyze test queries. This paper documents a study on the problem of determining visual query complexity as a measure for predicting image retrieval performance. We propose a quantitative metric for measuring complexity of image queries for content-based image search engine. A set of experiments are carried out …
