Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2013

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 31 - 60 of 252

Full-Text Articles in Databases and Information Systems

Social Sensing For Urban Crisis Management: The Case Of Singapore Haze, Philips Kokoh Prasetyo, Ming Gao, Ee Peng Lim, Christie N. Scollon Nov 2013

Social Sensing For Urban Crisis Management: The Case Of Singapore Haze, Philips Kokoh Prasetyo, Ming Gao, Ee Peng Lim, Christie N. Scollon

Research Collection School Of Computing and Information Systems

Sensing social media for trends and events has become possible as increasing number of users rely on social media to share information. In the event of a major disaster or social event, one can therefore study the event quickly by gathering and analyzing social media data. One can also design appropriate responses such as allocating resources to the affected areas, sharing event related information, and managing public anxiety. Past research on social event studies using social media often focused on one type of data analysis (e.g., hashtag clusters, diffusion of events, influential users, etc.) on a single social media data …


Predicting Best Answerers For New Questions: An Approach Leveraging Topic Modeling And Collaborative Voting, Yuan Tian, Pavneet Singh Kochhar, Ee Peng Lim, Feida Zhu, David Lo Nov 2013

Predicting Best Answerers For New Questions: An Approach Leveraging Topic Modeling And Collaborative Voting, Yuan Tian, Pavneet Singh Kochhar, Ee Peng Lim, Feida Zhu, David Lo

Research Collection School Of Computing and Information Systems

Community Question Answering (CQA) sites are becoming increasingly important source of information where users can share knowledge on various topics. Although these platforms bring new opportunities for users to seek help or provide solutions, they also pose many challenges with the ever growing size of the community. The sheer number of questions posted everyday motivates the problem of routing questions to the appropriate users who can answer them. In this paper, we propose an approach to predict the best answerer for a new question on CQA site. Our approach considers both user interest and user expertise relevant to the topics …


Covariance Selection By Thresholding The Sample Correlation Matrix, Binyan Jiang Nov 2013

Covariance Selection By Thresholding The Sample Correlation Matrix, Binyan Jiang

Research Collection School Of Computing and Information Systems

This article shows that when the nonzero coefficients of the population correlation matrix are all greater in absolute value than (C1logp/n)1/2 for some constant C1, we can obtain covariance selection consistency by thresholding the sample correlation matrix. Furthermore, the rate (logp/n)1/2 is shown to be optimal.


Predicting User's Political Party Using Ideological Stances, Swapna Gottopati, Minghui Qiu, Liu Yang, Feida Zhu, Jing Jiang Nov 2013

Predicting User's Political Party Using Ideological Stances, Swapna Gottopati, Minghui Qiu, Liu Yang, Feida Zhu, Jing Jiang

Research Collection School Of Computing and Information Systems

Predicting users political party in social media has important impacts on many real world applications such as targeted advertising, recommendation and personalization. Several political research studies on it indicate that political parties’ ideological beliefs on sociopolitical issues may influence the users political leaning. In our work, we exploit users’ ideological stances on controversial issues to predict political party of online users. We propose a collaborative filtering approach to solve the data sparsity problem of users stances on ideological topics and apply clustering method to group the users with the same party. We evaluated several state-of-the-art methods for party prediction task …


Upsizer: Synthetically Scaling An Empirical Relational Database, Y. C. Tay, Bing Tian Dai, Daniel T. Wang, Eldora Y. Sun, Yong Lin, Yuting Lin Nov 2013

Upsizer: Synthetically Scaling An Empirical Relational Database, Y. C. Tay, Bing Tian Dai, Daniel T. Wang, Eldora Y. Sun, Yong Lin, Yuting Lin

Research Collection School Of Computing and Information Systems

The TPC benchmarks have helped users evaluate database system performance at different scales. Although each benchmark is domain-specific, it is not equally relevant to different applications in the same domain. The present proliferation of applications also leaves many of them uncovered by the very limited number of current TPC benchmarks. There is therefore a need to develop tools for application-specific database benchmarking. This paper presents UpSizeR, a software that addresses the Dataset Scaling Problem: Given an empirical set of relational tables D and a scale factor s, generate a database state e D that is similar to D but s …


A Link-Bridged Topic Model For Cross-Domain Document Classification, Pei Yang, Wei Gao, Qi Tan, Kam-Fai Wong Nov 2013

A Link-Bridged Topic Model For Cross-Domain Document Classification, Pei Yang, Wei Gao, Qi Tan, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

Transfer learning utilizes labeled data available from some related domain (source domain) for achieving effective knowledge transformation to the target domain. However, most state-of-the-art cross-domain classification methods treat documents as plain text and ignore the hyperlink (or citation) relationship existing among the documents. In this paper, we propose a novel cross-domain document classification approach called Link-Bridged Topic model (LBT). LBT consists of two key steps. Firstly, LBT utilizes an auxiliary link network to discover the direct or indirect co-citation relationship among documents by embedding the background knowledge into a graph kernel. The mined co-citation relationship is leveraged to bridge the …


Social Listening For Customer Acquisition, Juan Du, Biying Tan, Feida Zhu, Ee-Peng Lim Nov 2013

Social Listening For Customer Acquisition, Juan Du, Biying Tan, Feida Zhu, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Social network analysis has received much attention from corporations recently. Corporations are trying to utilize social media platforms such as Twitter, Facebook and Sina Weibo to expand their own markets. Our system is an online tool to assist these corporations to 1) find potential customers, and 2) track a list of users by specific events from social networks. We employ both textual and network information, and thus produce a keyword-based relevance score for each user in pre-defined dimensions, which indicates the probability of the adoption of a product. Based on the score and its trend, out tool is able to …


Information Vs Interaction: An Alternative User Ranking Model For Social Networks, Wei Xie, Ai Phuong Hoang, Feida Zhu, Ee Peng Lim Nov 2013

Information Vs Interaction: An Alternative User Ranking Model For Social Networks, Wei Xie, Ai Phuong Hoang, Feida Zhu, Ee Peng Lim

Research Collection School Of Computing and Information Systems

The recent years have seen an unprecedented boom of social network services, such as Twitter, which boasts over 200 million users. In such big social platforms, the influential users are ideal targets for viral marketing to potentially reach an audience of maximal size. Most proposed algorithms rely on the linkage structure of the respective underlying network to determine the information flow and hence indicate a users influence. From social interaction perspective, we built a model based on the dynamic user interactions constantly taking place on top of these linkage structures. In particular, in the Twitter setting we supposed a principle …


Mining Fraudulent Patterns In Online Advertising, Richard J. Oentaryo, Ee-Peng Lim Nov 2013

Mining Fraudulent Patterns In Online Advertising, Richard J. Oentaryo, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Advances in web technologies have rendered onlineadvertising as an effective means for small and large businesses to target different market segments on the fly. Online advertising is a huge industry. According to Gartner Inc., worldwide online advertising revenue is projected tohit $11.4 billion in 2013, up from $9.6 billion in 2012. Global revenue will also reach $24.5 billion in 2016, with online advertising creating opportunities for app developers, advertising networks, and service providersin various regions. An online advertising ecosystem is typically coordinated by an advertising commissioner, acting as a broker between advertisers and content publishers. An advertiser plans a budget, …


Electroweak Measurements In Electron-Positron Collisions At W-Boson-Pair Energies At Lep, S. Schael, Manoj Thulasidas Nov 2013

Electroweak Measurements In Electron-Positron Collisions At W-Boson-Pair Energies At Lep, S. Schael, Manoj Thulasidas

Research Collection School Of Computing and Information Systems

Electroweak measurements performed with data taken at the electron–positron collider LEP at CERN from 1995 to 2000 are reported. The combined data set considered in this report corresponds to a total luminosity of about 3 fb −1 collected by the four LEP experiments ALEPH, DELPHI, L3 and OPAL, at centre-of-mass energies ranging from 130 GeV to 209 GeV. Combining the published results of the four LEP experiments, the measurements include total and differential cross-sections in photon-pair, fermion-pair and four-fermion production, the latter resulting from both double-resonant WW and ZZ production as well as singly resonant production. Total and differential cross-sections …


Second Order Online Collaborative Filtering, Jing Lu, Steven C. H. Hoi, Jialei Wang, Peilin Zhao Nov 2013

Second Order Online Collaborative Filtering, Jing Lu, Steven C. H. Hoi, Jialei Wang, Peilin Zhao

Research Collection School Of Computing and Information Systems

Collaborative Filtering (CF) is one of the most successful learning techniques in building real-world recommender systems. Traditional CF algorithms are often based on batch machine learning methods which suffer from several critical drawbacks, e.g., extremely expensive model retraining cost whenever new samples arrive, unable to capture the latest change of user preferences over time, and high cost and slow reaction to new users or products extension. Such limitations make batch learning based CF methods unsuitable for real-world online applications where data often arrives sequentially and user preferences may change dynamically and rapidly. To address these limitations, we investigate online collaborative …


Social Informatics, Adam Jatowt, Ee-Peng Lim, Ying Ding, Asako Miura, Taro Tezuka, Gael Dias, Katsumi Tanaka, Andrew J. Flanagin, Bing Tian Dai Nov 2013

Social Informatics, Adam Jatowt, Ee-Peng Lim, Ying Ding, Asako Miura, Taro Tezuka, Gael Dias, Katsumi Tanaka, Andrew J. Flanagin, Bing Tian Dai

Research Collection School Of Computing and Information Systems

This book constitutes the proceedings of the 5th International Conference on Social Informatics, SocInfo 2013, held in Kyoto, Japan, in November 2013. The 23 full papers, 15 short papers, and three poster papers included in this volume were carefully reviewed and selected from 103 submissions. The papers present original research work on studying the interplay between socially-centric platforms and social phenomena.


Using Micro-Reviews To Select An Efficient Set Of Reviews, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas Nov 2013

Using Micro-Reviews To Select An Efficient Set Of Reviews, Thanh-Son Nguyen, Hady W. Lauw, Panayiotis Tsaparas

Research Collection School Of Computing and Information Systems

Online reviews are an invaluable resource for web users trying to make decisions regarding products or services. However, the abundance of review content, as well as the unstructured, lengthy, and verbose nature of reviews make it hard for users to locate the appropriate reviews, and distill the useful information. With the recent growth of social networking and micro-blogging services, we observe the emergence of a new type of online review content, consisting of bite-sized, 140 character-long reviews often posted reactively on the spot via mobile devices. These micro-reviews are short, concise, and focused, nicely complementing the lengthy, elaborate, and verbose …


Efficient Index-Based Approaches For Skyline Queries In Location-Based Applications, Ken C. K. Lee, Baihua Zheng, Cindy Chen, Chi-Yin Chow Nov 2013

Efficient Index-Based Approaches For Skyline Queries In Location-Based Applications, Ken C. K. Lee, Baihua Zheng, Cindy Chen, Chi-Yin Chow

Research Collection School Of Computing and Information Systems

Enriching many location-based applications, various new skyline queries are proposed and formulated based on the notion of locational dominance, which extends conventional one by taking objects' nearness to query positions into account additional to objects' nonspatial attributes. To answer a representative class of skyline queries for location-based applications efficiently, this paper presents two index-based approaches, namely, augmented R-tree and dominance diagram. Augmented R-tree extends R-tree by including aggregated nonspatial attributes in index nodes to enable dominance checks during index traversal. Dominance diagram is a solution-based approach, by which each object is associated with a precomputed nondominance scope wherein query points …


Classification In P2p Networks With Cascade Support Vendor Machines, Hock Hee Ang, Vivekanand Gopalkrishnan, Steven C. H. Hoi, Wee-Keong Ng Nov 2013

Classification In P2p Networks With Cascade Support Vendor Machines, Hock Hee Ang, Vivekanand Gopalkrishnan, Steven C. H. Hoi, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

Classification in Peer-to-Peer (P2P) networks is important to many real applications, such as distributed intrusion detection, distributed recommendation systems, and distributed antispam detection. However, it is very challenging to perform classification in P2P networks due to many practical issues, such as scalability, peer dynamism, and asynchronism. This article investigates the practical techniques of constructing Support Vector Machine (SVM) classifiers in the P2P networks. In particular, we demonstrate how to efficiently cascade SVM in a P2P network with the use of reduced SVM. In addition, we propose to fuse the concept of cascade SVM with bootstrap aggregation to effectively balance the …


Multimedia Modeling, Chong-Wah Ngo, Klaus Schoeffmann, Yiannis Andreopoulos, Christian Breiteneder Nov 2013

Multimedia Modeling, Chong-Wah Ngo, Klaus Schoeffmann, Yiannis Andreopoulos, Christian Breiteneder

Research Collection School Of Computing and Information Systems

Multimedia modeling aims to study computational models for addressing real-world multimedia problems from various perspectives, including information fusion, perceptual understanding, performance evaluation and social media. The topic becomes increasingly important with the massive amount of data available over the Internet, representing different pieces of information in heterogeneous forms that need to be consolidated before being used for multimedia problems. On the other hand, the advancement in technologies such as mobile and sensing devices drive the needs for revisiting the existing models for not only dealing with audio-visual cues but also incorporating various sensory modalities that have potential in providing cheaper …


Why Do I Retweet It? An Information Propagation Model For Microblogs, Fabio Pezzoni, Jisun An, Andrea Passarella, Jon Crowcroft, Marco Conti Nov 2013

Why Do I Retweet It? An Information Propagation Model For Microblogs, Fabio Pezzoni, Jisun An, Andrea Passarella, Jon Crowcroft, Marco Conti

Research Collection School Of Computing and Information Systems

Microblogging platforms are Web 2.0 services that represent a suitable environment for studying how information is propagated in social networks and how users can become influential. In this work we analyse the impact of the network features and of the users' behaviour on the information diffusion. Our analysis highlights a strong relation between the level of visibility of a message in the flow of information seen by a user and the probability that the user further disseminates the message. In addition, we also highlight the existence of other latent factors that impact on the dissemination probability, correlated with the properties …


City Notifications As A Data Source For Traffic Management, Pramod Anantharam, Biplav Srivastava Oct 2013

City Notifications As A Data Source For Traffic Management, Pramod Anantharam, Biplav Srivastava

Kno.e.sis Publications

A common problem for cities of developing countries like India in managing traffic is the lack of basic automated instrumentation to track road conditions or vehicle locations. Still, to help their citizens make informed travel decisions based on changing city dynamics; many cities have an authorized, city-initiated, notification service in place to alert subscribing commuters about road conditions. Here, alternative means may be used to create informal textual notifications e.g., inputs from field personnel, citizen updates, and pre-authorized events from city calendar. In this paper, we show that collections of such notifications, when processed with information extraction techniques, can turn …


Suddenly...I'M Consulting On Data Management Plans! Data Management Plan Consultant Checklist, Kiyomi D. Deards Oct 2013

Suddenly...I'M Consulting On Data Management Plans! Data Management Plan Consultant Checklist, Kiyomi D. Deards

University of Nebraska-Lincoln Libraries: Presentations

This webinar will outline the most important questions to ask, and the best resources available, for those who "suddenly" will be consulting on data management plans.


Improving Reuse In Software Development For The Life Sciences, Nicholas Vincent Iannotti Oct 2013

Improving Reuse In Software Development For The Life Sciences, Nicholas Vincent Iannotti

Open Access Dissertations

The last several years have seen unprecedented advancements in the application of technology to the life sciences, particularly in the area of data generation. Novel scientific insights are now often driven primarily by software development supporting new multidisciplinary and increasingly multifaceted data analysis. However, despite the availability of tools such as best practice frameworks, the current rate of software development is not able to keep up with the needs of scientists. This bottleneck in software development is largely due to code reuse generally not being applied in practice.

This dissertation presents Legwork, a class library of reuse-optimized design pattern implementations …


Search Tool That Utilizes Scientific Metadata Matched Against User-Entered Parameters, Veronika Margaret Megler, David Maier Oct 2013

Search Tool That Utilizes Scientific Metadata Matched Against User-Entered Parameters, Veronika Margaret Megler, David Maier

Computer Science Faculty Publications and Presentations

A method for providing proximate dataset recommendations can begin with the creation of metadata records corresponding to datasets that represent scientific data by a scientific dataset search tool. The metadata records can conform to a standardized structural definition, and may be hierarchical. Values for the data elements of the metadata records can be contained within the datasets. Metadata records with a value that is proximate to a user-entered search parameter can be identified. A proximity score can be calculated for each identified metadata record. The proximity score can express a relevance of the corresponding dataset to the user-entered search parameters. …


Designing Mobile Educational Games On Voter‟S Education: A Tale Of Three Engines, Ma. Regina Justina E. Estuar, Nadia Rowena C. Leetian, Michael B. Syson Oct 2013

Designing Mobile Educational Games On Voter‟S Education: A Tale Of Three Engines, Ma. Regina Justina E. Estuar, Nadia Rowena C. Leetian, Michael B. Syson

Department of Information Systems & Computer Science Faculty Publications

The rapid growth of mobile learning is influenced by the ability to access learning content anytime and anywhere. The on demand capability is available because mobile devices allow for convergence of internet and communications technologies. At the same time, the availability of engines makes development of mobile applications faster and seamless. However, not all mobile development engines are alike. This paper discusses on the development of mobile learning applications using mobile development engines in teaching Filipinos on responsible voting. Specifically, this paper discusses how AndEngine, Ren’Py, and homegrown Usbong were used to develop a mobile board game and a mobile …


Modeling Interaction Features For Debate Side Clustering, Minghui Qiu, Liu Yang, Jing Jiang Oct 2013

Modeling Interaction Features For Debate Side Clustering, Minghui Qiu, Liu Yang, Jing Jiang

Research Collection School Of Computing and Information Systems

Online discussion forums are popular social media platforms for users to express their opinions and discuss controversial issues with each other. To automatically identify the sides/stances of posts or users from textual content in forums is an important task to help mine online opinions. To tackle the task, it is important to exploit user posts that implicitly contain support and dispute (interaction) information. The challenge we face is how to mine such interaction information from the content of posts and how to use them to help identify stances. This paper proposes a two-stage solution based on latent variable models: an …


On Effects Of Visual Query Complexity, Jialie Shen, Cheng Zhiyong Oct 2013

On Effects Of Visual Query Complexity, Jialie Shen, Cheng Zhiyong

Research Collection School Of Computing and Information Systems

As an effective technique to manage large scale image collections, content-based image retrieval (CBIR) has been received great attentions and became a very active research domain in recent years. While assessing system performance is one of the key factors for the related technological advancement, relatively little attention has been paid to model and analyze test queries. This paper documents a study on the problem of determining visual query complexity as a measure for predicting image retrieval performance. We propose a quantitative metric for measuring complexity of image queries for content-based image search engine. A set of experiments are carried out …


Consistent Stereo Image Editing, Tao Yan, Shengfeng He, Rynson W.H. Lau, Yun Xu Oct 2013

Consistent Stereo Image Editing, Tao Yan, Shengfeng He, Rynson W.H. Lau, Yun Xu

Research Collection School Of Computing and Information Systems

Stereo images and videos are very popular in recent years, and techniques for processing this media are attracting a lot of attention. In this paper, we extend the shift-map method for stereo image editing. Our method simultaneously processes the left and right images on pixel level using a global optimization algorithm. It enforces photo consistence between the two images and preserves 3D scene structures. It also addresses the occlusion and disocclusion problem, which may enable many stereo image editing functions, such as depth mapping, object depth adjustment and non-homogeneous image resizing. Our experiments show that the proposed method produces high …


Image Search By Graph-Based Label Propagation With Image Representation From Dnn, Yingwei Pan, Yao Ting, Kuiyuan Yang, Houqiang Li, Chong-Wah Ngo, Jingdong Wang, Tao Mei Oct 2013

Image Search By Graph-Based Label Propagation With Image Representation From Dnn, Yingwei Pan, Yao Ting, Kuiyuan Yang, Houqiang Li, Chong-Wah Ngo, Jingdong Wang, Tao Mei

Research Collection School Of Computing and Information Systems

Our objective is to estimate the relevance of an image to a query for image search purposes. We address two limitations of the existing image search engines in this paper. First, there is no straightforward way of bridging the gap between semantic textual queries as well as users’ search intents and image visual content. Image search engines therefore primarily rely on static and textual features. Visual features are mainly used to identify potentially useful recurrent patterns or relevant training examples for complementing search by image reranking. Second, image rankers are trained on query-image pairs labeled by human experts, making the …


Skyhunter: A Multi-Surface Environment For Supporting Oil And Gas Exploration, Teddy Seyed, Mario Costa Sousa, Frank Maurer, Anthony Tang Oct 2013

Skyhunter: A Multi-Surface Environment For Supporting Oil And Gas Exploration, Teddy Seyed, Mario Costa Sousa, Frank Maurer, Anthony Tang

Research Collection School Of Computing and Information Systems

The process of oil and gas exploration and its result, the decision to drill for oil in a specific location, relies on a number of distinct but related domains. These domains require effective collaboration to come to a decision that is both cost effective and maintains the integrity of the environment. As we show in this paper, many of the existing technologies and practices that support the oil and gas exploration process overlook fundamental user issues such as collaboration, interaction and visualization. The work presented in this paper is based upon a design process that involved expert users from an …


Online Multimodal Distance Metric Learning With Application To Image Retrieval, Pengcheng Wu, Steven C. H. Hoi, Hao Xia, Peilin Zhao, Dayong Wang, Chunyan Miao Oct 2013

Online Multimodal Distance Metric Learning With Application To Image Retrieval, Pengcheng Wu, Steven C. H. Hoi, Hao Xia, Peilin Zhao, Dayong Wang, Chunyan Miao

Research Collection School Of Computing and Information Systems

Recent years have witnessed extensive studies on distance metric learning (DML) for improving similarity search in multimedia information retrieval tasks. Despite their successes, most existing DML methods suffer from two critical limitations: (i) they typically attempt to learn a linear distance function on the input feature space, in which the assumption of linearity limits their capacity of measuring the similarity on complex patterns in real-world applications; (ii) they are often designed for learning distance metrics on uni-modal data, which may not effectively handle the similarity measures for multimedia objects with multimodal representations. To address these limitations, in this paper, we …


Predictive Handling Of Asynchronous Concept Drifts In Distributed Environments, Hock Hee Ang, Vivek Gopalkrishnan, Indre Zliobaite, Mykola Pechenizkiy, Steven C. H. Hoi Oct 2013

Predictive Handling Of Asynchronous Concept Drifts In Distributed Environments, Hock Hee Ang, Vivek Gopalkrishnan, Indre Zliobaite, Mykola Pechenizkiy, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

In a distributed computing environment, peers collaboratively learn to classify concepts of interest from each other. When external changes happen and their concepts drift, the peers should adapt to avoid increase in misclassification errors. The problem of adaptation becomes more difficult when the changes are asynchronous, i.e., when peers experience drifts at different times. We address this problem by developing an ensemble approach, PINE, that combines reactive adaptation via drift detection, and proactive handling of upcoming changes via early warning and adaptation across the peers. With empirical study on simulated and real-world data sets, we show that PINE handles asynchronous …


Online Multi-Task Collaborative Filtering For On-The-Fly Recommender Systems, Jialei Wang, Steven C. H. Hoi, Peilin Zhao, Zhi-Yong Liu Oct 2013

Online Multi-Task Collaborative Filtering For On-The-Fly Recommender Systems, Jialei Wang, Steven C. H. Hoi, Peilin Zhao, Zhi-Yong Liu

Research Collection School Of Computing and Information Systems

Traditional batch model-based Collaborative Filtering (CF) approaches typically assume a collection of users' rating data is given a priori for training the model. They suffer from a common yet critical drawback, i.e., the model has to be re-trained completely from scratch whenever new training data arrives, which is clearly non-scalable for large real recommender systems where users' rating data often arrives sequentially and frequently. In this paper, we investigate a novel efficient and scalable online collaborative filtering technique for on-the-fly recommender systems, which is able to effectively online update the recommendation model from a sequence of rating observations. Specifically, we …