Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2017

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 211 - 240 of 373

Full-Text Articles in Databases and Information Systems

Persona Generation From Aggregated Social Media Data, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Moeed Ahmad, Lene Nielsen, Bernard J. Jansen May 2017

Persona Generation From Aggregated Social Media Data, Soon-Gyo Jung, Jisun An, Haewoon Kwak, Moeed Ahmad, Lene Nielsen, Bernard J. Jansen

Research Collection School Of Computing and Information Systems

We develop a methodology for persona generation using real time social media data for the distribution of products via online platforms. From a large social media account containing more than 30 million interactions from users from 181 countries engaging with more than 4,200 digital products produced by a global media corporation, we demonstrate that our methodology can first identify both distinct and impactful user segments and then create persona descriptions by automatically adding pertinent features, such as names, photos, and personal attributes. We validate our approach by implementing the methodology into an actual working system that leverages large scale online …


The Creation Of A Building Map Application For A University Setting, William T. Whitesell Apr 2017

The Creation Of A Building Map Application For A University Setting, William T. Whitesell

Senior Honors Theses

The use of navigational technology in mobile and web devices has sharply increased in recent years. With the capability to create interactive maps now available, navigating in real time between locations has become possible. This is especially essential in areas and organizations experiencing rapid expansion like Liberty University (LU). Therefore, the author proposes a project to create an interactive map application (IMA) for LU’s academic buildings that is scalable and usable through both the university’s website and with a mobile application. There are several considerations that must be taken into account when creating the LU map application, such as development …


Mapping Community Space And Place In Mto Wa Mbu, Tanzania Through Surveys And Gis, Jessica Craigg Apr 2017

Mapping Community Space And Place In Mto Wa Mbu, Tanzania Through Surveys And Gis, Jessica Craigg

Georgia College Student Research Events

Cities throughout the African continent have been developing at an unprecedented pace, many of them due to the influence of the tourism industry. This is particularly true in Tanzania, a country famous for its national parks and their draw to tourists who help provide money for development. However, the only way to get the whole story on how to spend this money is through the experiences and needs of the people themselves. This study focuses on a small town in northeastern Tanzania, Mto wa Mbu, situated near Lake Manyara National Park, and its people’s perceptions of the park and community. …


Design And Implementation Of An Rfid-Based Customer Shopping Behavior Mining System, Zimu Zhou, Longfei Shangguan, Xiaolong Zheng, Lei Yang, Yunhao Liu Apr 2017

Design And Implementation Of An Rfid-Based Customer Shopping Behavior Mining System, Zimu Zhou, Longfei Shangguan, Xiaolong Zheng, Lei Yang, Yunhao Liu

Research Collection School Of Computing and Information Systems

Shopping behavior data is of great importance in understanding the effectiveness of marketing and merchandising campaigns. Online clothing stores are capable of capturing customer shopping behavior by analyzing the click streams and customer shopping carts. Retailers with physical clothing stores, however, still lack effective methods to comprehensively identify shopping behaviors. In this paper, we show that backscatter signals of passive RFID tags can be exploited to detect and record how customers browse stores, which garments they pay attention to, and which garments they usually pair up. The intuition is that the phase readings of tags attached to items will demonstrate …


Stream Data Quality Assessment Based On Distributed Computing Platforms, Wei Dai Apr 2017

Stream Data Quality Assessment Based On Distributed Computing Platforms, Wei Dai

Theses and Dissertations

In this era of big data, data quality will be increasingly important because people need high quality data to make decisions, analyze patterns, and discover knowledge. So, measuring data quality is a vital mission. In this thesis, Chapter 1 is the introduction, Chapter 2 is a literature review, Chapter 3 illustrates how to discover potentially important data based on a reference algorithm, a frequency algorithm, and an entropy algorithm, in Chapter 4, the author offers a concise five-layer data quality framework to measure stream data quality scorecards, in Chapter 5, the author shows how to visualize data quality scorecards through …


Blocking Strategies For Performing Entity Resolution In A Distributed Computing Environment, Pei Wang Apr 2017

Blocking Strategies For Performing Entity Resolution In A Distributed Computing Environment, Pei Wang

Theses and Dissertations

Entity resolution (ER) is an O(n2) problem where n is the number of records to be processed. The pair-wise nature of ER makes it impractical to perform on large datasets without the use of a technique called blocking. In blocking the records are separated into groups (called blocks) in such a way the records most likely to match are within the same block. The ER system only compares pairs of records within the same block, thus reducing the total number of pairs to match. Traditionally, blocking algorithms build inverted indices in memory to quickly locate potential matches. With the advent …


Viewability Prediction For Display Advertising, Chong Wang Apr 2017

Viewability Prediction For Display Advertising, Chong Wang

Dissertations

As a massive industry, display advertising delivers advertisers’ marketing messages to attract customers through graphic banners on webpages. Display advertising is also the most essential revenue source of online publishers. Currently, advertisers are charged by user response or ad serving. However, recent studies show that users barely click or convert display ads. Moreover, about half of the ads are actually never seen by users. In this case, advertisers cannot enhance their brand awareness and increase return on investment. Publishers also lose much revenue. Therefore, the ad pricing standards are shifting to a new model: ad impressions are paid if they …


Factored Similarity Models With Social Trust For Top-N Item Recommendation, Guibing Guo, Jie Zhang, Feida Zhu, Xingwei Wang Apr 2017

Factored Similarity Models With Social Trust For Top-N Item Recommendation, Guibing Guo, Jie Zhang, Feida Zhu, Xingwei Wang

Research Collection School of Computing and Information Systems

Trust-aware recommender systems have attracted much attention recently due to the prevalence of social networks. However, most existing trust-based approaches are designed for the recommendation task of rating prediction. Only few trust-aware methods have attempted to recommend users an ordered list of interesting items, i.e., item recommendation. In this article, we propose three factored similarity models with the incorporation of social trust for item recommendation based on implicit user feedback. Specifically, we introduce a matrix factorization technique to recover user preferences between rated items and unrated ones in the light of both user-user and item-item similarities. In addition, we claim …


A Proposed Frequency-Based Feature Selection Method For Cancer Classification, Yi Pan Apr 2017

A Proposed Frequency-Based Feature Selection Method For Cancer Classification, Yi Pan

Masters Theses & Specialist Projects

Feature selection method is becoming an essential procedure in data preprocessing step. The feature selection problem can affect the efficiency and accuracy of classification models. Therefore, it also relates to whether a classification model can have a reliable performance. In this study, we compared an original feature selection method and a proposed frequency-based feature selection method with four classification models and three filter-based ranking techniques using a cancer dataset. The proposed method was implemented in WEKA which is an open source software. The performance is evaluated by two evaluation methods: Recall and Receiver Operating Characteristic (ROC). Finally, we found the …


What Are People Tweeting About Zika? An Exploratory Study Concerning Its Symptoms, Treatment, Transmission, And Prevention, Michele Miller, Tanvi Banerjee, Roopteja Muppalla, William L. Romine, Amit Sheth Apr 2017

What Are People Tweeting About Zika? An Exploratory Study Concerning Its Symptoms, Treatment, Transmission, And Prevention, Michele Miller, Tanvi Banerjee, Roopteja Muppalla, William L. Romine, Amit Sheth

Kno.e.sis Publications

Background: In order to harness what people are tweeting about Zika, there needs to be a computational framework that leverages machine learning techniques to recognize relevant Zika tweets and, further, categorize these into disease-specific categories to address specific societal concerns related to the prevention, transmission, symptoms, and treatment of Zika virus.

Objective: The purpose of this study was to determine the relevancy of the tweets and what people were tweeting about the 4 disease characteristics of Zika: symptoms, transmission, prevention, and treatment.

Methods: A combination of natural language processing and machine learning techniques was used to determine what people were …


Eassistant: Cognitive Assistance For Identification And Auto-Triage Of Actionable Conversations, Hamid R. Motahari Nezhad, Kalpa Gunaratna, Juan Cappi Apr 2017

Eassistant: Cognitive Assistance For Identification And Auto-Triage Of Actionable Conversations, Hamid R. Motahari Nezhad, Kalpa Gunaratna, Juan Cappi

Kno.e.sis Publications

The browser and screen have been the main user interfaces of the Web and mobile apps. The notification mechanism is an evolution in the user interaction paradigm by keeping users updated without checking applications. Conversational agents are posed to be the next revolution in user interaction paradigms. However, without intelligence on the triage of content served by the interaction and content differentiation in applications, interaction paradigms may still place the burden of information overload on users. In this paper, we focus on the problem of intelligent identification of actionable information in the content served by applications, and in particular in …


On The Effectiveness Of Virtualization Based Memory Isolation On Multicore Platforms, Siqi Zhao, Xuhua Ding Apr 2017

On The Effectiveness Of Virtualization Based Memory Isolation On Multicore Platforms, Siqi Zhao, Xuhua Ding

Research Collection School Of Computing and Information Systems

Virtualization based memory isolation has beenwidely used as a security primitive in many security systems.This paper firstly provides an in-depth analysis of itseffectiveness in the multicore setting; a first in the literature.Our study reveals that memory isolation by itself is inadequatefor security. Due to the fundamental design choices inhardware, it faces several challenging issues including pagetable maintenance, address mapping validation and threadidentification. As demonstrated by our attacks implementedon XMHF and BitVisor, these issues undermine the security ofmemory isolation. Next, we propose a new isolation approachthat is immune to the aforementioned problems. In our design,the hypervisor constructs a fully isolated micro …


Discovering Anomalous Events From Urban Informatics Data, Kasthuri Jayarajah, Vigneshwaran Subbaraju, Dulanga Kaveesha Weerakoon Mudiyanselage, Archan Misra, La Thanh Tam, Noel Athaide Apr 2017

Discovering Anomalous Events From Urban Informatics Data, Kasthuri Jayarajah, Vigneshwaran Subbaraju, Dulanga Kaveesha Weerakoon Mudiyanselage, Archan Misra, La Thanh Tam, Noel Athaide

Research Collection School Of Computing and Information Systems

Singapore's "smart city" agenda is driving the government to provide public access to a broader variety of urban informatics sources, such as images from traffic cameras and information about buses servicing different bus stops. Such informatics data serves as probes of evolving conditions at different spatiotemporal scales. This paper explores how such multi-modal informatics data can be used to establish the normal operating conditions at different city locations, and then apply appropriate outlier-based analysis techniques to identify anomalous events at these selected locations. We will introduce the overall architecture of sociophysical analytics, where such infrastructural data sources can be combined …


Finding Causality And Responsibility For Probabilistic Reverse Skyline Query Non-Answers [Extended Abstract], Yunjun Gao, Qing Liu, Gang Chen, Linlin Zhou, Baihua Zheng Apr 2017

Finding Causality And Responsibility For Probabilistic Reverse Skyline Query Non-Answers [Extended Abstract], Yunjun Gao, Qing Liu, Gang Chen, Linlin Zhou, Baihua Zheng

Research Collection School Of Computing and Information Systems

This paper explores the causality and responsibility problem (CRP) for the non-answers to probabilistic reverse skyline queries (PRSQ). Towards this, we propose an efficient algorithm called CP to compute the causality and responsibility for the non-answers to PRSQ. CP first finds candidate causes, and then, it performs verification to obtain actual causes with their responsibilities, during which several strategies are used to boost efficiency. Extensive experiments using both real and synthetic data sets demonstrate the effectiveness and efficiency of the presented algorithms.


Neural Collaborative Filtering, Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, Tat-Seng Chua Apr 2017

Neural Collaborative Filtering, Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

In recent years, deep neural networks have yielded immense success on speech recognition, computer vision and natural language processing. However, the exploration of deep neural networks on recommender systems has received relatively less scrutiny. In this work, we strive to develop techniques based on neural networks to tackle the key problem in recommendation --- collaborative filtering --- on the basis of implicit feedback.Although some recent work has employed deep learning for recommendation, they primarily used it to model auxiliary information, such as textual descriptions of items and acoustic features of musics. When it comes to model the key factor in …


Online Growing Neural Gas For Anomaly Detection In Changing Surveillance Scenes, Qianru Sun, Hong Liu, Tatsuya Harada Apr 2017

Online Growing Neural Gas For Anomaly Detection In Changing Surveillance Scenes, Qianru Sun, Hong Liu, Tatsuya Harada

Research Collection School Of Computing and Information Systems

Anomaly detection is still a challenging task for video surveillance due to complex environments and unpredictable human behaviors. Most existing approaches train offline detectors using manually labeled data and predefined parameters, and are hard to model changing scenes. This paper introduces a neural network based model called online Growing Neural Gas (online GNG) to perform an unsupervised learning. Unlike a parameter-fixed GNG, our model updates learning parameters continuously, for which we propose several online neighbor-related strategies. Specific operations, namely neuron insertion, deletion, learning rate adaptation and stopping criteria selection, get upgraded to online modes. In the anomaly detection stage, the …


Learning Personalized Preference Of Strong And Weak Ties For Social Recommendation, Xin Wang, Steven C. H. Hoi, Martin Ester, Jiajun Bu, Chun Chen Apr 2017

Learning Personalized Preference Of Strong And Weak Ties For Social Recommendation, Xin Wang, Steven C. H. Hoi, Martin Ester, Jiajun Bu, Chun Chen

Research Collection School Of Computing and Information Systems

Recent years have seen a surge of research on social recommendation techniques for improving recommender systems due to the growing influence of social networks to our daily life. The intuition of social recommendation is that users tend to show affinities with items favored by their social ties due to social influence. Despite the extensive studies, no existing work has attempted to distinguish and learn the personalized preferences between strong and weak ties, two important terms widely used in social sciences, for each individual in social recommendation. In this paper, we first highlight the importance of different types of ties in …


On Analyzing User Topic-Specific Platform Preferences Across Multiple Social Media Sites, Roy Ka Wei Lee, Tuan Anh Hoang, Ee Peng Lim Apr 2017

On Analyzing User Topic-Specific Platform Preferences Across Multiple Social Media Sites, Roy Ka Wei Lee, Tuan Anh Hoang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Topic modeling has traditionally been studied for single text collections and applied to social media data represented in the form of text documents. With the emergence of many social media platforms, users find themselves using different social media for posting content and for social interaction. While many topics may be shared across social media platforms, users typically show preferences of certain social media platform(s) over others for certain topics. Such platform preferences may even be found at the individual level. To model social media topics as well as platform preferences of users, we propose a new topic model known as …


Now You See It, Now You Don't! A Study Of Content Modification Behavior In Facebook, Fuxiang Chen, Ee-Peng Lim Apr 2017

Now You See It, Now You Don't! A Study Of Content Modification Behavior In Facebook, Fuxiang Chen, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Social media, as a major platform to disseminate information, has changed the way users and communities contribute content. In this paper, we aim to study content modifications on public Facebook pages operated by news media, community groups, and bloggers. We also study the possible reasons behind them, and their effects on user interaction. We conducted a detailed study of Content Censorship (CC) and Content Edit (CE) in Facebook using a detailed longitudinal dataset consisting of 57 public Facebook pages over 3 weeks covering 145,955 posts and 9,379,200 comments. We detected many CC and CE activities between 28% and 56% of …


Machine Comprehension Using Match-Lstm And Answer Pointer, Shuohang Wang, Jing Jiang Apr 2017

Machine Comprehension Using Match-Lstm And Answer Pointer, Shuohang Wang, Jing Jiang

Research Collection School Of Computing and Information Systems

Machine comprehension of text is an important problem in natural language processing. A recently released dataset, the Stanford Question Answering Dataset (SQuAD), offers a large number of real questions and their answers created by humans through crowdsourcing. SQuAD provides a challenging testbed for evaluating machine comprehension algorithms, partly because compared with previous datasets, in SQuAD the answers do not come from a small set of candidate answers and they have variable lengths. We propose an end-to-end neural architecture for the task. The architecture is based on match-LSTM, a model we proposed previously for textual entailment, and Pointer Net, a sequence-to-sequence …


Modeling Topics And Behavior Of Microbloggers: An Integrated Approach, Tuan Anh Hoang, Ee-Peng Lim Apr 2017

Modeling Topics And Behavior Of Microbloggers: An Integrated Approach, Tuan Anh Hoang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

Microblogging encompasses both user-generated content and behavior. When modeling microblogging data, one has to consider personal and background topics, as well as how these topics generate the observed content and behavior. In this article, we propose the Generalized Behavior-Topic (GBT) model for simultaneously modeling background topics and users' topical interest in microblogging data. GBT considers multiple topical communities (or realms) with different background topical interests while learning the personal topics of each user and the user's dependence on realms to generate both content and behavior. This differentiates GBT from other previous works that consider either one realm only or content …


Assessing The Language Of Chat For Teamwork Dialogue, Antonette Shibani, Elizabeth Koh, Vivian Lai, Kyong Jin Shim Apr 2017

Assessing The Language Of Chat For Teamwork Dialogue, Antonette Shibani, Elizabeth Koh, Vivian Lai, Kyong Jin Shim

Research Collection School Of Computing and Information Systems

In technology enhanced language learning, many pedagogical activities involve students in online discussion such as synchronous chat, in order to help them practice their language skills. Besides developing the language competency of students, it is also crucial to nurture their teamwork competencies for today's global and complex environment. Language communication is an important glue of teamwork. In order to assess the language of chat for teamwork dimensions, several text mining methods are pos sible. However, difficulties arise such as pre-processing being a black box and classification approaches and algorithms being dependent on the context. To address these issues, the study …


Aspect Extraction From Product Reviews Using Category Hierarchy Information, Yifeng Yang, Chen Cen, Minghui Qiu, Forrest Sheng Bao Apr 2017

Aspect Extraction From Product Reviews Using Category Hierarchy Information, Yifeng Yang, Chen Cen, Minghui Qiu, Forrest Sheng Bao

Research Collection School Of Computing and Information Systems

Aspect extraction is a task to abstract the common properties of objects from corpora discussing them, such as reviews of products. Recent work on aspect extraction is leveraging the hierarchical relationship between products and their categories. However, such effort focuses on the aspects of child categories but ignores those from parent categories. Hence, we propose an LDA-based generative topic model inducing the two-layer categorical information (CAT-LDA), to balance the aspects of both a parent category and its child categories. Our hypothesis is that child categories inherit aspects from parent categories, controlled by the hierarchy between them. Experimental results on 5 …


Achievement And Friends: Key Factors Of Player Retention Vary Across Player Levels In Online Multiplayer Games, Korea Advanced Institute Of Science & Technology, Qatar Computing Research Institute, Haewoon Kwak Apr 2017

Achievement And Friends: Key Factors Of Player Retention Vary Across Player Levels In Online Multiplayer Games, Korea Advanced Institute Of Science & Technology, Qatar Computing Research Institute, Haewoon Kwak

Research Collection School Of Computing and Information Systems

Retaining players over an extended period of time is a long-standing challenge in game industry. Significant effort has been paid to understanding what motivates players enjoy games. While individuals may have varying reasons to play or abandon a game at different stages within the game, previous studies have looked at the retention problem from a snapshot view. This study, by analyzing in-game logs of 51,104 distinct individuals in an online multiplayer game, uniquely offers a multifaceted view of the retention problem over the players' virtual life phases. We find that key indicators of longevity change with the game level. Achievement …


A Compare-Aggregate Model For Matching Text Sequences, Shuohang Wang, Jing Jiang Apr 2017

A Compare-Aggregate Model For Matching Text Sequences, Shuohang Wang, Jing Jiang

Research Collection School Of Computing and Information Systems

Many NLP tasks including machine comprehension, answer selection and text entailment require the comparison between sequences. Matching the important units between sequences is a key to solve these problems. In this paper, we present a general "compare-aggregate" framework that performs word-level matching followed by aggregation using Convolutional Neural Networks. We particularly focus on the different comparison functions we can use to match two vectors. We use four different datasets to evaluate the model. We find that some simple comparison functions based on element-wise operations can work better than standard neural network and neural tensor network.


Collective Entity Linking In Tweets Over Space And Time, Wen Haw Chong, Ee-Peng Lim, William Cohen Apr 2017

Collective Entity Linking In Tweets Over Space And Time, Wen Haw Chong, Ee-Peng Lim, William Cohen

Research Collection School Of Computing and Information Systems

We propose collective entity linking over tweets that are close in space and time. This exploits the fact that events or geographical points of interest often result in related entities being mentioned in spatio-temporal proximity. Our approach directly applies to geocoded tweets. Where geocoded tweets are overly sparse among all tweets, we use a relaxed version of spatial proximity which utilizes both geocoded and non-geocoded tweets linked by common mentions. Entity linking is affected by noisy mentions extracted and incomplete knowledge bases. Moreover, to perform evaluation on the entity linking results, much manual annotation of mentions is often required. To …


Comparative Relation Generative Model, Maksim Tkachenko, Hady W. Lauw Apr 2017

Comparative Relation Generative Model, Maksim Tkachenko, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Online reviews are important decision aids to consumers. Other than helping users to evaluate individual products, reviews also support comparison shopping by comparing two (or more) products based on a specific aspect. However, making a comparison across two different reviews, written by different authors, is not always equitable due to the different standards and preferences of authors. Therefore, we focus on comparative sentences, whereby two products are compared directly by a review author within a sentence. We study the problem of comparative relation mining. Given a set of comparative sentences, each relating a pair of entities, our objective is three-fold: …


Understanding The Information-Based Transformation Of Strategy And Society, Eric K. Clemons, Rajiv M. Dewan, Robert J. Kauffman, Thomas A. Weber Apr 2017

Understanding The Information-Based Transformation Of Strategy And Society, Eric K. Clemons, Rajiv M. Dewan, Robert J. Kauffman, Thomas A. Weber

Research Collection School Of Computing and Information Systems

The world economy is undergoing dramatic changes, largely driven by the new availability of fine-grained information. Innovative ways of using data—large and small—have also prompted a rethinking of the boundaries for the combination and use of knowledge. The strategic design of information flows in the economy has the upside of higher economic rents and competitive advantage, as well as the downsides of wealth inequality and abuse of power. This has brought a wide range of regulatory challenges. To understand the nature of these sweeping changes, it is important to examine the new ways information is used, and how information flows …


I Would Not Plant Apple Trees If The World Will Be Wiped: Analyzing Hundreds Of Millions Of Behavioral Records Of Players During An Mmorpg Beta Test, Qatar Computing Research Institute, The State University Of New York University At Buffalo, Haewoon Kwak, Korea University Apr 2017

I Would Not Plant Apple Trees If The World Will Be Wiped: Analyzing Hundreds Of Millions Of Behavioral Records Of Players During An Mmorpg Beta Test, Qatar Computing Research Institute, The State University Of New York University At Buffalo, Haewoon Kwak, Korea University

Research Collection School Of Computing and Information Systems

In this work, we use player behavior during the closed beta test of the MMORPG ArcheAge as a proxy for an extreme situation: at the end of the closed beta test, all user data is deleted, and thus, the outcome (or penalty) of players' in-game behaviors in the last few days loses its meaning. We analyzed 270 million records of player behavior in the 4th closed beta test of ArcheAge. Our findings show that there are no apparent pandemic behavior changes, but some outlierswere more likely to exhibit anti-social behavior (e.g., player killing). We also found that contrary to the …


Harnessing Legal Complexity, Daniel Katz, J. Ruhl, M Bommarito Mar 2017

Harnessing Legal Complexity, Daniel Katz, J. Ruhl, M Bommarito

All Faculty Scholarship

No abstract provided.