Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2017

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 181 - 210 of 373

Full-Text Articles in Databases and Information Systems

Named Entity Recognition And Classification For Natural Language Inputs At Scale, Shreeraj Dabholkar May 2017

Named Entity Recognition And Classification For Natural Language Inputs At Scale, Shreeraj Dabholkar

Master's Projects

Natural language processing (NLP) is a technique by which computers can analyze, understand, and derive meaning from human language. Phrases in a body of natural text that represent names, such as those of persons, organizations or locations are referred to as named entities. Identifying and categorizing these named entities is still a challenging task, research on which, has been carried out for many years. In this project, we build a supervised learning based classifier which can perform named entity recognition and classification (NERC) on input text and implement it as part of a chatbot application. The implementation is then scaled …


Lightweight Data Aggregation Scheme Against Internal Attackers In Smart Grid Using Elliptic Curve Cryptography, Debiao He, Sherali Zeadally, Huaqun Wang, Qin Liu May 2017

Lightweight Data Aggregation Scheme Against Internal Attackers In Smart Grid Using Elliptic Curve Cryptography, Debiao He, Sherali Zeadally, Huaqun Wang, Qin Liu

Information Science Faculty Publications

Recent advances of Internet and microelectronics technologies have led to the concept of smart grid which has been a widespread concern for industry, governments, and academia. The openness of communications in the smart grid environment makes the system vulnerable to different types of attacks. The implementation of secure communication and the protection of consumers’ privacy have become challenging issues. The data aggregation scheme is an important technique for preserving consumers’ privacy because it can stop the leakage of a specific consumer’s data. To satisfy the security requirements of practical applications, a lot of data aggregation schemes were presented over the …


Software Development For Home Energy Audits: Reducing Energy Consumption In Harrisonburg Through Technology, Brantley E. Gilbert May 2017

Software Development For Home Energy Audits: Reducing Energy Consumption In Harrisonburg Through Technology, Brantley E. Gilbert

Senior Honors Projects, 2010-2019

Fossil fuels play a vital role in our daily lives. Oil, natural gas, and coal powers our cars, heats our homes and water, and are used by power companies to generate the massive amounts of electricity used every day by the United States. However, this reliance on a finite source of energy is not sustainable. Fossil fuels such as these are non-renewable resources whose production will eventually be unable to keep up with the rate of consumption. Furthermore, the extraction of the stored energy in these fuels through combustion releases harmful substances into the environment, including toxins and greenhouse gases …


Aspect Discovery From Product Reviews, Ying Ding May 2017

Aspect Discovery From Product Reviews, Ying Ding

Dissertations and Theses Collection

With the rapid development of online shopping sites and social media, product reviews are accumulating. These reviews contain information that is valuable to both businesses and customers. To businesses, companies can easily get a large number of feedback of their products, which is difficult to achieve by doing customer survey in the traditional way. To customers, they can know the products they are interested in better by reading reviews, which may be uneasy without online reviews. However, the accumulation has caused consuming all reviews impossible. It is necessary to develop automated techniques to efficiently process them. One of the most …


Mining Helpdesk Databases For Professional Development Topic Discovery, Joel T. Lowsky May 2017

Mining Helpdesk Databases For Professional Development Topic Discovery, Joel T. Lowsky

All Theses And Dissertations

This single-site, instrumental case study created and tested a methodological road map by which academic institutions can use text data mining techniques to derive technology skillset weaknesses and professional development topics from the site’s technical support helpdesk database. The methods employed were described in detail and applied to the helpdesk database of an independent, co-educational boarding high school in the northeastern United States. Standard text data mining procedures, including the formation of a wordlist (frequently occurring terms), and the creation and application of clustering (automated data grouping) and classification (automated data labeling) models generated meaningful and revealing themes from the …


Exploiting Semantic Distance In Linked Open Data For Recommendation, Sultan Dawood Alfarhood May 2017

Exploiting Semantic Distance In Linked Open Data For Recommendation, Sultan Dawood Alfarhood

Graduate Theses and Dissertations

The use of Linked Open Data (LOD) has been explored in recommender systems in different ways, primarily through its graphical representation. The graph structure of LOD is utilized to measure inter-resource relatedness via their semantic distance in the graph. The intuition behind this approach is that the more connected resources are to each other, the more related they are. One drawback of this approach is that it treats all inter-resource connections identically rather than prioritizing links that may be more important in semantic relatedness calculations. Another drawback of current approaches is that they only consider resources that are connected directly …


Country 2.0: Upgrading Cities With Smart Technologies, Steven M. Miller May 2017

Country 2.0: Upgrading Cities With Smart Technologies, Steven M. Miller

Asian Management Insights

Advancements in technology are being used to transform our cities into smart cities, but the process is not without its risks.


Real-Time Prediction Of Length Of Stay Using Passive Wi-Fi Sensing, Truc Viet Le, Baoyang Song, Laura Wynter May 2017

Real-Time Prediction Of Length Of Stay Using Passive Wi-Fi Sensing, Truc Viet Le, Baoyang Song, Laura Wynter

Research Collection School Of Computing and Information Systems

The proliferation of wireless technologies in today's everyday life is one of the key drivers of the Internet of Things (IoT). In addition to being an enabler of connectivity, the vast penetration of wireless devices today gives rise to a secondary functionality as a means of tracking and localization of the devices themselves. Indeed, in order to discover and automatically connect to known Wi-Fi networks, mobile devices have to scan and broadcast the so-called probe requests on all available channels, which can be captured and analyzed in a non-intrusive manner. Thus, one of the key applications of this feature is …


The Economics Of The Right To Be Forgotten, Byung-Cheol Kim, Jin Yeub Kim May 2017

The Economics Of The Right To Be Forgotten, Byung-Cheol Kim, Jin Yeub Kim

Department of Economics: Faculty Publications

Scholars and practitioners debate whether to expand the scope of the right to be forgotten—the right to have certain links removed from search results—to encompass global search results. The debate centers on the assumption that the expansion will increase the incidence of link removal, which reinforces privacy while hampering free speech. We develop a game-theoretic model to show that the expansion of the right to be forgotten can reduce the incidence of link removal. We also show that the expansion does not necessarily enhance the welfare of individuals who request removal and that it can either improve or reduce societal …


Robust Object Tracking Via Locality Sensitive Histograms, Shengfeng He, Rynson W.H Lau, Qingxiong Yang, Jiang Wang, Ming-Hsuan Yang May 2017

Robust Object Tracking Via Locality Sensitive Histograms, Shengfeng He, Rynson W.H Lau, Qingxiong Yang, Jiang Wang, Ming-Hsuan Yang

Research Collection School Of Computing and Information Systems

This paper presents a novel locality sensitive histogram (LSH) algorithm for visual tracking. Unlike the conventional image histogram that counts the frequency of occurrence of each intensity value by adding ones to the corresponding bin, an LSH is computed at each pixel location, and a floating-point value is added to the corresponding bin for each occurrence of an intensity value. The floating-point value exponentially reduces with respect to the distance to the pixel location where the histogram is computed. An efficient algorithm is proposed that enables the LSHs to be computed in time linear in the image size and the …


Joint Optimization Of Resource Provisioning In Cloud Computing, Jonathan David Chase, Dusit Niyato May 2017

Joint Optimization Of Resource Provisioning In Cloud Computing, Jonathan David Chase, Dusit Niyato

Research Collection School Of Computing and Information Systems

Cloud computing exploits virtualization to provision resources efficiently. Increasingly, Virtual Machines (VMs) have high bandwidth requirements; however, previous research does not fully address the challenge of both VM and bandwidth provisioning. To efficiently provision resources, a joint approach that combines VMs and bandwidth allocation is required. Furthermore, in practice, demand is uncertain. Service providers allow the reservation of resources. However, due to the dangers of over-and under-provisioning, we employ stochastic programming to account for this risk. To improve the efficiency of the stochastic optimization, we reduce the problem space with a scenario tree reduction algorithm, that significantly increases tractability, whilst …


Neural Correlates Of User Experience In Gaming, Y. Tejaswini, F. Nah, Keng Siau, L. Chen May 2017

Neural Correlates Of User Experience In Gaming, Y. Tejaswini, F. Nah, Keng Siau, L. Chen

Research Collection School Of Computing and Information Systems

The objective of this research is to understand the neural correlates of user states of experience in human-computer interaction using electroencephalogram (EEG). Such user states include flow, boredom, and anxiety that are experienced when a user interacts with a computer-based system. We propose using a within-subjects experiment to collect EEG data to assess and compare the neural correlates of three main states of user experience (i.e., flow, boredom, and anxiety) as well as compare them with the resting state as a baseline. We expect the findings from this research to contribute to an improved understanding of psychophysiological means of assessing …


Effects Of The Use Of Leaderboards In Education, Yu-Hsien Chiu, Fiona Fui-Hoon Nah May 2017

Effects Of The Use Of Leaderboards In Education, Yu-Hsien Chiu, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

Gamification has been used in education to increase student motivation and performance. In this research, we are interested to examine the effect of leaderboards on student motivation by assessing the interest of students to complete optional practice questions provided to them in a course. Based on goal setting theory and cognitive evaluation theory, we hypothesize that the use of leaderboards will lead to increased student motivation. We designed a within-subject experiment where leaderboards were not provided in the first half of the semester for the optional assignments comprising practice questions but were provided in the second half of the semester …


The Impact Of Monetary Value Gains And Losses On Cybersecurity Behavior, Samuel Noah Smith, Fiona Fui-Hoon Nah, Maggie Cheng, Santosh Kuma Ravindran May 2017

The Impact Of Monetary Value Gains And Losses On Cybersecurity Behavior, Samuel Noah Smith, Fiona Fui-Hoon Nah, Maggie Cheng, Santosh Kuma Ravindran

Research Collection School Of Computing and Information Systems

This research examines if users take more risky cybersecurity actions when presented with the possibility of losing monetary value rather than gaining monetary value. Prospect theory provides the theoretical foundation for the research. An experimental design is proposed to test the hypothesis for the research.


Real-Time Prediction Of Length Of Stay Using Passive Wi-Fi Sensing, Truc Viet Le, Baoyang Song, Laura Wynter May 2017

Real-Time Prediction Of Length Of Stay Using Passive Wi-Fi Sensing, Truc Viet Le, Baoyang Song, Laura Wynter

Research Collection School Of Computing and Information Systems

The proliferation of wireless technologies in today's everyday life is one of the key drivers of the Internet of Things (IoT). In addition to being an enabler of connectivity, the vast penetration of wireless devices today gives rise to a secondary functionality as a means of tracking and localization of the devices themselves. Indeed, in order to discover and automatically connect to known Wi-Fi networks, mobile devices have to scan and broadcast the so-called probe requests on all available channels, which can be captured and analyzed in a non-intrusive manner. Thus, one of the key applications of this feature is …


Lexicons In Sentiment Analytics, B. Yuan, Keng Siau May 2017

Lexicons In Sentiment Analytics, B. Yuan, Keng Siau

Research Collection School Of Computing and Information Systems

With the increasing amount of text data, sentiment analytics (SA) is becoming an important tool for text miners. An automated approach is needed to parse the online reviews and comments, and analyze their sentiments. Since lexicon is the most important component in SA, enhancing the quality of lexicons will improve the efficiency and accuracy of sentiment analysis. In this research, we study the effect of coupling a general lexicon with a specialized lexicon (for a specific domain) and its impact on sentiment analysis. Two special domains and one general domain were used. The two special domains are the petroleum domain …


Encrypted Data Processing With Homomorphic Re-Encryption, Wenxiu Ding, Zheng Yan, Robert H. Deng May 2017

Encrypted Data Processing With Homomorphic Re-Encryption, Wenxiu Ding, Zheng Yan, Robert H. Deng

Research Collection School Of Computing and Information Systems

Cloud computing offers various services to users by re-arranging storage and computing resources. In order to preserve data privacy, cloud users may choose to upload encrypted data rather than raw data to the cloud. However, processing and analyzing encrypted data are challenging problems, which have received increasing attention in recent years. Homomorphic Encryption (HE) was proposed to support computation on encrypted data and ensure data confidentiality simultaneously. However, a limitation of HE is it is a single user system, which means it only allows the party that owns a homomorphic decryption key to decrypt processed ciphertexts. Original HE cannot support …


Impact Of Artificial Intelligence, Robotics, And Machine Learning On Sales And Marketing, Keng Siau, Y. Yang May 2017

Impact Of Artificial Intelligence, Robotics, And Machine Learning On Sales And Marketing, Keng Siau, Y. Yang

Research Collection School Of Computing and Information Systems

AI, robotics, and machine learning are impacting the field of sales and marketing in an unprecedented way. A perfect storm is brewing! On one hand, online retail stores like Amazon are crushing the bricks and mortar stores. Sales and marketing professionals in bricks and mortar stores are facing a grim future. On the other hand, AI, robotics, and machine learning are replacing sales and marketing professionals in online stores. In fact, salespersons and marketers are predicted to be among the first to be replaced by robots. In a face-to-face environment, human may still prefer to interact with another human. In …


Machine Learning Approaches To Sentiment Analytics, W. Zhao, Keng Siau May 2017

Machine Learning Approaches To Sentiment Analytics, W. Zhao, Keng Siau

Research Collection School Of Computing and Information Systems

One key aspect of sentiment analytics is emotion classification. This research studies the use of machine learning approaches to classify human emotion. Two different machine learning approaches were compared in an experimental study. In one approach, emotions from both genders were used to train the machine. In another approach, genders were separated and two separate machines were used to learn the emotions of the two genders. We also manipulated the training sample sizes and study the effect of training sample sizes on the two machine learning approaches. Our preliminary results show that the approach where the genders were separated produces …


Exploiting Contextual Information For Fine-Grained Tweet Geolocation, Wen Haw Chong, Ee Peng Lim May 2017

Exploiting Contextual Information For Fine-Grained Tweet Geolocation, Wen Haw Chong, Ee Peng Lim

Research Collection School Of Computing and Information Systems

The problem of fine-grained tweet geolocation is to link tweets to their posting venues. We solve this in a learning to rank framework by ranking candidate venues given a test tweet. The problem is challenging as tweets are short and the vast majority are non-geocoded, meaning information is sparse for building models. Nonetheless, although only a small fraction of tweets are geocoded, we find that they are posted by a substantial proportion of users. Essentially, such users have location history data. Along with tweet posting time, these serve as additional contextual information for geolocation. In designing our geolocation models, we …


Collaborative Topic Regression For Online Recommender Systems: An Online And Bayesian Approach, Chenghao Liu, Tao Jin, Steven C. H. Hoi, Peilin Zhao, Jianling Sun May 2017

Collaborative Topic Regression For Online Recommender Systems: An Online And Bayesian Approach, Chenghao Liu, Tao Jin, Steven C. H. Hoi, Peilin Zhao, Jianling Sun

Research Collection School Of Computing and Information Systems

Collaborative Topic Regression (CTR) combines ideas of probabilistic matrix factorization (PMF) and topic modeling (such as LDA) for recommender systems, which has gained increasing success in many applications. Despite enjoying many advantages, the existing Batch Decoupled Inference algorithm for the CTR model has some critical limitations: First of all, it is designed to work in a batch learning manner, making it unsuitable to deal with streaming data or big data in real-world recommender systems. Secondly, in the existing algorithm, the item-specific topic proportions of LDA are fed to the downstream PMF but the rating information is not exploited in discovering …


Data-Driven Approach To Measuring The Level Of Press Freedom Using Media Attention Diversity From Unfiltered News, Jisun An, Haewoon Kwak May 2017

Data-Driven Approach To Measuring The Level Of Press Freedom Using Media Attention Diversity From Unfiltered News, Jisun An, Haewoon Kwak

Research Collection School Of Computing and Information Systems

Published by Reporters Without Borders every year, the Press Freedom Index (PFI) reflects the fear and tension in the newsroom pushed by the government and private sectors. While the PFI is invaluable in monitoring media environ- ments worldwide, the current survey-based method has in- herent limitations to updates in terms of cost and time. In this work, we introduce an alternative way to measure the level of press freedom using media attention diversity compiled from Unfiltered News.


Dpweka: Achieving Differential Privacy In Weka, Srinidhi Katla May 2017

Dpweka: Achieving Differential Privacy In Weka, Srinidhi Katla

Graduate Theses and Dissertations

Organizations belonging to the government, commercial, and non-profit industries collect and store large amounts of sensitive data, which include medical, financial, and personal information. They use data mining methods to formulate business strategies that yield high long-term and short-term financial benefits. While analyzing such data, the private information of the individuals present in the data must be protected for moral and legal reasons. Current practices such as redacting sensitive attributes, releasing only the aggregate values, and query auditing do not provide sufficient protection against an adversary armed with auxiliary information. In the presence of additional background information, the privacy protection …


Discovering Your Selling Points: Personalized Social Influential Tags Exploration, Yuchen Li, Kian-Lee Tan, Ju Fan, Dongxiang Zhang May 2017

Discovering Your Selling Points: Personalized Social Influential Tags Exploration, Yuchen Li, Kian-Lee Tan, Ju Fan, Dongxiang Zhang

Research Collection School Of Computing and Information Systems

Social influence has attracted significant attention owing to the prevalence of social networks (SNs). In this paper, we study a new social influence problem, called personalized social influential tags exploration (PITEX), to help any user in the SN explore how she influences the network. Given a target user, it finds a size-k tag set that maximizes this user’s social influence. We prove the problem is NP-hard to be approximated within any constant ratio. To solve it, we introduce a sampling-based framework, which has an approximation ratio of 1−ǫ 1+ǫ with high probabilistic guarantee. To speedup the computation, we devise more …


Dynamic Nearest Neighbor Queries In Euclidean Space, Sarana Nutanong, Mohammed Eunus Ali, Egemen Tanin, Kyriakos Mouratidis May 2017

Dynamic Nearest Neighbor Queries In Euclidean Space, Sarana Nutanong, Mohammed Eunus Ali, Egemen Tanin, Kyriakos Mouratidis

Research Collection School Of Computing and Information Systems

Given a query point q and a set D of data points, a nearest neighbor (NN) query returns the data point p in D that minimizes the distance DIST(q,p), where the distance function DIST(,) is the L2norm. One important variant of this query type is kNN query, which returns k data points with the minimum distances. When taking the temporal dimension into account, the k NN query result may change over a period of time due to changes in locations of the query point and/or data points.


Continuous Top-K Monitoring On Document Streams, Leong Hou U, Junjie Zhang, Kyriakos Mouratidis, Ye Li May 2017

Continuous Top-K Monitoring On Document Streams, Leong Hou U, Junjie Zhang, Kyriakos Mouratidis, Ye Li

Research Collection School Of Computing and Information Systems

The efficient processing of document streams plays an important role in many information filtering systems. Emerging applications, such as news update filtering and social network notifications, demand presenting end-users with the most relevant content to their preferences. In this work, user preferences are indicated by a set of keywords. A central server monitors the document stream and continuously reports to each user the top-k documents that are most relevant to her keywords. Our objective is to support large numbers of users and high stream rates, while refreshing the top-k results almost instantaneously. Our solution abandons the traditional frequency-ordered indexing approach. …


A Data-Driven Approach For Benchmarking Energy Efficiency Of Warehouse Buildings, Wee Leong Lee, Kar Way Tan, Zui Young Lim May 2017

A Data-Driven Approach For Benchmarking Energy Efficiency Of Warehouse Buildings, Wee Leong Lee, Kar Way Tan, Zui Young Lim

Research Collection School Of Computing and Information Systems

This study proposes adata-driven approach for benchmarking energy efficiency of warehouse buildings.Our proposed approach provides an alternative to the limitation of existingbenchmarking approaches where a theoretical energy-efficient warehouse was usedas a reference. Our approach starts by defining the questions needed to capturethe characteristics of warehouses relating to energy consumption. Using an existingdata set of warehouse building containing various attributes, we first cluster theminto groups by their characteristics. The warehouses characteristics derivedfrom the cluster assignments along with their past annual energy consumptionare subsequently used to train a decision tree model. The decision tree providesa classification of what factors contribute to different …


A Neural Network Model For Semi-Supervised Review Aspect Identification, Ying Ding, Changlong Yu, Jing Jiang May 2017

A Neural Network Model For Semi-Supervised Review Aspect Identification, Ying Ding, Changlong Yu, Jing Jiang

Research Collection School Of Computing and Information Systems

Aspect identification is an important problem in opinion mining. It is usually solved in an unsupervised manner, and topic models have been widely used for the task. In this work, we propose a neural network model to identify aspects from reviews by learning their distributional vectors. A key difference of our neural network model from topic models is that we do not use multinomial word distributions but instead embedding vectors to generate words. Furthermore, to leverage review sentences labeled with aspect words, a sequence labeler based on Recurrent Neural Networks (RNNs) is incorporated into our neural network. The resulting model …


Determining The Impact Regions Of Competing Options In Preference Space, Bo Tang, Kyriakos Mouratidis, Man Lung. Yiu May 2017

Determining The Impact Regions Of Competing Options In Preference Space, Bo Tang, Kyriakos Mouratidis, Man Lung. Yiu

Research Collection School Of Computing and Information Systems

In rank-aware processing, user preferences are typically represented by a numeric weight per data attribute, collectively forming a weight vector. The score of an option (data record) is defined as the weighted sum of its individual attributes. The highest-scoring options across a set of alternatives (dataset) are shortlisted for the user as the recommended ones. In that setting, the user input is a vector (equivalently, a point) in a d-dimensional preference space, where d is the number of data attributes. In this paper we study the problem of determining in which regions of the preference space the weight vector should …


Provably Secure Attribute Based Signcryption With Delegated Computation And Efficient Key Updating, Hanshu Hong, Yunhao Xia, Zhixin Sun, Ximeng Liu May 2017

Provably Secure Attribute Based Signcryption With Delegated Computation And Efficient Key Updating, Hanshu Hong, Yunhao Xia, Zhixin Sun, Ximeng Liu

Research Collection School Of Computing and Information Systems

Equipped with the advantages of flexible access control and fine-grained authentication, attribute based signcryption is diffusely designed for security preservation in many scenarios. However, realizing efficient key evolution and reducing the calculation costs are two challenges which should be given full consideration in attribute based cryptosystem. In this paper, we present a key-policy attribute based signcryption scheme (KP-ABSC) with delegated computation and efficient key updating. In our scheme, an access structure is embedded into user’s private key, while ciphertexts corresponds a target attribute set. Only the two are matched can a user decrypt and verify the ciphertexts. When the access …