Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2016

Discipline
Institution
Keyword
Publication
Publication Type

Articles 151 - 180 of 352

Full-Text Articles in Databases and Information Systems

Collaborative Development Of A Small Business Emergency Planning Model, Arthur Henry Hendela May 2016

Collaborative Development Of A Small Business Emergency Planning Model, Arthur Henry Hendela

Dissertations

Small businesses, which are defined by the US Small Business Administration as entities with less than 500 employees, suffer interruptions from diverse risks such as financial events, legal situations, or severe storms exemplified by Hurricane Sandy. Proper preparations can help lessen the length of the interruption and put employees and owners back to work. Large corporations generally have large budgets available for planning, business continuity, and disaster recovery. Small businesses must decide which risks are the most important and how best to mitigate those risks using minimal resources.

This research uses a series of surveys followed by mathematical modeling to …


Mediating Chance Encounters Through Opportunistic Social Matching, Julia M. Mayer May 2016

Mediating Chance Encounters Through Opportunistic Social Matching, Julia M. Mayer

Dissertations

Chance encounters, the unintended meeting between people unfamiliar with each other, serve as an important social lubricant helping people to create new social ties, such as making new friends or finding an activity, study or collaboration partner. Unfortunately, social barriers often prevent chance encounters in environments where people do not know each other and people have to rely on serendipity to meet or be introduced to interesting people around them. Little is known about the underlying dynamics of chance encounters and how systems could utilize contextual data to mediate chance encounters. This dissertation addresses this gap in research literature by …


Hybrid Similarity Function For Big Data Entity Matching With R-Swoosh, Vimal Chandra Gorijala May 2016

Hybrid Similarity Function For Big Data Entity Matching With R-Swoosh, Vimal Chandra Gorijala

Master's Projects

Entity Matching (EM) is the problem of determining if two entities in a data set refer to the same real-world object. For example, it decides if two given mentions in the data, such as “Helen Hunt” and “H. M. Hunt”, refer to the same real-world entity by using different similarity functions. This problem plays a key role in information integration, natural language understanding, information processing on the World-Wide Web, and on the emerging Semantic Web. This project deals with the similarity functions and thresholds utilized in them to determine the similarity of the entities. The work contains two major parts: …


Efficient Pair-Wise Similarity Computation Using Apache Spark, Parineetha Gandhi Tirumali May 2016

Efficient Pair-Wise Similarity Computation Using Apache Spark, Parineetha Gandhi Tirumali

Master's Projects

Entity matching is the process of identifying different manifestations of the same real world entity. These entities can be referred to as objects(string) or data instances. These entities are in turn split over several databases or clusters based on the signatures of the entities. When entity matching algorithms are performed on these databases or clusters, there is a high possibility that a particular entity pair is compared more than once. The number of comparison for any two entities depend on the number of common signatures or keys they possess. This effects the performance of any entity matching algorithm. This paper …


Library Writers Reward Project, Saravana Kumar Gajendran May 2016

Library Writers Reward Project, Saravana Kumar Gajendran

Master's Projects

Open-source library development exploits the distributed intelligence of participants in Internet communities. Nowadays, contribution to the open-source community is fading [16] (Stackalytics, 2016) as there is not much recognition for library writers. They can start exploring ways to generate revenue as they actively contribute to the open-source community.

This project helps library writers to generate revenue in the form of bitcoins for their contribution. Our solution to generate revenue for library writers is to integrate bitcoin mining with existing JavaScript libraries, such as jQuery. More use of the library leads to more revenue for the library writers. It uses the …


Processing Posting Lists Using Opencl, Radha Kotipalli May 2016

Processing Posting Lists Using Opencl, Radha Kotipalli

Master's Projects

One of the main requirements of internet search engines is the ability to retrieve relevant results with faster response times. Yioop is an open source search engine designed and developed in PHP by Dr. Chris Pollett. The goal of this project is to explore the possibilities of enhancing the performance of Yioop by substituting resource-intensive existing PHP functions with C based native PHP extensions and the parallel data processing technology OpenCL. OpenCL leverages the Graphical Processing Unit (GPU) of a computer system for performance improvements.

Some of the critical functions in search engines are resource-intensive in terms of processing power, …


The Mexican Water Forest: Benefits Of Using Remote Sensing Techniques To Assess Changes In Land Use And Land Cover, Maria F. Lopez Ornelas May 2016

The Mexican Water Forest: Benefits Of Using Remote Sensing Techniques To Assess Changes In Land Use And Land Cover, Maria F. Lopez Ornelas

Master's Projects and Capstones

In the past 30 years, anthropogenic activities like urbanization, agriculture, road fragmentation and deforestation have resulted in changes in the land use and land cover (LULC) in the Mexican Water Forest. Due to the important ecosystem services, and the natural resources this forest provides, in Mexico, it has become increasingly necessary to use new technologies and tools to support the planning, implementation and integration of forest management and conservation plans, as well as ecological and socioeconomic analysis of this ecosystem. Remote Sensing techniques and Geographic Information Systems (GIS) have been a true technological and methodological revolution in the acquisition, management …


Exploring Data Mining Techniques For Tree Species Classification Using Co-Registered Lidar And Hyperspectral Data, Julia K. Marrs May 2016

Exploring Data Mining Techniques For Tree Species Classification Using Co-Registered Lidar And Hyperspectral Data, Julia K. Marrs

Theses and Dissertations

NASA Goddard’s LiDAR, Hyperspectral, and Thermal imager provides co-registered remote sensing data on experimental forests. Data mining methods were used to achieve a final tree species classification accuracy of 68% using a combined LiDAR and hyperspectral dataset, and show promise for addressing deforestation and carbon sequestration on a species-specific level.


A Study Of Three Paradigms For Storing Geospatial Data: Distributed-Cloud Model, Relational Database, And Indexed Flat File, Matthew A. Toups May 2016

A Study Of Three Paradigms For Storing Geospatial Data: Distributed-Cloud Model, Relational Database, And Indexed Flat File, Matthew A. Toups

LSU New Orleans Theses and Dissertations

Geographic Information Systems (GIS) and related applications of geospatial data were once a small software niche; today nearly all Internet and mobile users utilize some sort of mapping or location-aware software. This widespread use reaches beyond mere consumption of geodata; projects like OpenStreetMap (OSM) represent a new source of geodata production, sometimes dubbed “Volunteered Geographic Information.” The volume of geodata produced and the user demand for geodata will surely continue to grow, so the storage and query techniques for geospatial data must evolve accordingly.

This thesis compares three paradigms for systems that manage vector data. Over the past few decades …


Identifying Relationships Between Scientific Datasets, Abdussalam Alawini May 2016

Identifying Relationships Between Scientific Datasets, Abdussalam Alawini

Dissertations and Theses

Scientific datasets associated with a research project can proliferate over time as a result of activities such as sharing datasets among collaborators, extending existing datasets with new measurements, and extracting subsets of data for analysis. As such datasets begin to accumulate, it becomes increasingly difficult for a scientist to keep track of their derivation history, which complicates data sharing, provenance tracking, and scientific reproducibility. Understanding what relationships exist between datasets can help scientists recall their original derivation history. For instance, if dataset A is contained in dataset B, then the connection between A and B could be that A was …


Cest: City Event Summarization Using Twitter, Deepa Mallela May 2016

Cest: City Event Summarization Using Twitter, Deepa Mallela

Computer Science Graduate Projects and Theses

Twitter, with 288 million active users, has become the most popular platform for continuous real-time discussions. This leads to huge amounts of information related to the real-world, which has attracted researchers from both academia and industry. Event detection on Twitter has gained attention as one of the most popular domains of interest within the research community. Unfortunately, existing event detection methodologies have yet to fully explore Twitter metadata and instead rely solely on identifying events based on prior information or focus on events that belong to specific categories. Given the heavy volume of tweets that discuss events, summarization techniques can …


#Greysanatomy Vs. #Yankees: Demographics And Hashtag Use On Twitter, Jisun An, Ingmar Weber May 2016

#Greysanatomy Vs. #Yankees: Demographics And Hashtag Use On Twitter, Jisun An, Ingmar Weber

Research Collection School Of Computing and Information Systems

Demographics, in particular, gender, age, and race, are a key predictor of human behavior. Despite the significant effect that demographics plays, most scientific studies using online social media do not consider this factor, mainly due to the lack of such information. In this work, we use state-of-the-art face analysis software to infer gender, age, and race from profile images of 350K Twitter users from New York. For the period from November 1, 2014 to October 31, 2015, we study which hashtags are used by different demographic groups. Though we find considerable overlap for the most popular hashtags, there are also …


Data Partitioning Methods To Process Queries On Encrypted Databases On The Cloud, Osama M. Omran May 2016

Data Partitioning Methods To Process Queries On Encrypted Databases On The Cloud, Osama M. Omran

Graduate Theses and Dissertations

Many features and advantages have been brought to organizations and computer users by Cloud computing. It allows different service providers to distribute many applications and services in an economical way. Consequently, many users and companies have begun using cloud computing. However, the users and companies are concerned about their data when data are stored and managed in the Cloud or outsourcing servers. The private data of individual users and companies is stored and managed by the service providers on the Cloud, which offers services on the other side of the Internet in terms of its users, and consequently results in …


Online Passive-Aggressive Active Learning, Jing Lu, Peilin Zhao, Steven C. H. Hoi May 2016

Online Passive-Aggressive Active Learning, Jing Lu, Peilin Zhao, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

We investigate online active learning techniques for online classification tasks. Unlike traditional supervised learning approaches, either batch or online learning, which often require to request class labels of each incoming instance, online active learning queries only a subset of informative incoming instances to update the classification model, aiming to maximize classification performance with minimal human labelling effort during the entire online learning task. In this paper, we present a new family of online active learning algorithms called Passive-Aggressive Active (PAA) learning algorithms by adapting the Passive-Aggressive algorithms in online active learning settings. Unlike conventional Perceptron-based approaches that employ only the …


Mining Social Ties Beyond Homophily, Hongwei Liang, Ke Wang, Feida Zhu May 2016

Mining Social Ties Beyond Homophily, Hongwei Liang, Ke Wang, Feida Zhu

Research Collection School Of Computing and Information Systems

Summarizing patterns of connections or social tiesin a social network, in terms of attributes information on nodesand edges, holds a key to the understanding of how the actorsinteract and form relationships. We formalize this problem asmining top-k group relationships (GRs), which captures strongsocial ties between groups of actors. While existing works focuson patterns that follow from the well known homophily principle,we are interested in social ties that do not follow from homophily,thus, provide new insights. Finding top-k GRs faces new challenges:it requires a novel ranking metric because traditionalmetrics favor patterns that are expected from the homophilyprinciple; it requires an innovative …


Modeling Human-Like Non-Rationality For Social Agents, Jaroslaw Kochanowicz, Ah-Hwee Tan, Daniel Thalmann May 2016

Modeling Human-Like Non-Rationality For Social Agents, Jaroslaw Kochanowicz, Ah-Hwee Tan, Daniel Thalmann

Research Collection School Of Computing and Information Systems

Humans are not rational beings. Deviations from rationality in human thinking are currently well documented [25] as non-reducible to rational pursuit of egoistic benefit or its occasional distortion with temporary emotional excitation, as it is often assumed. This occurs not only outside conceptual reasoning or rational goal realization but also subconsciously and often in certainty that they did not and could not take place ‘in my case’. Non-rationality can no longer be perceived as a rare affective abnormality in otherwise rational thinking, but as a systemic, permanent quality, ’a design feature’ of human cognition. While social psychology has systematically addressed …


Euclidean Co-Embedding Of Ordinal Data For Multi-Type Visualization, Dung D. Le, Hady W. Lauw May 2016

Euclidean Co-Embedding Of Ordinal Data For Multi-Type Visualization, Dung D. Le, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Embedding deals with reducing the high-dimensional representation of data into a low-dimensional representation. Previous work mostly focuses on preserving similarities among objects. Here, not only do we explicitly recognize multiple types of objects, but we also focus on the ordinal relationships across types. Collaborative Ordinal Embedding or COE is based on generative modelling of ordinal triples. Experiments show that COE outperforms the baselines on objective metrics, revealing its capacity for information preservation for ordinal data.


Fast Weighted Histograms For Bilateral Filtering And Nearest Neighbor Searching, Shengfeng He, Qingxiong Yang, Rynson W. H. Lau, Ming-Hsuan Yang May 2016

Fast Weighted Histograms For Bilateral Filtering And Nearest Neighbor Searching, Shengfeng He, Qingxiong Yang, Rynson W. H. Lau, Ming-Hsuan Yang

Research Collection School Of Computing and Information Systems

The locality sensitive histogram (LSH) injects spatial information into the local histogram in an efficient manner, and has been demonstrated to be very effective for visual tracking. In this paper, we explore the application of this efficient histogram in two important problems. We first extend the LSH to linear time bilateral filtering, and then propose a new type of histogram for efficiently computing edge-preserving nearest neighbor fields (NNFs). While the existing histogram-based bilateral filtering methods are the state of the art for efficient grayscale image processing, they are limited to box spatial filter kernels only. In our first application, we …


Temporal Kernel Descriptors For Learning With Time-Sensitive Patterns, Doyen Sahoo, Abhishek Sharma, Hoi, Steven C. H., Peilin Zhao May 2016

Temporal Kernel Descriptors For Learning With Time-Sensitive Patterns, Doyen Sahoo, Abhishek Sharma, Hoi, Steven C. H., Peilin Zhao

Research Collection School Of Computing and Information Systems

Detecting temporal patterns is one of the most prevalent challenges while mining data. Often, timestamps or information about when certain instances or events occurred can provide us with critical information to recognize temporal patterns. Unfortunately, most existing techniques are not able to fully extract useful temporal information based on the time (especially at different resolutions of time). They miss out on 3 crucial factors: (i) they do not distinguish between timestamp features (which have cyclical or periodic properties) and ordinary features; (ii) they are not able to detect patterns exhibited at different resolutions of time (e.g. different patterns at the …


Hdidx: High-Dimensional Indexing For Efficient Approximate Nearest Neighbor Search, Ji Wan, Sheng Tang, Yongdong Zhang, Jintao Li, Pengcheng Wu, Steven C. H. Hoi May 2016

Hdidx: High-Dimensional Indexing For Efficient Approximate Nearest Neighbor Search, Ji Wan, Sheng Tang, Yongdong Zhang, Jintao Li, Pengcheng Wu, Steven C. H. Hoi

Research Collection School Of Computing and Information Systems

Fast Nearest Neighbor (NN) search is a fundamental challenge in large-scale data processing and analytics, particularly for analyzing multimedia contents which are often of high dimensionality. Instead of using exact NN search, extensive research efforts have been focusing on approximate NN search algorithms. In this work, we present "HDIdx", an efficient high-dimensional indexing library for fast approximate NN search, which is open-source and written in Python. It offers a family of state-of-the-art algorithms that convert input high-dimensional vectors into compact binary codes, making them very efficient and scalable for NN search with very low space complexity.


Efspredictor: Predicting Configuration Bugs With Ensemble Feature Selection, Bowen Xu, David Lo, Xin Xia, Ashish Sureka, Shanping Li May 2016

Efspredictor: Predicting Configuration Bugs With Ensemble Feature Selection, Bowen Xu, David Lo, Xin Xia, Ashish Sureka, Shanping Li

Research Collection School Of Computing and Information Systems

The configuration of a system determines the system behavior and wrong configuration settings can adversely impact system's availability, performance, and correctness. We refer to these wrong configuration settings as configuration bugs. The importance of configuration bugs has prompted many researchers to study it, and past studies can be grouped into three categories: detection, localization, and fixing of configuration bugs. In the work, we focus on the detection of configuration bugs, in particular, we follow the line-of-work that tries to predict if a bug report is caused by a wrong configuration setting. Automatically prediction of whether a bug is a configuration …


Semantic Proximity Search On Graphs With Metagraph-Based Learning, Yuan Fang, Wenqing Lin, Vincent W. Zheng, Min Wu, Kevin Chen-Chuan Chang, Xiao-Li Li May 2016

Semantic Proximity Search On Graphs With Metagraph-Based Learning, Yuan Fang, Wenqing Lin, Vincent W. Zheng, Min Wu, Kevin Chen-Chuan Chang, Xiao-Li Li

Research Collection School Of Computing and Information Systems

Given ubiquitous graph data such as the Web and social networks, proximity search on graphs has been an active research topic. The task boils down to measuring the proximity between two nodes on a graph. Although most earlier studies deal with homogeneous or bipartite graphs only, many real-world graphs are heterogeneous with objects of various types, giving rise to different semantic classes of proximity. For instance, on a social network two users can be close for different reasons, such as being classmates or family members, which represent two distinct classes of proximity. Thus, it becomes inadequate to only measure a …


Learning To Query: Focused Web Page Harvesting For Entity Aspects, Yuan Fang, Vincent W. Zheng, Kevin Chen-Chuan Chang May 2016

Learning To Query: Focused Web Page Harvesting For Entity Aspects, Yuan Fang, Vincent W. Zheng, Kevin Chen-Chuan Chang

Research Collection School Of Computing and Information Systems

As the Web hosts rich information about real-world entities, our information quests become increasingly entity centric. In this paper, we study the problem of focused harvesting of Web pages for entity aspects, to support downstream applications such as business analytics and building a vertical portal. Given that search engines are the de facto gateways to assess information on the Web, we recognize the essence of our problem as Learning to Query (L2Q) - to intelligently select queries so that we can harvest pages, via a search engine, focused on an entity aspect of interest. Thus, it is crucial to quantify …


On Unravelling Opinions Of Issue Specific-Silent Users In Social Media, Wei Gong, Ee-Peng Lim, Feida Zhu, Pei Hua Cher May 2016

On Unravelling Opinions Of Issue Specific-Silent Users In Social Media, Wei Gong, Ee-Peng Lim, Feida Zhu, Pei Hua Cher

Research Collection School Of Computing and Information Systems

Social media has become a popular platform for people toshare opinions. Among the social media mining researchprojects that study user opinions and issues, most focus onanalyzing posted and shared content. They could run into thedanger of non-representative findings as the opinions of userswho do not post content are overlooked, which often happensin today’s marketing, recommendation, and social sensing research.For a more complete and representative profiling ofuser opinions on various topical issues, we need to investigatethe opinions of the users even when they stay silent onthese issues. We call these users the issue specific-silent users(i-silent users). To study them and their …


Joint Search By Social And Spatial Proximity [Extended Abstract], Kyriakos Mouratidis, Jing Li, Yu Tang, Nikos Mamoulis May 2016

Joint Search By Social And Spatial Proximity [Extended Abstract], Kyriakos Mouratidis, Jing Li, Yu Tang, Nikos Mamoulis

Research Collection School Of Computing and Information Systems

The diffusion of social networks introduces new challengesand opportunities for advanced services, especially so with their ongoingaddition of location-based features. We show how applications like company andfriend recommendation could significantly benefit from incorporating social andspatial proximity, and study a query type that captures these twofold semantics.We develop highly scalable algorithms for its processing, and use real socialnetwork data to empirically verify their efficiency and efficacy.


Efficient Verifiable Computation Of Linear And Quadratic Functions Over Encrypted Data, Ngoc Hieu Tran, Hwee Hwa Pang, Robert H. Deng May 2016

Efficient Verifiable Computation Of Linear And Quadratic Functions Over Encrypted Data, Ngoc Hieu Tran, Hwee Hwa Pang, Robert H. Deng

Research Collection School Of Computing and Information Systems

In data outsourcing, a client stores a large amount of data on an untrusted server; subsequently, the client can request the server to compute a function on any subset of the data. This setting naturally leads to two security requirements: confidentiality of input data, and authenticity of computations. Existing approaches that satisfy both requirements simultaneously are built on fully homomorphic encryption, which involves expensive computation on the server and client and hence is impractical. In this paper, we propose two verifiable homomorphic encryption schemes that do not rely on fully homomorphic encryption. The first is a simple and efficient scheme …


Online Sparse Passive Aggressive Learning With Kernels, Jing Lu, Peilin Zhao, Hoi, Steven C. H. May 2016

Online Sparse Passive Aggressive Learning With Kernels, Jing Lu, Peilin Zhao, Hoi, Steven C. H.

Research Collection School Of Computing and Information Systems

Conventional online kernel methods often yield an unboundedlarge number of support vectors, making them inefficient and non-scalable forlarge-scale applications. Recent studies on bounded kernel-based onlinelearning have attempted to overcome this shortcoming. Although they can boundthe number of support vectors at each iteration, most of them fail to bound thenumber of support vectors for the final output solution which is often obtainedby averaging the series of solutions over all the iterations. In this paper, wepropose a novel kernel-based online learning method, Sparse Passive Aggressivelearning (SPA), which can output a final solution with a bounded number ofsupport vectors. The key idea of …


A Key-Insulated Cp-Abe With Key Exposure Accountability For Secure Data Sharing In The Cloud, Hanshu Hong, Zhixin Sun, Ximeng Liu May 2016

A Key-Insulated Cp-Abe With Key Exposure Accountability For Secure Data Sharing In The Cloud, Hanshu Hong, Zhixin Sun, Ximeng Liu

Research Collection School Of Computing and Information Systems

ABE has become an effective tool for data protection in cloud computing. However, since users possessing the same attributes share the same private keys, there exist some malicious users exposing their private keys deliberately for illegal data sharing without being detected, which will threaten the security of the cloud system. Such issues remain in many current ABE schemes since the private keys are rarely associated with any user specific identifiers. In order to achieve user accountability as well as provide key exposure protection, in this paper, we propose a key-insulated ciphertext policy attribute based encryption with key exposure accountability (KI-CPABE-KEA). …


Are You Charlie Or Ahmed? Cultural Pluralism In Charlie Hebdo Response On Twitter, Jisun An, Haewoon Kwak, Yelena Mejova, Sonia Alonso Saenz De Oger, Braulio Gomez Fortes May 2016

Are You Charlie Or Ahmed? Cultural Pluralism In Charlie Hebdo Response On Twitter, Jisun An, Haewoon Kwak, Yelena Mejova, Sonia Alonso Saenz De Oger, Braulio Gomez Fortes

Research Collection School Of Computing and Information Systems

We study the response to the Charlie Hebdo shootings of January 7, 2015 on Twitter across the globe. We ask whether the stances on the issue of freedom of speech can be modeled using established sociological theories, including Huntington’s culturalist Clash of Civilizations, and those taking into consideration social context, including Density and Interdependence theories. We find support for Huntington’s culturalist explanation, in that the established traditions and norms of one’s “civilization” predetermine some of one’s opinion. However, at an individual level, we also find social context to play a significant role, with non-Arabs living in Arab countries using #JeSuisAhmed …


Learning Adversary Behavior In Security Games: A Pac Model Perspective, Arunesh Sinha, Debarun Kar, Milind Tambe May 2016

Learning Adversary Behavior In Security Games: A Pac Model Perspective, Arunesh Sinha, Debarun Kar, Milind Tambe

Research Collection School Of Computing and Information Systems

Recent applications of Stackelberg Security Games (SSG), from wildlife crime to urban crime, have employed machine learning tools to learn and predict adversary behavior using available data about defender-adversary interactions. Given these recent developments, this paper commits to an approach of directly learning the response function of the adversary. Using the PAC model, this paper lays a firm theoretical foundation for learning in SSGs (e.g., theoretically answer questions about the numbers of samples required to learn adversary behavior) and provides utility guarantees when the learned adversary model is used to plan the defender's strategy. The paper also aims to answer …