Exploring Data Mining Techniques For Tree Species Classification Using Co-Registered Lidar And Hyperspectral Data,
2016
CUNY Hunter College
Exploring Data Mining Techniques For Tree Species Classification Using Co-Registered Lidar And Hyperspectral Data, Julia K. Marrs
Theses and Dissertations
NASA Goddard’s LiDAR, Hyperspectral, and Thermal imager provides co-registered remote sensing data on experimental forests. Data mining methods were used to achieve a final tree species classification accuracy of 68% using a combined LiDAR and hyperspectral dataset, and show promise for addressing deforestation and carbon sequestration on a species-specific level.
A Study Of Three Paradigms For Storing Geospatial Data: Distributed-Cloud Model, Relational Database, And Indexed Flat File,
2016
University of New Orleans
A Study Of Three Paradigms For Storing Geospatial Data: Distributed-Cloud Model, Relational Database, And Indexed Flat File, Matthew A. Toups
LSU New Orleans Theses and Dissertations
Geographic Information Systems (GIS) and related applications of geospatial data were once a small software niche; today nearly all Internet and mobile users utilize some sort of mapping or location-aware software. This widespread use reaches beyond mere consumption of geodata; projects like OpenStreetMap (OSM) represent a new source of geodata production, sometimes dubbed “Volunteered Geographic Information.” The volume of geodata produced and the user demand for geodata will surely continue to grow, so the storage and query techniques for geospatial data must evolve accordingly.
This thesis compares three paradigms for systems that manage vector data. Over the past few decades …
Identifying Relationships Between Scientific Datasets,
2016
Portland State University
Identifying Relationships Between Scientific Datasets, Abdussalam Alawini
Dissertations and Theses
Scientific datasets associated with a research project can proliferate over time as a result of activities such as sharing datasets among collaborators, extending existing datasets with new measurements, and extracting subsets of data for analysis. As such datasets begin to accumulate, it becomes increasingly difficult for a scientist to keep track of their derivation history, which complicates data sharing, provenance tracking, and scientific reproducibility. Understanding what relationships exist between datasets can help scientists recall their original derivation history. For instance, if dataset A is contained in dataset B, then the connection between A and B could be that A was …
Cest: City Event Summarization Using Twitter,
2016
Boise State University
Cest: City Event Summarization Using Twitter, Deepa Mallela
Computer Science Graduate Projects and Theses
Twitter, with 288 million active users, has become the most popular platform for continuous real-time discussions. This leads to huge amounts of information related to the real-world, which has attracted researchers from both academia and industry. Event detection on Twitter has gained attention as one of the most popular domains of interest within the research community. Unfortunately, existing event detection methodologies have yet to fully explore Twitter metadata and instead rely solely on identifying events based on prior information or focus on events that belong to specific categories. Given the heavy volume of tweets that discuss events, summarization techniques can …
#Greysanatomy Vs. #Yankees: Demographics And Hashtag Use On Twitter,
2016
Singapore Management University
#Greysanatomy Vs. #Yankees: Demographics And Hashtag Use On Twitter, Jisun An, Ingmar Weber
Research Collection School Of Computing and Information Systems
Demographics, in particular, gender, age, and race, are a key predictor of human behavior. Despite the significant effect that demographics plays, most scientific studies using online social media do not consider this factor, mainly due to the lack of such information. In this work, we use state-of-the-art face analysis software to infer gender, age, and race from profile images of 350K Twitter users from New York. For the period from November 1, 2014 to October 31, 2015, we study which hashtags are used by different demographic groups. Though we find considerable overlap for the most popular hashtags, there are also …
Data Partitioning Methods To Process Queries On Encrypted Databases On The Cloud,
2016
University of Arkansas, Fayetteville
Data Partitioning Methods To Process Queries On Encrypted Databases On The Cloud, Osama M. Omran
Graduate Theses and Dissertations
Many features and advantages have been brought to organizations and computer users by Cloud computing. It allows different service providers to distribute many applications and services in an economical way. Consequently, many users and companies have begun using cloud computing. However, the users and companies are concerned about their data when data are stored and managed in the Cloud or outsourcing servers. The private data of individual users and companies is stored and managed by the service providers on the Cloud, which offers services on the other side of the Internet in terms of its users, and consequently results in …
Online Passive-Aggressive Active Learning,
2016
Singapore Management University
Online Passive-Aggressive Active Learning, Jing Lu, Peilin Zhao, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
We investigate online active learning techniques for online classification tasks. Unlike traditional supervised learning approaches, either batch or online learning, which often require to request class labels of each incoming instance, online active learning queries only a subset of informative incoming instances to update the classification model, aiming to maximize classification performance with minimal human labelling effort during the entire online learning task. In this paper, we present a new family of online active learning algorithms called Passive-Aggressive Active (PAA) learning algorithms by adapting the Passive-Aggressive algorithms in online active learning settings. Unlike conventional Perceptron-based approaches that employ only the …
Mining Social Ties Beyond Homophily,
2016
Singapore Management University
Mining Social Ties Beyond Homophily, Hongwei Liang, Ke Wang, Feida Zhu
Research Collection School Of Computing and Information Systems
Summarizing patterns of connections or social tiesin a social network, in terms of attributes information on nodesand edges, holds a key to the understanding of how the actorsinteract and form relationships. We formalize this problem asmining top-k group relationships (GRs), which captures strongsocial ties between groups of actors. While existing works focuson patterns that follow from the well known homophily principle,we are interested in social ties that do not follow from homophily,thus, provide new insights. Finding top-k GRs faces new challenges:it requires a novel ranking metric because traditionalmetrics favor patterns that are expected from the homophilyprinciple; it requires an innovative …
Modeling Human-Like Non-Rationality For Social Agents,
2016
Singapore Management University
Modeling Human-Like Non-Rationality For Social Agents, Jaroslaw Kochanowicz, Ah-Hwee Tan, Daniel Thalmann
Research Collection School Of Computing and Information Systems
Humans are not rational beings. Deviations from rationality in human thinking are currently well documented [25] as non-reducible to rational pursuit of egoistic benefit or its occasional distortion with temporary emotional excitation, as it is often assumed. This occurs not only outside conceptual reasoning or rational goal realization but also subconsciously and often in certainty that they did not and could not take place ‘in my case’. Non-rationality can no longer be perceived as a rare affective abnormality in otherwise rational thinking, but as a systemic, permanent quality, ’a design feature’ of human cognition. While social psychology has systematically addressed …
Euclidean Co-Embedding Of Ordinal Data For Multi-Type Visualization,
2016
Singapore Management University
Euclidean Co-Embedding Of Ordinal Data For Multi-Type Visualization, Dung D. Le, Hady W. Lauw
Research Collection School Of Computing and Information Systems
Embedding deals with reducing the high-dimensional representation of data into a low-dimensional representation. Previous work mostly focuses on preserving similarities among objects. Here, not only do we explicitly recognize multiple types of objects, but we also focus on the ordinal relationships across types. Collaborative Ordinal Embedding or COE is based on generative modelling of ordinal triples. Experiments show that COE outperforms the baselines on objective metrics, revealing its capacity for information preservation for ordinal data.
Fast Weighted Histograms For Bilateral Filtering And Nearest Neighbor Searching,
2016
Singapore Management University
Fast Weighted Histograms For Bilateral Filtering And Nearest Neighbor Searching, Shengfeng He, Qingxiong Yang, Rynson W. H. Lau, Ming-Hsuan Yang
Research Collection School Of Computing and Information Systems
The locality sensitive histogram (LSH) injects spatial information into the local histogram in an efficient manner, and has been demonstrated to be very effective for visual tracking. In this paper, we explore the application of this efficient histogram in two important problems. We first extend the LSH to linear time bilateral filtering, and then propose a new type of histogram for efficiently computing edge-preserving nearest neighbor fields (NNFs). While the existing histogram-based bilateral filtering methods are the state of the art for efficient grayscale image processing, they are limited to box spatial filter kernels only. In our first application, we …
Temporal Kernel Descriptors For Learning With Time-Sensitive Patterns,
2016
Singapore Management University
Temporal Kernel Descriptors For Learning With Time-Sensitive Patterns, Doyen Sahoo, Abhishek Sharma, Hoi, Steven C. H., Peilin Zhao
Research Collection School Of Computing and Information Systems
Detecting temporal patterns is one of the most prevalent challenges while mining data. Often, timestamps or information about when certain instances or events occurred can provide us with critical information to recognize temporal patterns. Unfortunately, most existing techniques are not able to fully extract useful temporal information based on the time (especially at different resolutions of time). They miss out on 3 crucial factors: (i) they do not distinguish between timestamp features (which have cyclical or periodic properties) and ordinary features; (ii) they are not able to detect patterns exhibited at different resolutions of time (e.g. different patterns at the …
Hdidx: High-Dimensional Indexing For Efficient Approximate Nearest Neighbor Search,
2016
Chinese Academy of Sciences
Hdidx: High-Dimensional Indexing For Efficient Approximate Nearest Neighbor Search, Ji Wan, Sheng Tang, Yongdong Zhang, Jintao Li, Pengcheng Wu, Steven C. H. Hoi
Research Collection School Of Computing and Information Systems
Fast Nearest Neighbor (NN) search is a fundamental challenge in large-scale data processing and analytics, particularly for analyzing multimedia contents which are often of high dimensionality. Instead of using exact NN search, extensive research efforts have been focusing on approximate NN search algorithms. In this work, we present "HDIdx", an efficient high-dimensional indexing library for fast approximate NN search, which is open-source and written in Python. It offers a family of state-of-the-art algorithms that convert input high-dimensional vectors into compact binary codes, making them very efficient and scalable for NN search with very low space complexity.
Efspredictor: Predicting Configuration Bugs With Ensemble Feature Selection,
2016
Zhejiang University
Efspredictor: Predicting Configuration Bugs With Ensemble Feature Selection, Bowen Xu, David Lo, Xin Xia, Ashish Sureka, Shanping Li
Research Collection School Of Computing and Information Systems
The configuration of a system determines the system behavior and wrong configuration settings can adversely impact system's availability, performance, and correctness. We refer to these wrong configuration settings as configuration bugs. The importance of configuration bugs has prompted many researchers to study it, and past studies can be grouped into three categories: detection, localization, and fixing of configuration bugs. In the work, we focus on the detection of configuration bugs, in particular, we follow the line-of-work that tries to predict if a bug report is caused by a wrong configuration setting. Automatically prediction of whether a bug is a configuration …
Semantic Proximity Search On Graphs With Metagraph-Based Learning,
2016
Singapore Management University
Semantic Proximity Search On Graphs With Metagraph-Based Learning, Yuan Fang, Wenqing Lin, Vincent W. Zheng, Min Wu, Kevin Chen-Chuan Chang, Xiao-Li Li
Research Collection School Of Computing and Information Systems
Given ubiquitous graph data such as the Web and social networks, proximity search on graphs has been an active research topic. The task boils down to measuring the proximity between two nodes on a graph. Although most earlier studies deal with homogeneous or bipartite graphs only, many real-world graphs are heterogeneous with objects of various types, giving rise to different semantic classes of proximity. For instance, on a social network two users can be close for different reasons, such as being classmates or family members, which represent two distinct classes of proximity. Thus, it becomes inadequate to only measure a …
Learning To Query: Focused Web Page Harvesting For Entity Aspects,
2016
Singapore Management University
Learning To Query: Focused Web Page Harvesting For Entity Aspects, Yuan Fang, Vincent W. Zheng, Kevin Chen-Chuan Chang
Research Collection School Of Computing and Information Systems
As the Web hosts rich information about real-world entities, our information quests become increasingly entity centric. In this paper, we study the problem of focused harvesting of Web pages for entity aspects, to support downstream applications such as business analytics and building a vertical portal. Given that search engines are the de facto gateways to assess information on the Web, we recognize the essence of our problem as Learning to Query (L2Q) - to intelligently select queries so that we can harvest pages, via a search engine, focused on an entity aspect of interest. Thus, it is crucial to quantify …
On Unravelling Opinions Of Issue Specific-Silent Users In Social Media,
2016
Singapore Management University
On Unravelling Opinions Of Issue Specific-Silent Users In Social Media, Wei Gong, Ee-Peng Lim, Feida Zhu, Pei Hua Cher
Research Collection School Of Computing and Information Systems
Social media has become a popular platform for people toshare opinions. Among the social media mining researchprojects that study user opinions and issues, most focus onanalyzing posted and shared content. They could run into thedanger of non-representative findings as the opinions of userswho do not post content are overlooked, which often happensin today’s marketing, recommendation, and social sensing research.For a more complete and representative profiling ofuser opinions on various topical issues, we need to investigatethe opinions of the users even when they stay silent onthese issues. We call these users the issue specific-silent users(i-silent users). To study them and their …
Joint Search By Social And Spatial Proximity [Extended Abstract],
2016
Singapore Management University
Joint Search By Social And Spatial Proximity [Extended Abstract], Kyriakos Mouratidis, Jing Li, Yu Tang, Nikos Mamoulis
Research Collection School Of Computing and Information Systems
The diffusion of social networks introduces new challengesand opportunities for advanced services, especially so with their ongoingaddition of location-based features. We show how applications like company andfriend recommendation could significantly benefit from incorporating social andspatial proximity, and study a query type that captures these twofold semantics.We develop highly scalable algorithms for its processing, and use real socialnetwork data to empirically verify their efficiency and efficacy.
Efficient Verifiable Computation Of Linear And Quadratic Functions Over Encrypted Data,
2016
Singapore Management University
Efficient Verifiable Computation Of Linear And Quadratic Functions Over Encrypted Data, Ngoc Hieu Tran, Hwee Hwa Pang, Robert H. Deng
Research Collection School Of Computing and Information Systems
In data outsourcing, a client stores a large amount of data on an untrusted server; subsequently, the client can request the server to compute a function on any subset of the data. This setting naturally leads to two security requirements: confidentiality of input data, and authenticity of computations. Existing approaches that satisfy both requirements simultaneously are built on fully homomorphic encryption, which involves expensive computation on the server and client and hence is impractical. In this paper, we propose two verifiable homomorphic encryption schemes that do not rely on fully homomorphic encryption. The first is a simple and efficient scheme …
Online Sparse Passive Aggressive Learning With Kernels,
2016
Singapore Management University
Online Sparse Passive Aggressive Learning With Kernels, Jing Lu, Peilin Zhao, Hoi, Steven C. H.
Research Collection School Of Computing and Information Systems
Conventional online kernel methods often yield an unboundedlarge number of support vectors, making them inefficient and non-scalable forlarge-scale applications. Recent studies on bounded kernel-based onlinelearning have attempted to overcome this shortcoming. Although they can boundthe number of support vectors at each iteration, most of them fail to bound thenumber of support vectors for the final output solution which is often obtainedby averaging the series of solutions over all the iterations. In this paper, wepropose a novel kernel-based online learning method, Sparse Passive Aggressivelearning (SPA), which can output a final solution with a bounded number ofsupport vectors. The key idea of …
