Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons™

Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 5281 - 5310 of 7258

Full-Text Articles in Computer Sciences

Fuzzy Matching Of Web Queries To Structured Data, Tao Cheng, Hady W. Lauw, Stelios Paparizos Mar 2010

Fuzzy Matching Of Web Queries To Structured Data, Tao Cheng, Hady W. Lauw, Stelios Paparizos

Research Collection School Of Computing and Information Systems

Recognizing the alternative ways people use to reference an entity, is important for many Web applications that query structured data. In such applications, there is often a mismatch between how content creators describe entities and how different users try to retrieve them. In this paper, we consider the problem of determining whether a candidate query approximately matches with an entity. We propose an off-line, data-driven, bottom-up approach that mines query logs for instances where Web content creators and Web users apply a variety of strings to refer to the same Web pages. This way, given a set of strings that …


Homophily In The Digital World: A Livejournal Case Study, Hady W. Lauw, John C. Shafer, Rakesh Agrawal, Alexandros Ntoulas Mar 2010

Homophily In The Digital World: A Livejournal Case Study, Hady W. Lauw, John C. Shafer, Rakesh Agrawal, Alexandros Ntoulas

Research Collection School Of Computing and Information Systems

Are two users more likely to be friends if they share common interests? Are two users more likely to share common interests if they're friends? The authors study the phenomenon of homophily in the digital world by answering these central questions. Unlike the physical world, the digital world doesn't impose any geographic or organizational constraints on friendships. So, although online friends might share common interests, a priori there's no reason to believe that two users with common interests are more likely to be friends. Using data from LiveJournal, the authors show that the answer to both questions is yes.


Information-Quality Aware Routing In Event-Driven Sensor Networks, Hwee Xian Tan, Mun-Choon Chan, Wendong Xiao, Peng-Yong Kong, Chen-Khong Tham Mar 2010

Information-Quality Aware Routing In Event-Driven Sensor Networks, Hwee Xian Tan, Mun-Choon Chan, Wendong Xiao, Peng-Yong Kong, Chen-Khong Tham

Research Collection School Of Computing and Information Systems

Upon the occurrence of a phenomenon of interest in a wireless sensor network, multiple sensors may be activated, leading to data implosion and redundancy. Data aggregation and/or fusion techniques exploit spatio-temporal correlation among sensory data to reduce traffic load and mitigate congestion. However, this is often at the expense of loss in Information Quality (IQ) of data that is collected at the fusion center. In this work, we address the problem of finding the least-cost routing tree that satisfies a given IQ constraint. We note that the optimal least-cost routing solution is a variation of the classical NP-hard Steiner tree …


Preference Queries In Large Multi-Cost Transportation Networks, Kyriakos Mouratidis, Yimin Lin, Man Lung Yiu Mar 2010

Preference Queries In Large Multi-Cost Transportation Networks, Kyriakos Mouratidis, Yimin Lin, Man Lung Yiu

Research Collection School Of Computing and Information Systems

Research on spatial network databases has so far considered that there is a single cost value associated with each road segment of the network. In most real-world situations, however, there may exist multiple cost types involved in transportation decision making. For example, the different costs of a road segment could be its Euclidean length, the driving time, the walking time, possible toll fee, etc. The relative significance of these cost types may vary from user to user. In this paper we consider such multi-cost transportation networks (MCN), where each edge (road segment) is associated with multiple cost values. We formulate …


Efficient Verification Of Shortest Path Search Via Authenticated Hints, Man Lung Yiu, Yimin Lin, Kyriakos Mouratidis Mar 2010

Efficient Verification Of Shortest Path Search Via Authenticated Hints, Man Lung Yiu, Yimin Lin, Kyriakos Mouratidis

Research Collection School Of Computing and Information Systems

Shortest path search in transportation networks is unarguably one of the most important online search services nowadays (e.g., Google Maps, MapQuest, etc), with applications spanning logistics, spatial optimization, or everyday driving decisions. Often times, the owner of the road network data (e.g., a transport authority) provides its database to third-party query services, which are responsible for answering shortest path queries posed by their clients. The issue arising here is that a query service might be returning sub-optimal paths either purposely (in order to serve its own purposes like computational savings or commercial reasons) or because it has been compromised by …


Top-K Aggregation Queries Over Large Networks, Xifeng Yan, Bin He, Feida Zhu, Jiawei Han Mar 2010

Top-K Aggregation Queries Over Large Networks, Xifeng Yan, Bin He, Feida Zhu, Jiawei Han

Research Collection School Of Computing and Information Systems

Searching and mining large graphs today is critical to a variety of application domains, ranging from personalized recommendation in social networks, to searches for functional associations in biological pathways. In these domains, there is a need to perform aggregation operations on large-scale networks. Unfortunately the existing implementation of aggregation operations on relational databases does not guarantee superior performance in network space, especially when it involves edge traversals and joins of gigantic tables. In this paper, we investigate the neighborhood aggregation queries: Find nodes that have top-k highest aggregate values over their h-hop neighbors. While these basic queries are common in …


K-Anonymity In The Presence Of External Databases, Dimitris Sacharidis, Kyriakos Mouratidis, Dimitris Papadias Mar 2010

K-Anonymity In The Presence Of External Databases, Dimitris Sacharidis, Kyriakos Mouratidis, Dimitris Papadias

Research Collection School Of Computing and Information Systems

The concept of k-anonymity has received considerable attention due to the need of several organizations to release microdata without revealing the identity of individuals. Although all previous k-anonymity techniques assume the existence of a public database (PD) that can be used to breach privacy, none utilizes PD during the anonymization process. Specifically, existing generalization algorithms create anonymous tables using only the microdata table (MT) to be published, independently of the external knowledge available. This omission leads to high information loss. Motivated by this observation we first introduce the concept of k-join-anonymity (KJA), which permits more effective generalization to reduce the …


Local Coordination Under Bounded Rationality: Coase Meets Simon, Finds Hayek, C. Jason Woodard Mar 2010

Local Coordination Under Bounded Rationality: Coase Meets Simon, Finds Hayek, C. Jason Woodard

Research Collection School Of Computing and Information Systems

This paper explores strategic behavior in a network of firms using an agent-based model. The model exhibits a tension between economic efficiency and the stability of the network in the face of incentives to change its configuration. This tension is to be expected because the conditions of the Coase theorem are violated: the boundedly rational firms in the model lack the ability to discover efficient network configurations or achieve them through collective action. In computational experiments, as predicted by theory, firms frequently became locked into inefficient outcomes or endless cycles of mutual frustration. However, simple institutional innovations such as property …


Survey Of Environmental Data Portals: Features, Characteristics, And Reliability Issues, Shahram Latifi, David Walker Feb 2010

Survey Of Environmental Data Portals: Features, Characteristics, And Reliability Issues, Shahram Latifi, David Walker

2010 Annual Nevada NSF EPSCoR Climate Change Conference

24 PowerPoint slides Session 2: Infrastructure Convener: Sergiu Dascalu, UNR Abstract: -What is a Data Portal? -Presents information from diverse sources in a unified way -Enables instant, reliable and secure exchange of information over the Web -The "portal" concept is to offer a single web page that aggregates content from several systems or servers.


Player Performance Prediction In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Richa Sharan, Jaideep Srivastava Feb 2010

Player Performance Prediction In Massively Multiplayer Online Role-Playing Games (Mmorpgs), Kyong Jin Shim, Richa Sharan, Jaideep Srivastava

Research Collection School Of Computing and Information Systems

Recent years have seen an ever increasing number of people interacting in the online space. Massively multiplayer online role-playing games (MMORPGs) are personal computer or console-based digital games where thousands of players can simultaneously sign on to the same online, persistent virtual world to interact and collaborate with each other through their in-game characters. In recent years, researchers have found virtual environments to be a sound venue for studying learning, collaboration, social participation, literacy in online space, and learning trajectory at the individual level as well as at the group level. While many games today provide web and GUI-based reports …


Estimating The Quality Of Postings In The Real-Time Web, Hady W. Lauw, Alexandros Ntoulas, Krishnaram Kenthapadi Feb 2010

Estimating The Quality Of Postings In The Real-Time Web, Hady W. Lauw, Alexandros Ntoulas, Krishnaram Kenthapadi

Research Collection School Of Computing and Information Systems

Millions of users are posting their status updates, interesting findings, news, ideas and observations in real-time on microblogging services such as Twitter, Jaiku and Plurk. This real-time Web can be a great resource of valuable timely information. Since the real-time Web is completely open and decentralized and anyone may post information at whim, distinguishing interesting and popular postings from the mundane ones is a challenging task. In this paper we study the problem of estimating the quality (or “interestingness”) of postings in the real-time Web. We identify several important factors that are indicative of the quality of postings, and present …


Privacy-Preserving Similarity-Based Text Retrieval, Hwee Hwa Pang, Jialie Shen, Ramayya Krishnan Feb 2010

Privacy-Preserving Similarity-Based Text Retrieval, Hwee Hwa Pang, Jialie Shen, Ramayya Krishnan

Research Collection School Of Computing and Information Systems

Users of online services are increasingly wary that their activities could disclose confidential information on their business or personal activities. It would be desirable for an online document service to perform text retrieval for users, while protecting the privacy of their activities. In this article, we introduce a privacy-preserving, similarity-based text retrieval scheme that (a) prevents the server from accurately reconstructing the term composition of queries and documents, and (b) anonymizes the search results from unauthorized observers. At the same time, our scheme preserves the relevance-ranking of the search server, and enables accounting of the number of documents that each …


Twitterrank: Finding Topic-Sensitive Influential Twitterers, Jianshu Weng, Ee Peng Lim, Jing Jiang, Qi He Feb 2010

Twitterrank: Finding Topic-Sensitive Influential Twitterers, Jianshu Weng, Ee Peng Lim, Jing Jiang, Qi He

Research Collection School Of Computing and Information Systems

This paper focuses on the problem of identifying influential users of micro-blogging services. Twitter, one of the most notable micro-blogging services, employs a social-networking model called "following", in which each user can choose who she wants to "follow" to receive tweets from without requiring the latter to give permission first. In a dataset prepared for this study, it is observed that (1) 72.4% of the users in Twitter follow more than 80% of their followers, and (2) 80.5% of the users have 80% of users they are following follow them back. Our study reveals that the presence of "reciprocity" can …


Efficient Valid Scope For Location-Dependent Spatial Queries In Mobile Environments, Ken C. K. Lee, Wang-Chien Lee, Hong Va Leong, Brandon Unger, Baihua Zheng Feb 2010

Efficient Valid Scope For Location-Dependent Spatial Queries In Mobile Environments, Ken C. K. Lee, Wang-Chien Lee, Hong Va Leong, Brandon Unger, Baihua Zheng

Research Collection School Of Computing and Information Systems

In mobile environments, mobile clients can access information with respect to their locations by submitting Location-Dependent Spatial Queries (LDSQs) to Location-Based Service (LBS) servers. Owing to scarce wireless channel bandwidth and limited client battery life, frequent LDSQ submission from clients must be avoided. Observing that LDSQs issued from a client located at nearby positions would likely return the same query results, we explore the idea of valid scope, which represents a spatial area in which a set of LDSQs will retrieve exactly the same set of query results. With a valid scope derived and an LDSQ result cached, a client …


Understanding Cognitive Differences In Processing Competing Visualizations Of Complex Systems, Madhavi Mukul Chakrabarty Jan 2010

Understanding Cognitive Differences In Processing Competing Visualizations Of Complex Systems, Madhavi Mukul Chakrabarty

Dissertations

Node-link diagrams are used represent systems having different elements and relationships among the elements. Representing the systems using visualizations like node-link diagrams provides cognitive aid to individuals in understanding the system and effectively managing these systems. Using appropriate visual tools aids in task completion by reducing the cognitive load of individuals in understanding the problems and solving them. However, the visualizations that are currently developed lack any cognitive processing based evaluation. Most of the evaluations (if any) are based on the result of tasks performed using these visualizations. Therefore, the evaluations do not provide any perspective from the point of …


Job Seeking And Job Application In Social Networking Sites : Predicting Job Seekers' Behavioral Intentions, Maria Marcella Plummer Jan 2010

Job Seeking And Job Application In Social Networking Sites : Predicting Job Seekers' Behavioral Intentions, Maria Marcella Plummer

Dissertations

Social networking sites (SNSs) are revolutionizing the way in which employers and job seekers connect and interact with each other. Despite the reported benefits of SNSs with respect to finding a job, there are issues such as privacy concerns that might be deterring job seekers from using these sites in their attempts to secure a job. It is therefore important to understand the factors that are salient in predicting job seekers' use of SNSs in applying for jobs.

In this research, a theoretical model was developed to explicate job seekers' intentions to use SNSs to apply for jobs. Two aspects …


Virtual World Commerce Adoption (Vwca) : A Case Study Of Second Life Investigating The Impacts Of Perceived Affordances, Trust, And Need Satisfaction, Kamolbhan Olapiriyakul Jan 2010

Virtual World Commerce Adoption (Vwca) : A Case Study Of Second Life Investigating The Impacts Of Perceived Affordances, Trust, And Need Satisfaction, Kamolbhan Olapiriyakul

Dissertations

Virtual worlds are computer-simulated worlds in which multi-players can simultaneously interact in a rich graphical environment. The development of virtual worlds, along with the massive growth of users, creates opportunities for business organizations. This dissertation involves many studies regarding virtual world adoption in business by virtual consumers.

Most of the research in Information Systems (IS) was conducted investigating factors influencing technology adoption, such as ease of use and usefulness, subjective norms and behavioral controls, self-efficacy, performance and effort expectancy, flow, etc. However, most of these research studies focused neither on design aspects related to affordances nor users' goal-oriented behaviors, such …


Skos And The Semantic Web: Knowledge Organization, Metadata, And Interoperability, Eric A. Robinson Jan 2010

Skos And The Semantic Web: Knowledge Organization, Metadata, And Interoperability, Eric A. Robinson

Other Topics

The Simplified Knowledge Organization System (SKOS) is a Semantic Web framework, based on the Resource Description Framework (RDF) for thesauri, classification schemes and simple ontologies. It allows for machine-actionable description of the structure of these knowledge organization systems (KOS) and provides an excellent tool for addressing interoperability and vocabulary control problems inherent to the rapidly expanding information environment of the Web. This paper discusses the foundations of the SKOS framework and reviews the literature on a variety of SKOS implementations. The limitations of SKOS that have been revealed through its broad application are addressed with brief attention to the proposed …


Understanding Events Through Analysis Of Social Media, Amit P. Sheth, Hemant Purohit, Ashutosh Sopan Jadhav, Pavan Kapanipathi, Lu Chen Jan 2010

Understanding Events Through Analysis Of Social Media, Amit P. Sheth, Hemant Purohit, Ashutosh Sopan Jadhav, Pavan Kapanipathi, Lu Chen

Kno.e.sis Publications

Users are sharing vast amounts of social data through social networking platforms accessible by Web and increasingly via mobile devices. This opens an exciting opportunity to extract social perceptions as well as obtain insights relevant to events around us. We discuss the significant need and opportunity for analyzing event-centric user generated content on social networks, present some of the technical challenges and our approach to address them. This includes aggregating social data related to events of interest, along with Web resources (news, Wikipedia pages, multimedia) related to an event of interest, and supporting analysis along spatial, temporal, thematic, and sentiment …


Mobicloud - Making Clouds Reachable: A Toolkit For Easy And Efficient Development Of Customized Cloud Mobile Hybrid Applications, Ashwin Manjunatha, Ajith Harshana Ranabahu, Amit P. Sheth, Krishnaprasad Thirunarayan Jan 2010

Mobicloud - Making Clouds Reachable: A Toolkit For Easy And Efficient Development Of Customized Cloud Mobile Hybrid Applications, Ashwin Manjunatha, Ajith Harshana Ranabahu, Amit P. Sheth, Krishnaprasad Thirunarayan

Kno.e.sis Publications

The advancements in computing have resulted in a boom of cheap, ubiquitous, connected mobile devices, as well as seemingly unlimited, utility style, pay as you go computing resources, commonly referred to as Cloud computing. However, taking full advantage of this mobile and cloud computing landscape, especially for the data intensive domains, has been hampered by the many heterogeneities that exist in the mobile space, as well as the Cloud space. Our research attempts to exploit the capabilities of the mobile and cloud landscape by introducing MobiCloud, an online toolkit to efficiently develop Cloud-mobile hybrid (CMH) applications. We define a CMH …


A Sketch-Based Language For Representing Uncertainty In The Locations Of Origin Of Herbarium Specimens, Barry Kronenfeld, Andrew Weeks Jan 2010

A Sketch-Based Language For Representing Uncertainty In The Locations Of Origin Of Herbarium Specimens, Barry Kronenfeld, Andrew Weeks

Faculty Research and Creative Activity

Uncertainty fields have been suggested as an appropriate model for retrospective georeferencing of herbarium specimens. Previous work has focused only on automated data capture methods, but techniques for manual data specification may be able to harness human spatial cognition skills to quickly interpret complex spatial propositions. This paper develops a formal modeling language by which location uncertainty fields can be derived from manually sketched features. The language consists of low-level specification of critical probability isolines from which a surface can be uniquely derived, and high-level specification of features and predicates from which low-level isolines can be derived. In a case …


Online Collaborative Whiteboard, Niranjana Kumari Subramanian Jan 2010

Online Collaborative Whiteboard, Niranjana Kumari Subramanian

Theses Digitization Project

The purpose of this project is to build a simple Online Collaborative Whiteboard (OCW) application to enable multiple users to brainstorm and share ideas interactively by using fun tools to start thinking together. Includes source code.


Nominal Schemas For Integrating Rules And Ontologies, Frederick Maier, Adila A. Krisnadhi, Pascal Hitzler Jan 2010

Nominal Schemas For Integrating Rules And Ontologies, Frederick Maier, Adila A. Krisnadhi, Pascal Hitzler

Computer Science and Engineering Faculty Publications

We propose a description-logic style extension of OWL DL, which includes DL-safe variable SWRL and seamlessly integrates datalog rules. Our language also sports a tractable fragment, which we call ELP 2, covering OWL EL, OWL RL, most of OWL QL, and variable restricted datalog.


Spreadsheets Grow Up: Three Spreadsheet Engineering Methodologies For Large Financial Planning Models, Thomas A. Grossman Jr., O Ozluk Jan 2010

Spreadsheets Grow Up: Three Spreadsheet Engineering Methodologies For Large Financial Planning Models, Thomas A. Grossman Jr., O Ozluk

Business Analytics and Information Systems

Many large financial planning models are written in a spreadsheet programming language (usually Microsoft Excel) and deployed as a spreadsheet application. Three groups, FAST Alliance, Operis Group, and BPM Analytics (under the name “Spreadsheet Standards Review Board”) have independently promulgated standardized processes for efficiently building such models. These spreadsheet engineering methodologies provide detailed guidance on design, construction process, and quality control. We summarize and compare these methodologies. They share many design practices, and standardized, mechanistic procedures to construct spreadsheets. We learned that a written book or standards document is by itself insufficient to understand a methodology. These methodologies represent a …


Scale: A Scalable Framework For Efficiently Clustering Transactional Data, Hua Yan, Keke Chen, Ling Liu, Zhang Yi Jan 2010

Scale: A Scalable Framework For Efficiently Clustering Transactional Data, Hua Yan, Keke Chen, Ling Liu, Zhang Yi

Kno.e.sis Publications

This paper presents SCALE, a fully automated transactional clustering framework. The SCALE design highlights three unique features. First, we introduce the concept of Weighted Coverage Density as a categorical similarity measure for efficient clustering of transactional datasets. The concept of weighted coverage density is intuitive and it allows the weight of each item in a cluster to be changed dynamically according to the occurrences of items. Second, we develop the weighted coverage density measure based clustering algorithm, a fast, memory-efficient, and scalable clustering algorithm for analyzing transactional data. Third, we introduce two clustering validation metrics and show that these domain …


From Questions To Effective Answers: On The Utility Of Knowledge-Driven Querying Systems For Life Sciences Data, Amir H. Asiaee, Prashant Doshi, Todd Minning, Satya S. Sahoo, Priti Parikh, Amit P. Sheth, Rick L. Tarleton Jan 2010

From Questions To Effective Answers: On The Utility Of Knowledge-Driven Querying Systems For Life Sciences Data, Amir H. Asiaee, Prashant Doshi, Todd Minning, Satya S. Sahoo, Priti Parikh, Amit P. Sheth, Rick L. Tarleton

Kno.e.sis Publications

We compare two distinct approaches for querying data in the context of the life sciences. The first approach utilizes conventional databases to store the data and intuitive form-based interfaces to facilitate easy querying of the data. These interfaces could be seen as implementing a set of 'pre-canned' queries commonly used by the life science researchers that we study. The second approach is based on semantic Web technologies and is knowledge (model) driven. It utilizes a large OWL ontology and same datasets as before but associated as RDF instances of the ontology concepts. An intuitive interface is provided that allows the …


Sensor Discovery On Linked Data, Josh Pschorr, Cory Andrew Henson, Harshal Kamlesh Patni, Amit P. Sheth Jan 2010

Sensor Discovery On Linked Data, Josh Pschorr, Cory Andrew Henson, Harshal Kamlesh Patni, Amit P. Sheth

Kno.e.sis Publications

There has been a drive recently to make sensor data accessible on the Web. However, because of the vast number of sensors collecting data about our environment, finding relevant sensors on the Web is a non-trivial challenge. In this paper, we present an approach to discovering sensors through a standard service interface over Linked Data. This is accomplished with a semantic sensor network middleware that includes a sensor registry on Linked Data and a sensor discovery service that extends the OGC Sensor Web Enablement. With this approach, we are able to access and discover sensors that are positioned near named-locations …


Sensor Data And Perception: Can Sensors Play 20 Questions, Cory Andrew Henson Jan 2010

Sensor Data And Perception: Can Sensors Play 20 Questions, Cory Andrew Henson

Kno.e.sis Publications

Currently, there are many sensors collecting information about our environment, leading to an overwhelming number of observations that must be analyzed and explained in order to achieve situation awareness. As perceptual beings, we are also constantly inundated with sensory data, yet we are able to make sense of our environment with relative ease. Why is the task of perception so easy for us, and so hard for machines; and could this have anything to do with how we play the game 20 Questions?


A Qualitative Examination Of Topical Tweet And Retweet Practices, Meenakshi Nagarajan, Hemant Purohit, Amit P. Sheth Jan 2010

A Qualitative Examination Of Topical Tweet And Retweet Practices, Meenakshi Nagarajan, Hemant Purohit, Amit P. Sheth

Kno.e.sis Publications

This work contributes to the study of retweet behavior on Twitter surrounding real-world events. We analyze over a million tweets pertaining to three events, present general tweet properties in such topical datasets and qualitatively analyze the properties of the retweet behavior surrounding the most tweeted/viral content pieces. Findings include a clear relationship between sparse/dense retweet patterns and the content and type of a tweet itself; suggesting the need to study content properties in link-based diffusion models.


A Sketch-Based Language For Representing Uncertainty In The Locations Of Origin Of Herbarium Specimens, Barry J. Kronenfeld, Andrew Weeks Jan 2010

A Sketch-Based Language For Representing Uncertainty In The Locations Of Origin Of Herbarium Specimens, Barry J. Kronenfeld, Andrew Weeks

Faculty Research and Creative Activity

Uncertainty fields have been suggested as an appropriate model for retrospective georeferencing of herbarium specimens. Previous work has focused only on automated data capture methods, but techniques for manual data specification may be able to harness human spatial cognition skills to quickly interpret complex spatial propositions. This paper develops a formal modeling language by which location uncertainty fields can be derived from manually sketched features. The language consists of low-level specification of critical probability isolines from which a surface can be uniquely derived, and high-level specification of features and predicates from which low-level isolines can be derived. In a case …