Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

2007

Discipline
Institution
Keyword
Publication
Publication Type

Articles 121 - 150 of 200

Full-Text Articles in Databases and Information Systems

Estimating The Cardinality Of Rdf Graph Patterns, Angela Maduko, Kemafor Anyanwu, Amit P. Sheth, Paul Schliekelman May 2007

Estimating The Cardinality Of Rdf Graph Patterns, Angela Maduko, Kemafor Anyanwu, Amit P. Sheth, Paul Schliekelman

Kno.e.sis Publications

Most RDF query languages allow for graph structure search through a conjunction of triples which is typically processed using join operations. A key factor in optimizing joins is determining the join order which depends on the expected cardinality of intermediate results. This work proposes a pattern-based summarization framework for estimating the cardinality of RDF graph patterns. We present experiments on real world and synthetic datasets which confirm the feasibility of our approach.


Altering Document Term Vectors For Classification - Ontologies As Expectations Of Co-Occurrence, Meenakshi Nagarajan, Amit P. Sheth, Marcos Aguilera, Kimberly Keeton, Arif Merchant, Mustafa Uysal May 2007

Altering Document Term Vectors For Classification - Ontologies As Expectations Of Co-Occurrence, Meenakshi Nagarajan, Amit P. Sheth, Marcos Aguilera, Kimberly Keeton, Arif Merchant, Mustafa Uysal

Kno.e.sis Publications

In this paper we extend the state-of-the-art in utilizing background knowledge for supervised classification by exploiting the semantic relationships between terms explicated in Ontologies. Preliminary evaluations indicate that the new approach generally improves precision and recall, more so for hard to classify cases and reveals patterns indicating the usefulness of such background knowledge.


Semantic Web: Technologies And Applications For The Real-World, Amit P. Sheth May 2007

Semantic Web: Technologies And Applications For The Real-World, Amit P. Sheth

Kno.e.sis Publications

No abstract provided.


Visualization Of Events In A Spatially And Multimedia Enriched Virtual Environment, Leonidas Deligiannidis, Farshad Hakimpour, Amit P. Sheth May 2007

Visualization Of Events In A Spatially And Multimedia Enriched Virtual Environment, Leonidas Deligiannidis, Farshad Hakimpour, Amit P. Sheth

Kno.e.sis Publications

Semantic Event Tracker (SET) is a highly interactive visualization tool for tracking and associating activities (events) in a spatially and Multimedia Enriched Virtual Environment. SET provides integrated views of information spaces while providing overview and detail to improve perception and evaluation of complex scenarios. We model an event as an object that describes an action and its location, time, and relations to other objects. Real world event information is extracted from Internet sources, then stored and processed using Semantic Web technologies that enable us to discover semantic associations between events. We use RDF graphs to represent semantic metadata and ontologies. …


An Experiment In Integrating Large Biomedical Knowledge Resources With Rdf: Application To Associating Genotype And Phenotype Information, Satya S. Sahoo, Olivier Bodenreider, Kelly Zeng, Amit P. Sheth May 2007

An Experiment In Integrating Large Biomedical Knowledge Resources With Rdf: Application To Associating Genotype And Phenotype Information, Satya S. Sahoo, Olivier Bodenreider, Kelly Zeng, Amit P. Sheth

Kno.e.sis Publications

Bridging between genotype and phenotype is generally achieved through the integration of knowledge sources such as Entrez Gene (EG), Online Mendelian Inheritance in Man (OMIM) and the Gene Ontology (GO). Traditionally, such integration implies manual effort or the development of customized software. In this paper, we demonstrate how the Resource Description Framework (RDF) can be used to represent and integrate these resources and support complex queries over the unified resource. We illustrate the effectiveness of our approach by answering a real-world biomedical query linking a specific molecular function, glycosyltransferase, to the disorder congenital muscular dystrophy, which potentially forms a new …


The Influence Of Transactive Memory On Mutual Knowledge In Virtual Teams: A Theoretical Proposal, Alanah Davis, Deepak Khazanchi May 2007

The Influence Of Transactive Memory On Mutual Knowledge In Virtual Teams: A Theoretical Proposal, Alanah Davis, Deepak Khazanchi

Information Systems and Quantitative Analysis Faculty Proceedings & Presentations

Advancements in information technologies (IT) have enabled the ability to exchange knowledge within and across organizations through virtual teams. However, the ability to effectively communicate and share knowledge in virtual settings can become a difficult task due to the complex nature of both the virtual context and the technology used to support them. This paper argues that transactive memory theory can explain how mutual knowledge enhances virtual team performance. We present a conceptual model and theoretical propositions for the study of the relationship between transactive memory and mutual knowledge in virtual teams.


The Effects Of Pairing Participants In Facilitated Group Support Systems Sessions, John D. Murphy, Deepak Khazanchi May 2007

The Effects Of Pairing Participants In Facilitated Group Support Systems Sessions, John D. Murphy, Deepak Khazanchi

Information Systems and Quantitative Analysis Faculty Proceedings & Presentations

Group Support Systems (GSS) have been used to support facilitated ideation sessions for years and have been studied from a number of different perspectives. Throughout this time the norm for running electronic brainstorming sessions has been for participants to work on their own workstations. A review of applicable literature suggests that pairing participants at GSS workstations could result in higher quality inputs and participant satisfaction. This proposition is examined with a lab experiment to test for differences between paired and unpaired facilitated GSS sessions. The results of the experiment suggest that pairing participants does yield higher quality ideas from facilitated …


Residual-Based Measurement Of Peer And Link Lifetimes In Gnutella Networks, Xiaoming Wang, Zhongmei Yao, Dmitri Loguinov May 2007

Residual-Based Measurement Of Peer And Link Lifetimes In Gnutella Networks, Xiaoming Wang, Zhongmei Yao, Dmitri Loguinov

Computer Science Faculty Publications

Existing methods of measuring lifetimes in P2P systems usually rely on the so-called create-based method (CBM), which divides a given observation window into two halves and samples users "created" in the first half every Delta time units until they die or the observation period ends. Despite its frequent use, this approach has no rigorous accuracy or overhead analysis in the literature. To shed more light on its performance, we flrst derive a model for CBM and show that small window size or large Delta may lead to highly inaccurate lifetime distributions. We then show that create-based sampling exhibits an inherent …


On Node Isolation Under Churn In Unstructured P2p Networks With Heavy-Tailed Lifetimes, Zhongmei Yao, Xiaoming Wang, Dmitri Loguinov May 2007

On Node Isolation Under Churn In Unstructured P2p Networks With Heavy-Tailed Lifetimes, Zhongmei Yao, Xiaoming Wang, Dmitri Loguinov

Computer Science Faculty Publications

Previous analytical studies [12], [18] of unstructured P2P resilience have assumed exponential user lifetimes and only considered age-independent neighbor replacement. In this paper, we overcome these limitations by introducing a general node-isolation model for heavy-tailed user lifetimes and arbitrary neighbor-selection algorithms. Using this model, we analyze two age-biased neighbor-selection strategies and show that they significantly improve the residual lifetimes of chosen users, which dramatically reduces the probability of user isolation and graph partitioning compared to uniform selection of neighbors. In fact, the second strategy based on random walks on age-weighted graphs demonstrates that for lifetimes with infinite variance, the system …


Employing Social Capital By Small & Medium Enterprises To Bear Fruit From Wireless Communications, Abdelnasser Abdelaal, Mehruz Kamal, Peter Wolcott May 2007

Employing Social Capital By Small & Medium Enterprises To Bear Fruit From Wireless Communications, Abdelnasser Abdelaal, Mehruz Kamal, Peter Wolcott

Information Systems and Quantitative Analysis Faculty Proceedings & Presentations

Wireless and mobile communications can save Small and Medium Enterprises (SMEs) significant time, money, and effort due to the mobility, flexibility, and ease of use mobile devices provide. SMEs that use such innovations can improve productivity, decrease costs, and enhance the quality of the business process. Lacking technical skills and financial resources, SMEs need special support from local communities and governments in order to survive the severe competition of big chain stores. This paper proposes a model for SMEs to adopt new innovations—those of wireless communications—by employing social capital. We have used a case study approach to show that social …


Gprune: A Constraint Pushing Framework For Graph Pattern Mining, Feida Zhu, Xifeng Yan, Jiawei Han, Philip S. Yu May 2007

Gprune: A Constraint Pushing Framework For Graph Pattern Mining, Feida Zhu, Xifeng Yan, Jiawei Han, Philip S. Yu

Research Collection School Of Computing and Information Systems

In graph mining applications, there has been an increasingly strong urge for imposing user-specified constraints on the mining results. However, unlike most traditional itemset constraints, structural constraints, such as density and diameter of a graph, are very hard to be pushed deep into the mining process. In this paper, we give the first comprehensive study on the pruning properties of both traditional and structural constraints aiming to reduce not only the pattern search space but the data search space as well. A new general framework, called gPrune, is proposed to incorporate all the constraints in such a way that they …


Analysis Of Topological Characteristics Of Huge Online Social Networking Services, Yong-Yeol Ahn, Seungyeop Han, Haewoon Kwak, Sue Moon, Hawoong Jeong May 2007

Analysis Of Topological Characteristics Of Huge Online Social Networking Services, Yong-Yeol Ahn, Seungyeop Han, Haewoon Kwak, Sue Moon, Hawoong Jeong

Research Collection School Of Computing and Information Systems

Social networking services are a fast-growing business in the Internet. However, it is unknown if online relationships and their growth patterns are the same as in real-life social networks. In this paper, we compare the structures of three online social networking services: Cyworld, MySpace, and orkut, each with more than 10 million users, respectively. We have access to complete data of Cyworld's ilchon (friend) relationships and analyze its degree distribution, clustering property, degree correlation, and evolution over time. We also use Cyworld data to evaluate the validity of snowball sampling method, which we use to crawl and obtain partial network …


Cognitive Evaluation Of Information Modeling Methods, Keng Siau, Yuan Wang May 2007

Cognitive Evaluation Of Information Modeling Methods, Keng Siau, Yuan Wang

Research Collection School Of Computing and Information Systems

In the field of information system engineering, information modeling method is a technique to capture user requirements and to understand system complexity. The importance of information modeling has been recognized by practitioners and researchers, but little has been explored to analyze the available information modeling methods or to evaluate them in terms of their strengths, weaknesses, and effectiveness. This research analyzes six information-modeling methods: use case diagram, rich picture diagram, entity-relationship diagram, Trochim’s concept mapping, repertory grid, and causal mapping. These information-modeling methods are analyzed from a cognitive perspective in order to better understand their nature, the assumptions, and the …


Learning To Classify E-Mail, Irena Koprinska, Josiah Poon, James Clark, Jason Yuk Hin Chan May 2007

Learning To Classify E-Mail, Irena Koprinska, Josiah Poon, James Clark, Jason Yuk Hin Chan

Research Collection School Of Computing and Information Systems

In this paper we study supervised and semi-supervised classification of e-mails. We consider two tasks: filing e-mails into folders and spam e-mail filtering. Firstly, in a supervised learning setting, we investigate the use of random forest for automatic e-mail filing into folders and spam e-mail filtering. We show that random forest is a good choice for these tasks as it runs fast on large and high dimensional databases, is easy to tune and is highly accurate, outperforming popular algorithms such as decision trees, support vector machines and naive Bayes. We introduce a new accurate feature selector with linear time complexity. …


Semantic Web Applications In Industry, Government, Health Care And Life Sciences, Amit P. Sheth Apr 2007

Semantic Web Applications In Industry, Government, Health Care And Life Sciences, Amit P. Sheth

Kno.e.sis Publications

No abstract provided.


Control Of The Electronic Management Of Information, A. Boone, R. Szatmary Jr. Apr 2007

Control Of The Electronic Management Of Information, A. Boone, R. Szatmary Jr.

Publications (YM)

This procedure establishes the responsibilities and provides direction for developing and evaluating the adequacy of process controls on specific uses of electronically stored information. These uses include, but are not limited to, information used in design input, developed as design output, or developed as input to or output from scientific investigation or performance assessment modeling and analysis. This pertains to information that resides in an electronic information management system or on electronic media.


Spatiotemporal And Thematic Semantic Analytics, Matthew Perry Apr 2007

Spatiotemporal And Thematic Semantic Analytics, Matthew Perry

Kno.e.sis Publications

No abstract provided.


Towards Attack-Resilient Geometric Data Perturbation, Keke Chen, Ling Liu Apr 2007

Towards Attack-Resilient Geometric Data Perturbation, Keke Chen, Ling Liu

Kno.e.sis Publications

Data perturbation is a popular technique for privacy-preserving data mining. The major challenge of data perturbation is balancing privacy protection and data quality, which are normally considered as a pair of contradictive factors. We propose that selectively preserving only the task/model specific information in perturbation would improve the balance. Geometric data perturbation, consisting of random rotation perturbation, random translation perturbation, and noise addition, aims at preserving the important geometric properties of a multidimensional dataset, while providing better privacy guarantee for data classification modeling. The preliminary study has shown that random geometric perturbation can well preserve model accuracy for several popular …


Mining Minimal Distinguishing Subsequence Patterns With Gap Constraints, Xiaonan Ji, James Bailey, Guozhu Dong Apr 2007

Mining Minimal Distinguishing Subsequence Patterns With Gap Constraints, Xiaonan Ji, James Bailey, Guozhu Dong

Kno.e.sis Publications

Discovering contrasts between collections of data is an important task in data mining. In this paper, we introduce a new type of contrast pattern, called a Minimal Distinguishing Subsequence (MDS). An MDS is a minimal subsequence that occurs frequently in one class of sequences and infrequently in sequences of another class. It is a natural way of representing strong and succinct contrast information between two sequential datasets and can be useful in applications such as protein comparison, document comparison and building sequential classification models. Mining MDS patterns is a challenging task and is significantly different from mining contrasts between relational/transactional …


Social Network Structures In Open Source Software Development Teams, Y. Long, Keng Siau Apr 2007

Social Network Structures In Open Source Software Development Teams, Y. Long, Keng Siau

Research Collection School Of Computing and Information Systems

Drawing on social network theories and previous studies, this research examines the dynamics of social network structures in open source software (OSS) teams. Three projects were selected from SourceForge.net in terms of their similarities as well as their differences. Monthly data were extracted from the bug tracking systems in order to achieve a longitudinal view of the interaction pattern of each project. Social network analysis was used to generate the indices of social structure. The finding suggests that the interaction pattern of OSS projects evolves from a single hub at the beginning to a corel periphery model as the projects …


A Multimodal And Multilevel Ranking Framework For Content-Based Video Retrieval, Steven C. H. Hoi, Michael R. Lyu Apr 2007

A Multimodal And Multilevel Ranking Framework For Content-Based Video Retrieval, Steven C. H. Hoi, Michael R. Lyu

Research Collection School Of Computing and Information Systems

One critical task in content-based video retrieval is to rank search results with combinations of multimodal resources effectively. This paper proposes a novel multimodal and multilevel ranking framework for content-based video retrieval. The main idea of our approach is to represent videos by graphs and learn harmonic ranking functions through fusing multimodal resources over these graphs smoothly. We further tackle the efficiency issue by a multilevel learning scheme, which makes the semi-supervised ranking method practical for large-scale applications. Our empirical evaluations on TRECVID 2005 dataset show that the proposed multimodal and multilevel ranking framework is effective and promising for content-based …


A Multimodal And Multilevel Ranking Framework For Content-Based Video Retrieval, Steven C. H. Hoi, Michael R. Lyu Apr 2007

A Multimodal And Multilevel Ranking Framework For Content-Based Video Retrieval, Steven C. H. Hoi, Michael R. Lyu

Research Collection School Of Computing and Information Systems

One critical task in content-based video retrieval is to rank search results with combinations of multimodal resources effectively. This paper proposes a novel multimodal and multilevel ranking framework for content-based video retrieval. The main idea of our approach is to represent videos by graphs and learn harmonic ranking functions through fusing multimodal resources over these graphs smoothly. We further tackle the efficiency issue by a multilevel learning scheme, which makes the semi-supervised ranking method practical for large-scale applications. Our empirical evaluations on TRECVID 2005 dataset show that the proposed multimodal and multilevel ranking framework is effective and promising for content-based …


Mining Colossal Frequent Patterns By Core Pattern Fusion, Feida Zhu, Xifeng Yan, Jiawei Han, Philip S. Yu, Hong Cheng Apr 2007

Mining Colossal Frequent Patterns By Core Pattern Fusion, Feida Zhu, Xifeng Yan, Jiawei Han, Philip S. Yu, Hong Cheng

Research Collection School Of Computing and Information Systems

Extensive research for frequent-pattern mining in the past decade has brought forth a number of pattern mining algorithms that are both effective and efficient. However, the existing frequent-pattern mining algorithms encounter challenges at mining rather large patterns, called colossal frequent patterns, in the presence of an explosive number of frequent patterns. Colossal patterns are critical to many applications, especially in domains like bioinformatics. In this study, we investigate a novel mining approach called Pattern-Fusion to efficiently find a good approximation to the colossal patterns. With Pattern-Fusion, a colossal pattern is discovered by fusing its small core patterns in one step, …


Summarizing Review Scores Of "Unequal" Reviewers, Hady W. Lauw, Ee Peng Lim, Ke Wang Apr 2007

Summarizing Review Scores Of "Unequal" Reviewers, Hady W. Lauw, Ee Peng Lim, Ke Wang

Research Collection School Of Computing and Information Systems

A frequently encountered problem in decision making is the following review problem: review a large number of objects and select a small number of the best ones. An example is selecting conference papers from a large number of submissions. This problem involves two sub-problems: assigning reviewers to each object, and summarizing reviewers ’ scores into an overall score that supposedly reflects the quality of an object. In this paper, we address the score summarization sub-problem for the scenario where a small number of reviewers evaluate each object. Simply averaging the scores may not work as even a single reviewer could …


Using Concept Maps To More Efficiently Create Intelligence Information Models, Christopher E. Coryell Mar 2007

Using Concept Maps To More Efficiently Create Intelligence Information Models, Christopher E. Coryell

Theses and Dissertations

Information models are a critical tool that enables intelligence customers to quickly and accurately comprehend U.S. intelligence agency products. The Knowledge Pre-positioning System (KPS) is the standard repository for information models at the National Air and Space Intelligence Center (NASIC). The current approach used by NASIC to build a KPS information model is laborious and costly. Intelligence analysts design an information model using a manual, butcher-paper-based process. The output of their work is then entered into KPS by either a single NASIC KPS "database modeler" or a contractor (at a cost of roughly $100K to the organization). This thesis proposes …


Multi-Dimensional Range Querying Using A Modification Of The Skip Graph, Gregory J. Brault Mar 2007

Multi-Dimensional Range Querying Using A Modification Of The Skip Graph, Gregory J. Brault

Theses and Dissertations

Skip graphs are an application layer-based distributed routing data structure that can be used in a sensor network to facilitate user queries of data collected by the sensor nodes. This research investigates the impact of a proposed modification to the skip graph proposed by Aspnes and Shah. Nodes contained in a standard skip graph are sorted by their key value into successively smaller groups based on random membership vectors computed locally at each node. The proposed modification inverts the node key and membership vector roles, where group membership is computed deterministically and node keys are computed randomly. Both skip graph …


Valuing Information Technology Infrastructures: A Growth Options Approach, Qizhi Dai, Robert J. Kauffman, Salvatore T. March Mar 2007

Valuing Information Technology Infrastructures: A Growth Options Approach, Qizhi Dai, Robert J. Kauffman, Salvatore T. March

Research Collection School Of Computing and Information Systems

Decisions to invest in information technology (IT) infrastructure are often made based on an assessment of its immediate value to the organization. However, an important source of value comes from the fact that such technologies have the potential to be leveraged in the development of future applications. From a real options perspective, IT infrastructure investments create growth options that can be exercised if and when an organization decides to develop systems to provide new or enhanced IT capabilities. We present an analytical model based on real options that shows the process by which this potential is converted into business value, …


Towards The Development Of A Defensive Cyber Damage And Mission Impact Methodology, Larry W. Fortson Jr. Mar 2007

Towards The Development Of A Defensive Cyber Damage And Mission Impact Methodology, Larry W. Fortson Jr.

Theses and Dissertations

The purpose of this research is to establish a conceptual methodological framework that will facilitate effective cyber damage and mission impact assessment and reporting following a cyber-based information incidents. Joint and service guidance requires mission impact reporting, but current efforts to implement such reporting have proven ineffective. This research seeks to understand the impediments existing in the current implementation and to propose an improved methodology. The research employed a hybrid historical analysis and case study methodology for data collection through extensive literature review, examination of existing case study research and interviews with Air Force members and civilian personnel employed as …


Tube (Text-Cube) For Discovering Documentary Evidence Of Associations Among Entities, Hady Lauw, Ee Peng Lim, Hwee Hwa Pang Mar 2007

Tube (Text-Cube) For Discovering Documentary Evidence Of Associations Among Entities, Hady Lauw, Ee Peng Lim, Hwee Hwa Pang

Research Collection School Of Computing and Information Systems

User-driven discovery of associations among entities, and documents that provide evidence for these associations, is an important search task conducted by researchers and do-main information specialists. Entities here refer to real or abstract objects such as people, organizations, ideologies, etc. Associations are the inter-relationships among entities. Most current works in query-driven document retrieval and finding representative subgraphs are ill-suited for the task as they lack an awareness of entity types as well as an intuitive representation of associations. We propose the TUBE model, a text cube approach for discovering associations and documentary evidence of these associations. The model consists of …


Automatic Composition Of Semantic Web Services Using Process And Data Mediation, Zixin Wu, Ajith H. Ranabahu, Karthik Gomadam, Amit P. Sheth, John A. Miller Feb 2007

Automatic Composition Of Semantic Web Services Using Process And Data Mediation, Zixin Wu, Ajith H. Ranabahu, Karthik Gomadam, Amit P. Sheth, John A. Miller

Kno.e.sis Publications

Web service composition has quickly become a key area of research in the services oriented architecture community. One of the challenges in composition is the existence of heterogeneities across independently created and autonomously managed Web service requesters and Web service providers. Previous work in this area either involved significant human effort or in cases of the efforts seeking to provide largely automated approaches, overlooked the problem of data heterogeneities, resulting in partial solutions that would not support executable workflow for real-world problems. In this paper, we present a planning-based approach to solve both the process heterogeneity and data heterogeneity problems. …