Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 751 - 780 of 809

Full-Text Articles in Numerical Analysis and Scientific Computing

Cache Invalidation And Replacement Strategies For Location-Dependent Data In Mobile Environments, Baihua Zheng, Jianliang Xu, Dik Lun Lee Oct 2002

Cache Invalidation And Replacement Strategies For Location-Dependent Data In Mobile Environments, Baihua Zheng, Jianliang Xu, Dik Lun Lee

Research Collection School Of Computing and Information Systems

Mobile location-dependent information services (LDISs) have become increasingly popular in recent years. However, data caching strategies for LDISs have thus far received little attention. In this paper, we study the issues of cache invalidation and cache replacement for location-dependent data under a geometric location model. We introduce a new performance criterion, called caching efficiency, and propose a generic method for location-dependent cache invalidation strategies. In addition, two cache replacement policies, PA and PAID, are proposed. Unlike the conventional replacement policies, PA and PAID take into consideration the valid scope area of a data value. We conduct a series of simulation …


Personalized Classification For Keyword-Based Category Profiles, Aixin Sun, Ee Peng Lim, Wee-Keong Ng Sep 2002

Personalized Classification For Keyword-Based Category Profiles, Aixin Sun, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

Personalized classification refers to allowing users to define their own categories and automating the assignment of documents to these categories. In this paper, we examine the use of keywords to define personalized categories and propose the use of Support Vector Machine (SVM) to perform personalized classification. Two scenarios have been investigated. The first assumes that the personalized categories are defined in a flat category space. The second assumes that each personalized category is defined within a pre-defined general category that provides a more specific context for the personalized category. The training documents for personalized categories are obtained from a training …


Hcl: A Specification Language For Hierarchical Text Classification, Aixin Sun, Ee Peng Lim, Wee-Keong Ng Aug 2002

Hcl: A Specification Language For Hierarchical Text Classification, Aixin Sun, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

Hierarchical text classification refers to assigning text documents to the categories in a given category tree based on their content. With large number of categories organized as a tree, hierarchical text classification helps users to find information more quickly and accurately. Nevertheless, hierarchical text classification methods in the past have often been constructed in a proprietary manner. The construction steps often involve human efforts and are not completely automated. In this paper, we therefore propose a specification language known as HCL (Hierarchical Classification Language) . HCL is designed to describe a hierarchical classification method including the definition of a category …


Fast Filter-And-Refine Algorithms For Subsequence Selection, Beng-Chin Ooi, Hwee Hwa Pang, Hao Wang, Limsoon Wong, Cui Yu Jul 2002

Fast Filter-And-Refine Algorithms For Subsequence Selection, Beng-Chin Ooi, Hwee Hwa Pang, Hao Wang, Limsoon Wong, Cui Yu

Research Collection School Of Computing and Information Systems

Large sequence databases, such as protein, DNA and gene sequences in biology, are becoming increasingly common. An important operation on a sequence database is approximate subsequence matching, where all subsequences that are within some distance from a given query string are retrieved. This paper proposes a filter-and-refine algorithm that enables efficient approximate subsequence matching in large DNA sequence databases. It employs a bitmap indexing structure to condense and encode each data sequence into a shorter index sequence. During query processing, the bitmap index is used to filter out most of the irrelevant subsequences, and false positives are removed in the …


Digital Libraries To Knowledge Portals: Towards A Global Knowledge Portal For Secondary Schools In Singapore, Yin-Leng Theng, Dion Hoe-Lian Goh, Chu Keong Lee, Ee Peng Lim, Zehua Liu Jul 2002

Digital Libraries To Knowledge Portals: Towards A Global Knowledge Portal For Secondary Schools In Singapore, Yin-Leng Theng, Dion Hoe-Lian Goh, Chu Keong Lee, Ee Peng Lim, Zehua Liu

Research Collection School Of Computing and Information Systems

For digital libraries to remain relevant in the new millennium where the ability to manage knowledge is critical, this paper explores how digital libraries could strategically be evolved into knowledge portals to encapsulate knowledge creation, management, sharing and reusability, features evidently lacking in most conventional digital libraries. Two digital library scenarios of use in education are described and implemented as knowledge portals using G-Portal and the Greenstone software. We hope that the initial work carried out on these two portal-like DLs will eventually form part of a Global Knowledge Portal for Secondary Schools in Singapore. Keywords Digital libraries, information portals, …


Product Schema Integration For Electronic Commerce: A Synonym Comparison Approach, Guanghao Yan, Wee-Keong Ng, Ee Peng Lim May 2002

Product Schema Integration For Electronic Commerce: A Synonym Comparison Approach, Guanghao Yan, Wee-Keong Ng, Ee Peng Lim

Research Collection School Of Computing and Information Systems

In any electronic commerce system, the heterogeneity of product descriptions is a critical impediment to efficient business information exchange. In the ABECOS electronic commerce system, buyer agents, seller agents, and directory agents liaise with one another in e-commerce activities. Only when agents have a common ontology of product descriptions (also called product schemas) are they able to interact seamlessly in e-commerce activities. This gives rise to the product schema integration problem (PSI); the problem of integrating heterogeneous schemas of a certain product into one globally compatible schema. We adopt an integration approach based on product attribute synonyms. We give a …


A Case For Analytical Customer Relationship Management, Jaideep Srivastava, Jau-Hwang Wang, Ee Peng Lim, San-Yih Hwang May 2002

A Case For Analytical Customer Relationship Management, Jaideep Srivastava, Jau-Hwang Wang, Ee Peng Lim, San-Yih Hwang

Research Collection School Of Computing and Information Systems

This paper describes how data analytics can be used to make various CRM functions like customer segmentation, communication targeting, retention, and loyalty much more effective. Also briefly describe the key technologies needed to implement analytical CRM, and are the organizational issues that must be carefully handled to make CRM a reality.


Mining Relationship Graphs For Effective Business Objectives, Kok-Leong Ong, Ee Peng Lim, Wee-Keong Ng May 2002

Mining Relationship Graphs For Effective Business Objectives, Kok-Leong Ong, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

Modern organization has two types of customer profiles: active and passive. Active customers contribute to the business goals of an organization, while passive customers are potential candidates that can be converted to active ones. Existing KDD techniques focused mainly on past data generated by active customers. The insights discovered apply well to active ones but may scale poorly with passive customers. This is because there is no attempt to generate know-how to convert passive customers into active ones. We propose an algorithm to discover relationship graphs using both types of profile. Using relationship graphs, an organization can be more effective …


An Intelligent Middleware For Linear Correlation Discovery, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim Mar 2002

An Intelligent Middleware For Linear Correlation Discovery, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Although it is widely accepted that research from data mining, knowledge discovery, and data warehousing should be synthesized, little research addresses the integration of existing data management and analysis software. We develop an intelligent middleware that facilitates linear correlation discovery, the discovery of associations between attributes and attribute groups. This middleware integrates data management and data analysis tools to improve traditional data analysis in three perspectives: (1) identify appropriate linear correlation functions to perform based on the semantics of a data set; (2) execute appropriate functions contained in the data analysis packages; and (3) derive useful knowledge from data analysis.


Hierarchical Text Classification And Evaluation, Aixin Sun, Ee Peng Lim Nov 2001

Hierarchical Text Classification And Evaluation, Aixin Sun, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Hierarchical Classification refers to assigning of one or more suitable categories from a hierarchical category space to a document. While previous work in hierarchical classification focused on virtual category trees where documents are assigned only to the leaf categories, we propose atop-down level-based classification method that can classify documents to both leaf and internal categories. As the standard performance measures assume independence between categories, they have not considered the documents incorrectly classified into categories that are similar or not far from the correct ones in the category tree. We therefore propose the Category-Similarity Measures and Distance-Based Measures to consider the …


A Review Of Data Mining Techniques, Sang Jun Lee, Keng Siau Oct 2001

A Review Of Data Mining Techniques, Sang Jun Lee, Keng Siau

Research Collection School Of Computing and Information Systems

Terabytes of data are generated everyday in many organizations. To extract hidden predictive information from large volumes of data, data mining (DM) techniques are needed. Organizations are starting to realize the importance of data mining in their strategic planning and successful application of DM techniques can be an enormous payoff for the organizations. This paper discusses the requirements and challenges of DM, and describes major DM techniques such as statistics, artificial intelligence, decision tree approach, genetic algorithm, and visualization.


Mining Multi-Level Rules With Recurrent Items Using Fp'-Tree, Kok-Leong Ong, Wee-Keong Ng, Ee Peng Lim Oct 2001

Mining Multi-Level Rules With Recurrent Items Using Fp'-Tree, Kok-Leong Ong, Wee-Keong Ng, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Association rule mining has received broad research in the academic and wide application in the real world. As a result, many variations exist and one such variant is the mining of multi-level rules. The mining of multi-level rules has proved to be useful in discovering important knowledge that conventional algorithms such as Apriori, SETM, DIC etc., miss. However, existing techniques for mining multi-level rules have failed to take into account the recurrence relationship that can occur in a transaction during the translation of an atomic item to a higher level representation. As a result, rules containing recurrent items go unnoticed. …


Vide: A Visual Data Extraction Environment For The Web, Yi Li, Wee-Keong Ng, Ee Peng Lim Sep 2001

Vide: A Visual Data Extraction Environment For The Web, Yi Li, Wee-Keong Ng, Ee Peng Lim

Research Collection School Of Computing and Information Systems

With the rapid growth of information on the Web, a means to combat information overload is critical. In this paper, we present ViDE (Visual Data Extraction), an interactive web data extraction environment that supports efficient hierarchical data wrapping of multiple web pages. ViDE has two unique features that differentiate it from other extraction mechanisms. First, data extraction rules can be easily specified in a graphical user interface that is seamlessly integrated with a web browser. Second, ViDE introduces the concept of grouping which unites the extraction rules for a set of documents with the navigational patterns that exist among them. …


Predictive Self-Organizing Networks For Text Categorization, Ah-Hwee Tan Apr 2001

Predictive Self-Organizing Networks For Text Categorization, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

This paper introduces a class of predictive self-organizing neural networks known as Adaptive Resonance Associative Map (ARAM) for classification of free-text documents. Whereas most sta- tistical approaches to text categorization derive classification knowledge based on training examples alone, ARAM performs supervised learn- ing and integrates user-defined classification knowledge in the form of IF-THEN rules. Through our experiments on the Reuters-21578 news database, we showed that ARAM performed reasonably well in mining categorization knowledge from sparse and high dimensional document feature space. In addition, ARAM predictive accuracy and learning efficiency can be improved by incorporating a set of rules derived from …


The Partial Evaluation Approach To Information Personalization, Naren Ramakrishnan, Saverio Perugini Jan 2001

The Partial Evaluation Approach To Information Personalization, Naren Ramakrishnan, Saverio Perugini

Computer Science Faculty Publications

Information personalization refers to the automatic adjustment of information content, structure, and presentation tailored to an individual user. By reducing information overload and customizing information access, personalization systems have emerged as an important segment of the Internet economy. This paper presents a systematic modeling methodology— PIPE (‘Personalization is Partial Evaluation’) — for personalization. Personalization systems are designed and implemented in PIPE by modeling an information-seeking interaction in a programmatic representation. The representation supports the description of information-seeking activities as partial information and their subsequent realization by partial evaluation, a technique for specializing programs. We describe the modeling methodology at a …


Incorporating Window-Based Passage-Level Evidence In Document Retrieval, Wensi Xi, Richard Xu-Rong, Christopher Soo Guan Khoo, Ee Peng Lim Jan 2001

Incorporating Window-Based Passage-Level Evidence In Document Retrieval, Wensi Xi, Richard Xu-Rong, Christopher Soo Guan Khoo, Ee Peng Lim

Research Collection School Of Computing and Information Systems

This study investigated whether document retrieval can be improved if documents are divided into smaller sub-documents or passages and the retrieval score for these passages are incorporated in the final retrieval score for the whole document. The documents were segmented by sliding a window of a certain size across the document and extracting the words displayed each time the window stopped. A retrieval score was calculated for each of the passages extracted and the highest score obtained by a passage of that size was taken as the document’s passage-level score for that window size. A range of window sizes was …


A Meta-Analysis On Relationship Modeling Accuracy: Comparing Relational And Semantic Models, Qing Cao, Fiona Fui-Hoon Nah, Keng Siau Aug 2000

A Meta-Analysis On Relationship Modeling Accuracy: Comparing Relational And Semantic Models, Qing Cao, Fiona Fui-Hoon Nah, Keng Siau

Research Collection School Of Computing and Information Systems

Semantic data modeling, such as entity-relationship (ER) modeling and extended/enhanced entity-relationship (EER) modeling, has emerged as an alternative to relational data modeling. The majority of research in data modeling suggests that the use of semantic data models leads to better performance. However the findings are not conclusive and sometimes inconsistent. In this research, we investigate modeling relationship correctness in relational and semantic models. The meta-analysis carried out in this research is an attempt to alleviate inconsistent results in previous studies.


Dtd-Miner: A Tool For Mining Dtds From Xml Documents, Moh Chuang Hue, Ee Peng Lim, Wee-Keong Ng Jun 2000

Dtd-Miner: A Tool For Mining Dtds From Xml Documents, Moh Chuang Hue, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

XML documents are semistructured and the structure of the documents is embedded in the tags. Although XML documents can be accompanied by a DTD that defines the structure of the documents, the presence of a DTD is not mandatory. The difficulty in deriving the DTD for XML documents lies in the fact that DTDs are of different syntax as XML and that prior knowledge of the structure of the documents is required. In this paper, we introduce DTD-Miner, an automatic structure-mining tool for XML documents. Using a Web-based interface, the user will be able to submit a set of similarly …


Re-Engineering Structures From Web Documents, Moh Chuang Hue, Ee Peng Lim, Wee-Keong Ng Jun 2000

Re-Engineering Structures From Web Documents, Moh Chuang Hue, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

To realize a wide range of applications (including digital libraries) on the Web, a more structured way of accessing the Web is required and such requirement can be facilitated by the use of XML standard. In this paper, we propose a general framework for reverse engineering (or re-engineering) the underlying structures i.e.,the DTD from a collection of similarly structured XML documents when they share some common but unknown DTDs. The essential data structures and algorithms for the DTD generation have been delveloped and experiments on real Web collections have been conducted to demonstrate their feasibilty. In addition, we also proposed …


Load Sharing In Distributed Multimedia-On-Demand Systems, Y. C. Tay, Hwee Hwa Pang May 2000

Load Sharing In Distributed Multimedia-On-Demand Systems, Y. C. Tay, Hwee Hwa Pang

Research Collection School Of Computing and Information Systems

Service providers have begun to offer multimedia-on-demand services to residential estates by installing isolated, small-scale multimedia servers at individual estates. Such an arrangement allows the service providers to operate without relying on a highspeed, large-capacity metropolitan area network, which is still not available in many countries. Unfortunately, installing isolated servers can incur very high server costs, as each server requires spare bandwidth to cope with fluctuations in user demand. The authors explore the feasibility of linking up several small multimedia servers to a (limited-capacity) network, and allowing servers with idle retrieval bandwidth to help out servers that are temporarily overloaded; …


Predictive Adaptive Resonance Theory And Knowledge Discovery In Databases, Ah-Hwee Tan, Hui-Shin Vivien Soon May 2000

Predictive Adaptive Resonance Theory And Knowledge Discovery In Databases, Ah-Hwee Tan, Hui-Shin Vivien Soon

Research Collection School Of Computing and Information Systems

This paper investigates the scalability of predictive Adaptive Resonance Theory (ART) networks for knowledge discovery in very large databases. Although predictive ART performs fast and incremental learning, the number of recognition categories or rules that it creates during learning may become substantially large and cause the learning speed to slow down. To tackle this problem, we introduce an on-line algorithm for evaluating and pruning categories during learning. Benchmark experiments on a large scale data set show that on-line pruning has been effective in reducing the number of the recognition categories and the time for convergence. Interestingly, the pruned networks also …


Personalizing The Gams Cross-Index, Saverio Perugini, Priya Lakshminarayanan, Naren Ramakrishnan Jan 2000

Personalizing The Gams Cross-Index, Saverio Perugini, Priya Lakshminarayanan, Naren Ramakrishnan

Computer Science Faculty Publications

The NIST Guide to Available Mathematical Software (GAMS) system at http://gams.nist .gov serves as the gateway to thousands of scientific codes and modules for numerical computation. We describe the PIPE personalization facility for GAMS, whereby content from the cross-index is specialized for a user desiring software recommendations for a specific problem instance. The key idea is to (i) mine structure, and (ii) exploit it in a programmatic manner to generate personalized web pages. Our approach supports both content-based and collaborative personalization and enables information integration from multiple (and complementary) web resources. We present case studies for the domain of linear, …


An Integrated Data Mining System To Automate Discovery, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim Jan 2000

An Integrated Data Mining System To Automate Discovery, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim

Research Collection School Of Computing and Information Systems

Many data analysts require tools which can integrate their database management packages (e.g. Microsoft Access) with their data analysis ones (e.g. SAS, SPSS), and provide guidance for the selection of appropriate mining algorithms. In addition, the analysts need to extract and validate statistical results to facilitate data mining. In this paper, we describe an integrated data mining system called the Linear Correlation Discovery System (LCDS) that meets the above requirement. LCDS consists of four major sub-components, two of which, the selection assistant and the statistics coupler, are discussed in this paper. The former examines the schema and instances to determine …


Zbroker: A Query Routing Broker For Z39.50 Databases, Yong Lin, Jian Xu, Ee Peng Lim, Wee-Keong Ng Nov 1999

Zbroker: A Query Routing Broker For Z39.50 Databases, Yong Lin, Jian Xu, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

A query routing broker is a software agent that determines from a large set of accessing information sources the ones most relevant to a user's information need. As the number of information sources on the Internet increases dramatically, future users will have to rely on query routing brokers to decide a small number of information sources to query without incurring too much query processing overheads. In this paper, we describe a query routing broker known as ZBroker developed for bibliographic database servers that support the Z39.50 protocol. ZBroker samples the content of each bibliographic database by using training queries and …


Wedagen: A Synthetic Web Database Generator, Pallavi Priyardarshini, Fengqiong Qin, Ee Peng Lim, Wee-Keong Ng Sep 1999

Wedagen: A Synthetic Web Database Generator, Pallavi Priyardarshini, Fengqiong Qin, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

At the Centre for Advanced Information Systems (CAIS), a Web warehousing system is being developed to store and manipulate Web information. The system named WHOWEDA (WareHouse Of WEb DAta) stores extracted Web information as Web tables and provides several Web operators, eg. Web join, Web select, global coupling, etc., to manipulate Web tables. During the implementation of WHOWEDA, it is necessary to perform systematic testing on the system and to evaluate its system performance. While it is possible for WHOWEDA to be tested or evaluated using actual Web pages downloaded from WWW, the amount of time required for such testing …


Locating Web Information Using Web Checkpoints, Aik Kee Luah, Wee-Keong Ng, Ee Peng Lim, Wee Peng Lee, Yinyan Cao Sep 1999

Locating Web Information Using Web Checkpoints, Aik Kee Luah, Wee-Keong Ng, Ee Peng Lim, Wee Peng Lee, Yinyan Cao

Research Collection School Of Computing and Information Systems

Conventional search engines locate information by letting users establish a single web checkpoint1. By specifying one or more keywords, users direct search engines to return a set of documents that contain those keywords. From the documents (links) returned by search engines, user proceed to further probe the WWW from there. Hence, these initial set of documents (contingent upon the occurrence of keyword(s)) serve as a web checkpoint. Generally, these links are numerous and may not result in much fruitful searches. By establishing multiple web checkpoints, a richer and controllable search procedure can be constructed to obtain more relevant Web information. …


Non-Repudiation In An Agent-Based E-Commerce System, Chin Chuan Liew, Wee-Keong Ng, Ee Peng Lim, Beng Suang Tan, Kok-Leong Ong Sep 1999

Non-Repudiation In An Agent-Based E-Commerce System, Chin Chuan Liew, Wee-Keong Ng, Ee Peng Lim, Beng Suang Tan, Kok-Leong Ong

Research Collection School Of Computing and Information Systems

Abecos is an agent-based e-commerce system under development at the Nanyang Technological University. A key factor in making this system usable in practice is strict security controls. One aspect of security is the provision of non-repudiation services. As protocols for non-repudiation have focused on -message non-repudiation, its adaptation to afford non-repudiation in a communication session for two agents in Abecos is inefficient. In this work, we investigate and propose a protocol for enforcing non-repudiation in a session. The protocol is believed to be applicable in any e-commerce system; agent- or not agent-based.


Cluster-Based Database Selection Techniques For Routing Bibliographic Queries, Jian Xu, Ee Peng Lim, Wee-Keong Ng Aug 1999

Cluster-Based Database Selection Techniques For Routing Bibliographic Queries, Jian Xu, Ee Peng Lim, Wee-Keong Ng

Research Collection School Of Computing and Information Systems

In this paper, we focus on the database selection problem in the context of a global digital library consisting of a large number of online bibliographic servers. Each server hosts a bibliographic database that contains bibliographic records each of which consists of text values for a number of pre-defined bibliographic attributes such as title, author, call number, subject, etc.. Each bibliographic database supports user queries on the bibliographic attributes. Since bibliographic records are relatively small in size, we only consider bibliographic databases that support boolean queries on the bibliographic attributes, and the query results are not ranked. While focusing on …


Resource Scheduling In A High-Performance Multimedia Server, Hwee Hwa Pang, Bobby Jose, M. S. Krishnan Mar 1999

Resource Scheduling In A High-Performance Multimedia Server, Hwee Hwa Pang, Bobby Jose, M. S. Krishnan

Research Collection School Of Computing and Information Systems

Supporting continuous media data-such as video and audio-imposes stringent demands on the retrieval performance of a multimedia server. In this paper, we propose and evaluate a set of data placement and retrieval algorithms to exploit the full capacity of the disks in a multimedia server. The data placement algorithm declusters every object over all of the disks in the server-using a time-based declustering unit-with the aim of balancing the disk load. As for runtime retrieval, the quintessence of the algorithm is to give each disk advance notification of the blocks that have to be fetched in the impending time periods, …


Harp: A Distributed Query System For Legacy Public Libraries And Structured Databases, Ee Peng Lim, Ying Lu Jan 1999

Harp: A Distributed Query System For Legacy Public Libraries And Structured Databases, Ee Peng Lim, Ying Lu

Research Collection School Of Computing and Information Systems

The main purpose of a digital library is to facilitate users easy access to enormous amount of globally networked information. Typically, this information includes preexisting public library catalog data, digitized document collections, and other databases. In this article, we describe the distributed query system of a digital library prototype system known as HARP. In the HARP project, we have designed and implemented a distributed query processor and its query front-end to support integrated queries to preexisting public library catalogs and structured databases. This article describes our experiences in the design of an extended Sequel (SQL) query language known as HarpSQL. …