Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Numerical Analysis and Scientific Computing (671)
- Social and Behavioral Sciences (374)
- Artificial Intelligence and Robotics (360)
- Graphics and Human Computer Interfaces (313)
- Business (253)
-
- Software Engineering (239)
- Communication (234)
- Social Media (202)
- Engineering (199)
- Theory and Algorithms (178)
- Computer Engineering (171)
- Information Security (148)
- OS and Networks (116)
- Programming Languages and Compilers (98)
- E-Commerce (84)
- Data Storage Systems (75)
- Medicine and Health Sciences (70)
- Public Affairs, Public Policy and Public Administration (60)
- Education (55)
- Management Information Systems (53)
- International and Area Studies (51)
- Asian Studies (50)
- Health Information Technology (48)
- Transportation (47)
- Finance and Financial Management (43)
- Digital Communications and Networking (32)
- Technology and Innovation (32)
- Keyword
-
- Social media (59)
- Machine learning (56)
- Online learning (46)
- Deep learning (43)
- Data mining (42)
-
- Artificial intelligence (36)
- Twitter (30)
- Query processing (29)
- Classification (26)
- Neural networks (25)
- Reinforcement learning (25)
- Deep Learning (24)
- Algorithms (23)
- Clustering (21)
- Social network (21)
- Algorithm (20)
- Graph neural networks (20)
- Machine Learning (20)
- Natural language processing (20)
- Recommender systems (20)
- Semantics (20)
- Task analysis (20)
- Anomaly detection (19)
- Cloud computing (19)
- Visualization (19)
- Image retrieval (18)
- Performance (18)
- Sentiment analysis (18)
- Singapore (18)
- Social networks (17)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (3441)
- Dissertations and Theses Collection (Open Access) (58)
- Research Collection Lee Kong Chian School Of Business (11)
- Asian Management Insights (8)
- Research Collection School Of Accountancy (7)
-
- Dissertations and Theses Collection (5)
- PhD Student’s Publications Collection (5)
- Research Collection College of Integrative Studies (5)
- Research Collection Yong Pung How School Of Law (5)
- MITB Thought Leadership Series (3)
- LARC Research Publications (2)
- Perspectives@SMU (2)
- Research Collection School of Computing and Information Systems (2)
- 2024 AI for Research Week (1)
- CCX Research (1)
- Research Collection School Of Economics (1)
- Research Collection School of Accountancy (1)
- Research Collection School of Social Sciences (1)
- Research@SMU Infographics (1)
- Publication Type
Articles 3181 - 3210 of 3560
Full-Text Articles in Databases and Information Systems
Linear Correlation Discovery In Databases: A Data Mining Approach, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim
Linear Correlation Discovery In Databases: A Data Mining Approach, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Very little research in knowledge discovery has studied how to incorporate statistical methods to automate linear correlation discovery (LCD). We present an automatic LCD methodology that adopts statistical measurement functions to discover correlations from databases’ attributes. Our methodology automatically pairs attribute groups having potential linear correlations, measures the linear correlation of each pair of attribute groups, and confirms the discovered correlation. The methodology is evaluated in two sets of experiments. The results demonstrate the methodology’s ability to facilitate linear correlation discovery for databases with a large amount of data.
Applying Scenario-Based Design And Claim Analysis To The Design Of A Digital Library Of Geography Examination Resources, Yin-Leng Theng, Dion Hoe-Lian Goh, Ee Peng Lim, Zehua Liu, Ming Yin, Natalie Lee-San Pang, Patricia Bao-Bao Wong
Applying Scenario-Based Design And Claim Analysis To The Design Of A Digital Library Of Geography Examination Resources, Yin-Leng Theng, Dion Hoe-Lian Goh, Ee Peng Lim, Zehua Liu, Ming Yin, Natalie Lee-San Pang, Patricia Bao-Bao Wong
Research Collection School Of Computing and Information Systems
This paper describes the application of Carroll’s scenario-based design and claims analysis as a means of refinement to the initial design of a digital library of geographical resources (GeogDL) to prepare Singapore students to take a national examination in geography. GeogDL is built on top of G-Portal, a digital library providing services over geospatial and georeferenced Web content. Beyond improving the initial design of GeogDL, a main contribution of the paper is making explicit the use of Carroll’s strong theory-based but undercapitalized scenario-based design and claims analysis that inspired recommendations for the refinement of GeogDL. The paper concludes with an …
Ontology-Assisted Mining Of Rdf Documents, Tao Jiang, Ah-Hwee Tan
Ontology-Assisted Mining Of Rdf Documents, Tao Jiang, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Resource description framework (RDF) is becoming a popular encoding language for describing and interchanging metadata of web resources. In this paper, we propose an Apriori-based algorithm for mining association rules (AR) from RDF documents. We treat relations (RDF statements) as items in traditional AR mining to mine associations among relations. The algorithm further makes use of a domain ontology to provide generalization of relations. To obtain compact rule sets, we present a generalized pruning method for removing uninteresting rules. We illustrate a potential usage of AR mining on RDF documents for detecting patterns of terrorist activities. Experiments conducted based on …
Exploring Bit-Difference For Approximate Knn Search In High-Dimensional Databases, Bin Cui, Heng Tao Shen, Jialie Shen, Kian-Lee Tan
Exploring Bit-Difference For Approximate Knn Search In High-Dimensional Databases, Bin Cui, Heng Tao Shen, Jialie Shen, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
In this paper, we develop a novel index structure to support effcient approximate k-nearest neighbor (KNN) query in high-dimensional databases. In high-dimensional spaces, the computational cost of the distance (e.g., Euclidean distance) between two points contributes a dominant portion of the overall query response time for memory processing. To reduce the distance computation, we first propose a structure (BID) using BIt-Difference to answer approximate KNN query. The BID employs one bit to represent each feature vector of point and the number of bit-difference is used to prune the further points. To facilitate real dataset which is typically skewed, we enhance …
Values Of Mobile Commerce To Customers, Keng Siau, H. Sheng, Fiona Fui-Hoon Nah
Values Of Mobile Commerce To Customers, Keng Siau, H. Sheng, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
This research studies the values of m-commerce using a qualitative means-ends approach, called Value-Focused Thinking. The conceptual foundation for this research is the Work System Framework. By interviewing both current and potential m-commerce users, we captured the values of m-commerce and develop a means-ends objective network to illustrate the relationships among these values. As one of the first empirical research to assess the values of m-commerce, this research contributes to an increased understanding of m-commerce. The means-ends objective network also serves as a theoretical foundation for future research in m-commerce. For practitioners, our findings highlight the concerns and issues of …
High Performance P300 Speller For Brain-Computer Interface, Cuntai Guan, Manoj Thulasidas, Jiankang Wu
High Performance P300 Speller For Brain-Computer Interface, Cuntai Guan, Manoj Thulasidas, Jiankang Wu
Research Collection School Of Computing and Information Systems
P300 speller is a communication tool with which one can input texts or commands to a computer by thought. The amplitude of the P300 evoked potential is inversely proportional to the probability of infrequent or task-related stimulus. In existing P300 spellers, rows and columns of a matrix are intensified successively and randomly, resulting in a stimulus frequency of 1/N (N is the number of rows or columns of the matrix). We propose a new paradigm to display each single character randomly and individually (therefore reducing the stimulus frequency to 1/(N*N)). On-line experiments showed that this new speller significantly improved the …
Method For Identifying Individuals, Manoj Thulasidas
Method For Identifying Individuals, Manoj Thulasidas
Research Collection School Of Computing and Information Systems
A method and system for identifying a subject comprises obtaining a digitised recording of an electrocardiogram measurement of the subject to be identified, the digitised recording being a cyclic waveform having a peak amplitude. The digitised recording is normalised to reduce variations due to physiological effects, and the normalised recording is processed to determine a feature vector in the frequency domain. The distance between the determined feature vector and a predetermined feature vector is measured to identify the subject.
Justilm: Few-Shot Justification Generation For Explainable Fact-Checking Of Real-World Claims, Fengzhu Zeng, Wei Gao
Justilm: Few-Shot Justification Generation For Explainable Fact-Checking Of Real-World Claims, Fengzhu Zeng, Wei Gao
Research Collection School Of Computing and Information Systems
Justification is an explanation that supports the verdict assigned to a claim in fact-checking. However, the task of justification generation is previously oversimplified as summarization of fact-check article authored by professional checkers. In this work, we propose a realistic approach to generate justification based on retrieved evidence. We present a new benchmark dataset called ExClaim for Explainable Claim verification, and introduce JustiLM, a novel few-shot retrieval-augmented language model to learn justification generation by leveraging fact-check articles as auxiliary resource during training. Our results show that JustiLM outperforms in-context learning (ICL)-enabled LMs including Flan-T5 and Llama2, and the retrieval-augmented model Atlas …
Factors Impacting E-Government Development, Keng Siau, Y. Long
Factors Impacting E-Government Development, Keng Siau, Y. Long
Research Collection School Of Computing and Information Systems
E-government is a way for governments to use new technologies such as the Internet to provide citizens with more convenient access to government information and services, to improve the quality of services, and to provide greater opportunities for citizens to participate in democratic institutions and processes. This research investigates factors impacting e-government development. Based on the growth theory and human capital theory from the economics literature, we hypothesize that information and computer technology, and human development are two factors impacting e-government development. The hypotheses were empirically tested using secondary data from the United Nations and the United Nations Development Programme. …
The Value Of Mobile Commerce To Customers, Keng Siau, H. Sheng, Fiona Fui-Hoon Nah
The Value Of Mobile Commerce To Customers, Keng Siau, H. Sheng, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
This research studies the values of m-commerce using a qualitative means-ends approach, called Value-Focused Thinking. The conceptual foundation for this research is the Work System Framework. By interviewing both current and potential m-commerce users, we captured the values of m-commerce and develop a means-ends objective network to illustrate the relationships among these values. As one of the first empirical research to assess the values of m-commerce, this research contributes to an increased understanding of mcommerce. The means-ends objective network also serves as a theoretical foundation for future research in mcommerce. For practitioners, our findings highlight the concerns and issues of …
Theoretical And Practical Complexity Of Unified Modeling Language: Delphi Study And Metrics Analyses, J. Erickson, Keng Siau
Theoretical And Practical Complexity Of Unified Modeling Language: Delphi Study And Metrics Analyses, J. Erickson, Keng Siau
Research Collection School Of Computing and Information Systems
Systems have become increasingly complex, and as a result development methods have become more complex as well. The unified modeling language (UML) has been criticized for the often cited and sometimes over- whelming complexity it presents to its users, and those seeking to learn to use it. Using Rossi and Brinkkemper’s (1996) complexity metrics, Siau and Cao (2001) completed a complexity analysis of UML and 36 other modeling techniques, finding that UML is indeed more complex than other techniques. Siau, Erickson and Lee (2002) proposed that Rossi and Brinkkemper’s metrics present the theoretical maximum complexity, known as theoretical complexity. This …
Supporting Field Study With Personalized Project Spaces In A Geographical Digital Library, Ee Peng Lim, Aixin Sun, Zehua Liu, John Hedberg, Chew-Hung Chang, Tiong-Sa Teh, Dion Hoe-Lian Goh, Yin-Leng Theng
Supporting Field Study With Personalized Project Spaces In A Geographical Digital Library, Ee Peng Lim, Aixin Sun, Zehua Liu, John Hedberg, Chew-Hung Chang, Tiong-Sa Teh, Dion Hoe-Lian Goh, Yin-Leng Theng
Research Collection School Of Computing and Information Systems
Digital libraries have been rather successful in supporting learning activities by providing learners with access to information and knowledge. However, this level of support is passive to learners and interactive and collaborative learning cannot be easily achieved. In this paper, we study how digital libraries could be extended to serve a more active role in collaborative learning activities. We focus on developing new services to support a common type of learning activity, field study, in a geospatial context. We propose the concept of personal project space that allows individuals to work in their personalized environment with a mix of private …
Design Lessons On Access Features In Paper, Yin-Leng Theng, Dion Hoe-Lian Goh, Ming Yin, Eng-Kai Suen, Ee Peng Lim
Design Lessons On Access Features In Paper, Yin-Leng Theng, Dion Hoe-Lian Goh, Ming Yin, Eng-Kai Suen, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Using Nielsen's Heuristic Evaluation, this paper reports a user study with six usability-trained subjects to evaluate PAPER's access features in assisting users to retrieve information efficiently, part of an on-going design partnership with stakeholders and designers/developers. PAPER (Personalised Adaptive Pathways for Exam Resources) is an improved version evolving from an earlier implementation of GeogDL built upon G-Portal, a geospatial digital library infrastructure. After two initial evaluations with student and teacher design partners, PAPER has evolved into a system containing a new bundle of personalized, interactive services with four modules: mock exam; personal coach (practice and review); trend analysis and performance …
U-Commerce: Emerging Trends And Research Issues, H. Galanxhi-Janaqi, Fiona Fui-Hoon Nah
U-Commerce: Emerging Trends And Research Issues, H. Galanxhi-Janaqi, Fiona Fui-Hoon Nah
Research Collection School Of Computing and Information Systems
Ubiquitous commerce or u-commerce is the combination of traditional e-commerce and wireless, television, voice and silent commerce. U-commerce implies ubiquity, universality, uniqueness and unison. It is not a replacement for other types of commerce, but an extension of them. While bringing many benefits, there are challenges and impediments to overcome. Research is needed to assess the value of u-commerce and to address its related issues and challenges. Questions that need to be addressed are: What is the value of u-commerce? What are the ways to maximize the benefits and value of u-commerce? Is it the right technology and what directions …
Accommodating Instance Heterogeneities In Database Integration, Ee Peng Lim, Roger Hsiang-Li Chiang
Accommodating Instance Heterogeneities In Database Integration, Ee Peng Lim, Roger Hsiang-Li Chiang
Research Collection School Of Computing and Information Systems
A complete data integration solution can be viewed as an iterative process that consists of three phases, namely analysis, derivation and evolution. The entire process is similar to a software development process with the target application being the derivation rules for the integrated databases. In many cases, data integration requires several iterations of refining the local-to-global database mapping rules before a stable set of rules can be obtained. In particular, the mapping rules, as well as the data model and query model for the integrated databases have to cope with poor data quality in local databases, ongoing local database updates …
On Semantic Caching And Query Scheduling For Mobile Nearest-Neighbor Search, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
On Semantic Caching And Query Scheduling For Mobile Nearest-Neighbor Search, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Location-based services have received increasing attention in recent years. In this paper, we address the performance issues of mobile nearest-neighbor search, in which the mobile user issues a query to retrieve stationary service objects nearest to him/her. An index based on Voronoi Diagram is used in the server to support such a search, while a semantic cache is proposed to enhance the access efficiency of the service. Cache replacement policies tailored for the proposed semantic cache are examined. Moreover, several query scheduling policies are proposed to address the inter-cell roaming issues in multi-cell environments. Simulations are conducted to evaluate the …
The D-Tree: An Index Structure For Planar Point Queries Location-Based Wireless Services, Jianliang Xu, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
The D-Tree: An Index Structure For Planar Point Queries Location-Based Wireless Services, Jianliang Xu, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Location-based services (LBSs), considered as a killer application in the wireless data market, provide information based on locations specified in the queries. In this paper, we examine the indexing issue for querying location-dependent data in wireless LBSs; in particular, we focus on an important class of queries, planar point queries. To address the issues of responsiveness, energy consumption, and bandwidth contention in wireless communications, an index has to minimize the search time and maintain a small storage overhead. It is shown that the traditional point-location algorithms and spatial index structures fail to achieve either objective or both. This paper proposes …
Spatial Queries In Wireless Broadcast Systems, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Spatial Queries In Wireless Broadcast Systems, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Owing to the advent of wireless networking and personal digital devices, information systems in the era of mobile computing are expected to be able to handle a tremendous amount of traffic and service requests from the users. Wireless data broadcast, thanks to its high scalability, is particularly suitable for meeting such a challenge. Indexing techniques have been developed for wireless data broadcast systems in order to conserve the scarce power resources in mobile clients. However, most of the previous studies do not take into account the impact of location information of users. In this paper, we address the issues of …
Finding Constrained Frequent Episodes Using Minimal Occurrences, Xi Ma, Hwee Hwa Pang, Kian-Lee Tan
Finding Constrained Frequent Episodes Using Minimal Occurrences, Xi Ma, Hwee Hwa Pang, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
Recurrent combinations of events within an event sequence, known as episodes, often reveal useful information. Most of the proposed episode mining algorithms adopt an apriori-like approach that generates candidates and then calculates their support levels. Obviously, such an approach is computationally expensive. Moreover, those algorithms are capable of handling only a limited range of constraints. In this paper, we introduce two mining algorithms - episode prefix tree (EPT) and position pairs set (PPS) - based on a prefix-growth approach to overcome the above limitations. Both algorithms push constraints systematically into the mining process. Performance study shows that the proposed algorithms …
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
One common approach in hierarchical text classification involves associating classifiers with nodes in the category tree and classifying text documents in a top-down manner. Classification methods using this top-down approach can scale well and cope with changes to the category trees. However, all these methods suffer from blocking which refers to documents wrongly rejected by the classifiers at higher-levels and cannot be passed to the classifiers at lower-levels. We propose a classifier-centric performance measure known as blocking factor to determine the extent of the blocking. Three methods are proposed to address the blocking problem, namely, threshold reduction, restricted voting, and …
Improving Transliteration With Precise Alignment Of Phoneme Chunks And Using Contextual Features, Wei Gao, Kam-Fai Wong, Wai Lam
Improving Transliteration With Precise Alignment Of Phoneme Chunks And Using Contextual Features, Wei Gao, Kam-Fai Wong, Wai Lam
Research Collection School Of Computing and Information Systems
Automatic transliteration of foreign names is basically regarded as a diminutive clone of the machine translation (MT) problem. It thus follows IBM’s conventional MT models under the sourcechannel framework. Nonetheless, some parameters of this model dealing with zero-fertility words in the target sequences, can negatively impact transliteration effectiveness because of the inevitable inverted conditional probability estimation. Instead of source-channel, this paper presents a direct probabilistic transliteration model using contextual features of phonemes with a tailored alignment scheme for phoneme chunks. Experiments demonstrate superior performance over the source-channel for the task of English-Chinese transliteration.
A Novel Log-Based Relevance Feedback Technique In Content-Based Image Retrieval, Steven C. H. Hoi, Michael R. Lyu
A Novel Log-Based Relevance Feedback Technique In Content-Based Image Retrieval, Steven C. H. Hoi, Michael R. Lyu
Research Collection School Of Computing and Information Systems
Relevance feedback has been proposed as an important technique to boost the retrieval performance in content-based image retrieval (CBIR). However, since there exists a semantic gap between low-level features and high-level semantic concepts in CBIR, typical relevance feedback techniques need to perform a lot of rounds of feedback for achieving satisfactory results. These procedures are time-consuming and may make the users bored in the retrieval tasks. For a long-term study purpose in CBIR, we notice that the users' feedback logs can be available and employed for helping the retrieval tasks in CBIR systems. In this paper, we propose a novel …
A Survey Of Online E-Banking Retail Initiatives, P. Southard, Keng Siau
A Survey Of Online E-Banking Retail Initiatives, P. Southard, Keng Siau
Research Collection School Of Computing and Information Systems
Customer demand is forcing banks to provide their services online. There are two successful paths they can take: to grow, or to specialize in providing localized services and information.
Clip-Based Similarity Measure For Hierarchical Video Retrieval, Yuxin Peng, Chong-Wah Ngo
Clip-Based Similarity Measure For Hierarchical Video Retrieval, Yuxin Peng, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
This paper proposes a new approach and algorithm for the similarity measure of video clips. The similarity is mainly based on two bipartite graph matching algorithms: maximum matching (MM) and optimal matching (OM). MM is able to rapidly filter irrelevant video clips, while OM is capable of ranking the similarity of clips according to the visual and granularity factors. Based on MM and OM, a hierarchical video retrieval framework is constructed for the approximate matching of video clips. To allow the matching between a query and a long video, an online clip segmentation algorithm is also proposed to rapidly locate …
Are Use Case And Class Diagrams Complementary In Requirements Analysis? An Experimental Study On Use Case And Class Diagrams In Uml, Keng Siau, Lihyunn Lee
Are Use Case And Class Diagrams Complementary In Requirements Analysis? An Experimental Study On Use Case And Class Diagrams In Uml, Keng Siau, Lihyunn Lee
Research Collection School Of Computing and Information Systems
Despite the status of united modeling language (UML) as the de facto standard for object oriented modeling, it has received controversial reviews. The most controversial diagram in UML is the use case diagram. Some practitioners claim that use case diagrams are not valuable in requirements analysis and some have even argued that use case diagrams should not be part of UML. This research examined the values of use case diagram in interpreting requirements when use case diagrams are used in conjunction with class diagrams. In other words, the study investigated the possible synergetic values and relationships between the use case …
Shared-Storage Auction Ensures Data Availability, Hady W. Lauw, Siu-Cheung Hui, Edmund M. K. Lai
Shared-Storage Auction Ensures Data Availability, Hady W. Lauw, Siu-Cheung Hui, Edmund M. K. Lai
Research Collection School Of Computing and Information Systems
Most current e-auction systems are based on the client-server architecture. Such centralized systems provide a single point of failure and control. In contrast, peer-to-peer systems permit distributed control and minimize individual node and link failures' impact on the system. The shared-storage-based auction model described decentralizes services among peers to share the required processing load and aggregates peers' resources for common use. The model is based on the principles of local computation at each peer, direct inter-peer communication, and a shared storage space.
A Spectroscopy Of Texts For Effective Clustering, Wenyuan Li, Wee-Keong Ng, Kok-Leong Ong, Ee Peng Lim
A Spectroscopy Of Texts For Effective Clustering, Wenyuan Li, Wee-Keong Ng, Kok-Leong Ong, Ee Peng Lim
Research Collection School Of Computing and Information Systems
For many clustering algorithms, such as k-means, EM, and CLOPE, there is usually a requirement to set some parameters. Often, these parameters directly or indirectly control the number of clusters to return. In the presence of different data characteristics and analysis contexts, it is often difficult for the user to estimate the number of clusters in the data set. This is especially true in text collections such as Web documents, images or biological data. The fundamental question this paper addresses is: ldquoHow can we effectively estimate the natural number of clusters in a given text collection?rdquo. We propose to use …
Robust Classification Of Event-Related Potential For Brain-Computer Interface, Manoj Thulasidas
Robust Classification Of Event-Related Potential For Brain-Computer Interface, Manoj Thulasidas
Research Collection School Of Computing and Information Systems
We report the implementation of a text input application (speller) based on the P300 event related potential. We obtain high accuracies by using an SVM classifier and a novel feature. These techniques enable us to maintain fast performance without sacrificing the accuracy, thus making the speller usable in an online mode. In order to further improve the usability, we perform various studies on the data with a view to minimizing the training time required. We present data collected from nine healthy subjects, along with the high accuracies (of the order of 95% or more) measured online. We show that the …
Sclope: An Algorithm For Clustering Data Streams Of Categorical Attributes, Kok-Leong Ong, Wenyuan Li, Wee-Keong Ng, Ee Peng Lim
Sclope: An Algorithm For Clustering Data Streams Of Categorical Attributes, Kok-Leong Ong, Wenyuan Li, Wee-Keong Ng, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Clustering is a difficult problem especially when we consider the task in the context of a data stream of categorical attributes. In this paper, we propose SCLOPE, a novel algorithm based on CLOPE's intuitive observation about cluster histograms. Unlike CLOPE however, our algorithm is very fast and operates within the constraints of a data stream environment. In particular, we designed SCLOPE according to the recent CluStream framework. Our evaluation of SCLOPE shows very promising results. It consistently outperforms CLOPE in speed and scalability tests on our data sets while maintaining high cluster purity; it also supports cluster analysis that other …
Towards Personalised Web Intelligence, Ah-Hwee Tan, Hwee-Leng Ong, Hong Pan, Jamie Ng, Qiu-Xiang Li
Towards Personalised Web Intelligence, Ah-Hwee Tan, Hwee-Leng Ong, Hong Pan, Jamie Ng, Qiu-Xiang Li
Research Collection School Of Computing and Information Systems
The Flexible Organizer for Competitive Intelligence (FOCI) is a personalised web intelligence system that provides an integrated platform for gathering, organising, tracking, and disseminating competitive information on the web. FOCI builds personalised information portfolios through a novel method called User-Configurable Clustering, which allows a user to personalise his/her portfolios in terms of the content as well as the organisational structure. This paper outlines the key challenges we face in personalised information management and gives a detailed account of FOCI’s underlying personalisation mechanism. For a quantitative evaluation of the system’s performance, we propose a set of performance indices based on information …