Open Access. Powered by Scholars. Published by Universities.®
Numerical Analysis and Scientific Computing Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (671)
- University of Dayton (31)
- Central Bank of Nigeria (20)
- University of Arkansas, Fayetteville (7)
- University of Nebraska - Lincoln (6)
-
- LSU New Orleans (4)
- California Polytechnic State University, San Luis Obispo (3)
- City University of New York (CUNY) (3)
- Montclair State University (3)
- Purdue University (3)
- San Jose State University (3)
- Technological University Dublin (3)
- Bryant University (2)
- Chapman University (2)
- Clemson University (2)
- East Tennessee State University (2)
- Embry-Riddle Aeronautical University (2)
- Georgia Southern University (2)
- Old Dominion University (2)
- Portland State University (2)
- Southern Methodist University (2)
- The University of Akron (2)
- University of Kentucky (2)
- University of Nevada, Las Vegas (2)
- California State University, San Bernardino (1)
- Claremont Colleges (1)
- Columbus State University (1)
- DePaul University (1)
- Eastern Washington University (1)
- Fort Hays State University (1)
- Keyword
-
- Data mining (25)
- Social media (20)
- Query processing (19)
- Classification (17)
- Online learning (16)
-
- Twitter (15)
- Machine learning (13)
- Algorithms (12)
- Machine Learning (12)
- Neural networks (11)
- Algorithm (10)
- Natural language processing (10)
- Sentiment analysis (9)
- Artificial intelligence (8)
- Deep Learning (8)
- Spatial databases (8)
- Data structures (7)
- Database (7)
- Location-based services (7)
- Recommender systems (7)
- Spatial database (7)
- Text mining (7)
- Data models (6)
- Deep learning (6)
- Feature extraction (6)
- Information retrieval (6)
- Online Learning (6)
- Road network (6)
- Semantics (6)
- Social network (6)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (661)
- Computer Science Faculty Publications (33)
- CBN Journal of Applied Statistics (JAS) (20)
- Dissertations and Theses Collection (Open Access) (6)
- Graduate Theses and Dissertations (5)
-
- LSU New Orleans Theses and Dissertations (4)
- Theses and Dissertations (4)
- Department of Computer Science Faculty Scholarship and Creative Works (3)
- All Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Department of Agricultural and Biological Systems Engineering: Dissertations, Theses, and Student Research (2)
- Electronic Theses and Dissertations (2)
- Honors Projects in Information Systems and Analytics (2)
- Research Collection Lee Kong Chian School Of Business (2)
- SMU Data Science Review (2)
- STAR Program Research Presentations (2)
- SWITCH (2)
- School of Computing: Dissertations, Theses, and Student Research (2)
- The Summer Undergraduate Research Fellowship (SURF) Symposium (2)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (2)
- Williams Honors College, Honors Research Projects (2)
- Asian Management Insights (1)
- Bulletin of TUIT: Management and Communication Technologies (1)
- CMC Senior Theses (1)
- College of Computing and Digital Media Dissertations (1)
- Computer Science and Computer Engineering Faculty Publications and Presentations (1)
- Computer Science and Computer Engineering Undergraduate Honors Theses (1)
- Computer Science and Software Engineering (1)
- Conference papers (1)
- Department of Computer Science Publications (1)
- Publication Type
Articles 691 - 720 of 809
Full-Text Articles in Numerical Analysis and Scientific Computing
Verifying Completeness Of Relational Query Results In Data Publishing, Hwee Hwa Pang, Arpit Jain, Krithi Ramamritham, Kian-Lee Tan
Verifying Completeness Of Relational Query Results In Data Publishing, Hwee Hwa Pang, Arpit Jain, Krithi Ramamritham, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
In data publishing, the owner delegates the role of satisfying user queries to a third-party publisher. As the publisher may be untrusted or susceptible to attacks, it could produce incorrect query results. In this paper, we introduce a scheme for users to verify that their query results are complete (i.e., no qualifying tuples are omitted) and authentic (i.e., all the result values originated from the owner). The scheme supports range selection on key and non-key attributes, project as well as join queries on relational databases. Moreover, the proposed scheme complies with access control policies, is computationally secure, and can be …
Aggregate Nearest Neighbor Queries In Spatial Databases, Dimitris Papadias, Yufei Tao, Kyriakos Mouratidis, Chun Kit Hui
Aggregate Nearest Neighbor Queries In Spatial Databases, Dimitris Papadias, Yufei Tao, Kyriakos Mouratidis, Chun Kit Hui
Research Collection School Of Computing and Information Systems
Given two spatial datasets P (e.g., facilities) and Q (queries), an aggregate nearest neighbor (ANN) query retrieves the point(s) of P with the smallest aggregate distance(s) to points in Q. Assuming, for example, n users at locations q1,...qn, an ANN query outputs the facility p belongs to P that minimizes the sum of distances |pqi| for 1 is less than or equal to i is less than or equal to n that the users have to travel in order to meet there. Similarly, another ANN query may report the point p belongs to P that minimizes the maximum distance that …
Conceptual Partitioning: An Efficient Method For Continuous Nearest Neighbor Monitoring, Kyriakos Mouratidis, Marios Hadjieleftheriou, Dimitris Papadias
Conceptual Partitioning: An Efficient Method For Continuous Nearest Neighbor Monitoring, Kyriakos Mouratidis, Marios Hadjieleftheriou, Dimitris Papadias
Research Collection School Of Computing and Information Systems
Given a set of objects P and a query point q, a k nearest neighbor (k-NN) query retrieves the k objects in P that lie closest to q. Even though the problem is well-studied for static datasets, the traditional methods do not extend to highly dynamic environments where multiple continuous queries require real-time results, and both objects and queries receive frequent location updates. In this paper we propose conceptual partitioning (CPM), a comprehensive technique for the efficient monitoring of continuous NN queries. CPM achieves low running time by handling location updates only from objects that fall in the vicinity of …
Event-Driven Document Selection For Terrorism, Zhen Sun, Ee Peng Lim, Kuiyu Chang, Teng-Kwee Ong, Rohan Kumar Gunaratna
Event-Driven Document Selection For Terrorism, Zhen Sun, Ee Peng Lim, Kuiyu Chang, Teng-Kwee Ong, Rohan Kumar Gunaratna
Research Collection School Of Computing and Information Systems
In this paper, we examine the task of extracting information about terrorism related events hidden in a large document collection. The task assumes that a terrorism related event can be described by a set of entity and relation instances. To reduce the amount of time and efforts in extracting these event related instances, one should ideally perform the task on the relevant documents only. We have therefore proposed some document selection strategies based on information extraction (IE) patterns. Each strategy attempts to select one document at a time such that the gain of event related instance information is maximized. Our …
Information Dissemination Via Wireless Broadcast, Baihua Zheng, Dik Lun Lee
Information Dissemination Via Wireless Broadcast, Baihua Zheng, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Unrestricted mobility adds a new dimension to data access methodology--- one that must be addressed before true ubiquity can be realized.
Dynamically Optimized Context In Recommender Systems, Ghim-Eng Yap, Ah-Hwee Tan, Hwee Hwa Pang
Dynamically Optimized Context In Recommender Systems, Ghim-Eng Yap, Ah-Hwee Tan, Hwee Hwa Pang
Research Collection School Of Computing and Information Systems
Traditional approaches to recommender systems have not taken into account situational information when making recommendations, and this seriously limits the relevance of the results. This paper advocates context-awareness as a promising approach to enhance the performance of recommenders, and introduces a mechanism to realize this approach. We present a framework that separates the contextual concerns from the actual recommendation module, so that contexts can be readily shared across applications. More importantly, we devise a learning algorithm to dynamically identify the optimal set of contexts for a specific recommendation task and user. An extensive series of experiments has validated that our …
Tosa: A Near-Optimal Scheduling Algorithm For Multi-Channel Data Broadcast, Baihua Zheng, Xia Xu, Xing Jin, Dik Lun Lee
Tosa: A Near-Optimal Scheduling Algorithm For Multi-Channel Data Broadcast, Baihua Zheng, Xia Xu, Xing Jin, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Wireless broadcast is very suitable for delivering information to a large user population. In this paper, we concentrate on data allocation methods for multiple broadcast channels. To the best of our knowledge, this is the first allocation model that takes into the consideration of items' access frequencies, items' lengths. and bandwidth of different channels. We first derive the optimal average expected delay for multiple channels for the general case where data access frequencies, data sizes, and channel bandwidths can all be non-uniform. Second, we develop TOSA, a multi-channel allocation method that does not assume a uniform broadcast schedule for data …
Mining Social Network From Spatio-Temporal Events, Hady Wirawan Lauw, Ee Peng Lim, Teck Tim Tan, Hwee Hwa Pang
Mining Social Network From Spatio-Temporal Events, Hady Wirawan Lauw, Ee Peng Lim, Teck Tim Tan, Hwee Hwa Pang
Research Collection School Of Computing and Information Systems
Knowing patterns of relationship in a social network is very useful for law enforcement agencies to investigate collaborations among criminals, for businesses to exploit relationships to sell products, or for individuals who wish to network with others. After all, it is not just what you know, but also whom you know, that matters. However, finding out who is related to whom on a large scale is a complex problem. Asking every single individual would be impractical, given the huge number of individuals and the changing dynamics of relationships. Recent advancement in technology has allowed more data about activities of individuals …
Proactive Caching For Spatial Queries In Mobile Environments, Haibo Hu, Jianliang Xu, Wing Sing Wong, Baihua Zheng, Dik Lun Lee, Wang-Chien Lee
Proactive Caching For Spatial Queries In Mobile Environments, Haibo Hu, Jianliang Xu, Wing Sing Wong, Baihua Zheng, Dik Lun Lee, Wang-Chien Lee
Research Collection School Of Computing and Information Systems
Semantic caching enables mobile clients to answer spatial queries locally by storing the query descriptions together with the results. However, it supports only a limited number of query types, and sharing results among these types is difficult. To address these issues, we propose a proactive caching model which caches the result objects as well as the index that supports these objects as the results. The cached index enables the objects to be reused for all common types of queries. We also propose an adaptive scheme to cache such an index, which further optimizes the query response time for the best …
Scheduling Queries To Improve The Freshness Of A Website, Haifeng Liu, Wee-Keong Ng, Ee Peng Lim
Scheduling Queries To Improve The Freshness Of A Website, Haifeng Liu, Wee-Keong Ng, Ee Peng Lim
Research Collection School Of Computing and Information Systems
The World Wide Web is a new advertising medium that corporations use to increase their exposure to consumers. Very large websites whose content is derived from a source database need to maintain a freshness that reflects changes that are made to the base data. This issue is particularly significant for websites that present fast-changing information such as stock-exchange information and product information. In this article, we formally define and study the freshness of a website that is refreshed by a scheduled set of queries that fetch fresh data from the databases. We propose several online-scheduling algorithms and compare the performance …
Recommender Systems Research, Saverio Perugini
Recommender Systems Research, Saverio Perugini
Computer Science Faculty Publications
We outline the history of recommender systems from their roots in information retrieval and filtering to their role in today’s Internet economy. Recommender systems attempt to reduce information overload and retain customers by selecting a subset of items from a universal set based on user preferences. Research in recommender systems lies at the intersection of several areas of computer science, such as artificial intelligence and human-computer interaction, and has progressed to an important research area of its own. It is important to note that recommendations are not delivered within a vacuum, but rather cast within an informal community of users …
The Good, Bad And The Indifferent: Explorations In Recommender System Health, Benjamin J. Keller, Sun-Mi Kim, N. Srinivas Vemuri, Naren Ramakrishnan, Saverio Perugini
The Good, Bad And The Indifferent: Explorations In Recommender System Health, Benjamin J. Keller, Sun-Mi Kim, N. Srinivas Vemuri, Naren Ramakrishnan, Saverio Perugini
Computer Science Faculty Publications
Our work is based on the premise that analysis of the connections exploited by a recommender algorithm can provide insight into the algorithm that could be useful to predict its performance in a fielded system. We use the jumping connections model defined by Mirza et al. [6], which describes the recommendation process in terms of graphs. Here we discuss our work that has come out of trying to understand algorithm behavior in terms of these graphs. We start by describing a natural extension of the jumping connections model of Mirza et al., and then discuss observations that have come from …
Linear Correlation Discovery In Databases: A Data Mining Approach, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim
Linear Correlation Discovery In Databases: A Data Mining Approach, Cecil Chua, Roger Hsiang-Li Chiang, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Very little research in knowledge discovery has studied how to incorporate statistical methods to automate linear correlation discovery (LCD). We present an automatic LCD methodology that adopts statistical measurement functions to discover correlations from databases’ attributes. Our methodology automatically pairs attribute groups having potential linear correlations, measures the linear correlation of each pair of attribute groups, and confirms the discovered correlation. The methodology is evaluated in two sets of experiments. The results demonstrate the methodology’s ability to facilitate linear correlation discovery for databases with a large amount of data.
Applying Scenario-Based Design And Claim Analysis To The Design Of A Digital Library Of Geography Examination Resources, Yin-Leng Theng, Dion Hoe-Lian Goh, Ee Peng Lim, Zehua Liu, Ming Yin, Natalie Lee-San Pang, Patricia Bao-Bao Wong
Applying Scenario-Based Design And Claim Analysis To The Design Of A Digital Library Of Geography Examination Resources, Yin-Leng Theng, Dion Hoe-Lian Goh, Ee Peng Lim, Zehua Liu, Ming Yin, Natalie Lee-San Pang, Patricia Bao-Bao Wong
Research Collection School Of Computing and Information Systems
This paper describes the application of Carroll’s scenario-based design and claims analysis as a means of refinement to the initial design of a digital library of geographical resources (GeogDL) to prepare Singapore students to take a national examination in geography. GeogDL is built on top of G-Portal, a digital library providing services over geospatial and georeferenced Web content. Beyond improving the initial design of GeogDL, a main contribution of the paper is making explicit the use of Carroll’s strong theory-based but undercapitalized scenario-based design and claims analysis that inspired recommendations for the refinement of GeogDL. The paper concludes with an …
Justilm: Few-Shot Justification Generation For Explainable Fact-Checking Of Real-World Claims, Fengzhu Zeng, Wei Gao
Justilm: Few-Shot Justification Generation For Explainable Fact-Checking Of Real-World Claims, Fengzhu Zeng, Wei Gao
Research Collection School Of Computing and Information Systems
Justification is an explanation that supports the verdict assigned to a claim in fact-checking. However, the task of justification generation is previously oversimplified as summarization of fact-check article authored by professional checkers. In this work, we propose a realistic approach to generate justification based on retrieved evidence. We present a new benchmark dataset called ExClaim for Explainable Claim verification, and introduce JustiLM, a novel few-shot retrieval-augmented language model to learn justification generation by leveraging fact-check articles as auxiliary resource during training. Our results show that JustiLM outperforms in-context learning (ICL)-enabled LMs including Flan-T5 and Llama2, and the retrieval-augmented model Atlas …
Design Lessons On Access Features In Paper, Yin-Leng Theng, Dion Hoe-Lian Goh, Ming Yin, Eng-Kai Suen, Ee Peng Lim
Design Lessons On Access Features In Paper, Yin-Leng Theng, Dion Hoe-Lian Goh, Ming Yin, Eng-Kai Suen, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Using Nielsen's Heuristic Evaluation, this paper reports a user study with six usability-trained subjects to evaluate PAPER's access features in assisting users to retrieve information efficiently, part of an on-going design partnership with stakeholders and designers/developers. PAPER (Personalised Adaptive Pathways for Exam Resources) is an improved version evolving from an earlier implementation of GeogDL built upon G-Portal, a geospatial digital library infrastructure. After two initial evaluations with student and teacher design partners, PAPER has evolved into a system containing a new bundle of personalized, interactive services with four modules: mock exam; personal coach (practice and review); trend analysis and performance …
Supporting Field Study With Personalized Project Spaces In A Geographical Digital Library, Ee Peng Lim, Aixin Sun, Zehua Liu, John Hedberg, Chew-Hung Chang, Tiong-Sa Teh, Dion Hoe-Lian Goh, Yin-Leng Theng
Supporting Field Study With Personalized Project Spaces In A Geographical Digital Library, Ee Peng Lim, Aixin Sun, Zehua Liu, John Hedberg, Chew-Hung Chang, Tiong-Sa Teh, Dion Hoe-Lian Goh, Yin-Leng Theng
Research Collection School Of Computing and Information Systems
Digital libraries have been rather successful in supporting learning activities by providing learners with access to information and knowledge. However, this level of support is passive to learners and interactive and collaborative learning cannot be easily achieved. In this paper, we study how digital libraries could be extended to serve a more active role in collaborative learning activities. We focus on developing new services to support a common type of learning activity, field study, in a geospatial context. We propose the concept of personal project space that allows individuals to work in their personalized environment with a mix of private …
On Semantic Caching And Query Scheduling For Mobile Nearest-Neighbor Search, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
On Semantic Caching And Query Scheduling For Mobile Nearest-Neighbor Search, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Location-based services have received increasing attention in recent years. In this paper, we address the performance issues of mobile nearest-neighbor search, in which the mobile user issues a query to retrieve stationary service objects nearest to him/her. An index based on Voronoi Diagram is used in the server to support such a search, while a semantic cache is proposed to enhance the access efficiency of the service. Cache replacement policies tailored for the proposed semantic cache are examined. Moreover, several query scheduling policies are proposed to address the inter-cell roaming issues in multi-cell environments. Simulations are conducted to evaluate the …
Spatial Queries In Wireless Broadcast Systems, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Spatial Queries In Wireless Broadcast Systems, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Owing to the advent of wireless networking and personal digital devices, information systems in the era of mobile computing are expected to be able to handle a tremendous amount of traffic and service requests from the users. Wireless data broadcast, thanks to its high scalability, is particularly suitable for meeting such a challenge. Indexing techniques have been developed for wireless data broadcast systems in order to conserve the scarce power resources in mobile clients. However, most of the previous studies do not take into account the impact of location information of users. In this paper, we address the issues of …
The D-Tree: An Index Structure For Planar Point Queries Location-Based Wireless Services, Jianliang Xu, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
The D-Tree: An Index Structure For Planar Point Queries Location-Based Wireless Services, Jianliang Xu, Baihua Zheng, Wang-Chien Lee, Dik Lun Lee
Research Collection School Of Computing and Information Systems
Location-based services (LBSs), considered as a killer application in the wireless data market, provide information based on locations specified in the queries. In this paper, we examine the indexing issue for querying location-dependent data in wireless LBSs; in particular, we focus on an important class of queries, planar point queries. To address the issues of responsiveness, energy consumption, and bandwidth contention in wireless communications, an index has to minimize the search time and maintain a small storage overhead. It is shown that the traditional point-location algorithms and spatial index structures fail to achieve either objective or both. This paper proposes …
Finding Constrained Frequent Episodes Using Minimal Occurrences, Xi Ma, Hwee Hwa Pang, Kian-Lee Tan
Finding Constrained Frequent Episodes Using Minimal Occurrences, Xi Ma, Hwee Hwa Pang, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
Recurrent combinations of events within an event sequence, known as episodes, often reveal useful information. Most of the proposed episode mining algorithms adopt an apriori-like approach that generates candidates and then calculates their support levels. Obviously, such an approach is computationally expensive. Moreover, those algorithms are capable of handling only a limited range of constraints. In this paper, we introduce two mining algorithms - episode prefix tree (EPT) and position pairs set (PPS) - based on a prefix-growth approach to overcome the above limitations. Both algorithms push constraints systematically into the mining process. Performance study shows that the proposed algorithms …
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
One common approach in hierarchical text classification involves associating classifiers with nodes in the category tree and classifying text documents in a top-down manner. Classification methods using this top-down approach can scale well and cope with changes to the category trees. However, all these methods suffer from blocking which refers to documents wrongly rejected by the classifiers at higher-levels and cannot be passed to the classifiers at lower-levels. We propose a classifier-centric performance measure known as blocking factor to determine the extent of the blocking. Three methods are proposed to address the blocking problem, namely, threshold reduction, restricted voting, and …
Shared-Storage Auction Ensures Data Availability, Hady W. Lauw, Siu-Cheung Hui, Edmund M. K. Lai
Shared-Storage Auction Ensures Data Availability, Hady W. Lauw, Siu-Cheung Hui, Edmund M. K. Lai
Research Collection School Of Computing and Information Systems
Most current e-auction systems are based on the client-server architecture. Such centralized systems provide a single point of failure and control. In contrast, peer-to-peer systems permit distributed control and minimize individual node and link failures' impact on the system. The shared-storage-based auction model described decentralizes services among peers to share the required processing load and aggregates peers' resources for common use. The model is based on the principles of local computation at each peer, direct inter-peer communication, and a shared storage space.
Recommender Systems Research: A Connection-Centric Survey, Saverio Perugini, Marcos André Gonçalves, Edward A. Fox
Recommender Systems Research: A Connection-Centric Survey, Saverio Perugini, Marcos André Gonçalves, Edward A. Fox
Computer Science Faculty Publications
Recommender systems attempt to reduce information overload and retain customers by selecting a subset of items from a universal set based on user preferences. While research in recommender systems grew out of information retrieval and filtering, the topic has steadily advanced into a legitimate and challenging research area of its own. Recommender systems have traditionally been studied from a content-based filtering vs. collaborative design perspective. Recommendations, however, are not delivered within a vacuum, but rather cast within an informal community of users and social context. Therefore, ultimately all recommender systems make connections among people and thus should be surveyed from …
A Spectroscopy Of Texts For Effective Clustering, Wenyuan Li, Wee-Keong Ng, Kok-Leong Ong, Ee Peng Lim
A Spectroscopy Of Texts For Effective Clustering, Wenyuan Li, Wee-Keong Ng, Kok-Leong Ong, Ee Peng Lim
Research Collection School Of Computing and Information Systems
For many clustering algorithms, such as k-means, EM, and CLOPE, there is usually a requirement to set some parameters. Often, these parameters directly or indirectly control the number of clusters to return. In the presence of different data characteristics and analysis contexts, it is often difficult for the user to estimate the number of clusters in the data set. This is especially true in text collections such as Web documents, images or biological data. The fundamental question this paper addresses is: ldquoHow can we effectively estimate the natural number of clusters in a given text collection?rdquo. We propose to use …
Ltam: A Location-Temporal Authorization Model, Hai Yu, Ee Peng Lim
Ltam: A Location-Temporal Authorization Model, Hai Yu, Ee Peng Lim
Research Collection School Of Computing and Information Systems
This paper describes an authorization model for specifying access privileges of users who make requests to access a set of locations in a building or more generally a physical or virtual infrastructure. In the model, primitive locations can be grouped into composite locations and the connectivities among locations are represented in a multilevel location graph. Authorizations are defined with temporal constraints on the time to enter and leave a location and constraints on the number of times users can access a location. Access control enforcement is conducted by monitoring user movement and checking access requests against an authorization database. The …
A Support-Ordered Trie For Fast Frequent Itemset Discovery, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
A Support-Ordered Trie For Fast Frequent Itemset Discovery, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
Research Collection School Of Computing and Information Systems
The importance of data mining is apparent with the advent of powerful data collection and storage tools; raw data is so abundant that manual analysis is no longer possible. Unfortunately, data mining problems are difficult to solve and this prompted the introduction of several novel data structures to improve mining efficiency. Here, we critically examine existing preprocessing data structures used in association rule mining for enhancing performance in an attempt to understand their strengths and weaknesses. Our analyses culminate in a practical structure called the SOTrielT (support-ordered trie itemset) and two synergistic algorithms to accompany it for the fast discovery …
Steganographic Schemes For File System And B-Tree, Hwee Hwa Pang, Kian-Lee Tan, Xuan Zhou
Steganographic Schemes For File System And B-Tree, Hwee Hwa Pang, Kian-Lee Tan, Xuan Zhou
Research Collection School Of Computing and Information Systems
While user access control and encryption can protect valuable data from passive observers, these techniques leave visible ciphertexts that are likely to alert an active adversary to the existence of the data. We introduce StegFD, a steganographic file driver that securely hides user-selected files in a file system so that, without the corresponding access keys, an attacker would not be able to deduce their existence. Unlike other steganographic schemes proposed previously, our construction satisfies the prerequisites of a practical file system in ensuring the integrity of the files and maintaining efficient space utilization. We also propose two schemes for implementing …
Efficient Group Pattern Mining Using Data Summarization, Yida Wang, Ee Peng Lim, San-Yih Hwang
Efficient Group Pattern Mining Using Data Summarization, Yida Wang, Ee Peng Lim, San-Yih Hwang
Research Collection School Of Computing and Information Systems
In group pattern mining, we discover group patterns from a given user movement database based on their spatio-temporal distances. When both the number of users and the logging duration are large, group pattern mining algorithms become very inefficient. In this paper, we therefore propose a spherical location summarization method to reduce the overhead of mining valid 2-groups. In our experiments, we show that our group mining algorithm using summarized data may require much less execution time than that using non-summarized data.
Group Nearest Neighbor Queries, Dimitris Papadias, Qiongmao Shen, Yufei Tao, Kyriakos Mouratidis
Group Nearest Neighbor Queries, Dimitris Papadias, Qiongmao Shen, Yufei Tao, Kyriakos Mouratidis
Research Collection School Of Computing and Information Systems
Given two sets of points P and Q, a group nearest neighbor (GNN) query retrieves the point(s) of P with the smallest sum of distances to all points in Q. Consider, for instance, three users at locations q1 , q2 and q3 that want to find a meeting point (e.g., a restaurant); the corresponding query returns the data point p that minimizes the sum of Euclidean distances |pqi| for 1 ≤i ≤3. Assuming that Q fits in memory and P is indexed by an R-tree, we propose several algorithms for finding the group nearest neighbors efficiently. As a second step, …