Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- China Simulation Federation (3880)
- Singapore Management University (1060)
- University of Nebraska - Lincoln (747)
- Missouri University of Science and Technology (104)
- Old Dominion University (51)
-
- University of Arkansas, Fayetteville (48)
- University of Dayton (36)
- California Polytechnic State University, San Luis Obispo (33)
- Chapman University (28)
- City University of New York (CUNY) (28)
- Purdue University (25)
- Embry-Riddle Aeronautical University (24)
- Illinois State University (23)
- University of Kentucky (23)
- Central Bank of Nigeria (21)
- Southern Methodist University (20)
- Technological University Dublin (19)
- University of New Mexico (16)
- University of Montana (15)
- University of Nebraska at Omaha (15)
- Air Force Institute of Technology (13)
- Kennesaw State University (13)
- Claremont Colleges (12)
- LSU New Orleans (12)
- University of Nevada, Las Vegas (12)
- Virginia Commonwealth University (11)
- Central Washington University (10)
- Columbus State University (10)
- Loyola University Chicago (10)
- Edith Cowan University (9)
- Keyword
-
- Simulation (162)
- Deep learning (76)
- Path planning (76)
- Machine learning (75)
- Genetic algorithm (51)
-
- Data mining (47)
- Numerical simulation (47)
- Reinforcement learning (47)
- Modeling (46)
- Machine Learning (44)
- Virtual reality (44)
- Digital twin (43)
- Multi-objective optimization (42)
- Modeling and simulation (39)
- Particle swarm optimization (37)
- Optimization (35)
- Visualization (35)
- Social media (34)
- Deep reinforcement learning (30)
- Attention mechanism (29)
- Fault diagnosis (29)
- Neural network (29)
- Twitter (28)
- Feature extraction (27)
- UAV (27)
- Classification (26)
- Clustering (26)
- Neural networks (25)
- Simulation model (25)
- Artificial intelligence (24)
- Publication Year
- Publication
-
- Journal of System Simulation (3880)
- Research Collection School Of Computing and Information Systems (1024)
- The R Journal (708)
- Computer Science Faculty Publications (52)
- Physics Faculty Research & Creative Works (44)
-
- Theses and Dissertations (41)
- CBN Journal of Applied Statistics (JAS) (21)
- Annual Symposium on Biomathematics and Ecology Education and Research (20)
- Geosciences and Geological and Petroleum Engineering Faculty Research & Creative Works (19)
- Graduate Theses and Dissertations (19)
- Dissertations (15)
- Dissertations and Theses Collection (Open Access) (15)
- Electronic Theses and Dissertations (15)
- Graduate Student Theses, Dissertations, & Professional Papers (15)
- Chemistry Faculty Research & Creative Works (14)
- Computer Science and Computer Engineering Undergraduate Honors Theses (14)
- Master's Theses (14)
- The Summer Undergraduate Research Fellowship (SURF) Symposium (14)
- Doctoral Dissertations and Master's Theses (12)
- LSU New Orleans Theses and Dissertations (12)
- Dissertations, Theses, and Capstone Projects (11)
- Electrical & Computer Engineering Theses & Dissertations (11)
- Holland Computing Center: Faculty Publications (10)
- SMU Data Science Review (10)
- Computer Science Theses & Dissertations (9)
- Computer Science: Faculty Publications and Other Works (9)
- Interdisciplinary Informatics Faculty Proceedings & Presentations (9)
- Publications and Research (9)
- STAR Program Research Presentations (9)
- Williams Honors College, Honors Research Projects (9)
- Publication Type
- File Type
Articles 6511 - 6540 of 6663
Full-Text Articles in Computer Sciences
Finding Constrained Frequent Episodes Using Minimal Occurrences, Xi Ma, Hwee Hwa Pang, Kian-Lee Tan
Finding Constrained Frequent Episodes Using Minimal Occurrences, Xi Ma, Hwee Hwa Pang, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
Recurrent combinations of events within an event sequence, known as episodes, often reveal useful information. Most of the proposed episode mining algorithms adopt an apriori-like approach that generates candidates and then calculates their support levels. Obviously, such an approach is computationally expensive. Moreover, those algorithms are capable of handling only a limited range of constraints. In this paper, we introduce two mining algorithms - episode prefix tree (EPT) and position pairs set (PPS) - based on a prefix-growth approach to overcome the above limitations. Both algorithms push constraints systematically into the mining process. Performance study shows that the proposed algorithms …
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
One common approach in hierarchical text classification involves associating classifiers with nodes in the category tree and classifying text documents in a top-down manner. Classification methods using this top-down approach can scale well and cope with changes to the category trees. However, all these methods suffer from blocking which refers to documents wrongly rejected by the classifiers at higher-levels and cannot be passed to the classifiers at lower-levels. We propose a classifier-centric performance measure known as blocking factor to determine the extent of the blocking. Three methods are proposed to address the blocking problem, namely, threshold reduction, restricted voting, and …
Shared-Storage Auction Ensures Data Availability, Hady W. Lauw, Siu-Cheung Hui, Edmund M. K. Lai
Shared-Storage Auction Ensures Data Availability, Hady W. Lauw, Siu-Cheung Hui, Edmund M. K. Lai
Research Collection School Of Computing and Information Systems
Most current e-auction systems are based on the client-server architecture. Such centralized systems provide a single point of failure and control. In contrast, peer-to-peer systems permit distributed control and minimize individual node and link failures' impact on the system. The shared-storage-based auction model described decentralizes services among peers to share the required processing load and aggregates peers' resources for common use. The model is based on the principles of local computation at each peer, direct inter-peer communication, and a shared storage space.
Recommender Systems Research: A Connection-Centric Survey, Saverio Perugini, Marcos André Gonçalves, Edward A. Fox
Recommender Systems Research: A Connection-Centric Survey, Saverio Perugini, Marcos André Gonçalves, Edward A. Fox
Computer Science Faculty Publications
Recommender systems attempt to reduce information overload and retain customers by selecting a subset of items from a universal set based on user preferences. While research in recommender systems grew out of information retrieval and filtering, the topic has steadily advanced into a legitimate and challenging research area of its own. Recommender systems have traditionally been studied from a content-based filtering vs. collaborative design perspective. Recommendations, however, are not delivered within a vacuum, but rather cast within an informal community of users and social context. Therefore, ultimately all recommender systems make connections among people and thus should be surveyed from …
A Spectroscopy Of Texts For Effective Clustering, Wenyuan Li, Wee-Keong Ng, Kok-Leong Ong, Ee Peng Lim
A Spectroscopy Of Texts For Effective Clustering, Wenyuan Li, Wee-Keong Ng, Kok-Leong Ong, Ee Peng Lim
Research Collection School Of Computing and Information Systems
For many clustering algorithms, such as k-means, EM, and CLOPE, there is usually a requirement to set some parameters. Often, these parameters directly or indirectly control the number of clusters to return. In the presence of different data characteristics and analysis contexts, it is often difficult for the user to estimate the number of clusters in the data set. This is especially true in text collections such as Web documents, images or biological data. The fundamental question this paper addresses is: ldquoHow can we effectively estimate the natural number of clusters in a given text collection?rdquo. We propose to use …
Ltam: A Location-Temporal Authorization Model, Hai Yu, Ee Peng Lim
Ltam: A Location-Temporal Authorization Model, Hai Yu, Ee Peng Lim
Research Collection School Of Computing and Information Systems
This paper describes an authorization model for specifying access privileges of users who make requests to access a set of locations in a building or more generally a physical or virtual infrastructure. In the model, primitive locations can be grouped into composite locations and the connectivities among locations are represented in a multilevel location graph. Authorizations are defined with temporal constraints on the time to enter and leave a location and constraints on the number of times users can access a location. Access control enforcement is conducted by monitoring user movement and checking access requests against an authorization database. The …
A Support-Ordered Trie For Fast Frequent Itemset Discovery, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
A Support-Ordered Trie For Fast Frequent Itemset Discovery, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
Research Collection School Of Computing and Information Systems
The importance of data mining is apparent with the advent of powerful data collection and storage tools; raw data is so abundant that manual analysis is no longer possible. Unfortunately, data mining problems are difficult to solve and this prompted the introduction of several novel data structures to improve mining efficiency. Here, we critically examine existing preprocessing data structures used in association rule mining for enhancing performance in an attempt to understand their strengths and weaknesses. Our analyses culminate in a practical structure called the SOTrielT (support-ordered trie itemset) and two synergistic algorithms to accompany it for the fast discovery …
Tournament Versus Fitness Uniform Selection, Shane Legg, Marcus Hutter, Akshat Kumar
Tournament Versus Fitness Uniform Selection, Shane Legg, Marcus Hutter, Akshat Kumar
Research Collection School Of Computing and Information Systems
In evolutionary algorithms a critical parameter that must be tuned is that of selection pressure. If it is set too low then the rate of convergence towards the optimum is likely to be slow. Alternatively if the selection pressure is set too high the system is likely to become stuck in a local optimum due to a loss of diversity in the population. The recent Fitness Uniform Selection Scheme (FUSS) is a conceptually simple but somewhat radical approach to addressing this problem - rather than biasing the selection towards higher fitness, FUSS biases selection towards sparsely populated fitness levels. In …
Steganographic Schemes For File System And B-Tree, Hwee Hwa Pang, Kian-Lee Tan, Xuan Zhou
Steganographic Schemes For File System And B-Tree, Hwee Hwa Pang, Kian-Lee Tan, Xuan Zhou
Research Collection School Of Computing and Information Systems
While user access control and encryption can protect valuable data from passive observers, these techniques leave visible ciphertexts that are likely to alert an active adversary to the existence of the data. We introduce StegFD, a steganographic file driver that securely hides user-selected files in a file system so that, without the corresponding access keys, an attacker would not be able to deduce their existence. Unlike other steganographic schemes proposed previously, our construction satisfies the prerequisites of a practical file system in ensuring the integrity of the files and maintaining efficient space utilization. We also propose two schemes for implementing …
Parameter Inference Of Queueing Models For It Systems Using End-To-End Measurements, Laura Wynter, Cathy H. Xia, Fan Zhang
Parameter Inference Of Queueing Models For It Systems Using End-To-End Measurements, Laura Wynter, Cathy H. Xia, Fan Zhang
Research Collection School Of Computing and Information Systems
he scope of available applications, IT systems increase at a fascinating rate in both size and complexity. For example, today, a typical Web service hosting center may have hundreds of nodes and dozens of different applications simultaneously running on it. Each of the nodes in turn has often multiple processors and layered caches. These nodes make use of both local and shared storage systems. The size and complexity of such systems make performance modeling much more difficult, if at all tractable. Detailed modeling, fine tuning and accurate analysis can be carried out only on very small IT systems or very …
Performance Planning, Quality-Of-Service, And Pricing Under Competition, Corinne Touati, Parijat Dube, Laura Wynter
Performance Planning, Quality-Of-Service, And Pricing Under Competition, Corinne Touati, Parijat Dube, Laura Wynter
Research Collection School Of Computing and Information Systems
In this work we model the relationship between the capacity and the Quality of Service (QoS) offered by the firm in a competitive scenario of two firm’s working to maximize their profits. Using simple queueing theoretic models we study the sensitivity of a firm’s market share to price, capacity and market size. Our preliminary studies yield important properties of the equilibrium solution which may further provide important “engineering” guidelines for performance planning and pricing strategies.
Dynamic Offloading In A Multi-Provider Environment: A Behavioral Framework For Use In Influencing Peering, Zhen Liu, Vishal Misra, Laura Wynter
Dynamic Offloading In A Multi-Provider Environment: A Behavioral Framework For Use In Influencing Peering, Zhen Liu, Vishal Misra, Laura Wynter
Research Collection School Of Computing and Information Systems
We pose the question of how to encourage the resource sharing in a distributed, multi-provider environment, where each node, or provider, has local work but is able to accept additional work from other nodes/providers if there is available capacity. An instance of such an environment is found in content delivery, where. numerous, competing providers can work together if enough benefit is to be gained from doing so. We model individual provider behavior as essentially selfish, and then propose pricing schemes to exploit the selfishness to achieve system wide performance gains. We employ a game theoretic framework to analyze the problem, …
Efficient Group Pattern Mining Using Data Summarization, Yida Wang, Ee Peng Lim, San-Yih Hwang
Efficient Group Pattern Mining Using Data Summarization, Yida Wang, Ee Peng Lim, San-Yih Hwang
Research Collection School Of Computing and Information Systems
In group pattern mining, we discover group patterns from a given user movement database based on their spatio-temporal distances. When both the number of users and the logging duration are large, group pattern mining algorithms become very inefficient. In this paper, we therefore propose a spherical location summarization method to reduce the overhead of mining valid 2-groups. In our experiments, we show that our group mining algorithm using summarized data may require much less execution time than that using non-summarized data.
Group Nearest Neighbor Queries, Dimitris Papadias, Qiongmao Shen, Yufei Tao, Kyriakos Mouratidis
Group Nearest Neighbor Queries, Dimitris Papadias, Qiongmao Shen, Yufei Tao, Kyriakos Mouratidis
Research Collection School Of Computing and Information Systems
Given two sets of points P and Q, a group nearest neighbor (GNN) query retrieves the point(s) of P with the smallest sum of distances to all points in Q. Consider, for instance, three users at locations q1 , q2 and q3 that want to find a meeting point (e.g., a restaurant); the corresponding query returns the data point p that minimizes the sum of Euclidean distances |pqi| for 1 ≤i ≤3. Assuming that Q fits in memory and P is indexed by an R-tree, we propose several algorithms for finding the group nearest neighbors efficiently. As a second step, …
Natural Xml For Data Binding, Processing, And Persistence, George K. Thiruvathukal, Konstantin Läufer
Natural Xml For Data Binding, Processing, And Persistence, George K. Thiruvathukal, Konstantin Läufer
Computer Science: Faculty Publications and Other Works
The article explains what you need to do to incorporate XML directly into your computational science application. The exploration involves the use of a standard parser to automatically build object trees entirely from application-specific classes. This discussion very much focuses on object-oriented programming languages such as Java and Python, but it can work for non-object-oriented languages as well. The ideas in the article provide a glimpse into the Natural XML research project.
Spatial Queries In The Presence Of Obstacles, Jun Zhang, Dimitris Papadias, Kyriakos Mouratidis, Manli Zhu
Spatial Queries In The Presence Of Obstacles, Jun Zhang, Dimitris Papadias, Kyriakos Mouratidis, Manli Zhu
Research Collection School Of Computing and Information Systems
Despite the existence of obstacles in many database applications, traditional spatial query processing utilizes the Euclidean distance metric assuming that points in space are directly reachable. In this paper, we study spatial queries in the presence of obstacles, where the obstructed distance between two points is defined as the length of the shortest path that connects them without crossing any obstacles. We propose efficient algorithms for the most important query types, namely, range search, nearest neighbors, e-distance joins and closest pairs, considering that both data objects and obstacles are indexed by R-trees. The effectiveness of the proposed solutions is verified …
An Automated Algorithm For Extracting Website Skeleton, Zehua Liu, Wee-Keong Ng, Ee Peng Lim
An Automated Algorithm For Extracting Website Skeleton, Zehua Liu, Wee-Keong Ng, Ee Peng Lim
Research Collection School Of Computing and Information Systems
The huge amount of information available on the Web has attracted many research efforts into developing wrappers that extract data from webpages. However, as most of the systems for generating wrappers focus on extracting data at page-level, data extraction at site-level remains a manual or semi-automatic process. In this paper, we study the problem of extracting website skeleton, i.e. extracting the underlying hyperlink structure that is used to organize the content pages in a given website. We propose an automated algorithm, called the Sew algorithm, to discover the skeleton of a website. Given a page, the algorithm examines hyperlinks in …
Authenticating Query Results In Edge Computing, Hwee Hwa Pang, Kian-Lee Tan
Authenticating Query Results In Edge Computing, Hwee Hwa Pang, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
Edge computing pushes application logic and the underlying data to the edge of the network, with the aim of improving availability and scalability. As the edge servers are not necessarily secure, there must be provisions for validating their outputs. This paper proposes a mechanism that creates a verification object (VO) for checking the integrity of each query result produced by an edge server - that values in the result tuples are not tampered with, and that no spurious tuples are introduced. The primary advantages of our proposed mechanism are that the VO is independent of the database size, and that …
Hiding Data Accesses In Steganographic File System, Xuan Zhou, Hwee Hwa Pang, Kian-Lee Tan
Hiding Data Accesses In Steganographic File System, Xuan Zhou, Hwee Hwa Pang, Kian-Lee Tan
Research Collection School Of Computing and Information Systems
To support ubiquitous computing, the underlying data have to be persistent and available anywhere-anytime. The data thus have to migrate from devices local to individual computers, to shared storage volumes that are accessible over open network. This potentially exposes the data to heightened security risks. We propose two mechanisms, in the context of a steganographic file system, to mitigate the risk of attacks initiated through analyzing data accesses from user applications. The first mechanism is intended to counter attempts to locate data through updates in between snapshots - in short, update analysis. The second mechanism prevents traffic analysis - identifying …
Automatically Generating Interfaces For Personalized Interaction With Digital Libraries, Saverio Perugini, Naren Ramakrishnan, Edward A. Fox
Automatically Generating Interfaces For Personalized Interaction With Digital Libraries, Saverio Perugini, Naren Ramakrishnan, Edward A. Fox
Computer Science Faculty Publications
We present an approach to automatically generate interfaces supporting personalized interaction with digital libraries; these interfaces augment the user-DL dialog by empowering the user to (optionally) supply out-of-turn information during an interaction, flatten or restructure the dialog, and inquire about dialog options. Interfaces generated using this approach for CITIDEL are described.
Incremental Genetic K-Means Algorithm And Its Application In Gene Expression Data Analysis, Yi Lu, Shiyong Lu, Farshad Fotouhi, Youping Deng, Susan J. Brown
Incremental Genetic K-Means Algorithm And Its Application In Gene Expression Data Analysis, Yi Lu, Shiyong Lu, Farshad Fotouhi, Youping Deng, Susan J. Brown
Wayne State University Associated BioMed Central Scholarship
Abstract
Background
In recent years, clustering algorithms have been effectively applied in molecular biology for gene expression data analysis. With the help of clustering algorithms such as K-means, hierarchical clustering, SOM, etc, genes are partitioned into groups based on the similarity between their expression profiles. In this way, functionally related genes are identified. As the amount of laboratory data in molecular biology grows exponentially each year due to advanced technologies such as Microarray, new efficient and effective methods for clustering must be developed to process this growing amount of biological data.
Results
In this paper, we propose a new clustering …
Web Usage Mining: Algorithms And Results, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
Web Usage Mining: Algorithms And Results, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
Research Collection School Of Computing and Information Systems
No abstract provided.
Xml In Computational Science, George K. Thiruvathukal
Xml In Computational Science, George K. Thiruvathukal
Computer Science: Faculty Publications and Other Works
In this first article in a series about XML in computational science, I present some background and lightweight examples of XML usage, describe some XML component frameworksalong with their purpose and applicability to computational science, and discuss some technical obstacles to overcome for the language to be taken seriously in computational science.
Analysis Of A Yield Management Model For On Demand Computing Centers, Yezekael Hayel, Laura Wynter, Parijat Dube
Analysis Of A Yield Management Model For On Demand Computing Centers, Yezekael Hayel, Laura Wynter, Parijat Dube
Research Collection School Of Computing and Information Systems
The concept of yield management for IT infrastructures, and in particular for on demand IT utilities was recently introduced in [17]. The present paper provides a detailed analysis of that model, both in simplified cases where an analytical analysis is possible, and numerically on larger problem instances, and confirms the significant revenue benefit that can accrue through use of yield management in an IT on demand operating environment.
Development Of A Systems Engineering Model For Chemical Separation Process, Lijian Sun
Development Of A Systems Engineering Model For Chemical Separation Process, Lijian Sun
UNLV Theses, Dissertations, Professional Papers, and Capstones
This thesis is concerned with the efforts to develop a general-purpose systems engineering model software TRPSEMPro1 that can be used to improve productivity in the design process. Different features of TRPSEMPro will be presented in this thesis. First, Systems Engineering technology is presented, followed by the exposition of different numerical optimization technologies and DOE (Design of Experiments) study technologies. Second, the detailed software process, Object-Oriented Analysis and Design (OOA&D) for the TRPSEMPro is presented. All the design data models are expressed by using Unified Modeling Language (UML).
AMUSESimulator is another software package which has been designed and implemented in order …
Predicting Nonlinear Network Traffic Using Fuzzy Neural Network, Zhaoxia Wang, Tingzhu Hao, Zengqiang Chen, Zhuzhi Yuan
Predicting Nonlinear Network Traffic Using Fuzzy Neural Network, Zhaoxia Wang, Tingzhu Hao, Zengqiang Chen, Zhuzhi Yuan
Research Collection School Of Computing and Information Systems
Network traffic is a complex and nonlinear process, which is significantly affected by immeasurable parameters and variables. This paper addresses the use of the five-layer fuzzy neural network (FNN) for predicting the nonlinear network traffic. The structure of this system is introduced in detail. Through training the FNN using back-propagation algorithm with inertia] terms the traffic series can be well predicted by this FNN system. We analyze the performance of the FNN in terms of prediction ability as compared with solely neural network. The simulation demonstrates that the proposed FNN is superior to the solely neural network systems. In addition, …
Xstamps: A Multiversion Timestamps Concurrency Control Protocol For Xml Data, Khin-Myo Win, Wee-Keong Ng, Ee Peng Lim
Xstamps: A Multiversion Timestamps Concurrency Control Protocol For Xml Data, Khin-Myo Win, Wee-Keong Ng, Ee Peng Lim
Research Collection School Of Computing and Information Systems
With the tremendous growth of XML data over the Web, efficient management of such data becomes a new challenge for database community. Several data management solutions, proposed in recent years, extend the capability of traditional database systems to meet the needs of XML data while alternative approaches introduce new generation databases, named as native XML database management systems. Although traditional databases have mature transaction management and concurrency control techniques, there is still a need to tailor techniques for native XML databases in order to deal with distinct characteristics of XML. In this paper, we propose XStamps, a multiversion timestamps concurrency …
Paper For An Educational Digital Library, Dion Hoe-Lian Goh, Yin-Leng Theng, Ming Yin, Ee Peng Lim
Paper For An Educational Digital Library, Dion Hoe-Lian Goh, Yin-Leng Theng, Ming Yin, Ee Peng Lim
Research Collection School Of Computing and Information Systems
GeogDL is a digital library of geography examination resources designed to assist students in revising for a national geography examination in Singapore. As part of an iterative design process, we carried out participatory design and brainstorming with student and teacher design partners. The first study involved prospective student design partners. In response to the first study, we describe in this paper an implementation of PAPER – Personalised Adaptive Pathways for Exam Resources – a new bundle of personalized, interactive services containing a mock exam and a personal coach. The mock exam provides a simulation of the actual geography examination while …
Pricing And Qos Of Information Services In A Competitive Market, Zhen Liu, Laura Wynter, Cathy Xia
Pricing And Qos Of Information Services In A Competitive Market, Zhen Liu, Laura Wynter, Cathy Xia
Research Collection School Of Computing and Information Systems
Design of e-commerce services that are competitive in a quickly responding market requires the analyses of prices and price structures. We develop a general model of an e-commerce market that allows us to analyze optimal price structures, both flat and usage-based. Based on the price structure of a major web hosting provider, we consider single-tier and two-tier (burst-rate) pricing, and our result suggests that the more complex two-tier structure may not be worth the marketing effort, as the firm's equilibrium profits will not increase through the use of this structure. An essential feature of our approach is that we model …
Web Unit Mining: Finding And Classifying Subgraphs Of Web Pages, Aixin Sun, Ee Peng Lim
Web Unit Mining: Finding And Classifying Subgraphs Of Web Pages, Aixin Sun, Ee Peng Lim
Research Collection School Of Computing and Information Systems
In web classification, most researchers assume that the objects to classify are individual web pages from one or more web sites. In practice, the assumption is too restrictive since a web page itself may not always correspond to a concept instance of some semantic concept (or category) given to the classification task. In this paper, we want to relax this assumption and allow a concept instance to be represented by a subgraph of web pages or a set of web pages. We identify several new issues to be addressed when the assumption is removed, and formulate the web unit mining …