Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (90)
- Engineering (60)
- Numerical Analysis and Scientific Computing (47)
- Computer Engineering (46)
- Artificial Intelligence and Robotics (36)
-
- Social and Behavioral Sciences (32)
- Software Engineering (24)
- Business (22)
- Electrical and Computer Engineering (22)
- Mathematics (19)
- Data Science (17)
- Systems Architecture (17)
- Theory and Algorithms (17)
- Life Sciences (16)
- Statistics and Probability (16)
- Logic and Foundations (13)
- Communication (11)
- Information Security (11)
- Medicine and Health Sciences (11)
- Operations Research, Systems Engineering and Industrial Engineering (11)
- Other Computer Sciences (11)
- Bioinformatics (9)
- Management Information Systems (9)
- Social Media (9)
- Systems Science (8)
- Education (7)
- Library and Information Science (6)
- OS and Networks (6)
- Institution
-
- Singapore Management University (70)
- Portland State University (28)
- New Jersey Institute of Technology (16)
- TÜBİTAK (16)
- University at Albany, State University of New York (9)
-
- Zayed University (9)
- Air Force Institute of Technology (8)
- China Simulation Federation (8)
- Louisiana State University (8)
- Old Dominion University (8)
- Institute of Business Administration (7)
- Louisiana Tech University (7)
- University of Nebraska - Lincoln (7)
- Edith Cowan University (6)
- Technological University Dublin (6)
- University of Texas at Arlington (6)
- Brigham Young University (4)
- Kennesaw State University (4)
- MBZUAI (4)
- University of Louisville (4)
- University of Nevada, Las Vegas (4)
- Western Kentucky University (4)
- Central Washington University (3)
- Claremont Colleges (3)
- Clemson University (3)
- Embry-Riddle Aeronautical University (3)
- Nova Southeastern University (3)
- Purdue University (3)
- San Jose State University (3)
- University of Kentucky (3)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (64)
- Complex Systems Faculty Publications and Presentations (22)
- Dissertations (16)
- Turkish Journal of Electrical Engineering and Computer Sciences (16)
- Theses and Dissertations (11)
-
- All Works (9)
- Legacy Theses & Dissertations (2009 - 2024) (9)
- Doctoral Dissertations (8)
- Journal of System Simulation (8)
- International Conference on Information and Communication Technologies (7)
- Computer Science Faculty Publications (6)
- Electronic Theses and Dissertations (6)
- Computer Science Faculty Publications and Presentations (5)
- Faculty Publications (5)
- LSU Doctoral Dissertations (5)
- Computer Science and Engineering Theses - Archive (4)
- Dissertations and Theses Collection (Open Access) (4)
- Faculty Articles (4)
- Theses (4)
- All Faculty Scholarship for the College of the Sciences (3)
- CCAC Theses and Dissertations (3)
- CGU Faculty Publications and Research (3)
- Conference papers (3)
- LSU Master's Theses (3)
- Machine Learning Faculty Publications (3)
- Masters Theses & Specialist Projects (3)
- Theses and Dissertations--Computer Science (3)
- UNLV Theses, Dissertations, Professional Papers, and Capstones (3)
- All Dissertations (2)
- All Graduate Theses, Dissertations, and Other Capstone Projects (2)
- Publication Type
Articles 211 - 240 of 337
Full-Text Articles in Computer Sciences
Using Machine Learning Techniques To Customize The User's Profile, Helps Intelligent Tv Decoder’S Design, Alketa Hyso, Roneda Mucaj
Using Machine Learning Techniques To Customize The User's Profile, Helps Intelligent Tv Decoder’S Design, Alketa Hyso, Roneda Mucaj
UBT International Conference
In today's society due to the increase of the quantity of information is becoming more difficult to find the information we search. "Data mining" offers us the most important methods and techniques in data analysis. Through this work, we aim to study the several data mining techniques, methods and applications in specific areas. We experiment with an “open software" WEKA, to perform some data analysis, presenting the reliability and advantages of data mining classification technique. We use the decision trees technique to achieve the task of classification, to customize user profiles based on their requirements and needs. This paper presents …
Mining Branching-Time Scenarios, Dirk Fahland, David Lo, Shahar Maoz
Mining Branching-Time Scenarios, Dirk Fahland, David Lo, Shahar Maoz
Research Collection School Of Computing and Information Systems
Specification mining extracts candidate specification from existing systems, to be used for downstream tasks such as testing and verification. Specifically, we are interested in the extraction of behavior models from execution traces. In this paper we introduce mining of branching-time scenarios in the form of existential, conditional Live Sequence Charts, using a statistical data-mining algorithm. We show the power of branching scenarios to reveal alternative scenario-based behaviors, which could not be mined by previous approaches. The work contrasts and complements previous works on mining linear-time scenarios. An implementation and evaluation over execution trace sets recorded from several real-world applications shows …
Modeling Interaction Features For Debate Side Clustering, Minghui Qiu, Liu Yang, Jing Jiang
Modeling Interaction Features For Debate Side Clustering, Minghui Qiu, Liu Yang, Jing Jiang
Research Collection School Of Computing and Information Systems
Online discussion forums are popular social media platforms for users to express their opinions and discuss controversial issues with each other. To automatically identify the sides/stances of posts or users from textual content in forums is an important task to help mine online opinions. To tackle the task, it is important to exploit user posts that implicitly contain support and dispute (interaction) information. The challenge we face is how to mine such interaction information from the content of posts and how to use them to help identify stances. This paper proposes a two-stage solution based on latent variable models: an …
Automated Library Recommendation, Ferdian Thung, David Lo, Julia Lawall
Automated Library Recommendation, Ferdian Thung, David Lo, Julia Lawall
Research Collection School Of Computing and Information Systems
Many third party libraries are available to be downloaded and used. Using such libraries can reduce development time and make the developed software more reliable. However, developers are often unaware of suitable libraries to be used for their projects and thus they miss out on these benefits. To help developers better take advantage of the available libraries, we propose a new technique that automatically recommends libraries to developers. Our technique takes as input the set of libraries that an application currently uses, and recommends other libraries that are likely to be relevant. We follow a hybrid approach that combines association …
A Data Mining Approach To Market Power Analysis, Nghia Tran, Dr. Sean Warnick
A Data Mining Approach To Market Power Analysis, Nghia Tran, Dr. Sean Warnick
Journal of Undergraduate Research
The question about market power is interesting both for business managers and market regulators. Market regulators are interested in maximizing total social welfare in a market by protecting market competition. Business managers want to know the power of different firms or parts of a firm to know their true value in acquisition. The measurement of market power turns out to computation-intensive and requires a very large amount of market data. The well-known measure, the Herfindahl-Hirshman Index, requires on large amount of scanner data to determine the market share of individual firms in markets. This measure also relies on legal definitions …
Generative Models For Item Adoptions Using Social Correlation, Freddy Chong Tat Chua, Hady Wirawan Lauw, Ee Peng Lim
Generative Models For Item Adoptions Using Social Correlation, Freddy Chong Tat Chua, Hady Wirawan Lauw, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Users face many choices on the Web when it comes to choosing which product to buy, which video to watch, etc. In making adoption decisions, users rely not only on their own preferences, but also on friends. We call the latter social correlation which may be caused by the homophily and social influence effects. In this paper, we focus on modeling social correlation on users’ item adoptions. Given a user-user social graph and an item-user adoption graph, our research seeks to answer the following questions: whether the items adopted by a user correlate to items adopted by her friends, and …
Knowledge Extraction From Survey Data Using Neural Networks, Imran Ahmed Khan
Knowledge Extraction From Survey Data Using Neural Networks, Imran Ahmed Khan
Computer Science Theses
Surveys are an important tool for researchers. Survey attributes are typically discrete data measured on a Likert scale. Collected responses from the survey contain an enormous amount of data. It is increasingly important to develop powerful means for clustering such data and knowledge extraction that could help in decision-making. The process of clustering becomes complex if the number of survey attributes is large. Another major issue in Likert-Scale data is the uniqueness of tuples. A large number of unique tuples may result in a large number of patterns and that may increase the complexity of the knowledge extraction process. Also, …
Data Near Here: Bringing Relevant Data Closer To Scientists, Veronika M. Megler, David Maier
Data Near Here: Bringing Relevant Data Closer To Scientists, Veronika M. Megler, David Maier
Computer Science Faculty Publications and Presentations
Large scientific repositories run the risk of losing value as their holdings expand, if it means increased effort for a scientist to locate particular datasets of interest. We discuss the challenges that scientists face in locating relevant data, and present our work in applying Information Retrieval techniques to dataset search, as embodied in the Data Near Here application.
Document Collection Visualization And Clustering Using An Atom Metaphor For Display And Interaction, Khanh V. Nghi
Document Collection Visualization And Clustering Using An Atom Metaphor For Display And Interaction, Khanh V. Nghi
Theses and Dissertations - UTB/UTPA
Visual Data Mining have proven to be of high value in exploratory data analysis and data mining because it provides an intuitive feedback on data analysis and support decision-making activities. Several visualization techniques have been developed for cluster discovery such as Grand Tour, HD-Eye, Star Coordinates, etc. They are very useful tool which are visualized in 2D or 3D; however, they have not simple for users who are not trained. This thesis proposes a new approach to build a 3D clustering visualization system for document clustering by using k-mean algorithm. A cluster will be represented by a neutron (centroid) and …
Disclosing Climate Change Patterns Using An Adaptive Markov Chain Pattern Detection Method, Zhaoxia Wang, Gary Lee, Hoong Maeng Chan, Reuben Li, Xiuju Fu, Rick Goh, Pauline A. W. Poh Kim, Martin L. Hibberd, Hoong Chor Chin
Disclosing Climate Change Patterns Using An Adaptive Markov Chain Pattern Detection Method, Zhaoxia Wang, Gary Lee, Hoong Maeng Chan, Reuben Li, Xiuju Fu, Rick Goh, Pauline A. W. Poh Kim, Martin L. Hibberd, Hoong Chor Chin
Research Collection School Of Computing and Information Systems
This paper proposes an adaptive Markov chain pattern detection (AMCPD) method for disclosing the climate change patterns of Singapore through meteorological data mining. Meteorological variables, including daily mean temperature, mean dew point temperature, mean visibility, mean wind speed, maximum sustained wind speed, maximum temperature and minimum temperature are simultaneously considered for identifying climate change patterns in this study. The results depict various weather patterns from 1962 to 2011 in Singapore, based on the records of the Changi Meteorological Station. Different scenarios with varied cluster thresholds are employed for testing the sensitivity of the proposed method. The robustness of the proposed …
Predicting Sql Injection And Cross Site Scripting Vulnerabilities Through Mining Input Sanitization Patterns, Lwin Khin Shar, Hee Beng Kuan Tan
Predicting Sql Injection And Cross Site Scripting Vulnerabilities Through Mining Input Sanitization Patterns, Lwin Khin Shar, Hee Beng Kuan Tan
Research Collection School Of Computing and Information Systems
ContextSQL injection (SQLI) and cross site scripting (XSS) are the two most common and serious web application vulnerabilities for the past decade. To mitigate these two security threats, many vulnerability detection approaches based on static and dynamic taint analysis techniques have been proposed. Alternatively, there are also vulnerability prediction approaches based on machine learning techniques, which showed that static code attributes such as code complexity measures are cheap and useful predictors. However, current prediction approaches target general vulnerabilities. And most of these approaches locate vulnerable code only at software component or file levels. Some approaches also involve process attributes that …
New Algorithms For Frequent Sequential Pattern And Itemset Data Mining In Certain And Uncertain Databases, Erich Allen Peterson
New Algorithms For Frequent Sequential Pattern And Itemset Data Mining In Certain And Uncertain Databases, Erich Allen Peterson
Theses and Dissertations
items) which do not appear in the database. Given this taxonomy, there exist generalized items which do not appear in the uncertain database, but nevertheless, can be probabilistically frequent within. Thus, a new method is introduced which calculates the probability of a generalized item appearing within a transaction, and thus, one can then mine for PGFIs in an uncertain database.
Data Mining The Functional Characterizations Of Proteins To Predict Their Cancer-Relatedness, Peter Revesz, Christopher Assi
Data Mining The Functional Characterizations Of Proteins To Predict Their Cancer-Relatedness, Peter Revesz, Christopher Assi
School of Computing: Faculty Publications
This paper considers two types of protein data. First, data about protein function described in a number of ways, such as, GO terms and PFAM families. Second, data about whether individual proteins are experimentally associated with cancer by an anomalous elevation or lowering of their expressions within cancerous cells. We combine these two types of protein data and test whether the first type of data, that is, the functional descriptors, can predict the second type of data, that is, cancer-relatedness. By using data mining and machine learning, we derive a classifier algorithm that using only GO term and PFAM family …
A Novel Computational Framework For Transcriptome Analysis With Rna-Seq Data, Yin Hu
A Novel Computational Framework For Transcriptome Analysis With Rna-Seq Data, Yin Hu
Theses and Dissertations--Computer Science
The advance of high-throughput sequencing technologies and their application on mRNA transcriptome sequencing (RNA-seq) have enabled comprehensive and unbiased profiling of the landscape of transcription in a cell. In order to address the current limitation of analyzing accuracy and scalability in transcriptome analysis, a novel computational framework has been developed on large-scale RNA-seq datasets with no dependence on transcript annotations. Directly from raw reads, a probabilistic approach is first applied to infer the best transcript fragment alignments from paired-end reads. Empowered by the identification of alternative splicing modules, this framework then performs precise and efficient differential analysis at automatically detected …
A Rule Induction Algorithm For Knowledge Discovery And Classification, Ömer Akgöbek
A Rule Induction Algorithm For Knowledge Discovery And Classification, Ömer Akgöbek
Turkish Journal of Electrical Engineering and Computer Sciences
Classification and rule induction are key topics in the fields of decision making and knowledge discovery. The objective of this study is to present a new algorithm developed for automatic knowledge acquisition in data mining. The proposed algorithm has been named RES-2 (Rule Extraction System). It aims at eliminating the pitfalls and disadvantages of the techniques and algorithms currently in use. The proposed algorithm makes use of the direct rule extraction approach, rather than the decision tree. For this purpose, it uses a set of examples to induce general rules. In this study, 15 datasets consisting of multiclass values with …
On Identifying Critical Nuggets Of Information During Classification Task, David Sathiaraj
On Identifying Critical Nuggets Of Information During Classification Task, David Sathiaraj
LSU Doctoral Dissertations
In large databases, there may exist critical nuggets - small collections of records or instances that contain domain-specific important information. This information can be used for future decision making such as labeling of critical, unlabeled data records and improving classification results by reducing false positive and false negative errors. In recent years, data mining efforts have focussed on pattern and outlier detection methods. However, not much effort has been dedicated to finding critical nuggets within a data set. This work introduces the idea of critical nuggets, proposes an innovative domain-independent method to measure criticality, suggests a heuristic to reduce the …
Exploring The Learnability Of Numeric Datasets, Di Lin
Exploring The Learnability Of Numeric Datasets, Di Lin
LSU Doctoral Dissertations
When doing classification, it has often been observed that datasets may exhibit different levels of difficulty with respect to how accurately they can be classified. That is, there are some datasets which can be classified very accurately by many classification algorithms, and there also exist some other datasets that no classifier can classify them with high accuracy. Based on this observation, we try to address the following problems: a)what are the factors that make a dataset easy or difficult to be accurately classified? b) how to use such factors to predict the difficulties of unclassified datasets? and c) how to …
Human-Readable Real-Time Classifications Of Malicious Executables, Anselm Teh, Arran Stewart
Human-Readable Real-Time Classifications Of Malicious Executables, Anselm Teh, Arran Stewart
Australian Information Security Management Conference
Shafiq et al. (2009a) propose a non–signature-based technique for detecting malware which applies data mining techniques to features extracted from executable files. Their technique has a high level of accuracy, a low false positive rate, and a speed on par with commercial anti-virus products. One portion of their technique uses a multi-layer perceptron as a classifier, which provides little insight into the reasons for classification. Our experience is that network security analysts prefer tools which provide human-comprehensible reasons for a classification, rather than operating as “black boxes”. We therefore build on the results of Shafiq et al. by demonstrating a …
Data Mining Of Pancreatic Cancer Protein Databases, Peter Revesz, Christopher Assi
Data Mining Of Pancreatic Cancer Protein Databases, Peter Revesz, Christopher Assi
School of Computing: Conference and Workshop Papers
Data mining of protein databases poses special challenges because many protein databases are non- relational whereas most data mining and machine learning algorithms assume the input data to be a type of rela- tional database that is also representable as an ARFF file. We developed a method to restructure protein databases so that they become amenable for various data mining and machine learning tools. Our restructuring method en- abled us to apply both decision tree and support vector machine classifiers to a pancreatic protein database. The SVM classifier that used both GO term and PFAM families to characterize proteins gave …
Adaptive Grid Based Localized Learning For Multidimensional Data, Sheetal Saini
Adaptive Grid Based Localized Learning For Multidimensional Data, Sheetal Saini
Doctoral Dissertations
Rapid advances in data-rich domains of science, technology, and business has amplified the computational challenges of "Big Data" synthesis necessary to slow the widening gap between the rate at which the data is being collected and analyzed for knowledge. This has led to the renewed need for efficient and accurate algorithms, framework, and algorithmic mechanisms essential for knowledge discovery, especially in the domains of clustering, classification, dimensionality reduction, feature ranking, and feature selection. However, data mining algorithms are frequently challenged by the sparseness due to the high dimensionality of the datasets in such domains which is particularly detrimental to the …
Building A Computer Program To Support Children, Parents, And Distraction During Healthcare Procedures, Kirsten Hanrahan, Ann Marie Mccarthy, Charmaine Kleiber, Kaan Ataman, W. Nick Street, M. Bridget Zimmerman, Annel L. Ersig
Building A Computer Program To Support Children, Parents, And Distraction During Healthcare Procedures, Kirsten Hanrahan, Ann Marie Mccarthy, Charmaine Kleiber, Kaan Ataman, W. Nick Street, M. Bridget Zimmerman, Annel L. Ersig
Business Faculty Articles and Research
This secondary data analysis used data mining methods to develop predictive models of child risk for distress during a healthcare procedure. Data used came from a study that predicted factors associated with children's responses to an intravenous catheter insertion while parents provided distraction coaching. From the 255 items used in the primary study, 44 predictive items were identified through automatic feature selection and used to build support vector machine regression models. Models were validated using multiple cross-validation tests and by comparing variables identified as explanatory in the traditional versus support vector machine regression. Rule-based approaches were applied to the model …
Semi-Automatic Simulation Initialization By Mining Structured And Unstructured Data Formats From Local And Web Data Sources, Olcay Sahin
Computational Modeling & Simulation Engineering Theses & Dissertations
Initialization is one of the most important processes for obtaining successful results from a simulation. However, initialization is a challenge when 1) a simulation requires hundreds or even thousands of input parameters or 2) re-initializing the simulation due to different initial conditions or runtime errors. These challenges lead to the modeler spending more time initializing a simulation and may lead to errors due to poor input data.
This thesis proposes two semi-automatic simulation initialization approaches that provide initialization using data mining from structured and unstructured data formats from local and web data sources. First, the System Initialization with Retrieval (SIR) …
A Confidence-Prioritization Approach To Data Processing In Noisy Data Sets And Resulting Estimation Models For Predicting Streamflow Diel Signals In The Pacific Northwest, Nathaniel Lee Gustafson
A Confidence-Prioritization Approach To Data Processing In Noisy Data Sets And Resulting Estimation Models For Predicting Streamflow Diel Signals In The Pacific Northwest, Nathaniel Lee Gustafson
Theses and Dissertations
Streams in small watersheds are often known to exhibit diel fluctuations, in which streamflow oscillates on a 24-hour cycle. Streamflow diel fluctuations, which we investigate in this study, are an informative indicator of environmental processes. However, in Environmental Data sets, as well as many others, there is a range of noise associated with individual data points. Some points are extracted under relatively clear and defined conditions, while others may include a range of known or unknown confounding factors, which may decrease those points' validity. These points may or may not remain useful for training, depending on how much uncertainty they …
From Clickstreams To Searchstreams: Search Network Graph Evidence From A B2b E-Market, Mei Lin, M. F. Lin, Robert J. Kauffman
From Clickstreams To Searchstreams: Search Network Graph Evidence From A B2b E-Market, Mei Lin, M. F. Lin, Robert J. Kauffman
Research Collection School Of Computing and Information Systems
Consumers in e-commerce acquire information through search engines, yet to date there has been little empirical study on how users interact with the results produced by search engines. This is analogous to, but different from, the ever-expanding research on clickstreams, where users interact with static web pages. We propose a new network approach to analyzing search engine server log data. We call this searchstream data. We create graph representations based on the web pages that users traverse as they explore the search results that their use of search engines generates. We then analyze the graph-level properties of these search network …
Data Mining Of Protein Databases, Christopher Assi
Data Mining Of Protein Databases, Christopher Assi
School of Computing: Dissertations, Theses, and Student Research
Data mining of protein databases poses special challenges because many protein databases are non-relational whereas most data mining and machine learning algorithms assume the input data to be a relational database. Protein databases are non-relational mainly because they often contain set data types. We developed new data mining algorithms that can restructure non-relational protein databases so that they become relational and amenable for various data mining and machine learning tools. We applied the new restructuring algorithms to a pancreatic protein database. After the restructuring, we also applied two classification methods, such as decision tree and SVM classifiers and compared their …
Mining Input Sanitization Patterns For Predicting Sql Injection And Cross Site Scripting Vulnerabilities, Lwin Khin Shar, Hee Beng Kuan Tan
Mining Input Sanitization Patterns For Predicting Sql Injection And Cross Site Scripting Vulnerabilities, Lwin Khin Shar, Hee Beng Kuan Tan
Research Collection School Of Computing and Information Systems
Static code attributes such as lines of code and cyclomatic complexity have been shown to be useful indicators of defects in software modules. As web applications adopt input sanitization routines to prevent web security risks, static code attributes that represent the characteristics of these routines may be useful for predicting web application vulnerabilities. In this paper, we classify various input sanitization methods into different types and propose a set of static code attributes that represent these types. Then we use data mining methods to predict SQL injection and cross site scripting vulnerabilities in web applications. Preliminary experiments show that our …
Data Mining Of Tetraloop-Tetraloop Receptors In Rna Xml Files, Sinan Ramazanoglu
Data Mining Of Tetraloop-Tetraloop Receptors In Rna Xml Files, Sinan Ramazanoglu
Theses
RNA (Ribonucleic acid) Motifs are tertiary structures that play an important role in the folding mechanism of the RNA molecule. The overall function of a RNA Motif depends on its specific bp (base pairs) sequence that constitutes the secondary structure. Data mining is a novel method in both discovering potential tertiary structures within DNA (Deoxyribonucleic acid), RNA, and protein molecules and storing the information in databases. The RNA Motif of interest is the tetraloop-tetraloop receptor, which is composed of a highly conserved 11 nt (nucleotide) sequence and a tetraloop with the generic form of GNRA (where N = any base …
Ensemble Of Feature Selection Techniques For High Dimensional Data, Sri Harsha Vege
Ensemble Of Feature Selection Techniques For High Dimensional Data, Sri Harsha Vege
Masters Theses & Specialist Projects
Data mining involves the use of data analysis tools to discover previously unknown, valid patterns and relationships from large amounts of data stored in databases, data warehouses, or other information repositories. Feature selection is an important preprocessing step of data mining that helps increase the predictive performance of a model. The main aim of feature selection is to choose a subset of features with high predictive information and eliminate irrelevant features with little or no predictive information. Using a single feature selection technique may generate local optima.
In this thesis we propose an ensemble approach for feature selection, where multiple …
Analysis And Characterization Of Author Contribution Patterns In Open Source Software Development, Quinn Carlson Taylor
Analysis And Characterization Of Author Contribution Patterns In Open Source Software Development, Quinn Carlson Taylor
Theses and Dissertations
Software development is a process fraught with unpredictability, in part because software is created by people. Human interactions add complexity to development processes, and collaborative development can become a liability if not properly understood and managed. Recent years have seen an increase in the use of data mining techniques on publicly-available repository data with the goal of improving software development processes, and by extension, software quality. In this thesis, we introduce the concept of author entropy as a metric for quantifying interaction and collaboration (both within individual files and across projects), present results from two empirical observational studies of open-source …
Medical Data Analysis Method For Epilepsy, Ameen Eetemadi
Medical Data Analysis Method For Epilepsy, Ameen Eetemadi
Wayne State University Theses
Applying data mining techniques on medical databases which contain un-structured and semi-structured data is a challenging task. It is not only due to the complexity of such databases but also due to the characteristics of the medical domain. This thesis describes how multiple layers of data mining techniques have been applied to a Human Brain Image Database system. It starts with data preparation which paves the way for conventional data analysis techniques to be applied to the data. A similarity based patient retrieval tool has been designed and developed to assist in treatment planning and outcome estimation for epileptic patients. …