Open Access. Powered by Scholars. Published by Universities.®
Databases and Information Systems Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Numerical Analysis and Scientific Computing (25)
- Social and Behavioral Sciences (18)
- Business (15)
- Theory and Algorithms (9)
- Communication (8)
-
- Engineering (7)
- Management Information Systems (7)
- Computer Engineering (6)
- Social Media (6)
- Statistics and Probability (6)
- Medicine and Health Sciences (5)
- Software Engineering (5)
- Data Science (4)
- Data Storage Systems (4)
- Life Sciences (4)
- Artificial Intelligence and Robotics (3)
- Communication Technology and New Media (3)
- Health Information Technology (3)
- Information Security (3)
- Medical Specialties (3)
- Public Affairs, Public Policy and Public Administration (3)
- Transportation (3)
- Bioinformatics (2)
- Computational Linguistics (2)
- E-Commerce (2)
- Education (2)
- Library and Information Science (2)
- Institution
-
- Singapore Management University (42)
- Portland State University (5)
- University of Nebraska - Lincoln (5)
- Institute of Business Administration (4)
- New Jersey Institute of Technology (4)
-
- Claremont Colleges (3)
- Western Kentucky University (3)
- California State University, San Bernardino (2)
- City University of New York (CUNY) (2)
- Embry-Riddle Aeronautical University (2)
- Kennesaw State University (2)
- Chapman University (1)
- Dartmouth College (1)
- Edith Cowan University (1)
- Georgia Southern University (1)
- Kutztown University (1)
- Louisiana Tech University (1)
- Old Dominion University (1)
- Purdue University (1)
- Technological University Dublin (1)
- University of Arkansas Little Rock (1)
- University of Arkansas, Fayetteville (1)
- University of Missouri, St. Louis (1)
- University of Nevada, Las Vegas (1)
- University of South Florida (1)
- Ursinus College (1)
- Wright State University (1)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (39)
- Complex Systems Faculty Publications and Presentations (4)
- Dissertations (4)
- International Conference on Information and Communication Technologies (4)
- CGU Faculty Publications and Research (3)
-
- Dissertations and Theses Collection (2)
- Faculty Articles (2)
- Masters Theses & Specialist Projects (2)
- Theses Digitization Project (2)
- Theses and Dissertations (2)
- Beyond: Undergraduate Research Journal (1)
- Business Faculty Articles and Research (1)
- College of Graduate Studies: Theses & Dissertations (1)
- Computer Science Faculty Publications (1)
- Computer Science Faculty Publications and Presentations (1)
- Computer Science Summer Fellows (1)
- Computer Science and Information Technology Faculty (1)
- Dartmouth Scholarship (1)
- Department of Agricultural Economics: Dissertations, Theses, and Student Research (1)
- Dissertations and Theses Collection (Open Access) (1)
- Doctoral (1)
- Doctoral Dissertations (1)
- Economics Faculty Publications (1)
- Graduate Theses and Dissertations (1)
- Kno.e.sis Publications (1)
- Open Access Theses (1)
- Publications (1)
- Publications and Research (1)
- Research outputs 2014 to 2021 (1)
- School of Computing: Conference and Workshop Papers (1)
- Publication Type
Articles 61 - 90 of 90
Full-Text Articles in Databases and Information Systems
Artificial Intelligence – I: Adaptive Automated Teller Machines — Part I, Ghulam Mujtaba, Tariq Mahmood
Artificial Intelligence – I: Adaptive Automated Teller Machines — Part I, Ghulam Mujtaba, Tariq Mahmood
International Conference on Information and Communication Technologies
During the past few years, the banking sector has started providing a variety of services to its customers. One of the most significant of such services has been the introduction of the Automated Teller Machines (ATMs) for providing online support to bank customers. The use of ATMs has reached its zenith in every developed country, and thousands of ATM transactions are occurring on a daily basis. In order to increase the customers' satisfaction and to provide them with more user-friendly ATM interfaces, it becomes important to mine the ATM transactions to discover useful patterns about the customers' interacting behaviors. In …
Empirical Methods For Predicting Student Retention- A Summary From The Literature, Matt Bogard
Empirical Methods For Predicting Student Retention- A Summary From The Literature, Matt Bogard
Economics Faculty Publications
The vast majority of the literature related to the empirical estimation of retention models includes a discussion of the theoretical retention framework established by Bean, Braxton, Tinto, Pascarella, Terenzini and others (see Bean, 1980; Bean, 2000; Braxton, 2000; Braxton et al, 2004; Chapman and Pascarella, 1983; Pascarell and Ternzini, 1978; St. John and Cabrera, 2000; Tinto, 1975) This body of research provides a starting point for the consideration of which explanatory variables to include in any model specification, as well as identifying possible data sources. The literature separates itself into two major camps including research related to the hypothesis testing …
Efficient Schema Extraction From A Collection Of Xml Documents, Vijayeandra Parthepan
Efficient Schema Extraction From A Collection Of Xml Documents, Vijayeandra Parthepan
Masters Theses & Specialist Projects
The eXtensible Markup Language (XML) has become the standard format for data exchange on the Internet, providing interoperability between different business applications. Such wide use results in large volumes of heterogeneous XML data, i.e., XML documents conforming to different schemas. Although schemas are important in many business applications, they are often missing in XML documents. In this thesis, we present a suite of algorithms that are effective in extracting schema information from a large collection of XML documents. We propose using the cost of NFA simulation to compute the Minimum Length Description to rank the inferred schema. We also studied …
An Approach To Nearest Neighboring Search For Multi-Dimensional Data, Yong Shi, Li Zhang, Lei Zhu
An Approach To Nearest Neighboring Search For Multi-Dimensional Data, Yong Shi, Li Zhang, Lei Zhu
Faculty Articles
Finding nearest neighbors in large multi-dimensional data has always been one of the research interests in data mining field. In this paper, we present our continuous research on similarity search problems. Previously we have worked on exploring the meaning of K nearest neighbors from a new perspective in PanKNN [20]. It redefines the distances between data points and a given query point Q, efficiently and effectively selecting data points which are closest to Q. It can be applied in various data mining fields. A large amount of real data sets have irrelevant or obstacle information which greatly affects the effectiveness …
Combining Natural Language Processing And Statistical Text Mining: A Study Of Specialized Versus Common Languages, Jay Jarman
USF Tampa Graduate Theses and Dissertations
This dissertation focuses on developing and evaluating hybrid approaches for analyzing free-form text in the medical domain. This research draws on natural language processing (NLP) techniques that are used to parse and extract concepts based on a controlled vocabulary. Once important concepts are extracted, additional machine learning algorithms, such as association rule mining and decision tree induction, are used to discover classification rules for specific targets. This multi-stage pipeline approach is contrasted with traditional statistical text mining (STM) methods based on term counts and term-by-document frequencies. The aim is to create effective text analytic processes by adapting and combining individual …
Stevent: Spatio-Temporal Event Model For Social Network Discovery, Hady W. Lauw, Ee Peng Lim, Hwee Hwa Pang, Teck-Tim Tan
Stevent: Spatio-Temporal Event Model For Social Network Discovery, Hady W. Lauw, Ee Peng Lim, Hwee Hwa Pang, Teck-Tim Tan
Research Collection School Of Computing and Information Systems
Spatio-temporal data concerning the movement of individuals over space and time contains latent information on the associations among these individuals. Sources of spatio-temporal data include usage logs of mobile and Internet technologies. This article defines a spatio-temporal event by the co-occurrences among individuals that indicate potential associations among them. Each spatio-temporal event is assigned a weight based on the precision and uniqueness of the event. By aggregating the weights of events relating two individuals, we can determine the strength of association between them. We conduct extensive experimentation to investigate both the efficacy of the proposed model as well as the …
Effects Of Similarity Metrics On Document Clustering, Rushikesh Veni
Effects Of Similarity Metrics On Document Clustering, Rushikesh Veni
UNLV Theses, Dissertations, Professional Papers, and Capstones
Document clustering or unsupervised document classification is an automated process of grouping documents with similar content. A typical technique uses a similarity function to compare documents. In the literature, many similarity functions such as dot product or cosine measures are proposed for the comparison operator.
For the thesis, we evaluate the effects a similarity function may have on clustering. We start by representing a document and a query, both as a vector of high-dimensional space corresponding to the keywords followed by using an appropriate distance measure in k-means to compute similarity between the document vector and the query vector to …
The Impact Of Directionality In Predications On Text Mining, Gondy Leroy, Marcelo Fiszman, Thomas C. Rindflesch
The Impact Of Directionality In Predications On Text Mining, Gondy Leroy, Marcelo Fiszman, Thomas C. Rindflesch
CGU Faculty Publications and Research
The number of publications in biomedicine is increasing enormously each year. To help researchers digest the information in these documents, text mining tools are being developed that present co-occurrence relations between concepts. Statistical measures are used to mine interesting subsets of relations. We demonstrate how directionality of these relations affects interestingness. Support and confidence, simple data mining statistics, are used as proxies for interestingness metrics. We first built a test bed of 126,404 directional relations extracted from biomedical abstracts, which we represent as graphs containing a central starting concept and 2 rings of associated relations. We manipulated directionality in four …
Mobile Semantic Computing, Karthik Gomadam, Anupam Joshi, Amit P. Sheth
Mobile Semantic Computing, Karthik Gomadam, Anupam Joshi, Amit P. Sheth
Kno.e.sis Publications
We propose to organize a special session on research in the intersection of mobile computing, the Semantic Web and Web services.
This session will examine how the research in these areas can serve as a foundation for new architectural and communication paradigms that can enhance service creation, distribution, discovery, integration and utilization in distributed and ubiquitous environments. Some of the initial areas that our early research have highlighted are :
- Semantic annotation of data in bandwidth constrained environments such as mobile networks to promote efficient bandwidth utilization
- Possibilities of using microformats such as RDFa and opportunities that can be explored …
Bias And Controversy: Beyond The Statistical Deviation, Hady W. Lauw, Ee Peng Lim, Ke Wang
Bias And Controversy: Beyond The Statistical Deviation, Hady W. Lauw, Ee Peng Lim, Ke Wang
Research Collection School Of Computing and Information Systems
In this paper, we investigate how deviation in evaluation activities may reveal bias on the part of reviewers and controversy on the part of evaluated objects. We focus on a 'data-centric approach' where the evaluation data is assumed to represent the ground truth'. The standard statistical approaches take evaluation and deviation at face value. We argue that attention should be paid to the subjectivity of evaluation, judging the evaluation score not just on 'what is being said' (deviation), but also on 'who says it' (reviewer) as well as on 'whom it is said about' (object). Furthermore, we observe that bias …
Sgpm: Static Group Pattern Mining Using Apriori-Like Sliding Window, John Goh, David Taniar, Ee Peng Lim
Sgpm: Static Group Pattern Mining Using Apriori-Like Sliding Window, John Goh, David Taniar, Ee Peng Lim
Research Collection School Of Computing and Information Systems
Mobile user data mining is a field that focuses on extracting interesting pattern and knowledge out from data generated by mobile users. Group pattern is a type of mobile user data mining method. In group pattern mining, group patterns from a given user movement database is found based on spatio-temporal distances. In this paper, we propose an improvement of efficiency using area method for locating mobile users and using sliding window for static group pattern mining. This reduces the complexity of valid group pattern mining problem. We support the use of static method, which uses areas and sliding windows instead …
Fisa: Feature-Based Instance Selection For Imbalanced Text Classification, Aixin Sun, Ee Peng Lim, Boualem Benatallah, Mahbub Hassan
Fisa: Feature-Based Instance Selection For Imbalanced Text Classification, Aixin Sun, Ee Peng Lim, Boualem Benatallah, Mahbub Hassan
Research Collection School Of Computing and Information Systems
Support Vector Machines (SVM) classifiers are widely used in text classification tasks and these tasks often involve imbalanced training. In this paper, we specifically address the cases where negative training documents significantly outnumber the positive ones. A generic algorithm known as FISA (Feature-based Instance Selection Algorithm), is proposed to select only a subset of negative training documents for training a SVM classifier. With a smaller carefully selected training set, a SVM classifier can be more efficiently trained while delivering comparable or better classification accuracy. In our experiments on the 20-Newsgroups dataset, using only 35% negative training examples and 60% learning …
Text Mining With Exploitation Of User's Background Knowledge : Discovering Novel Association Rules From Text, Xin Chen
Dissertations
The goal of text mining is to find interesting and non-trivial patterns or knowledge from unstructured documents. Both objective and subjective measures have been proposed in the literature to evaluate the interestingness of discovered patterns. However, objective measures alone are insufficient because such measures do not consider knowledge and interests of the users. Subjective measures require explicit input of user expectations which is difficult or even impossible to obtain in text mining environments.
This study proposes a user-oriented text-mining framework and applies it to the problem of discovering novel association rules from documents. The developed system, uMining, consists of two …
Data Mining Techniques To Study Therapy Success With Autistic Children, Gondy A. Leroy, Annika Irmscher, Marjorie H. Charlop
Data Mining Techniques To Study Therapy Success With Autistic Children, Gondy A. Leroy, Annika Irmscher, Marjorie H. Charlop
CGU Faculty Publications and Research
Autism spectrum disorder has become one of the most prevalent developmental disorders, characterized by a wide variety of symptoms. Many children need extensive therapy for years to improve their behavior and facilitate integration in society. However, few systematic evaluations are done on a large scale that can provide insights into how, where, and how therapy has an impact. We describe how data mining techniques can be used to provide insights into behavioral therapy as well as its effect on participants. To this end, we are developing a digital library of coded video segments that contains data on appropriate and inappropriate …
Keynote: The Use Of Meta-Heuristic Algorithms For Data Mining, Dr. Beatrize De La Iglesia, A. Reynolds
Keynote: The Use Of Meta-Heuristic Algorithms For Data Mining, Dr. Beatrize De La Iglesia, A. Reynolds
International Conference on Information and Communication Technologies
In this paper we explore the application of powerful optimisers known as metaheuristic algorithms to problems within the data mining domain. We introduce some well-known data mining problems, and show how they can be formulated as optimisation problems. We then review the use of metaheuristics in this context. In particular, we focus on the task of partial classification and show how multi-objective metaheuristics have produced results that are comparable to the best known techniques but more scalable to large databases. We conclude by reinforcing the importance of research on the areas of metaheuristics for optimisation and data mining. The combination …
A Dynamic Weight Assignment Approach For Ir Systems, M. Shoaib, Prof Dr. Abad Ali Shah, A. Vashishta
A Dynamic Weight Assignment Approach For Ir Systems, M. Shoaib, Prof Dr. Abad Ali Shah, A. Vashishta
International Conference on Information and Communication Technologies
Weights are assigned to the extracted keywords for partial matching and computing ranking in an IR system. Weight assignment technique is suggested by the IR model that is used for an IR system. Currently suggested weight assignment techniques are static which means that once weight is assigned a keyword it remains unchanged during life-span of an IR system. In this paper, we suggest a dynamic weight assignment technique. This technique can be used by any IR model that supports partial matching.
Social Network Discovery By Mining Spatio-Temporal Events, Hady Lauw, Ee Peng Lim, Hwee Hwa Pang, Teck-Tim Tan
Social Network Discovery By Mining Spatio-Temporal Events, Hady Lauw, Ee Peng Lim, Hwee Hwa Pang, Teck-Tim Tan
Research Collection School Of Computing and Information Systems
Knowing patterns of relationship in a social network is very useful for law enforcement agencies to investigate collaborations among criminals, for businesses to exploit relationships to sell products, or for individuals who wish to network with others. After all, it is not just what you know, but also whom you know, that matters. However, finding out who is related to whom on a large scale is a complex problem. Asking every single individual would be impractical, given the huge number of individuals and the changing dynamics of relationships. Recent advancement in technology has allowed more data about activities of individuals …
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Blocking Reduction Strategies In Hierarchical Text Classification, Ee Peng Lim, Aixin Sun, Wee-Keong Ng, Jaideep Srivastava
Research Collection School Of Computing and Information Systems
One common approach in hierarchical text classification involves associating classifiers with nodes in the category tree and classifying text documents in a top-down manner. Classification methods using this top-down approach can scale well and cope with changes to the category trees. However, all these methods suffer from blocking which refers to documents wrongly rejected by the classifiers at higher-levels and cannot be passed to the classifiers at lower-levels. We propose a classifier-centric performance measure known as blocking factor to determine the extent of the blocking. Three methods are proposed to address the blocking problem, namely, threshold reduction, restricted voting, and …
A Support-Ordered Trie For Fast Frequent Itemset Discovery, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
A Support-Ordered Trie For Fast Frequent Itemset Discovery, Ee Peng Lim, Yew-Kwong Woon, Wee-Keong Ng
Research Collection School Of Computing and Information Systems
The importance of data mining is apparent with the advent of powerful data collection and storage tools; raw data is so abundant that manual analysis is no longer possible. Unfortunately, data mining problems are difficult to solve and this prompted the introduction of several novel data structures to improve mining efficiency. Here, we critically examine existing preprocessing data structures used in association rule mining for enhancing performance in an attempt to understand their strengths and weaknesses. Our analyses culminate in a practical structure called the SOTrielT (support-ordered trie itemset) and two synergistic algorithms to accompany it for the fast discovery …
Customer Relationship Management For Banking System, Pingyu Hou
Customer Relationship Management For Banking System, Pingyu Hou
Theses Digitization Project
The purpose of this project is to design, build, and implement a Customer Relationship Management (CRM) system for a bank. CRM BANKING is an online application that caters to strengthening and stabilizing customer relationships in a bank.
Reconstructability Analysis With Fourier Transforms, Martin Zwick
Reconstructability Analysis With Fourier Transforms, Martin Zwick
Complex Systems Faculty Publications and Presentations
Fourier methods used in two‐ and three‐dimensional image reconstruction can be used also in reconstructability analysis (RA). These methods maximize a variance‐type measure instead of information‐theoretic uncertainty, but the two measures are roughly collinear and the Fourier approach yields results close to that of standard RA. The Fourier method, however, does not require iterative calculations for models with loops. Moreover, the error in Fourier RA models can be assessed without actually generating the full probability distributions of the models; calculations scale with the size of the data rather than the state space. State‐based modeling using the Fourier approach is also …
Directed Extended Dependency Analysis For Data Mining, Thaddeus T. Shannon, Martin Zwick
Directed Extended Dependency Analysis For Data Mining, Thaddeus T. Shannon, Martin Zwick
Complex Systems Faculty Publications and Presentations
Extended dependency analysis (EDA) is a heuristic search technique for finding significant relationships between nominal variables in large data sets. The directed version of EDA searches for maximally predictive sets of independent variables with respect to a target dependent variable. The original implementation of EDA was an extension of reconstructability analysis. Our new implementation adds a variety of statistical significance tests at each decision point that allow the user to tailor the algorithm to a particular objective. It also utilizes data structures appropriate for the sparse data sets customary in contemporary data mining problems. Two examples that illustrate different approaches …
An Overview Of Reconstructability Analysis, Martin Zwick
An Overview Of Reconstructability Analysis, Martin Zwick
Complex Systems Faculty Publications and Presentations
This paper is an overview of reconstructability analysis (RA), a discrete multivariate modeling methodology developed in the systems literature; an earlier version of this tutorial is Zwick (2001). RA was derived from Ashby (1964), and was developed by Broekstra, Cavallo, Cellier Conant, Jones, Klir, Krippendorff, and others (Klir, 1986, 1996). RA resembles and partially overlaps log‐line (LL) statistical methods used in the social sciences (Bishop et al., 1978; Knoke and Burke, 1980). RA also resembles and overlaps methods used in logic design and machine learning (LDL) in electrical and computer engineering (e.g. Perkowski et al., 1997). Applications of RA, like …
Genescene: Biomedical Text And Data Mining, Gondy Leroy, Hsinchun Chen, Jesse D. Martinez, Shauna Eggers, Ryan R. Falsey, Kerri L. Kislin, Zan Huang, Jiexun Li, Jie Xu, Daniel M. Mcdonald, Gavin Ng
Genescene: Biomedical Text And Data Mining, Gondy Leroy, Hsinchun Chen, Jesse D. Martinez, Shauna Eggers, Ryan R. Falsey, Kerri L. Kislin, Zan Huang, Jiexun Li, Jie Xu, Daniel M. Mcdonald, Gavin Ng
CGU Faculty Publications and Research
To access the content of digital texts efficiently, it is necessary to provide more sophisticated access than keyword based searching. GeneScene provides biomedical researchers with research findings and background relations automatically extracted from text and experimental data. These provide a more detailed overview of the information available. The extracted relations were evaluated by qualified researchers and are precise. A qualitative ongoing evaluation of the current online interface indicates that this method to search the literature is more useful and efficient than keyword based searching.
Health Care Informatics, Keng Siau
Health Care Informatics, Keng Siau
Research Collection School Of Computing and Information Systems
The health care industry is currently experiencing a fundamental change. Health care organizations are reorganizing their processes to reduce costs, be more competitive, and provide better and more personalized customer care. This new business strategy requires health care organizations to implement new technologies, such as Internet applications, enterprise systems, and mobile technologies in order to achieve their desired business changes. This article offers a conceptual model for implementing new information systems, integrating internal data, and linking suppliers and patients.
Data Warehouse Applications In Modern Day Business, Carla Mounir Issa
Data Warehouse Applications In Modern Day Business, Carla Mounir Issa
Theses Digitization Project
Data warehousing provides organizations with strategic tools to achieve the competitive advantage that organazations are constantly seeking. The use of tools such as data mining, indexing and summaries enables management to retrieve information and perform thorough analysis, planning and forcasting to meet the changes in the market environment. in addition, The data warehouse is providing security measures that, if properly implemented and planned, are helping organizations ensure that their data quality and validity remain intact.
A Review Of Data Mining Techniques, Sang Jun Lee, Keng Siau
A Review Of Data Mining Techniques, Sang Jun Lee, Keng Siau
Research Collection School Of Computing and Information Systems
Terabytes of data are generated everyday in many organizations. To extract hidden predictive information from large volumes of data, data mining (DM) techniques are needed. Organizations are starting to realize the importance of data mining in their strategic planning and successful application of DM techniques can be an enormous payoff for the organizations. This paper discusses the requirements and challenges of DM, and describes major DM techniques such as statistics, artificial intelligence, decision tree approach, genetic algorithm, and visualization.
Predictive Self-Organizing Networks For Text Categorization, Ah-Hwee Tan
Predictive Self-Organizing Networks For Text Categorization, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
This paper introduces a class of predictive self-organizing neural networks known as Adaptive Resonance Associative Map (ARAM) for classification of free-text documents. Whereas most sta- tistical approaches to text categorization derive classification knowledge based on training examples alone, ARAM performs supervised learn- ing and integrates user-defined classification knowledge in the form of IF-THEN rules. Through our experiments on the Reuters-21578 news database, we showed that ARAM performed reasonably well in mining categorization knowledge from sparse and high dimensional document feature space. In addition, ARAM predictive accuracy and learning efficiency can be improved by incorporating a set of rules derived from …
Knowledge Discovery As An Aid To Organizational Creativity, Keng Siau
Knowledge Discovery As An Aid To Organizational Creativity, Keng Siau
Research Collection School Of Computing and Information Systems
Computers can play an important role in the creative process. With the abundance of data and increasing speed of computers, creativity can now be stimulated and enhanced with knowledge mined from available data. The process is known as knowledge discovery or data mining. Knowledge discovery is the process of discovering interesting associations among data in the database. Users in the creativity process can feed on the discovered associations to generate creative solutions. The objective of this paper is to present knowledge discovery as an aid to creativity. The paper first presents the concept of knowledge discovery and then discusses the …
Knowledge Discovery In Biological Databases : A Neural Network Approach, Qicheng Ma
Knowledge Discovery In Biological Databases : A Neural Network Approach, Qicheng Ma
Dissertations
Knowledge discovery, in databases, also known as data mining, is aimed to find significant information from a set of data. The knowledge to be mined from the dataset may refer to patterns, association rules, classification and clustering rules, and so forth. In this dissertation, we present a neural network approach to finding knowledge in biological databases. Specifically, we propose new methods to process biological sequences in two case studies: the classification of protein sequences and the prediction of E. Coli promoters in DNA sequences. Our proposed methods, based oil neural network architectures combine techniques ranging from Bayesian inference, coding theory, …