Open Access. Powered by Scholars. Published by Universities.®

Databases and Information Systems Commons™

Open Access. Powered by Scholars. Published by Universities.®

Information retrieval

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 45 of 45

Full-Text Articles in Databases and Information Systems

Fisa: Feature-Based Instance Selection For Imbalanced Text Classification, Aixin Sun, Ee Peng Lim, Boualem Benatallah, Mahbub Hassan Apr 2006

Fisa: Feature-Based Instance Selection For Imbalanced Text Classification, Aixin Sun, Ee Peng Lim, Boualem Benatallah, Mahbub Hassan

Research Collection School Of Computing and Information Systems

Support Vector Machines (SVM) classifiers are widely used in text classification tasks and these tasks often involve imbalanced training. In this paper, we specifically address the cases where negative training documents significantly outnumber the positive ones. A generic algorithm known as FISA (Feature-based Instance Selection Algorithm), is proposed to select only a subset of negative training documents for training a SVM classifier. With a smaller carefully selected training set, a SVM classifier can be more efficiently trained while delivering comparable or better classification accuracy. In our experiments on the 20-Newsgroups dataset, using only 35% negative training examples and 60% learning …


Opems: Online Peer-To-Peer Expertise Matching System, Sharifullah Khan, S.M. Nabeel Aug 2005

Opems: Online Peer-To-Peer Expertise Matching System, Sharifullah Khan, S.M. Nabeel

International Conference on Information and Communication Technologies

Internet is a vital source to disseminate and share information to the masses. This has made information available in abundance on the Web. However, finding relevant information is difficult if not impossible. This difficulty is bilateral between information providers and seekers in terms of information presentation and accessibility respectively. This paper proposed an online Peer-to-Peer Expertise Matching system. The approach provides a highly scalable and self-organizing system and helps individuals in presenting and accessing the information in a consistent format on the Web. This makes the sharing of information among the autonomous organizations successful.


Visualization Of Retrieved Positive Data Using Blending Function, Muhammad Shoaib, Habib -Ur- Rehman, Dr. Abad Ali Shah Aug 2005

Visualization Of Retrieved Positive Data Using Blending Function, Muhammad Shoaib, Habib -Ur- Rehman, Dr. Abad Ali Shah

International Conference on Information and Communication Technologies

Data visualization is an important technique used in data mining. We present the retrieved data into visual format to discover features and trends inherent to the data. Some features of the data to be retrieved are already known to us. Visualization should preserve these known features inherent to the data. Positivity is one such known feature that is inherent to most of the scientific and business data sets. For example, mass, volume and percentage concentration are meaningful only when they are positive values. However certain visualization techniques do not guarantee to preserve this feature while constructing visualization of retrieved data …


Improving Document Representation By Accumulating Relevance Feedback : The Relevance Feedback Accumulation (Rfa) Algorithm, Razvan Stefan Bot May 2005

Improving Document Representation By Accumulating Relevance Feedback : The Relevance Feedback Accumulation (Rfa) Algorithm, Razvan Stefan Bot

Dissertations

Document representation (indexing) techniques are dominated by variants of the term-frequency analysis approach, based on the assumption that the more occurrences a term has throughout a document the more important the term is in that document. Inherent drawbacks associated with this approach include: poor index quality, high document representation size and the word mismatch problem. To tackle these drawbacks, a document representation improvement method called the Relevance Feedback Accumulation (RFA) algorithm is presented. The algorithm provides a mechanism to continuously accumulate relevance assessments over time and across users. It also provides a document representation modification function, or document representation learning …


Integrating User Feedback Log Into Relevance Feedback By Coupled Svm For Content-Based Image Retrieval, Steven C. H. Hoi, Michael R. Lyu, Rong Jin Apr 2005

Integrating User Feedback Log Into Relevance Feedback By Coupled Svm For Content-Based Image Retrieval, Steven C. H. Hoi, Michael R. Lyu, Rong Jin

Research Collection School Of Computing and Information Systems

Relevance feedback has been shown as an important tool to boost the retrieval performance in content-based image retrieval. In the past decade, various algorithms have been proposed to formulate relevance feedback in contentbased image retrieval. Traditional relevance feedback techniques mainly carry out the learning tasks by focusing lowlevel visual features of image content with little consideration on log information of user feedback. However, from a long-term learning perspective, the user feedback log is one of the most important resources to bridge the semantic gap problem in image retrieval. In this paper we propose a novel technique to integrate the log …


An Exploration Of Cultural Factors Affecting Use Of Communities Of Practice, Peter L. Hinrichsen Mar 2004

An Exploration Of Cultural Factors Affecting Use Of Communities Of Practice, Peter L. Hinrichsen

Theses and Dissertations

On-line communities of practice are potentially powerful social learning networks that can improve organizational performance. Unfortunately, administrators of on-line communities of practice report that community members do not take full advantage of this potential. This study used Shaw and Tuggle's (2003) factors of knowledge management (KM) culture affecting organizational acceptance of a knowledge management initiative to explore this issue. It was hypothesized that respondents whose communities of practice possessed higher average community use per member would rate KM culture variables higher than respondents whose communities possessed a lower average community use. An analysis of survey data collected from Air Force …


Genescene: Biomedical Text And Data Mining, Gondy Leroy, Hsinchun Chen, Jesse D. Martinez, Shauna Eggers, Ryan R. Falsey, Kerri L. Kislin, Zan Huang, Jiexun Li, Jie Xu, Daniel M. Mcdonald, Gavin Ng May 2003

Genescene: Biomedical Text And Data Mining, Gondy Leroy, Hsinchun Chen, Jesse D. Martinez, Shauna Eggers, Ryan R. Falsey, Kerri L. Kislin, Zan Huang, Jiexun Li, Jie Xu, Daniel M. Mcdonald, Gavin Ng

CGU Faculty Publications and Research

To access the content of digital texts efficiently, it is necessary to provide more sophisticated access than keyword based searching. GeneScene provides biomedical researchers with research findings and background relations automatically extracted from text and experimental data. These provide a more detailed overview of the information available. The extracted relations were evaluated by qualified researchers and are precise. A qualitative ongoing evaluation of the current online interface indicates that this method to search the literature is more useful and efficient than keyword based searching.


Federating Heterogeneous Digital Libraries By Metadata Harvesting, Xiaoming Liu Jan 2002

Federating Heterogeneous Digital Libraries By Metadata Harvesting, Xiaoming Liu

Computer Science Theses & Dissertations

This dissertation studies the challenges and issues faced in federating heterogeneous digital libraries (DLs) by metadata harvesting. The objective of federation is to provide high-level services (e.g. transparent search across all DLs) on the collective metadata from different digital libraries. There are two main approaches to federate DLs: distributed searching approach and harvesting approach. As the distributed searching approach replies on executing queries to digital libraries in real time, it has problems with scalability. The difficulty of creating a distributed searching service for a large federation is the motivation behind Open Archives Initiatives Protocols for Metadata Harvesting (OAI-PMH). OAI-PMH supports …


The Partial Evaluation Approach To Information Personalization, Naren Ramakrishnan, Saverio Perugini Jan 2001

The Partial Evaluation Approach To Information Personalization, Naren Ramakrishnan, Saverio Perugini

Computer Science Faculty Publications

Information personalization refers to the automatic adjustment of information content, structure, and presentation tailored to an individual user. By reducing information overload and customizing information access, personalization systems have emerged as an important segment of the Internet economy. This paper presents a systematic modeling methodology— PIPE (‘Personalization is Partial Evaluation’) — for personalization. Personalization systems are designed and implemented in PIPE by modeling an information-seeking interaction in a programmatic representation. The representation supports the description of information-seeking activities as partial information and their subsequent realization by partial evaluation, a technique for specializing programs. We describe the modeling methodology at a …


A Web Based Healthcare Management System, Farid Abdullah Muhammad Jan 2001

A Web Based Healthcare Management System, Farid Abdullah Muhammad

Student Works (2000-2009)

The project generally involved around the basic information about medical, health, drugs and diseases. This kind of information is vital yet has a very limited access. The tradition that when we get sick then we go to see doctor and ate whatever medicine that the doctor gave well, does not apply here anymore. People should know at least a simple thing about their health or medical status. They should know what they eat, what cause the illness or how to prevent it. This is possible as the system target on medical student and the public to make the system useful …


Integrated Student Management Information System, Subramaniam Melissa Malathy Jan 2001

Integrated Student Management Information System, Subramaniam Melissa Malathy

Student Works (2000-2009)

ISMIS is a an integrated system designed to provide students information in order to facilitate schools administration, decision-making and monitoring of students' development in curricular and co-curriculum activities. Specially for teachers and school administrators, the purpose of this project is to develop a student information system that contains the information of the students in a school and provide a system that will enable the retrieval of these information that include students' personal particulars, academic and non-academic achievements as well as performance records and all other items pertaining to student matters. The development of ISMIS is divided into three sections, which …


Financial Information System (Fis), Chien Wee Kuan Jan 2001

Financial Information System (Fis), Chien Wee Kuan

Student Works (2000-2009)

Financial Information System for the Faculty of Computer Science and Information Technology (FCSIT-FIS) is a clien/server system introduced as a sub-system in the e-Faculty project. Its main purpose is to create a paperless, web-based system by providing on-line functions that allow users to access to FCSIT financial information. The main functions of FCSIT-FIS are: on-line spending request, financial information retrieval and account maintenance. FCSIT -FIS will be developing using ASP technology, VBScript, JavaScript and HTML language. Microsoft SQL Server will be use as the database server of the system.


The Personal Library And Indexing Management Information System (Plimis), Wai Kit Choong Jan 2001

The Personal Library And Indexing Management Information System (Plimis), Wai Kit Choong

Student Works (2000-2009)

The Personal Library and Indexing Management Information System (PLIMIS) is an information system used to organize and manage personal collections of reading materials such as books, journals, journal articles, magazines, articles, dictionaries and CD-ROMs. The system allows easy searching and retrieving of the items in the collections as it adopts a systematic classification and indexing for the reading materials. This will save the time in searching a specific item. PLIMIS is designed for personal use and contains most of the features in the current library information systems. It provides user friendly interface and help to the users for easy using …


On Integrating Existing Bibliographic Databases And Structured Databases, Ying Lu, Ee Peng Lim Aug 1996

On Integrating Existing Bibliographic Databases And Structured Databases, Ying Lu, Ee Peng Lim

Research Collection School Of Computing and Information Systems

It is widely accepted that future digital library applications have to be built upon different kinds of database servers to draw different forms of data from them. These data include bibliographic data, text data, multimedia data, and structured data. We address the problem of integrating existing bibliographic and structured databases which reside at different locations in the network. To integrate bibliographic data and structured data, we extended the well-known SQL model to represent bibliographic related attributes and queries. In particular, we have added a new data type to model attributes in the bibliographic database. We have also designed specialized predicates …


The Effect Of Domain Knowledge On Elementary School Children's Search Behavior On An Information Retrieval System: The Science Library Catalog, Sandra Hirsh May 1995

The Effect Of Domain Knowledge On Elementary School Children's Search Behavior On An Information Retrieval System: The Science Library Catalog, Sandra Hirsh

Faculty Publications

Few information retrieval systems are designed with children’s special needs and capabilities in mind. We need to learn more about children’s information-seeking behavior in order to provide them with information-based tools which support exploratory learning. This dissertation examines children’s search behavior on a hypertext-based automated library catalog designed for elementary school children. The focus of this research is on the effect of domain knowledge on children’s search performance, search behavior, and learning as they look for science books on this system. Reseaxch has shown that level of domain knowledge in~luences the way people search for information. Data was collected through …