Open Access. Powered by Scholars. Published by Universities.®
- Discipline
- Institution
Articles 1 - 4 of 4
Full-Text Articles in Computational Linguistics
Computational Linguistics For Metadata Building: Aggregating Text Processing Technologies For Enhanced Image Access, Judith Klavans, Carolyn Sheffield, Eileen Abels, Joan E. Beaudoin, Laura Jenemann, Jimmy Lin, Tom Lippincott, Rebecca Passonneau, Tandeep Sidhu, Dagobert Soergel, Tae Yano
Computational Linguistics For Metadata Building: Aggregating Text Processing Technologies For Enhanced Image Access, Judith Klavans, Carolyn Sheffield, Eileen Abels, Joan E. Beaudoin, Laura Jenemann, Jimmy Lin, Tom Lippincott, Rebecca Passonneau, Tandeep Sidhu, Dagobert Soergel, Tae Yano
School of Information Sciences Faculty Research Publications
We present a system which applies text mining using computational linguistic techniques to automatically extract, categorize, disambiguate and filter metadata for image access. Candidate subject terms are identified through standard approaches; novel semantic categorization using machine learning and disambiguation using both WordNet and a domain specific thesaurus are applied. The resulting metadata can be manually edited by image catalogers or filtered by semi-automatic rules. We describe the implementation of this workbench created for, and evaluated by, image catalogers. We discuss the system's current functionality, developed under the Computational Linguistics for Metadata Building (CLiMB) research project. The CLiMB Toolkit has been …
The Impact Of Directionality In Predications On Text Mining, Gondy Leroy, Marcelo Fiszman, Thomas C. Rindflesch
The Impact Of Directionality In Predications On Text Mining, Gondy Leroy, Marcelo Fiszman, Thomas C. Rindflesch
CGU Faculty Publications and Research
The number of publications in biomedicine is increasing enormously each year. To help researchers digest the information in these documents, text mining tools are being developed that present co-occurrence relations between concepts. Statistical measures are used to mine interesting subsets of relations. We demonstrate how directionality of these relations affects interestingness. Support and confidence, simple data mining statistics, are used as proxies for interestingness metrics. We first built a test bed of 126,404 directional relations extracted from biomedical abstracts, which we represent as graphs containing a central starting concept and 2 rings of associated relations. We manipulated directionality in four …
Tagset Design, Inflected Languages, And N-Gram Tagging, Anna Feldman
Tagset Design, Inflected Languages, And N-Gram Tagging, Anna Feldman
Department of Linguistics Faculty Scholarship and Creative Works
This paper explores the relationship between the tagset design and linguistic properties of inflected languages for the task of morphosyntactic tagging. Some information theoretic measures and statistics on these languages are reported which show, unsurprisingly, that the tagsets for morphologically rich languages are larger than tagsets for English and the average tag/token ambiguity is higher. The surprising outcome of the experiments is that for Catalan, Czech, Polish, Portuguese,and Russian – which are considered to be “word order” free languages (to various degrees) – the knowledge about the preceding tag reduces the uncertainty about the tag in question if the detailed …
Referring Expression Generation Challenge 2008 Dit System Descriptions (Dit-Fbi, Dit-Tvas, Dit-Cbsr, Dit-Rbr, Dit-Fbi-Cbsr, Dit-Tvas-Rbr), John D. Kelleher, Brian Mac Namee
Referring Expression Generation Challenge 2008 Dit System Descriptions (Dit-Fbi, Dit-Tvas, Dit-Cbsr, Dit-Rbr, Dit-Fbi-Cbsr, Dit-Tvas-Rbr), John D. Kelleher, Brian Mac Namee
Conference papers
This papers desibes a set of systems developed at DIT for the Referring Expression Generation challenage at INLG 2008.In Proceedings of the 5th International Natural Language Generation Conference (INLG-08)