Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 211 - 240 of 250

Full-Text Articles in Computational Linguistics

Using Textual Features To Predict Popular Content On Digg, Paul H. Miller Apr 2011

Using Textual Features To Predict Popular Content On Digg, Paul H. Miller

Department of English: Dissertations, Theses, and Student Research

Over the past few years, collaborative rating sites, such as Netflix, Digg and Stumble, have become increasingly prevalent sites for users to find trending content.  I used various data mining techniques to study Digg, a social news site, to examine the influence of content on popularity.  What influence does content have on popularity, and what influence does content have on users’ decisions?  Overwhelmingly, prior studies have consistently shown that predicting popularity based on content is difficult and maybe even inherently impossible.  The same submission can have multiple outcomes and content neither determines popularity, nor individual user decisions.  My results show …


Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. Mcdonald, Jerry Schnepp Jan 2011

Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. Mcdonald, Jerry Schnepp

Visual Communications and Technology Education Faculty Publications

Computer-generated three-dimensional animation holds great promise for synthesizing utterances in American Sign Language (ASL) that are not only grammatical, but well tolerated by members of the Deaf community. Unfortunately, animation poses several challenges stemming from the necessity of grappling with massive amounts of data. However, the linguistics of ASL can aid in surmounting the challenge by providing structure and rules for organizing animation data. An exploration of the linguistic and extra linguistic behavior of the brows from an animator’s viewpoint yields a new approach for synthesizing nonmanuals that differs from the conventional animation of anatomy and instead offers a different …


Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner Jan 2011

Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner

Department of Linguistics Faculty Scholarship and Creative Works

The Penn Forced Aligner automates the alignment process using the Hidden Markov Model Toolkit (HTK). The core of Prosodylab-Aligner is align.py, a script which performs acoustic model training and alignment. This script automates calls to HTK and SoX, an open-source command-line tool which is capable of resampling audio. The included README file provides instructions for installing HTK and SoX on Linux and Mac OS X, and can also be run on Windows. During training, the model is initialized with flat-start monophones, which are then submitted to a single round of model estimation. Then, a tied-state 'small pause' model is inserted …


Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman Dec 2010

Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

Text classification is a widely studied topic in the area of machine learning. A number of techniques have been developed to represent and classify text documents. Most of the techniques try to achieve good classification performance while taking a document only by its words (e.g. statistical analysis on word frequency and distribution patterns). One of the recent trends in text classification research is to incorporate more semantic interpretation in text classification, especially by using Wikipedia. This paper introduces a technique for incorporating the vast amount of human knowledge accumulated in Wikipedia into text representation and classification. The aim is to …


Study Of Stemming Algorithms, Savitha Kodimala Dec 2010

Study Of Stemming Algorithms, Savitha Kodimala

UNLV Theses, Dissertations, Professional Papers, and Capstones

Automated stemming is the process of reducing words to their roots. The stemmed words are typically used to overcome the mismatch problems associated with text searching.


In this thesis, we report on the various methods developed for stemming. In particular, we show the effectiveness of n-gram stemming methods on a collection of documents.


Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee Nov 2010

Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee

Conference papers

Dialogues between humans and robots are necessarily situated and so, often, a shared visual context is present. Exophoric references are very frequent in situated dialogues, and are particularly important in the presence of a shared visual context - for example when a human is verbally guiding a tele-operated mobile robot. We present an approach to automatically resolving exophoric referring expressions in a situated dialogue based on the visual salience of possible referents. We evaluate the effectiveness of this approach and a range of different salience metrics using data from the SCARE corpus which we have augmented with visual information. The …


Situating Spatial Templates For Human-Robot Interaction, John D. Kelleher, Robert J. Ross, Brian Mac Namee, Colm Sloan Nov 2010

Situating Spatial Templates For Human-Robot Interaction, John D. Kelleher, Robert J. Ross, Brian Mac Namee, Colm Sloan

Conference papers

People often refer to objects by describing the object's spatial location relative to another object. Due to their ubiquity in situated discourse, the ability to use 'locative expressions' is fundamental to human-robot dialogue systems. A key component of this ability are computational models of spatial term semantics. These models bridge the grounding gap between spatial language and sensor data. Within the Artificial Intelligence and Robotics communities, spatial template based accounts, such as the Attention Vector Sum model (Regier and Carlson, 2001), have found considerable application in mediating situated human-machine communication (Gorniak, 2004; Brenner et a., 2007; Kelleher and Costello, 2009). …


New Trends In Automatic Assessment: Ontology Matching, Maria Mitina, Patricia Magee, John Cardiff Oct 2010

New Trends In Automatic Assessment: Ontology Matching, Maria Mitina, Patricia Magee, John Cardiff

Conference Papers

Instant individual feedback represents a result of assessment which allows for considerable improvements in both teaching and learning. In this paper we present the application of ontology matching techniques in automatic correction of students’ answers for SQL tests, which will provide teachers with instant feedback to facilitate manual correction and marking and which they can pass to the students. Students experience many problems learning SQL due to the necessity to memorise database schemas, unclear feedback from the database engine on the execution of the query, etc. The program environment utilising the described approach is designed to solve the abovementioned problems …


Topology In Composite Spatial Terms, John D. Kelleher, Robert J. Ross Aug 2010

Topology In Composite Spatial Terms, John D. Kelleher, Robert J. Ross

Conference papers

People often refer to objects by describing the object's spatial location relative to another object, e.g. the book on the right of the table. This type of referring expression is called a spatial locative expression. Spatial locatives have three major components: (1) the target object that is being located (the book), (2) the landmark object relative to which the target is being located (the table), and (3) the description of the spatial relationship that exists between the target and the landmark (on the right of ). In English spatial relationships are often described using spatial prepositions. The set of English …


Proceedings Of The Sixth International Natural Language Generation Conference (Inlg 2010)., John D. Kelleher, Brian Mac Namee, Ielka Van Der Sluis Jul 2010

Proceedings Of The Sixth International Natural Language Generation Conference (Inlg 2010)., John D. Kelleher, Brian Mac Namee, Ielka Van Der Sluis

Conference papers

No abstract provided.


Personal Sense And Idiolect: Combining Authorship Attribution And Opinion Analysis, Polina Panicheva, John Cardiff, Paolo Rosso May 2010

Personal Sense And Idiolect: Combining Authorship Attribution And Opinion Analysis, Polina Panicheva, John Cardiff, Paolo Rosso

Conference Papers

Subjectivity analysis and authorship attribution are very popular areas of research. However, work in these two areas has been done separately. We believe that by combining information about subjectivity in texts and authorship, the performance of both tasks can be improved. In the paper a personalized approach to opinion mining is presented, in which the notions of personal sense and idiolect are introduced and used for polarity classification task. The results of applying the personalized approach to opinion mining are presented, confirming that the approach increases the performance of the opinion mining task. Automatic authorship attribution is further applied to …


A Computer-Aided Error Analysis Of Multi-Word Units In A Malaysian Learner English Corpus., Lee Yit Sim Jan 2010

A Computer-Aided Error Analysis Of Multi-Word Units In A Malaysian Learner English Corpus., Lee Yit Sim

Student Works (2010-2019)

This is a study on multi-word unit (MWU) errors made by Malaysian English learners. The major approaches used in the analysis of learners’ writings are: error analysis and corpus-based approach. The findings from this study reveal that learners have problems with these MWUs: modal structures , infinitive structures , ‘adjective + noun’ collocations , and connectors . An analysis of these structures shows that the common erroneous patterns in these structures can generally be categorized as ‘overgeneralization’, ‘misformation’, ‘misselection’, and ‘distortion’. The causes of such errors are both interlingual as well as intralingual. Besides these factors, this study also highlighted …


Applying Computational Models Of Spatial Prepositions To Visually Situated Dialog, John D. Kelleher, Fintan Costello Jun 2009

Applying Computational Models Of Spatial Prepositions To Visually Situated Dialog, John D. Kelleher, Fintan Costello

Articles

This article describes the application of computational models of spatial prepositions to visually situated dialog systems. In these dialogs, spatial prepositions are important because people often use them to refer to entities in the visual context of a dialog. We first describe a generic architecture for a visually situated dialog system and highlight the interactions between the spatial cognition module, which provides the interface to the models of prepositional semantics, and the other components in the architecture. Following this, we present two new computational models of topological and projective spatial prepositions. The main novelty within these models is the fact …


Computational Linguistics For Metadata Building: Aggregating Text Processing Technologies For Enhanced Image Access, Judith Klavans, Carolyn Sheffield, Eileen Abels, Joan E. Beaudoin, Laura Jenemann, Jimmy Lin, Tom Lippincott, Rebecca Passonneau, Tandeep Sidhu, Dagobert Soergel, Tae Yano Aug 2008

Computational Linguistics For Metadata Building: Aggregating Text Processing Technologies For Enhanced Image Access, Judith Klavans, Carolyn Sheffield, Eileen Abels, Joan E. Beaudoin, Laura Jenemann, Jimmy Lin, Tom Lippincott, Rebecca Passonneau, Tandeep Sidhu, Dagobert Soergel, Tae Yano

School of Information Sciences Faculty Research Publications

We present a system which applies text mining using computational linguistic techniques to automatically extract, categorize, disambiguate and filter metadata for image access. Candidate subject terms are identified through standard approaches; novel semantic categorization using machine learning and disambiguation using both WordNet and a domain specific thesaurus are applied. The resulting metadata can be manually edited by image catalogers or filtered by semi-automatic rules. We describe the implementation of this workbench created for, and evaluated by, image catalogers. We discuss the system's current functionality, developed under the Computational Linguistics for Metadata Building (CLiMB) research project. The CLiMB Toolkit has been …


The Impact Of Directionality In Predications On Text Mining, Gondy Leroy, Marcelo Fiszman, Thomas C. Rindflesch Jan 2008

The Impact Of Directionality In Predications On Text Mining, Gondy Leroy, Marcelo Fiszman, Thomas C. Rindflesch

CGU Faculty Publications and Research

The number of publications in biomedicine is increasing enormously each year. To help researchers digest the information in these documents, text mining tools are being developed that present co-occurrence relations between concepts. Statistical measures are used to mine interesting subsets of relations. We demonstrate how directionality of these relations affects interestingness. Support and confidence, simple data mining statistics, are used as proxies for interestingness metrics. We first built a test bed of 126,404 directional relations extracted from biomedical abstracts, which we represent as graphs containing a central starting concept and 2 rings of associated relations. We manipulated directionality in four …


Tagset Design, Inflected Languages, And N-Gram Tagging, Anna Feldman Jan 2008

Tagset Design, Inflected Languages, And N-Gram Tagging, Anna Feldman

Department of Linguistics Faculty Scholarship and Creative Works

This paper explores the relationship between the tagset design and linguistic properties of inflected languages for the task of morphosyntactic tagging. Some information theoretic measures and statistics on these languages are reported which show, unsurprisingly, that the tagsets for morphologically rich languages are larger than tagsets for English and the average tag/token ambiguity is higher. The surprising outcome of the experiments is that for Catalan, Czech, Polish, Portuguese,and Russian – which are considered to be “word order” free languages (to various degrees) – the knowledge about the preceding tag reduces the uncertainty about the tag in question if the detailed …


Referring Expression Generation Challenge 2008 Dit System Descriptions (Dit-Fbi, Dit-Tvas, Dit-Cbsr, Dit-Rbr, Dit-Fbi-Cbsr, Dit-Tvas-Rbr), John D. Kelleher, Brian Mac Namee Jan 2008

Referring Expression Generation Challenge 2008 Dit System Descriptions (Dit-Fbi, Dit-Tvas, Dit-Cbsr, Dit-Rbr, Dit-Fbi-Cbsr, Dit-Tvas-Rbr), John D. Kelleher, Brian Mac Namee

Conference papers

This papers desibes a set of systems developed at DIT for the Referring Expression Generation challenage at INLG 2008.In Proceedings of the 5th International Natural Language Generation Conference (INLG-08)


Statistical Machine Translation Of Japanese, Erik A. Chapla Mar 2007

Statistical Machine Translation Of Japanese, Erik A. Chapla

Theses and Dissertations

The purpose of this research was to find ways to improve the performance of a statistical machine translation system that translates text from Japanese to English. Methods included altering the training and test data by adding a prior linguistic knowledge, altering sentence structures, and looking for better ways to statistically alter the way words align between the two languages. In addition, methods for properly segmenting words in Japanese text through statistical methods were examined. Finally, experiments were conducted on Japanese speech to produce the best text transcription of the speech. The best statistical machine translation methods implemented resulted in improvements …


A Classifier To Evaluate Language Specificity In Medical Documents, Trudi Miller '08, Gondy A. Leroy, Samir Chatterjee, Jie Fan, Brian Thoms '09 Jan 2007

A Classifier To Evaluate Language Specificity In Medical Documents, Trudi Miller '08, Gondy A. Leroy, Samir Chatterjee, Jie Fan, Brian Thoms '09

CGU Faculty Publications and Research

Consumer health information written by health care professionals is often inaccessible to the consumers it is written for. Traditional readability formulas examine syntactic features like sentence length and number of syllables, ignoring the target audience's grasp of the words themselves. The use of specialized vocabulary disrupts the understanding of patients with low reading skills, causing a decrease in comprehension. A naive Bayes classifier for three levels of increasing medical terminology specificity (consumer/patient, novice health learner, medical professional) was created with a lexicon generated from a representative medical corpus. Ninety-six percent accuracy in classification was attained. The classifier was then applied …


Frequency Based Incremental Attribute Selection For Gre., John D. Kelleher Jan 2007

Frequency Based Incremental Attribute Selection For Gre., John D. Kelleher

Conference papers

The DIT system uses an incremental greedy search to generate descriptions, similar to the incremental algorithm described in (Dale and Reiter, 1995). The selection of the next attribute to be tested for inclusion in the description is ordered by the absolute frequency of each attribute in the training corpus. Attributes are selected in descending order of frequency (i.e. the attribute that occurred most frequently in the training corpus is selected first). Where two or more attributes have the same frequency of occurrence the first attribute found with that frequency is selected. The type attribute is always included in the description. …


Proceedings Of The 4th Acl-Sigsem Workshop On Prepositions At Acl-2007., Fintan Costello, John D. Kelleher, Martin Volk Jan 2007

Proceedings Of The 4th Acl-Sigsem Workshop On Prepositions At Acl-2007., Fintan Costello, John D. Kelleher, Martin Volk

Conference papers

This volume contains the papers presented at the Fourth ACL-SIGSEM Workshop on Prepositions. This workshop is endorsed by the ACL Special Interest Group on Semantics (ACL-SIGSEM), and is hosted in conjunction with ACL 2007, taking place on 28th June, 2007 in Prague, the Czech Republic.


Active Learning For Part-Of-Speech Tagging: Accelerating Corpus Annotation, Deryle W. Lonsdale, Eric K. Ringger, Peter J. Mcclanahan, Robbie A. Haertel, George Busby, Marc A. Carmen, James Carroll, Kevin Seppi Jan 2007

Active Learning For Part-Of-Speech Tagging: Accelerating Corpus Annotation, Deryle W. Lonsdale, Eric K. Ringger, Peter J. Mcclanahan, Robbie A. Haertel, George Busby, Marc A. Carmen, James Carroll, Kevin Seppi

Faculty Publications

In the construction of a part-of-speech annotated corpus, we are constrained by a fixed budget. A fully annotated corpus is required, but we can afford to label only a subset. We train a Maximum Entropy Markov Model tagger from a labeled subset and automatically tag the remainder. This paper addresses the question of where to focus our manual tagging efforts in order to deliver an annotation of highest quality. In this context, we find that active learning is always helpful. We focus on Query by Uncertainty (QBU) and Query by Committee (QBC) and report on experiments with several baselines and …


Multilingual Phoneme Models For Rapid Speech Processing System Development, Eric G. Hansen Sep 2006

Multilingual Phoneme Models For Rapid Speech Processing System Development, Eric G. Hansen

Theses and Dissertations

Current speech recognition systems tend to be developed only for commercially viable languages. The resources needed for a typical speech recognition system include hundreds of hours of transcribed speech for acoustic models and 10 to 100 million words of text for language models; both of these requirements can be costly in time and money. The goal of this research is to facilitate rapid development of speech systems to new languages by using multilingual phoneme models to alleviate requirements for large amounts of transcribed speech. The Global Phone database, winch contains transcribed speech from 15 languages, is used as source data …


Xnl-Soar, Incremental Parsing, And The Minimalist Program, Deryle W. Lonsdale, Lareina Hingson, Jamison Cooper-Leavitt, David W. Casbeer, Rebecca Madsen Mar 2006

Xnl-Soar, Incremental Parsing, And The Minimalist Program, Deryle W. Lonsdale, Lareina Hingson, Jamison Cooper-Leavitt, David W. Casbeer, Rebecca Madsen

Faculty Publications

Minimalist Principles (Chomsky 1995)

Hierarchy of Projections (Adger 2003)

Features play a central role

NP, VP symmetry including shells


Speech Recognition Using The Mellin Transform, Jesse R. Hornback Mar 2006

Speech Recognition Using The Mellin Transform, Jesse R. Hornback

Theses and Dissertations

The purpose of this research was to improve performance in speech recognition. Specifically, a new approach was investigating by applying an integral transform known as the Mellin transform (MT) on the output of an auditory model to improve the recognition rate of phonemes through the scale-invariance property of the Mellin transform. Scale-invariance means that as a time-domain signal is subjected to dilations, the distribution of the signal in the MT domain remains unaffected. An auditory model was used to transform speech waveforms into images representing how the brain "sees" a sound. The MT was applied and features were extracted. The …


Incremental Generation Of Spatial Referring Expressions In Situated Dialogue, John D. Kelleher, Geert-Jan Kruijff Jan 2006

Incremental Generation Of Spatial Referring Expressions In Situated Dialogue, John D. Kelleher, Geert-Jan Kruijff

Conference papers

This paper presents an approach to incrementally generating locative expressions. It addresses the issue of combinatorial explosion inherent in the construction of relational context models by: (a) contextually defining the set of objects in the context that may function as a landmark, and (b) sequencing the order in which spatial relations are considered using a cognitively motivated hierarchy of relations, and visual and discourse salience.


A Computational Model Of The Referential Semantics Of Projective Prepositions, John D. Kelleher, Josef Van Genabith Jan 2006

A Computational Model Of The Referential Semantics Of Projective Prepositions, John D. Kelleher, Josef Van Genabith

Conference papers

In this paper we present a framework for interpreting locative expressions containing the prepositions in front of and behind. These prepositions have different semantics in the viewer-centred and intrinsic frames of reference (Vandeloise, 1991). We define a model of their semantics in each frame of reference. The basis of these models is a novel parameterized continuum function that creates a 3-D spatial template. In the intrinsic frame of reference the origin used by the continuum function is assumed to be known a priori and object occlusion does not impact on the applicability rating of a point in the spatial template. …


Proximity In Context: An Empirically Grounded Computational Model Of Proximity For Processing Topological Spatial Expression., John D. Kelleher, Geert-Jan Kruijff, Fintan Costello Jan 2006

Proximity In Context: An Empirically Grounded Computational Model Of Proximity For Processing Topological Spatial Expression., John D. Kelleher, Geert-Jan Kruijff, Fintan Costello

Conference papers

The paper presents a new model for context-dependent interpretation of linguistic expressions about spatial proximity between objects in a natural scene. The paper discusses novel psycholinguistic experimental data that tests and verifies the model. The model has been implemented, and enables a conversational robot to identify objects in a scene through topological spatial relations (e.g. ''X near Y''). The model can help motivate the choice between topological and projective prepositions.


An Operator-Based Account Of Semantic Processing, Deryle W. Lonsdale, C. Anton Rytting Jan 2006

An Operator-Based Account Of Semantic Processing, Deryle W. Lonsdale, C. Anton Rytting

Faculty Publications

This paper explores issues of psychological plausibility in modeling natural language understanding within Soar, a symbolic cognitive model. It focuses on constructing syntactic and semantic representations in simulated real time, with particular emphasis on word sense disambiguation (WSD). We discuss (i) what level of WSD should be modeled and (ii) how to use resources such as WordNet to inform these models. A preliminary model of coarse-grained WSD is included to show how syntactic, semantic, and other knowledge sources interact in Soar. Finally, we explore issues of interleaving, learning, and integrating other WSD approaches with Soar's native model of learning.


Automatic Creation Of Web Services From Extraction Ontologies, Deryle W. Lonsdale, Cui Tao, Yihong Ding Jan 2006

Automatic Creation Of Web Services From Extraction Ontologies, Deryle W. Lonsdale, Cui Tao, Yihong Ding

Faculty Publications

The Semantic Web promises to provide timely, targeted access to user-specified information online. Though standardized services exist for performing this work, specifying these services is too complex for most people. Annotating these services is also problematic. A similar situation exists for traditional information extraction, where ontologies are increasingly used to specify information used by various extraction methods. The approach we introduce in this paper involves converting such ontologies into executable Java code. These APIs act individually or compositionally as services for Semantic Web extraction.