Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

250 Full-Text Articles 420 Authors 150,666 Downloads 65 Institutions

All Articles in Computational Linguistics

Faceted Search

250 full-text articles. Page 11 of 12.

Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman 2014 Charles University, Prague

Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman

Department of Linguistics Faculty Scholarship and Creative Works

The paper describes a corpus of texts produced by non-native speakersof Czech. We discuss its annotation scheme, consisting of three interlinked tiers,designed to handle a wide range of error types present in the input. Each tier correctsdifferent types of errors; links between the tiers allow capturing errors in word orderand complex discontinuous expressions. Errors are not only corrected, but alsoclassified. The annotation scheme is tested on a data set including approx. 175,000words with fair inter-annotator agreement results. We also explore the possibility ofapplying automated linguistic annotation tools (taggers, spell checkers and grammarcheckers) to the learner text to support or even …


Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova 2014 Montclair State University

Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova

Department of Computer Science Faculty Scholarship and Creative Works

We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that words in a given text segment, such as a paragraph, that are highranking representatives of a common topic of discussion are less likely to be a part of an idiomatic expression. Our additional hypothesis is that contexts in which idioms occur, typically, are more affective and therefore, we incorporate a simple analysis of the intensity of the emotions expressed by the contexts. We investigate the bag of words topic representation of one to three paragraphs containing an expression that should be classified as …


Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee 2014 Technological University Dublin

Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee

Articles

In a situated dialogue, misunderstandings may arise if the participants perceive or interpret the environment in different ways. In human-computer dialogue this may be due the sensor errors. We present an experiment system and a series of experiments in which we investigate this problem.


Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale 2014 University of Kentucky

Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale

Theses and Dissertations--Linguistics

In Paradigm Function Morphology, it is usual to model affix position classes with an ordered sequence of inflectional rule blocks. Each rule block determines how (or whether) a particular affix position is filled. In this model, competition among inflectional rules is assumed to be limited to members of the same rule block; thus, the appearance of an affix in one position cannot be precluded by the appearance of an affix in another position. I present evidence that apparently disconfirms this restriction and suggests that a more general conception of rule competition is necessary. The data appear to imply that an …


Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes 2013 California Polytechnic State University, San Luis Obispo

Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes

Master's Theses

In the field of speech recognition, an algorithm must learn to tell the difference between "a nice rock" and "a gneiss rock". These identical-sounding phrases are called oronyms. Word frequency dictionaries are often used by speech recognition systems to help resolve phonetic sequences with more than one possible orthographic phrase interpretation, by looking up which oronym of the root phonetic sequence contains the most-common words.

Our paper demonstrates a technique used to validate word frequency dictionary values. We chose to use frequency values from the UNISYN dictionary, which tallies each word on a per-occurance basis, using a proprietary text corpus, …


Csc Senior Project: Nlpstats, Michael Mease 2013 California Polytechnic State University - San Luis Obispo

Csc Senior Project: Nlpstats, Michael Mease

Computer Science and Software Engineering

Natural Language Processing has recently increased in popularity. The field of authorship analysis, specifically, uses various characteristics of text quantified by markers. NLPStats serves as a tool designed to streamline marker extraction based on user needs. A flexible query system allows for custom marker requests, adjustment of result formatting, and preprocessing options. Furthermore, an efficiently designed structure ensures that users retrieve information quickly. As a whole, NLPStats enables anyone, regardless of NLP experience, to extract important information about the text of a document.


Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana 2013 Montclair State University

Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana

Department of Computer Science Faculty Scholarship and Creative Works

The goal of this study is to investigate whether learners’ written data in highly inflectional Czech can suggest a consistent set of clues for automatic identification of the learners’ L1 background. For our experiments, we use texts written by learners of Czech, which have been automatically and manually annotated for errors. We define two classes of learners: speakers of Indo-European languages and speakers of non-Indo-European languages. We use an SVM classifier to perform the binary classification. We show that non-content based features perform well on highly inflectional data. In particular, features reflecting errors in orthography are the most useful, yielding …


Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez 2013 CUNY New York City College of Technology

Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez

Publications and Research

The aim of this paper is to compare the rhetorical functions gathered from the citations of (14) fourteen master´s theses written by seven Spanish and seven Philippine authors. A typology of nine categories was used in order to identify the cultural rhetorical differences that exist in the use of citation from the contrast between contrasting this element in the Philippine and Spanish cultures. The methodology used is textual analysis of the linguistic context of these citations and its subsequent classification within these nine categories. The results show that there are quantitative and qualitative differences between the cultural conventions of citations …


Lexicalization And De-Lexicalization Processes In Sign Languages: Comparing Depicting Constructions And Viewpoint Gestures, Kearsy Cormier, David Quinto-Pozos, Zed Sehyr, Adam Schembri 2012 University College London

Lexicalization And De-Lexicalization Processes In Sign Languages: Comparing Depicting Constructions And Viewpoint Gestures, Kearsy Cormier, David Quinto-Pozos, Zed Sehyr, Adam Schembri

Communication Sciences and Disorders Faculty Articles and Research

In this paper, we compare so-called “classifier” constructions in signed languages (which we refer to as “depicting constructions”) with comparable iconic gestures produced by non-signers. We show clear correspondences between entity constructions and observer viewpoint gestures on the one hand, and handling constructions and character viewpoint gestures on the other. Such correspondences help account for both lexicalisation and de-lexicalisation processes in signed languages and how these processes are influenced by viewpoint. Understanding these processes is crucial when coding and annotating natural sign language data.


Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier 2012 CUNY York College

Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier

Publications and Research

We investigate how novel English-derived words (anglicisms) are used in a German-language Internet hip hop forum, and what factors contribute to their uptake.


Using Textual Features To Predict Popular Content On Digg, Paul H. Miller 2011 University of Nebraska-Lincoln

Using Textual Features To Predict Popular Content On Digg, Paul H. Miller

Department of English: Dissertations, Theses, and Student Research

Over the past few years, collaborative rating sites, such as Netflix, Digg and Stumble, have become increasingly prevalent sites for users to find trending content.  I used various data mining techniques to study Digg, a social news site, to examine the influence of content on popularity.  What influence does content have on popularity, and what influence does content have on users’ decisions?  Overwhelmingly, prior studies have consistently shown that predicting popularity based on content is difficult and maybe even inherently impossible.  The same submission can have multiple outcomes and content neither determines popularity, nor individual user decisions.  My results show …


Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. McDonald, Jerry Schnepp 2011 DePaul University

Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. Mcdonald, Jerry Schnepp

Visual Communications and Technology Education Faculty Publications

Computer-generated three-dimensional animation holds great promise for synthesizing utterances in American Sign Language (ASL) that are not only grammatical, but well tolerated by members of the Deaf community. Unfortunately, animation poses several challenges stemming from the necessity of grappling with massive amounts of data. However, the linguistics of ASL can aid in surmounting the challenge by providing structure and rules for organizing animation data. An exploration of the linguistic and extra linguistic behavior of the brows from an animator’s viewpoint yields a new approach for synthesizing nonmanuals that differs from the conventional animation of anatomy and instead offers a different …


Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner 2011 University of Pennsylvania

Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner

Department of Linguistics Faculty Scholarship and Creative Works

The Penn Forced Aligner automates the alignment process using the Hidden Markov Model Toolkit (HTK). The core of Prosodylab-Aligner is align.py, a script which performs acoustic model training and alignment. This script automates calls to HTK and SoX, an open-source command-line tool which is capable of resampling audio. The included README file provides instructions for installing HTK and SoX on Linux and Mac OS X, and can also be run on Windows. During training, the model is initialized with flat-start monophones, which are then submitted to a single round of model estimation. Then, a tied-state 'small pause' model is inserted …


Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman 2010 Montclair State University

Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

Text classification is a widely studied topic in the area of machine learning. A number of techniques have been developed to represent and classify text documents. Most of the techniques try to achieve good classification performance while taking a document only by its words (e.g. statistical analysis on word frequency and distribution patterns). One of the recent trends in text classification research is to incorporate more semantic interpretation in text classification, especially by using Wikipedia. This paper introduces a technique for incorporating the vast amount of human knowledge accumulated in Wikipedia into text representation and classification. The aim is to …


Study Of Stemming Algorithms, Savitha Kodimala 2010 University of Nevada, Las Vegas

Study Of Stemming Algorithms, Savitha Kodimala

UNLV Theses, Dissertations, Professional Papers, and Capstones

Automated stemming is the process of reducing words to their roots. The stemmed words are typically used to overcome the mismatch problems associated with text searching.


In this thesis, we report on the various methods developed for stemming. In particular, we show the effectiveness of n-gram stemming methods on a collection of documents.


Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee 2010 Technological University Dublin

Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee

Conference papers

Dialogues between humans and robots are necessarily situated and so, often, a shared visual context is present. Exophoric references are very frequent in situated dialogues, and are particularly important in the presence of a shared visual context - for example when a human is verbally guiding a tele-operated mobile robot. We present an approach to automatically resolving exophoric referring expressions in a situated dialogue based on the visual salience of possible referents. We evaluate the effectiveness of this approach and a range of different salience metrics using data from the SCARE corpus which we have augmented with visual information. The …


Situating Spatial Templates For Human-Robot Interaction, John D. Kelleher, Robert J. Ross, Brian Mac Namee, Colm Sloan 2010 Technological University Dublin

Situating Spatial Templates For Human-Robot Interaction, John D. Kelleher, Robert J. Ross, Brian Mac Namee, Colm Sloan

Conference papers

People often refer to objects by describing the object's spatial location relative to another object. Due to their ubiquity in situated discourse, the ability to use 'locative expressions' is fundamental to human-robot dialogue systems. A key component of this ability are computational models of spatial term semantics. These models bridge the grounding gap between spatial language and sensor data. Within the Artificial Intelligence and Robotics communities, spatial template based accounts, such as the Attention Vector Sum model (Regier and Carlson, 2001), have found considerable application in mediating situated human-machine communication (Gorniak, 2004; Brenner et a., 2007; Kelleher and Costello, 2009). …


New Trends In Automatic Assessment: Ontology Matching, Maria Mitina, Patricia Magee, John Cardiff 2010 Technological University Dublin

New Trends In Automatic Assessment: Ontology Matching, Maria Mitina, Patricia Magee, John Cardiff

Conference Papers

Instant individual feedback represents a result of assessment which allows for considerable improvements in both teaching and learning. In this paper we present the application of ontology matching techniques in automatic correction of students’ answers for SQL tests, which will provide teachers with instant feedback to facilitate manual correction and marking and which they can pass to the students. Students experience many problems learning SQL due to the necessity to memorise database schemas, unclear feedback from the database engine on the execution of the query, etc. The program environment utilising the described approach is designed to solve the abovementioned problems …


Topology In Composite Spatial Terms, John D. Kelleher, Robert J. Ross 2010 Technological University Dublin

Topology In Composite Spatial Terms, John D. Kelleher, Robert J. Ross

Conference papers

People often refer to objects by describing the object's spatial location relative to another object, e.g. the book on the right of the table. This type of referring expression is called a spatial locative expression. Spatial locatives have three major components: (1) the target object that is being located (the book), (2) the landmark object relative to which the target is being located (the table), and (3) the description of the spatial relationship that exists between the target and the landmark (on the right of ). In English spatial relationships are often described using spatial prepositions. The set of English …


Proceedings Of The Sixth International Natural Language Generation Conference (Inlg 2010)., John D. Kelleher, Brian Mac Namee, Ielka van der Sluis 2010 Technological University Dublin

Proceedings Of The Sixth International Natural Language Generation Conference (Inlg 2010)., John D. Kelleher, Brian Mac Namee, Ielka Van Der Sluis

Conference papers

No abstract provided.


Digital Commons powered by bepress