Evaluating And Automating The Annotation Of A Learner Corpus,
2014
Charles University, Prague
Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman
Department of Linguistics Faculty Scholarship and Creative Works
The paper describes a corpus of texts produced by non-native speakersof Czech. We discuss its annotation scheme, consisting of three interlinked tiers,designed to handle a wide range of error types present in the input. Each tier correctsdifferent types of errors; links between the tiers allow capturing errors in word orderand complex discontinuous expressions. Errors are not only corrected, but alsoclassified. The annotation scheme is tested on a data set including approx. 175,000words with fair inter-annotator agreement results. We also explore the possibility ofapplying automated linguistic annotation tools (taggers, spell checkers and grammarcheckers) to the learner text to support or even …
Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions,
2014
Montclair State University
Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova
Department of Computer Science Faculty Scholarship and Creative Works
We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that words in a given text segment, such as a paragraph, that are highranking representatives of a common topic of discussion are less likely to be a part of an idiomatic expression. Our additional hypothesis is that contexts in which idioms occur, typically, are more affective and therefore, we incorporate a simple analysis of the intensity of the emotions expressed by the contexts. We investigate the bag of words topic representation of one to three paragraphs containing an expression that should be classified as …
Perception Based Misunderstandings In Human-Computer Dialogues,
2014
Technological University Dublin
Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee
Articles
In a situated dialogue, misunderstandings may arise if the participants perceive or interpret the environment in different ways. In human-computer dialogue this may be due the sensor errors. We present an experiment system and a series of experiments in which we investigate this problem.
Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions,
2014
University of Kentucky
Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale
Theses and Dissertations--Linguistics
In Paradigm Function Morphology, it is usual to model affix position classes with an ordered sequence of inflectional rule blocks. Each rule block determines how (or whether) a particular affix position is filled. In this model, competition among inflectional rules is assumed to be limited to members of the same rule block; thus, the appearance of an affix in one position cannot be precluded by the appearance of an affix in another position. I present evidence that apparently disconfirms this restriction and suggests that a more general conception of rule competition is necessary. The data appear to imply that an …
Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries,
2013
California Polytechnic State University, San Luis Obispo
Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes
Master's Theses
In the field of speech recognition, an algorithm must learn to tell the difference between "a nice rock" and "a gneiss rock". These identical-sounding phrases are called oronyms. Word frequency dictionaries are often used by speech recognition systems to help resolve phonetic sequences with more than one possible orthographic phrase interpretation, by looking up which oronym of the root phonetic sequence contains the most-common words.
Our paper demonstrates a technique used to validate word frequency dictionary values. We chose to use frequency values from the UNISYN dictionary, which tallies each word on a per-occurance basis, using a proprietary text corpus, …
Csc Senior Project: Nlpstats,
2013
California Polytechnic State University - San Luis Obispo
Csc Senior Project: Nlpstats, Michael Mease
Computer Science and Software Engineering
Natural Language Processing has recently increased in popularity. The field of authorship analysis, specifically, uses various characteristics of text quantified by markers. NLPStats serves as a tool designed to streamline marker extraction based on user needs. A flexible query system allows for custom marker requests, adjustment of result formatting, and preprocessing options. Furthermore, an efficiently designed structure ensures that users retrieve information quickly. As a whole, NLPStats enables anyone, regardless of NLP experience, to extract important information about the text of a document.
Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech,
2013
Montclair State University
Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana
Department of Computer Science Faculty Scholarship and Creative Works
The goal of this study is to investigate whether learners’ written data in highly inflectional Czech can suggest a consistent set of clues for automatic identification of the learners’ L1 background. For our experiments, we use texts written by learners of Czech, which have been automatically and manually annotated for errors. We define two classes of learners: speakers of Indo-European languages and speakers of non-Indo-European languages. We use an SVM classifier to perform the binary classification. We show that non-content based features perform well on highly inflectional data. In particular, features reflecting errors in orthography are the most useful, yielding …
Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria,
2013
CUNY New York City College of Technology
Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez
Publications and Research
The aim of this paper is to compare the rhetorical functions gathered from the citations of (14) fourteen master´s theses written by seven Spanish and seven Philippine authors. A typology of nine categories was used in order to identify the cultural rhetorical differences that exist in the use of citation from the contrast between contrasting this element in the Philippine and Spanish cultures. The methodology used is textual analysis of the linguistic context of these citations and its subsequent classification within these nine categories. The results show that there are quantitative and qualitative differences between the cultural conventions of citations …
Lexicalization And De-Lexicalization Processes In Sign Languages: Comparing Depicting Constructions And Viewpoint Gestures,
2012
University College London
Lexicalization And De-Lexicalization Processes In Sign Languages: Comparing Depicting Constructions And Viewpoint Gestures, Kearsy Cormier, David Quinto-Pozos, Zed Sehyr, Adam Schembri
Communication Sciences and Disorders Faculty Articles and Research
In this paper, we compare so-called “classifier” constructions in signed languages (which we refer to as “depicting constructions”) with comparable iconic gestures produced by non-signers. We show clear correspondences between entity constructions and observer viewpoint gestures on the one hand, and handling constructions and character viewpoint gestures on the other. Such correspondences help account for both lexicalisation and de-lexicalisation processes in signed languages and how these processes are influenced by viewpoint. Understanding these processes is crucial when coding and annotating natural sign language data.
Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum,
2012
CUNY York College
Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier
Publications and Research
We investigate how novel English-derived words (anglicisms) are used in a German-language Internet hip hop forum, and what factors contribute to their uptake.
Using Textual Features To Predict Popular Content On Digg,
2011
University of Nebraska-Lincoln
Using Textual Features To Predict Popular Content On Digg, Paul H. Miller
Department of English: Dissertations, Theses, and Student Research
Over the past few years, collaborative rating sites, such as Netflix, Digg and Stumble, have become increasingly prevalent sites for users to find trending content. I used various data mining techniques to study Digg, a social news site, to examine the influence of content on popularity. What influence does content have on popularity, and what influence does content have on users’ decisions? Overwhelmingly, prior studies have consistently shown that predicting popularity based on content is difficult and maybe even inherently impossible. The same submission can have multiple outcomes and content neither determines popularity, nor individual user decisions. My results show …
Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language,
2011
DePaul University
Linguistics As Structure In Computer Animation: Toward A More Effective Synthesis Of Brow Motion In American Sign Language, Rosalee Wolfe, Peter Cook, John C. Mcdonald, Jerry Schnepp
Visual Communications and Technology Education Faculty Publications
Computer-generated three-dimensional animation holds great promise for synthesizing utterances in American Sign Language (ASL) that are not only grammatical, but well tolerated by members of the Deaf community. Unfortunately, animation poses several challenges stemming from the necessity of grappling with massive amounts of data. However, the linguistics of ASL can aid in surmounting the challenge by providing structure and rules for organizing animation data. An exploration of the linguistic and extra linguistic behavior of the brows from an animator’s viewpoint yields a new approach for synthesizing nonmanuals that differs from the conventional animation of anatomy and instead offers a different …
Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech,
2011
University of Pennsylvania
Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner
Department of Linguistics Faculty Scholarship and Creative Works
The Penn Forced Aligner automates the alignment process using the Hidden Markov Model Toolkit (HTK). The core of Prosodylab-Aligner is align.py, a script which performs acoustic model training and alignment. This script automates calls to HTK and SoX, an open-source command-line tool which is capable of resampling audio. The included README file provides instructions for installing HTK and SoX on Linux and Mac OS X, and can also be run on Windows. During training, the model is initialized with flat-start monophones, which are then submitted to a single round of model estimation. Then, a tied-state 'small pause' model is inserted …
Semantic Enrichment Of Text Representation With Wikipedia For Text Classification,
2010
Montclair State University
Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Text classification is a widely studied topic in the area of machine learning. A number of techniques have been developed to represent and classify text documents. Most of the techniques try to achieve good classification performance while taking a document only by its words (e.g. statistical analysis on word frequency and distribution patterns). One of the recent trends in text classification research is to incorporate more semantic interpretation in text classification, especially by using Wikipedia. This paper introduces a technique for incorporating the vast amount of human knowledge accumulated in Wikipedia into text representation and classification. The aim is to …
Study Of Stemming Algorithms,
2010
University of Nevada, Las Vegas
Study Of Stemming Algorithms, Savitha Kodimala
UNLV Theses, Dissertations, Professional Papers, and Capstones
Automated stemming is the process of reducing words to their roots. The stemmed words are typically used to overcome the mismatch problems associated with text searching.
In this thesis, we report on the various methods developed for stemming. In particular, we show the effectiveness of n-gram stemming methods on a collection of documents.
Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation.,
2010
Technological University Dublin
Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee
Conference papers
Dialogues between humans and robots are necessarily situated and so, often, a shared visual context is present. Exophoric references are very frequent in situated dialogues, and are particularly important in the presence of a shared visual context - for example when a human is verbally guiding a tele-operated mobile robot. We present an approach to automatically resolving exophoric referring expressions in a situated dialogue based on the visual salience of possible referents. We evaluate the effectiveness of this approach and a range of different salience metrics using data from the SCARE corpus which we have augmented with visual information. The …
Situating Spatial Templates For Human-Robot Interaction,
2010
Technological University Dublin
Situating Spatial Templates For Human-Robot Interaction, John D. Kelleher, Robert J. Ross, Brian Mac Namee, Colm Sloan
Conference papers
People often refer to objects by describing the object's spatial location relative to another object. Due to their ubiquity in situated discourse, the ability to use 'locative expressions' is fundamental to human-robot dialogue systems. A key component of this ability are computational models of spatial term semantics. These models bridge the grounding gap between spatial language and sensor data. Within the Artificial Intelligence and Robotics communities, spatial template based accounts, such as the Attention Vector Sum model (Regier and Carlson, 2001), have found considerable application in mediating situated human-machine communication (Gorniak, 2004; Brenner et a., 2007; Kelleher and Costello, 2009). …
New Trends In Automatic Assessment: Ontology Matching,
2010
Technological University Dublin
New Trends In Automatic Assessment: Ontology Matching, Maria Mitina, Patricia Magee, John Cardiff
Conference Papers
Instant individual feedback represents a result of assessment which allows for considerable improvements in both teaching and learning. In this paper we present the application of ontology matching techniques in automatic correction of students’ answers for SQL tests, which will provide teachers with instant feedback to facilitate manual correction and marking and which they can pass to the students. Students experience many problems learning SQL due to the necessity to memorise database schemas, unclear feedback from the database engine on the execution of the query, etc. The program environment utilising the described approach is designed to solve the abovementioned problems …
Topology In Composite Spatial Terms,
2010
Technological University Dublin
Topology In Composite Spatial Terms, John D. Kelleher, Robert J. Ross
Conference papers
People often refer to objects by describing the object's spatial location relative to another object, e.g. the book on the right of the table. This type of referring expression is called a spatial locative expression. Spatial locatives have three major components: (1) the target object that is being located (the book), (2) the landmark object relative to which the target is being located (the table), and (3) the description of the spatial relationship that exists between the target and the landmark (on the right of ). In English spatial relationships are often described using spatial prepositions. The set of English …
Proceedings Of The Sixth International Natural Language Generation Conference (Inlg 2010).,
2010
Technological University Dublin
Proceedings Of The Sixth International Natural Language Generation Conference (Inlg 2010)., John D. Kelleher, Brian Mac Namee, Ielka Van Der Sluis
Conference papers
No abstract provided.
