Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 4 of 4

Full-Text Articles in Computational Linguistics

Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes Jun 2013

Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes

Master's Theses

In the field of speech recognition, an algorithm must learn to tell the difference between "a nice rock" and "a gneiss rock". These identical-sounding phrases are called oronyms. Word frequency dictionaries are often used by speech recognition systems to help resolve phonetic sequences with more than one possible orthographic phrase interpretation, by looking up which oronym of the root phonetic sequence contains the most-common words.

Our paper demonstrates a technique used to validate word frequency dictionary values. We chose to use frequency values from the UNISYN dictionary, which tallies each word on a per-occurance basis, using a proprietary text corpus, …


Csc Senior Project: Nlpstats, Michael Mease Mar 2013

Csc Senior Project: Nlpstats, Michael Mease

Computer Science and Software Engineering

Natural Language Processing has recently increased in popularity. The field of authorship analysis, specifically, uses various characteristics of text quantified by markers. NLPStats serves as a tool designed to streamline marker extraction based on user needs. A flexible query system allows for custom marker requests, adjustment of result formatting, and preprocessing options. Furthermore, an efficiently designed structure ensures that users retrieve information quickly. As a whole, NLPStats enables anyone, regardless of NLP experience, to extract important information about the text of a document.


Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana Jan 2013

Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana

Department of Computer Science Faculty Scholarship and Creative Works

The goal of this study is to investigate whether learners’ written data in highly inflectional Czech can suggest a consistent set of clues for automatic identification of the learners’ L1 background. For our experiments, we use texts written by learners of Czech, which have been automatically and manually annotated for errors. We define two classes of learners: speakers of Indo-European languages and speakers of non-Indo-European languages. We use an SVM classifier to perform the binary classification. We show that non-content based features perform well on highly inflectional data. In particular, features reflecting errors in orthography are the most useful, yielding …


Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez Jan 2013

Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez

Publications and Research

The aim of this paper is to compare the rhetorical functions gathered from the citations of (14) fourteen master´s theses written by seven Spanish and seven Philippine authors. A typology of nine categories was used in order to identify the cultural rhetorical differences that exist in the use of citation from the contrast between contrasting this element in the Philippine and Spanish cultures. The methodology used is textual analysis of the linguistic context of these citations and its subsequent classification within these nine categories. The results show that there are quantitative and qualitative differences between the cultural conventions of citations …