Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 181 - 210 of 250

Full-Text Articles in Computational Linguistics

Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri Jun 2016

Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri

Dissertations, Theses, and Capstone Projects

Technology to automatically synthesize linguistically accurate and natural-looking animations of American Sign Language (ASL) would make it easier to add ASL content to websites and media, thereby increasing information accessibility for many people who are deaf and have low English literacy skills. State-of-art sign language animation tools focus mostly on accuracy of manual signs rather than on the facial expressions. We are investigating the synthesis of syntactic ASL facial expressions, which are grammatically required and essential to the meaning of sentences. In this thesis, we propose to: (1) explore the methodological aspects of evaluating sign language animations with facial expressions, …


Cest: City Event Summarization Using Twitter, Deepa Mallela May 2016

Cest: City Event Summarization Using Twitter, Deepa Mallela

Computer Science Graduate Projects and Theses

Twitter, with 288 million active users, has become the most popular platform for continuous real-time discussions. This leads to huge amounts of information related to the real-world, which has attracted researchers from both academia and industry. Event detection on Twitter has gained attention as one of the most popular domains of interest within the research community. Unfortunately, existing event detection methodologies have yet to fully explore Twitter metadata and instead rely solely on identifying events based on prior information or focus on events that belong to specific categories. Given the heavy volume of tweets that discuss events, summarization techniques can …


Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan Mar 2016

Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan

Purdue Linguistics, Literature, and Second Language Studies Conference

The synthesized paper covers the topics of “corpus linguistics” and “language instruction and pedagogies”. I would like to do a presentation to highlight the key points in my paper.


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang Feb 2016

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim Jan 2016

A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim

All Graduate Theses, Dissertations, and Other Capstone Projects

This work explores the trade-offs between time and frequency information during the feature extraction process of an automatic speech recognition (ASR) system using wavelet transform (WT) features instead of Mel-frequency cepstral coefficients (MFCCs) and the benefits of combining the WTs and the MFCCs as inputs to an ASR system. A virtual machine from the Speech Recognition Virtual Kitchen resource (www.speechkitchen.org) is used as the context for implementing a wavelet signal processing module in a speech recognition system. Contributions include a comparison of MFCCs and WT features on small and large vocabulary tasks, application of combined MFCC and WT features on …


Experiments In Idiom Recognition, Jing Peng, Anna Feldman Jan 2016

Experiments In Idiom Recognition, Jing Peng, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

Some expressions can be ambiguous between idiomatic and literal interpretations depending on the context they occur in, e.g., sales hit the roof vs. hit the roof of the car. We present a novel method of classifying whether a given instance is literal or idiomatic, focusing on verb-noun constructions. We report state-of-the-art results on this task using an approach based on the hypothesis that the distributions of the contexts of the idiomatic phrases will be different from the contexts of the literal usages. We measure contexts by using projections of the words into vector space. For comparison, we implement Fazly et …


The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner Jan 2016

The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner

Theses and Dissertations--Linguistics

This study examines the relationship between types of gestures and instances of self-initiated self-repair (SISR) used by persons with non-fluent aphasia (NFA), which is a type of aphasia characterized by stilted speech or signing (Papathanasiou et al., 2013), in interactions with clinicians. Conversation repairs in this study are assessed using the framework of Conversation Analysis (CA), which is an approach for describing, analyzing, and understanding social interaction (Sidnell, 2010). Previous linguistic studies have demonstrated a distinct preference for the use of gesture during a repair by persons with aphasia (Goodwin, 1995; Klippi, 2015; Wilkinson, 2013). This study draws more conclusive …


In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman Jan 2016

In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

Expressions, such as add fuel to the fire, can be interpreted literally or idiomatically depending on the context they occur in. Many Natural Language Processing applications could improve their performance if idiom recognition were improved. Our approach is based on the idea that idioms violate cohesive ties in local contexts, while literal expressions do not. We propose two approaches: 1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) …


Emotional Facial Expressions In Synthesised Sign Language Avatars: A Manual Evaluation., Robert G Smith, Brian Nolan Oct 2015

Emotional Facial Expressions In Synthesised Sign Language Avatars: A Manual Evaluation., Robert G Smith, Brian Nolan

Other Resources

This research explores and evaluates the contribution that facial expressions might have regarding improved comprehension and acceptability in sign language avatars. Focusing specifically on Irish sign language (ISL), the Deaf (the uppercase ‘‘D’’ in the word ‘‘Deaf’’ indicates Deaf as a culture as opposed to ‘‘deaf’’ as a medical condition) community’s responsiveness to sign language avatars is examined. The hypothesis of this is as follows: augmenting an existing avatar with the seven widely accepted universal emotions identified by Ekman (Basic emotions: handbook of cognition and emotion. Wiley, London, 2005) to achieve underlying facial expressions will make that avatar more human-like …


A Computational Translation Of The Phaistos Disk, Peter Revesz Oct 2015

A Computational Translation Of The Phaistos Disk, Peter Revesz

School of Computing: Conference and Workshop Papers

For over a century the text of the Phaistos Disk remained an enigma without a convincing translation. This paper presents a novel semi-automatic translation method that uses for the first time a recently discovered connection between the Phaistos Disk symbols and other ancient scripts, including the Old Hungarian alphabet. The connection between the Phaistos Disk script and the Old Hungarian alphabet suggested the possibility that the Phaistos Disk language may be related to Proto-Finno-Ugric, Proto-Ugric, or Proto-Hungarian. Using words and suffixes from those languages, it is possible to translate the Phaistos Disk text as an ancient sun hymn, possibly connected …


A Computational Study Of The Evolution Of Cretan And Related Scripts, Peter Revesz Oct 2015

A Computational Study Of The Evolution Of Cretan And Related Scripts, Peter Revesz

School of Computing: Conference and Workshop Papers

Crete was the birthplace of several ancient writings, including the Cretan Hieroglyphs, the Linear A and the Linear B scripts. Out of these three only Linear B is deciphered. The sound values of the Cretan Hieroglyph and the Linear A symbols are unknown and attempts to reconstruct them based on Linear B have not been fruitful. In this paper, we compare the ancient Cretan scripts with four other Mediterranean and Black Sea scripts, namely Phoenician, South Arabic, Greek and Old Hungarian. We provide a computational study of the evolution of the three Cretan and four other scripts. This study encompasses …


Providing Objective Metrics Of Team Communication Skills Via Interpersonal Coordination Mechanisms, Celine De Looze, Brian Vaughan, Finnian Kelly, Alison Kay Sep 2015

Providing Objective Metrics Of Team Communication Skills Via Interpersonal Coordination Mechanisms, Celine De Looze, Brian Vaughan, Finnian Kelly, Alison Kay

Conference Papers

Being able to communicate efficiently has been acknowledged as a vital skill in many different domains. In particular, team communication skills are of key importance in the operation of complex machinery such as aircrafts, maritime vessels and such other, highly-specialized, civilian or military vehicles, as well as the performance of complex tasks in the medical domain. In this paper, we propose to use prosodic accommodation and turn- taking organisation to provide objective metrics of communica- tion skills. To do this, human-factors evaluations, via a coordi- nation Demand Analysis (CDA), were used in conjunction with a dynamic model of prosodic accommodation …


A Study On The Efficacy Of Sentiment Analysis In Author Attribution, Michael J. Schneider Aug 2015

A Study On The Efficacy Of Sentiment Analysis In Author Attribution, Michael J. Schneider

Electronic Theses and Dissertations

The field of authorship attribution seeks to characterize an author’s writing style well enough to determine whether he or she has written a text of interest. One subfield of authorship attribution, stylometry, seeks to find the necessary literary attributes to quantify an author’s writing style. The research presented here sought to determine the efficacy of sentiment analysis as a new stylometric feature, by comparing its performance in attributing authorship against the performance of traditional stylometric features. Experimentation, with a corpus of sci-fi texts, found sentiment analysis to have a much lower performance in assigning authorship than the traditional stylometric features.


Towards News Verification: Deception Detection Methods For News Discourse, Yimin Chen, Victoria L. Rubin, Niall Conroy Jan 2015

Towards News Verification: Deception Detection Methods For News Discourse, Yimin Chen, Victoria L. Rubin, Niall Conroy

FIMS Presentations

News verification is a process of determining whether a particular news report is truthful or deceptive. Deliberately deceptive (fabricated) news creates false conclusions in the readers’ minds. Truthful (authentic) news matches the writer’s knowledge. How do you tell the difference between the two in an automated way? To investigate this question, we analyzed rhetorical structures, discourse constituent parts and their coherence relations in deceptive and truthful news sample from NPR’s “Bluff the Listener”. Subsequently, we applied a vector space model to cluster the news by discourse feature similarity, achieving 63% accuracy. Our predictive model is not significantly better than chance …


Classifying Idiomatic And Literal Expressions Using Vector Space Representations, Jing Peng, Anna Feldman, Hamza Jazmati Jan 2015

Classifying Idiomatic And Literal Expressions Using Vector Space Representations, Jing Peng, Anna Feldman, Hamza Jazmati

Department of Computer Science Faculty Scholarship and Creative Works

We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that idioms and literal expressions occur in different contexts. Idioms tend to violate cohesive ties in local contexts, while literals are expected to fit in. Our goal is to capture this intuition using a vector representation of words. We propose two approaches: (1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) Compute …


Metadata And Linked Data In Word Sense Disambiguation, Matthew Corsmeier Jan 2015

Metadata And Linked Data In Word Sense Disambiguation, Matthew Corsmeier

Library Philosophy and Practice (e-journal)

Word Sense Disambiguation (WSD) can be assisted by taking advantage of the metadata embedded in the various ontologies, lexica, databases, etc… that exist in the Semantic Web. Automated processes that exploit the links already present in the Semantic Web can strengthen parsing of word senses by using user-contributed and semantically-linked data. These processes are only possible because of a commitment to interoperability and the creation of shared standards. This paper will review some of the most heavily used Linguistic Linked Open Data (LLOD) tools and models which show the most promise for using metadata to alleviate problems caused by polysemous …


An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler Dec 2014

An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler

LSU New Orleans Theses and Dissertations

This thesis performs an empirical analysis of Word2Vec by comparing its output to WordNet, a well-known, human-curated lexical database. It finds that Word2Vec tends to uncover more of certain types of semantic relations than others -- with Word2Vec returning more hypernyms, synonomyns and hyponyms than hyponyms or holonyms. It also shows the probability that neighbors separated by a given cosine distance in Word2Vec are semantically related in WordNet. This result both adds to our understanding of the still-unknown Word2Vec and helps to benchmark new semantic tools built from word vectors.


The Role Of Emotion And Facial Expression In Synthesised Sign Language Avatars, Robert G Smith Sep 2014

The Role Of Emotion And Facial Expression In Synthesised Sign Language Avatars, Robert G Smith

Other Resources

This thesis explores the role that underlying emotional facial expressions might have in regards to understandability in sign language avatars. Focusing specifically on Irish Sign Language (ISL), we examine the Deaf community’s requirement for a visual-gestural language as well as some linguistic attributes of ISL which we consider fundamental to this research. Unlike spoken language, visual-gestural languages such as ISL have no standard written representation. Given this, we compare current methods of written representation for signed languages as we consider: which, if any, is the most suitable transcription method for the medical receptionist dialogue corpus. A growing body of work …


The Effect Of Sensor Errors In Situated Human-Computer Dialogue, Niels Schütte, John D. Kelleher, Brian Mac Namee Aug 2014

The Effect Of Sensor Errors In Situated Human-Computer Dialogue, Niels Schütte, John D. Kelleher, Brian Mac Namee

Conference papers

Errors in perception are a problem for computer systems that use sensors to perceive the environment. If a computer system is engaged in dialogue with a human user, these problems in perception lead to problems in the dialogue. We present two experiments, one in which participants interact through dialogue with a robot with perfect perception to fulfil a simple task, and a second one in which the robot is affected by sensor errors and compare the resulting dialogues to determine whether the sensor problems have an impact on dialogue success.


Predicting Music Genre Preferences Based On Online Comments, Andrew J. Sinclair Jun 2014

Predicting Music Genre Preferences Based On Online Comments, Andrew J. Sinclair

Master's Theses

Communication Accommodation Theory (CAT) states that individuals adapt to each other’s communicative behaviors. This adaptation is called “convergence.” In this work we explore the convergence of writing styles of users of the online music distribution plat- form SoundCloud.com. In order to evaluate our system we created a corpus of over 38,000 comments retrieved from SoundCloud in April 2014. The corpus represents comments from 8 distinct musical genres: Classical, Electronic, Hip Hop, Jazz, Country, Metal, Folk, and World. Our corpus contains: short comments, frequent misspellings, little sentence struc- ture, hashtags, emoticons, and URLs. We adapt techniques used by researchers analyzing other …


Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman Mar 2014

Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman

Department of Linguistics Faculty Scholarship and Creative Works

The paper describes a corpus of texts produced by non-native speakersof Czech. We discuss its annotation scheme, consisting of three interlinked tiers,designed to handle a wide range of error types present in the input. Each tier correctsdifferent types of errors; links between the tiers allow capturing errors in word orderand complex discontinuous expressions. Errors are not only corrected, but alsoclassified. The annotation scheme is tested on a data set including approx. 175,000words with fair inter-annotator agreement results. We also explore the possibility ofapplying automated linguistic annotation tools (taggers, spell checkers and grammarcheckers) to the learner text to support or even …


Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova Jan 2014

Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova

Department of Computer Science Faculty Scholarship and Creative Works

We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that words in a given text segment, such as a paragraph, that are highranking representatives of a common topic of discussion are less likely to be a part of an idiomatic expression. Our additional hypothesis is that contexts in which idioms occur, typically, are more affective and therefore, we incorporate a simple analysis of the intensity of the emotions expressed by the contexts. We investigate the bag of words topic representation of one to three paragraphs containing an expression that should be classified as …


Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale Jan 2014

Position Class Preclusion: A Computational Resolution Of Mutually Exclusive Affix Positions, Rebecca O. Hale

Theses and Dissertations--Linguistics

In Paradigm Function Morphology, it is usual to model affix position classes with an ordered sequence of inflectional rule blocks. Each rule block determines how (or whether) a particular affix position is filled. In this model, competition among inflectional rules is assumed to be limited to members of the same rule block; thus, the appearance of an affix in one position cannot be precluded by the appearance of an affix in another position. I present evidence that apparently disconfirms this restriction and suggests that a more general conception of rule competition is necessary. The data appear to imply that an …


Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee Jan 2014

Perception Based Misunderstandings In Human-Computer Dialogues, Niels Schütte, John D. Kelleher, Brian Mac Namee

Articles

In a situated dialogue, misunderstandings may arise if the participants perceive or interpret the environment in different ways. In human-computer dialogue this may be due the sensor errors. We present an experiment system and a series of experiments in which we investigate this problem.


Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes Jun 2013

Misheard Me Oronyminator: Using Oronyms To Validate The Correctness Of Frequency Dictionaries, Jennifer G. Hughes

Master's Theses

In the field of speech recognition, an algorithm must learn to tell the difference between "a nice rock" and "a gneiss rock". These identical-sounding phrases are called oronyms. Word frequency dictionaries are often used by speech recognition systems to help resolve phonetic sequences with more than one possible orthographic phrase interpretation, by looking up which oronym of the root phonetic sequence contains the most-common words.

Our paper demonstrates a technique used to validate word frequency dictionary values. We chose to use frequency values from the UNISYN dictionary, which tallies each word on a per-occurance basis, using a proprietary text corpus, …


Csc Senior Project: Nlpstats, Michael Mease Mar 2013

Csc Senior Project: Nlpstats, Michael Mease

Computer Science and Software Engineering

Natural Language Processing has recently increased in popularity. The field of authorship analysis, specifically, uses various characteristics of text quantified by markers. NLPStats serves as a tool designed to streamline marker extraction based on user needs. A flexible query system allows for custom marker requests, adjustment of result formatting, and preprocessing options. Furthermore, an efficiently designed structure ensures that users retrieve information quickly. As a whole, NLPStats enables anyone, regardless of NLP experience, to extract important information about the text of a document.


Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana Jan 2013

Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana

Department of Computer Science Faculty Scholarship and Creative Works

The goal of this study is to investigate whether learners’ written data in highly inflectional Czech can suggest a consistent set of clues for automatic identification of the learners’ L1 background. For our experiments, we use texts written by learners of Czech, which have been automatically and manually annotated for errors. We define two classes of learners: speakers of Indo-European languages and speakers of non-Indo-European languages. We use an SVM classifier to perform the binary classification. We show that non-content based features perform well on highly inflectional data. In particular, features reflecting errors in orthography are the most useful, yielding …


Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez Jan 2013

Aplicabilidad De La Tipología De Funciones Retóricas De Las Citas Al Género De La Memoria De Máster En Un Contexto Transcultural De Enseñanza Universitaria, David Sánchez-Jiménez

Publications and Research

The aim of this paper is to compare the rhetorical functions gathered from the citations of (14) fourteen master´s theses written by seven Spanish and seven Philippine authors. A typology of nine categories was used in order to identify the cultural rhetorical differences that exist in the use of citation from the contrast between contrasting this element in the Philippine and Spanish cultures. The methodology used is textual analysis of the linguistic context of these citations and its subsequent classification within these nine categories. The results show that there are quantitative and qualitative differences between the cultural conventions of citations …


Lexicalization And De-Lexicalization Processes In Sign Languages: Comparing Depicting Constructions And Viewpoint Gestures, Kearsy Cormier, David Quinto-Pozos, Zed Sehyr, Adam Schembri Nov 2012

Lexicalization And De-Lexicalization Processes In Sign Languages: Comparing Depicting Constructions And Viewpoint Gestures, Kearsy Cormier, David Quinto-Pozos, Zed Sehyr, Adam Schembri

Communication Sciences and Disorders Faculty Articles and Research

In this paper, we compare so-called “classifier” constructions in signed languages (which we refer to as “depicting constructions”) with comparable iconic gestures produced by non-signers. We show clear correspondences between entity constructions and observer viewpoint gestures on the one hand, and handling constructions and character viewpoint gestures on the other. Such correspondences help account for both lexicalisation and de-lexicalisation processes in signed languages and how these processes are influenced by viewpoint. Understanding these processes is crucial when coding and annotating natural sign language data.


Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier Jan 2012

Beefmoves: Dissemination, Diversity, And Dynamics Of English Borrowings In A German Hip Hop Forum, Matt Garley, Julia Hockenmaier

Publications and Research

We investigate how novel English-derived words (anglicisms) are used in a German-language Internet hip hop forum, and what factors contribute to their uptake.