Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

250 Full-Text Articles 420 Authors 150,666 Downloads 65 Institutions

All Articles in Computational Linguistics

Faceted Search

250 full-text articles. Page 10 of 12.

Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri 2016 CUNY Graduate Center

Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri

Dissertations, Theses, and Capstone Projects

Technology to automatically synthesize linguistically accurate and natural-looking animations of American Sign Language (ASL) would make it easier to add ASL content to websites and media, thereby increasing information accessibility for many people who are deaf and have low English literacy skills. State-of-art sign language animation tools focus mostly on accuracy of manual signs rather than on the facial expressions. We are investigating the synthesis of syntactic ASL facial expressions, which are grammatically required and essential to the meaning of sentences. In this thesis, we propose to: (1) explore the methodological aspects of evaluating sign language animations with facial expressions, …


Cest: City Event Summarization Using Twitter, Deepa Mallela 2016 Boise State University

Cest: City Event Summarization Using Twitter, Deepa Mallela

Computer Science Graduate Projects and Theses

Twitter, with 288 million active users, has become the most popular platform for continuous real-time discussions. This leads to huge amounts of information related to the real-world, which has attracted researchers from both academia and industry. Event detection on Twitter has gained attention as one of the most popular domains of interest within the research community. Unfortunately, existing event detection methodologies have yet to fully explore Twitter metadata and instead rely solely on identifying events based on prior information or focus on events that belong to specific categories. Given the heavy volume of tweets that discuss events, summarization techniques can …


Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan 2016 Purdue University

Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan

Purdue Linguistics, Literature, and Second Language Studies Conference

The synthesized paper covers the topics of “corpus linguistics” and “language instruction and pedagogies”. I would like to do a presentation to highlight the key points in my paper.


Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang 2016 Fox Chase Cancer Center

Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang

COBRA Preprint Series

Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …


A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim 2016 Minnesota State University Mankato

A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim

All Graduate Theses, Dissertations, and Other Capstone Projects

This work explores the trade-offs between time and frequency information during the feature extraction process of an automatic speech recognition (ASR) system using wavelet transform (WT) features instead of Mel-frequency cepstral coefficients (MFCCs) and the benefits of combining the WTs and the MFCCs as inputs to an ASR system. A virtual machine from the Speech Recognition Virtual Kitchen resource (www.speechkitchen.org) is used as the context for implementing a wavelet signal processing module in a speech recognition system. Contributions include a comparison of MFCCs and WT features on small and large vocabulary tasks, application of combined MFCC and WT features on …


Experiments In Idiom Recognition, Jing Peng, Anna Feldman 2016 Montclair State University

Experiments In Idiom Recognition, Jing Peng, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

Some expressions can be ambiguous between idiomatic and literal interpretations depending on the context they occur in, e.g., sales hit the roof vs. hit the roof of the car. We present a novel method of classifying whether a given instance is literal or idiomatic, focusing on verb-noun constructions. We report state-of-the-art results on this task using an approach based on the hypothesis that the distributions of the contexts of the idiomatic phrases will be different from the contexts of the literal usages. We measure contexts by using projections of the words into vector space. For comparison, we implement Fazly et …


The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner 2016 University of Kentucky

The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner

Theses and Dissertations--Linguistics

This study examines the relationship between types of gestures and instances of self-initiated self-repair (SISR) used by persons with non-fluent aphasia (NFA), which is a type of aphasia characterized by stilted speech or signing (Papathanasiou et al., 2013), in interactions with clinicians. Conversation repairs in this study are assessed using the framework of Conversation Analysis (CA), which is an approach for describing, analyzing, and understanding social interaction (Sidnell, 2010). Previous linguistic studies have demonstrated a distinct preference for the use of gesture during a repair by persons with aphasia (Goodwin, 1995; Klippi, 2015; Wilkinson, 2013). This study draws more conclusive …


In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman 2016 Montclair State University

In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman

Department of Computer Science Faculty Scholarship and Creative Works

Expressions, such as add fuel to the fire, can be interpreted literally or idiomatically depending on the context they occur in. Many Natural Language Processing applications could improve their performance if idiom recognition were improved. Our approach is based on the idea that idioms violate cohesive ties in local contexts, while literal expressions do not. We propose two approaches: 1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) …


Emotional Facial Expressions In Synthesised Sign Language Avatars: A Manual Evaluation., Robert G Smith, Brian Nolan 2015 Technological University Dublin

Emotional Facial Expressions In Synthesised Sign Language Avatars: A Manual Evaluation., Robert G Smith, Brian Nolan

Other Resources

This research explores and evaluates the contribution that facial expressions might have regarding improved comprehension and acceptability in sign language avatars. Focusing specifically on Irish sign language (ISL), the Deaf (the uppercase ‘‘D’’ in the word ‘‘Deaf’’ indicates Deaf as a culture as opposed to ‘‘deaf’’ as a medical condition) community’s responsiveness to sign language avatars is examined. The hypothesis of this is as follows: augmenting an existing avatar with the seven widely accepted universal emotions identified by Ekman (Basic emotions: handbook of cognition and emotion. Wiley, London, 2005) to achieve underlying facial expressions will make that avatar more human-like …


A Computational Translation Of The Phaistos Disk, Peter Revesz 2015 University of Nebraska-Lincoln

A Computational Translation Of The Phaistos Disk, Peter Revesz

School of Computing: Conference and Workshop Papers

For over a century the text of the Phaistos Disk remained an enigma without a convincing translation. This paper presents a novel semi-automatic translation method that uses for the first time a recently discovered connection between the Phaistos Disk symbols and other ancient scripts, including the Old Hungarian alphabet. The connection between the Phaistos Disk script and the Old Hungarian alphabet suggested the possibility that the Phaistos Disk language may be related to Proto-Finno-Ugric, Proto-Ugric, or Proto-Hungarian. Using words and suffixes from those languages, it is possible to translate the Phaistos Disk text as an ancient sun hymn, possibly connected …


A Computational Study Of The Evolution Of Cretan And Related Scripts, Peter Revesz 2015 University of Nebraska-Lincoln

A Computational Study Of The Evolution Of Cretan And Related Scripts, Peter Revesz

School of Computing: Conference and Workshop Papers

Crete was the birthplace of several ancient writings, including the Cretan Hieroglyphs, the Linear A and the Linear B scripts. Out of these three only Linear B is deciphered. The sound values of the Cretan Hieroglyph and the Linear A symbols are unknown and attempts to reconstruct them based on Linear B have not been fruitful. In this paper, we compare the ancient Cretan scripts with four other Mediterranean and Black Sea scripts, namely Phoenician, South Arabic, Greek and Old Hungarian. We provide a computational study of the evolution of the three Cretan and four other scripts. This study encompasses …


Providing Objective Metrics Of Team Communication Skills Via Interpersonal Coordination Mechanisms, Celine De Looze, Brian Vaughan, Finnian Kelly, Alison Kay 2015 Trinity College Dublin

Providing Objective Metrics Of Team Communication Skills Via Interpersonal Coordination Mechanisms, Celine De Looze, Brian Vaughan, Finnian Kelly, Alison Kay

Conference Papers

Being able to communicate efficiently has been acknowledged as a vital skill in many different domains. In particular, team communication skills are of key importance in the operation of complex machinery such as aircrafts, maritime vessels and such other, highly-specialized, civilian or military vehicles, as well as the performance of complex tasks in the medical domain. In this paper, we propose to use prosodic accommodation and turn- taking organisation to provide objective metrics of communica- tion skills. To do this, human-factors evaluations, via a coordi- nation Demand Analysis (CDA), were used in conjunction with a dynamic model of prosodic accommodation …


A Study On The Efficacy Of Sentiment Analysis In Author Attribution, Michael J. Schneider 2015 East Tennessee State University

A Study On The Efficacy Of Sentiment Analysis In Author Attribution, Michael J. Schneider

Electronic Theses and Dissertations

The field of authorship attribution seeks to characterize an author’s writing style well enough to determine whether he or she has written a text of interest. One subfield of authorship attribution, stylometry, seeks to find the necessary literary attributes to quantify an author’s writing style. The research presented here sought to determine the efficacy of sentiment analysis as a new stylometric feature, by comparing its performance in attributing authorship against the performance of traditional stylometric features. Experimentation, with a corpus of sci-fi texts, found sentiment analysis to have a much lower performance in assigning authorship than the traditional stylometric features.


Towards News Verification: Deception Detection Methods For News Discourse, Yimin Chen, Victoria L. Rubin, Niall Conroy 2015 Western University

Towards News Verification: Deception Detection Methods For News Discourse, Yimin Chen, Victoria L. Rubin, Niall Conroy

FIMS Presentations

News verification is a process of determining whether a particular news report is truthful or deceptive. Deliberately deceptive (fabricated) news creates false conclusions in the readers’ minds. Truthful (authentic) news matches the writer’s knowledge. How do you tell the difference between the two in an automated way? To investigate this question, we analyzed rhetorical structures, discourse constituent parts and their coherence relations in deceptive and truthful news sample from NPR’s “Bluff the Listener”. Subsequently, we applied a vector space model to cluster the news by discourse feature similarity, achieving 63% accuracy. Our predictive model is not significantly better than chance …


Classifying Idiomatic And Literal Expressions Using Vector Space Representations, Jing Peng, Anna Feldman, Hamza Jazmati 2015 Montclair State University

Classifying Idiomatic And Literal Expressions Using Vector Space Representations, Jing Peng, Anna Feldman, Hamza Jazmati

Department of Computer Science Faculty Scholarship and Creative Works

We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that idioms and literal expressions occur in different contexts. Idioms tend to violate cohesive ties in local contexts, while literals are expected to fit in. Our goal is to capture this intuition using a vector representation of words. We propose two approaches: (1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) Compute …


Metadata And Linked Data In Word Sense Disambiguation, Matthew Corsmeier 2015 San Jose State University

Metadata And Linked Data In Word Sense Disambiguation, Matthew Corsmeier

Library Philosophy and Practice (e-journal)

Word Sense Disambiguation (WSD) can be assisted by taking advantage of the metadata embedded in the various ontologies, lexica, databases, etc… that exist in the Semantic Web. Automated processes that exploit the links already present in the Semantic Web can strengthen parsing of word senses by using user-contributed and semantically-linked data. These processes are only possible because of a commitment to interoperability and the creation of shared standards. This paper will review some of the most heavily used Linguistic Linked Open Data (LLOD) tools and models which show the most promise for using metadata to alleviate problems caused by polysemous …


An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler 2014 LSU New Orleans

An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler

LSU New Orleans Theses and Dissertations

This thesis performs an empirical analysis of Word2Vec by comparing its output to WordNet, a well-known, human-curated lexical database. It finds that Word2Vec tends to uncover more of certain types of semantic relations than others -- with Word2Vec returning more hypernyms, synonomyns and hyponyms than hyponyms or holonyms. It also shows the probability that neighbors separated by a given cosine distance in Word2Vec are semantically related in WordNet. This result both adds to our understanding of the still-unknown Word2Vec and helps to benchmark new semantic tools built from word vectors.


The Role Of Emotion And Facial Expression In Synthesised Sign Language Avatars, Robert G Smith 2014 Technological University Dublin

The Role Of Emotion And Facial Expression In Synthesised Sign Language Avatars, Robert G Smith

Other Resources

This thesis explores the role that underlying emotional facial expressions might have in regards to understandability in sign language avatars. Focusing specifically on Irish Sign Language (ISL), we examine the Deaf community’s requirement for a visual-gestural language as well as some linguistic attributes of ISL which we consider fundamental to this research. Unlike spoken language, visual-gestural languages such as ISL have no standard written representation. Given this, we compare current methods of written representation for signed languages as we consider: which, if any, is the most suitable transcription method for the medical receptionist dialogue corpus. A growing body of work …


The Effect Of Sensor Errors In Situated Human-Computer Dialogue, Niels Schütte, John D. Kelleher, Brian Mac Namee 2014 Technological University Dublin

The Effect Of Sensor Errors In Situated Human-Computer Dialogue, Niels Schütte, John D. Kelleher, Brian Mac Namee

Conference papers

Errors in perception are a problem for computer systems that use sensors to perceive the environment. If a computer system is engaged in dialogue with a human user, these problems in perception lead to problems in the dialogue. We present two experiments, one in which participants interact through dialogue with a robot with perfect perception to fulfil a simple task, and a second one in which the robot is affected by sensor errors and compare the resulting dialogues to determine whether the sensor problems have an impact on dialogue success.


Predicting Music Genre Preferences Based On Online Comments, Andrew J. Sinclair 2014 California Polytechnic State University, San Luis Obispo

Predicting Music Genre Preferences Based On Online Comments, Andrew J. Sinclair

Master's Theses

Communication Accommodation Theory (CAT) states that individuals adapt to each other’s communicative behaviors. This adaptation is called “convergence.” In this work we explore the convergence of writing styles of users of the online music distribution plat- form SoundCloud.com. In order to evaluate our system we created a corpus of over 38,000 comments retrieved from SoundCloud in April 2014. The corpus represents comments from 8 distinct musical genres: Classical, Electronic, Hip Hop, Jazz, Country, Metal, Folk, and World. Our corpus contains: short comments, frequent misspellings, little sentence struc- ture, hashtags, emoticons, and URLs. We adapt techniques used by researchers analyzing other …


Digital Commons powered by bepress