Open Access. Powered by Scholars. Published by Universities.®
- Discipline
- Keyword
-
- Prosody (2)
- Accession rate (1)
- Acoustic prominence (1)
- Average ambiguity (1)
- Computational linguistics (1)
-
- Corpus lingusitics (1)
- Czech (1)
- Distributional semantics (1)
- Ensemble (1)
- Error annotation (1)
- Euphemisms (1)
- Idiom recognition (1)
- Inflected languages (1)
- Learner corpus (1)
- Machine learning (1)
- Morphosyntax (1)
- NLP (1)
- Politeness (1)
- Prominence (1)
- Second language acquisition (1)
- Semantics (1)
- Speech processing (1)
- Tagset (1)
- Text classification (1)
- Text representation (1)
- Vector space models (1)
- Voting (1)
- Web as corpus (1)
- Wikipedia (1)
- Word embeddings (1)
- Publication Year
- Publication
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Computational Linguistics
Metaphor Detection In Poems In Misurata Arabic Sub-Dialect : An Lstm Model, Azza Abugharsa
Metaphor Detection In Poems In Misurata Arabic Sub-Dialect : An Lstm Model, Azza Abugharsa
Theses, Dissertations and Culminating Projects
Natural Language Processing (NLP) in Arabic is witnessing an increasing interest in investigating different topics in the field. One of the topics that have drawn attention is the automatic processing of Arabic figurative language. The focus in previous projects is on detecting and interpreting metaphors in comments from social media as well as phrases and/or headlines from news articles. The current project focuses on metaphor detection in poems written in the Misurata Arabic sub-dialect spoken in Misurata, located in the North African region. The dataset is initially annotated by a group of linguists, and their annotation is treated as the …
Searching For Pets: Using Distributional And Sentiment-Based Methods To Find Potentially Euphemistic Terms, Patrick Lee, Martha Gavidia, Anna Feldman, Jing Peng
Searching For Pets: Using Distributional And Sentiment-Based Methods To Find Potentially Euphemistic Terms, Patrick Lee, Martha Gavidia, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
This paper presents a linguistically driven proof of concept for finding potentially euphemistic terms, or PETs. Acknowledging that PETs tend to be commonly used expressions for a certain range of sensitive topics, we make use of distributional similarities to select and filter phrase candidates from a sentence and rank them using a set of simple sentiment-based metrics. We present the results of our approach tested on a corpus of sentences containing euphemisms, demonstrating its efficacy for detecting single and multi-word PETs from a broad range of topics. We also discuss future potential for sentiment-based methods on this task.
A Report On The Euphemisms Detection Shared Task, Patrick Lee, Anna Feldman, Jing Peng
A Report On The Euphemisms Detection Shared Task, Patrick Lee, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
This paper presents The Shared Task on Euphemism Detection for the Third Workshop on Figurative Language Processing (FigLang 2022) held in conjunction with EMNLP 2022. Participants were invited to investigate the euphemism detection task: given input text, identify whether it contains a euphemism. The input data is a corpus of sentences containing potentially euphemistic terms (PETs) collected from the GloWbE corpus (Davies and Fuchs, 2015), and are human-annotated as containing either a euphemistic or literal usage of a PET. In this paper, we present the results and analyze the common themes, methods and findings of the participating teams.
Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng
Cats Are Fuzzy Pets: A Corpus And Analysis Of Potentially Euphemistic Terms, Martha Gavidia, Patrick Lee, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
Euphemisms have not received much attention in natural language processing, despite being an important element of polite and figurative language. Euphemisms prove to be a difficult topic, not only because they are subject to language change, but also because humans may not agree on what is a euphemism and what is not. Nonetheless, the first step to tackling the issue is to collect and analyze examples of euphemisms. We present a corpus of potentially euphemistic terms (PETs) along with example texts from the GloWbE corpus. Additionally, we present a subcorpus of texts where these PETs are not being used euphemistically, …
Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman
Findings Of The Nlp4if-2021 Shared Tasks On Fighting The Covid-19 Infodemic And Censorship Detection, Shaden Shaar, Firoj Alam, Giovanni Da San Martino, Alex Nikolov, Wajdi Zaghouani, Preslav Nakov, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
We present the results and the main findings of the NLP4IF-2021 shared tasks. Task 1 focused on fighting the COVID-19 infodemic in social media, and it was offered in Arabic, Bulgarian, and English. Given a tweet, it asked to predict whether that tweet contains a verifiable claim, and if so, whether it is likely to be false, is of general interest, is likely to be harmful, and is worthy of manual fact-checking; also, whether it is harmful to society, and whether it requires the attention of policy makers. Task 2 focused on censorship detection, and was offered in Chinese. A …
You Don’T Say... Linguistic Features In Sarcasm Detection, Martina Ducret, Lauren Kruse, Carlos Martinez, Anna Feldman, Jing Peng
You Don’T Say... Linguistic Features In Sarcasm Detection, Martina Ducret, Lauren Kruse, Carlos Martinez, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
We explore linguistic features that contribute to sarcasm detection. The linguistic features that we investigate are a combination of text and word complexity, stylistic and psychological features. We experiment with sarcastic tweets with and without context. The results of our experiments indicate that contextual information is crucial for sarcasm prediction. One important observation is that sarcastic tweets are typically incongruent with their context in terms of sentiment or emotional load.
Leveraging Nlp And Social Network Analytic Techniques To Detect Censored Keywords: System Design And Experiments, Christopher S. Leberknight, Anna Feldman
Leveraging Nlp And Social Network Analytic Techniques To Detect Censored Keywords: System Design And Experiments, Christopher S. Leberknight, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Internet regulation in the form of online censorship and Internet shutdowns have been increasing over recent years. This paper presents a natural language processing (NLP) application for performing cross country probing that conceals the exact location of the originating request. A detailed discussion of the application aims to stimulate further investigation into new methods for measuring and quantifying Internet censorship practices around the world. In addition, results from two experiments involving search engine queries of banned keywords demonstrates censorship practices vary across different search engines. These results suggest opportunities for developing circumvention technologies that enable open and free access to …
Automatic Idiom Recognition With Word Embeddings, Jing Peng, Anna Feldman
Automatic Idiom Recognition With Word Embeddings, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Expressions, such as add fuel to the fire, can be interpreted literally or idiomatically depending on the context they occur in. Many Natural Language Processing applications could improve their performance if idiom recognition were improved. Our approach is based on the idea that idioms and their literal counterparts do not appear in the same contexts. We propose two approaches: (1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) …
Acoustic Classification Of Focus: On The Web And In The Lab, Jonathan Howell, Mats Rooth, Michael Wagner
Acoustic Classification Of Focus: On The Web And In The Lab, Jonathan Howell, Mats Rooth, Michael Wagner
Department of Linguistics Faculty Scholarship and Creative Works
We present a new methodological approach which combines both naturally-occurring speech harvested on the web and speech data elicited in the laboratory. This proof-of-concept study examines the phenomenon of focus sensitivity in English, in which the interpretation of particular grammatical constructions (e.g., the comparative) is sensitive to the location of prosodic prominence. Machine learning algorithms (support vector machines and linear discriminant analysis) and human perception experiments are used to cross-validate the web-harvested and lab-elicited speech. Results con rm the theoretical predictions for location of prominence in comparative clauses and the advantages using both web-harvested and lab-elicited speech. The most robust …
Experiments In Idiom Recognition, Jing Peng, Anna Feldman
Experiments In Idiom Recognition, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Some expressions can be ambiguous between idiomatic and literal interpretations depending on the context they occur in, e.g., sales hit the roof vs. hit the roof of the car. We present a novel method of classifying whether a given instance is literal or idiomatic, focusing on verb-noun constructions. We report state-of-the-art results on this task using an approach based on the hypothesis that the distributions of the contexts of the idiomatic phrases will be different from the contexts of the literal usages. We measure contexts by using projections of the words into vector space. For comparison, we implement Fazly et …
In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman
In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Expressions, such as add fuel to the fire, can be interpreted literally or idiomatically depending on the context they occur in. Many Natural Language Processing applications could improve their performance if idiom recognition were improved. Our approach is based on the idea that idioms violate cohesive ties in local contexts, while literal expressions do not. We propose two approaches: 1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) …
Classifying Idiomatic And Literal Expressions Using Vector Space Representations, Jing Peng, Anna Feldman, Hamza Jazmati
Classifying Idiomatic And Literal Expressions Using Vector Space Representations, Jing Peng, Anna Feldman, Hamza Jazmati
Department of Computer Science Faculty Scholarship and Creative Works
We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that idioms and literal expressions occur in different contexts. Idioms tend to violate cohesive ties in local contexts, while literals are expected to fit in. Our goal is to capture this intuition using a vector representation of words. We propose two approaches: (1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) Compute …
Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman
Evaluating And Automating The Annotation Of A Learner Corpus, Alexandr Rosen, Jirka Hana, Barbora Stindlova, Anna Feldman
Department of Linguistics Faculty Scholarship and Creative Works
The paper describes a corpus of texts produced by non-native speakersof Czech. We discuss its annotation scheme, consisting of three interlinked tiers,designed to handle a wide range of error types present in the input. Each tier correctsdifferent types of errors; links between the tiers allow capturing errors in word orderand complex discontinuous expressions. Errors are not only corrected, but alsoclassified. The annotation scheme is tested on a data set including approx. 175,000words with fair inter-annotator agreement results. We also explore the possibility ofapplying automated linguistic annotation tools (taggers, spell checkers and grammarcheckers) to the learner text to support or even …
Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova
Classifying Idiomatic And Literal Expressions Using Topic Models And Intensity Of Emotions, Jing Peng, Anna Feldman, Ekaterina Vylomova
Department of Computer Science Faculty Scholarship and Creative Works
We describe an algorithm for automatic classification of idiomatic and literal expressions. Our starting point is that words in a given text segment, such as a paragraph, that are highranking representatives of a common topic of discussion are less likely to be a part of an idiomatic expression. Our additional hypothesis is that contexts in which idioms occur, typically, are more affective and therefore, we incorporate a simple analysis of the intensity of the emotions expressed by the contexts. We investigate the bag of words topic representation of one to three paragraphs containing an expression that should be classified as …
Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana
Automatic Identification Of Learners’ Language Background Based On Their Writing In Czech, Katsiaryna Aharodnik, Marco Chang, Anna Feldman, Jirka Hana
Department of Computer Science Faculty Scholarship and Creative Works
The goal of this study is to investigate whether learners’ written data in highly inflectional Czech can suggest a consistent set of clues for automatic identification of the learners’ L1 background. For our experiments, we use texts written by learners of Czech, which have been automatically and manually annotated for errors. We define two classes of learners: speakers of Indo-European languages and speakers of non-Indo-European languages. We use an SVM classifier to perform the binary classification. We show that non-content based features perform well on highly inflectional data. In particular, features reflecting errors in orthography are the most useful, yielding …
Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner
Prosodylab-Aligner: A Tool For Forced Alignment Of Laboratory Speech, Kyle Gorman, Jonathan Howell, Michael Wagner
Department of Linguistics Faculty Scholarship and Creative Works
The Penn Forced Aligner automates the alignment process using the Hidden Markov Model Toolkit (HTK). The core of Prosodylab-Aligner is align.py, a script which performs acoustic model training and alignment. This script automates calls to HTK and SoX, an open-source command-line tool which is capable of resampling audio. The included README file provides instructions for installing HTK and SoX on Linux and Mac OS X, and can also be run on Windows. During training, the model is initialized with flat-start monophones, which are then submitted to a single round of model estimation. Then, a tied-state 'small pause' model is inserted …
Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman
Semantic Enrichment Of Text Representation With Wikipedia For Text Classification, Hiroki Yamakawa, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Text classification is a widely studied topic in the area of machine learning. A number of techniques have been developed to represent and classify text documents. Most of the techniques try to achieve good classification performance while taking a document only by its words (e.g. statistical analysis on word frequency and distribution patterns). One of the recent trends in text classification research is to incorporate more semantic interpretation in text classification, especially by using Wikipedia. This paper introduces a technique for incorporating the vast amount of human knowledge accumulated in Wikipedia into text representation and classification. The aim is to …
Tagset Design, Inflected Languages, And N-Gram Tagging, Anna Feldman
Tagset Design, Inflected Languages, And N-Gram Tagging, Anna Feldman
Department of Linguistics Faculty Scholarship and Creative Works
This paper explores the relationship between the tagset design and linguistic properties of inflected languages for the task of morphosyntactic tagging. Some information theoretic measures and statistics on these languages are reported which show, unsurprisingly, that the tagsets for morphologically rich languages are larger than tagsets for English and the average tag/token ambiguity is higher. The surprising outcome of the experiments is that for Catalan, Czech, Polish, Portuguese,and Russian – which are considered to be “word order” free languages (to various degrees) – the knowledge about the preceding tag reduces the uncertainty about the tag in question if the detailed …