Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (9)
- Physical Sciences and Mathematics (9)
- Discourse and Text Linguistics (5)
- Communication (4)
- Applied Linguistics (3)
-
- Databases and Information Systems (3)
- Semantics and Pragmatics (3)
- Social Media (3)
- Theory and Algorithms (3)
- Artificial Intelligence and Robotics (2)
- Communication Technology and New Media (2)
- Medicine and Health Sciences (2)
- Other Computer Sciences (2)
- Applied Mathematics (1)
- Astrophysics and Astronomy (1)
- Biochemistry, Biophysics, and Structural Biology (1)
- Bioinformatics (1)
- Biomechanics (1)
- Biometry (1)
- Biostatistics (1)
- Categorical Data Analysis (1)
- Communication Sciences and Disorders (1)
- Computational Biology (1)
- Computational Neuroscience (1)
- Electrical and Computer Engineering (1)
- Engineering (1)
- Environmental Sciences (1)
- Institution
- Keyword
-
- Cyberbullying (2)
- Detection (2)
- Accessibility (1)
- Algorithm (1)
- Aphasia (1)
-
- Authorship Attribution (1)
- Authorship attribution (1)
- Bioinformatics (1)
- Biomedical signal processing (1)
- Biometrics (1)
- Classification (1)
- Cognitive science (1)
- Communication disorders (1)
- Computational Linguistics (1)
- Computational linguistics (1)
- Conversation Analysis (1)
- Corpus linguistics (1)
- Data mining (1)
- Dialogue (1)
- Dialogue Systems (1)
- Dictionary (1)
- Electrical engineering (1)
- Event extraction (1)
- Eye Tracking (1)
- Facial Expression Modeling (1)
- Frame of Reference (1)
- Gesture (1)
- High dimensional data (1)
- High-performance computing (1)
- High-throughput genomics (1)
- Publication
- Publication Type
Articles 1 - 18 of 18
Full-Text Articles in Computational Linguistics
Towards Multipurpose Readability Assessment, Ion Madrazo
Towards Multipurpose Readability Assessment, Ion Madrazo
Boise State University Theses and Dissertations
Readability refers to the ease with which a reader can understand a text. Automatic readability assessment has been widely studied over the past 50 years. However, most of the studies focus on the development of tools that apply either to a single language, domain, or document type. This supposes duplicate efforts for both developers, who need to integrate multiple tools in their systems, and final users, who have to deal with incompatibilities among the readability scales of different tools. In this manuscript, we present MultiRead, a multipurpose readability assessment tool capable of predicting the reading difficulty of texts of varied …
Towards A Computational Model Of Frame Of Reference Alignment In Swedish Dialogue, Simon Dobnik, Christine Howes, Kim Demaret, John D. Kelleher
Towards A Computational Model Of Frame Of Reference Alignment In Swedish Dialogue, Simon Dobnik, Christine Howes, Kim Demaret, John D. Kelleher
Conference papers
In this paper we examine how people negotiate, interpret and repair the frame of reference (FoR) in online text based dialogues discussing spatial scenes in Swedish. We describe work-in-progress in which participants are given different perspectives of the same scene and asked to locate several objects that are only shown on one of their pictures. This task requires participants to coordinate on FoR in order to identify the missing objects. This study has implications for situated dialogue systems.
Lexical Variation, Lexical Innovation, And Speaker Motivations: A Historical Psycholinguistic Approach, Jason Timm Dr.
Lexical Variation, Lexical Innovation, And Speaker Motivations: A Historical Psycholinguistic Approach, Jason Timm Dr.
Linguistics ETDs
Speakers commonly re-purpose existing forms in the mental lexicon to create novel form-meaning. Contemporary evidence that such innovation processes have occurred historically is attested in varying degrees of polysemy in the mental lexicon. This dissertation considers speaker motivations underlying these innnovation processes historically. Strong synchronic relationships between frequency and degree of polysemy, on one hand, and frequency and lexical access, on the other hand, have traditionally been interpreted as evidence for the primacy of economic motivations in processes of lexical innovation. In contrast, the cognitive processes that most commonly facilitate innovation, metaphor and metonymy, have largely been described as processes …
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz
Dissertations, Theses, and Capstone Projects
In recent years, Twitter has become a popular testing ground for techniques in authorship attribution. This is due to both the ease of building large corpora as well as the challenges associated with the character limit imposed by the service and the writing styles that have developed as a result. As both false and genuine claims of hacked Twitter accounts have made international news, there is an increasing need for this type of work. For newer Twitter accounts, however, there is little training data. Thus, this study looks to lay the groundwork for cross-domain authorship attribution: training on one source …
An Evaluation Of Pos Taggers For The Childes Corpus, Rui Huang
An Evaluation Of Pos Taggers For The Childes Corpus, Rui Huang
Dissertations, Theses, and Capstone Projects
This project evaluates four mainstream taggers on a representative collection of child-adult’s dialogues from Child Language Data Exchange System. The nine children’s files from Valian corpora and part of Eve corpora have been manually labeled, and rewrote with LARC tagset. They served as gold standard corpora in the training and testing process. Four taggers: CLAN MOR tagger, ACOPOST trigram tagger, Stanford parser, and Ver. 1.14 of Brill tagger have been tested by 10-fold cross validation. By analyzing what kinds of assumptions the tagger made about category assignment lead to failing, we identify several problematic cases of tagging. By comparing the …
Latent Semantic Indexing In The Discovery Of Cyber-Bullying In Online Text, Jacob L. Bigelow
Latent Semantic Indexing In The Discovery Of Cyber-Bullying In Online Text, Jacob L. Bigelow
Computer Science Summer Fellows
The rise in the use of social media and particularly the rise of adolescent use has led to a new means of bullying. Cyber-bullying has proven consequential to youth internet users causing a need for a response. In order to effectively stop this problem we need a verified method of detecting cyber-bullying in online text; we aim to find that method. For this project we look at thirteen thousand labeled posts from Formspring and create a bank of words used in the posts. First the posts are cleaned up by taking out punctuation, normalizing emoticons, and removing high and low …
Detection Of Cyberbullying In Sms Messaging, Bryan W. Bradley
Detection Of Cyberbullying In Sms Messaging, Bryan W. Bradley
Computer Science Summer Fellows
Cyberbullying is a type of bullying that uses technology such as cell phones to harass or malign another person. To detect acts of cyberbullying, we are developing an algorithm that will detect cyberbullying in SMS (text) messages. Over 80,000 text messages have been collected by software installed on cell phones carried by participants in our study. This paper describes the development of the algorithm to detect cyberbullying messages, using the cell phone data collected previously. The algorithm works by first separating the messages into conversations in an automated way. The algorithm then analyzes the conversations and scores the severity and …
Event Parsing In Narrative: Trials And Tribulations Of Archaic English Fairy Tales, Rebecca Lovering
Event Parsing In Narrative: Trials And Tribulations Of Archaic English Fairy Tales, Rebecca Lovering
Dissertations, Theses, and Capstone Projects
While event extraction and automatic summarization have taken great strides in the realm of news stories, fictional narratives like fairy tales have not been so fortunate. A number of challenges arise from the literary elements present in fairy tales that are not found in more straightforward corpora of natural language, such as archaic expressions and sentence structures. To aid in summarization of fictional texts, I created an class - a template for a digital object, in this case a semantic and story event - that captures elements predicted to help classify events as important for inclusion. I wrote a processor …
Nondescript: A Web Tool To Aid Subversion Of Authorship Attribution, Robin Davis
Nondescript: A Web Tool To Aid Subversion Of Authorship Attribution, Robin Davis
Dissertations, Theses, and Capstone Projects
A person’s writing style is uniquely quantifiable and can serve reliably as a biometric. A writer who wishes to remain anonymous can use a number of privacy technologies but can still be identified simply by the words they choose to use — how frequently they use common words like “of,” for instance. Nondescript is a web tool designed first to identify the user’s writing style in terms of word frequency from a given writing sample and document, then to suggest how the author can change their document to lessen its probability of being attributed to them. While Nondescript does not …
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Utilizing Linguistic Context To Improve Individual And Cohort Identification In Typed Text, Adam Goodkind
Dissertations, Theses, and Capstone Projects
The process of producing written text is complex and constrained by pressures that range from physical to psychological. In a series of three sets of experiments, this thesis demonstrates the effects of linguistic context on the timing patterns of the production of keystrokes. We elucidate the effect of linguistic context at three different levels of granularity: The first set of experiments illustrate how the nontraditional syntax of a single linguistic construct, the multi-word expression, can create significant changes in keystroke production patterns. This set of experiments is followed by a set of experiments that test the hypothesis on the entire …
Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri
Data-Driven Synthesis And Evaluation Of Syntactic Facial Expressions In American Sign Language Animation, Hernisa Kacorri
Dissertations, Theses, and Capstone Projects
Technology to automatically synthesize linguistically accurate and natural-looking animations of American Sign Language (ASL) would make it easier to add ASL content to websites and media, thereby increasing information accessibility for many people who are deaf and have low English literacy skills. State-of-art sign language animation tools focus mostly on accuracy of manual signs rather than on the facial expressions. We are investigating the synthesis of syntactic ASL facial expressions, which are grammatically required and essential to the meaning of sentences. In this thesis, we propose to: (1) explore the methodological aspects of evaluating sign language animations with facial expressions, …
Cest: City Event Summarization Using Twitter, Deepa Mallela
Cest: City Event Summarization Using Twitter, Deepa Mallela
Computer Science Graduate Projects and Theses
Twitter, with 288 million active users, has become the most popular platform for continuous real-time discussions. This leads to huge amounts of information related to the real-world, which has attracted researchers from both academia and industry. Event detection on Twitter has gained attention as one of the most popular domains of interest within the research community. Unfortunately, existing event detection methodologies have yet to fully explore Twitter metadata and instead rely solely on identifying events based on prior information or focus on events that belong to specific categories. Given the heavy volume of tweets that discuss events, summarization techniques can …
Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan
Using Corpus To Facilitate Vocabulary Teaching In The Data-Driven Learning Classroom, Ge Lan
Purdue Linguistics, Literature, and Second Language Studies Conference
The synthesized paper covers the topics of “corpus linguistics” and “language instruction and pedagogies”. I would like to do a presentation to highlight the key points in my paper.
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
Hpcnmf: A High-Performance Toolbox For Non-Negative Matrix Factorization, Karthik Devarajan, Guoli Wang
COBRA Preprint Series
Non-negative matrix factorization (NMF) is a widely used machine learning algorithm for dimension reduction of large-scale data. It has found successful applications in a variety of fields such as computational biology, neuroscience, natural language processing, information retrieval, image processing and speech recognition. In bioinformatics, for example, it has been used to extract patterns and profiles from genomic and text-mining data as well as in protein sequence and structure analysis. While the scientific performance of NMF is very promising in dealing with high dimensional data sets and complex data structures, its computational cost is high and sometimes could be critical for …
A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim
A Wavelet Transform Module For A Speech Recognition Virtual Machine, Euisung Kim
All Graduate Theses, Dissertations, and Other Capstone Projects
This work explores the trade-offs between time and frequency information during the feature extraction process of an automatic speech recognition (ASR) system using wavelet transform (WT) features instead of Mel-frequency cepstral coefficients (MFCCs) and the benefits of combining the WTs and the MFCCs as inputs to an ASR system. A virtual machine from the Speech Recognition Virtual Kitchen resource (www.speechkitchen.org) is used as the context for implementing a wavelet signal processing module in a speech recognition system. Contributions include a comparison of MFCCs and WT features on small and large vocabulary tasks, application of combined MFCC and WT features on …
Experiments In Idiom Recognition, Jing Peng, Anna Feldman
Experiments In Idiom Recognition, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Some expressions can be ambiguous between idiomatic and literal interpretations depending on the context they occur in, e.g., sales hit the roof vs. hit the roof of the car. We present a novel method of classifying whether a given instance is literal or idiomatic, focusing on verb-noun constructions. We report state-of-the-art results on this task using an approach based on the hypothesis that the distributions of the contexts of the idiomatic phrases will be different from the contexts of the literal usages. We measure contexts by using projections of the words into vector space. For comparison, we implement Fazly et …
The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner
The Use Of Gesture In Self-Initiated Self-Repair Sequences By Persons With Non-Fluent Aphasia, Eleanor M. Feltner
Theses and Dissertations--Linguistics
This study examines the relationship between types of gestures and instances of self-initiated self-repair (SISR) used by persons with non-fluent aphasia (NFA), which is a type of aphasia characterized by stilted speech or signing (Papathanasiou et al., 2013), in interactions with clinicians. Conversation repairs in this study are assessed using the framework of Conversation Analysis (CA), which is an approach for describing, analyzing, and understanding social interaction (Sidnell, 2010). Previous linguistic studies have demonstrated a distinct preference for the use of gesture during a repair by persons with aphasia (Goodwin, 1995; Klippi, 2015; Wilkinson, 2013). This study draws more conclusive …
In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman
In God We Trust. All Others Must Bring Data. - W. Edwards Deming Using Word Embeddings To Recognize Idioms, Jing Peng, Anna Feldman
Department of Computer Science Faculty Scholarship and Creative Works
Expressions, such as add fuel to the fire, can be interpreted literally or idiomatically depending on the context they occur in. Many Natural Language Processing applications could improve their performance if idiom recognition were improved. Our approach is based on the idea that idioms violate cohesive ties in local contexts, while literal expressions do not. We propose two approaches: 1) Compute inner product of context word vectors with the vector representing a target expression. Since literal vectors predict well local contexts, their inner product with contexts should be larger than idiomatic ones, thereby telling apart literals from idioms; and (2) …