Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (8)
- Physical Sciences and Mathematics (8)
- Arts and Humanities (3)
- Phonetics and Phonology (3)
- Software Engineering (3)
-
- Syntax (3)
- Semantics and Pragmatics (2)
- Theory and Algorithms (2)
- Applied Linguistics (1)
- Art and Design (1)
- Behavioral Economics (1)
- Book and Paper (1)
- Cognition and Perception (1)
- Communication (1)
- Communication Technology and New Media (1)
- Comparative and Historical Linguistics (1)
- Composition (1)
- Computational Engineering (1)
- Creative Writing (1)
- Databases and Information Systems (1)
- Digital Humanities (1)
- Disability Studies (1)
- Discourse and Text Linguistics (1)
- Economics (1)
- Education (1)
- Educational Methods (1)
- Engineering (1)
- Institution
- Keyword
-
- Computational linguistics (3)
- Natural language processing (2)
- Natural language processing (Computer science) (2)
- Analogy (1)
- Automatic Speech Recognition (1)
-
- CAPT (1)
- CS1 (1)
- Children’s books (1)
- Cognition (1)
- College Apps (1)
- Collocation (1)
- Computational Phonology (1)
- Computer (1)
- Corpus (1)
- Cs education (1)
- English songs (1)
- Grants (1)
- Human (1)
- Information retrieval (1)
- Information storage and retrieval systems (1)
- Interactions (1)
- Interface design (1)
- KATR (1)
- Language acquisition (1)
- Language and gender (1)
- Language resource (1)
- Learnability (1)
- Linguistics (1)
- Machine Learning (1)
- Machine learning (1)
- Publication
- Publication Type
Articles 1 - 21 of 21
Full-Text Articles in Computational Linguistics
A Critical Genre Analysis Of The User Requirement Specification Report In The Information Technology Industry, Shahizan Shaharuddin
A Critical Genre Analysis Of The User Requirement Specification Report In The Information Technology Industry, Shahizan Shaharuddin
Student Works (2020-2029)
While it has been established that discursive practices involving writing are highly situated and context-bound, few local studies have examined the processes involved in the construction of texts and the shaping of the intended text. Recent perspectives on writing as ways of working and acting, and texts as social action suggest that studying the processes behind the construction of a text could provide a clearer understanding of the discursive practices involved in the shaping of the text. This is seen as relevant to professional writing context in the present study as texts are often shaped to facilitate the social action …
Mitigating Gender Bias In Neural Machine Translation Using Counterfactual Data, Alan Wong
Mitigating Gender Bias In Neural Machine Translation Using Counterfactual Data, Alan Wong
Dissertations, Theses, and Capstone Projects
Recent advances in deep learning have greatly improved the ability of researchers to develop effective machine translation systems. In particular, the application of modern neural architectures, such as the Transformer, has achieved state-of-the-art BLEU scores in many translation tasks. However, it has been found that even state-of-the-art neural machine translation models can suffer from certain implicit biases, such as gender bias (Lu et al., 2019). In response to this issue, researchers have proposed various potential solutions: some have proposed approaches that inject missing gender information into models, while others have attempted modifying the training data itself. We focus on mitigating …
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Dissertations, Theses, and Capstone Projects
This thesis presents experiments with using representation learning to explore how neural networks learn. Neural networks which take text as input create internal representations of the text during their training. Recent work has found that these representations can be used to perform other downstream linguistic tasks, such as part-of-speech (POS) tagging. This demonstrates that the neural networks are learning linguistic information and storing this information in the representations. We focus on the representations created by neural machine translation (NMT) models and whether they can be used in POS tagging. We train 5 NMT models including an auto-encoder. We extract the …
A Semantic Prosody Analysis Of Swear Words In A Corpus Of English Songs, Hashim Hazri Shahreen
A Semantic Prosody Analysis Of Swear Words In A Corpus Of English Songs, Hashim Hazri Shahreen
Student Works (2020-2029)
This study looks at the semantic prosody of swear words found in a corpus of English songs by looking at their collocations using a corpus software; AntConc. A total of 545 songs were chosen based on the Billboard year-end chart from 2011 to 2016 with 243,689 number of tokens and 8,139 number of word types. Word lists and concordance lines were generated from the corpus for the data analysis. The analysis on the concordance lines showed that not all swear words with negative-based meaning possessed negative semantic prosody. Despite possessing negative semantic prosody, 10 out of 15 swear words possessed …
Lulling Waters: A Poetry Reading For Real-Time Music Generation Through Emotion Mapping, Ashley Muniz, Toshihisa Tsuruoka
Lulling Waters: A Poetry Reading For Real-Time Music Generation Through Emotion Mapping, Ashley Muniz, Toshihisa Tsuruoka
Electronic Literature Organization Conference 2020
Through a poetic narrative, “Lulling Waters” tells the story of a whale overcoming the loss of his mother, who passed away from ingesting plastic, as he attempts to escape from the polluted oceanic world. The live performance of this poem utilizes a software system called Soundwriter, which was developed with the goal of enriching the oral storytelling experience through music. This video demonstrates how Soundwriter’s real-time hybrid system was able to analyze “Lulling Waters” through its lexical and auditory features. Emotionally salient words were given ratings based on arousal, valence, and dominance while the emotionally charged prosodic features of the …
Poetry For Seers Or The Peruvian Visual Poetic Tradition In Front Of New Media, Michael Hurtado, Pamela Medina, Enrique García, Michael Prado
Poetry For Seers Or The Peruvian Visual Poetic Tradition In Front Of New Media, Michael Hurtado, Pamela Medina, Enrique García, Michael Prado
Electronic Literature Organization Conference 2020
Since the first decades of the twentieth century, Peruvian poetic tradition has been characterized by experimental uses of language. Among these possibilities, some records tensioned this medium from the link with the plastic arts, as in the case of the poetry of José María Eguren, while others opted for the playing with the spatiality and visuality of the blank sheet, such as in the case of the work of Carlos Oquendo de Amat. However, it is not until the appearance of the poetry of César Vallejo, specifically with a poems like Trilce in 1922, that these breakages force us to …
Identifying Facets Of Reader-Generated Online Reviews Of Children’S Books Based On A Textual Analysis Approach, Yunseon Choi, Soohyung Joo
Identifying Facets Of Reader-Generated Online Reviews Of Children’S Books Based On A Textual Analysis Approach, Yunseon Choi, Soohyung Joo
Information Science Faculty Publications
With the increasing popularity of social media, online reviews have become one of the primary information sources for book selection. Prior studies have analyzed online reviews, mostly in the domain of business. However, little research has examined the content of online book reviews of children’s books. Book reviews generated by book readers contain different aspects of information, such as opinions, feedback, or emotional responses, from the perspectives of readers. This study explores what aspects of the books are addressed in readers’ reviews, and then it intends to identify categorical features or facets of online book reviews of children’s books. We …
Empirical Analysis Of Cbow And Skip Gram Nlp Models, Tejas Menon
Empirical Analysis Of Cbow And Skip Gram Nlp Models, Tejas Menon
University Honors Theses
CBOW and Skip Gram are two NLP techniques to produce word embedding models that are accurate and performant. They were invented in the seminal paper by T. Mikolov et al. and have since observed optimizations such as negative sampling and subsampling. This paper implements a fully-optimized version of these models using Py-Torch and runs them through a toy sentiment/subject analysis. It is weakly observed that different corpus types affect the skew of word embeddings such that fictional corpus are better suited for sentiment analysis and non-fictional for subject analysis.
Automatic Keyphrase Extraction From Russian-Language Scholarly Papers In Computational Linguistics, Yves Wienecke
Automatic Keyphrase Extraction From Russian-Language Scholarly Papers In Computational Linguistics, Yves Wienecke
University Honors Theses
The automatic extraction of keyphrases from scholarly papers is a necessary step for many Natural Language Processing (NLP) tasks, including text retrieval, machine translation, and text summarization. However, due to the different grammatical and semantic intricacies of languages, this is a highly language-dependent task. Many free and open source implementations of state-of-the-art keyphrase extraction techniques exist, but they are not adapted for processing Russian text. Furthermore, the multi-linguistic character of scholarly papers in the field of Russian computational linguistics and NLP introduces additional complexity to keyphrase extraction. This paper describes a free and open source program as a proof of …
Enhancing The Performance Of Ir-Based Traceability Recovery Of Requirement Artifacts Using Noun Phrases, Dafaalla Abdelrahman Mashahi Khalafalla
Enhancing The Performance Of Ir-Based Traceability Recovery Of Requirement Artifacts Using Noun Phrases, Dafaalla Abdelrahman Mashahi Khalafalla
Student Works (2020-2029)
Requirement traceability can be considered as a measure of software quality to help achieve validation, verification, and reusability. Neglecting traceability leads to less maintainable software. Creating traceability links after-the-fact, known as traceability recovery, is a tedious and time-consuming process when it is done manually. Therefore, information retrieval (IR) methods have been used to automatically identify traceability links between the artifacts. However, as a result of limitations of the software engineer and the IR techniques, the performance of the IR methods is negatively affected. There is no IR method that is able to recover traceability links between artifacts with high precision …
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Inferring Research Fields In Administrative Records Using Text Data, Ekaterina Levitskaya
Dissertations, Theses, and Capstone Projects
The UMETRICS database (Universities: Measuring the Effects of Research on Innovation, Competitiveness, and Science) contains rich information on grants from sponsored federal and non-federal research for 32 universities over a 15-year period. It is hosted at IRIS (Institute for Research on Innovation and Science, University of Michigan) and serves as a rich source of university administrative data; however, it does not contain information on research fields. Categorizing grants data by research field can help to measure results of investment in research and science and provide evidence for the data-driven policy-making; yet administrative data often lacks this type of categorization. In …
Genderlects In Social Media, Alina Korovatskaya
Genderlects In Social Media, Alina Korovatskaya
Dissertations, Theses, and Capstone Projects
Many studies have found significant differences in ways men and women use language; some argue that these differences occur as a result of culture differences, and others suggest that they are influenced by differences in social status and power between the genders. However, some of the major studies were concluded decades ago and do not reflect changes in gender relations in recent years. In this study, we analyze modern conversations using two social media platforms, Twitter and Reddit, to determine whether substantial differences between men and women’s use of language were preserved between the genders.
Doing Away With Defaults: Motivation For A Gradient Parameter Space, Katherine Howitt
Doing Away With Defaults: Motivation For A Gradient Parameter Space, Katherine Howitt
Dissertations, Theses, and Capstone Projects
In this thesis, I propose a reconceptualization of the traditional syntactic parameter space of the principles and parameters framework (Chomsky, 1981). In lieu of binary parameter settings, parameter values exist on a gradient plane where a learner’s knowledge of their language is encoded in their confidence that a particular parametric target value, and thus grammatical construction of an encountered sentence, is likely to be licensed by their target grammar. First, I discuss other learnability models in the classic parameter space which lack either psychological plausibility, theoretical consistency, or some combination of the two. Then, I argue for the Gradient Parameter …
English Wordnet Taxonomic Random Walk Pseudo-Corpora, Filip Klubicka, Alfredo Maldonado, Abhijit Mahalunkar, John D. Kelleher
English Wordnet Taxonomic Random Walk Pseudo-Corpora, Filip Klubicka, Alfredo Maldonado, Abhijit Mahalunkar, John D. Kelleher
Conference papers
This is a resource description paper that describes the creation and properties of a set of pseudo-corpora generated artificially from a random walk over the English WordNet taxonomy. Our WordNet taxonomic random walk implementation allows the exploration of different random walk hyperparameters and the generation of a variety of different pseudo-corpora. We find that different combinations of the walk’s hyperparameters result in varying statistical properties of the generated pseudo-corpora. We have published a total of 81 pseudo-corpora that we have used in our previous research, but have not exhausted all possible combinations of hyperparameters, which is why we have also …
Chaprates, Brinly Xavier, Micole Amanda Marietta, Nidhi Vedantam
Chaprates, Brinly Xavier, Micole Amanda Marietta, Nidhi Vedantam
Student Scholar Symposium Abstracts and Posters
On the Chapman campus, through taking and choosing various classes, there is a significant need for communication and feedback between students and peers, professors, tutors, and study groups. With this, we wanted to create an application that enables users from various majors to not only easily and effectively communicate with various people in their field, but one that also enables them to give and receive feedback on various classes through a rating system. We believe that the application will aid students in a myriad of specific ways, including being involved in study groups and getting tutoring help, determining which classes …
Computational Approaches To The Syntax–Prosody Interface: Using Prosody To Improve Parsing, Hussein M. Ghaly
Computational Approaches To The Syntax–Prosody Interface: Using Prosody To Improve Parsing, Hussein M. Ghaly
Dissertations, Theses, and Capstone Projects
Prosody has strong ties with syntax, since prosody can be used to resolve some syntactic ambiguities. Syntactic ambiguities have been shown to negatively impact automatic syntactic parsing, hence there is reason to believe that prosodic information can help improve parsing. This dissertation considers a number of approaches that aim to computationally examine the relationship between prosody and syntax of natural languages, while also addressing the role of syntactic phrase length, with the ultimate goal of using prosody to improve parsing.
Chapter 2 examines the effect of syntactic phrase length on prosody in double center embedded sentences in French. Data collected …
Phonologically-Informed Speech Coding For Automatic Speech Recognition-Based Foreign Language Pronunciation Training, Anthony J. Vicario
Phonologically-Informed Speech Coding For Automatic Speech Recognition-Based Foreign Language Pronunciation Training, Anthony J. Vicario
Dissertations, Theses, and Capstone Projects
Automatic speech recognition (ASR) and computer-assisted pronunciation training (CAPT) systems used in foreign-language educational contexts are often not developed with the specific task of second-language acquisition in mind. Systems that are built for this task are often excessively targeted to one native language (L1) or a single phonemic contrast and are therefore burdensome to train. Current algorithms have been shown to provide erroneous feedback to learners and show inconsistencies between human and computer perception. These discrepancies have thus far hindered more extensive application of ASR in educational systems.
This thesis reviews the computational models of the human perception of American …
Ghost Peppers: Using Ensemble Models To Detect Professor Attractiveness Commentary On Ratemyprofessors.Com, Angie Waller
Ghost Peppers: Using Ensemble Models To Detect Professor Attractiveness Commentary On Ratemyprofessors.Com, Angie Waller
Dissertations, Theses, and Capstone Projects
In June 2018, RateMyProfessors.com (RMP), a popular website for students to leave professor reviews, removed a controversial feature known as the “chili pepper” which allowed students to rate their professors as “hot” or “not hot.” Though past research has rigorously analyzed the correlation of the chili pepper with higher ratings in other categories (Felton, Mitchell, and Stinson, 2004; Felton et al., 2008), none has measured the effect of the removal of the chili pepper on the text content submitted by students. While it is a positive step that the chili pepper has been removed, text commentary on teacher attractiveness persists …
You Don’T Say... Linguistic Features In Sarcasm Detection, Martina Ducret, Lauren Kruse, Carlos Martinez, Anna Feldman, Jing Peng
You Don’T Say... Linguistic Features In Sarcasm Detection, Martina Ducret, Lauren Kruse, Carlos Martinez, Anna Feldman, Jing Peng
Department of Computer Science Faculty Scholarship and Creative Works
We explore linguistic features that contribute to sarcasm detection. The linguistic features that we investigate are a combination of text and word complexity, stylistic and psychological features. We experiment with sarcastic tweets with and without context. The results of our experiments indicate that contextual information is crucial for sarcasm prediction. One important observation is that sarcastic tweets are typically incongruent with their context in terms of sentiment or emotional load.
Pmkns For Pie: Parsed Morphological Katr Networks Of Sanskrit For Proto-Indo-European, Ryan Mark Mcdonald
Pmkns For Pie: Parsed Morphological Katr Networks Of Sanskrit For Proto-Indo-European, Ryan Mark Mcdonald
Theses and Dissertations--Linguistics
In this thesis, I construct two computational networks for Sanskrit to test theories of nominal accentuation as a way of examining the simplicity of each theory. I will be examining the Paradigmatic Approach and the Compositional Approach to nominal accentuation. For the Paradigmatic Approach, nominals are categorized into mobile and static categories based on how the accent appears in the paradigm (Fortson 2010). For the Compositional Approach, accent mobility is a result of the combination of morphemes and their inherent accent states (Kirparsky 2010). To construct these networks, I use the KATR extension to the DATR language for lexical knowledge …
The Stained Glass Of Knowledge: On Understanding Novice Mental Models Of Computing, Briana Christina Bettin
The Stained Glass Of Knowledge: On Understanding Novice Mental Models Of Computing, Briana Christina Bettin
Dissertations, Master's Theses and Master's Reports
Learning to program can be a novel experience. The rigidity of programming can be at odds with beginning programmer's existing perceptions, and the concepts can feel entirely unfamiliar. These observations motivated this research, which explores two major questions: What factors influence how novices learn programming? and How can analogy by more appropriately leveraged in programming education?
This dissertation investigates the factors influencing novice programming through multiple methods. The CS1 classroom is observed as a "whole system", with consideration to the factors present in it that can influence the learning process. Learning's cognitive processes are elaborated to ground exploration into specifically …