Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Computer Sciences (4)
- Physical Sciences and Mathematics (4)
- Engineering (3)
- Artificial Intelligence and Robotics (2)
- Computer Engineering (2)
-
- Library and Information Science (2)
- Applied Mathematics (1)
- Arts and Humanities (1)
- Business (1)
- Cataloging and Metadata (1)
- Classical Literature and Philology (1)
- Classics (1)
- Computational Engineering (1)
- Computer and Systems Architecture (1)
- Data Science (1)
- Data Storage Systems (1)
- Databases and Information Systems (1)
- Electrical and Computer Engineering (1)
- Labor and Employment Law (1)
- Law (1)
- Medicine and Health Sciences (1)
- Mental and Social Health (1)
- Numerical Analysis and Computation (1)
- Phonetics and Phonology (1)
- Psychiatric and Mental Health (1)
- Psycholinguistics and Neurolinguistics (1)
- Science and Technology Studies (1)
- Institution
- Publication
- Publication Type
Articles 1 - 12 of 12
Full-Text Articles in Computational Linguistics
Analysis Of The Inter-Annotators Agreement And Its Effect On The System Of Anti-Asian Hate Crime Detection On Twitter During Covid-19, Amir Toliyat
Dissertations, Theses, and Capstone Projects
Coronavirus disease 2019 (COVID-19) started in Wuhan, China, in late 2019, and after being utterly contagious in Asian countries, it rapidly spread to other countries. This disease caused governments worldwide to declare a public health crisis with severe measures taken to reduce the speed of the spread of the disease. This pandemic affected the lives of millions of people. Many citizens that lost their loved ones and jobs experienced a wide range of emotions, such as disbelief, shock, concerns about health, fear about food supplies, anxiety, and panic. All of the aforementioned phenomena led to the spread of racism and …
Evaluation Of Different Machine Learning, Deep Learning And Text Processing Techniques For Hate Speech Detection, Nabil Shawkat
Evaluation Of Different Machine Learning, Deep Learning And Text Processing Techniques For Hate Speech Detection, Nabil Shawkat
Graduate Theses/Dissertations
Social media has become a domain that involves a lot of hate speech. Some users feel entitled to engage in abusive conversations by sending abusive messages, tweets, or photos to other users. It is critical to detect hate speech and prevent innocent users from becoming victims. In this study, I explore the effectiveness and performance of various machine learning methods employing text processing techniques to create a robust system for hate speech identification. I assess the performance of Naïve Bayes, Support Vector Machines, Decision Trees, Random Forests, Logistic Regression, and K Nearest Neighbors using three distinct datasets sourced from social …
Creating Data From Unstructured Text With Context Rule Assisted Machine Learning (Craml), Stephen Meisenbacher, Peter Norlander
Creating Data From Unstructured Text With Context Rule Assisted Machine Learning (Craml), Stephen Meisenbacher, Peter Norlander
School of Business: Faculty Publications and Other Works
Popular approaches to building data from unstructured text come with limitations, such as scalability, interpretability, replicability, and real-world applicability. These can be overcome with Context Rule Assisted Machine Learning (CRAML), a method and no-code suite of software tools that builds structured, labeled datasets which are accurate and reproducible. CRAML enables domain experts to access uncommon constructs within a document corpus in a low-resource, transparent, and flexible manner. CRAML produces document-level datasets for quantitative research and makes qualitative classification schemes scalable over large volumes of text. We demonstrate that the method is useful for bibliographic analysis, transparent analysis of proprietary data, …
A Machine Learning Approach To Text-Based Sarcasm Detection, Lara I. Novic
A Machine Learning Approach To Text-Based Sarcasm Detection, Lara I. Novic
Dissertations, Theses, and Capstone Projects
Sarcasm and indirect language are commonplace for humans to produce and recognize but difficult for machines to detect. While artificial intelligence can accurately analyze sentiment and emotion in speech and text, it may struggle with insincere and sardonic content, although it is possible to train a machine to identify uttered and written sarcasm. This paper aims to detect sarcasm using logistic regression and a support vector machine (SVM) and compare their results to a baseline.
The models are trained on headlines from a Kaggle dataset containing headlines from the satirical news website The Onion and serious news website Huffpost (formerly …
Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian
Toward Suicidal Ideation Detection With Lexical Network Features And Machine Learning, Ulya Bayram, William Lee, Daniel Santel, Ali Minai, Peggy Clark, Tracy Glauser, John Pestian
Northeast Journal of Complex Systems (NEJCS)
In this study, we introduce a new network feature for detecting suicidal ideation from clinical texts and conduct various additional experiments to enrich the state of knowledge. We evaluate statistical features with and without stopwords, use lexical networks for feature extraction and classification, and compare the results with standard machine learning methods using a logistic classifier, a neural network, and a deep learning method. We utilize three text collections. The first two contain transcriptions of interviews conducted by experts with suicidal (n=161 patients that experienced severe ideation) and control subjects (n=153). The third collection consists of interviews conducted by experts …
Computational Methods For Comparative Analyses Of Discourses, Zachary Stine
Computational Methods For Comparative Analyses Of Discourses, Zachary Stine
Theses and Dissertations
An underexplored aspect of the abundant linguistic data available for computational analysis is the potential for conducting large-scale comparative analyses between different discourses as a more rigorous complement to existing methods. A significant obstacle in comparative studies is the requirement that a researcher distinguish between discursive distinctions that are superficial and those that reflect deeper, structural relationships between the discourses being compared. In this dissertation, I describe two methodologies for comparing discourses in terms of their underlying structures, drawing on distributional semantic models and information theory. These methodologies are explored within three case studies in which the discourses of several …
Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale
Label Imputation For Homograph Disambiguation: Theoretical And Practical Approaches, Jennifer M. Seale
Dissertations, Theses, and Capstone Projects
This dissertation presents the first implementation of label imputation for the task of homograph disambiguation using 1) transcribed audio, and 2) parallel, or translated, corpora. For label imputation from parallel corpora, a hypothesis of interlingual alignment between homograph pronunciations and text word forms is developed and formalized. Both audio and parallel corpora label imputation techniques are tested empirically in experiments that compare homograph disambiguation model performance using: 1) hand-labeled training data, and 2) hand-labeled training data augmented with label-imputed data. Regularized, multinomial logistic regression and pre-trained ALBERT, BERT, and XLNet language models fine-tuned as token classifiers are developed for homograph …
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Does The Word "Chien" Bark? Representation Learning In Neural Machine Translation Encoders, Emily Campbell
Dissertations, Theses, and Capstone Projects
This thesis presents experiments with using representation learning to explore how neural networks learn. Neural networks which take text as input create internal representations of the text during their training. Recent work has found that these representations can be used to perform other downstream linguistic tasks, such as part-of-speech (POS) tagging. This demonstrates that the neural networks are learning linguistic information and storing this information in the representations. We focus on the representations created by neural machine translation (NMT) models and whether they can be used in POS tagging. We train 5 NMT models including an auto-encoder. We extract the …
Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin
Detecting Clickbait: Here’S How To Do It, Christopher Brogly, Victoria Rubin
Data and Test Instruments
Automatic clickbait detection is a relatively novel task in natural language processing (NLP) and machine learning (ML). “Clickbait” is a hyperlink created primarily to attract attention to its target content. This article introduces a binary classifier, the Language and Information Technology Research Lab (LiT.RL, pronounced “literal”) Clickbait Detector, which automatically distinguishes clickbait from nonclickbait. We used NLP and ML for 38 textual features, contrasting clickbait with “headlinese.” When tested on 11,000 hyperlinks, it achieves 94 per cent accuracy using a support vector machine. Integrated with the LiT.RL News Verification Browser, a downloadable stand-alone research tool, the Clickbait Detector user interface …
Back To The Future: Logic And Machine Learning, Simon Dobnik, John D. Kelleher
Back To The Future: Logic And Machine Learning, Simon Dobnik, John D. Kelleher
Conference papers
In this paper we argue that since the beginning of the natural language processing or computational linguistics there has been a strong connection between logic and machine learning. First of all, there is something logical about language or linguistic about logic. Secondly, we argue that rather than distinguishing between logic and machine learning, a more useful distinction is between top-down approaches and data-driven approaches. Examining some recent approaches in deep learning we argue that they incorporate both properties and this is the reason for their very successful adoption to solve several problems within language technology.
Quantitative Criticism Of Literary Relationships, Joseph P. Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James Brofos, Jorge A. Bonilla Lopez, Lea Schroeder
Quantitative Criticism Of Literary Relationships, Joseph P. Dexter, Theodore Katz, Nilesh Tripuraneni, Tathagata Dasgupta, Ajay Kannan, James Brofos, Jorge A. Bonilla Lopez, Lea Schroeder
Dartmouth Scholarship
Authors often convey meaning by referring to or imitating prior works of literature, a process that creates complex networks of literary relationships (“intertextuality”) and contributes to cultural evolution. In this paper, we use techniques from stylometry and machine learning to address subjective literary critical questions about Latin literature, a corpus marked by an extraordinary concentration of intertextuality. Our work, which we term “quantitative criticism,” focuses on case studies involving two influential Roman authors, the playwright Seneca and the historian Livy. We find that four plays related to but distinct from Seneca’s main writings are differentiated from the rest of the …
Acoustic Classification Of Focus: On The Web And In The Lab, Jonathan Howell, Mats Rooth, Michael Wagner
Acoustic Classification Of Focus: On The Web And In The Lab, Jonathan Howell, Mats Rooth, Michael Wagner
Department of Linguistics Faculty Scholarship and Creative Works
We present a new methodological approach which combines both naturally-occurring speech harvested on the web and speech data elicited in the laboratory. This proof-of-concept study examines the phenomenon of focus sensitivity in English, in which the interpretation of particular grammatical constructions (e.g., the comparative) is sensitive to the location of prosodic prominence. Machine learning algorithms (support vector machines and linear discriminant analysis) and human perception experiments are used to cross-validate the web-harvested and lab-elicited speech. Results con rm the theoretical predictions for location of prominence in comparative clauses and the advantages using both web-harvested and lab-elicited speech. The most robust …