Open Access. Powered by Scholars. Published by Universities.®

Computational Linguistics Commons

Open Access. Powered by Scholars. Published by Universities.®

Natural Language Processing

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 19 of 19

Full-Text Articles in Computational Linguistics

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma May 2026

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

This presentation introduces Propasafe-Hybrid, a hybrid system for sentence-level propaganda detection that combines offline transformer-based classification with selective large language model (LLM) explainability. The system employs a two-stage pipeline in which a local BERT-based classifier evaluates all input text and filters non-propagandistic content, while only high-confidence candidates are forwarded to an LLM for rhetorical technique labeling and explanation. This design enables cost-aware, privacy-conscious, and scalable analysis by reducing unnecessary reliance on external models.

Propasafe-Hybrid identifies propagandistic techniques such as loaded language, obfuscation, and appeal to fear, and generates concise natural language rationales that make these techniques interpretable to users. By …


From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie Apr 2025

From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie

Undergraduate Theses

Adversarial attacks pose a significant threat to the reliability of machine learning-based spam detection systems in social media. This undergraduate thesis, "From Adversarial Attacks to Robust Classifiers: A Study in Social Media Spam Detection – Black Box & White Box," systematically examines the impact of both black-box and white-box adversarial attacks on a range of spam classifiers, including Logistic Regression, Decision Trees, Random Forests, K-Nearest Neighbors, Bagging, Gradient Boosting, and Support Vector Machines. Leveraging a novel dataset derived from Twitter spam messages and enhanced with adversarial perturbations such as synonym replacement and character-level modifications, this study evaluates classifier performance under …


Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri Jan 2024

Streamlining Public Engagement In Transportation Projects Using Text Analytics, Alireza Shamshiri

Civil Engineering Dissertations - Archive

Infrastructure projects impact a broad range of stakeholders, particularly local communities, whose engagement is critical for successful outcomes. Despite the importance of public engagement in these projects, traditional methods of capturing and analyzing public opinion often fail to fully represent the diverse, genuine perspectives involved. This has led to conflicts between community members and project sponsors. On the other hand, despite advancements in text analytics, including natural language processing (NLP) and its subfields such as topic modeling, sentiment analysis, and neural networks, their functionalities and effectiveness in analyzing public opinion in the domain of infrastructure projects have not been fully …


Ai Approaches To Understand Human Deceptions, Perceptions, And Perspectives In Social Media, Chih-Yuan Li May 2023

Ai Approaches To Understand Human Deceptions, Perceptions, And Perspectives In Social Media, Chih-Yuan Li

Dissertations

Social media platforms have created virtual space for sharing user generated information, connecting, and interacting among users. However, there are research and societal challenges: 1) The users are generating and sharing the disinformation 2) It is difficult to understand citizens' perceptions or opinions expressed on wide variety of topics; and 3) There are overloaded information and echo chamber problems without overall understanding of the different perspectives taken by different people or groups.

This dissertation addresses these three research challenges with advanced AI and Machine Learning approaches. To address the fake news, as deceptions on the facts, this dissertation presents Machine …


Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian) Mar 2023

Chatgpt As Metamorphosis Designer For The Future Of Artificial Intelligence (Ai): A Conceptual Investigation, Amarjit Kumar Singh (Library Assistant), Dr. Pankaj Mathur (Deputy Librarian)

Library Philosophy and Practice (e-journal)

Abstract

Purpose: The purpose of this research paper is to explore ChatGPT’s potential as an innovative designer tool for the future development of artificial intelligence. Specifically, this conceptual investigation aims to analyze ChatGPT’s capabilities as a tool for designing and developing near about human intelligent systems for futuristic used and developed in the field of Artificial Intelligence (AI). Also with the helps of this paper, researchers are analyzed the strengths and weaknesses of ChatGPT as a tool, and identify possible areas for improvement in its development and implementation. This investigation focused on the various features and functions of ChatGPT that …


Analysis Of The Inter-Annotators Agreement And Its Effect On The System Of Anti-Asian Hate Crime Detection On Twitter During Covid-19, Amir Toliyat Feb 2023

Analysis Of The Inter-Annotators Agreement And Its Effect On The System Of Anti-Asian Hate Crime Detection On Twitter During Covid-19, Amir Toliyat

Dissertations, Theses, and Capstone Projects

Coronavirus disease 2019 (COVID-19) started in Wuhan, China, in late 2019, and after being utterly contagious in Asian countries, it rapidly spread to other countries. This disease caused governments worldwide to declare a public health crisis with severe measures taken to reduce the speed of the spread of the disease. This pandemic affected the lives of millions of people. Many citizens that lost their loved ones and jobs experienced a wide range of emotions, such as disbelief, shock, concerns about health, fear about food supplies, anxiety, and panic. All of the aforementioned phenomena led to the spread of racism and …


Simulating The Machine Translation Of Low-Resource Languages By Designing A Translator Between English And An Artificially Constructed Language, Michaela Snyder Jan 2023

Simulating The Machine Translation Of Low-Resource Languages By Designing A Translator Between English And An Artificially Constructed Language, Michaela Snyder

Mahurin Honors College Capstone Experience/Thesis Projects

Natural language processing (NLP), or the use of computers to analyze natural language, is a field that relies heavily on syntax. It would seem intuitive that computers would thrive in this area due to their strict syntax requirements, but the syntax of natural languages leaves them unable to properly parse and generate sentences that seem normal to the average speaker. A subfield of NLP, machine translation, works mainly to computerize translation between different languages. Unfortunately, such translation is not without its weaknesses; language documentation is not created equal, and many low-resource languages—languages with relatively few kinds of documentation, most often …


N-Gram Text Classification On Standard Croatian, Bosnian And Serbian, Kegan Messmer Jan 2023

N-Gram Text Classification On Standard Croatian, Bosnian And Serbian, Kegan Messmer

Honors Theses and Capstones

This study attempts to use three different kinds of n-gram text classification models to differentiate the standard forms of Croatian, Bosnian, and Serbian. These three languages, along with Montenegrin, were once considered one language, collectively termed “Serbo-Croatian”. These languages share a common South Slavic ancestry, and there is an argument to be made that their novel status as distinct languages is due to non-linguistic factors, such as culture and politics. This study uses 300,000 sentences from each language, sourced from Wikipedia pages written in the standard forms of each language. Three different classifiers were used: one unigram, one bigram, and …


Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell Oct 2022

Applying Positive Psychology’S Subjective Well-Being To Online Interactions, Nicole Delellis, Dominique Kelly, Liu Yifan, Alex Mayhew, Yimin Chen, Victoria Rubin, Sarah Cornwell

Data and Test Instruments

This paper outlines the complexity of the psychological construct of individuals' subjective well-being (SWB) and argues for the importance of examining behaviours and linguistic expression of individuals online social interactions in relation to self-reported SWB. This paper calls for a systematic review of the psychology research which examines SWB and its association with various character strengths, personality traits, and behaviours. While the Big Five personality traits (OCEAN) have an underlying neuropsychological basis and are considered as universal dimensions of personality along which humans differ one from another, minimal research has attempted to evaluate the relationship between personality traits, SWB, and …


Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin Jun 2022

Misinformation And Disinformation: Detecting Fakes With The Eye And Ai, Victoria Rubin

Data and Test Instruments

How do we detect, deter, and prevent the spread of mis- and disinformationwith the human eye and AI? How does theory inform the practice, and how do theevidence-based research and best practices in lie-catching and truth-seekingprofessions—inform AI? The book looks into well-established human practicessuch as the routines and processes used in detective work, journalism, and scientificinquiry, and how they contribute toward innovative AI solutions. The book explainsthe principles, inner workings, and recent evolution of five types of state-of-the-artAI technologies suitable for curtailing the spread of mis- and disinformation:automated deception detectors, clickbait detectors, satirical fake detectors, rumordebunkers, and computational fact-checking tools.


Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby Jun 2021

Predicting Stock Price Movements Using Sentiment And Subjectivity Analyses, Andrew Kirby

Dissertations, Theses, and Capstone Projects

In a quick search online, one can find many tools which use information from news headlines to make predictions concerning the trajectory of a given stock. But what if we went further, looking instead into the text of the article, to extract this and other information? Here, the goal is to extract the sentence in which a stock ticker symbol is mentioned from a news article, then determine sentiment and subjectivity values from that sentence, and finally make a prediction on whether or not the value of that stock will go up or not in a 24-hour timespan. Bloomberg News …


Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos May 2021

Plprepare: A Grammar Checker For Challenging Cases, Jacob Hoyos

Electronic Theses and Dissertations

This study investigates one of the Polish language’s most arbitrary cases: the genitive masculine inanimate singular. It collects and ranks several guidelines to help language learners discern its proper usage and also introduces a framework to provide detailed feedback regarding arbitrary cases. The study tests this framework by implementing and evaluating a hybrid grammar checker called PLPrepare. PLPrepare performs similarly to other grammar checkers and is able to detect genitive case usages and provide feedback based on a number of error classifications.


Quantifying Coherence In A Transdiagnostic Sample: A Methodological Investigation Of Computationally-Derived Coherence Using Ambulatory Assessment, Taylor L. Fedechko Mar 2019

Quantifying Coherence In A Transdiagnostic Sample: A Methodological Investigation Of Computationally-Derived Coherence Using Ambulatory Assessment, Taylor L. Fedechko

LSU Master's Theses

Schizophrenia is a clinical diagnosis assigned to individuals that experience positive (e.g., hallucinations and delusions), negative (e.g., blunted affect), and disorganized (e.g., incoherent speech) symptoms. One particularly disabling symptom is incoherence, which is defined as the meaning-based relationship between ideas. This symptom can drastically affect an individual’s quality of life by affecting areas such as social and occupational functioning. Currently, the mechanism behind this symptom is unknown and requires further study. One way to examine incoherence is to understand its level of expression in other clinical populations. With the advent of computationally-derived natural language processing (NLP), coherence can be quantified …


Multimodal Depression Detection: An Investigation Of Features And Fusion Techniques For Automated Systems, Michelle Renee Morales May 2018

Multimodal Depression Detection: An Investigation Of Features And Fusion Techniques For Automated Systems, Michelle Renee Morales

Dissertations, Theses, and Capstone Projects

Depression is a serious illness that affects a large portion of the world’s population. Given the large effect it has on society, it is evident that depression is a serious health issue. This thesis evaluates, at length, how technology may aid in assessing depression. We present an in-depth investigation of features and fusion techniques for depression detection systems. We also present OpenMM: a novel tool for multimodal feature extraction. Lastly, we present novel techniques for multimodal fusion. The contributions of this work add considerably to our knowledge of depression detection systems and have the potential to improve future systems by …


An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz Sep 2016

An Examination Of Cross-Domain Authorship Attribution Techniques, Maxwell B. Schwartz

Dissertations, Theses, and Capstone Projects

In recent years, Twitter has become a popular testing ground for techniques in authorship attribution. This is due to both the ease of building large corpora as well as the challenges associated with the character limit imposed by the service and the writing styles that have developed as a result. As both false and genuine claims of hacked Twitter accounts have made international news, there is an increasing need for this type of work. For newer Twitter accounts, however, there is little training data. Thus, this study looks to lay the groundwork for cross-domain authorship attribution: training on one source …


An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler Dec 2014

An Empirical Study Of Semantic Similarity In Wordnet And Word2vec, Abram Handler

LSU New Orleans Theses and Dissertations

This thesis performs an empirical analysis of Word2Vec by comparing its output to WordNet, a well-known, human-curated lexical database. It finds that Word2Vec tends to uncover more of certain types of semantic relations than others -- with Word2Vec returning more hypernyms, synonomyns and hyponyms than hyponyms or holonyms. It also shows the probability that neighbors separated by a given cosine distance in Word2Vec are semantically related in WordNet. This result both adds to our understanding of the still-unknown Word2Vec and helps to benchmark new semantic tools built from word vectors.


Predicting Music Genre Preferences Based On Online Comments, Andrew J. Sinclair Jun 2014

Predicting Music Genre Preferences Based On Online Comments, Andrew J. Sinclair

Master's Theses

Communication Accommodation Theory (CAT) states that individuals adapt to each other’s communicative behaviors. This adaptation is called “convergence.” In this work we explore the convergence of writing styles of users of the online music distribution plat- form SoundCloud.com. In order to evaluate our system we created a corpus of over 38,000 comments retrieved from SoundCloud in April 2014. The corpus represents comments from 8 distinct musical genres: Classical, Electronic, Hip Hop, Jazz, Country, Metal, Folk, and World. Our corpus contains: short comments, frequent misspellings, little sentence struc- ture, hashtags, emoticons, and URLs. We adapt techniques used by researchers analyzing other …


Csc Senior Project: Nlpstats, Michael Mease Mar 2013

Csc Senior Project: Nlpstats, Michael Mease

Computer Science and Software Engineering

Natural Language Processing has recently increased in popularity. The field of authorship analysis, specifically, uses various characteristics of text quantified by markers. NLPStats serves as a tool designed to streamline marker extraction based on user needs. A flexible query system allows for custom marker requests, adjustment of result formatting, and preprocessing options. Furthermore, an efficiently designed structure ensures that users retrieve information quickly. As a whole, NLPStats enables anyone, regardless of NLP experience, to extract important information about the text of a document.


Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee Nov 2010

Visual Salience And Reference Resolution In Situated Dialogues: A Corpus-Based Evaluation., Niels Schütte, John D. Kelleher, Brian Mac Namee

Conference papers

Dialogues between humans and robots are necessarily situated and so, often, a shared visual context is present. Exophoric references are very frequent in situated dialogues, and are particularly important in the presence of a shared visual context - for example when a human is verbally guiding a tele-operated mobile robot. We present an approach to automatically resolving exophoric referring expressions in a situated dialogue based on the visual salience of possible referents. We evaluate the effectiveness of this approach and a range of different salience metrics using data from the SCARE corpus which we have augmented with visual information. The …