Open Access. Powered by Scholars. Published by Universities.®

Natural Language Processing

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 90

Full-Text Articles in Artificial Intelligence and Robotics

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma May 2026

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool (Ila 2026 Presentation), Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

This presentation introduces Propasafe-Hybrid, a hybrid system for sentence-level propaganda detection that combines offline transformer-based classification with selective large language model (LLM) explainability. The system employs a two-stage pipeline in which a local BERT-based classifier evaluates all input text and filters non-propagandistic content, while only high-confidence candidates are forwarded to an LLM for rhetorical technique labeling and explanation. This design enables cost-aware, privacy-conscious, and scalable analysis by reducing unnecessary reliance on external models.

Propasafe-Hybrid identifies propagandistic techniques such as loaded language, obfuscation, and appeal to fear, and generates concise natural language rationales that make these techniques interpretable to users. By …


Towards Multi-Hop Retrieval Using Bipartite Question-Oriented Graphs, Micah Mccollum May 2026

Towards Multi-Hop Retrieval Using Bipartite Question-Oriented Graphs, Micah Mccollum

Electrical Engineering and Computer Science Undergraduate Honors Theses

Accurately answering multi-hop questions requires full retrieval of multiple, interdependent passages and is a long-standing problem in the area of natural language question answering (QA). While retrieval-augmented generation (RAG) helps address single-hop questions, many retrievers presently focus on semantic similarity in a dense vector space, which is insufficient for handling multi-hop questions specifically. To ameliorate this, we propose constructing a bipartite question- oriented graph composed of hypothetically generated questions connected to passages at index time. The construction of the graph is guided by a large language model (LLM) to prioritize the formation of edges that signal whether a question can …


Using Ai For Data Loss Prevention, Camden A. Wright May 2026

Using Ai For Data Loss Prevention, Camden A. Wright

Theses/Capstones/Creative Projects

Data Loss Prevention (DLP) systems play a critical role in protecting modern systems that handle sensitive information from both accidental and malicious exposure. Traditional DLP approaches often rely on static rules and methods that can struggle to adapt to complex and evolving data patterns. This paper presents a hybrid DPL system that integrates machine learning-based message classification, rule based policy enforcement, and context-aware access control to improve both detection accuracy and decision reliability. In addition, the system introduces a second stage access control model that evaluates user context, including role of clearance level and job title to determine whether access …


Synthetic Dataset For Understanding Negation In Text-Guided Image Editing, Nhat-Tan Bui Dec 2025

Synthetic Dataset For Understanding Negation In Text-Guided Image Editing, Nhat-Tan Bui

Graduate Theses and Dissertations

Negation is a fundamental linguistic concept used by humans to convey information that they do not desire. Despite this, minimal research has focused on negation within text-guided image editing. This lack of research means that vision-language models (VLMs) for image editing may struggle to understand negation, implying that they struggle to provide accurate results. One barrier to achieving human-level intelligence is the lack of a standard collection by which research into negation can be evaluated. This thesis presents the first large-scale dataset, Negative Instruction (NeIn), for studying negation within instruction-based image editing. Our dataset comprises 366,957 quintuplets, i.e., source image, …


When It Comes To Scientific Information Extraction And Llms, Less Is More, Sameer Shaik Nov 2025

When It Comes To Scientific Information Extraction And Llms, Less Is More, Sameer Shaik

Theses and Dissertations from DePaul University

The scientific literature continues to expand rapidly, making manual extraction of structured scientific facts increasingly impractical. Traditional Machine Learning and Natural Language Processing (NLP) pipelines require large expert-annotated datasets, which are costly to produce. Novel Large Language Models (LLMs) face challenges in long-context scientific reasoning, hallucinations, and entity linking. This thesis investigates ELSIE-Blob, a domain-aware preprocessing method that segments scientific articles into compact text “blobs” containing components of entity relations (here, polymer names, melting point indicators, and numerical values). We test whether blob-based input allows lightweight, consumer-hardware-accessible LLMs to extract polymer–melting point (polymer–Tm) pairs accurately without training data. Experiments using …


The Challenge Of Achieving Attributability In Multilingual Table-To-Text Generation With Question-Answer Blueprints, Aden Haussmann Oct 2025

The Challenge Of Achieving Attributability In Multilingual Table-To-Text Generation With Question-Answer Blueprints, Aden Haussmann

International Journal of Undergraduate Research and Creative Activities

Generating faithful text descriptions from data tables is a significant challenge in Natural Language Processing (NLP), especially for the world’s many low-resource languages. This paper investigates whether Question-Answer (QA) blueprints—an intermediate planning step where a model first asks and answers questions about the data—can improve the factual accuracy of multilingual table-to-text generation. This novel approach is tested on the TaTA dataset, which includes several African languages, by finetuning models with and without these blueprints.

The results show a key distinction: while the QA blueprint method improves performance for English-only models, these gains disappear in the multilingual setting. This paper’s analysis …


A Web Application For Generating Argument Maps For Essays Using Llms, Alexis R. Chalmers Jul 2025

A Web Application For Generating Argument Maps For Essays Using Llms, Alexis R. Chalmers

Theses

Argumentative writing is a critical skill that strengthens students’ reasoning, communication, and analytical abilities. However, maintaining a clear and organized argument structure while writing can be challenging. Argument maps — visual diagrams which explicitly show an argument’s structure — have been shown to improve students’ writing, but are rarely used outside of the planning stage of an essay due to the time and effort required to create them. Automatically generating argument maps from student essays helps students to evaluate the structure of their argument as they write and makes identifying unsupported claims visible. To evaluate whether large language models (LLMs) …


Sepsis: I Can Catch Your Lies – A New Paradigm For Deception Detection, Anku Rani, Dwip Dalal, Shreya Gautam, Pankaj Gupta, Vinija Jain, Aman Chadha, Amitava Das, Amit P. Sheth Jul 2025

Sepsis: I Can Catch Your Lies – A New Paradigm For Deception Detection, Anku Rani, Dwip Dalal, Shreya Gautam, Pankaj Gupta, Vinija Jain, Aman Chadha, Amitava Das, Amit P. Sheth

Publications

Deception is the intentional practice of twisting information. It is a nuanced societal practice deeply intertwined with human societal evolution, characterized by a multitude of facets. This research explores the problem of deception through the lens of psychology, employing a framework that categorizes deception into three forms: lies of omission, lies of commission, and lies of influence. The primary focus of this study is specifically on investigating only lies of omission. We propose a novel framework for deception detection leveraging NLP techniques. We curated an annotated dataset of 876,784 samples by amalgamating a popular large-scale fake news dataset and scraped …


Optimizing Small Ai Models For Biomedical Tasks Through Efficient Knowledge Transfer From Large Domain Models, Girish Sundaram May 2025

Optimizing Small Ai Models For Biomedical Tasks Through Efficient Knowledge Transfer From Large Domain Models, Girish Sundaram

Theses and Dissertations

The PICO (Population, Intervention, Comparison, Outcome) framework is a widely adopted methodology for structuring clinical research questions and extracting relevant information from unstructured medical texts. However, traditional approaches for PICO classification demand computationally expensive domain-specific language models, such as BioBERT and ClinicalBERT, which require extensive training and large annotated datasets. This dissertation introduces Distilled Rapid Embedding Transfer (DRET), a novel knowledge transfer method designed to enable resource-constrained domain adaptation. DRET aims to efficiently transfer biomedical domain knowledge from large, specialized models to a compact, general-purpose model, DistilBERT, thereby enhancing its ability to perform domain-specific tasks without access to the original …


Using Natural Language Processing And Machine Learning To Detect Online Radicalisation In The Maldivian Language, Dhivehi, Hussain Ibrahim, Ahmed Ibrahim, Michael N. Johnstone May 2025

Using Natural Language Processing And Machine Learning To Detect Online Radicalisation In The Maldivian Language, Dhivehi, Hussain Ibrahim, Ahmed Ibrahim, Michael N. Johnstone

Research outputs 2022 to 2026

Early detection of online radical content is important for intelligence services to combat radicalisation and terrorism. The motivation for this research was the lack of language tools in the detection of radicalisation in the Maldivian language, Dhivehi. This research applied Machine Learning and Natural Language Processing (NLP) to detect online radicalisation content in Dhivehi, with the incorporation of domain-specific knowledge. The research used Machine Learning to evaluate the most effective technique for detection of radicalisation text in Dhivehi and used interviews with Subject Matter Experts and self-deradicalised individuals to validate the results, add contextual information and improve recognition accuracy. The …


Investigating Key Structures In Protective Scenes For Llms, Eben M. Weisman Mar 2025

Investigating Key Structures In Protective Scenes For Llms, Eben M. Weisman

University Honors Theses

This research delves into the realm of "protective scenes" within Large Language Models (LLMs), exploring their impact on bias mitigation, deception, and context preservation. The study investigates the use of roleplay prompting human-like behavior and reasoning in LLMs, focusing on the Character-LLM framework's concept of protective scenes with graduated levels of protection. By combining insights from psychology, cognitive science, and computational analysis, this research aims to develop a framework for understanding how protective scenes influence roleplay performance in LLMs, ultimately contributing to the development of more reliable and ethical AI systems.


Language Processing: The Precedence Of Neural Networks On The Account Of Hidden Markov Models, Dia Eddin Abuzeina Mar 2025

Language Processing: The Precedence Of Neural Networks On The Account Of Hidden Markov Models, Dia Eddin Abuzeina

An-Najah University Journal for Research - B (Humanities)

Background: since its discovery at the beginning of the last century, Markov models gain a great popularity, and have been widely used in different domains. However, the most prominent use was in computational linguistics, or what is known as natural language processing (NLP). Abstractly, Markov models are nothing but a statistical representation of a particular system. The mathematical statistical representation of a given system is the heart of Markov theory. Markov models characterized by solid mathematical representation, which significantly promotes using it. No doubt, Markov models are mainly used in prediction and classification, to serve computational linguistics as well as …


Performing Requirements Specification And Analysis Through Open Generative Pre-Trained Transformers, Harvey J. Hurst Mar 2025

Performing Requirements Specification And Analysis Through Open Generative Pre-Trained Transformers, Harvey J. Hurst

Theses and Dissertations

Every acquisition program begins with a requirement, and for those programs to succeed, robust requirements engineering (RE) must be implemented. RE encompasses eliciting, analyzing, specifying, and validating requirements—a critical process throughout a program's lifecycle. Despite its importance, RE faces challenges such as scope creep, ambiguity, redundancy, and inadequate automation support, often exacerbated by reliance on historical data. To address these issues, this thesis leverages advancements in Generative Technology, particularly large language models (LLMs) such as Generative Pre-Trained Transformers (GPTs). This research developed two GPT-based tools: the Single Requirement Analysis Tool and the Set of Requirements Analysis Tool. These tools were …


Incorporating Visual Information Into Natural Language Processing, Maxwell Mbabilla Aladago Jan 2025

Incorporating Visual Information Into Natural Language Processing, Maxwell Mbabilla Aladago

Dartmouth College Ph.D Dissertations

Natural language describes entities in the world, some real and some abstract. It is also common practice to complement human learning of natural language with visual cues. This is evident in the heavily graphical nature of children’s literature which underscores the importance of visual cues in language acquisition. Similarly, the notion of “visual learners” is well recognized, reflecting the understanding that visual signals such as illustrations, gestures, and depictions effectively supplement language. In machine learning, two primary paradigms have emerged for training systems involving natural language. The first paradigm encompasses setups where pre-training and downstream tasks are exclusively in natural …


Intergenerational Classification Of Reddit Comments Based On Slang And Emoji Usage, James T. Dracup Jan 2025

Intergenerational Classification Of Reddit Comments Based On Slang And Emoji Usage, James T. Dracup

West Chester University Master’s Theses

The rapid evolution of language, driven by technological advancements, has created notable cultural gaps between generations, particularly in how they communicate. This gap is most apparent in the growing use of slang and emojis among younger generations. This study aims to explore whether Reddit comments can be classified by generation based on the usage of slang and emojis, the frequency of their use across generations, and how such features (slang and emojis) might influence the meaning of traditional language. Using Reddit’s API, we collected comments from four generational subreddits and applied various machine learning models, Naïve Bayes, Neural Networks, and …


Pca Text Sentiment Analysis Tool, Luke Gegick Jan 2025

Pca Text Sentiment Analysis Tool, Luke Gegick

Williams Honors College, Honors Research Projects

This project applies principal component analysis (PCA) to sentiment analysis of text to identify complex emotional responses from plain text. Existing sentiment analysis tools often rely on large language models or struggle to achieve high accuracy when processing large collections of short inputs, such as social media comments. By contrast, this project uses PCA as a lightweight, mathematically grounded alternative that can scale efficiently while still capturing meaningful emotional structure in text data.

PCA has shown strong effectiveness in text analysis, particularly when supported by a sufficiently large dataset and a robust preprocessing pipeline. To create consistent, information-rich input vectors, …


Closed Domain Question Answering With Language Models: Application Of Retrieval-Augmented Generation And Parameter Efficient Fine-Tuning In Healthcare, Aaron Cummings Dec 2024

Closed Domain Question Answering With Language Models: Application Of Retrieval-Augmented Generation And Parameter Efficient Fine-Tuning In Healthcare, Aaron Cummings

Master's Theses

Dementia care presents significant challenges for informal caregivers, particularly in managing behavioral symptoms that affect over 90% of individuals with Alzheimer’s Disease and Related Dementias (ADRD) during the moderate-to-severe stages. These symptoms, including agitation, wandering, and repetitive activities, impose emotional and physical burdens on caregivers, often exacerbated by a lack of reliable, accessible, and personalized resources. Non-pharmacological interventions, while evidence-based, are underutilized due to knowledge gaps and the inefficiency of traditional training and information retrieval methods.

This research explores the adaptation of large language models (LLMs) to address these challenges by developing a framework for closed-domain Question Answering (QA) systems, …


Pixel: Ai Chatbot For Clear And Effective Senior Design Assistance, Asmin Pothula Dec 2024

Pixel: Ai Chatbot For Clear And Effective Senior Design Assistance, Asmin Pothula

2024 Fall Honors Capstone Projects - Archive

This research explores the development of an AI-driven chatbot named Pixel, specifically designed to assist Computer Science and Engineering Senior Design students by providing immediate, clear, and accurate responses to project-related queries. While my Senior Design project focuses on developing a "Senior Design Project Management Tool," my honors capstone project centers on developing Pixel and integrating it into both the project management tool and the CSE Senior Design Knowledge Base. Pixel leverages this knowledge base to offer guidance on tasks such as using lab equipment, performing technical procedures, and troubleshooting common issues, ensuring that students have swift access to relevant …


Enhancing Low-Resource Language Performance In Multilingual Large Language Models, Mingqi Li Dec 2024

Enhancing Low-Resource Language Performance In Multilingual Large Language Models, Mingqi Li

All Dissertations

The large language models play an important role in many natural language tasks. However, training these models requires large amounts of data, which is not available for many languages. A noticeable performance gap exists between English and other languages, with low-resource languages showcasing this gap prominently. Therefore, it becomes imperative to improve large language models for low-resource languages. To address these challenges, we developed knowledge distillation and strategic prompt-learning, and attention alignment methods to improve the representation capabilities of large language models for low-resource language, and then enhanced their performance in downstream tasks.

In our first study, we developed a …


Enhancing Password Security And Memorability Using Machine Learning And Linguistic Patterns, Jared Wise Dec 2024

Enhancing Password Security And Memorability Using Machine Learning And Linguistic Patterns, Jared Wise

LSU New Orleans Theses and Dissertations

In the digital age, text-based passwords remain a primary method for securing online accounts. Yet, users frequently face a dilemma between creating passwords that are easy to remember and sufficiently secure against cyberattacks. This research introduces an approach to password generation that bridges this gap by utilizing linguistic patterns, particularly song lyrics, to develop highly secure and naturally memorable passwords. Using large lyric datasets gained from web scrapes from popular song lyric websites (AZ Lyrics, Genius), features are extracted from a corpus of over 5 million lyrics using sentence structure and natural language processing in a novel way. In using …


Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi Nov 2024

Vision-Language Integration For Enhanced Locomotion Mode Prediction, Ehsan Ahmadi

LSU Master's Theses

Wearable exoskeletons offer significant potential in enhancing human mobility in industrial environments. However, their adaptability to dynamic, task-intensive settings presents challenges, especially in accurately predicting locomotion modes such as ladder climbing, stair navigation, low-space movement, and obstacle navigation. This research proposes a multimodal framework that integrates visual data and speech commands to improve locomotion mode prediction in unpredictable environments. Multimodal data was collected using smart glasses, capturing both the user’s perspective (field-of-view, FOV) and voice during locomotion tasks. State-of-the-art models—CLIP, ImageBind, and GPT-4o—process these visual and linguistic inputs to predict locomotion activities. The models were evaluated in zero-shot and fine-tuned …


Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen May 2024

Representation Learning For Generative Models With Applications To Healthcare, Astronautics, And Aviation, Van Minh Nguyen

Theses and Dissertations

This dissertation explores applications of representation learning and generative models to challenges in healthcare, astronautics, and aviation.

The first part investigates the use of Generative Adversarial Networks (GANs) to synthesize realistic electronic health record (EHR) data. An initial attempt at training a GAN on the MIMIC-IV dataset encountered stability and convergence issues, motivating a deeper study of 1-Lipschitz regularization techniques for Auxiliary Classifier GANs (AC-GANs). An extensive ablation study on the CIFAR-10 dataset found that Spectral Normalization is key for AC-GAN stability and performance, while Weight Clipping fails to converge without Spectral Normalization. Analysis of the training dynamics provided further …


Performing Information Extraction For Mission Engineering Applications, Samuel R. Koski Apr 2024

Performing Information Extraction For Mission Engineering Applications, Samuel R. Koski

Engineering Management & Systems Engineering Theses & Dissertations

The process of extracting structured data from unstructured and semi-structured text is manual, time consuming and error prone. Current natural language processing approaches for automating this process are difficult to verify for non-trivial and context-sensitive corpora. Large Language Models (LLMs) like ChatGPT have become a subject of considerable interest, opening a promising avenue of exploration. However, there is limited evidence on the performance of LLMs for information extraction.

In this dissertation, an approach is proposed to evaluate the accuracy of Stanford OpenIE and OpenAI's ChatGPT for this purpose. This includes comparing Resource Description Framework (RDF) triples extracted by each of …


Artificial General Intelligence And The Mind-Body Problem: Exploring The Computability Of Simulated Human Intelligence In Light Of The Immaterial Mind, Caleb Parks Apr 2024

Artificial General Intelligence And The Mind-Body Problem: Exploring The Computability Of Simulated Human Intelligence In Light Of The Immaterial Mind, Caleb Parks

Senior Honors Theses

In this thesis I explore whether achieving artificial general intelligence (AGI) through simulating the human brain is theoretically possible. Because of the scientific community’s predominantly physicalist outlook on the mind-body problem, AGI research may be limited by erroneous foundational presuppositions. Arguments from linguistics and mathematics demonstrate that the human intellect is partially immaterial, opening the door for novel analysis of the mind’s simulability. I categorize mind-body problem philosophies in a manner relevant to computer science based upon state transitions, and determine their ramifications on mind-simulation. Finally, I demonstrate how classical architectures cannot resolve so-called Gödel statements, discuss why this inability …


Improving Automatic Transcription Using Natural Language Processing, Anna Kiefer Mar 2024

Improving Automatic Transcription Using Natural Language Processing, Anna Kiefer

Master's Theses

Digital Democracy is a CalMatters and California Polytechnic State University initia-
tive to promote transparency in state government by increasing access to the Califor-
nia legislature. While Digital Democracy is made up of many resources, one founda-
tional step of the project is obtaining accurate, timely transcripts of California Senate
and Assembly hearings. The information extracted from these transcripts provides
crucial data for subsequent steps in the pipeline. In the context of Digital Democracy,
upleveling is when humans verify, correct, and annotate the transcript results after
the legislative hearings have been automatically transcribed. The upleveling process
is done with the …


Using Natural Language Processing To Identify Mental Health Indicators In Aviation Voluntary Safety Reports, Michael Sawyer, Katherine Berry, Amelia Kinsella, R Jordan Hinson, Edward Bynum Feb 2024

Using Natural Language Processing To Identify Mental Health Indicators In Aviation Voluntary Safety Reports, Michael Sawyer, Katherine Berry, Amelia Kinsella, R Jordan Hinson, Edward Bynum

National Training Aircraft Symposium (NTAS)

Voluntary Safety Reporting Programs (VSRPs) are a critical tool in the aviation industry for monitoring safety issues observed by the frontline workforce. While VSRPs primarily focus on operational safety, report narratives often describe factors such as fatigue, workload, culture, staffing, and health, directly or indirectly impacting mental health. These reports can provide individual and organizational insights into aviation personnel's physical and psychological well-being. This poster introduces the AVIation Analytic Neural network for Safety events (AVIAN-S) model as a potential tool to extract and monitor these insights. AVIAN-S is a novel machine-learning model that leverages natural language processing (NLP) to analyze …


Kodai: Framework Towards Data Augmentation Of Large Language Models In Machine Learning, Rick Rejeleene Feb 2024

Kodai: Framework Towards Data Augmentation Of Large Language Models In Machine Learning, Rick Rejeleene

Theses and Dissertations

Machine Learning is rapidly advancing at an incredible pace due to an increase in computational size, and availability of data. Data is the cornerstone of machine learning algorithms. Information quality has a profound impact on the performance of machine learning systems. Supervised, Unsupervised, and Reinforcement learning are three major ways of performing machine learning. Self-supervised learning, part of unsupervised learning, has made breakthroughs in engineering and research by using Large Language Models (LLM). LLM are a type of neural network, used for language understanding and generation. Recently, LLM has taken a distinct lead by demonstrating state-of-the-art capabilities in natural language …


Using Natural Language Processing And Patient Journey Clustering For Temporal Phenotyping Of Antimicrobial Therapies For Cat Bite Abscesses, Brian Hur, Karin M. Verspoor, Timothy Baldwin, Laura Y. Hardefeldt, Caitlin Pfeiffer, Caroline Mansfield, Riati Scarborough, James R. Gilkerson Feb 2024

Using Natural Language Processing And Patient Journey Clustering For Temporal Phenotyping Of Antimicrobial Therapies For Cat Bite Abscesses, Brian Hur, Karin M. Verspoor, Timothy Baldwin, Laura Y. Hardefeldt, Caitlin Pfeiffer, Caroline Mansfield, Riati Scarborough, James R. Gilkerson

Natural Language Processing Faculty Publications

Background: Temporal phenotyping of patient journeys, which capture the common sequence patterns of interventions in the treatment of a specific condition, is useful to support understanding of antimicrobial usage in veterinary patients. Identifying and describing these phenotypes can inform antimicrobial stewardship programs designed to fight antimicrobial resistance, a major health crisis affecting both humans and animals, in which veterinarians have an important role to play. Objective: This research proposes a framework for extracting temporal phenotypes of patient journeys from clinical practice data through the application of natural language processing (NLP) and unsupervised machine learning (ML) techniques, using cat bite abscesses …


Using Ai For Qualitative Labeling: Consistency And Comparisons, James Mcintyre Jan 2024

Using Ai For Qualitative Labeling: Consistency And Comparisons, James Mcintyre

Honors Program Theses

This paper details a research study evaluating AI's ability to perform qualitative deductive coding. Multiple AI models were utilized and compared against three human coders and one expert coder. A series of 107 statements were sourced from a group discussion for a qualitative impact assessment of an organization. The AI models were provided these statements and directed to code them using the Community Capitals Framework. Two generations of AI models were evaluated. Overall, the AI achieved a fair level of agreement with the human annotators, but the alignment was far from perfect. Newer AI models did not increase agreement with …


Language Models For Rare Disease Information Extraction: Empirical Insights And Model Comparisons, Shashank Gupta Jan 2024

Language Models For Rare Disease Information Extraction: Empirical Insights And Model Comparisons, Shashank Gupta

Theses and Dissertations--Computer Science

End-to-end relation extraction (E2ERE) is a crucial task in natural language processing (NLP) that involves identifying and classifying semantic relationships between entities in text. This thesis compares three paradigms for end-to-end relation extraction (E2ERE) in biomedicine, focusing on rare diseases with discontinuous and nested entities. We evaluate Named Entity Recognition (NER) to Relation Extraction (RE) pipelines, sequence-to-sequence models, and generative pre-trained transformer (GPT) models using the RareDis information extraction dataset. Our findings indicate that pipeline models are the most effective, followed closely by sequence-to-sequence models. GPT models, despite having eight times as many parameters, perform worse than sequence-to-sequence models and …