Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

2025

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 211 - 240 of 504

Full-Text Articles in Data Science

Compartmental Disaggregation: Bridging Simulation And Sampling Methods For Synthetic Population Data Generation, Dylan Mack May 2025

Compartmental Disaggregation: Bridging Simulation And Sampling Methods For Synthetic Population Data Generation, Dylan Mack

McKelvey School of Engineering Graduate Student Theses & Dissertations

As agent-based models (ABMs) grow increasingly widespread in public health, their associated challenges have become all the more significant. Lauded for their ability to capture population heterogeneity, nonlinear dynamics, and emergent behaviors, disease ABMs are also computationally expensive and often require detailed inputs that describe each agent at the individual-level, known as synthetic population data. Current approaches for synthetic population data generation generally fall into one of two categories: sampling or simulation. These methods are both feasible only under restricted conditions and suffer from challenges surrounding data availability and computing power. This thesis proposes compartmental disaggregation, an intermediate method for …


Exploring The Evolution Of Global Warming Discourse: A Twitter-Based Analysis Across The United States, United Kingdom, And India, Trevor Christensen Apr 2025

Exploring The Evolution Of Global Warming Discourse: A Twitter-Based Analysis Across The United States, United Kingdom, And India, Trevor Christensen

Honors Projects in Data Science

Global warming has gained increasing attention over the past decade, with public discourse intensifying on social media platforms, particularly on Twitter. This increased discussion stems from political controversies surrounding climate change and the rise in extreme weather events. This study explores the evolution of global warming discourse on Twitter, with a focus on the United States, United Kingdom, and India. This research used a dataset of historical tweets ranging from 2010 to 2023 containing the keyword "global warming." Using natural language processing (NLP) techniques such as emotion analysis, word cloud visualization, and topic modeling (LDA), approximately twenty-eight million tweets were …


Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins Apr 2025

Mortgage Default Classification Modeling For Variable Analysis, Brendan R. Goggins

Honors College Theses

The financial crisis of the early 2000’s is a prime example of the severe consequences that mortgage default and borrower insolvency can have on economies at large. Mortgage default specifically is a prime case with the popularization of mortgage backed securities and the commonality of this loan structure. Multiple hypotheses and models have been formed to understand the reasons, causes, and consequences of mortgage default. This paper uses both machine learning and statistical classification models to inform an understanding of the variables most significant and impactful to the default outcome of mortgages. Consideration is given to both loan-level microeconomic variables …


Rewriting War: Improving The Mlb’S Go-To Advanced Metric, Kolin Atwood Apr 2025

Rewriting War: Improving The Mlb’S Go-To Advanced Metric, Kolin Atwood

Honors Projects

Traditional Wins Above Replacement (WAR) metrics have long served as a cornerstone of player evaluation in Major League Baseball, offering a context-neutral summary of offensive, defensive, and baserunning contributions. However, this neutrality often overlooks critical factors such as game situation, lineup strength, and advanced baserunning impact. This project proposes an enhanced model, WAR-PC (Wins Above Replacement – Plus Context), that integrates three key improvements: context-dependent batting value (RE24), clutch performance (Win Probability Added, WPA), and Statcast-based baserunning metrics. Using R, player logs, and modern baseball data sources, WAR-PC was calculated for eight players from the 2023 MLB season. The revised …


The Impact Of Institutional Features On Student Retention Rates Using Regression And Random Forest Modeling, Grayson Tvrdik Apr 2025

The Impact Of Institutional Features On Student Retention Rates Using Regression And Random Forest Modeling, Grayson Tvrdik

Honors Projects

Student retention is a focus for higher education institutions aiming to improve student outcomes and institutional success. While previous research has often relied on qualitative assessments of college related factors, this project applies quantitative techniques at a national scale. Random forest and beta regression models were used to predict retention rates for public colleges based on institutional characteristics such as financial variables, enrollment patterns, and demographic metrics. The random forest models demonstrated higher accuracy than the beta regression models, leading us to find that financial variables and student integration factors are significant predictors of retention. Beta regression models, though less …


A Comparison Of The Bacterial Communities Of The Yamuna River (India) And Mississippi River (Usa), Jacob Gareis, Osvaldo Martinez, Silas Bergen Apr 2025

A Comparison Of The Bacterial Communities Of The Yamuna River (India) And Mississippi River (Usa), Jacob Gareis, Osvaldo Martinez, Silas Bergen

Research & Creative Achievement Day

This study compared the bacteria in the Yamuna River in India and the Mississippi River in the USA. Water samples were taken from 11 sites (2 in the Mississippi River and 9 in the Yamuna River). The bacteria were identified using 16S rRNA gene sequencing. Some phyla such as Proteobacteria and Bacteroidetes were seen in abundance regardless of country. However, principal-component analysis showed three distinct groupings: Mississippi/Ganga/Tons and other Yamuna River locations (7 sites), Yamuna River at Delhi (2 sites), Yamuna River below Glacier (2 sites) Additionally, the Yamuna River below Glacier had the highest bacterial diversity, with a Shannon …


Evaluating The Causal Effect Of Receiving Transthoracic Echocardiography On 28-Day Mortality In Mimic-Iii Intensive Care Unit Patients With Sepsis, Monserrath Velez, Jack O'Connor, Nicholas Della Pesca, Rio Baliga Apr 2025

Evaluating The Causal Effect Of Receiving Transthoracic Echocardiography On 28-Day Mortality In Mimic-Iii Intensive Care Unit Patients With Sepsis, Monserrath Velez, Jack O'Connor, Nicholas Della Pesca, Rio Baliga

Research & Creative Achievement Day

❏ Sepsis (an infection associated with vital organ dysfunction) is an emergent

medical condition estimated to occur in about 30% of intensive care unit

(ICU) patients[1] and is responsible for 20% of all deaths worldwide[2]

❏ Transthoracic echocardiography (TTE; an imaging modality that uses

ultrasound technology to record heart structure and function) is widely used in

medical treatment [3]

❏ 28-day mortality (whether or not a patient dies within 28 days after receiving a

treatment) is a measure considered to closely approximate ICU mortality[4]

Objectives

❏ Assess and quantify the causal effect of receiving a TTE on …


Analyzing The Sentiment Of Feminist And Non-Feminist Works, Jasmine Borie, Megan G. Falschlehner Apr 2025

Analyzing The Sentiment Of Feminist And Non-Feminist Works, Jasmine Borie, Megan G. Falschlehner

Mathematics, Computer Science & Statistics Presentations

This presentation focuses on a group of texts that advocate for a change in the current belief system. These texts are the Feminist Manifesto, Sojourner Truth: Ain’t I a Woman?, and Civilization and Its Discontents. These first two texts advocate for women’s rights, while Freud’s book is focused on civilization’s decline and how our understanding of community can affect this. Through our presentation, we want to examine the differences in sentiment and language between the feminist texts and Freud’s texts to pinpoint whether or not sentiment changes when advocating for different beliefs.


Analyzing Cie Texts Through History Using R, Rachel A. Hart, Aaron Ditto Apr 2025

Analyzing Cie Texts Through History Using R, Rachel A. Hart, Aaron Ditto

Mathematics, Computer Science & Statistics Presentations

In this presentation, we analyzed three separate CIE texts from different time periods. First, “The Allegory of the Cave” from 380 BC, then “The Declaration of Independence” from 1776, and lastly “The Lottery” from 1948. We compared them using tidy text techniques like sentiment lexicons, creating word clouds, and bigram analysis to see if the types of words and sentiments used have changed over time in these short texts.


A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn Apr 2025

A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn

Mathematics, Computer Science & Statistics Presentations

The purpose of this project was to discover similarities between sentiments in Old Testament and New Testament books of the Bible, track emotional valence and find the most common words and sentiments in the books. Text analysis was performed on Genesis, Exodus, Matthew and Luke. Word clouds were also created for these texts.


Application For Prediction Of Heart Failure; The Next Step In Machine Learning For Healthcare, Amy Adyanthaya, Dawn Bowerman, Rachel Liercke, Robert Slater Apr 2025

Application For Prediction Of Heart Failure; The Next Step In Machine Learning For Healthcare, Amy Adyanthaya, Dawn Bowerman, Rachel Liercke, Robert Slater

SMU Data Science Review

Heart failure (HF) is a serious medical condition affecting approximately 6.7 million U.S. adults and is expected to impact 8.5 million Americans by 2030 [1]. Heart failure is a complicated clinical ailment and characterizes the final course of numerous heart diseases [2]. This paper introduces a machine-learning-based application that utilizes Support Vector Machine (SVM), Multi-Layer Perceptron (MLP), and XGBoost models, implemented through the Python Flask framework, to predict HF risk using clinical data. The results indicate high model performance, with precision and recall metrics underscoring the application’s reliability in identifying at-risk patients. By providing real-time, accessible insights, this tool aims …


Multi-Agent Translation Team (Matt): Enhancing Low-Resource Language Translation Through Multi-Agent Workflow, Anishka Peter, Mai Dang, Michael Liu, Joaquin Dominguez, Nibhrat Lohia Apr 2025

Multi-Agent Translation Team (Matt): Enhancing Low-Resource Language Translation Through Multi-Agent Workflow, Anishka Peter, Mai Dang, Michael Liu, Joaquin Dominguez, Nibhrat Lohia

SMU Data Science Review

Like humans, large language models (LLMs) benefit from revision and refinement, especially for complex tasks requiring critical thinking. Inspired by human collaborative problem-solving, this study introduces a novel multi-agent workflow designed to enhance LLM translations from English to low-resource languages. Multi-Agent Translation Team (MATT) involves the collaboration of agents that are assigned specific roles, such as translator, evaluation coordinator, and various levels of editing, to refine the initial translation into the most desired version possible. The agents work collaboratively in an iterative loop until the translation loss meets a satisfactory threshold. It stands out from other multi-agent workflows by combining …


Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun Apr 2025

Enhancing Animal Shelter Operations With Time Series And Machine Learning, Sakava L. Kiv, Donald L. Anderson, Shivam Negi, Jacquelyn Cheun

SMU Data Science Review

Enhancing animal shelter operations through machine learning involves employing a variety of advanced techniques aimed at increasing efficiency, promoting animal welfare, and optimizing resource allocation. This paper explores predictive analytics for adoption rates using regression models to estimate the likelihood of adoption based on historical data, encompassing variables such as breed, health status, and previous adoption trends. Additionally, classification algorithms are utilized to categorize animals by adoption probability, facilitating better resources and marketing prioritization. Clustering algorithms are employed to group animals according to behavior patterns and/or physical health, enabling tailored medical care and enrichment activities that improve their mental and …


Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan Apr 2025

Enhancing Network Security Through Dual-Layer Log Analysis: Integrating Machine Learning Classifiers With Large Language Models For Intelligent Anomaly Detection, Anthony Burton-Cordova, O'Neil Gray, Mohammad Al Rousan

SMU Data Science Review

This paper presents an innovative approach to enhancing network security by integrating machine learning algorithms with fine-tuned large language models (LLMs) to provide an expert assistant querying. The proposed method utilizes machine learning for efficient preprocessing and feature extraction from log data, followed by the application of a fine-tuned LLM to analyze and interpret anomalies with greater accuracy. This dual-layer detection system is designed to improve the identification of subtle and sophisticated security threats. The research team’s extensive evaluation using real-world log datasets indicates that the combined approach increases detection rates and communicates results in an understandable manner, demonstrating its …


Statistical Study Of Solar Wind Conditions Prior To Substorm Onsets, Luke H. Francis Apr 2025

Statistical Study Of Solar Wind Conditions Prior To Substorm Onsets, Luke H. Francis

Doctoral Dissertations and Master's Theses

Due to complex, multi-region, coupled plasma systems, auroral substorm onsets have been historically difficult to predict. The northward turning of the interplanetary magnetic field was considered the primary candidate as an external triggering mechanism for substorm onsets. However, that was later shown to be coincidental in nature. This study is motivated by recent multi-spacecraft observations that show how several magnetosheath jets at the bow shock were heavily correlated to substorm onsets, indicated by a strongly radial IMF interval. In the past, studies have looked at small samples of substorms in order to make large-scale predictions. However in this study, a …


Analyzing Musical Emotions: A Multi-Dataset Approach To Sentiment And Mood Classification In Songs, Mahad Syed Apr 2025

Analyzing Musical Emotions: A Multi-Dataset Approach To Sentiment And Mood Classification In Songs, Mahad Syed

SPARK Symposium Presentations

Music evokes a wide range of emotions, yet most music recommendation systems focus on sound and listening patterns rather than the meaning of lyrics. This project enhances lyric-based emotion recognition by applying Natural Language Processing (NLP) and Machine Learning (ML) to classify song lyrics into emotional categories.

I used eight datasets from Kaggle, including collections of lyrics, emotion labels, and audio features, providing a strong foundation for analysis. Our approach combines traditional NLP techniques (like TF-IDF and Word2Vec) with advanced deep learning models (such as BERT and XLNet) to classify lyrics into categories like happy, sad, angry, calm, romantic, and …


Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul Apr 2025

Diabetes: Non-Invasive Blood Glucose Monitoring Using Federated Learning With Biosensor Signals, Narmatha Chellamani, Saleh Ali Albelwi, Manimurugan Shanmuganathan, Palanisamy Amirthalingam, Anand Paul

School of Public Health Faculty Publications

Diabetes is a growing global health concern, affecting millions and leading to severe complications if not properly managed. The primary challenge in diabetes management is maintaining blood glucose levels (BGLs) within a safe range to prevent complications such as renal failure, cardiovascular disease, and neuropathy. Traditional methods, such as finger-prick testing, often result in low patient adherence due to discomfort, invasiveness, and inconvenience. Consequently, there is an increasing need for non-invasive techniques that provide accurate BGL measurements. Photoplethysmography (PPG), a photosensitive method that detects blood volume variations, has shown promise for non-invasive glucose monitoring. Deep neural networks (DNNs) applied to …


Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh Apr 2025

Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh

Computer Science ETDs

Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …


Data Science For Engineers, Heidi Moulton Apr 2025

Data Science For Engineers, Heidi Moulton

Student Research Symposium

30% of USU undergraduate students participate in some sort of research, and for engineering students this often means generating large amounts of data.

Data Science for Engineers is a series of four modules that introduce students to data processing, visualization, and graphing in the Python programming language using Pandas DataFrames and Juypter Notebooks.

The modules are intended for students with a basic understanding of programming in Python, specifically those who have taken CS 1400 Introduction to Computer Science.


A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss Apr 2025

A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss

Campus Research Month

We developed a machine-learning tool-supported methodology for modeling the nonprofit donor relationship. This approach was demonstrated in the case of a US-based nonprofit. Conclusions were drawn from this example and tool-support provided for use by other nonprofits.


From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie Apr 2025

From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie

Undergraduate Theses

Adversarial attacks pose a significant threat to the reliability of machine learning-based spam detection systems in social media. This undergraduate thesis, "From Adversarial Attacks to Robust Classifiers: A Study in Social Media Spam Detection – Black Box & White Box," systematically examines the impact of both black-box and white-box adversarial attacks on a range of spam classifiers, including Logistic Regression, Decision Trees, Random Forests, K-Nearest Neighbors, Bagging, Gradient Boosting, and Support Vector Machines. Leveraging a novel dataset derived from Twitter spam messages and enhanced with adversarial perturbations such as synonym replacement and character-level modifications, this study evaluates classifier performance under …


36 - Investigation Of The Digital Footprint Of Scientific Research In Social Media – Preliminary Findings, Lee Logan, Dominik Soos, Sean Baker, Jian Wu Apr 2025

36 - Investigation Of The Digital Footprint Of Scientific Research In Social Media – Preliminary Findings, Lee Logan, Dominik Soos, Sean Baker, Jian Wu

Undergraduate Research Symposium

Title: Investigation of The Digital Footprint of Scientific Research in Social Media – Preliminary Findings

Authors: Lee Logan, Sean Baker, Dominik Soos, Jian Wu

The spread of scientific information and research beyond the confines of academic institutions plays a central role in how the public understands and trusts modern sciences. Social media has become an essential means of dissemination for scholarly news, papers, and other forms of engagement. This research aims to explore how scientific research is disseminated over social media to understand its role as a bridge between peer-reviewed research and the public's overall understanding. To support the research …


Opioid Vs. Money Choice Preference Patterns In Regular Heroin Users, Amolak S. Jhand, Mark Greenwald Apr 2025

Opioid Vs. Money Choice Preference Patterns In Regular Heroin Users, Amolak S. Jhand, Mark Greenwald

Medical Student Research Symposium

About two-thirds of people treated for opioid use disorder (OUD) return to opioid use within the first-year post-treatment, and about 10% report use while on agonist therapy. Understanding determinants of opioid-seeking is vital to reducing recurrence and its risks. We assessed individual differences in effortful choices between opioid and money amounts, modeling real-world choices.

Our lab conducted studies in which regular heroin-users were stabilized on buprenorphine to suppress withdrawal. Within experimental sessions, the participant could choose repeatedly across 12 trials between units of hydromorphone (HYD, 1 or 2 mg IM) vs. money ($2 or $4); HYD and money amounts differed …


Identifying Saharan Air Layer Events And Their Relationship To Instability And Rainfall Patterns In The Tropical North Atlantic Ocean, Charles H. Dolce Apr 2025

Identifying Saharan Air Layer Events And Their Relationship To Instability And Rainfall Patterns In The Tropical North Atlantic Ocean, Charles H. Dolce

LSU Master's Theses

Every year, strong North African winds across the Sahara loft dust particles and advect them westward within the Saharan Air Layer (SAL). These plumes of dust often reach the Caribbean Sea and can heavily affect regional weather patterns by suppressing convection. When concentrations are high, decreased rainfall can yield negative societal and ecological impacts, potentially leading to drought. The Puerto Rican early rainfall season (ERS), spanning from 1 April through 31 July, overlaps with Saharan dust migration periods, and systematically detecting instances of SAL activity near the island is important for understanding longer-term drought-forcing trends in the region.

Corridors of …


Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova Apr 2025

Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova

Engineering Faculty Articles and Research

Fire Influence on Regional to Global Environments and Air Quality (FIREX-AQ) was a field campaign aimed at better understanding the impact of wildfires and agricultural fires on air quality and climate. The FIREX-AQ campaign took place in August 2019 and involved two aircraft and multiple coordinated satellite observations. This study applied and evaluated a self-supervised machine learning (ML) method for the active fire and smoke plume identification and tracking in the satellite and sub-orbital remote sensing datasets collected during the campaign. Our unique methodology combines remote sensing observations with different spatial and spectral resolutions. With as much as a 10% …


Skating For A Payday: Analyzing Nhl Player Performance In Contract Years, Mike Anthony Dibenedetto Apr 2025

Skating For A Payday: Analyzing Nhl Player Performance In Contract Years, Mike Anthony Dibenedetto

Honors Projects in Economics

This research investigates the contract year phenomenon in the National Hockey League (NHL) by analyzing player performance during contract years rather than salary outcomes. Through an extensive literature review and empirical analysis, I examine whether NHL players exhibit statistically significant differences in performance during the final year of their contracts. Drawing from a broad range of studies on performance motivation, team systems, free agency, arbitration rights, and the role of advanced analytics, this paper synthesizes existing theories with new regression-based evidence.

Using a dataset of 2,717 player-season observations, I evaluate the effect of contract year status on key performance metrics, …


Tackling Crime: A Data-Driven Comparison Of Nfl Players With The General Population, Ryan Piersza Apr 2025

Tackling Crime: A Data-Driven Comparison Of Nfl Players With The General Population, Ryan Piersza

Honors Projects in Data Science

This research investigates the frequency of arrests among NFL players compared to the general population and analyze the types of crimes committed. A main theme from the literature is that previous groups of NFL players were found to be arrested at a rate lower than the general population. However, there has been no analysis on how the COVID-19 pandemic impacted the crime rates and types. Data obtained from the FBI’s Uniform Crime Report, The U.S Census Bureau, an NFL arrest database, and player statistics from Pro Football Reference is used for the analysis. The data is analyzed with a rate …


Resurrecting The Past: Engineering Ancestral Enzymes For Modern Bacterial Applications, Gregory M. Gladkowski Iii, Masa Watanabe Apr 2025

Resurrecting The Past: Engineering Ancestral Enzymes For Modern Bacterial Applications, Gregory M. Gladkowski Iii, Masa Watanabe

SACAD: Scholarly Activities

Ancestral Sequence Reconstruction (ASR) is a computational technique that infers and resurrects ancient protein sequences to explore molecular evolution and enable protein engineering. By integrating phylogenetics and statistical modeling, ASR produces stable, mutation-tolerant enzymes ideal for directed evolution. This poster presents a workflow for reconstructing enzymes from pathogenic bacteria and highlights ASR’s value in uncovering novel functions and advancing applications in drug discovery and synthetic biology.


Cogprog: Utilizing Large Language Models To Forecast In-The-Moment Health Assessment, Gina Sprint, Maureen Schmitter-Edgecombe, Raven Weaver, Lisa Wiese, Diane Cook Apr 2025

Cogprog: Utilizing Large Language Models To Forecast In-The-Moment Health Assessment, Gina Sprint, Maureen Schmitter-Edgecombe, Raven Weaver, Lisa Wiese, Diane Cook

Computer Science Faculty Scholarship

Forecasting future health status is beneficial for understanding health patterns and providing anticipatory support for cognitive and physical health difficulties. In recent years, generative Large Language Models (LLMs) have shown promise as forecasters. Though not traditionally considered strong candidates for numeric tasks, LLMs demonstrate emerging abilities to address various forecasting problems. They also provide the ability to incorporate unstructured information and explain their reasoning process. In this article, we explore whether LLMs can effectively forecast future self-reported health state. To do this, we utilized in-the-moment assessments of mental sharpness, fatigue, and stress from multiple studies, utilizing daily responses (N = …


Extending Feature-Based Detection For Artificial Intelligence, Kayla Ahrndt Apr 2025

Extending Feature-Based Detection For Artificial Intelligence, Kayla Ahrndt

SPARK Symposium Presentations

AI text generation is rapidly developing, and, as a result, it is becoming increasingly difficult to differentiate it from human written text. Our base study by Leon Fröhling et al. proposed a feature-based detection model trained on GPT2, GPT3, and Grover data, as well as human-generated text. Our work extends their research by training a modified model with four neural networks on word embeddings, select features from the original study, as well as updated data (GPT3, GPT4, and Grover).