Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

NLP

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 30

Full-Text Articles in Data Science

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma May 2026

Propasafe-Hybrid: A Text-Based Hybrid Propaganda Detection Tool, Thomas Kimmeth, Avijit Roy, Vivek Sharma

Publications and Research

Propagandistic content increasingly circulates through online news and social media, where readers often encounter it with limited scrutiny, highlighting the need for reliable and fine-grained detection. This paper introduces Propasafe-Hybrid, a sentence-level system that integrates a fine-tuned transformer classifier with LLM-based technique classification to identify, label, and explain specific propaganda strategies. The pipeline generates actionable outputs, including highlighted sentences, technique assignments, and concise rationales, so users can immediately understand why a sentence was flagged and how each label was determined. To control inference cost, Propasafe-Hybrid employs a cost-aware pre-filtering stage that forwards only high-likelihood sentences to LLMs, reducing token usage …


Nlp Bias And African American English, Kenya Roy, Faizan Javed Mar 2026

Nlp Bias And African American English, Kenya Roy, Faizan Javed

SMU Data Science Review

African American English (AAE), also referred to as African American Vernacular English (AAVE), is widely used on social media, but most sentiment analysis tools are trained only on Standard American English (SAE). This mismatch can cause models to misclassify dialectal expressions—especially by labeling neutral or positive AAE as negative or toxic. These errors matter, since Natural Language Processing (NLP) systems are now central to content moderation and brand monitoring. This research will evaluate the VADER, RoBERTa, GPT-OSS, and Gemma’s handling of AAE in comparison to SAE using the TwitterAAE corpus, a public dataset of tweets with estimated AAVE usage. The …


Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh Mar 2026

Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh

SMU Data Science Review

Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …


Customer Service Support. Utilizing Machine Learning To Classify, Prioritize And Summarize Issues., Hoai Nhan Nguyen Jan 2025

Customer Service Support. Utilizing Machine Learning To Classify, Prioritize And Summarize Issues., Hoai Nhan Nguyen

ICT

This capstone project investigates the application of machine learning and natural language processing (NLP) to enhance customer support operations through automated ticket classification, prioritization, and summarization. Using the multilingual Customer Support Emails dataset from Kaggle, the project follows the CRISP-DM methodology, performing extensive data cleaning, preprocessing, feature engineering, and class balancing. Five machine learning models—Decision Tree, KNN, LinearSVC, Naive Bayes, and Random Forest—were evaluated using hyperparameter tuning, cross-validation, confusion matrix analysis, and learning curves. LinearSVC demonstrated the strongest performance for both queue and priority classification, achieving accuracies of 89.8% and 81.2% respectively, with consistent generalization across folds. For summarization, extractive …


Intelligent Solutions For Retroactive Anomaly Detection And Resolution With Log File Systems, Derek G. Rogers, Chanvo Nguyen, Abhay Sharma May 2024

Intelligent Solutions For Retroactive Anomaly Detection And Resolution With Log File Systems, Derek G. Rogers, Chanvo Nguyen, Abhay Sharma

SMU Data Science Review

This paper explores the intricate challenges log files pose from data science and machine learning perspectives. Drawing inspiration from existing methods, LAnoBERT, PULL, LLMs, and the breadth of recent research, this paper aims to push the boundaries of machine learning for log file systems. Our study comprehensively examines the unique challenges presented in our problem setup, delineates the limitations of existing methods, and introduces innovative solutions. These contributions are organized to offer valuable insights, predictions, and actionable recommendations tailored for Microsoft's engineers working on log data analysis.


Game Recommendation Analysis Using Steam Profiles And Reviews, Robert Blue, Luis Garcia, Jacob Turner May 2024

Game Recommendation Analysis Using Steam Profiles And Reviews, Robert Blue, Luis Garcia, Jacob Turner

SMU Data Science Review

Smaller game studios are at a disadvantage when it comes to getting their product noticed by users. This study aims to provide insights on how recommendation engines work so that these smaller studios can have their games noticed on Steam. Steam is one of the largest video game distribution services and they have a recommendation engine which promotes games to its user base. This study utilized user information such as number of games played, the type of games, and the hours played and created recommendation engines to identify the qualities in the game that are driving recommendations.


Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn May 2024

Leveraging Transformer Models For Genre Classification, Andreea C. Craus, Ben Berger, Yves Hughes, Hayley Horn

SMU Data Science Review

As the digital music landscape continues to expand, the need for effective methods to understand and contextualize the diverse genres of lyrical content becomes increasingly critical. This research focuses on the application of transformer models in the domain of music analysis, specifically in the task of lyric genre classification. By leveraging the advanced capabilities of transformer architectures, this project aims to capture intricate linguistic nuances within song lyrics, thereby enhancing the accuracy and efficiency of genre classification. The relevance of this project lies in its potential to contribute to the development of automated systems for music recommendation and genre-based playlist …


A Nlp Approach To Automating The Generation Of Surveys For Market Research, Anav Chug May 2024

A Nlp Approach To Automating The Generation Of Surveys For Market Research, Anav Chug

Honors College Theses

Market Research is vital but includes activities that are often laborious and time consuming. Survey questionnaires are one possible output of the process and market researchers spend a lot of time manually developing questions for focus groups. The proposed research aims to develop a software prototype that utilizes Natural Language Processing (NLP) to automate the process of generating survey questions for market research. The software uses a pre-trained Open AI language model to generate multiple choice survey questions based on a given product prompt, send it to a targeted email list, and also provides a real-time analysis of the responses …


Low-Resource Icd Coding Of Discharge Summaries, Ashton Williamson May 2024

Low-Resource Icd Coding Of Discharge Summaries, Ashton Williamson

All Theses

Medical coding is the process by which standardized medical codes are assigned to patient health records. This is a complex and challenging task that typically requires an expert human coder to review health records and assign codes from a classification system based on a standard set of rules. Considering the downstream use of these codes in statistical analysis, billing, and patient care, improving the accuracy and efficiency of the medical coding process through automation could have a far-reaching impact on the healthcare domain. Since health records typically consist of a large proportion of free-text documents, this problem has traditionally been …


A Holistic And Collaborative Behavioral Health Detection Framework Using Sensitive Police Narratives, Martin Keagan Wynne Brown Apr 2024

A Holistic And Collaborative Behavioral Health Detection Framework Using Sensitive Police Narratives, Martin Keagan Wynne Brown

Dissertations

Identifying behavioral health is paramount for law enforcement officers to provide appropriate follow-up community care. In the current practice, law enforcement offices manually identify these behavioral health cases to allow the designation of the relevant follow-up resources. Police reports generated by officers' response to 911 calls remain an untapped resource for identifying such incidents. Therefore, we advocate for the incorporation of manual annotations from experts, natural language processing (NLP), active learning, advanced machine learning, and ensemble techniques to detect behavioral health cases within police reports. In this dissertation, we develop tools and frameworks to automatically detect behavioral health cases from …


Cannabidiol Tweet Miner: A Framework For Identifying Misinformation In Cbd Tweets., Jason Turner Aug 2023

Cannabidiol Tweet Miner: A Framework For Identifying Misinformation In Cbd Tweets., Jason Turner

Electronic Theses and Dissertations

As regulations surrounding cannabis continue to develop, the demand for cannabis-based products is on the rise. Despite not producing the psychoactive effects commonly associated with THC, products containing cannabidiol (CBD) have gained immense popularity in recent years as a potential treatment option for a range of conditions, particularly those associated with pain or sleep disorders. However, due to current federal policies, these products have yet to undergo comprehensive safety and efficacy testing. Fortunately, utilizing advanced natural language processing (NLP) techniques, data harvested from social networks have been employed to investigate various social trends within healthcare, such as disease tracking and …


Identifying Features And Predicting Consumer Helpfulness Of Product Reviews, Triston Hudgins, Shijo Joseph, Douglas Yip, Gaston Besanson Apr 2023

Identifying Features And Predicting Consumer Helpfulness Of Product Reviews, Triston Hudgins, Shijo Joseph, Douglas Yip, Gaston Besanson

SMU Data Science Review

Major corporations utilize data from online platforms to make user product or service recommendations. Companies like Netflix, Amazon, Yelp, and Spotify rely on purchasing trends, user reviews, and helpfulness votes to make content recommendations. This strategy can increase user engagement on a company's platform. However, misleading and/or spam reviews significantly hinder the success of these recommendation strategies. The rise of social media has made it increasingly difficult to distinguish between authentic content and advertising, leading to a burst of deceptive reviews across the marketplace. The helpfulness of the review is subjective to a voting system. As such, this study aims …


Question Answering With Distilled Bert Models: A Case Study For Biomedical Data, Brittany Lewandowski, Rayon Morris, Pearly Merin Paul, Robert Slater Apr 2023

Question Answering With Distilled Bert Models: A Case Study For Biomedical Data, Brittany Lewandowski, Rayon Morris, Pearly Merin Paul, Robert Slater

SMU Data Science Review

In the healthcare industry today, 80% of data is unstructured (Razzak et al., 2019). The challenge this imposes on healthcare providers is that they rely on unstructured data to inform their decision-making. Although Electronic Health Records (EHRs) exist to integrate patient data, healthcare providers are still challenged with searching for information and answers contained within unstructured data. Prior NLP and Deep Learning research has shown that these methods can improve information extraction on unstructured medical documents. This research expands upon those studies by developing a Question Answering system using distilled BERT models. Healthcare providers can use this system on their …


Using Nlp To Model U.S. Supreme Court Cases, Katherine Lockard, Robert Slater, Brandon Sucrese Apr 2023

Using Nlp To Model U.S. Supreme Court Cases, Katherine Lockard, Robert Slater, Brandon Sucrese

SMU Data Science Review

The advantages of employing text analysis to uncover policy positions, generate legal predictions, and inform or evaluate reform practices are multifold. Given the far-reaching effects of legislation at all levels of society these insights and their continued improvement are impactful. This research explores the use of natural language processing (NLP) and machine learning to predictively model U.S. Supreme Court case outcomes based on textual case facts. The final model achieved an F1-score of .324 and an AUC of .68. This suggests that the model can distinguish between the two target classes; however, further research is needed before machine learning models …


Beyond News Values On Twitter: Predicting Factors That Drive User Engagement In News, Zhiyan Zhong Apr 2023

Beyond News Values On Twitter: Predicting Factors That Drive User Engagement In News, Zhiyan Zhong

Dartmouth College Master’s Theses

When deciding on what news stories to cover, traditional journalism determines news values by following several elements of newsworthiness, such as impact, timeliness, and prominence. However, these guidelines do not always seem to correspond with the success of content on social media. As people are increasingly turning to social media for news, our research aims to understand and predict factors that drive user engagement for news on social media. In this study, we analyze news content published on Twitter, and examine a diverse set of characteristics like metrics retrieved from the Twitter API and semantics by natural language processing, including …


Application Of Big Data Technology, Text Classification, And Azure Machine Learning For Financial Risk Management Using Data Science Methodology, Oluwaseyi A. Ijogun Jan 2023

Application Of Big Data Technology, Text Classification, And Azure Machine Learning For Financial Risk Management Using Data Science Methodology, Oluwaseyi A. Ijogun

College of Graduate Studies: Theses & Dissertations

Data science plays a crucial role in enabling organizations to optimize data-driven opportunities within financial risk management. It involves identifying, assessing, and mitigating risks, ultimately safeguarding investments, reducing uncertainty, ensuring regulatory compliance, enhancing decision-making, and fostering long-term sustainability. This thesis explores three facets of Data Science projects: enhancing customer understanding, fraud prevention, and predictive analysis, with the goal of improving existing tools and enabling more informed decision-making. The first project examined leveraged big data technologies, such as Hadoop and Spark, to enhance financial risk management by accurately predicting loan defaulters and their repayment likelihood. In the second project, we investigated …


Natural Language Processing In Legal Tech, Jens Frankenreiter, Julian Nyarko Jan 2023

Natural Language Processing In Legal Tech, Jens Frankenreiter, Julian Nyarko

Scholarship@WashULaw

Natural language processing techniques promise to automate an activity that lies at the core of many tasks performed by lawyers, namely the extraction and processing of information from unstructured text. The relevant methods are thought to be a key ingredient for both current and future legal tech applications. This chapter provides a non-technical overview of the current state of NLP techniques, focusing on their promise and potential pitfalls in the context of legal tech applications. It argues that, while NLP-powered legal tech can be expected to outperform humans in specific categories of tasks that play to the strengths of current …


Phishing Detection Using Natural Language Processing And Machine Learning, Apurv Mittal, Dr Daniel Engels, Harsha Kommanapalli, Ravi Sivaraman, Taifur Chowdhury Sep 2022

Phishing Detection Using Natural Language Processing And Machine Learning, Apurv Mittal, Dr Daniel Engels, Harsha Kommanapalli, Ravi Sivaraman, Taifur Chowdhury

SMU Data Science Review

Phishing emails are a primary mode of entry for attackers into an organization. A successful phishing attempt leads to unauthorized access to sensitive information and systems. However, automatically identifying phishing emails is often difficult since many phishing emails have composite features such as body text and metadata that are nearly indistinguishable from valid emails. This paper presents a novel machine learning-based framework, the DARTH framework, that characterizes and combines multiple models, with one model for each composite feature, that enables the accurate identification of phishing emails. The framework analyses each composite feature independently utilizing a multi-faceted approach using Natural Language …


Using Natural Language Processing To Increase Modularity And Interpretability Of Automated Essay Evaluation And Student Feedback, Chris Roche, Nathan Deinlein, Darryl Dawkins, Faizan Javed Sep 2022

Using Natural Language Processing To Increase Modularity And Interpretability Of Automated Essay Evaluation And Student Feedback, Chris Roche, Nathan Deinlein, Darryl Dawkins, Faizan Javed

SMU Data Science Review

For English teachers and students who are dissatisfied with the one-size-fits-all approach of current Automated Essay Scoring (AES) systems, this research uses Natural Language Processing (NLP) techniques that provide a focus on configurability and interpretability. Unlike traditional AES models which are designed to provide an overall score based on pre-trained criteria, this tool allows teachers to tailor feedback based upon specific focus areas. The tool implements a user-interface that serves as a customizable rubric. Students’ essays are inputted into the tool either by the student or by the teacher via the application’s user-interface. Based on the rubric settings, the tool …


Stock Forecasts With Lstm And Web Sentiment, Michael Burgess, Faizan Javed, Nnenna Okpara, Chance Robinson Sep 2022

Stock Forecasts With Lstm And Web Sentiment, Michael Burgess, Faizan Javed, Nnenna Okpara, Chance Robinson

SMU Data Science Review

Traditional time-series techniques, such as auto-regressive and moving average models, can have difficulties when applied to stock data due to the randomness inherent to the markets. In this study, Long Short-Term Memory Recurrent Neural Networks, or LSTMs, have been applied to pricing data along with sentiment scores derived from web sources such as Twitter and other financial media outlets. The project team utilized this approach to complement the technical indicators observed at the end of each trading day for three stocks from the NASDAQ stock exchange over a 12-year span. A common benchmark to assess model performance on time series …


Web Page Multiclass Classification, Brian Gaither, Antonio Debouse, Catherine Huang Jun 2022

Web Page Multiclass Classification, Brian Gaither, Antonio Debouse, Catherine Huang

SMU Data Science Review

As the internet age evolves, the volume of content hosted on the Web is rapidly expanding. With this ever-expanding content, the capability to accurately categorize web pages is a current challenge to serve many use cases. This paper proposes a variation in the approach to text preprocessing pipeline whereby noun phrase extraction is performed first followed by lemmatization, contraction expansion, removing special characters, removing extra white space, lower casing, and removal of stop words. The first step of noun phrase extraction is aimed at reducing the set of terms to those that best describe what the web pages are about …


Leveraging Context Patterns For Medical Entity Classification, Garrett Johnston Jun 2022

Leveraging Context Patterns For Medical Entity Classification, Garrett Johnston

Computer Science Senior Theses

The ability of patients to understand health-related text is important for optimal health outcomes. A system that can automatically annotate medical entities could help patients better understand health-related text. Such a system would also accelerate manual data annotation for this low-resource domain as well as assist in down- stream medical NLP tasks such as finding textual similarity, identifying conflicting medical advice, and aspect-based sentiment analysis. In this work, we investigate a state-of-the-art entity set expansion model, BootstrapNet, for the task of medical entity classification on a new dataset of medical advice text. We also propose EP SBERT, a simple model …


Reading Level Identification Using Natural Language Processing Techniques, William Arnost, Ellen Lull, Joseph Schueder, Joseph Engler Dec 2021

Reading Level Identification Using Natural Language Processing Techniques, William Arnost, Ellen Lull, Joseph Schueder, Joseph Engler

SMU Data Science Review

This paper investigates using the Bidirectional Encoder Representations from Transformers (BERT) algorithm and lexical-syntactic features to measure readability. Readability is important in many disciplines, for functions such as selecting passages for school children, assessing the complexity of publications, and writing documentation. Text at an appropriate reading level will help make communication clear and effective. Readability is primarily measured using well-established statistical methods. Recent advances in Natural Language Processing (NLP) have had mixed success incorporating higher-level text features in a way that consistently beats established metrics. This paper contributes a readability method using a modern transformer technique and compares the results …


Qualitative Leveraging Natural Language Processing To Establish Judge Incrimination Statistics To Educate Voters In Re-Elections, Aurian Ghaemmaghami, Paul Huggins, Grace Lang, Julia Layne, Robert Slater Dec 2021

Qualitative Leveraging Natural Language Processing To Establish Judge Incrimination Statistics To Educate Voters In Re-Elections, Aurian Ghaemmaghami, Paul Huggins, Grace Lang, Julia Layne, Robert Slater

SMU Data Science Review

The prevalence of data has given consumers the power to make informed choices based off reviews, ratings, and descriptive statistics. However, when a local judge is coming up for re-election there is not any available data that aids voters in making data-driven decision on their vote. Currently court docket data is stored in text or PDFs with very little uniformity. Scaling the collection of this information could prove to be complicated and tiresome. There is a demand for an automated, intelligent system that can extract and organize useful information from the datasets. This paper covers the process of web scraping …


Detecting Stance On Covid-19 Vaccine In A Polarized Media, Rodica Ceslov Sep 2021

Detecting Stance On Covid-19 Vaccine In A Polarized Media, Rodica Ceslov

Dissertations, Theses, and Capstone Projects

The growing polarization in the United States has been widely reported. Media coverage plays an important role in shaping public opinion and influences public debates on complex and unfamiliar topics. There are some benefits to individuals and society from political polarization and conflict between opposing viewpoints. However, recent research has primarily highlighted the negative consequences of polarization which reached an all-time high. One such topic is the Covid-19 vaccine which was developed in record time, and the public learned about its safety and possible risks through the media coverage.

In this capstone, we examine U.S. news media coverage on the …


Exploring Media Portrayals Of People With Mental Disorders Using Nlp, Swapna Gottipati, Mark Chong, Andrew Wei Kiat Lim, Benny Haryanto Kawidiredjo Feb 2021

Exploring Media Portrayals Of People With Mental Disorders Using Nlp, Swapna Gottipati, Mark Chong, Andrew Wei Kiat Lim, Benny Haryanto Kawidiredjo

Research Collection School Of Computing and Information Systems

Media plays an important role in creating an impact in society. Several studies show that news media and entertainment channels, at times may create overwhelming images of the mental illness that emphasize criminality and dangerousness. The consequences of such negative impact may impact the audience with stigma and on the other hand, they impair the self-esteem and help-seeking behavior of the people with mental disorders. This is the first study to examine the Singapore media’s portrayal of persons with mental disorders (MDs) using text analytics and natural language processing. To date, most studies on media portrayal of people with MDs …


Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater Jan 2021

Multi-Modal Classification Using Images And Text, Stuart J. Miller, Justin Howard, Paul Adams, Mel Schwan, Robert Slater

SMU Data Science Review

This paper proposes a method for the integration of natural language understanding in image classification to improve classification accuracy by making use of associated metadata. Traditionally, only image features have been used in the classification process; however, metadata accompanies images from many sources. This study implemented a multi-modal image classification model that combines convolutional methods with natural language understanding of descriptions, titles, and tags to improve image classification. The novelty of this approach was to learn from additional external features associated with the images using natural language understanding with transfer learning. It was found that the combination of ResNet-50 image …


Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi Jan 2021

Improving Space Efficiency Of Deep Neural Networks, Aliakbar Panahi

Theses and Dissertations

Language models employ a very large number of trainable parameters. Despite being highly overparameterized, these networks often achieve good out-of-sample test performance on the original task and easily fine-tune to related tasks. Recent observations involving, for example, intrinsic dimension of the objective landscape and the lottery ticket hypothesis, indicate that often training actively involves only a small fraction of the parameter space. Thus, a question remains how large a parameter space needs to be in the first place — the evidence from recent work on model compression, parameter sharing, factorized representations, and knowledge distillation increasingly shows that models can be …


Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam Jan 2021

Analyzing Tweets On New Norm: Work From Home During Covid-19 Outbreak, Swapna Gottipati, Kyong Jin Shim, Hui Hian Teo, Karthik Nityanand, Shreyansh Shivam

Research Collection School Of Computing and Information Systems

The COVID-19 pandemic triggered a large-scale work-from-home trend globally in recent months. In this paper, we study the phenomenon of “work-from-home” (WFH) by performing social listening. We propose an analytics pipeline designed to crawl social media data and perform text mining analyzes on textual data from tweets scrapped based on hashtags related to WFH in COVID-19 situation. We apply text mining and NLP techniques to analyze the tweets for extracting the WFH themes and sentiments (positive and negative). Our Twitter theme analysis adds further value by summarizing the common key topics, allowing employers to gain more insights on areas of …


Health-Aware Food Planner: A Personalized Recipe Generation Approach Based On Gpt-2, Bushra Aljbawi Jan 2020

Health-Aware Food Planner: A Personalized Recipe Generation Approach Based On Gpt-2, Bushra Aljbawi

Theses and Dissertations (Comprehensive)

"What to eat today?" With the flourish of Internet, more and more people nowadays are inclined to find an answer to this most problematic question online. The recent explosion of food networks; however, produces large volumes of recipes, making it even harder to make an informed decision. This yields the need for advanced decision-making algorithms and efficient recommendation systems. Conventional recommender systems are not feasible anymore as food is a complicated feature that presents unique challenges and is less studied. For example, it can be one of the main reasons for obesity and many other chronic diseases. Food recommender system …