Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

3,231 Full-Text Articles 9,305 Authors 1,316,836 Downloads 221 Institutions

All Articles in Data Science

Faceted Search

3,231 full-text articles. Page 8 of 155.

Nlp Bias And African American English, Kenya Roy, Faizan Javed 2026 Southern Methodist University

Nlp Bias And African American English, Kenya Roy, Faizan Javed

SMU Data Science Review

African American English (AAE), also referred to as African American Vernacular English (AAVE), is widely used on social media, but most sentiment analysis tools are trained only on Standard American English (SAE). This mismatch can cause models to misclassify dialectal expressions—especially by labeling neutral or positive AAE as negative or toxic. These errors matter, since Natural Language Processing (NLP) systems are now central to content moderation and brand monitoring. This research will evaluate the VADER, RoBERTa, GPT-OSS, and Gemma’s handling of AAE in comparison to SAE using the TwitterAAE corpus, a public dataset of tweets with estimated AAVE usage. The …


Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh 2026 Southern Methodist University

Automating Cardiff Model Data Capture In Emergency Departments: Ambient Nlp Integration With Oracle-Cerner Fhir Systems, Simi Augustine, Marco A. Lopez, Jacquelyn Cheun, Chris Papesh

SMU Data Science Review

Violence and overdose events in Las Vegas occur at rates above the national average, with fewer than half of violent injuries reported to law enforcement [2,7]. The Cardiff Model offers a proven framework for standardized data collection and sharing between hospitals and public safety partners, yet many implementations still rely on manual entry. We propose an ambient triage pipeline integrated with Oracle-Cerner electronic health record systems to listen to nurse–patient dialogue, convert speech to text, extract Cardiff fields, and write standards-based FHIR Bundles for analytics. Using SMART on FHIR standards and Cerner Millennium APIs, the study evaluates whether ambient capture …


Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal 2026 Southern Methodist University

Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal

SMU Data Science Review

Large-scale software systems produce vast volumes of logs and telemetry, making manual incident triage slow and error prone. This study presents an unsupervised anomaly detection pipeline that fuses logs, metrics, and traces through late fusion. Using Hybrid Ensemble modeling with Isolation Forest, and Long Short-Term Memory (LSTM) Deep Learning model, the system detects cross-service anomalies producing and assigning a composite triage score reflecting severity and impact. Ranked alerts are categorized into Critical, High, or Medium priorities for review. A retrieval-augmented generation (RAG) layer enriches results with contextual summaries for explainable triage. Evaluated on synthetic multi-service datasets, the pipeline …


A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia 2026 Southern Methodist University

A Modular Framework For Cost-Efficient Aspect-Based Sentiment Analysis Using Small Language Models, Senthil Kumar, Nibhrat Lohia

SMU Data Science Review

Aspect-based sentiment analysis (ABSA) links opinions in text to specific product attributes (for example, battery life, screen quality, or delivery speed) rather than only assigning an overall star rating. This level of detail is important in domains such as e-commerce, where teams need to know which features customers praised and which they criticized. Traditional ABSA pipelines have relied on large language models (LLMs), which achieved high quality but were expensive to run and difficult to scale. This study evaluated whether small language models (SLMs) in the 1–3 billion parameter range could serve as a lower-cost alternative. We implemented a modular …


Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy McSimov, Trevor S. Kunz, Jeffrey Billo 2026 Southern Methodist University

Availability Model To Evaluate Ai Data Centers’ Role In Grid Stability, Troy Mcsimov, Trevor S. Kunz, Jeffrey Billo

SMU Data Science Review

The United States has made it clear; it is imperative that the US wins the global AI race. This paper focuses on one of the most challenging puzzle pieces surfaced at the POWER Data Center conference (San Antonio, Sept. 30.); for Electric Reliability Council of Texas (ERCOT) the limiting factor is not generation alone but the need to balance generation and load to preserve grid reliability.

The regulatory landscape fundamentally changed with the passage of Texas Senate Bill 6 in June 2025, which mandates new large loads must "contribute to the recovery of the interconnecting electric utility’s costs" (Texas Legislature, …


Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn 2026 SMU

Application Of Open-Source Small Large Language Models For Finance Report Analysis, Tue Vu, Mark Austin, Marcel Tuijn

SMU Data Science Review

The rapid integration of generative AI in finance introduces both opportunities and challenges, particularly when analyzing sensitive data such as Securities and Exchange Commission (SEC) filings. This study investigates the use of open-source Small Large Language Models (SLLMs), deployed locally through the Ollama and LangChain frameworks, combined with Retrieval-Augmented Generation (RAG) for extracting financial insights relevant to index performance and reporting quality. Two key objectives guide this work: (1) benchmarking multiple open-source SLLMs for sentiment analysis, multiple-choice reasoning, and financial question answering, and (2) assessing the feasibility of locally deployed SLLMs for domain-specific financial queries. A standardized set of 50 …


Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler 2026 Southern Methodist University

Reducing Range Anxiety Through Predictive Modeling Of Ev Battery Degradation, Caleb Thornsbury, Christian Castro, Bivin Sadler

SMU Data Science Review

Electric Vehicles (EV) range anxiety remains one of the top barriers for broader adoption. Range anxiety can be attributed to battery pack age and degradation over time. This paper plans to explore how to address this issue by creating a machine learning model that can predict degradation based on usage, temperature, battery chemistry, charging habits and exploring whether other factors tie into range degradation. This research will be using real world charging data along with lab tested chemistry data to build a model that can be chemistry specific for degradation. This paper will help perspective used-EV buyers learn about battery …


“It’S A Lot More To It Than Just Research”: Integrating Critical Data Literacy And Reasoning With Data Into A Stem Summer Camp, Marc T. Sager, Saki L. Milton, Candace Walkington 2026 Southern Methodist University

“It’S A Lot More To It Than Just Research”: Integrating Critical Data Literacy And Reasoning With Data Into A Stem Summer Camp, Marc T. Sager, Saki L. Milton, Candace Walkington

Publications

Purpose: This study explores how middle-grade girls from predominantly underrepresented and underserved racially and ethnically minoritized (UUREM) backgrounds developed critical data literacy (CDL) through participation in a week-long residential STEM camp. Given the increasing importance of data science education in a data-driven world, this research examines how informal learning environments can support CDL development among youth from historically marginalized groups.

Design/Methodology/Approach: The study draws on qualitative interview data from eleven participants, and their group-produced artifacts to investigate how the camp experience supported engagement with data and the development of CDL. Interviews explored participants' experiences with data collection, organization, analysis, and …


Using Ensemble Disagreement To Stabilize Conformal Prediction Under Distribution Shift, Patrick D. Murphy 2026 California Polytechnic State University, San Luis Obispo

Using Ensemble Disagreement To Stabilize Conformal Prediction Under Distribution Shift, Patrick D. Murphy

Master's Theses

Semantic segmentation of eelgrass from drone imagery is crucial for coastal habitat monitoring, restoration, and management, as these habitats continue to see rapid changes due to climate change and human influence. However, the reliability of generalizing a deployed classification model relies on both high-accuracy segmentation as well as robust uncertainty quantification that holds up when conditions change over years or locations. Conformal prediction (CP) is a method that converts a classifier's output into prediction sets with a guaranteed average coverage level for in-distribution data. However, the “vanilla” conformal score can often under-cover in hard or out-of-distribution (OOD) regions under drift. …


A Unified Methodological Framework For Generating Digital Twins Of Multi Class Uncrewed Systems (Uxs), Sai Raghava Pathuri 2026 University of South Alabama

A Unified Methodological Framework For Generating Digital Twins Of Multi Class Uncrewed Systems (Uxs), Sai Raghava Pathuri

Shelby Hall Graduate Research Forum Presentations

No abstract provided.


Evaluating The Effects Of Anti-Forensic Activities In Additive Manufacturing Devices, Daniel B. Miller 2026 University of South Alabama

Evaluating The Effects Of Anti-Forensic Activities In Additive Manufacturing Devices, Daniel B. Miller

Shelby Hall Graduate Research Forum Posters

Additive Manufacturing (AM) is a newer famlily of production tecchologies that constructs objects by fusing layers of material into the desired shape. Methods for achieving this as described in Gibson et al. [1] are varied and include Fused Filament Deposition, Selective Laser Sintering, Stereolithography (SLA), and Powder Bed Fusion. Computers are integral to the processes being responsible for creating and decoding design instructions, collecting and processing sensor data, and, ultimately, directing the activity of the machines which implement the process. Additionally, the AM industry is rapidly expanding, worth an estimated $23 billion in 2023 and projected to reach $88 billion …


Cellebrite Reliability In Digital Forensics, Christina Huynh 2026 University of South Alabama

Cellebrite Reliability In Digital Forensics, Christina Huynh

Shelby Hall Graduate Research Forum Posters

Forensic tools like Cellebrite are commonly used in court to gather and interpret raw data for evidence. Cellebrite does not only collect data but creates and interprets the artifacts of data to create a scene of the process it has been through. With this, evidence can be influenced by software designs and not just the data on the mobile device. Courts and police use Cellebrite to gather evidence and reconstruct it to create an easily readable dataset. These tools lack reproducibility, transparency, integrity, and chain of evidence command. Cellebrite is often used in court and by police without further vetting …


Evaluating Regularized Logistic Regression And K-Nn On Mnist Under Increasing Random Missingness, Daniel Markwei 2026 University of Central Florida

Evaluating Regularized Logistic Regression And K-Nn On Mnist Under Increasing Random Missingness, Daniel Markwei

Data Science and Data Mining

This paper investigates the effect of random missingness on the performance of regularized multinomial logistic regression and the k-nearest neighbors (k-NN) classifier for handwritten digit recognition on the MNIST dataset. In particular, we study L1-regularized (LASSO) logistic regression and L2-regularized (Ridge) logistic regression alongside k-NN. Varying percentages of random missingness were introduced into the original dataset, and each model was evaluated in terms of its classification performance. The results show that random missingness degrades the performance of all three classifiers. Overall, k-NN consistently achieves higher accuracy than both L1- and L2-regularized logistic regression across all missingness levels; however, its performance …


Optimization Of Image Quality Of Simulated Multiple Detectors Computed Tomography Acquisition Parameters Using Machine Learning And Catphan Phantom, Ali O. Masoud, Najat K. Mohammed, Khamis O. Amour, Ahmed M. Jusabani, Denise Mwalongo, Mwingereza John Kumwenda 2026 Directorate of Radiation Control Unit, Tanzania Atomic Energy Commission (TAEC) P, O Box 1585, Dodoma, Tanzania

Optimization Of Image Quality Of Simulated Multiple Detectors Computed Tomography Acquisition Parameters Using Machine Learning And Catphan Phantom, Ali O. Masoud, Najat K. Mohammed, Khamis O. Amour, Ahmed M. Jusabani, Denise Mwalongo, Mwingereza John Kumwenda

Tanzania Journal of Science

The study successfully employed Monte Carlo (MC) simulation and a Machine Learning (ML) approach using a Random Forest Regression (RFR) model to develop optimized Multi-Detector CT (MDCT) protocols that significantly reduce radiation dose while maintaining diagnostic image quality. The MC engine accurately modeled X-ray spectra, and the RFR model demonstrated high predictive power for key metrics, achieving R2 scores of 0.97 for CTDIvol and over 0.92 for image quality metrics (Noise, CNR). Through multi-objective optimization guided by the RFR, the final protocol (Optimization-3) was found on the Pareto front, achieving a notable 35% dose reduction (from 15.5 mGy to 9.9 …


Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell 2026 The University of Alabama, Tuscaloosa

Cerebral Documents And Algorithmic Sensemaking: Searching For Expressions In Human And Artificial Cognitive Collaborations, Rebekah L. Cowell

Proceedings from the Document Academy

Generative Artificial Intelligences (AIs) and current advanced large language models (LLMs) are algorithmically designed to generate text-based conversations as conversational agents (CAs), by replicating human language and conversational communication. Pairing human cognition with generative computationally coded cognition. We have never been here before: cerebral and artificial information collaborations and processing producing expressions that may or may not become visible as second-hand/secondary source documents.

Sensemaking or sense(un)making is a unique autonomous human drive cognitively, our information processing is sensemaking in action and expressions and articulations are evidence of the sensemaking cycle. Documentation [expressed or articulated through various mediums] are a product …


Effective Wordle Heuristics, Ronald I. Greenberg 2026 Loyola University Chicago

Effective Wordle Heuristics, Ronald I. Greenberg

Computer Science: Faculty Publications and Other Works

While previous researchers have performed an exhaustive search to determine an optimal Wordle strategy, that computation is very time consuming and produced a strategy using words that are unfamiliar to most people. With Wordle solutions being gradually eliminated (with a new puzzle each day and no reuse), an improved strategy could be generated each day, but the computation time makes a daily exhaustive search impractical. This paper shows that simple heuristics allow for fast generation of effective strategies and that little is lost by guessing only words that are possible solution words rather than more obscure words.


Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra 2026 Smith College

Tidychangepoint: A Unified Framework For Analyzing Changepoint Detection In Univariate Time Series, Ben Baumer, Biviana Marcela Suárez Sierra

Statistical and Data Sciences: Faculty Publications

We present tidychangepoint, a new R package for changepoint detection analysis. Most R packages for segmenting univariate time series focus on providing one or two algorithms for changepoint detection that work with a small set of models and penalized objective functions, and all of them return a custom, nonstandard object type. This makes comparing results across various algorithms, models, and penalized objective functions unnecessarily difficult. tidychangepoint solves this problem by wrapping functions from a variety of existing packages and storing the results in a common S3 class called tidycpt. The package then provides functionality for easily extracting comparable numeric or …


Bibliography For Love Data Week 2026, Annikah Carpio, Sally Park 2026 Chapman University

Bibliography For Love Data Week 2026, Annikah Carpio, Sally Park

Library Displays and Bibliographies

A bibliography created to support a display about research data and Love Data Week during February 2026 at the Leatherby Libraries at Chapman University.


Typeface: Machine-Viewing Gentrification On Storefront Imagery In Bedford-Stuyvesant, Brooklyn, Alexander McQuilkin 2026 CUNY Graduate Center

Typeface: Machine-Viewing Gentrification On Storefront Imagery In Bedford-Stuyvesant, Brooklyn, Alexander Mcquilkin

Dissertations, Theses, and Capstone Projects

Gentrification—broadly, the replacement of a less powerful group by a more powerful one in an urban context—is oft-discussed in the popular press, but its definition is much-debated in the urban planning literature. Furthermore, academic treatments of displacement understandably focus on measurable yet fairly abstract indicators like changes in rent or income, whereas neighborhood change is often registered by residents on the ground using visual, but difficult-to-quantify markers like retail turnover. This project uses image recognition technology on a set of storefront photos to index the visual streetscape of a neighborhood, as well as to track changes to that portrait over …


Courts Of New York: A Visual Atlas Of The City’S Public Basketball Spaces, Nathaniel Rattner 2026 CUNY Graduate Center

Courts Of New York: A Visual Atlas Of The City’S Public Basketball Spaces, Nathaniel Rattner

Dissertations, Theses, and Capstone Projects

Basketball courts in New York City are recreation facilities, community anchors and part of the city’s cultural image. In the basketball capital of the world, New Yorkers are rarely more than a few blocks away from a court. The visual diversity of these courts, however, is not widely documented in systematic ways.

This project makes that diversity visible to the public, combining open data, aerial imagery and computational analysis to document this important public space across the five boroughs. It is a narrative story and digital atlas of New York City’s public basketball courts, using surface color as a way …


Digital Commons powered by bepress