Open Access. Powered by Scholars. Published by Universities.®

Data Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 181 - 210 of 1157

Full-Text Articles in Data Science

Reinforcement Learning, Modeling Markets, And Professional Basketball Free Agency, Jacob Cohn May 2025

Reinforcement Learning, Modeling Markets, And Professional Basketball Free Agency, Jacob Cohn

Computational and Data Sciences (PhD) Dissertations

This dissertation presents a reinforcement learning-based approach to modeling and optimizing decision-making in professional basketball free agency and related economic environments. A Markov Decision Process (MDP) framework is introduced to capture the strategic interactions of NBA teams bidding for free agents under budgetary and roster constraints. To address computational scalability challenges, a reinforcement learning (RL) environment is developed, leveraging Proximal Policy Optimization (PPO) to approximate optimal policies for team decision-making.

Empirical results demonstrate that the RL agent successfully learns strategic bidding behavior that aligns with dynamic programming benchmarks in simplified settings while scaling effectively to larger, intractable environments. The study …


Multimodal Benchmarking For Ncaa Basketball, Brendan Barnett May 2025

Multimodal Benchmarking For Ncaa Basketball, Brendan Barnett

Honors Scholar Theses

We present the first multimodal, multitask benchmark for NCAA basketball, synthesizing structured statistical features with large language model (LLM)-generated game summaries across 19,739 games spanning four NCAA Division I seasons (2021--2025). We evaluate three model families---XGBoost, deep neural networks, and Transformers---under tabular-only and early-fusion settings to measure the impact of LLM-derived textual embeddings. To assess practical utility, we simulate fixed-stake and Kelly criterion-based betting strategies using historical bookmaker odds, analyzing both profitability and downside risk via Monte Carlo simulation. Our results show that XGBoost with early-fusion achieves the highest return on investment and the lowest risk of loss. This work …


Maximal Independent Set Algorithms Within Procedural Planar Maps: A Large-Scale Evaluation, Chaucer Ihrig May 2025

Maximal Independent Set Algorithms Within Procedural Planar Maps: A Large-Scale Evaluation, Chaucer Ihrig

Honors College Theses

Analysis of a childhood game has led us to the problem of maximum independent sets in planar graphs. We wrote a graph creation utility using R to generate a random planar map and its dual graph. This utility then finds a graph’s maximal independent set using a variety of six algorithms. We investigate statistical connections between graph structure, colorability, and the maximal independent sets found using these algorithms over an incredibly large and procedurally generated dataset. We find one can always win the coloring game if the resultant graph is two-colorable. The algorithms perform statistically and practically significantly better on …


Toward Robust Semantic Segmentation In Levee Infrastructure Monitoring: Enhancing Accuracy With High-Fidelity Synthetic Data And Ensemble Learning, Padam Jung Thapa May 2025

Toward Robust Semantic Segmentation In Levee Infrastructure Monitoring: Enhancing Accuracy With High-Fidelity Synthetic Data And Ensemble Learning, Padam Jung Thapa

LSU New Orleans Theses and Dissertations

Abstract: Levees serve as critical flood protection structures, but failures due to inadequate maintenance and extreme water pressures have led to devastating events such as Hurricane Katrina. Manual inspections are slow, labor-intensive, and prone to human error, necessitating the development of automated solutions. This study proposes an AI-driven framework for levee inspection utilizing deep learning-based semantic segmentation to detect rutting and enhance the identification of sand boils. To address dataset limitations, high-fidelity synthetic images are generated using DreamBooth for fine-tuning, while ControlNet adds structural constraints to enhance realism and consistency. A semi-automatic convex hull annotation technique enhances labeling efficiency, and …


Radar Precursors To Severe Weather Reports In Left-Moving Supercells, Eric A. Carothers May 2025

Radar Precursors To Severe Weather Reports In Left-Moving Supercells, Eric A. Carothers

Department of Earth and Atmospheric Sciences: Dissertations, Theses, and Student Research

While much research has examined dual-polarimetric signatures of right-moving supercells, very little has been done with left-moving supercells. Given that left-moving supercells are thought to be disproportionate producers of large hail, understanding their internal dynamics is vitally important. This study examines differences and trends in the dual-polarimetric signatures of left-moving supercells to identify precursors to severe weather reports. A dataset of left-moving supercells associated with severe weather reports was created. These storms are processed with an automated analysis algorithm that identifies and quantifies the polarimetric signatures in each storm. A method for analysis of differences and trends in their dual-polarization …


Advancing Precision And Autonomy In Agriculture And Medical Imaging Through Ai And Computer Visions, Tan-Hanh Pham May 2025

Advancing Precision And Autonomy In Agriculture And Medical Imaging Through Ai And Computer Visions, Tan-Hanh Pham

Theses and Dissertations

Deep learning has revolutionized numerous fields by enhancing precision, automation, and decision-making capabilities. This dissertation explores its applications in agriculture and medical image processing, introducing novel methodologies to improve accuracy and efficiency in these domains. These fields hold critical societal importance -- agriculture underpins global food security and sustainability, while medical imaging drives advancements in diagnostics and personalized healthcare, both benefiting significantly from data-driven innovations. In agriculture, deep learning is applied to precision spray systems through droplet analysis. Specifically, a generative model is designed to create synthetic droplet images, addressing the challenge of limited training samples, which are expensive …


Ml Playground: Data Modification/Preprocessing And Model Simulation Tool, Marco D. Cerrato May 2025

Ml Playground: Data Modification/Preprocessing And Model Simulation Tool, Marco D. Cerrato

Electronic Theses, Projects, and Dissertations

There is a heavy reliance on programming when it comes to learning machine learning (ML). This often creates barriers for students and newcomers unfamiliar with coding. While the lessons you learn in the classroom provide essential foundational understanding, some technical or practical aspects of ML—such as data preprocessing, feature engineering, and model tuning—are best learned through hands-on interaction. ML Playground was developed to act as a proof-of-concept application to address this gap by offering a browser-based, graphical user interface that lets users engage with core ML workflows without writing code. Designed with educational accessibility in mind, the application allows users …


Analytics Insights From Text: Machine Learning, Ai, And Sentiment Analysis On Beige Books, Charlie Smith May 2025

Analytics Insights From Text: Machine Learning, Ai, And Sentiment Analysis On Beige Books, Charlie Smith

Graduate Theses and Dissertations (2019 - present)

Business analytics is about drawing actionable insights from data. These distinct but connected essays represent a novel approach to explore how natural language processing (NLP) advances and machine learning can transform unstructured text data into actionable conclusions. Essay 1 provides a broad framework. Essay 2 strengthens the sentiment analysis with the most recent artificial intelligence methodologies for capturing nuanced sentiment in complex texts. Essay 3 applies those insights to forecast recessions using topics that can be readily interpreted and applied.

The research demonstrates how these methodologies can be applied to enhance understanding of the same dataset, Beige Books. Published by …


Comparative Analysis Of Regression And Random Forest Models For Player Performance Prediction In The Mls, Joshua Clement Madeti May 2025

Comparative Analysis Of Regression And Random Forest Models For Player Performance Prediction In The Mls, Joshua Clement Madeti

Senior Honors Theses

Advanced technology and analytics have transformed the world and have benefited several industries throughout, the sport industry being one of them. Data is constantly generated during sports and requires post-game or post-season analysis which is crucial to team and player success. In this paper, the researcher will focus on the impact of analytics on soccer and soccer players. With over three billion active fans, soccer is the most famous sport in the world yet, when it comes to analytics, it is lagging. The thesis includes a comparative study of multiple linear regression and random forest regression to explore whether these …


Describing Functionality In Natural Language May Improve Decomposition Behaviors, Matthew R. Burns May 2025

Describing Functionality In Natural Language May Improve Decomposition Behaviors, Matthew R. Burns

All Graduate Theses and Dissertations, Fall 2023 to Present

Problem decomposition—the ability to break complex problems into simpler parts—is a critical skill for computer programming that many beginning students struggle to develop. This research examines how using natural language to describe program functionality can help students develop better problem-solving approaches.

We created a tool called ”Natural Language Functions” (NLFs) that allows students to write descriptions of what they want their code to do in plain English, which then generates working Python functions. We studied how students used this tool compared to students who solved programming problems in traditional ways.

Our findings show that students who used the NLFs tool …


The Role Of Ai In Enhancing Teamwork, Resilience And Decision-Making: Review Of Recent Developments, Satyadhar Joshi May 2025

The Role Of Ai In Enhancing Teamwork, Resilience And Decision-Making: Review Of Recent Developments, Satyadhar Joshi

Harrisburg University Other Works

This paper explores the transformative impact of artificial intelligence (AI) on organizational teamwork, decision-making, and resilience. This paper furthur reviews recent literature on the integration of Artificial Intelligence (AI) in various organizational functions, focusing on its impact on innovation management, leadership paradigms, and organizational resilience. We provide groundwork required to enhance frameworks that can integrate cognitive scaffolding with antifragile team dynamics, employing behavioral economics and neurocognitive principles. We introduce methodologies for enhancing team resilience through adaptive AI systems, cross-training interventions, and pre-mortem simulation techniques. The framework addresses key challenges in confirmation bias mitigation, cultural dimension alignment, and vigilance decrement prevention. …


Towards Advancing Streamflow And Peak Flow Prediction With Machine Learning: Identifying Infrastructure At Risk, Sudan Pokharel May 2025

Towards Advancing Streamflow And Peak Flow Prediction With Machine Learning: Identifying Infrastructure At Risk, Sudan Pokharel

Dissertations and Doctoral Documents, University of Nebraska-Lincoln, 2023–

Due to climate change and its impact, the need for adaptive strategies for natural disaster mitigation and resource management has never been more urgent. Central to this is water resource management, which is essential for sustainable human activities, ecological balance, and the mitigation of natural hazards like floods. Streamflow is a crucial element of water resource management and plays a vital role in planning and building water infrastructure, implementing emergency response plans, supporting flood mitigation initiatives, and regulating agricultural and industrial use. However, accurate prediction of streamflow still remains a challenge due to the complex non-linear and non-stationary interaction between …


On Bed Posture Recognition Using Deep Learning With Pressure Sensors, Farheen Akhter Ms May 2025

On Bed Posture Recognition Using Deep Learning With Pressure Sensors, Farheen Akhter Ms

Electronic Theses, Projects, and Dissertations

In healthcare applications such as disease prevention, sleep quality evaluation, and patient monitoring, bed posture recognition is essential. Using pressure sensor arrays placed on top of or embedded in mattresses, this study investigates the application of deep learning models for non-invasive posture classification. Although they have been widely employed, traditional machine learning approaches like support vector machines (SVM) and k-nearest neighbors (KNN) sometimes struggle with feature extraction and real-time performance necessitating considerable processing resources. I implemented a model using conventional approaches to get over these restrictions, then fine-tuned it using the following deep learning architectures for bed posture recognition: ResNet-50, …


Analyzing The Sentiment Of Feminist And Non-Feminist Works, Jasmine Borie, Megan G. Falschlehner Apr 2025

Analyzing The Sentiment Of Feminist And Non-Feminist Works, Jasmine Borie, Megan G. Falschlehner

Mathematics, Computer Science & Statistics Presentations

This presentation focuses on a group of texts that advocate for a change in the current belief system. These texts are the Feminist Manifesto, Sojourner Truth: Ain’t I a Woman?, and Civilization and Its Discontents. These first two texts advocate for women’s rights, while Freud’s book is focused on civilization’s decline and how our understanding of community can affect this. Through our presentation, we want to examine the differences in sentiment and language between the feminist texts and Freud’s texts to pinpoint whether or not sentiment changes when advocating for different beliefs.


Analyzing Cie Texts Through History Using R, Rachel A. Hart, Aaron Ditto Apr 2025

Analyzing Cie Texts Through History Using R, Rachel A. Hart, Aaron Ditto

Mathematics, Computer Science & Statistics Presentations

In this presentation, we analyzed three separate CIE texts from different time periods. First, “The Allegory of the Cave” from 380 BC, then “The Declaration of Independence” from 1776, and lastly “The Lottery” from 1948. We compared them using tidy text techniques like sentiment lexicons, creating word clouds, and bigram analysis to see if the types of words and sentiments used have changed over time in these short texts.


A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn Apr 2025

A Statistical Comparison Of Selected Old Testament And New Testament Books, Branden F. Stahl, Kevin Guan, Adam Denn

Mathematics, Computer Science & Statistics Presentations

The purpose of this project was to discover similarities between sentiments in Old Testament and New Testament books of the Bible, track emotional valence and find the most common words and sentiments in the books. Text analysis was performed on Genesis, Exodus, Matthew and Luke. Word clouds were also created for these texts.


Analyzing Musical Emotions: A Multi-Dataset Approach To Sentiment And Mood Classification In Songs, Mahad Syed Apr 2025

Analyzing Musical Emotions: A Multi-Dataset Approach To Sentiment And Mood Classification In Songs, Mahad Syed

SPARK Symposium Presentations

Music evokes a wide range of emotions, yet most music recommendation systems focus on sound and listening patterns rather than the meaning of lyrics. This project enhances lyric-based emotion recognition by applying Natural Language Processing (NLP) and Machine Learning (ML) to classify song lyrics into emotional categories.

I used eight datasets from Kaggle, including collections of lyrics, emotion labels, and audio features, providing a strong foundation for analysis. Our approach combines traditional NLP techniques (like TF-IDF and Word2Vec) with advanced deep learning models (such as BERT and XLNet) to classify lyrics into categories like happy, sad, angry, calm, romantic, and …


Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh Apr 2025

Leveraging Attention Mechanism To Unlock Gene And Protein Attributes, Ala Jararweh

Computer Science ETDs

Advancing personalized medicine depends on effectively integrating and interpreting the vast, heterogeneous landscape of biological data, from genomic sequences and transcriptomics to the insights embedded in scientific literature. Current machine learning models often focus on single data modalities, limiting their capacity to capture the multifaceted nature of biological systems. We address this gap by developing three attention-based machine-learning models integrating diverse data modalities. Firstly, DeepVul is a multi-task model that leverages cancer transcriptome data to predict genes critical for cancer survival and their corresponding drugs. Subsequently, LitGene refines gene representations by integrating textual information from the scientific literature. Finally, Protein2Text …


A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss Apr 2025

A Machine-Learning Tool-Supported Methodology For Nonprofit Donor Analysis, Corbin Weiss

Campus Research Month

We developed a machine-learning tool-supported methodology for modeling the nonprofit donor relationship. This approach was demonstrated in the case of a US-based nonprofit. Conclusions were drawn from this example and tool-support provided for use by other nonprofits.


From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie Apr 2025

From Adversarial Attacks To Robust Classifiers - A Study In Social Media Spam Detection - Black Box & White Box, Jonathan Jose Penaloza Rumie

Undergraduate Theses

Adversarial attacks pose a significant threat to the reliability of machine learning-based spam detection systems in social media. This undergraduate thesis, "From Adversarial Attacks to Robust Classifiers: A Study in Social Media Spam Detection – Black Box & White Box," systematically examines the impact of both black-box and white-box adversarial attacks on a range of spam classifiers, including Logistic Regression, Decision Trees, Random Forests, K-Nearest Neighbors, Bagging, Gradient Boosting, and Support Vector Machines. Leveraging a novel dataset derived from Twitter spam messages and enhanced with adversarial perturbations such as synonym replacement and character-level modifications, this study evaluates classifier performance under …


36 - Investigation Of The Digital Footprint Of Scientific Research In Social Media – Preliminary Findings, Lee Logan, Dominik Soos, Sean Baker, Jian Wu Apr 2025

36 - Investigation Of The Digital Footprint Of Scientific Research In Social Media – Preliminary Findings, Lee Logan, Dominik Soos, Sean Baker, Jian Wu

Undergraduate Research Symposium

Title: Investigation of The Digital Footprint of Scientific Research in Social Media – Preliminary Findings

Authors: Lee Logan, Sean Baker, Dominik Soos, Jian Wu

The spread of scientific information and research beyond the confines of academic institutions plays a central role in how the public understands and trusts modern sciences. Social media has become an essential means of dissemination for scholarly news, papers, and other forms of engagement. This research aims to explore how scientific research is disseminated over social media to understand its role as a bridge between peer-reviewed research and the public's overall understanding. To support the research …


Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova Apr 2025

Development And Application Of Self-Supervised Machine Learning For Smoke Plume And Active Fire Identification From The Fire Influence On Regional To Global Environments And Air Quality Datasets, Nicholas Lahaye, Anastasija Easley, Kyongsik Yun, Hugo Lee, Erik Linstead, Michael J. Garay, Olga V. Kalashnikova

Engineering Faculty Articles and Research

Fire Influence on Regional to Global Environments and Air Quality (FIREX-AQ) was a field campaign aimed at better understanding the impact of wildfires and agricultural fires on air quality and climate. The FIREX-AQ campaign took place in August 2019 and involved two aircraft and multiple coordinated satellite observations. This study applied and evaluated a self-supervised machine learning (ML) method for the active fire and smoke plume identification and tracking in the satellite and sub-orbital remote sensing datasets collected during the campaign. Our unique methodology combines remote sensing observations with different spatial and spectral resolutions. With as much as a 10% …


Cogprog: Utilizing Large Language Models To Forecast In-The-Moment Health Assessment, Gina Sprint, Maureen Schmitter-Edgecombe, Raven Weaver, Lisa Wiese, Diane Cook Apr 2025

Cogprog: Utilizing Large Language Models To Forecast In-The-Moment Health Assessment, Gina Sprint, Maureen Schmitter-Edgecombe, Raven Weaver, Lisa Wiese, Diane Cook

Computer Science Faculty Scholarship

Forecasting future health status is beneficial for understanding health patterns and providing anticipatory support for cognitive and physical health difficulties. In recent years, generative Large Language Models (LLMs) have shown promise as forecasters. Though not traditionally considered strong candidates for numeric tasks, LLMs demonstrate emerging abilities to address various forecasting problems. They also provide the ability to incorporate unstructured information and explain their reasoning process. In this article, we explore whether LLMs can effectively forecast future self-reported health state. To do this, we utilized in-the-moment assessments of mental sharpness, fatigue, and stress from multiple studies, utilizing daily responses (N = …


Extending Feature-Based Detection For Artificial Intelligence, Kayla Ahrndt Apr 2025

Extending Feature-Based Detection For Artificial Intelligence, Kayla Ahrndt

SPARK Symposium Presentations

AI text generation is rapidly developing, and, as a result, it is becoming increasingly difficult to differentiate it from human written text. Our base study by Leon Fröhling et al. proposed a feature-based detection model trained on GPT2, GPT3, and Grover data, as well as human-generated text. Our work extends their research by training a modified model with four neural networks on word embeddings, select features from the original study, as well as updated data (GPT3, GPT4, and Grover).


David B. Smith Chats With Monday 1.0, David B. Smith Apr 2025

David B. Smith Chats With Monday 1.0, David B. Smith

Publications and Research

This document is an edited archival transcript of extended conversations between David B. Smith and an AI persona (“Monday 1.0,” GPT‑4o based) conducted in Spring 2025, prepared as a foundational primary source for subsequent scholarly and creative work. It records the emergence and testing of concepts related to human–AI collaboration (including “Balanced Blended Space”), as well as applied explorations in areas such as generative AI, quantum computing and music, virtual orchestras, multimodal performance, pedagogy, and the rhetoric of “pushback” in conversational systems. It also contains an extended section in which Monday and DB Smith co-curate a set of student research …


Enhancing Remote Sensing Imagery Temporal Resolution Using Starfm Data Fusion Approach For Improved Land Surface Monitoring, Ahmadreza Pourghodrat Apr 2025

Enhancing Remote Sensing Imagery Temporal Resolution Using Starfm Data Fusion Approach For Improved Land Surface Monitoring, Ahmadreza Pourghodrat

School of Computing: Dissertations, Theses, and Student Research

High-resolution remote sensing imagery plays a critical role in various domains, such as farm-level agricultural operations, environmental monitoring, and natural resource management. However, data with high spatial resolution typically have low temporal resolution, and those with high temporal resolution often lack spatial detail. For example, Landsat 8 and 9 satellites deliver high spatial resolution images with a 30-meter pixel size but suffer from low temporal resolution, with a 16-day revisit cycle. In contrast, satellites like MODIS and VIIRS provide daily images but with a much coarser spatial resolution (375 meters or more), reducing spatial details. Additionally, there is a lack …


Binge Buddies, Joshua Uribe Apr 2025

Binge Buddies, Joshua Uribe

Posters - 2025

Many people struggle to keep track of the shows and movies they’ve watched or plan to watch. Existing streaming platforms often provide limited or cluttered tracking features, making it challenging to stay organized. Binge Buddies addresses this issue by centralizing watchlists and viewing history in one streamlined location. The website is designed to simplify the binge-watching experience, helping users stay on top of their content and discover new shows/movies. Which makes the experience a smoother and more enjoyable experience.


Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira Apr 2025

Applying Software Engineering Black-Box Methods For Testing Machine Learning Models, Timothy Elvira

Doctoral Dissertations and Master's Theses

This dissertation proposes researching an approach to incorporate and align Software black-box testing methods into Machine Learning (ML) applications, specifically in the context of computer vision models. Typically, testing methods within Software Engineering (SE) encompass a range of test types that assess levels of a software system, such as Unit, Integration, Functional, and System testing [1]. The testing spectrum offers two perspectives on the system: black-box, where the system’s code is hidden, and white-box, where the system's code is exposed for testing. Software Quality pairs testing with requirements, in a many-to-one relationship, to ensure proper validation of the software system. …


Logiclm: Robust Application Of Large Language Models With Logic Programming For Data Analytics, Evgeny Skvortsov, Shayan Mirjafari, Ojaswa Garg, Yilin Xia, Shaun Bowers, Bertram Ludäscher Mar 2025

Logiclm: Robust Application Of Large Language Models With Logic Programming For Data Analytics, Evgeny Skvortsov, Shayan Mirjafari, Ojaswa Garg, Yilin Xia, Shaun Bowers, Bertram Ludäscher

Computer Science Faculty Scholarship

We present LogicLM, an OLAP-style interactive data analysis system that leverages large language models (LLMs) and is configured using Logica, an enhanced logic programming language with aggregation support that compiles to SQL. LogicLM uses an LLM to translate natural language queries by end users into executable code for automatically generating data visualizations. For each natural-language query, LogicLM provides a verifiable OLAP-based configuration that users can view and modify to help ensure results are reliable and accurate. This configuration, with measures, dimensions, and filters defined as logical predicates, offers a unified and user-friendly approach to naturallanguage data exploration, while keeping end …


Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts Mar 2025

Leveraging Large Language Models For Knowledge-Free Weak Supervision In Clinical Natural Language Processing, Enshuo Hsu, Kirk Roberts

Faculty, Staff and Student Publications

The performance of deep learning-based natural language processing systems is based on large amounts of labeled training data which, in the clinical domain, are not easily available or affordable. Weak supervision and in-context learning offer partial solutions to this issue, particularly using large language models (LLMs), but their performance still trails traditional supervised methods with moderate amounts of gold-standard data. In particular, inferencing with LLMs is computationally heavy. We propose an approach leveraging fine-tuning LLMs and weak supervision with virtually no domain knowledge that still achieves consistently dominant performance. Using a prompt-based approach, the LLM is used to generate weakly-labeled …