Open Access. Powered by Scholars. Published by Universities.®

Software Engineering

Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 1 - 30 of 373

Full-Text Articles in Artificial Intelligence and Robotics

Mutation-Based Multi-Agent Test Case Update, Dawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang, Jianlei Chi, Jun Sun, Xiaohong Su Oct 2026

Mutation-Based Multi-Agent Test Case Update, Dawei Tian, Jiakun Liu, Yun Peng, Yichen Zhang, Jianlei Chi, Jun Sun, Xiaohong Su

Research Collection School Of Computing and Information Systems

Modern software systems evolve rapidly under CI/CD practices, where tests are critical for quality. However, substantial code changes often render existing test cases obsolete, causing pipeline disruptions, reduced productivity, and compromised quality. Recent automatic test update approaches leverage LLMs to refine test cases via execution feedback and exact-matching context retrieval, prioritizing executability and line coverage but suffering three limitations: (1) neglecting test assertion adequacy, weakening fault detection; (2) relying on coarse line coverage instead of specific uncovered lines/branches; (3) using exact-matching retrieval, which fails for LLM hallucinated queries. To address these, we propose MuMuTestUp, a mutation-guided multi-agent framework with three …


Ddor: Delta Debugging For Explainable Overrefusal Testing And Repair, Qinyan Zhou, Peixin Zhang, Jun Sun, Haonan Zhang, Dongxia Wang Oct 2026

Ddor: Delta Debugging For Explainable Overrefusal Testing And Repair, Qinyan Zhou, Peixin Zhang, Jun Sun, Haonan Zhang, Dongxia Wang

Research Collection School Of Computing and Information Systems

While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwarranted rejection of benign queries that merely appear risky. We present DDOR (Delta Debugging for OverRefusal), a fully automated and explainable framework for overrefusal testing and repair in a black-box setting, where only model inputs and outputs are accessible and internal safety mechanisms remain opaque. DDOR applies delta debugging to localize minimal refusal-triggering fragments (mRTFs) that provide phrase-level, explainable evidence for why a refusal occurs. Conditioned on these mRTFs, DDOR generates diverse, context-rich prompts and performs multi-oracle validation to filter intrinsically …


A Hybrid Llm-Srgm Framework For Ai-Enabled Reliability Assessment In Safety-Critical Software Systems, Caleb Stone, Shrenik Jadhav Aug 2026

A Hybrid Llm-Srgm Framework For Ai-Enabled Reliability Assessment In Safety-Critical Software Systems, Caleb Stone, Shrenik Jadhav

Discovery Day - Daytona Beach

Ensuring the reliability of software intensive and safety critical systems is a persistent challenge across aerospace, defense, transportation, and other mis- sion focused domains. Traditional software relia- bility growth models (SRGM) provide useful quanti- tative insight into defect discovery trends, but they rely mostly only on numerical failure data and do not use the rich contextual information contained in test logs, anomaly reports, and engineering notes. This paper presents a hybrid framework that com- bines semantic features extracted by a large lan- guage model (LLM) with a non-homogeneous Pois- son process (NHPP) based software reliability growth model. The LLM analyzes …


Smart Atm, Majid Hakeem Jul 2026

Smart Atm, Majid Hakeem

Systems Manuals - 2026

The Smart ATM is a new system that goal is to create a new experience for ATMs users by creating a different way of using the ATM. It would protect users from germs and bacteria. Instead of using buttons and touch screens, the main input for this system will be the motion sensor. To specify, the system uses Microsoft Kinect, which is the motion sensor for the Xbox One.

This document is a user guide, which is intended to give assistance to users for using the Smart ATM. This document contains an application overview, list of user requirements, the design …


Solving Rubik's Cube By Using Artificial Intelligence, Polat Coban, Seth Reed Jul 2026

Solving Rubik's Cube By Using Artificial Intelligence, Polat Coban, Seth Reed

Systems Manuals - 2026

This document will detail a proposal to build a Rubik’s cube simulator, and a Rubik’s cube solver. It is broken into several sections which in turn are broken into subsections.


Improving Urban Search And Rescue Team Coordination Through Adaptive Context Awareness, Daniel Reyes Duran Jul 2026

Improving Urban Search And Rescue Team Coordination Through Adaptive Context Awareness, Daniel Reyes Duran

Doctoral Dissertations and Master's Theses

Modern multi-agent Urban Search and Rescue (USAR) operations heavily rely on mobile geospatial Common Operating Pictures (COPs) to maintain team coordination and Situational Awareness (SA). However, the proliferation of high-frequency sensor telemetry at the tactical edge has introduced a data saturation paradox challenge: while information theoretically drives informed decision-making, unmanaged data surges induce increased operator cognitive overload and alert fatigue on mobile End-User Devices (EUDs), while downstream data-broadcasting models inherently strain edge processing and viewport environments.

To resolve these constraints, this dissertation presents a context-aware Value of Information (VoI) data-management framework integrated directly with a custom, event-driven Android Team Awareness …


Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo Jul 2026

Survey On Learning-Based Dynamic Fault Localization: From Traditional Machine Learning To Large Language Models, Chunyan Liu, Yan Lei, Huan Xie, Jinping Wang, Yue Yu, David Lo

Research Collection School Of Computing and Information Systems

Learning-based dynamic fault localization techniques play a crucial role in the field of software engineering. These techniques dynamically execute test cases to meticulously extract useful knowledge from the execution information in the program, with the aim of identifying fault locations by leveraging machine learning, deep learning, and large language models. Currently, there is already a flourishing body of research that is intensely focused on learning-based dynamic fault localization. Research literature can be categorized into two main aspects for learning-based dynamic fault localization: data-based enhancements (i.e., the datasets) and model-based enhancements (i.e., the suspiciousness algorithms). Thus, we conduct an extensive literature …


Accountable Agents In Software Engineering: An Analysis Of Terms Of Service And A Research Roadmap, Christoph Treude Jul 2026

Accountable Agents In Software Engineering: An Analysis Of Terms Of Service And A Research Roadmap, Christoph Treude

Research Collection School Of Computing and Information Systems

AI coding assistants and autonomous agents are becoming integral to software development workflows, reshaping how code is produced, reviewed, and maintained. While recent research has focused mainly on the capabilities and impacts of productivity of these systems, much less attention has been paid to accountability: who is responsible when agents generate, modify, or recommend code? In practice, accountability is defined through the Terms of Service (ToS) and related policy documents that govern the use of AI-powered development tools.In this vision paper, we present a comparative analysis of the Terms of Service for widely used AI coding assistants and agent-enabled development …


Configuring Agentic Ai Coding Tools: An Exploratory Study, Matthias Galster, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes Jul 2026

Configuring Agentic Ai Coding Tools: An Exploratory Study, Matthias Galster, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes

Research Collection School Of Computing and Information Systems

Agentic AI coding tools increasingly automate software development tasks. Developers can configure these tools through versioned repository-level artifacts such as Markdown and JSON files. We present a systematic analysis of configuration mechanisms for agentic AI coding tools, covering Claude Code, GitHub Copilot, Cursor, Gemini, and Codex. We identify eight configuration mechanisms spanning from static context to executable and external integrations and, in an empirical study of 2,853 GitHub repositories, examine whether and how they are adopted, with a detailed analysis of Context Files, Skills, and Subagents. First, Context Files dominate the configuration landscape and are often the sole mechanism in …


Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong Jul 2026

Operationalizing Ethics For Ai Agents: How Developers Encode Values Into Repository Context Files, Christoph Treude, Sebastian Baltes, Marc Cheong

Research Collection School Of Computing and Information Systems

As AI coding agents become embedded in software development workflows, developers are beginning to operationalize ethical principles by encoding behavioral rules into repository-level context files for AI agents, such as AGENTS.md files. Rather than examining the ethics of AI agents in the abstract, this vision paper investigates how ethics and values are already being translated for AI agents into actionable instructions that shape agent behavior. Through a preliminary investigation, we find that developers are already embedding guidance related to fairness, accessibility, sustainability, tone, and privacy. These artifacts function as a developer-authored governance layer, translating abstract principles into situated, natural-language directives …


A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes Jul 2026

A Dataset Of Agentic Ai Coding Tool Configurations, Matthias Galster, Seyedmoein Mohsenimofidi, Levi Böhme, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes

Research Collection School Of Computing and Information Systems

Agentic AI coding tools such as Claude Code and OpenAI Codex execute multi-step coding tasks with limited human oversight. To steer these tools, developers create repository-level configuration artifacts (e.g., Markdown files) for configuration mechanisms such as Context Files, Skills, Rules, and Hooks. There is no curated dataset yet that captures these configurations at scale. This dataset, collected from open-source GitHub repositories, fills that gap. We selected 40,585 actively maintained repositories through metadata filtering, classified them using GPT-5.2 to identify 36,710 as belonging to engineered software projects, and systematically detected configuration artifacts in these repositories. The dataset covers 4,738 repositories across …


Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen Jul 2026

Train In Vain: Functionality-Preserving Poisoning To Prevent Unauthorized Use Of Code Datasets, Yuan Xiao, Yuchen Chen, Jiaming Wang, Wei Song, Jun Sun, Shiqing Ma, Yanzhou Mu, Juan Zhai, Chunrong Fang, Jin Song Dong, Zhenyu Chen

Research Collection School Of Computing and Information Systems

The widespread availability of large-scale code datasets has accelerated the development of code large language models (CodeLLMs), raising concerns about unauthorized dataset usage. Dataset poisoning offers a proactive defense by reducing the utility of such unauthorized training. However, existing poisoning methods often require full-dataset poisoning and introduce transformations that break code compilability. In this paper, we introduce FunPoison, a functionality-preserving poisoning approach that injects short, compilable weak-use fragments into executed code paths. FunPoison leverages reusable statement-level templates with automatic repair and conservative safety checking to ensure side-effect freedom, while a type-aware synthesis module preserves type correctness, suppresses static-analysis warnings, and …


Robust Real-Time Uav Target Tracking With Onboard Vision-Based Yaw Control, Rylan Malarchick, Jose Castelblanco, Enrique Amaya, Carmen Dimario, Graysen Brinkman, Chirag Kumar, Kiwon Yoon, Sajid Berhane Jun 2026

Robust Real-Time Uav Target Tracking With Onboard Vision-Based Yaw Control, Rylan Malarchick, Jose Castelblanco, Enrique Amaya, Carmen Dimario, Graysen Brinkman, Chirag Kumar, Kiwon Yoon, Sajid Berhane

Beyond: Undergraduate Research Journal

Autonomous tracking of agile unmanned aerial vehicles (UAVs) presents significant challenges for real-time perception and control systems. This work presents AIRHOUND (Autonomous Intelligent Rotorcraft for Hostile Object Unified Navigation and Detection), a UAV platform implementing vision-based yaw tracking through a modular ROS2 software architecture. The system employs YOLOv8 object detection optimized with NVIDIA TensorRT for embedded deployment on an NVIDIA Jetson Orin companion computer. Detected targets are processed through a geometric tracking module that converts pixel coordinates to angular yaw errors using pinhole camera intrinsics, with a proportional controller generating rate-limited yaw commands. These commands are streamed to a PX4 …


Evaluation And Distillation Of Source Code Generation Tasks By Large Language Models, Danny Brahman Jun 2026

Evaluation And Distillation Of Source Code Generation Tasks By Large Language Models, Danny Brahman

Electronic Theses and Dissertations

Large Language Models (LLMs) are predominantly assessed based on their common sense reasoning, language comprehension, and logical reasoning abilities. While models trained in specialized domains like mathematics or coding have demonstrated remarkable advancements in logical reasoning, there remains a significant gap in evaluating their code generation capabilities. Existing benchmark datasets fall short in pinpointing specific strengths and weaknesses, impeding targeted enhancements in models’ reasoning abilities to synthesize code.

To bridge this gap, this thesis introduces two novel contributions: CodeEval and CodeQual. CodeEval is an innovative, pedagogical benchmarking method that mirrors the evaluation processes encountered in academic programming courses. It comprises …


Ai Interview Helper: A Tool For Assisting Search And Rescue Long-Profile Interviews, Dylan P. Starink Jun 2026

Ai Interview Helper: A Tool For Assisting Search And Rescue Long-Profile Interviews, Dylan P. Starink

Master's Theses

In Search and Rescue (SAR) operations, time pressure and limited interviewer experience can lead to missed opportunities when interviewing a missing person’s friends and family. This thesis presents a real-time, end-to-end system that provides context-aware follow-up question suggestions as interviews unfold. Leveraging large language models (LLMs) and agentic design patterns, the system is intended to support interviewers by helping them identify relevant follow-up questions and pursue potentially overlooked lines of inquiry.

The system was evaluated through three mock interviews with two SAR interviewer participants across two events. Given the limited sample size, the results provide early insights into the feasibility …


The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala Jun 2026

The Security Of Llm-Generated Code, Christopher Brian Gonzalez Ayala

Student Theses

The rapid adoption of Large Language Models (LLMs) in software development has transformed coding practices by enabling automated code generation, completion, and optimization. Despite these advantages, concerns persist regarding the security and reliability of LLM-generated code. This study presents a comprehensive evaluation of both the functional correctness and security of code produced by three prominent LLMs as of early 2026. A total of 4,800 code snippets were generated using 100 security-focused programming prompts derived from the OWASP Top 10:2025, translated across eight natural languages and two phrasing styles (literal and natural developer-oriented prompts). To assess performance, a multi-stage experimental framework …


Towards Auto-Evaluation For Large Language Models, Jiahao Ying Jun 2026

Towards Auto-Evaluation For Large Language Models, Jiahao Ying

Dissertations and Theses Collection (Open Access)

The rapid advancement of large language models (LLMs) has created an urgent need for evaluation methodologies that are timely, scalable, reliable, and informative. Conventional evaluation benchmarks, although essential for measuring model capabilities and guiding model development, are often constructed and maintained through labor-intensive human annotation. As LLMs continue to improve through increases in model scale, training data, and computational resources, static benchmarks may quickly lose discriminative power. Moreover, the growing use of large and diverse training corpora increases the risk of benchmark leakage, which can inflate evaluation results and obscure the true capabilities of models. These challenges call for a …


Political Inconsistency Detection Across Legislative Speech And Public Communications, Scott M. Pramuk Jun 2026

Political Inconsistency Detection Across Legislative Speech And Public Communications, Scott M. Pramuk

Master's Theses

Political actors communicate about legislation across multiple contexts, including committee hearings, recorded votes, and public-facing press releases. Differences between these forms of communication can provide useful signals for journalists and researchers seeking to understand how legislators present policy positions to different audiences.

This thesis extends the Digital Democracy Project, a legislative transparency initiative that provides access to California state legislative hearing transcripts, voting records, and related legislative data. Specifically, this work incorporates publicly accessible, legislator-authored news releases into the Digital Democracy Database and develops a pipeline for analyzing legislative communication across multiple sources. The system collects news releases from California …


“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt Jun 2026

“Alexa, Do Not Say That In Front Of My Boss!” A Cross-Cultural Comparison Of User And Ai Preferences For Privacy-Aware Smart Speaker Interactions Across Contexts, Lynne Warin, Emily Aurelia, Anthony Tang, Emily Aurelia, Delphine Reinhardt

Research Collection School Of Computing and Information Systems

Due to their limited ability to reason about the social context in which they are used, smart speakers pose significant privacy risks by responding in ways that may violate people's implicit social boundaries. We conducted a cross-cultural vignette study (N = 944) in Germany and Singapore to investigate how situational factors—specifically social context (bystander relationships and closeness), physical context (location), and interaction context (topic and deceptive intent)—regulate user preferences for smart speaker responses. Our results demonstrate that these factors are superior predictors of response preferences than dispositional user traits (i.e., intrinsic personal traits). We identify two distinct social dynamics: a …


How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández Jun 2026

How Do Machine Learning Models Change?, Joel Castaño, Rafael Cabañas, Antonio Salmerón, David Lo, Silverio Martínez-Fernández

Research Collection School Of Computing and Information Systems

The proliferation of Machine Learning (ML) models and their open source implementations has transformed AI research and applications. Platforms like Hugging Face (HF) enable this evolving ecosystem, yet a large-scale longitudinal study of how these models change is lacking. This study addresses this gap by analyzing over 680,000 commits from 100,000 models and 2,251 releases from 202 of these models on HF using repository mining and longitudinal methods. We apply an extended ML change taxonomy to classify commits and use Bayesian networks to model temporal patterns in commit and release activities. Our findings show that commit activities align with established …


Synthergy: Social Deduction And Deception In Llm-Powered Agents, Lauren Campbell, Andrew Forney May 2026

Synthergy: Social Deduction And Deception In Llm-Powered Agents, Lauren Campbell, Andrew Forney

Honors Thesis

Synthergy is an online social deduction game designed to enable comparative analysis of how large language model-powered agents engage in social deduction and deception under conditions of asymmetric information. Inspired by social deduction games such as Town of Salem, Throne of Lies, and Mafia, the game consists of two factions, Harmony and Discord, to which agents are secretly assigned. Agents must infer others’ affiliations through dialogue, in-game abilities, and voting behavior. To evaluate agent behavior, we conducted 100 simulated games across six agent types: a random baseline agent (RandomSynth), an LLM-based agent (Synth), a chain-of-thought agent (CoT Synth), a Bayesian …


Developing Narrative-Based Stem Learning Tool For K-6 Visually Impaired Students, Daniel Tsivkovski, Dylan Ravel, Jeffrey Kraskouskas, Brandon Foley, Maryam Etezad, Franceli Cibrian, Rajeev Joshi, Ariel Han May 2026

Developing Narrative-Based Stem Learning Tool For K-6 Visually Impaired Students, Daniel Tsivkovski, Dylan Ravel, Jeffrey Kraskouskas, Brandon Foley, Maryam Etezad, Franceli Cibrian, Rajeev Joshi, Ariel Han

Student Scholar Symposium Abstracts and Posters

This research develops a free, accessible web application that enables K-6 students who are blind or visually impaired (BVI) to learn STEM concepts using refreshable braille displays. Currently, most online learning tools are not designed for BVI students, creating a significant educational barrier.

The application interfaces with commercial braille displays and uses narrative-based learning to make STEM content approachable and engaging. By presenting material as personalized interactive stories generated with the help of Artificial Intellligence (AI), students can connect with concepts while developing braille reading skills. The curriculum design prioritizes accessibility through the Accessible Rich Internet Applications (ARIA) standards and …


Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder May 2026

Pyspqr: A Python Package For Density Estimation Using Deep Learning, Cameron Eddy, Reetam Majumder

Electrical Engineering and Computer Science Undergraduate Honors Theses

Splines are used for representing complex functions. In statistics, splines can be used for distributional shapes that are difficult to model by traditional parametric approaches. Ramsay (1) uses M-Spline bases to estimate continuous distributions. Semi-Parametric Quantile Regression (SPQR), developed by Xu and Reich (2), models conditional distributions where a neural network is used to estimate the basis function weights that depend on covariates. (3) implements a package for SPQR in R. We build on this by implementing a version of SPQR in Python with PyTorch. By using PyTorch, we can use more sophisticated deep learning architectures than those available in …


Sql Query Optimization - Human Vs. Chatgpt, Hailey Dennis May 2026

Sql Query Optimization - Human Vs. Chatgpt, Hailey Dennis

All Graduate Reports and Creative Projects, Fall 2023 to Present

Large Language Models (LLMs) such as ChatGPT have become ubiquitous tools for working professionals in the software industry. Many engineers are finding new ways to increase productivity by offloading tasks onto LLMs, while others are finding it difficult to trust code produced artificially, even after review. Taking a look at both perspectives, this study aims to compare a human’s ability to optimize SQL queries to that of an LLM and assess the experience using both methods.

Manual query optimization is a tedious task that relies heavily on statistics, heuristics, and good intuition. The SQL developer must search for the optimal …


The Quality Assurance Machine – A Software Quality Assurance Architecture For Ml-Enabled Systems, Shane E. Downing May 2026

The Quality Assurance Machine – A Software Quality Assurance Architecture For Ml-Enabled Systems, Shane E. Downing

All-Inclusive List of Electronic Theses and Dissertations

This dissertation evaluates whether a reusable assurance architecture, the Quality Assurance Machine (QAM), can provide effective product and process quality assurance for ML-enabled software platforms. The QAM is a system-level SQA architecture that turns plans and policies into versioned configurations, executes them in controlled environments, and produces preserved run evidence that supports traceability, auditability, and controlled change. The study follows Design Science Research and evaluates the instantiated artifact using eight assurance requirements (AR1–AR8) synthesized from standards-based guidance, including IEEE 730 and ISO/IEC/IEEE 15026. A four-year longitudinal evaluation combines two methods. First, operational evidence from routine regression and release-validation runs, defect …


Autonomous Deficiency Detection And Vision-Language Summarization For Underground Infrastructure On Embedded Edge Systems, Johny Lopez May 2026

Autonomous Deficiency Detection And Vision-Language Summarization For Underground Infrastructure On Embedded Edge Systems, Johny Lopez

LSU New Orleans Theses and Dissertations

Aging underground infrastructure poses significant risks to public health and environmental safety, yet structural condition assessment remains bottlenecked by labor-intensive manual CCTV inspections. This thesis proposes a comprehensive algorithmic framework enabling fully autonomous, real-time deficiency detection, geometric assessment, and natural language reporting on resource- constrained edge computing platforms. Three core components address this challenge. First, RAPID-SCAN, a novel semantic segmentation architecture utilizing a Dynamic Feature Pyramid Network and Channel-Spatial Attention, achieves real-time, pixel-precise defect localization with dramatically reduced parameters. Second, an Edge-Optimized Vision-Language Model pipeline employing LoRA and 4-bit QLoRA quantization compresses Phi-3.5 for local deployment, en- abling autonomous technical …


Selective Concolic Testing, Guofeng Zhang, Zhenbang Chen, Ziqi Shuai, Jun Sun, Weijiang Hong, Yufeng Zhang, Ji Wang, Yang Liu May 2026

Selective Concolic Testing, Guofeng Zhang, Zhenbang Chen, Ziqi Shuai, Jun Sun, Weijiang Hong, Yufeng Zhang, Ji Wang, Yang Liu

Research Collection School Of Computing and Information Systems

The principled combination of symbolic execution and random testing lacks a formal foundation, especially in deciding which inputs to symbolize. We propose selective concolic testing, a cost-aware framework that formulates this choice as an optimized policy problem of a MDP (Markov Decision Process). We model program exploration over a finite control-flow graph, where MDP states represent covered statements, actions partition path constraints into symbolic and random fragments, rewards reflect coverage gain, and costs account for SMT solving effort and sampling inefficiency. Our framework yields the first formal characterization of selective symbolization as policy synthesis in a probabilistic system. We prove …


Gencode: A Generic Data Augmentation Framework For Boosting Deep Learning-Based Code Understanding, Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Yves Le Traon, Jianjun Zhao May 2026

Gencode: A Generic Data Augmentation Framework For Boosting Deep Learning-Based Code Understanding, Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy, Mike Papadakis, Yves Le Traon, Jianjun Zhao

Research Collection School Of Computing and Information Systems

Pre-trained code models lead the era of code intelligence, with multiple models designed with impressive performance. However, one important problem, data augmentation for code data that automatically helps developers prepare training data lacks study in this field. In this paper, we introduce a generic data augmentation framework, GenCode, to enhance the training of code understanding models. Simply speaking, GenCode follows a generation-and-selection paradigm to prepare useful training code data. Specifically, it employs code augmentation techniques to generate new code candidates first and then identifies important ones as the training data by influence scores. To evaluate the effectiveness of GenCode, we …


Llm-Based Stock Sentiment And Market Intelligence Platform, Joshua Thrower, Andrew Pinkerton, Ian Duggan, Wyatt Lester Apr 2026

Llm-Based Stock Sentiment And Market Intelligence Platform, Joshua Thrower, Andrew Pinkerton, Ian Duggan, Wyatt Lester

ATU Scholars Symposium

Financial markets increasingly react to social media discourse, yet investors lack tools to translate this unstructured commentary into measurable indicators. Platforms such as YouTube host extensive discussions about publicly traded equities, but extracting reliable sentiment trends from high-volume, noisy comment streams remains technically challenging. This project develops a stock sentiment and market intelligence platform that transforms YouTube comment data into aggregated sentiment indicators aligned to specific equities. Comments are mapped to equities using ticker specific keyword identification combined with contextual filtering to reduce false associations from ambiguous or off-topic mentions. The system assigns numerical sentiment scores to individual comments and …


Towards Physics-Informed Neural Networks For Simulating Multiphase Geothermal Convection​*, Daniel C. Patton, Andrew Harrison Eno Apr 2026

Towards Physics-Informed Neural Networks For Simulating Multiphase Geothermal Convection​*, Daniel C. Patton, Andrew Harrison Eno

Campus Research Month

Water and steam flow through porous rock, transferring heat via conduction and buoyancy-driven convection caused by density differences. Traditional numerical methods (finite-volume/finite-element) model this well but can become memory-intensive and unstable for long, high-detail simulations. This work demonstrates a Physics-Informed Neural Network (PINN) using a finite-difference approach within the NVIDIA PhysicsNeMo framework to simulate magma chambers in 2D. Tested on the Rio Pisco pluton in Peru, results are compared with the USGS HYDROTHERM model. PINNs learn from physical laws, offering accurate, flexible solutions with less data and development effort.