Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Institution
-
- Singapore Management University (137)
- City University of New York (CUNY) (18)
- Chapman University (17)
- California Polytechnic State University, San Luis Obispo (16)
- Western University (12)
-
- University of Arkansas, Fayetteville (11)
- University of Malaya (11)
- Loyola University Chicago (9)
- St. Mary's University (8)
- Embry-Riddle Aeronautical University (7)
- Old Dominion University (5)
- The University of Akron (5)
- California State University, San Bernardino (4)
- Claremont Colleges (4)
- Kennesaw State University (4)
- Montclair State University (4)
- San Jose State University (4)
- University of Kentucky (4)
- LSU New Orleans (3)
- Portland State University (3)
- Southern Methodist University (3)
- University of Texas at Arlington (3)
- West Virginia University (3)
- Arkansas Tech University (2)
- Belmont University (2)
- Bucknell University (2)
- Dartmouth College (2)
- East Tennessee State University (2)
- Edith Cowan University (2)
- Fort Hays State University (2)
- Keyword
-
- Machine learning (29)
- Deep learning (23)
- Machine Learning (20)
- Software engineering (19)
- Deep Learning (15)
-
- Artificial Intelligence (14)
- Artificial intelligence (10)
- Computer Science (10)
- Large language models (10)
- Artificial Intelligence (AI) (9)
- Computer Vision (9)
- Refactoring (8)
- ChatGPT (7)
- Computational thinking (7)
- Computer science (7)
- Cybersecurity (7)
- Generative AI (7)
- Imperative programs (7)
- Large Language Models (7)
- Large language model (7)
- Software Engineering (7)
- Computer vision (6)
- Python (6)
- AI (5)
- Equity (5)
- Large Language Model (5)
- Natural Language Processing (5)
- Robotics (5)
- Empirical studies (4)
- Hybrid programming paradigms (4)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (129)
- Journal of Computer Science Integration (16)
- Publications and Research (13)
- Electrical and Computer Engineering Publications (12)
- Master's Theses (11)
-
- Computer Science: Faculty Publications and Other Works (9)
- Computer Science and Computer Engineering Undergraduate Honors Theses (8)
- Student Works (2000-2009) (8)
- Electronic Theses and Dissertations (5)
- Williams Honors College, Honors Research Projects (5)
- CMC Senior Theses (4)
- Dissertations and Theses Collection (Open Access) (4)
- Posters - 2026 (4)
- College of Engineering Summer Undergraduate Research Program (3)
- Department of Computer Science Faculty Scholarship and Creative Works (3)
- Dissertations (3)
- Doctoral Dissertations and Master's Theses (3)
- Electronic Theses, Projects, and Dissertations (3)
- Graduate Theses, Dissertations, and Problem Reports (ETD) (3)
- Honors Theses (3)
- LSU New Orleans Theses and Dissertations (3)
- Open Educational Resources (3)
- Student Works (2020-2029) (3)
- Theses and Dissertations (3)
- Theses and Dissertations--Computer Science (3)
- ATU Scholars Symposium (2)
- Computer Engineering (2)
- Computer Science and Engineering Theses and Dissertations (2)
- Dartmouth College Master’s Theses (2)
- Engineering Management & Systems Engineering Faculty Publications (2)
- Publication Type
- File Type
Articles 31 - 60 of 373
Full-Text Articles in Artificial Intelligence and Robotics
Match-A-Fit, Brianna Mendoza, Adan Diaz De Leon, Pedro Jacobo, Juan Marco Saca Dada
Match-A-Fit, Brianna Mendoza, Adan Diaz De Leon, Pedro Jacobo, Juan Marco Saca Dada
Posters - 2026
Welcome to Match-a-Fit! Match-a-Fit is an iOS application that allows the user to create a digital closet by uploading images of their clothing items. With AI, the program can generate outfits based on the digital closet, the time, and the occasion. Match-a-Fit’s purpose is designed to help users who struggle to get ready, run out of time, or can’t decide on an outfit, by easily generating outfit options based on the occasion.
A.I.R.E., Laurene Robinson
A.I.R.E., Laurene Robinson
Presentations - 2026
•Cybersecurity analysts rely on reverse engineering to understand suspicious software. •Ghidra can surface decompiled code, but it does not fully explain function purpose, behavioral meaning, or analyst priority. •When symbols are stripped and context is weak, analysts must still reconstruct intent manually from low-level output. •That process is Time-consuming , complex and , operationally costly
A.I.R.E. - Ai-Assisted Reverse Engineering, Laurene Robinson
A.I.R.E. - Ai-Assisted Reverse Engineering, Laurene Robinson
Posters - 2026
Reverse engineering plays a vital role in cybersecurity by helping analysts examine unknown binaries, investigate malware, identify vulnerabilities, and better protect sensitive systems. However, once a program is compiled and stripped, the meaningful names that describe its behavior are lost, leaving behind generic function labels like FUN_00401a30. Analysts must then manually interpret decompiled code, trace call chains, and infer program behavior function by function, which is slow and mentally demanding on large binaries. To address this challenge, this project introduces A.I.R.E., a local Ghidra extension that extracts contextual evidence from stripped functions and uses a locally hosted language model to …
Sentra, David Aguilar
Sentra, David Aguilar
Posters - 2026
In the fast-paced space of event organization, fostering continuous collaboration among participants is essential. However, organizers often lose valuable time monitoring multiple, disconnected systems once an event is underway. Enter Sentra: an all-in-one Discord bot tailored specifically for weekend events like hackathons. Sentra bridges the gap between participants and organizers by consolidating seamless team matchmaking, robust support ticketing, and automated AI moderation into a single, unified interface.
Synapse, Nicolas Diaz, Alexander Murphy, Sonia Cerrillo, Naomi Ramirez, Jesse Kemmer
Synapse, Nicolas Diaz, Alexander Murphy, Sonia Cerrillo, Naomi Ramirez, Jesse Kemmer
Posters - 2026
People tend to accumulate a great deal of notes throughout their lives with no coherent way to organize them. Even with the built-in notes app, the notes eventually accumulate until it becomes borderline impossible to find what is needed. Our proposed solution is Synapse, an LLM powered notes app with a tagging system that allows notes to be sorted by topic. The LLM will be able to read the user's notes and recommend tags
Match-A-Fit, Adan Diaz De Leon, Juan Marco Saca Dada, Brianna Mendoza, Arsalan Kataneh, Theophile Nsabimana, Pedro Jacobo
Match-A-Fit, Adan Diaz De Leon, Juan Marco Saca Dada, Brianna Mendoza, Arsalan Kataneh, Theophile Nsabimana, Pedro Jacobo
Presentations - 2026
Welcome to Match-a-Fit! Match-a-Fit is an iOS application that allows the user to create a digital closet by uploading images of their clothing items. With AI, the program can generate outfits based on the digital closet, the time, and the occasion. Match-a-Fit’s purpose is designed to help users who struggle to get ready, run out of time, or can’t decide on an outfit, by easily generating outfit options based on the occasion.
Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang
Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang
Research Collection School Of Computing and Information Systems
Since the launch of ChatGPT, a powerful AI Chatbot developed by OpenAI, large language models (LLMs) have made significant advancements in both academia and industry, bringing about a fundamental engineering paradigm shift in many areas. While LLMs are powerful, it is also crucial to best use their power where “prompt” plays a core role. However, the booming LLMs themselves, including excellent APIs like ChatGPT, have several inherent limitations: (1) temporal lag of training data, and (2) the lack of physical capabilities to perform external actions. Recently, we have observed the trend of utilizing prompt-based tools to better utilize the power …
Bridging Bug Localization And Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models, Jianming Chang, Xin Zhou, Lulu Wang, David Lo, Bixin Li
Bridging Bug Localization And Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models, Jianming Chang, Xin Zhou, Lulu Wang, David Lo, Bixin Li
Research Collection School Of Computing and Information Systems
Automated issue fixing is a critical task in software debugging and has recently garnered significant attention from academia and industry. However, existing fixing techniques predominantly focus on the repair phase, often overlooking the importance of improving the preceding bug localization phase. As a foundational step in issue fixing, bug localization plays a pivotal role in determining the overall effectiveness of the entire process. To enhance the precision of issue fixing by accurately identifying bug locations in large-scale projects, this paper presents BugCerberus, the first hierarchical bug localization framework powered by three customized large language models. First, BugCerberus analyzes intermediate representations …
Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo
Research Collection School Of Computing and Information Systems
Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with Large Language Model (LLM)-powered agents, but existing approaches either rely on a single generic agent that struggles in complex scenarios or narrowly specialized agents that cannot adapt to diverse vulnerability types. We therefore introduce PenForge, a framework that dynamically constructs expert agents during testing rather than relying on those prepared beforehand. By integrating automated reconnaissance of potential attack surfaces with agents instantiated on the fly for context-aware exploitation, PenForge achieves a 30.0% exploit success rate (12/40) …
Patchgpt: Multi-Agent Patch Backporting Without Model Fine-Tuning, Ye Liu, Ruidong Han, Chengyan Ma, Yuqing Niu, David Lo
Patchgpt: Multi-Agent Patch Backporting Without Model Fine-Tuning, Ye Liu, Ruidong Han, Chengyan Ma, Yuqing Niu, David Lo
Research Collection School Of Computing and Information Systems
Patch backporting is crucial and prevalent in the maintenance of modern open-source software such as Linux kernels and forked repositories. However, porting patches across program versions remains a challenging problem due to the complexity of synergizing diverse patches with divergent program versions. In this paper, we propose PatchGPT, an agentic patch backporting framework for fine-grained patch generation. PatchGPT encompasses three agents: Miner for decomposing a sequence of atomic change steps as the original patch plan, Adapter for adapting the patch plan, and Executor for executing the adapted patch plan according to predefined change semantics. We conduct experiments on the PPatHF’s …
Large-Scale File Fragment Classification Via Multi-View Learning, Samuel Hildebrand
Large-Scale File Fragment Classification Via Multi-View Learning, Samuel Hildebrand
LSU Master's Theses
File reassembly is one of the most fundamental tasks in digital forensics, enabling recovery of data from potentially damaged storage media even when file system metadata is unavailable. This thesis reviews more than two decades of work in the realm of file carving, with a particular focus on fragmented file carving, which remains a focus of research, and file fragment classification, a principal component of fragmented file carving. This thesis serves a literature review of both file carving and fragmented file carving, surveys the massive amounts of data needed for the task of fragment classification and the datasets that serve …
Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal
Anomaly Detection For Multi-System Bug Triage, Gibran Miguel Zavala Gamero, Hayoung Cheon, Mustafa Iqbal
SMU Data Science Review
Large-scale software systems produce vast volumes of logs and telemetry, making manual incident triage slow and error prone. This study presents an unsupervised anomaly detection pipeline that fuses logs, metrics, and traces through late fusion. Using Hybrid Ensemble modeling with Isolation Forest, and Long Short-Term Memory (LSTM) Deep Learning model, the system detects cross-service anomalies producing and assigning a composite triage score reflecting severity and impact. Ranked alerts are categorized into Critical, High, or Medium priorities for review. A retrieval-augmented generation (RAG) layer enriches results with contextual summaries for explainable triage. Evaluated on synthetic multi-service datasets, the pipeline …
Optimizing And Fortifying Ai Software Through The Lens Of Artifact Synthesis, Jieke Shi
Optimizing And Fortifying Ai Software Through The Lens Of Artifact Synthesis, Jieke Shi
Dissertations and Theses Collection (Open Access)
Artificial Intelligence (AI) has transformed the software landscape, ushering in a new era of intelligent systems that increasingly shape our daily lives. This transformation is evident in various domains, including Software Engineering (SE), where Large Language Models (LLMs) support many development tools, and control systems, where self-driving cars and autonomous drones rely on deep learning models for real-time decision-making. These AI systems are collectively referred to as AI software, with the former categorized as AI4SE software (AI for Software Engineering) and the latter as AI4Control software (AI for Control). As AI software becomes central to modern computing infrastructure, its reliability …
How Agile Became The Design Philosophy Of Ai Fishbowl Under Real-World Constraints, Jad Saad
How Agile Became The Design Philosophy Of Ai Fishbowl Under Real-World Constraints, Jad Saad
University Honors Theses
This capstone review examines the development of AI Fishbowl, a public-facing, interactive artificial intelligence system, as a case study in how Agile methods evolve from a project management tool into a design philosophy under real-world constraints. Although the project adopted an Agile workflow early on through a Kanban-style task management approach, the initial system design and architecture were still shaped by a largely plan-first mindset. This created a mismatch between flexible process and rigid design assumptions, which became increasingly apparent as the team moved from high-level architecture into implementation.
A critical turning point occurred when early architectural plans proved difficult …
Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui
Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui
Research Collection School Of Computing and Information Systems
Evaluating the alignment of large language models (LLMs) with user-defined coding preferences is a challenging endeavor that requires a deep assessment of LLMs' outputs. Existing methods and benchmarks rely primarily on automated metrics and static analysis tools, which often fail to capture the nuances of user instructions and LLM outputs. To address this gap, we introduce the LLM-as-a-Judge evaluation framework and present CodeUltraFeedback, a comprehensive dataset for assessing and improving LLM alignment with coding preferences. CodeUltraFeedback consists of 10,000 coding instructions, each annotated with four responses generated from a diverse pool of 14 LLMs. These responses are annotated using GPT-3.5 …
Identifying And Mitigating Api Misuse In Large Language Models, Terry Yue Zhuo, Junda He, Jiamou Sun, Zhenchang Xing, David Lo, John Grundy, Xiaoning Du
Identifying And Mitigating Api Misuse In Large Language Models, Terry Yue Zhuo, Junda He, Jiamou Sun, Zhenchang Xing, David Lo, John Grundy, Xiaoning Du
Research Collection School Of Computing and Information Systems
API misuse in code generated by large language models (LLMs) presents a serious and growing challenge in software development. While LLMs demonstrate impressive code generation capabilities, their interactions with complex library APIs are often error-prone, potentially leading to software failures and vulnerabilities. In this paper, we conduct a large-scale study of API misuse patterns in LLM-generated code, analyzing both method selection and parameter usage across Python and Java, using three representative LLMs (StarCoder-7B, Qwen2.5-Coder-7B, and GitHub Copilot). Based on extensive manual annotation of 3,209 method-level and 3,492 parameter-level misuses, we identify and categorize four recurring misuse types by building on …
Large Language Model Communication And Data Transfer Across A Simulated Telephone Line, Jared Reyes
Large Language Model Communication And Data Transfer Across A Simulated Telephone Line, Jared Reyes
Dissertations and Theses
This thesis presents a proof-of-concept communication system in which two local large language model endpoints communicate across a simulated analog telephone line using legacy USB modems. The project combines modem voice mode, modem data mode, local speech processing, and structured machine messaging into a single staged session. During the voice phase, one endpoint places a call, the other answers, speech generated by a locally hosted LLaMA-family model is synthesized with Piper, transmitted through the modem voice path, captured on the remote side, and transcribed with Faster-Whisper to drive the next response. After the voice exchange, the system transitions to a …
Distilling The Complexity Of Agent-Based Simulations Into Textual Explanations Via Large Language Models, Noé Y. Flandre, Philippe J. Giabbanelli
Distilling The Complexity Of Agent-Based Simulations Into Textual Explanations Via Large Language Models, Noé Y. Flandre, Philippe J. Giabbanelli
VMASC Publications
Communicating the design and results of agent-based models (ABMs) to subject matter experts is challenging, which hinders participation and limits trust in simulation-based decision support. Large language models (LLMs) can communicate ABMs as textual summaries, thus complementing traditional disclosure through statistical and visualization techniques. While prior work translated the structure of conceptual models into narratives via LLMs, our extension covers the dynamics of simulation models via an automated simulation-to-text method that extracts contextual information from NetLogo ABMs, performs repeated simulations, and generates narrative descriptions (including the model’s purpose, parameters, and simulation dynamics) using mutimodal LLMs. Furthermore, four summarization algorithms spanning …
A Web-Based Wizard-Of-Oz Platform For Collaborative And Reproducible Human-Robot Interaction Research, Sean O'Connor
A Web-Based Wizard-Of-Oz Platform For Collaborative And Reproducible Human-Robot Interaction Research, Sean O'Connor
Honors Theses
The Wizard-of-Oz (WoZ) technique is widely used in Human-Robot Interaction (HRI) research, but two persistent problems limit its effectiveness: existing tools impose technical barriers that exclude non-engineering domain experts (the Accessibility Problem), and the fragmented landscape of robot-specific implementations makes interaction scripts difficult to port across platforms (the Reproducibility Problem- concerning execution consistency and portability, not third-party replication). Through a literature review, I identified three design principles to address both: a hierarchical specification model, an event-driven execution model, and a plugin architecture that decouples experiment logic from robot-specific implementations. I realized these principles in HRIStudio, an open-source, web-based platform providing …
An Empirical Framework For Evaluating Semantic Preservation Using Hugging Face, Nan Jia, Anita Raja, Raffi Khatchadourian
An Empirical Framework For Evaluating Semantic Preservation Using Hugging Face, Nan Jia, Anita Raja, Raffi Khatchadourian
Publications and Research
As machine learning (ML) becomes an integral part of high-autonomy systems, it is critical to ensure the trustworthiness of learning-enabled software systems (LESS). Yet, the nondeterministic and run-time-defined semantics of ML complicate traditional software refactoring. We define semantic preservation in LESS as the property that optimizations of intelligent components do not alter the system's overall functional behavior. This paper introduces an empirical framework to evaluate semantic preservation in LESS by mining model evolution data from HuggingFace. We extract commit histories, $\textit{Model Cards}$, and performance metrics from a large number of models. To establish baselines, we conducted case studies in three …
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha
Algorithmic Trading In Idiosyncratic-Payoff Markets: A Multi-Agent System For On-Chain Prediction Contracts, Saif Aldeen A.K. Agha
CMC Senior Theses
This thesis documents the design, deployment, and forward-test evaluation of an evolutionary multi-agent algorithmic trading system on Polymarket, the largest decentralized prediction market. The system pairs a locally-hosted 72-billion-parameter language model with a gradient-boosted statistical filter and an evolutionary selection mechanism that maintains a population of approximately 500 autonomous trading agents. Each agent generates a probability estimate for an event, compares it to the prevailing market price, and trades the resulting disagreement.
The central empirical exercise estimates a panel regression of trade-level profit on the absolute disagreement between the agent's probability estimate and the market price, controlling for agent identity, …
Error-Driven Density Control For Compact Gaussian Splatting Under Sparse Supervision, Abdelrhman Elrawy
Error-Driven Density Control For Compact Gaussian Splatting Under Sparse Supervision, Abdelrhman Elrawy
Theses and Dissertations (Comprehensive)
This thesis studies efficiency and stability challenges in Gaussian-splatting-based reconstruction under sparse supervision. In few-shot novel view synthesis, standard 3D Gaussian Splatting (3DGS) can overfit the limited training views and grow an unnecessarily large number of primitives due to limitations in its Adaptive Density Control (ADC) mechanism. This thesis introduces an error-driven reformulation of ADC that triggers densification using opacity gradients as a lightweight proxy for rendering error, and shows that such aggressive densification must be paired with delayed and conservative pruning to prevent destructive create--destroy cycles. When combined with depth-based geometric regularization, the resulting framework produces substantially more compact …
Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail
Geovig And Purevig: Geometry-Aware Architectures For Efficient Computer Vision, Omar Ismail
Theses and Dissertations (Comprehensive)
Deploying deep learning models for medical image analysis on mobile devices requires a balance between inference latency, memory footprint, and delineating anatomical boundaries with high accuracy. While Convolutional Neural Networks (CNNs) and mobile Vision Transformers (ViTs) offer efficiency, they often struggle to model the irregular, non-local geometric structures inherent in biological tissues without incurring prohibitive computational costs. In this thesis, we introduce GeoViG (Geometric Vision Graph), an architecture that bridges the gap between efficient grid-based processing and explicit Geometric Deep Learning. GeoViG introduces a novel transition from high-resolution pixel grids to low-resolution dynamic graphs via a SpreadEdgePool operator, a geometry-aware …
Llm-Driven Weekly Newsletter To Assess Open Source Software Project Github Health, Christian Novalski, Christopher Chavez, Ghalian Fayyadh, Kostadin Damevski
Llm-Driven Weekly Newsletter To Assess Open Source Software Project Github Health, Christian Novalski, Christopher Chavez, Ghalian Fayyadh, Kostadin Damevski
UROP Posters
Open Source Software (OSS) projects increasingly depend on a diverse set of contributors, including episodic participants who contribute intermittently. Episodic contributors represent a large portion of OSS communities, yet projects often struggle to retain them, leading to decreased project health and continuity. While dashboards and real-time communication tools support continuously active contributors, they often fail to serve the unique needs of episodic participants, who may struggle to remain informed and re-engage with project activity after periods of absence. In this study, we examine the effect of a weekly, email-based newsletter intervention designed to improve awareness and engagement among episodic OSS …
Benchmarking Gaslighting Negation Attacks Against Reasoning Models, Bin Zhu, Hailong Yin, Jingjing Chen, Yu Gang Jiang
Benchmarking Gaslighting Negation Attacks Against Reasoning Models, Bin Zhu, Hailong Yin, Jingjing Chen, Yu Gang Jiang
Research Collection School Of Computing and Information Systems
Recent advances in reasoning-centric models promise improved robustness through mechanisms such as chain-of-thought prompting and test-time scaling. However, their ability to withstand gaslighting negation attacks—adversarial prompts that confidently deny correct answers—remains underexplored. In this paper, we conduct a systematic evaluation of three state-of-the-art reasoning models, i.e., OpenAI’s o4-mini, Claude-3.7-Sonnet and Gemini-2.5-Flash, across three multimodal benchmarks: MMMU, MathVista, and CharXiv. Our evaluation reveals significant accuracy drops (25–29% on average) following gaslighting negation attacks, indicating that even top-tier reasoning models struggle to preserve correct answers under manipulative user feedback. Built upon the insights of the evaluation and to further probe this vulnerability, …
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Research Collection School Of Computing and Information Systems
The rapid integration of Large Language Models (LLMs) into software engineering (SE) has revolutionized tasks from code generation to program repair, producing a massive volume of software artifacts. This surge in automated creation has exposed a critical bottleneck: the lack of scalable and reliable methods to evaluate the quality of these outputs. Human evaluation, while effective, is very costly and time-consuming. Traditional automated metrics like BLEU rely on high-quality references and struggle to capture nuanced aspects of software quality, such as readability and usefulness. In response, the LLM-as-a-Judge paradigm, which employs LLMs for automated evaluation, has emerged. This approach leverages …
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Face Value: A Computational Approach To Subjective Impressions Of Faces, Kevin Kpankou
Undergraduate Research Symposium
Various computational models of first impressions have been developed to uncover the mechanisms driving these judgments. However, the implicit notion of a singular ``human'' often overlooks meaningful individual differences in beliefs, attitudes, and associations, as well as culturally grounded group-level constructs. In this paper, we extend Cultural Consensus Theory (CCT) to estimate culturally shared beliefs about faces by incorporating latent constructs structured around interpretable facial features extracted via computer vision algorithms. We apply our model to a large-scale dataset of people’s first impressions of faces. Our approach reveals a robust mapping between facial features and culturally constructed impressions, allowing us …
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
Developing An Ai-Assisted Grading System Using Large Language Models, Andrei Modiga
MS in Computer Science Project Reports
We present a grading system that accelerates evaluation of open-ended student work across scanned and digital workflows. The system crops answer regions from PDFs, assigns submissions via OCR on identity regions only, and groups answers by visual semantics using a vision LLM. Instructors review and edit groups, apply rubric items once per group, and export grades from an on-screen table. The solution integrates Ghostscript rasterization, PdfPig page orchestration, SkiaSharp region extraction, Tesseract identity OCR, and GPT-4o Vision for grouping. We detail the architecture, token-budgeted batching strategy, and persistence design, then describe testing results for grouping quality, time-on-task, and usability. The …
Challenges In Engineering Machine Learning (Software) Systems, Raffi T. Khatchadourian Ph.D.
Challenges In Engineering Machine Learning (Software) Systems, Raffi T. Khatchadourian Ph.D.
Open Educational Resources
Lecture slides on the software engineering challenges unique to machine learning systems, for an undergraduate software engineering course. After contrasting traditional programming with machine learning, the deck examines why the usual tools for managing complexity—abstraction, reuse, and composition—are harder to apply to ML, given the lack of clear specifications and modularity. It covers concept drift, feedback loops (illustrated with crime-prediction and recommendation examples), and the accumulation of technical debt in ML systems, including the role and pitfalls of notebooks in moving from experimentation to production. Based on "Machine Learning in Production/AI Engineering" by Christian Kaestner and Eunsuk Kang (Carnegie Mellon …
A Predictive Model For Multi- Week Respiratory Risk From Red Tide On Florida’S Gulf Coast., Elmer S. Ochaeta
A Predictive Model For Multi- Week Respiratory Risk From Red Tide On Florida’S Gulf Coast., Elmer S. Ochaeta
Computer Science and Engineering Faculty Publications
Florida’s Gulf Coast red tide (Karenia brevis) can put toxins into the air, making people cough, irritating the throat, and worsening asthma or other breathing problems especially when winds blow from the ocean toward the beach. Right now, most public updates don’t really help with the question people actually ask when planning a weekend or vacation: “Will going to or close to the beach be risky in the next few weeks?”.
In this project, I build a weekly early warning system that estimates respiratory risk for specific beaches and predicts that risk 2 to 4 weeks ahead. The study covers …