Open Access. Powered by Scholars. Published by Universities.®

Software Engineering Commons™

Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type

Articles 31 - 60 of 2211

Full-Text Articles in Software Engineering

Causality-Aware Safety Testing For Autonomous Driving Systems, Wenbing Tang, Mingfei Cheng, Renzhi Wang, Yuan Zhou, Chengwei Liu, Yang Liu, Zuohua Ding Apr 2026

Causality-Aware Safety Testing For Autonomous Driving Systems, Wenbing Tang, Mingfei Cheng, Renzhi Wang, Yuan Zhou, Chengwei Liu, Yang Liu, Zuohua Ding

Research Collection School Of Computing and Information Systems

Simulation-based testing is essential for evaluating the safety of Autonomous Driving Systems (ADSs). Comprehensive evaluation requires testing across diverse scenarios that can trigger various types of violations under different conditions. While existing methods typically focus on individual diversity metrics, such as input scenarios, ADS-generated motion commands, and system violations, they often fail to capture the complex interrelationships among these elements. For instance, identical motion commands can produce different collision risks in varying scenes, and the same collision may result from different commands under different scenarios. This oversight leads to gaps in testing coverage, potentially missing critical issues in the ADS …


Patchgpt: Multi-Agent Patch Backporting Without Model Fine-Tuning, Ye Liu, Ruidong Han, Chengyan Ma, Yuqing Niu, David Lo Apr 2026

Patchgpt: Multi-Agent Patch Backporting Without Model Fine-Tuning, Ye Liu, Ruidong Han, Chengyan Ma, Yuqing Niu, David Lo

Research Collection School Of Computing and Information Systems

Patch backporting is crucial and prevalent in the maintenance of modern open-source software such as Linux kernels and forked repositories. However, porting patches across program versions remains a challenging problem due to the complexity of synergizing diverse patches with divergent program versions. In this paper, we propose PatchGPT, an agentic patch backporting framework for fine-grained patch generation. PatchGPT encompasses three agents: Miner for decomposing a sequence of atomic change steps as the original patch plan, Adapter for adapting the patch plan, and Executor for executing the adapted patch plan according to predefined change semantics. We conduct experiments on the PPatHF’s …


Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo Apr 2026

Penforge: On-The-Fly Expert Agent Construction For Automated Penetration Testing, Huihui Huang, Jieke Shi, Junkai Chen, Ting Zhang, Yikun Li, Chengran Yang, Eng Lieh Ouh, Lwin Khin Shar, David Lo

Research Collection School Of Computing and Information Systems

Penetration testing is essential for identifying vulnerabilities in web applications before real adversaries can exploit them. Recent work has explored automating this process with Large Language Model (LLM)-powered agents, but existing approaches either rely on a single generic agent that struggles in complex scenarios or narrowly specialized agents that cannot adapt to diverse vulnerability types. We therefore introduce PenForge, a framework that dynamically constructs expert agents during testing rather than relying on those prepared beforehand. By integrating automated reconnaissance of potential attack surfaces with agents instantiated on the fly for context-aware exploitation, PenForge achieves a 30.0% exploit success rate (12/40) …


Finding Missing Input Validation In Tees Via Llm-Assisted Symbolic Execution, Chengyan Ma, Jieke Shi, Ruidong Han, Ye Liu, Yuqing Niu, David Lo Apr 2026

Finding Missing Input Validation In Tees Via Llm-Assisted Symbolic Execution, Chengyan Ma, Jieke Shi, Ruidong Han, Ye Liu, Yuqing Niu, David Lo

Research Collection School Of Computing and Information Systems

Trusted Execution Environments (TEEs) provide hardware-enforced isolation that protects sensitive code and data from untrusted software. Despite their strong security guarantees, analyzing TEE applications remains challenging due to the high cost and complexity of configuring complete TEE build and runtime environments, as well as the limited observability imposed by hardware isolation. This paper presents SymTEE, a novel large language model (LLM)-assisted symbolic execution framework for detecting missing input validation issues in TEE applications without requiring real TEE setups. SymTEE begins by leveraging Abstract Syntax Tree (AST) analysis to extract TEE code slices that may lack sufficient input validation, and then …


Agentspec: Customizable Runtime Enforcement For Safe And Reliable Llm Agents, Haoyu Wang, Christopher M. Poskitt, Jun Sun Apr 2026

Agentspec: Customizable Runtime Enforcement For Safe And Reliable Llm Agents, Haoyu Wang, Christopher M. Poskitt, Jun Sun

Research Collection School Of Computing and Information Systems

Agents built on LLMs are increasingly deployed across diverse domains, automating complex decision-making and task execution. However, their autonomy introduces safety risks, including security vulnerabilities, legal violations, and unintended harmful actions. Existing mitigation methods, such as model-based safeguards and early enforcement strategies, fall short in robustness, interpretability, and adaptability. To address these challenges, we propose AgentSpec, a lightweight domain-specific language for specifying and enforcing runtime constraints on LLM agents. With AgentSpec, users define structured rules that incorporate triggers, predicates, and enforcement mechanisms, ensuring agents operate within predefined safety boundaries. We implement AgentSpec across multiple domains, including code execution, embodied agents, …


Conflogger: Enhance Systems’ Configuration Diagnosability Through Configuration Logging, Shiwen Shan, Yintong Huo, Yuxin Su, Zhining Wang, Dan Li, Zibin Zheng Apr 2026

Conflogger: Enhance Systems’ Configuration Diagnosability Through Configuration Logging, Shiwen Shan, Yintong Huo, Yuxin Su, Zhining Wang, Dan Li, Zibin Zheng

Research Collection School Of Computing and Information Systems

Modern configurable systems offer customization via intricate configuration spaces, yet such flexibility introduces pervasive configuration-related issues such as misconfigurations and latent softwarebugs. Existing diagnosability supports focus on post-failure analysis of software behavior to identify configuration issues, but none of these approaches look into whether the software clue sufficient failure information for diagnosis. To fill in the blank, we propose the idea of configuration logging to enhance existing logging practices at the source code level. We develop ConfLogger, the first tool that unifies configuration-aware static taint analysis with LLM-based log generation to enhance software configuration diagnosability. Specifically, our method 1) identifies …


Managing Reproducibility Debt In Scientific Software: A Practical Framework, Zara Hassan, Christoph Treude, Graham Williams, Michael Norrish, Alex Potanin Apr 2026

Managing Reproducibility Debt In Scientific Software: A Practical Framework, Zara Hassan, Christoph Treude, Graham Williams, Michael Norrish, Alex Potanin

Research Collection School Of Computing and Information Systems

Scientific software includes end-user applications, modelling tools, research software for publications, and production systems for real users. It plays a key role across various scientific disciplines by enabling large-scale computation, simulation, and data analysis. Unlike commercial software, scientific software is often developed in dynamic research environments with limited engineering practices, documentation, or testing. This makes it fragile and difficult to reproduce results, even when code and data are available, conditions in which Reproducibility Debt (RpD) accumulates. This paper presents the Reproducibility Debt Management Framework (RpD-MF), which is grounded in evidence from a systematic literature review, practitioner interviews, and a global …


Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang Apr 2026

Prompting Frameworks For Large Language Models: A Survey, Xiaoxia Liu, Jingyi Wang, Jun Sun, Xiaohan Yuan, Guoliang Dong, Peng Di, Wenhai Wang, Dongxia Wang

Research Collection School Of Computing and Information Systems

Since the launch of ChatGPT, a powerful AI Chatbot developed by OpenAI, large language models (LLMs) have made significant advancements in both academia and industry, bringing about a fundamental engineering paradigm shift in many areas. While LLMs are powerful, it is also crucial to best use their power where “prompt” plays a core role. However, the booming LLMs themselves, including excellent APIs like ChatGPT, have several inherent limitations: (1) temporal lag of training data, and (2) the lack of physical capabilities to perform external actions. Recently, we have observed the trend of utilizing prompt-based tools to better utilize the power …


Bridging Bug Localization And Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models, Jianming Chang, Xin Zhou, Lulu Wang, David Lo, Bixin Li Apr 2026

Bridging Bug Localization And Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models, Jianming Chang, Xin Zhou, Lulu Wang, David Lo, Bixin Li

Research Collection School Of Computing and Information Systems

Automated issue fixing is a critical task in software debugging and has recently garnered significant attention from academia and industry. However, existing fixing techniques predominantly focus on the repair phase, often overlooking the importance of improving the preceding bug localization phase. As a foundational step in issue fixing, bug localization plays a pivotal role in determining the overall effectiveness of the entire process. To enhance the precision of issue fixing by accurately identifying bug locations in large-scale projects, this paper presents BugCerberus, the first hierarchical bug localization framework powered by three customized large language models. First, BugCerberus analyzes intermediate representations …


Visual Loop: Bridging The Cognitive Gap In Software Development Through Visual-Ai Collaboration, Luis Filipe Fernandes Gomes, Xin Zhou, David Lo, Rui Abreu Apr 2026

Visual Loop: Bridging The Cognitive Gap In Software Development Through Visual-Ai Collaboration, Luis Filipe Fernandes Gomes, Xin Zhou, David Lo, Rui Abreu

Research Collection School Of Computing and Information Systems

Software development remains predominantly text-centric, despite decades of evidence showing that developers think and communicate visually. While sketches and diagrams externalize developers’ mental models, they remain disconnected from source code and quickly become outdated. Recent advances in foundation models, capable of both code and visual reasoning, create an opportunity to unify these representations. In this vision paper, we introduce Visual Loop, a continuous visual development environment that keeps code and informal sketches in bidirectional synchronization. Our prototype connects a code editor with a tablet-based visualization workspace, allowing developers to explore, annotate, and modify systems through freehand sketches interpreted by multimodal …


Understanding Codebase Like A Professional! Human-Ai Collaboration For Code Comprehension, Jie Gao, Yue Xue, Xiaofei Xie, Junming Cao, Soemin Thant, Erika Lee, Bowen Xu Apr 2026

Understanding Codebase Like A Professional! Human-Ai Collaboration For Code Comprehension, Jie Gao, Yue Xue, Xiaofei Xie, Junming Cao, Soemin Thant, Erika Lee, Bowen Xu

Research Collection School Of Computing and Information Systems

Understanding an unfamiliar codebase is an essential task for developers in various scenarios, such as during the onboarding process. Especially when the codebase is large and time is limited, achieving a decent level of comprehension remains challenging for both experienced and novice developers, even with the assistance of large language models (LLMs). Existing studies have shown that LLMs often fail to support users in understanding code structures or to provide user-centered, adaptive, and dynamic assistance in real-world settings.To address this, we propose learning from the perspective of a unique role, code auditors, whose work often requires them to quickly familiarize …


Autologger: A Multi-Agent Framework For The End-To-End Automated Logging, Renyi Zhong, Yintong Huo, Wenwei Gu, Yichen Li, Michael R. Lyu Apr 2026

Autologger: A Multi-Agent Framework For The End-To-End Automated Logging, Renyi Zhong, Yintong Huo, Wenwei Gu, Yichen Li, Michael R. Lyu

Research Collection School Of Computing and Information Systems

Software logging is critical for system observability, yet developers face a dual crisis of costly overlogging and risky underlogging. Existing automated logging tools often overlook the fundamental whether-to-log decision and struggle with the composite nature of logging. In this paper, we propose AutoLogger, a novel hybrid framework that addresses the complete the end-to-end logging pipeline. AutoLogger first employs a fine-tuned classifier, the Judger, to accurately determine if a method requires new logging statements. If logging is needed, a multi-agent system is activated. The system includes specialized agents: a Locator dedicated to determining where to log, and a Generator focused on …


Context Engineering For Ai Agents In Open-Source Software, Seyedmoein Mohsenimofidi, Matthias Galster, Christoph Treude, Sebastian Baltes Apr 2026

Context Engineering For Ai Agents In Open-Source Software, Seyedmoein Mohsenimofidi, Matthias Galster, Christoph Treude, Sebastian Baltes

Research Collection School Of Computing and Information Systems

GenAI-based coding assistants have disrupted software development. The next generation of these tools is agent-based, operating with more autonomy and potentially without human oversight. Like human developers, AI agents require contextual information to develop solutions that are in line with the standards, policies, and workflows of the software projects they operate in. Vendors of popular agentic tools (e.g., Claude Code) recommend maintaining version-controlled Markdown files that describe aspects such as the project structure, code style, or building and testing. The content of these files is then automatically added to each prompt. Recently, AGENTS.md has emerged as a potential standard that …


On Autopilot? An Empirical Study Of Human-Ai Teaming And Review Practices In Open Source, Haoyu Gao, Peerachai Banyongrakkul, Hao Guan, Mansooreh Zahedi, Christoph Treude Apr 2026

On Autopilot? An Empirical Study Of Human-Ai Teaming And Review Practices In Open Source, Haoyu Gao, Peerachai Banyongrakkul, Hao Guan, Mansooreh Zahedi, Christoph Treude

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) increasingly automate software engineering tasks. While recent studies highlight the accelerated adoption of “AI as a teammate” in Open Source Software (OSS), developer interaction patterns remain under-explored. In this work, we investigated project-level guidelines and developers’ interactions with AI-assisted pull requests (PRs) by expanding the AIDev dataset to include finer-grained contributor code ownership and a comparative baseline of human-created PRs. We found that over 67.5% of AI-co-authored PRs originate from contributors without prior code ownership. Despite this, the majority of repositories lack guidelines for AI-coding agent usage. Notably, we observed a distinct interaction pattern: AI-co-authored PRs …


Who Said Cve? How Vulnerability Identifiers Are Mentioned By Humans, Bots, And Agents In Pull Requests, Pien Rooijendijk, Christoph Treude, Mairieli Wessel Apr 2026

Who Said Cve? How Vulnerability Identifiers Are Mentioned By Humans, Bots, And Agents In Pull Requests, Pien Rooijendijk, Christoph Treude, Mairieli Wessel

Research Collection School Of Computing and Information Systems

Vulnerability identifiers such as CVE, CWE, and GHSA are standardised references to known software security issues, yet their use in practice is not well understood. This paper compares vulnerability ID use in GitHub pull requests authored by autonomous agents, bots, and human developers. Using the AIDev pop dataset and an augmented set of pull requests from the same repositories, we analyse who mentions vulnerability identifiers and where they appear. Bots account for around 69.1% of all mentions, usually adding few identifiers in pull request descriptions, while human and agent mentions are rarer but span more locations. Qualitative analysis shows that …


Optimizing And Fortifying Ai Software Through The Lens Of Artifact Synthesis, Jieke Shi Mar 2026

Optimizing And Fortifying Ai Software Through The Lens Of Artifact Synthesis, Jieke Shi

Dissertations and Theses Collection (Open Access)

Artificial Intelligence (AI) has transformed the software landscape, ushering in a new era of intelligent systems that increasingly shape our daily lives. This transformation is evident in various domains, including Software Engineering (SE), where Large Language Models (LLMs) support many development tools, and control systems, where self-driving cars and autonomous drones rely on deep learning models for real-time decision-making. These AI systems are collectively referred to as AI software, with the former categorized as AI4SE software (AI for Software Engineering) and the latter as AI4Control software (AI for Control). As AI software becomes central to modern computing infrastructure, its reliability …


Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui Mar 2026

Codeultrafeedback: An Llm-As-A-Judge Dataset For Aligning Large Language Models To Coding Preferences, Martin Weyssow, Aton Kamanda, Xin Zhou, Houari Sahraoui

Research Collection School Of Computing and Information Systems

Evaluating the alignment of large language models (LLMs) with user-defined coding preferences is a challenging endeavor that requires a deep assessment of LLMs' outputs. Existing methods and benchmarks rely primarily on automated metrics and static analysis tools, which often fail to capture the nuances of user instructions and LLM outputs. To address this gap, we introduce the LLM-as-a-Judge evaluation framework and present CodeUltraFeedback, a comprehensive dataset for assessing and improving LLM alignment with coding preferences. CodeUltraFeedback consists of 10,000 coding instructions, each annotated with four responses generated from a diverse pool of 14 LLMs. These responses are annotated using GPT-3.5 …


Invert Your Prompt: Editing-Aware Diffusion Inversion, Yangyang Xu, Wenqi Shao, Yong Du, Haiming Zhu, Yang Zhou, Jiayuan Xie, Ping Luo, Shengfeng He Mar 2026

Invert Your Prompt: Editing-Aware Diffusion Inversion, Yangyang Xu, Wenqi Shao, Yong Du, Haiming Zhu, Yang Zhou, Jiayuan Xie, Ping Luo, Shengfeng He

Research Collection School Of Computing and Information Systems

Recent advancements in text-guided diffusion models have enabled powerful image manipulation capabilities. However, balancing reconstruction fidelity and editability for real images remains a significant challenge. In this work, we introduce Editing Inversion (EditInv), a novel framework that inverts and edits real images for specific editing tasks by optimizing specific prompt embeddings within the extended  space. By leveraging distinct embeddings across different U-Net layers and time steps, EditInv seamlessly integrates inversion and editing through reciprocal optimization, ensuring both high fidelity and precise editability. This hierarchical editing mechanism classifies tasks into structure, appearance, and global edits, optimizing only those embeddings that are …


Identifying And Mitigating Api Misuse In Large Language Models, Terry Yue Zhuo, Junda He, Jiamou Sun, Zhenchang Xing, David Lo, John Grundy, Xiaoning Du Mar 2026

Identifying And Mitigating Api Misuse In Large Language Models, Terry Yue Zhuo, Junda He, Jiamou Sun, Zhenchang Xing, David Lo, John Grundy, Xiaoning Du

Research Collection School Of Computing and Information Systems

API misuse in code generated by large language models (LLMs) presents a serious and growing challenge in software development. While LLMs demonstrate impressive code generation capabilities, their interactions with complex library APIs are often error-prone, potentially leading to software failures and vulnerabilities. In this paper, we conduct a large-scale study of API misuse patterns in LLM-generated code, analyzing both method selection and parameter usage across Python and Java, using three representative LLMs (StarCoder-7B, Qwen2.5-Coder-7B, and GitHub Copilot). Based on extensive manual annotation of 3,209 method-level and 3,492 parameter-level misuses, we identify and categorize four recurring misuse types by building on …


Exploring Neural Network Structure Code Reuse In The Open-Source Community For Improving Maintenance, Xiaoning Ren, Yuekun Wang, Chongyang Liu, Yueming Wu, Qiang Hu, Lijun Zhang, Yinxing Xue Mar 2026

Exploring Neural Network Structure Code Reuse In The Open-Source Community For Improving Maintenance, Xiaoning Ren, Yuekun Wang, Chongyang Liu, Yueming Wu, Qiang Hu, Lijun Zhang, Yinxing Xue

Research Collection School Of Computing and Information Systems

Neural networks (NNs) have rapidly advanced, demonstrating exceptional performance across various fields, leading to a surge in open-source NN projects. The complexity and rapid growth of these projects pose significant challenges for maintenance within the open-source community. Given that NN architecture code is the core asset of NN projects, understanding its reuse in the open-source community is essential for effective maintenance, such as reducing redundancy and identifying potential intellectual property violations. While prior studies have examined code reuse in open-source projects, they have two key limitations: They do not specifically address NN structure code, and they rely on manually selected …


Less Is More: Docstring Compression In Code Generation, Guang Yang, Yu Zhou, Wei Cheng, Xiangyu Zhang, Xiang Chen, Terry Yue Zhuo, Xin Zhou, Ke Liu, David Lo, Taolue Chen Feb 2026

Less Is More: Docstring Compression In Code Generation, Guang Yang, Yu Zhou, Wei Cheng, Xiangyu Zhang, Xiang Chen, Terry Yue Zhuo, Xin Zhou, Ke Liu, David Lo, Taolue Chen

Research Collection School Of Computing and Information Systems

The widespread use of Large Language Models (LLMs) in software engineering has intensified the need for improved model and resource efficiency. In particular, for neural code generation, LLMs are used to translate function/method signature and DocString to executable code. DocStrings, which capture user requirements for the code and are typically used as the prompt for LLMs, often contain redundant information. Recent advancements in prompt compression have shown promising results in Natural Language Processing (NLP), but their applicability to code generation remains uncertain. Our empirical study shows that the state-ofthe-art prompt compression methods achieve only about 10% reduction, as further reductions …


Defending Code Language Models Against Backdoor Attacks With Deceptive Cross-Entropy Loss, Guang Yang, Yu Zhou, Xiangyu Zhang, Xiang Chen, Terry Yue Zhuo, David Lo, Taolue Chen Feb 2026

Defending Code Language Models Against Backdoor Attacks With Deceptive Cross-Entropy Loss, Guang Yang, Yu Zhou, Xiangyu Zhang, Xiang Chen, Terry Yue Zhuo, David Lo, Taolue Chen

Research Collection School Of Computing and Information Systems

Code Language Models (CLMs), particularly those leveraging deep learning, have achieved significant success in code intelligence domain. However, the issue of security, particularly backdoor attacks, is often overlooked in this process. The previous research has focused on designing backdoor attacks for CLMs, but effective defenses have not been adequately addressed. In particular, existing defense methods from natural language processing, when directly applied to CLMs, are not effective enough and lack generality, working well in some models and scenarios but failing in others, thus fall short in consistently mitigating backdoor attacks. To bridge this gap, we first confirm the phenomenon of …


Fcghunter: Towards Evaluating Robustness Of Graph-Based Android Malware Detection, Shiwen Song, Xiaofei Xie, Ruitao Feng, Qi Guo, Sen Chen Feb 2026

Fcghunter: Towards Evaluating Robustness Of Graph-Based Android Malware Detection, Shiwen Song, Xiaofei Xie, Ruitao Feng, Qi Guo, Sen Chen

Research Collection School Of Computing and Information Systems

Graph-based detection methods leveraging Function Call Graph (FCG) have shown promise for Android malware detection (AMD) due to their semantic insights. However, the deployment of malware detectors in dynamic and hostile environments raises significant concerns about their robustness. While recent approaches evaluate the robustness of FCG-based detectors using adversarial attacks, their effectiveness is constrained by the vast perturbation space, particularly across diverse models and features. To address these challenges, we introduce FCGHunter, a novel robustness testing framework for FCG-based AMD systems. Specifically, FCGHunter employs innovative techniques to enhance exploration and exploitation within this huge search space. Initially, it identifies critical …


Fortifying The Seams Between C/C++ And Rust: Characterizing Bugs In Interop Tools, Xuemeng Cai, Jiakun Liu, Cunyang Liu, Lingfeng Bao, Yijun Yu, Lingxiao Jiang Feb 2026

Fortifying The Seams Between C/C++ And Rust: Characterizing Bugs In Interop Tools, Xuemeng Cai, Jiakun Liu, Cunyang Liu, Lingfeng Bao, Yijun Yu, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

Rust has become increasingly popular in recent years due to its safety and high performance. Despite these advantages, Rust projects rarely start from scratch in practice, and many Rust-based systems instead use hybrid programming, where Rust interoperates with existing C/C++ code. To reduce the manual effort involved in this interoperation (interop) process, several interop tools have been proposed to facilitate hybrid programming between Rust and C/C++. However, the challenges and limitations of these tools remain largely unexplored, leaving developers unclear about the future directions and users unclear about the appropriate usage scenarios. To fill the gap, we mined 320 bugs …


Vercation: Precise Vulnerable Open-Source Software Version Identification Based On Static Analysis And Llm, Yiran Cheng, Ting Zhang, Lwin Khin Shar, Shouguo Yang, Chaopeng Dong, David Lo, Shichao Lv, Zhiqiang Shi, Limin Sun Feb 2026

Vercation: Precise Vulnerable Open-Source Software Version Identification Based On Static Analysis And Llm, Yiran Cheng, Ting Zhang, Lwin Khin Shar, Shouguo Yang, Chaopeng Dong, David Lo, Shichao Lv, Zhiqiang Shi, Limin Sun

Research Collection School Of Computing and Information Systems

Open-source software (OSS) has experienced a surge in popularity, attributed to its collaborative development model and cost-effective nature. However, the adoption of specific software versions in development projects may introduce security risks when these versions bring along vulnerabilities. Current methods of identifying vulnerable versions typically analyze and extract the code features involved in vulnerability patches using static analysis with pre-defined rules. They then use code clone detection to identify the vulnerable versions. These methods are hindered by imprecision due to (1) the exclusion of vulnerability- irrelevant code in the analysis and (2) the inadequacy of code clone detection. This paper …


The Feelit System: Application Content-Aware Perspectives And Challenges On Understanding User Likes In Social Network Posts, Konstantinos Theocharidis, Hady W. Lauw, Panagiotis Karras Feb 2026

The Feelit System: Application Content-Aware Perspectives And Challenges On Understanding User Likes In Social Network Posts, Konstantinos Theocharidis, Hady W. Lauw, Panagiotis Karras

Research Collection School Of Computing and Information Systems

In a series of our prior works, we study influence and subscription maximization problems in social networks that are based on posts having influential content; as content we consider a set of features where each feature corresponds to a specific social network page, whereas influence and subscription relate to gaining the postlike and subscription-to-brand page of targeted users, respectively; subscription is conceptually achieved as repetitive influence on users. So, both influence and subscription depend on content that gains the likes of users; however, to be realistic, modeling and estimating such likes is a complex problem that has not been adequately …


Efficient Function Orchestration For Large Language Models, Xiaoxia Liu, Peng Di, Cong Li, Jun Sun, Jingyi Wang Feb 2026

Efficient Function Orchestration For Large Language Models, Xiaoxia Liu, Peng Di, Cong Li, Jun Sun, Jingyi Wang

Research Collection School Of Computing and Information Systems

Function calling is a fundamental capability of today's large language models, but sequential function calling posed efficiency problems. Recent studies have proposed to request function calls with parallelism support in order to alleviate this issue. However, they either delegate the concurrent function calls to users for execution which are conversely executed sequentially, or overlook the relations among various function calls, rending limited efficiency. This paper introduces LLMOrch, an advanced framework for automated, parallel function calling in large language models. The key principle behind LLMOrch is to identify an available processor to execute a function call while preventing any single processor …


Exploring Jvm Garbage Collector Testing With Event-Coverage, Kai Zheng, Yingquan Zhao, Junjie Chen, Hanmo You, Haoyu Wang, Haoyu Wang, Tianchang Gao Feb 2026

Exploring Jvm Garbage Collector Testing With Event-Coverage, Kai Zheng, Yingquan Zhao, Junjie Chen, Hanmo You, Haoyu Wang, Haoyu Wang, Tianchang Gao

Research Collection School Of Computing and Information Systems

Garbage Collection (GC) in the Java Virtual Machine (JVM) serves as an automatic memory management mechanism, efficiently reclaiming unused memory space in different production scenarios. To optimize JVM performance, developers typically fine-tune the garbage collector by identifying an optimal set of GC configurations for specific scenarios. Despite the sophisticated design of garbage collectors, they still have the potential for bugs in different settings, and these bugs can result in more severe consequences. Hence, comprehensive testing of these garbage collectors is imperative before their release. Code coverage criteria are typically employed to assess the comprehensiveness of a test suite. However, traditional …


Zero-Shot Video Translation Via Token Warping, Haiming Zhu, Yangyang Xu, Jun Yu, Shengfeng He Feb 2026

Zero-Shot Video Translation Via Token Warping, Haiming Zhu, Yangyang Xu, Jun Yu, Shengfeng He

Research Collection School Of Computing and Information Systems

With the revolution of generative AI, video-related tasks have been widely studied. However, current state-of-the-art video models still lag behind image models in visual quality and user control over generated content. In this paper, we introduce TokenWarping, a novel framework for temporally coherent video translation. Existing diffusion-based video editing approaches rely solely on key and value patches in self-attention to ensure temporal consistency, often sacrificing the preservation of local and structural regions. Critically, these methods overlook the significance of the query patches in achieving accurate feature aggregation and temporal coherence. In contrast, TokenWarping leverages complementary token priors by constructing temporal …


Do Comments And Expertise Still Matter? An Experiment On Programmers’ Adoption Of Ai-Generated Javascript Code, Changwen Li, Christoph Treude, Ofir Turel Jan 2026

Do Comments And Expertise Still Matter? An Experiment On Programmers’ Adoption Of Ai-Generated Javascript Code, Changwen Li, Christoph Treude, Ofir Turel

Research Collection School Of Computing and Information Systems

This paper investigates the factors influencing programmers’ adoption of AI-generated JavaScript code recommendations within the context of lightweight, function-level programming tasks. It extends prior research by (1) utilizing objective (as opposed to the typically self-reported) measurements for programmers’ adoption of AI-generated code and (2) examining whether AI-generated comments added to code recommendations and development expertise drive AI-generated code adoption. We tested these potential drivers in an online experiment with 173 programmers. Participants were asked to answer some questions to demonstrate their level of development expertise. Then, they were asked to solve a LeetCode problem without AI support. After attempting to …