Open Access. Powered by Scholars. Published by Universities.®

Software Engineering Commons™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year

Articles 121 - 150 of 2149

Full-Text Articles in Software Engineering

Unambiguous Granularity Distillation For Asymmetric Image Retrieval, Hongrui Zhang, Yi Xie, Haoquan Zhang, Cheng Xu, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng Ann Heng, Shengfeng He Jul 2025

Unambiguous Granularity Distillation For Asymmetric Image Retrieval, Hongrui Zhang, Yi Xie, Haoquan Zhang, Cheng Xu, Xuandi Luo, Donglei Chen, Xuemiao Xu, Huaidong Zhang, Pheng Ann Heng, Shengfeng He

Research Collection School Of Computing and Information Systems

Previous asymmetric image retrieval methods based on knowledge distillation have primarily focused on aligning the global features of two networks to transfer global semantic information from the gallery network to the query network. However, these methods often fail to effectively transfer local semantic information, limiting the fine-grained alignment of feature representation spaces between the two networks. To overcome this limitation, we propose a novel approach called Layered-Granularity Localized Distillation (GranDist). GranDist constructs layered feature representations that balance the richness of contextual information with the granularity of local features. As we progress through the layers, the contextual information becomes more detailed, …


Llm-Based Multi-Agent Systems For Software Engineering: Literature Review, Vision And The Road Ahead, Junda He, Christoph Treude, David Lo Jul 2025

Llm-Based Multi-Agent Systems For Software Engineering: Literature Review, Vision And The Road Ahead, Junda He, Christoph Treude, David Lo

Research Collection School Of Computing and Information Systems

Integrating Large Language Models (LLMs) into autonomous agents marks a significant shift in the research landscape by offering cognitive abilities that are competitive with human planning and reasoning. This paper explores the transformative potential of integrating Large Language Models into Multi-Agent (LMA) systems for addressing complex challenges in software engineering (SE). By leveraging the collaborative and specialized abilities of multiple agents, LMA systems enable autonomous problem-solving, improve robustness, and provide scalable solutions for managing the complexity of real-world software projects. In this paper, we conduct a systematic review of recent primary studies to map the current landscape of LMA applications …


Runtime Anomaly Detection For Drones: An Integrated Rule-Mining And Unsupervised Learning Approach, Ivan Wei Han Tan, Wei Minn, Christopher M. Poskitt, Lwin Khin Shar, Lingxiao Jiang Jul 2025

Runtime Anomaly Detection For Drones: An Integrated Rule-Mining And Unsupervised Learning Approach, Ivan Wei Han Tan, Wei Minn, Christopher M. Poskitt, Lwin Khin Shar, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

Unmanned Aerial Vehicles (UAVs), commonly referred to as drones, have witnessed a remarkable surge in popularity due to their versatile applications. These cyber-physical systems depend on multiple sensor inputs, such as cameras, GPS receivers, accelerometers, and gyroscopes, with faults potentially leading to physical instability and serious safety concerns. To mitigate such risks, anomaly detection has emerged as a crucial safeguarding mechanism, capable of identifying the physical manifestations of emerging issues and allowing operators to take preemptive action at runtime. Recent anomaly detection methods based on LSTM neural networks have shown promising results, but three challenges persist: the need for models …


Explaining Explanations: An Empirical Study Of Explanations In Code Reviews, Ratnadira Widyasari, Ting Zhang, Abir Bouraffa, Walid Maalej, David Lo Jul 2025

Explaining Explanations: An Empirical Study Of Explanations In Code Reviews, Ratnadira Widyasari, Ting Zhang, Abir Bouraffa, Walid Maalej, David Lo

Research Collection School Of Computing and Information Systems

Code reviews are central for software quality assurance. Ideally, reviewers should explain their feedback to enable authors of code changes to understand the feedback and act accordingly. Different developers might need different explanations in different contexts. Therefore, assisting this process first requires understanding the types of explanations reviewers usually provide. The goal of this article is to study the types of explanations used in code reviews and explore the potential of Large Language Models (LLMs), specifically ChatGPT, in generating these specific types. We extracted 793 code review comments from Gerrit and manually labeled them based on whether they contained a …


How Are We Detecting Inconsistent Method Names? An Empirical Study From Code Review Perspective, Kisub Kim, Xin Zhou, Dongsun Kim, Julia Lawall, Kui Liu, Tegawendé F. Bissyandé, Jacques Klein, Jaekwon Lee, David Lo Jul 2025

How Are We Detecting Inconsistent Method Names? An Empirical Study From Code Review Perspective, Kisub Kim, Xin Zhou, Dongsun Kim, Julia Lawall, Kui Liu, Tegawendé F. Bissyandé, Jacques Klein, Jaekwon Lee, David Lo

Research Collection School Of Computing and Information Systems

Proper naming of methods can make program code easier to understand, and thus enhance software maintainability. Yet, developers may use inconsistent names due to poor communication or a lack of familiarity with conventions within the software development lifecycle. To address this issue, much research effort has been invested into building automatic tools that can check for method name inconsistency and recommend consistent names. However, existing datasets generally do not provide precise details about why a method name was deemed improper and required to be changed. Such information can give useful hints on how to improve the recommendation of adequate method …


Enhancing Project-Specific Code Completion By Inferring Internal Api Information, Le Deng, Xiaoxia Ren, Chao Ni, Ming Liang, David Lo, Zhongxin Liu Jul 2025

Enhancing Project-Specific Code Completion By Inferring Internal Api Information, Le Deng, Xiaoxia Ren, Chao Ni, Ming Liang, David Lo, Zhongxin Liu

Research Collection School Of Computing and Information Systems

Project-specific code completion, which aims to complete code based on the context of the project, is an important and practical software engineering task. The state-of-the-art approaches employ the retrieval-augmented generation (RAG) paradigm and prompt large language models (LLMs) with information retrieved from the target project for project-specific code completion. In practice, developers always define and use custom functionalities, namely internal APIs, to facilitate the implementation of specific project requirements. Thus, it is essential to consider internal API information for accurate project-specific code completion. However, existing approaches either retrieve similar code snippets, which do not necessarily contain related internal API information, …


On-Demand Scenario Generation For Testing Automated Driving Systems, Songyang Yan, Xiaodong Zhang, Kunkun Hao, Haojie Xin, Yonggang Luo, Jucheng Yang, Ming Fan, Chao Yang, Jun Sun, Zijiang Yang Jun 2025

On-Demand Scenario Generation For Testing Automated Driving Systems, Songyang Yan, Xiaodong Zhang, Kunkun Hao, Haojie Xin, Yonggang Luo, Jucheng Yang, Ming Fan, Chao Yang, Jun Sun, Zijiang Yang

Research Collection School Of Computing and Information Systems

The safety and reliability of Automated Driving Systems (ADS) are paramount, necessitating rigorous testing methodologies to uncover potential failures before deployment. Traditional testing approaches often prioritize either natural scenario sampling or safety-critical scenario generation, resulting in overly simplistic or unrealistic hazardous tests. In practice, the demand for natural scenarios (e.g., when evaluating the ADS's reliability in real-world conditions), critical scenarios (e.g., when evaluating safety in critical situations), or somewhere in between (e.g., when testing the ADS in regions with less civilized drivers) varies depending on the testing objectives. To address this issue, we propose the On-demand Scenario Generation (OSG) Framework, …


Irhunter: Universal Detection Of Instruction Reordering Vulnerabilities For Enhanced Concurrency In Distributed And Parallel Systems, Guohua Xin, Guangquan Xu, Yao Zhang, Cheng Wen, Cen Zhang, Xiaofei Xie, Neal N. Xiong, Shaoying Liu, Pan Gao Jun 2025

Irhunter: Universal Detection Of Instruction Reordering Vulnerabilities For Enhanced Concurrency In Distributed And Parallel Systems, Guohua Xin, Guangquan Xu, Yao Zhang, Cheng Wen, Cen Zhang, Xiaofei Xie, Neal N. Xiong, Shaoying Liu, Pan Gao

Research Collection School Of Computing and Information Systems

Instruction reordering is an essential optimization technique used in both compilers and multi-core processors to enhance parallelism and resource utilization. Although the original intent of this technique is to benefit the program, some improper reordering can significantly impact the program correctness, which we call instruction reordering vulnerability (IRV). However, existing methods detect IRV by defining CPU instruction reordering rules to schedule execution paths while neglecting compiler reordering, and thus generate false positives that require manual filtering and resulting in inefficiency. To bridge this gap, in this paper, we propose the IRV detection method, , which analyzes IRV characteristics and extracts …


Contested: Consistency-Aided Tested Code Generation With Llm, Jinhao Dong, Jun Sun, Wenjie Zhang, Jinsong Dong, Dan Hao Jun 2025

Contested: Consistency-Aided Tested Code Generation With Llm, Jinhao Dong, Jun Sun, Wenjie Zhang, Jinsong Dong, Dan Hao

Research Collection School Of Computing and Information Systems

Recent advancements in large language models (LLMs) have significantly improved code generation, which generates code snippets automatically based on natural language requirements. Despite achieving state-of-the-art performance, LLMs often struggle to generate accurate and reliable code, requiring developers to spend substantial effort debugging and evaluating the generated output. Researchers have proposed leveraging Consistency to select code that passes more tests (inter-consistency) and demonstrates consistent behavior across more counterparts (intra-consistency). However, since the tests themselves are also generated by LLMs, relying on majority voting based on incorrect tests leads to unreliable results. To address this, we propose a lightweight interaction framework that …


Enhancing Vulnerability Detection Via Inter-Procedural Semantic Completion, Bozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao, Jun Sun, Shang-Wei Lin Jun 2025

Enhancing Vulnerability Detection Via Inter-Procedural Semantic Completion, Bozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao, Jun Sun, Shang-Wei Lin

Research Collection School Of Computing and Information Systems

Inspired by advances in deep learning, numerous learning-based approaches for vulnerability detection have emerged, primarily operating at the function level for scalability. However, this design choice has a critical limitation: many vulnerabilities span multiple functions, causing function-level approaches to lose the semantics of called functions and fail to capture true vulnerability patterns. To address this issue, we propose VulnSC, a novel framework designed to enhance learning-based approaches by complementing inter-procedural semantics. VulnSC retrieves the source code of called functions for datasets and leverages large language models (LLMs) with well-designed prompts to generate summaries for these functions. The datasets, enhanced with …


Reaccept: Automated Co-Evolution Of Production And Test Code Based On Dynamic Validation And Large Language Models, Jianlei Chi, Xiaotian Wang, Yuhan Huang, Lechen Yu, Di Cui, Jianguo Sun, Jun Sun Jun 2025

Reaccept: Automated Co-Evolution Of Production And Test Code Based On Dynamic Validation And Large Language Models, Jianlei Chi, Xiaotian Wang, Yuhan Huang, Lechen Yu, Di Cui, Jianguo Sun, Jun Sun

Research Collection School Of Computing and Information Systems

Synchronizing production and test code, known as PT co-evolution, is critical for software quality. Given the significant manual effort involved, researchers have tried automating PT co-evolution using predefined heuristics and machine learning models. However, existing solutions are still incomplete. Most approaches only detect and flag obsolete test cases, leaving developers to manually update them. Meanwhile, existing solutions may suffer from low accuracy, especially when applied to real-world software projects. In this paper, we propose ReAccept, a novel approach leveraging large language models (LLMs), retrievalaugmented generation (RAG), and dynamic validation to fully automate PT co-evolution with high accuracy. ReAccept employs an …


De-Duplicating Silent Compiler Bugs Via Deep Semantic Representation, Junjie Chen, Xingyu Fan, Chen Yang, Shuang Liu, Jun Sun Jun 2025

De-Duplicating Silent Compiler Bugs Via Deep Semantic Representation, Junjie Chen, Xingyu Fan, Chen Yang, Shuang Liu, Jun Sun

Research Collection School Of Computing and Information Systems

The compiler bug duplication problem (where many test failures are caused by the same compiler bug) can lead to huge waste of time and resource in diagnosing test failures produced by compiler testing. It is particularly challenging with regard to the silent compiler bugs that do not produce any error messages. To address this problem, multiple white-box techniques were proposed, but they are inapplicable in many practical scenarios. Black-box techniques are more practical, but the existing ones are less effective as they often rely on irrelevant syntactic information. To bridge this gap, we propose a novel black-box technique (BLADE), which …


A Comprehensive Study Of Oop-Related Bugs In C++ Compilers, Bo Wang, Chong Chen, Junjie Chen, Bowen Xu, Chen Ye, Youfang Lin, Guoliang Dong, Jun Sun Jun 2025

A Comprehensive Study Of Oop-Related Bugs In C++ Compilers, Bo Wang, Chong Chen, Junjie Chen, Bowen Xu, Chen Ye, Youfang Lin, Guoliang Dong, Jun Sun

Research Collection School Of Computing and Information Systems

Modern C++, a programming language characterized by its extensive use of object-oriented programming (OOP) features, is widely used for system programming. However, C++ compilers often struggle to correctly handle these sophisticated OOP features, resulting in numerous high-profile compiler bugs that can lead to crashes or miscompilation. Despite the significance of OOP-related bugs, existing studies largely overlook OOP features, hindering their ability to discover such bugs. To assist both compiler fuzzer designers and compiler developers, we conduct a comprehensive study of the compiler bugs caused by incorrectly handling C++ OOP-related features. First, we systematically extract 788 OOP-related C++ compiler bugs from …


Demystifying Memorization In Llm-Based Program Repair Via A General Hypothesis Testing Framework, Jiaolong Kong, Xiaofei Xie, Shangqing Liu Jun 2025

Demystifying Memorization In Llm-Based Program Repair Via A General Hypothesis Testing Framework, Jiaolong Kong, Xiaofei Xie, Shangqing Liu

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have achieved remarkable success in various applications, particularly in code-related tasks such as code generation and program repair, setting new performance benchmarks. However, the extensive use of large training corpora raises concerns about whether these achievements stem from genuine understanding or mere memorization of training data—a question often overlooked in current research. This paper aims to study the memorization issue within LLM-based program repair by investigating whether the correct patches generated by LLMs are the result of memorization. The key challenge lies in the absence of ground truth for confirming memorization, leading to various ad-hoc methods …


Cashift: Benchmarking Log-Based Cloud Attack Detection Under Normality Shift, Jiongchi Yu, Xiaofei Xie, Qiang Hu, Bowen Zhang, Ziming Zhao, Yun Lin, Lei Ma, Ruitao Feng, Frank Liau Jun 2025

Cashift: Benchmarking Log-Based Cloud Attack Detection Under Normality Shift, Jiongchi Yu, Xiaofei Xie, Qiang Hu, Bowen Zhang, Ziming Zhao, Yun Lin, Lei Ma, Ruitao Feng, Frank Liau

Research Collection School Of Computing and Information Systems

With the rapid advancement of cloud-native computing, securing cloud environments has become an important task. Log-based Anomaly Detection (LAD) is the most representative technique used in different systems for attack detection and safety guarantee, where multiple LAD methods and relevant datasets have been proposed. However, even though some of these datasets are specifically prepared for cloud systems, they only cover limited cloud behaviors and lack information from a whole-system perspective. Another critical issue to consider is normality shift, which implies that the test distribution could differ from the training distribution and highly affect the performance of LAD. Unfortunately, existing works …


Regtrieve: Reducing System-Level Regression Errors For Machine Learning Systems Via Retrieval-Enhanced Ensemble, Junming Cao, Xuwen Xiang, Mingfei Cheng, Bihuan Chen, Xinyan Wang, You Lu, Chaofeng Sha, Xiaofei Xie, Xin Peng Jun 2025

Regtrieve: Reducing System-Level Regression Errors For Machine Learning Systems Via Retrieval-Enhanced Ensemble, Junming Cao, Xuwen Xiang, Mingfei Cheng, Bihuan Chen, Xinyan Wang, You Lu, Chaofeng Sha, Xiaofei Xie, Xin Peng

Research Collection School Of Computing and Information Systems

Multiple machine learning (ML) models are often incorporated into real-world ML systems. However, updating an individual model in these ML systems frequently results in regression errors, where the new model performs worse than the old model for some inputs. While model-level regression errors have been widely studied, little is known about how regression errors propagate at system level. To address this gap, we propose RegTrieve, a novel retrieval-enhanced ensemble approach to reduce regression errors at both model and system level. Our evaluation across various model update scenarios shows that RegTrieve reduces system-level regression errors with almost no impact on system …


Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang Jun 2025

Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang

Research Collection School Of Computing and Information Systems

Low-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable inten sity. The former enforces …


Why Does My Transaction Fail? A First Look At Failed Transactions On The Solana Blockchain, Xiaoye Zheng, Zhiyuan Wan, David Lo, Difan Xie, Xiaohu Yang Jun 2025

Why Does My Transaction Fail? A First Look At Failed Transactions On The Solana Blockchain, Xiaoye Zheng, Zhiyuan Wan, David Lo, Difan Xie, Xiaohu Yang

Research Collection School Of Computing and Information Systems

Solana is an emerging blockchain platform, recognized for its high throughput and low transaction costs, positioning it as a preferred infrastructure for Decentralized Finance (DeFi), Non-Fungible Tokens (NFTs), and other Web 3.0 applications. In the Solana ecosystem, transaction initiators submit various instructions to interact with a diverse range of Solana smart contracts, among which are decentralized exchanges (DEXs) that utilize automated market makers (AMMs), allowing users to trade cryptocurrencies directly on the blockchain without the need for intermediaries. Despite the high throughput and low transaction costs of Solana, the advantages have exposed Solana to bot spamming for financial exploitation, resulting …


Less Is More: On The Importance Of Data Quality For Unit Test Generation, Junwei Zhang, Xing Hu, Shan Gao, Xin Xia, David Lo, Shanping Li Jun 2025

Less Is More: On The Importance Of Data Quality For Unit Test Generation, Junwei Zhang, Xing Hu, Shan Gao, Xin Xia, David Lo, Shanping Li

Research Collection School Of Computing and Information Systems

Unit testing is crucial for software development and maintenance. Effective unit testing ensures and improves software quality, but writing unit tests is time-consuming and labor-intensive. Recent studies have proposed deep learning (DL) techniques or large language models (LLMs) to automate unit test generation. These models are usually trained or fine-tuned on large-scale datasets. Despite growing awareness of the importance of data quality, there has been limited research on the quality of datasets used for test generation. To bridge this gap, we systematically examine the impact of noise on the performance of learning-based test generation models. We first apply the open …


Moditector: Module-Directed Testing For Autonomous Driving Systems, Renzhi Wang, Mingfei Cheng, Xiaofei Xie, Yuan Zhou, Lei Ma Jun 2025

Moditector: Module-Directed Testing For Autonomous Driving Systems, Renzhi Wang, Mingfei Cheng, Xiaofei Xie, Yuan Zhou, Lei Ma

Research Collection School Of Computing and Information Systems

Testing Autonomous Driving Systems (ADSs) is crucial for ensuring their safety, reliability, and performance. Despite numerous testing methods available that can generate diverse and challenging scenarios to uncover potential vulnerabilities, these methods often treat ADS as a black-box, primarily focusing on identifying system-level failures like collisions or near-misses without pinpointing the specific modules responsible for these failures. This lack of root causes understanding for the failures hinders effective debugging and subsequent system repair. Furthermore, current approaches often fall short in generating violations that adequately test the individual modules of an ADS from a system-level perspective, such as perception, prediction, planning, …


Large Language Model For Vulnerability Detection And Repair: Literature Review And The Road Ahead, Xin Zhou, Sicong Cao, Xiaobing Sun, David Lo Jun 2025

Large Language Model For Vulnerability Detection And Repair: Literature Review And The Road Ahead, Xin Zhou, Sicong Cao, Xiaobing Sun, David Lo

Research Collection School Of Computing and Information Systems

The significant advancements in Large Language Models (LLMs) have resulted in their widespread adoption across various tasks within Software Engineering (SE), including vulnerability detection and repair. Numerous studies have investigated the application of LLMs to enhance vulnerability detection and repair tasks. Despite the increasing research interest, there is currently no existing survey that focuses on the utilization of LLMs for vulnerability detection and repair. In this paper, we aim to bridge this gap by offering a systematic literature review of approaches aimed at improving vulnerability detection and repair through the utilization of LLMs. The review encompasses research work from leading …


Ntire 2025 Challenge On Event-Based Image Deblurring: Methods And Results, Lei Sun, Et. Al. Jun 2025

Ntire 2025 Challenge On Event-Based Image Deblurring: Methods And Results, Lei Sun, Et. Al.

Research Collection School Of Computing and Information Systems

This paper presents an overview of NTIRE 2025, the First Challenge on Event-Based Image Deblurring, detailing the proposed methodologies and corresponding results. The primary goal of the challenge is to design an event-based method that achieves high-quality image deblurring, with performance quantitatively assessed using Peak Signal-toNoise Ratio (PSNR). Notably, there are no restrictions on computational complexity or model size. The task focuses on leveraging both events and images as inputs for singleimage deblurring. A total of 199 participants registered, among whom 15 teams successfully submitted valid results, offering valuable insights into the current state of eventbased image deblurring. We anticipate …


A Knowledge Enhanced Large Language Model For Bug Localization, Yue Li, Bohan Liu, Ting Zhang, Zhiqi Wang, David Lo, Lanxin Yang, Jun Lyu, He Zhang Jun 2025

A Knowledge Enhanced Large Language Model For Bug Localization, Yue Li, Bohan Liu, Ting Zhang, Zhiqi Wang, David Lo, Lanxin Yang, Jun Lyu, He Zhang

Research Collection School Of Computing and Information Systems

A significant number of bug reports are generated every day as software systems continue to develop. Large Language Models (LLMs) have been used to correlate bug reports with source code to locate bugs automatically. The existing research has shown that LLMs are effective for bug localization and can increase software development efficiency. However, these studies still have two limitations. First, these models fail to capture context information about bug reports and source code. Second, these models are unable to understand the domain-specific expertise inherent to particular projects, such as version information in projects that are composed of alphanumeric characters without …


Efficient And Green Large Language Models For Software Engineering: Literature Review, Vision, And The Road Ahead, Jieke Shi, Zhou Yang, David Lo Jun 2025

Efficient And Green Large Language Models For Software Engineering: Literature Review, Vision, And The Road Ahead, Jieke Shi, Zhou Yang, David Lo

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have recently shown remarkable capabilities in various software engineering tasks, spurring the rapid growth of the Large Language Models for Software Engineering (LLM4SE) area. However, limited attention has been paid to developing efficient LLM4SE techniques that demand minimal computational cost, time, and memory resources, as well as green LLM4SE solutions that reduce energy consumption, water usage, and carbon emissions. This article aims to redirect the focus of the research community toward the efficiency and greenness of LLM4SE, while also sharing potential research directions to achieve this goal. It commences with a brief overview of the significance …


Deepvec: State-Vector Aware Test Case Selection For Enhancing Recurrent Neural Network, Zhonghao Jiang, Meng Yan, Li Huang, Weifeng Sun, Chao Liu, Song Sun, David Lo Jun 2025

Deepvec: State-Vector Aware Test Case Selection For Enhancing Recurrent Neural Network, Zhonghao Jiang, Meng Yan, Li Huang, Weifeng Sun, Chao Liu, Song Sun, David Lo

Research Collection School Of Computing and Information Systems

Deep Neural Networks (DNN) have realized significant achievements across various application domains. There is no doubt that testing and enhancing a pre-trained DNN that has been deployed in an application scenario is crucial, because it can reduce the failures of the DNN. DNN-driven software testing and enhancement require large amounts of labeled data. The high cost and inefficiency caused by the large volume of data of manual labeling, and the time consumption of testing all cases in real scenarios are unacceptable. Therefore, test case selection technologies are proposed to reduce the time cost by selecting and only labeling representative test …


On Lexicographic Proof Rules For Probabilistic Termination, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Jiří Zárevucký, Dorde Zikelic Jun 2025

On Lexicographic Proof Rules For Probabilistic Termination, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Jiří Zárevucký, Dorde Zikelic

Research Collection School Of Computing and Information Systems

We consider the almost-sure (a.s.) termination problem for probabilistic programs, which are a stochastic extension of classical imperative programs. Lexicographic ranking functions provide a sound and practical approach for termination of non-probabilistic programs, and their extension to probabilistic programs is achieved via lexicographic ranking supermartingales (LexRSMs). However, LexRSMs introduced in the previous work have a limitation that impedes their automation: all of their components have to be non-negative in all reachable states. This might result in a LexRSM not existing even for simple terminating programs. Our contributions are twofold. First, we introduce a generalization of LexRSMs that allows for some …


Dissecting Global Search: A Simple Yet Effective Method To Boost Individual Discrimination Testing And Repair, Lili Quan, Tianlin Li, Xiaofei Xie, Zhenpeng Chen, Sen Chen, Lingxiao Jiang, Xiaohong Li May 2025

Dissecting Global Search: A Simple Yet Effective Method To Boost Individual Discrimination Testing And Repair, Lili Quan, Tianlin Li, Xiaofei Xie, Zhenpeng Chen, Sen Chen, Lingxiao Jiang, Xiaohong Li

Research Collection School Of Computing and Information Systems

Deep Learning (DL) has achieved significant success in socially critical decision-making applications but often exhibits unfair behaviors, raising social concerns. Among these unfair behaviors, individual discrimination-examining inequalities between instance pairs with identical profiles differing only in sensitive attributes such as gender, race, and age-is extremely socially impactful. Existing methods have made significant and commendable efforts in testing individual discrimination before deployment. However, their efficiency and effectiveness remain limited, particularly when evaluating relatively fairer models. It remains unclear which phase of the existing testing framework (global or local) is the primary bottleneck limiting performance. Facing the above issues, we first identify …


Enriching Automatic Test Case Generation By Extracting Relevant Test Inputs From Bug Reports, Wendkuuni C. Ouedraogo, Laura Plein, Kader Kabore, Andrew Habib, Jacques Klein, David Lo, Tegawende F. Bissyande May 2025

Enriching Automatic Test Case Generation By Extracting Relevant Test Inputs From Bug Reports, Wendkuuni C. Ouedraogo, Laura Plein, Kader Kabore, Andrew Habib, Jacques Klein, David Lo, Tegawende F. Bissyande

Research Collection School Of Computing and Information Systems

The quality of software is closely tied to the effectiveness of the tests it undergoes. Manual test writing, though crucial for bug detection, is time-consuming, which has driven significant research into automated test case generation. However, current methods often struggle to generate relevant inputs, limiting the effectiveness of the tests produced. To address this, we introduce BRMiner, a novel approach that leverages Large Language Models (LLMs) in combination with traditional techniques to extract relevant inputs from bug reports, thereby enhancing automated test generation tools. In this study, we evaluate BRMiner using the Defects4J benchmark and test generation tools such as …


Enhancing Deliberativeness: Evaluating The Impact Of Multimodal Reflection Nudges, Shun Yi Yeo, Zhuoqun Jiang, Anthony Tang, Simon Tangi Perrault May 2025

Enhancing Deliberativeness: Evaluating The Impact Of Multimodal Reflection Nudges, Shun Yi Yeo, Zhuoqun Jiang, Anthony Tang, Simon Tangi Perrault

Research Collection School Of Computing and Information Systems

Nudging participants with text-based reflective nudges enhances deliberation quality on online deliberation platforms. The effectiveness of multimodal reflective nudges, however, remains largely unexplored. Given the multi-sensory nature of human perception, incorporating diverse modalities into self-reflection mechanisms has the potential to better support various reflective styles. This paper explores how presenting reflective nudges of different types (direct: persona and indirect: storytelling) in different modalities (text, image, video and audio) affects deliberation quality. We conducted two user studies with 20 and 200 participants respectively. The first study identifies the preferred modality for each type of reflective nudges, revealing that text is most …


Prompting An Embodied Ai Agent: How Embodiment And Multimodal Signaling Affects Prompting Behaviour, Tianyi Zhang, Colin Au Yeung, Emily Aurelia, Yuki Onishi, Neil Chulpongsatorn, Jiannan Li, Anthony Tang May 2025

Prompting An Embodied Ai Agent: How Embodiment And Multimodal Signaling Affects Prompting Behaviour, Tianyi Zhang, Colin Au Yeung, Emily Aurelia, Yuki Onishi, Neil Chulpongsatorn, Jiannan Li, Anthony Tang

Research Collection School Of Computing and Information Systems

Current voice agents wait for a user to complete their verbal instruction before responding; yet, this is misaligned with how humans engage in everyday conversational interaction, where interlocutors use multimodal signaling (e.g. nodding, grunting, or looking at referred to objects) to ensure conversational grounding. We designed an embodied VR agent that exhibits multimodal signaling behaviors in response to situated prompts, by turning its head, or by visually highlighting objects being discussed or referred to. We explore how people prompt this agent to design and manipulate the objects in a VR scene. Through a Wizard of Oz study, we found that …