Open Access. Powered by Scholars. Published by Universities.®

Physical Sciences and Mathematics Commons

Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 12931 - 12960 of 291657

Full-Text Articles in Physical Sciences and Mathematics

On Lexicographic Proof Rules For Probabilistic Termination, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Jiří Zárevucký, Dorde Zikelic Jun 2025

On Lexicographic Proof Rules For Probabilistic Termination, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Petr Novotný, Jiří Zárevucký, Dorde Zikelic

Research Collection School Of Computing and Information Systems

We consider the almost-sure (a.s.) termination problem for probabilistic programs, which are a stochastic extension of classical imperative programs. Lexicographic ranking functions provide a sound and practical approach for termination of non-probabilistic programs, and their extension to probabilistic programs is achieved via lexicographic ranking supermartingales (LexRSMs). However, LexRSMs introduced in the previous work have a limitation that impedes their automation: all of their components have to be non-negative in all reachable states. This might result in a LexRSM not existing even for simple terminating programs. Our contributions are twofold. First, we introduce a generalization of LexRSMs that allows for some …


Community Detection In Heterogeneous Information Networks Without Materialization, Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu Jun 2025

Community Detection In Heterogeneous Information Networks Without Materialization, Jiaxin Jiang, Siyuan Yao, Yuhang Chen, Bingsheng He, Yudong Niu, Yuchen Li, Shixuan Sun, Yongchao Liu

Research Collection School Of Computing and Information Systems

Community detection in heterogeneous information networks (HINs) poses significant challenges due to the diversity of entity types and the complexity of their interrelations. While traditional algorithms may perform adequately in some scenarios, many struggle with the high memory usage and computational demands of large-scale HINs. To address these challenges, we introduce a novel framework, SCAR, which efficiently uncovers community structures in HINs without requiring network materialization. SCAR leverages insights from meta-paths to interpret multi-relational data through compact vertex-based sketches, significantly reducing computational overhead and materialization overhead. We propose a sketch-based technique for estimating changes in modularity, improving both the precision …


Large Language Model For Vulnerability Detection And Repair: Literature Review And The Road Ahead, Xin Zhou, Sicong Cao, Xiaobing Sun, David Lo Jun 2025

Large Language Model For Vulnerability Detection And Repair: Literature Review And The Road Ahead, Xin Zhou, Sicong Cao, Xiaobing Sun, David Lo

Research Collection School Of Computing and Information Systems

The significant advancements in Large Language Models (LLMs) have resulted in their widespread adoption across various tasks within Software Engineering (SE), including vulnerability detection and repair. Numerous studies have investigated the application of LLMs to enhance vulnerability detection and repair tasks. Despite the increasing research interest, there is currently no existing survey that focuses on the utilization of LLMs for vulnerability detection and repair. In this paper, we aim to bridge this gap by offering a systematic literature review of approaches aimed at improving vulnerability detection and repair through the utilization of LLMs. The review encompasses research work from leading …


Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu Jun 2025

Keep The Balance: A Parameter-Efficient Symmetrical Framework For Rgb+X Semantic Segmentation, Jiaxin Cai, Jingze Su, Qi Li, Wenjie Yang, Shu Wang, Tiesong Zhao, Shengfeng He, Wenxi Liu

Research Collection School Of Computing and Information Systems

Multimodal semantic segmentation is a critical challenge in computer vision, with early methods suffering from high computational costs and limited transferability due to full fine-tuning of RGB-based pre-trained parameters. Recent studies, while leveraging additional modalities as supplementary prompts to RGB, still predominantly rely on RGB, which restricts the full potential of other modalities. To address these issues, we propose a novel symmetric parameter-efficient fine-tuning framework for multimodal segmentation, featuring with a modality-aware prompting and adaptation scheme, to simultaneously adapt the capabilities of a powerful pre-trained model to both RGB and X modalities. Furthermore, prevalent approaches use the global cross-modality correlations …


A Knowledge Enhanced Large Language Model For Bug Localization, Yue Li, Bohan Liu, Ting Zhang, Zhiqi Wang, David Lo, Lanxin Yang, Jun Lyu, He Zhang Jun 2025

A Knowledge Enhanced Large Language Model For Bug Localization, Yue Li, Bohan Liu, Ting Zhang, Zhiqi Wang, David Lo, Lanxin Yang, Jun Lyu, He Zhang

Research Collection School Of Computing and Information Systems

A significant number of bug reports are generated every day as software systems continue to develop. Large Language Models (LLMs) have been used to correlate bug reports with source code to locate bugs automatically. The existing research has shown that LLMs are effective for bug localization and can increase software development efficiency. However, these studies still have two limitations. First, these models fail to capture context information about bug reports and source code. Second, these models are unable to understand the domain-specific expertise inherent to particular projects, such as version information in projects that are composed of alphanumeric characters without …


Human-Computer Interaction And Artificial Intelligence For Ageing Population, Keng Siau, Hailiang Wang, Fiona Fui-Hoon Nah, Runyu Wang, Ruitong Che, Can Liu Jun 2025

Human-Computer Interaction And Artificial Intelligence For Ageing Population, Keng Siau, Hailiang Wang, Fiona Fui-Hoon Nah, Runyu Wang, Ruitong Che, Can Liu

Research Collection School Of Computing and Information Systems

As the global population ages rapidly, the field of human-computer interaction (HCI) is in urgent need of innovation, redesign, and reengineering to meet the evolving needs of older adults. The older demographic faces a range of challenges—including physical limitations, cognitive decline, reduced social in-tegration, and varying levels of technological literacy—that can hinder effective engagement with digital technologies. In response to these challenges, research-ers and designers are using inclusive and adaptive approaches to enhance acces-sibility, usability, and emotional well-being. This paper reviews key design prin-ciples in HCI for the ageing population and discusses how artificial intelligence (AI) tools, such as voice …


Unlocking The Power Of Socio-Knowledge Association For Enterprise Risk Identification In Stock Market, Zhenghao Liu, Keng Siau, Shaochen Yang, Feicheng Ma Jun 2025

Unlocking The Power Of Socio-Knowledge Association For Enterprise Risk Identification In Stock Market, Zhenghao Liu, Keng Siau, Shaochen Yang, Feicheng Ma

Research Collection School Of Computing and Information Systems

Potential risk signals reflected in supply chain and equity connections between enterprises and social connections between investors are becoming crucial to identifying enterprise risks in addition to basic financial indicators. Traditional risk management systems face challenges in adapting to these complexities, highlighting the need for a proactive paradigm shift in risk management. Leveraging graph models such as social networks and knowledge graphs offers a promising approach to identifying and managing potential associated risks effectively. To bridge existing research gaps, a novel risk identification framework driven by social-knowledge graphs has been proposed, integrating graph deep learning and reinforcement learning techniques guided …


Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang Jun 2025

Hvi: A New Color Space For Low-Light Image Enhancement, Qingsen Yan, Yixu Feng, Cheng Zhang, Guansong Pang, Kangbiao Shi, Peng Wu, Wei Dong, Jinqiu Sun, Yanning Zhang

Research Collection School Of Computing and Information Systems

Low-Light Image Enhancement (LLIE) is a crucial computer vision task that aims to restore detailed visual information from corrupted low-light images. Many existing LLIE methods are based on standard RGB (sRGB) space, which often produce color bias and brightness artifacts due to inherent high color sensitivity in sRGB. While converting the images using Hue, Saturation and Value (HSV) color space helps resolve the brightness issue, it introduces significant red and black noise artifacts. To address this issue, we propose a new color space for LLIE, namely Horizontal/Vertical-Intensity (HVI), defined by polarized HS maps and learnable inten sity. The former enforces …


Why Does My Transaction Fail? A First Look At Failed Transactions On The Solana Blockchain, Xiaoye Zheng, Zhiyuan Wan, David Lo, Difan Xie, Xiaohu Yang Jun 2025

Why Does My Transaction Fail? A First Look At Failed Transactions On The Solana Blockchain, Xiaoye Zheng, Zhiyuan Wan, David Lo, Difan Xie, Xiaohu Yang

Research Collection School Of Computing and Information Systems

Solana is an emerging blockchain platform, recognized for its high throughput and low transaction costs, positioning it as a preferred infrastructure for Decentralized Finance (DeFi), Non-Fungible Tokens (NFTs), and other Web 3.0 applications. In the Solana ecosystem, transaction initiators submit various instructions to interact with a diverse range of Solana smart contracts, among which are decentralized exchanges (DEXs) that utilize automated market makers (AMMs), allowing users to trade cryptocurrencies directly on the blockchain without the need for intermediaries. Despite the high throughput and low transaction costs of Solana, the advantages have exposed Solana to bot spamming for financial exploitation, resulting …


Efficient And Green Large Language Models For Software Engineering: Literature Review, Vision, And The Road Ahead, Jieke Shi, Zhou Yang, David Lo Jun 2025

Efficient And Green Large Language Models For Software Engineering: Literature Review, Vision, And The Road Ahead, Jieke Shi, Zhou Yang, David Lo

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have recently shown remarkable capabilities in various software engineering tasks, spurring the rapid growth of the Large Language Models for Software Engineering (LLM4SE) area. However, limited attention has been paid to developing efficient LLM4SE techniques that demand minimal computational cost, time, and memory resources, as well as green LLM4SE solutions that reduce energy consumption, water usage, and carbon emissions. This article aims to redirect the focus of the research community toward the efficiency and greenness of LLM4SE, while also sharing potential research directions to achieve this goal. It commences with a brief overview of the significance …


Ef21 With Bells & Whistles: Six Algorithmic Extensions Of Modern Error Feedback, Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov, Zhize Li, Peter Richtarik Jun 2025

Ef21 With Bells & Whistles: Six Algorithmic Extensions Of Modern Error Feedback, Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov, Zhize Li, Peter Richtarik

Research Collection School Of Computing and Information Systems

First proposed by Seide (2014) as a heuristic, error feedback (EF) is a very popular mechanism for enforcing convergence of distributed gradient-based optimization methods enhanced with communication compression strategies based on the application of contractive compression operators. However, existing theory of EF relies on very strong assumptions (e.g., bounded gradients), and provides pessimistic convergence rates (e.g., while the best known rate for EF in the smooth nonconvex regime, and when full gradients are compressed, is O(1/T2/3), the rate of gradient descent in the same regime is O(1/T)). Recently, Richtàrik et al. (2021) proposed a new error feedback mechanism, EF21, based …


Less Is More: On The Importance Of Data Quality For Unit Test Generation, Junwei Zhang, Xing Hu, Shan Gao, Xin Xia, David Lo, Shanping Li Jun 2025

Less Is More: On The Importance Of Data Quality For Unit Test Generation, Junwei Zhang, Xing Hu, Shan Gao, Xin Xia, David Lo, Shanping Li

Research Collection School Of Computing and Information Systems

Unit testing is crucial for software development and maintenance. Effective unit testing ensures and improves software quality, but writing unit tests is time-consuming and labor-intensive. Recent studies have proposed deep learning (DL) techniques or large language models (LLMs) to automate unit test generation. These models are usually trained or fine-tuned on large-scale datasets. Despite growing awareness of the importance of data quality, there has been limited research on the quality of datasets used for test generation. To bridge this gap, we systematically examine the impact of noise on the performance of learning-based test generation models. We first apply the open …


Group-And-Match Vs. Route-Then-Insert: Order Dispatching In Vehicle-Based Dual Services (Vedus), Yue Lin, Hai Yang, Hai Wang Jun 2025

Group-And-Match Vs. Route-Then-Insert: Order Dispatching In Vehicle-Based Dual Services (Vedus), Yue Lin, Hai Yang, Hai Wang

Research Collection School Of Computing and Information Systems

Rapid urban transportation and delivery demand and relevant resource constraints have driven the need for more efficient vehicle utilization. An innovative concept, “Vehicle-based MultiServices” (VeMuS), is a service model in which a single vehicle offers multiple services simultaneously in an urban mobility system. Similarly, “Vehicle-based Dual Services” (VeDuS) refers to a vehicle that provides two services simultaneously (Sun et al., 2023).


Ntire 2025 Challenge On Event-Based Image Deblurring: Methods And Results, Lei Sun, Et. Al. Jun 2025

Ntire 2025 Challenge On Event-Based Image Deblurring: Methods And Results, Lei Sun, Et. Al.

Research Collection School Of Computing and Information Systems

This paper presents an overview of NTIRE 2025, the First Challenge on Event-Based Image Deblurring, detailing the proposed methodologies and corresponding results. The primary goal of the challenge is to design an event-based method that achieves high-quality image deblurring, with performance quantitatively assessed using Peak Signal-toNoise Ratio (PSNR). Notably, there are no restrictions on computational complexity or model size. The task focuses on leveraging both events and images as inputs for singleimage deblurring. A total of 199 participants registered, among whom 15 teams successfully submitted valid results, offering valuable insights into the current state of eventbased image deblurring. We anticipate …


A Multimodal Fusion Model Leveraging Mlp Mixer And Handcrafted Features-Based Deep Learning Networks For Facial Palsy Detection, Heng Yim Nicole Oo, Min Hun Lee, Jeong Hoon Lim Jun 2025

A Multimodal Fusion Model Leveraging Mlp Mixer And Handcrafted Features-Based Deep Learning Networks For Facial Palsy Detection, Heng Yim Nicole Oo, Min Hun Lee, Jeong Hoon Lim

Research Collection School Of Computing and Information Systems

Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessments by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes an MLP mixer-based model to process unstructured data (i.e. RGB images or images with facial line segments) and a feed-forward neural network to process structured data (i.e. facial landmark coordinates, features of facial expressions, or handcrafted features) for detecting facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of …


Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al. Jun 2025

Event-Based Eye Tracking: Event-Based Vision Workshop 2025, Qinyu Chen, Et. Al.

Research Collection School Of Computing and Information Systems

No abstract provided.


Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He Jun 2025

Modfinity: Unsupervised Domain Adaptation With Multimodal Information Flow Intertwining, Shanglin Liu, Jianming Lv, Jingdan Kang, Huaidong Zhang, Zequan Liang, Shengfeng He

Research Collection School Of Computing and Information Systems

Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of different modalities through pseudo-labeling and feature-level exchange, varying sample quality across modalities can lead to the propagation of inaccurate information, resulting in error accumulation. To address this, we propose Modal-Affinity Multimodal Domain Adaptation (MODfinity), a method that dynamically manages multimodal information flow through fine-grained control over teacher model selection, guiding information intertwining at both feature and label levels. By treating labels as an independent modality, MODfinity enables balanced performance assessment across modalities, employing a novel …


Large Language Models For Logical Fallacy Detection, Nicole Anne Hui-Ying Teo, Donghao Huang, Erik Cambria, Zhaoxia Wang Jun 2025

Large Language Models For Logical Fallacy Detection, Nicole Anne Hui-Ying Teo, Donghao Huang, Erik Cambria, Zhaoxia Wang

Research Collection School Of Computing and Information Systems

Identifying logical fallacies is essential for maintaining log-ical reasoning and reducing false information in a variety of domains, such as the media, law, and education. We present an extensive study on the use of large language models (LLMs) for logical fallacy detection and provide a comparative overview of model performance across various fallacy classes. We evaluate the logical fallacy detection capabilities of multiple state-of-the-art models (LLaMA, Qwen, Gemma, Phi) utilizing accuracy, precision, recall, and F1-score as assessment measures. Accord-ing to our findings, our models do well on simple fallacies like “circular reasoning,” but they have trouble with more interpretive reasoning …


Verify All Traffic: Towards Zero-Trust In-Network Intrusion Detection Against Multipath Routing, Ziming Zhao, Zhaoxuan Li, Xiaofei Xie, Zhipeng Liu, Tingting Li, Jiongchi Yu, Fan Zhang, Binbin Chen Jun 2025

Verify All Traffic: Towards Zero-Trust In-Network Intrusion Detection Against Multipath Routing, Ziming Zhao, Zhaoxuan Li, Xiaofei Xie, Zhipeng Liu, Tingting Li, Jiongchi Yu, Fan Zhang, Binbin Chen

Research Collection School Of Computing and Information Systems

With the popularity of encryption protocols, machine learning (ML)-based traffic analysis technologies have attracted widespread attention. To adapt to modern high-speed bandwidth, recent research is dedicated to advancing zero-trust intrusion detection by offloading feature extraction and model inference into the network dataplane. Especially, with the rise of programmable switches, achieving line-speed ML inference becomes promising. However, existing research only considers a single switch node as a relay to conduct evaluation. This is far from real-world deployments involving multiple switches (given that zero-trust security assumes that threats can originate from anywhere, including within the network), particularly the multipath routing phenomenon that …


Learning Spatio-Temporal Dynamics For Trajectory Recovery Via Time-Aware Transformer, Tian Sun, Yuqi Chen, Baihua Zheng, Weiwei Sun Jun 2025

Learning Spatio-Temporal Dynamics For Trajectory Recovery Via Time-Aware Transformer, Tian Sun, Yuqi Chen, Baihua Zheng, Weiwei Sun

Research Collection School Of Computing and Information Systems

In real-world applications, GPS trajectories often suffer from low sampling rates, with large and irregular intervals between consecutive GPS points. This sparse characteristic presents challenges for their direct use in GPS-based systems. This paper addresses the task of map-constrained trajectory recovery, aiming to enhance trajectory sampling rates of GPS trajectories. Previous studies commonly adopt a sequence-to-sequence framework, where an encoder captures the trajectory patterns and a decoder reconstructs the target trajectory. Within this framework, effectively representing the road network and extracting relevant trajectory features are crucial for overall performance. Despite advancements in these models, they fail to fully leverage the …


Building Narratives And Probing Concepts: Preparing Materials For Co-Design With Autistic Livestreamers, Terrance Mok, Tyson Hartley, Anthony Tang, Adam Mccrimmon, Lora Oehlberg Jun 2025

Building Narratives And Probing Concepts: Preparing Materials For Co-Design With Autistic Livestreamers, Terrance Mok, Tyson Hartley, Anthony Tang, Adam Mccrimmon, Lora Oehlberg

Research Collection School Of Computing and Information Systems

Based on ten semi-structured interviews with autistic Twitch streamers, we introduce a series of scenario-based design narratives coupled with technology design concepts as a starting point for co-design discussion about autistic streaming. This work builds on prior thematic analysis of the unique intersection between autism and livestreaming. Our user-centered scenarios highlight the needs, goals, and challenges of autistic individuals in livestreaming contexts. By using evocative narratives, the scenarios serve to facilitate empathy and deeper engagement with the needs of autistic users, and help facilitate and support co-creative dialogues and discussions about new technology designs. We contribute this starting point for …


Meta-Learning Hyperparameters For Foundation Model Adaptation In Remote-Sensing Imagery, Zichen Tian, Yaoyao Liu, Qianru Sun Jun 2025

Meta-Learning Hyperparameters For Foundation Model Adaptation In Remote-Sensing Imagery, Zichen Tian, Yaoyao Liu, Qianru Sun

Research Collection School Of Computing and Information Systems

Training large foundation models of remote-sensing (RS) images is almost impossible due to the limited and long-tailed data problems. Fine-tuning natural image pre-trained models on RS images is a straightforward solution. To reduce computational costs and improve performance on tail classes, existing methods apply parameter-efficient fine-tuning (PEFT) techniques, such as LoRA and AdaptFormer. However, we observe that fixed hyperparameters -- such as intra-layer positions, layer depth, and scaling factors, can considerably hinder PEFT performance, as fine-tuning on RS images proves highly sensitive to these settings. To address this, we propose MetaPEFT, a method incorporating adaptive scalers that dynamically adjust module …


Irhunter: Universal Detection Of Instruction Reordering Vulnerabilities For Enhanced Concurrency In Distributed And Parallel Systems, Guohua Xin, Guangquan Xu, Yao Zhang, Cheng Wen, Cen Zhang, Xiaofei Xie, Neal N. Xiong, Shaoying Liu, Pan Gao Jun 2025

Irhunter: Universal Detection Of Instruction Reordering Vulnerabilities For Enhanced Concurrency In Distributed And Parallel Systems, Guohua Xin, Guangquan Xu, Yao Zhang, Cheng Wen, Cen Zhang, Xiaofei Xie, Neal N. Xiong, Shaoying Liu, Pan Gao

Research Collection School Of Computing and Information Systems

Instruction reordering is an essential optimization technique used in both compilers and multi-core processors to enhance parallelism and resource utilization. Although the original intent of this technique is to benefit the program, some improper reordering can significantly impact the program correctness, which we call instruction reordering vulnerability (IRV). However, existing methods detect IRV by defining CPU instruction reordering rules to schedule execution paths while neglecting compiler reordering, and thus generate false positives that require manual filtering and resulting in inefficiency. To bridge this gap, in this paper, we propose the IRV detection method, , which analyzes IRV characteristics and extracts …


On-Demand Heterogeneous Drone Delivery Problem, Xupeng Wen, Zhiguang Cao, Shu Xu, Dapeng Ren, Guohua Wu, Yaoxin Wu Jun 2025

On-Demand Heterogeneous Drone Delivery Problem, Xupeng Wen, Zhiguang Cao, Shu Xu, Dapeng Ren, Guohua Wu, Yaoxin Wu

Research Collection School Of Computing and Information Systems

In the on-demand problem domain, actual demand frequently deviates from the expected demand. This paper intricately delves into the exploration of on-demand heterogeneous multi-drone routing problem (ODHDRP), in which a transport drone carries multiple terminal drones to subregions in the first echelon, and the terminal drones deliver parcels during a flight trip to customers with demands in subregions to maintain economies of scale in the second echelon. We formulate the customer demands using a normal distribution, and exploit a reliability model of customer demands with chance constraints. To solve the ODHDRP efficiently, we propose a hybrid iterative optimisation heuristic (HIOH) …


Alayadb: The Data Foundation For Efficient And Effective Long-Context Llm Inference, Yangshen Deng, Zhengxin You, Long Xiang, Qilong Li, Peiqi Yuan, Zhaoyang Hong, Yitao Zheng, Wanting Li, Runzhong Li, Haotian Liu, Kyriakos Mouratidis, Man Lung Yiu, Huan Li, Qiaomu Shen, Rui Mao, Bo Tang Jun 2025

Alayadb: The Data Foundation For Efficient And Effective Long-Context Llm Inference, Yangshen Deng, Zhengxin You, Long Xiang, Qilong Li, Peiqi Yuan, Zhaoyang Hong, Yitao Zheng, Wanting Li, Runzhong Li, Haotian Liu, Kyriakos Mouratidis, Man Lung Yiu, Huan Li, Qiaomu Shen, Rui Mao, Bo Tang

Research Collection School Of Computing and Information Systems

AlayaDB is a cutting-edge vector database system natively architected for efficient and effective long-context inference for Large Language Models (LLMs) at AlayaDB AI. Specifically, it decouples the KV cache and attention computation from the LLM inference systems, and encapsulates them into a novel vector database system. For the Model as a Service providers (MaaS), AlayaDB consumes fewer hardware resources and offers higher generation quality for various workloads with different kinds of Service Level Objectives (SLOs), when compared with the existing alternative solutions (e.g., KV cache disaggregation, retrieval-based sparse attention). The crux of AlayaDB is that it abstracts the attention computation …


Enhancing Vulnerability Detection Via Inter-Procedural Semantic Completion, Bozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao, Jun Sun, Shang-Wei Lin Jun 2025

Enhancing Vulnerability Detection Via Inter-Procedural Semantic Completion, Bozhi Wu, Chengjie Liu, Zhiming Li, Yushi Cao, Jun Sun, Shang-Wei Lin

Research Collection School Of Computing and Information Systems

Inspired by advances in deep learning, numerous learning-based approaches for vulnerability detection have emerged, primarily operating at the function level for scalability. However, this design choice has a critical limitation: many vulnerabilities span multiple functions, causing function-level approaches to lose the semantics of called functions and fail to capture true vulnerability patterns. To address this issue, we propose VulnSC, a novel framework designed to enhance learning-based approaches by complementing inter-procedural semantics. VulnSC retrieves the source code of called functions for datasets and leverages large language models (LLMs) with well-designed prompts to generate summaries for these functions. The datasets, enhanced with …


Reaccept: Automated Co-Evolution Of Production And Test Code Based On Dynamic Validation And Large Language Models, Jianlei Chi, Xiaotian Wang, Yuhan Huang, Lechen Yu, Di Cui, Jianguo Sun, Jun Sun Jun 2025

Reaccept: Automated Co-Evolution Of Production And Test Code Based On Dynamic Validation And Large Language Models, Jianlei Chi, Xiaotian Wang, Yuhan Huang, Lechen Yu, Di Cui, Jianguo Sun, Jun Sun

Research Collection School Of Computing and Information Systems

Synchronizing production and test code, known as PT co-evolution, is critical for software quality. Given the significant manual effort involved, researchers have tried automating PT co-evolution using predefined heuristics and machine learning models. However, existing solutions are still incomplete. Most approaches only detect and flag obsolete test cases, leaving developers to manually update them. Meanwhile, existing solutions may suffer from low accuracy, especially when applied to real-world software projects. In this paper, we propose ReAccept, a novel approach leveraging large language models (LLMs), retrievalaugmented generation (RAG), and dynamic validation to fully automate PT co-evolution with high accuracy. ReAccept employs an …


On-Demand Scenario Generation For Testing Automated Driving Systems, Songyang Yan, Xiaodong Zhang, Kunkun Hao, Haojie Xin, Yonggang Luo, Jucheng Yang, Ming Fan, Chao Yang, Jun Sun, Zijiang Yang Jun 2025

On-Demand Scenario Generation For Testing Automated Driving Systems, Songyang Yan, Xiaodong Zhang, Kunkun Hao, Haojie Xin, Yonggang Luo, Jucheng Yang, Ming Fan, Chao Yang, Jun Sun, Zijiang Yang

Research Collection School Of Computing and Information Systems

The safety and reliability of Automated Driving Systems (ADS) are paramount, necessitating rigorous testing methodologies to uncover potential failures before deployment. Traditional testing approaches often prioritize either natural scenario sampling or safety-critical scenario generation, resulting in overly simplistic or unrealistic hazardous tests. In practice, the demand for natural scenarios (e.g., when evaluating the ADS's reliability in real-world conditions), critical scenarios (e.g., when evaluating safety in critical situations), or somewhere in between (e.g., when testing the ADS in regions with less civilized drivers) varies depending on the testing objectives. To address this issue, we propose the On-demand Scenario Generation (OSG) Framework, …


De-Duplicating Silent Compiler Bugs Via Deep Semantic Representation, Junjie Chen, Xingyu Fan, Chen Yang, Shuang Liu, Jun Sun Jun 2025

De-Duplicating Silent Compiler Bugs Via Deep Semantic Representation, Junjie Chen, Xingyu Fan, Chen Yang, Shuang Liu, Jun Sun

Research Collection School Of Computing and Information Systems

The compiler bug duplication problem (where many test failures are caused by the same compiler bug) can lead to huge waste of time and resource in diagnosing test failures produced by compiler testing. It is particularly challenging with regard to the silent compiler bugs that do not produce any error messages. To address this problem, multiple white-box techniques were proposed, but they are inapplicable in many practical scenarios. Black-box techniques are more practical, but the existing ones are less effective as they often rely on irrelevant syntactic information. To bridge this gap, we propose a novel black-box technique (BLADE), which …


Demystifying Memorization In Llm-Based Program Repair Via A General Hypothesis Testing Framework, Jiaolong Kong, Xiaofei Xie, Shangqing Liu Jun 2025

Demystifying Memorization In Llm-Based Program Repair Via A General Hypothesis Testing Framework, Jiaolong Kong, Xiaofei Xie, Shangqing Liu

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have achieved remarkable success in various applications, particularly in code-related tasks such as code generation and program repair, setting new performance benchmarks. However, the extensive use of large training corpora raises concerns about whether these achievements stem from genuine understanding or mere memorization of training data—a question often overlooked in current research. This paper aims to study the memorization issue within LLM-based program repair by investigating whether the correct patches generated by LLMs are the result of memorization. The key challenge lies in the absence of ground truth for confirming memorization, leading to various ad-hoc methods …