Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (239)
- Programming Languages and Compilers (198)
- Artificial Intelligence and Robotics (142)
- Engineering (142)
- Computer Engineering (126)
-
- Information Security (109)
- OS and Networks (80)
- Graphics and Human Computer Interfaces (78)
- Social and Behavioral Sciences (76)
- Numerical Analysis and Scientific Computing (59)
- Theory and Algorithms (46)
- Business (43)
- Computer and Systems Architecture (39)
- Digital Communications and Networking (39)
- Medicine and Health Sciences (33)
- Communication (30)
- Education (28)
- Systems Architecture (21)
- Health Information Technology (20)
- Sociology (18)
- Finance and Financial Management (15)
- Gerontology (15)
- Public Affairs, Public Policy and Public Administration (15)
- Higher Education (13)
- Social Media (13)
- Transportation (12)
- Technology and Innovation (11)
- Keyword
-
- Deep learning (54)
- Software engineering (49)
- Empirical study (43)
- Machine learning (35)
- Software (30)
-
- Android (29)
- Model Check (29)
- Collaboration (26)
- Deep Learning (24)
- GitHub (22)
- Stack Overflow (22)
- Fuzzing (21)
- Security (21)
- Testing (20)
- Data mining (19)
- Large language models (18)
- Programming (18)
- Software Engineering (18)
- Code search (17)
- Computer bugs (17)
- Codes (16)
- Empirical Study (16)
- Information retrieval (16)
- Large Language Models (16)
- Large language model (16)
- Software testing (15)
- Large Language Model (14)
- Linear Temporal Logic (14)
- Software maintenance (14)
- Vulnerability detection (14)
- Publication Year
- Publication
- Publication Type
Articles 181 - 210 of 2211
Full-Text Articles in Software Engineering
Can Llms Replace Manual Annotation Of Software Engineering Artifacts?, Toufique Ahmed, Premkumar Devanbu, Christoph Treude, Michael Pradel
Can Llms Replace Manual Annotation Of Software Engineering Artifacts?, Toufique Ahmed, Premkumar Devanbu, Christoph Treude, Michael Pradel
Research Collection School Of Computing and Information Systems
Experimental evaluations of software engineering innovations, e.g., tools and processes, often include human-subject studies as a component of a multi-pronged strategy to obtain greater generalizability of the findings. However, human-subject studies in our field are challenging, due to the cost and difficulty of finding and employing suitable subjects, ideally, professional programmers with varying degrees of experience. Meanwhile, large language models (LLMs) have recently started to demonstrate human-level performance in several areas. This paper explores the possibility of substituting costly human subjects with much cheaper LLM queries in evaluations of code and code-related artifacts. We study this idea by applying six …
A Functional Software Reference Architecture For Llm-Integrated Systems, Alessio Bucaioni, Martin Weyssow, Junda He, Yunbo Lyu, David Lo
A Functional Software Reference Architecture For Llm-Integrated Systems, Alessio Bucaioni, Martin Weyssow, Junda He, Yunbo Lyu, David Lo
Research Collection School Of Computing and Information Systems
The integration of large language models into software systems is transforming capabilities such as natural language understanding, decision-making, and autonomous task execution. However, the absence of a commonly accepted software reference architecture hinders systematic reasoning about their design and quality attributes. This gap makes it challenging to address critical concerns like privacy, security, modularity, and interoperability, which are increasingly important as these systems grow in complexity and societal impact. In this paper, we describe our emerging results for a preliminary functional reference architecture as a conceptual framework to address these challenges and guide the design, evaluation, and evolution of large …
Bigcodebench: Benchmarking Code Generation With Diverse Function Calls And Complex Instructions, T.Y. Zhuo, M.C. Vu, J. Chim, ..., David Lo
Bigcodebench: Benchmarking Code Generation With Diverse Function Calls And Complex Instructions, T.Y. Zhuo, M.C. Vu, J. Chim, ..., David Lo
Research Collection School Of Computing and Information Systems
Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks ranging from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing diverse function calls as tools to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately …
Characterising Reproducibility Debt In Scientific Software: A Systematic Literature Review, Zara Hassan, Christoph Treude, Michael Norrish, Graham Williams, Alex Potanin
Characterising Reproducibility Debt In Scientific Software: A Systematic Literature Review, Zara Hassan, Christoph Treude, Michael Norrish, Graham Williams, Alex Potanin
Research Collection School Of Computing and Information Systems
Context: In scientific software, the inability to reproduce results is often due to technical issues and challenges in recreating the full computational workflow from the original analysis. We conceptualise this problem as Reproducibility Debt (RpD). Much research has been performed to propose solutions to tackle these issues across various computational science disciplines. It is essential to identify and accumulate existing knowledge on reproducibility issues and state-of-the-art solutions so as to provide researchers and practitioners with information that enables further research activities and RpD management in practice. Objective: In the context of scientific software, we aim to characterise RpD by providing …
Prioritizing Speech Test Cases, Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, Bowen Xu, Xin Zhou, Donggyun Han, David Lo
Prioritizing Speech Test Cases, Zhou Yang, Jieke Shi, Muhammad Hilmi Asyrofi, Bowen Xu, Xin Zhou, Donggyun Han, David Lo
Research Collection School Of Computing and Information Systems
As Automated Speech Recognition (ASR) systems gain widespread acceptance, there is a pressing need to rigorously test and enhance their performance. Nonetheless, the process of collecting and executing speech test cases is typically both costly and time-consuming. This presents a compelling case for the strategic prioritization of speech test cases, which consist of a piece of audio and the corresponding reference text. The central question we address is: In what sequence should speech test cases be collected and executed to identify the maximum number of errors at the earliest stage? In this study, we introduce PRiOritizing sPeecH tEsT …
Neurovig: Integrating Event Cameras For Resource-Efficient Video Grounding, Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo Hwee Lim, Archan Misra
Neurovig: Integrating Event Cameras For Resource-Efficient Video Grounding, Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo Hwee Lim, Archan Misra
Research Collection School Of Computing and Information Systems
Spatio-Temporal Video Grounding (STVG) - the task of identifying the target object in the field-of-view that the language instruction refers to - is a fundamental vision-language task. Current STVG approaches typically utilize feeds from an RGB camera that is assumed to be always-on and process the video frames using complex neural network pipelines. As a result they often impose prohibitive system overheads (energy latency) on pervasive devices. To address this we propose NeuroViG with two key innovations: (a) leveraging on event streams from a low-power neuromorphic event camera sensor to perform selective triggering of the more energy-hungry RGB camera for …
Understanding The Oss Communities Of Deep Learning Frameworks: A Comparative Case Study Of Pytorch And Tensorflow, Yunqi Chen, Zhiyuan Wan, Yifei Zhuang, Ning Liu, David Lo, Xiaohu Yang
Understanding The Oss Communities Of Deep Learning Frameworks: A Comparative Case Study Of Pytorch And Tensorflow, Yunqi Chen, Zhiyuan Wan, Yifei Zhuang, Ning Liu, David Lo, Xiaohu Yang
Research Collection School Of Computing and Information Systems
Over the past two decades, deep learning has received tremendous success in developing software systems across various domains. Deep learning frameworks have been proposed to facilitate the development of such software systems, among which, PyTorch and TensorFlow stand out as notable examples. Considerable attention focuses on exploring software engineering practices and addressing diverse technical aspects in developing and deploying deep learning frameworks and software systems. Despite these efforts, little is known about the open source software communities involved in the development of deep learning frameworks. In this article, we perform a comparative investigation into the open source software communities of …
Cachealarm: Monitoring Sensitive Behaviors Of Android Apps Using Cache Side Channel, Jianwen Tian, Haoyu Ma, Debin Gao, Xiaohui Kuang
Cachealarm: Monitoring Sensitive Behaviors Of Android Apps Using Cache Side Channel, Jianwen Tian, Haoyu Ma, Debin Gao, Xiaohui Kuang
Research Collection School Of Computing and Information Systems
Malware attack has been a serious threat to the security and privacy of both individual and corporation users of the Android platform. Business entities seek to protect themselves by means of monitoring privacy-related sensitive behaviors conducted on company-issued Android devices. However, due to Android’s own access control and privacy protection policies, this is difficult to be done with third-party apps using only normal privileges. Existing works proposed using side-channel readings from leaky APIs and system virtual files to speculate runtime app behaviors, which could be unreliable due to future system updates (that ban exploited resources), hardware jittering, etc. In this …
Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang
Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) aims to enhance software reliability by automatically generating bug-fixing patches. Recent work has improved the state-of-the-art of APR by fine-tuning pre-trained large language models (LLMs), such as CodeT5, for APR. However, the effectiveness of fine-tuning be-comes weakened in data scarcity scenarios, and data scarcity can be a common issue in practice, limiting fine-tuning performance. To alleviate this limitation, this paper adapts prompt tuning for enhanced APR and conducts a comprehensive study to evaluate its effectiveness in data scarcity scenarios, using three LLMs of different sizes and six diverse datasets across four programming languages. Prompt tuning rewrites …
Adaptive Deviation Learning For Visual Anomaly Detection With Data Contamination, Aanindya Sundar Das, Guansong Pang, Monowar Bhuyan
Adaptive Deviation Learning For Visual Anomaly Detection With Data Contamination, Aanindya Sundar Das, Guansong Pang, Monowar Bhuyan
Research Collection School Of Computing and Information Systems
Visual anomaly detection targets to detect images that notably differ from normal pattern, and it has found extensive application in identifying defective parts within the manufacturing industry. These anomaly detection paradigms predominantly focus on training detection models using only clean, unlabeled normal samples, assuming an absence of contamination; a condition often unmet in real-world scenarios. The performance of these methods significantly depends on the quality of the data and usually decreases when exposed to noise. We introduce a systematic adaptive method that employs deviation learning to compute anomaly scores end-to-end while addressing data contamination by assigning relative importance to the …
Revisiting Sentiment Analysis For Software Engineering In The Era Of Large Language Models, Ting Zhang, Ivana Clairine Irsan, Thung Ferdian, David Lo
Revisiting Sentiment Analysis For Software Engineering In The Era Of Large Language Models, Ting Zhang, Ivana Clairine Irsan, Thung Ferdian, David Lo
Research Collection School Of Computing and Information Systems
Software development involves collaborative interactions where stakeholders express opinions across various platforms. Recognizing the sentiments conveyed in these interactions is crucial for the effective development and ongoing maintenance of software systems. For software products, analyzing the sentiment of user feedback, e.g., reviews, comments, and forum posts can provide valuable insights into user satisfaction and areas for improvement. This can guide the development of future updates and features. However, accurately identifying sentiments in software engineering datasets remains challenging.This study investigates bigger large language models (bLLMs) in addressing the labeled data shortage that hampers fine-tuned smaller large language models (sLLMs) in software …
Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang
Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
In recent years, AI-based software engineering has progressed from pre-trained models to advanced agentic workflows, with Software Development Agents representing the next major leap. These agents, capable of reasoning, planning, and interacting with external environments, offer promising solutions to complex software engineering tasks. However, while much research has evaluated code generated by large language models (LLMs), comprehensive studies on agent-generated patches, particularly in real-world settings, are lacking. This study addresses that gap by evaluating 4,892 patches from 10 top-ranked agents on 500 real-world GitHub issues from SWE-Bench Verified, focusing on their impact on code quality. Our analysis shows no single …
Density Boosts Everything: A One-Stop Strategy For Improving Performance, Robustness, And Sustainability Of Malware Detectors, Jianwen Tian, Wei Kong, Debin Gao, Tong Wang, Taotao Gu, Kefan Qiu, Zhi Wang, Xiaohui Kuang
Density Boosts Everything: A One-Stop Strategy For Improving Performance, Robustness, And Sustainability Of Malware Detectors, Jianwen Tian, Wei Kong, Debin Gao, Tong Wang, Taotao Gu, Kefan Qiu, Zhi Wang, Xiaohui Kuang
Research Collection School Of Computing and Information Systems
In the contemporary landscape of cybersecurity, AI-driven detectors have emerged as pivotal in the realm of malware detection. However, existing AI-driven detectors encounter a myriad of challenges, including poisoning attacks, evasion attacks, and concept drift, which stem from the inherent characteristics of AI methodologies. While numerous solutions have been proposed to address these issues, they often concentrate on isolated problems, neglecting the broader implications for other facets of malware detection. This paper diverges from the conventional approach by not targeting a singular issue but instead identifying one of the fundamental causes of these challenges, sparsity. Sparsity refers to a scenario …
The Role Of Surprisal In Issue Trackers, James Caddy, Christoph Treude, Markus Wagner, Earl T. Barr
The Role Of Surprisal In Issue Trackers, James Caddy, Christoph Treude, Markus Wagner, Earl T. Barr
Research Collection School Of Computing and Information Systems
Context: Software development creates and relies on a large volume of information, yet the volume of this information can make it challenging for developers to maintain an overview of all goings-on that a team and external actors contribute to a project. We posit that unexpected or “surprising” events could serve as important signposts amidst this information overload. These unexpected events may indicate underlying anomalies or emergent situations that require immediate attention. To explore this premise, our study leverages the concept of ‘surprisal’ from information theory to identify and quantify these unusual occurrences from the issues and pull requests of popular …
Bridging Expert Knowledge With Deep Learning Techniques For Just-In-Time Defect Prediction, Xin Zhou, Donggyun Han, David Lo
Bridging Expert Knowledge With Deep Learning Techniques For Just-In-Time Defect Prediction, Xin Zhou, Donggyun Han, David Lo
Research Collection School Of Computing and Information Systems
Just-In-Time (JIT) defect prediction aims to automatically predict whether a commit is defective or not, and has been widely studied in recent years. In general, most studies can be classified into two categories: 1) simple models using traditional machine learning classifiers with hand-crafted features, and 2) complex models using deep learning techniques to automatically extract features from commit contents. Hand-crafted features used by simple models are based on expert knowledge but may not fully represent the semantic meaning of the commits. On the other hand, deep learning-based features used by complex models represent the semantic meaning of commits but may …
Wf-Ppg: A Wrist-Finger Dual-Channel Dataset For Studying The Impact Of Contact Pressure On Ppg Morphology, Matthew Yiwen Ho, Hung Manh Pham, Aaqib Saeed, Dong Ma
Wf-Ppg: A Wrist-Finger Dual-Channel Dataset For Studying The Impact Of Contact Pressure On Ppg Morphology, Matthew Yiwen Ho, Hung Manh Pham, Aaqib Saeed, Dong Ma
Research Collection School Of Computing and Information Systems
Photoplethysmography (PPG) is a simple optical technique widely used in wearable devices for continuous cardiac health monitoring. However, the quality of PPG signals, particularly their morphology, is influenced by the contact pressure between the skin and the sensor. This variability in signal quality complicates complex tasks that rely on high-quality signals, such as blood pressure and heart rate variability estimation, making them less reliable or even impossible. To address this issue, we present a novel dataset (termed WF-PPG) comprising PPG signals from the wrist measured under varying contact pressures, along with high-quality PPG signals from the fingertip captured simultaneously. Data …
Learning An Interpretable Stylized Subspace For 3d-Aware Animatable Artforms, Chenxi Zheng, Bangzhen Liu, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Learning An Interpretable Stylized Subspace For 3d-Aware Animatable Artforms, Chenxi Zheng, Bangzhen Liu, Xuemiao Xu, Huaidong Zhang, Shengfeng He
Research Collection School Of Computing and Information Systems
Throughout history, static paintings have captivated viewers within display frames, yet the possibility of making these masterpieces vividly interactive remains intriguing. This research paper introduces 3DArtmator, a novel approach that aims to represent artforms in a highly interpretable stylized space, enabling 3D-aware animatable reconstruction and editing. Our rationale is to transfer the interpretability and 3D controllability of the latent space in a 3D-aware GAN to a stylized sub-space of a customized GAN, revitalizing the original artforms. To this end, the proposed two-stage optimization framework of 3DArtmator begins with discovering an anchor in the original latent space that accurately mimics the …
Towards Resource-Efficient Reactive And Proactive Auto-Scaling For Microservice Architectures, Hussain Ahmad, Christoph Treude, Markus Wagner, Claudia Szabo
Towards Resource-Efficient Reactive And Proactive Auto-Scaling For Microservice Architectures, Hussain Ahmad, Christoph Treude, Markus Wagner, Claudia Szabo
Research Collection School Of Computing and Information Systems
Microservice architectures have become increasingly popular in both academia and industry, providing enhanced agility, elasticity, and maintainability in software development and deployment. To simplify scaling operations in microservice architectures, container orchestration platforms such as Kubernetes feature Horizontal Pod Auto-scalers (HPAs) designed to adjust the resources of microservices to accommodate fluctuating workloads. However, existing HPAs are not suitable for resource-constrained environments, as they make scaling decisions based on the individual resource capacities of microservices, leading to service unavailability, resource mismanagement, and financial losses. Furthermore, the inherent delay in initializing and terminating microservice pods hinders HPAs from timely responding to workload fluctuations, …
Ptm4tag+: Tag Recommendation Of Stack Overflow Posts With Pre-Trained Models, Junda He, Bowen Xu, Zhou Yang, Donggyun Han, Chengran Yang, Jiakun Liu, Zhipeng Zhao, David Lo
Ptm4tag+: Tag Recommendation Of Stack Overflow Posts With Pre-Trained Models, Junda He, Bowen Xu, Zhou Yang, Donggyun Han, Chengran Yang, Jiakun Liu, Zhipeng Zhao, David Lo
Research Collection School Of Computing and Information Systems
Stack Overflow is one of the most influential Software Question & Answer (SQA) websites, hosting millions of programming-related questions and answers. Tags play a critical role in efficiently organizing the contents on Stack Overflow and are vital to support various site operations, such as querying relevant content. Poorly chosen tags often lead to issues such as tag ambiguity and tag explosion. Therefore, a precise and accurate automated tag recommendation technique is needed. Inspired by the recent success of pre-trained models (PTMs) in natural language processing (NLP), we present PTM4Tag+, a tag recommendation framework for Stack Overflow posts that utilize PTMs …
Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang
Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Deep Neural Network-based models have demonstrated high accuracy for semantic code clone detection. However, the lack of generalization poses a threat to the trustworthiness and reliability of these models. Furthermore, the black-box nature of these models makes interpreting the model’s decisions very challenging. Currently, there is only a limited understanding of the semantic code clone detection behavior of existing models. There is a lack of transparency in understanding how a model identifies semantic code clones and the exact code components influencing its prediction. In this paper, we introduce the use of a causal interpretation framework based on the Neyman-Rubin causal …
Demo2test: Transfer Testing Of Agent In Competitive Environment With Failure Demonstrations, Jianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie, Dandan Wang, Qing Wang, Fanjiang Xu
Demo2test: Transfer Testing Of Agent In Competitive Environment With Failure Demonstrations, Jianming Chen, Yawen Wang, Junjie Wang, Xiaofei Xie, Dandan Wang, Qing Wang, Fanjiang Xu
Research Collection School Of Computing and Information Systems
The competitive game between agents exists in many critical applications, such as military unmanned aerial vehicles. It is urgent to test these agents to reduce the significant losses caused by their failures. Existing studies mainly are to construct a testing agent that competes with the target agent to induce its failures. These approaches usually focus on a single task, requiring much more time for multi-task testing. However, if the previously tested tasks (source tasks) and the task to be tested (target task) share similar agents or task objectives, the transferable knowledge in source tasks can potentially increase the effectiveness of …
The Gender Wage Gap In An Online Labor Market: The Cost Of Interruptions, Abi Adams, Kotaro Hara, Kristy Milland, Chris Callison-Burch
The Gender Wage Gap In An Online Labor Market: The Cost Of Interruptions, Abi Adams, Kotaro Hara, Kristy Milland, Chris Callison-Burch
Research Collection School Of Computing and Information Systems
This paper analyses gender differences in working patterns and wages on Amazon Mechanical Turk, a popular online labour platform. Using information on 2 million tasks, we find no gender differences in task selection nor experience. Nonetheless, women earn 20% less per hour on average. Gender differences in working patterns are a significant driver of this wage gap. Women are more likely to interrupt their working time on the platform with consequences for their task completion speed. A follow-up survey shows that the gender differences in working patterns and hourly wages are concentrated amongst workers with children.
Neuron Semantic-Guided Test Generation For Deep Neural Networks Fuzzing, Li Huang, Weifeng Sun, Meng Yan, Zhongxin Liu, Yan Lei, David Lo
Neuron Semantic-Guided Test Generation For Deep Neural Networks Fuzzing, Li Huang, Weifeng Sun, Meng Yan, Zhongxin Liu, Yan Lei, David Lo
Research Collection School Of Computing and Information Systems
In recent years, significant progress has been made in testing methods for deep neural networks (DNNs) to ensure their correctness and robustness. Coverage-guided criteria, such as neuron-wise, layer-wise, and path-/trace-wise, have been proposed for DNN fuzzing. However, existing coverage-based criteria encounter performance bottlenecks for several reasons: Testing Adequacy: Partial neural coverage criteria have been observed to achieve full coverage using only a small number of test inputs. In this case, increasing the number of test inputs does not consistently improve the quality of models. Interpretability: The current coverage criteria lack interpretability. Consequently, testers are unable to identify and understand which …
Adapting Installation Instructions In Rapidly Evolving Software Ecosystems, Haoyu Gao, Christoph Treude, Mansooreh Zahedi
Adapting Installation Instructions In Rapidly Evolving Software Ecosystems, Haoyu Gao, Christoph Treude, Mansooreh Zahedi
Research Collection School Of Computing and Information Systems
files play an important role in providing installation-related instructions to software users and are widely used in open source software systems on platforms such as GitHub. Software projects evolve rapidly alongside their dependencies in dynamic software ecosystems, requiring frequent updates to installation instructions. These instructions are crucial for users to start with a software project. Despite their significance, there is a lack of systematic understanding regarding the documentation efforts invested in README files and the triggers behind them. To fill the research gap, we conducted a qualitative study, investigating 400 GitHub repositories with 1,163 README commits that focused on updates …
More Effective Javascript Breaking Change Detection Via Dynamic Object Relation Graph, Dezhen Kong, Jiakun Liu, Chao Ni, David Lo, Lingfeng Bao
More Effective Javascript Breaking Change Detection Via Dynamic Object Relation Graph, Dezhen Kong, Jiakun Liu, Chao Ni, David Lo, Lingfeng Bao
Research Collection School Of Computing and Information Systems
JavaScript libraries are characterized by their widespread use, frequent code changes, and a high tolerance for backward incompatible changes. Awareness of such breaking changes can help developers adapt to version updates and avoid negative impacts. Several tools have been targeted to or can be used to detect breaking change detection in the JavaScript community. However, these tools detect breaking changes using different ways, and there are currently no systematic reviews of these approaches. From a preliminary study on popular JavaScript libraries, we find that existing approaches, including simple regression testing, model-based testing and type differencing cannot detect many breaking changes …
Don’T Complete It! Preventing Unhelpful Code Completion For Productive And Sustainable Neural Code Completion Systems, Zhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang, Mingze Ni, Li Li, David Lo
Don’T Complete It! Preventing Unhelpful Code Completion For Productive And Sustainable Neural Code Completion Systems, Zhensu Sun, Xiaoning Du, Fu Song, Shangwen Wang, Mingze Ni, Li Li, David Lo
Research Collection School Of Computing and Information Systems
Currently, large pre-trained language models are widely applied in neural code completion systems. Though large code models significantly outperform their smaller counterparts, around 70% of displayed code completions from Github Copilot are not accepted by developers. Being reviewed but not accepted, their help to developer productivity is considerably limited and may conversely aggravate the workload of developers, as the code completions are automatically and actively generated in state-of-the-art code completion systems as developers type out once the service is enabled. Even worse, considering the high cost of the large code models, it is a huge waste of computing resources and …
Automated Program Refinement: Guide And Verify Code Large Language Model With Refinement Calculus, Yufan Cai, Zhe Hou, David Sanan, Xiaokun Luan, Yun Lin, Jun Sun, Jin Song Dong
Automated Program Refinement: Guide And Verify Code Large Language Model With Refinement Calculus, Yufan Cai, Zhe Hou, David Sanan, Xiaokun Luan, Yun Lin, Jun Sun, Jin Song Dong
Research Collection School Of Computing and Information Systems
Recently, the rise of code-centric large language models (LLMs) appears to have reshaped the software engineering world with low-barrier tools like Copilot that can generate code easily. However, there is no correctness guarantee for the code generated by LLMs, which suffer from the hallucination problem, and their output is fraught with risks. Besides, the end-to-end process from specification to code through LLMs is a non-transparent and uncontrolled black box. This opacity makes it difficult for users to understand and trust the generated code. Addressing these challenges is both necessary and critical. In contrast, program refinement transforms high-level specification statements into …
Performance Evaluation Of Newsql Databases In A Distributed Architecture, Zhiyao Zhang, Alan @ Ali Madjelisi Megargel, Lingxiao Jiang
Performance Evaluation Of Newsql Databases In A Distributed Architecture, Zhiyao Zhang, Alan @ Ali Madjelisi Megargel, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
In the last decade, application architectures have evolved drastically, moving from monolithic architectures to distributed architectures where deployment has shifted from dedicated on-premises servers to the cloud. Distributed architectures and cloud computing has enabled businesses to scale their application components across different geographical locations. While it is easy to scale the application layer, scaling its database layer that relies on traditional SQL databases is challenging and often is a common source of bottlenecks when it comes to application performance. This paper evaluates the performance characteristics between two NewSQL databases solutions, MySQL NDB Cluster vs. TIBCO ActiveSpaces IMDG. Serving as an …
Triadic Temporal-Semantic Alignment For Weakly-Supervised Video Moment Retrieval, Jin Liu, Jialong Xie, Fengyu Zhou, Shengfeng He
Triadic Temporal-Semantic Alignment For Weakly-Supervised Video Moment Retrieval, Jin Liu, Jialong Xie, Fengyu Zhou, Shengfeng He
Research Collection School Of Computing and Information Systems
Video Moment Retrieval (VMR) aims to identify specific event moments within untrimmed videos based on natural language queries. Existing VMR methods have been criticized for relying heavily on moment annotation bias rather than true multi-modal alignment reasoning. Weakly supervised VMR approaches inherently overcome this issue by training without precise temporal location information. However, they struggle with fine-grained semantic alignment and often yield multiple speculative predictions with prolonged video spans. In this paper, we take a step forward in the context of weakly supervised VMR by proposing a triadic temporalsemantic alignment model. Our proposed approach augments weak supervision by comprehensively addressing …
Ali-Agent: Assessing Llms’ Alignment With Human Values Via Agent-Based Evaluation, Jingnan Zheng, Han Wang, Tai D. Nguyen, An Zhang, Jun Sun, Tat-Seng Chua
Ali-Agent: Assessing Llms’ Alignment With Human Values Via Agent-Based Evaluation, Jingnan Zheng, Han Wang, Tai D. Nguyen, An Zhang, Jun Sun, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Large Language Models (LLMs) can elicit unintended and even harmful content when misaligned with human values, posing severe risks to users and society. To mitigate these risks, current evaluation benchmarks predominantly employ expertdesigned contextual scenarios to assess how well LLMs align with human values. However, the labor-intensive nature of these benchmarks limits their test scope, hindering their ability to generalize to the extensive variety of open-world use cases and identify rare but crucial long-tail risks. Additionally, these static tests fail to adapt to the rapid evolution of LLMs, making it hard to evaluate timely alignment issues. To address these challenges, …