Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (4208)
- Computer Sciences (4205)
- Databases and Information Systems (1558)
- Software Engineering (1251)
- Artificial Intelligence and Robotics (743)
-
- Information Security (727)
- Graphics and Human Computer Interfaces (489)
- Numerical Analysis and Scientific Computing (454)
- Engineering (344)
- Programming Languages and Compilers (269)
- Computer Engineering (231)
- Social and Behavioral Sciences (213)
- Theory and Algorithms (212)
- Business (202)
- OS and Networks (148)
- Communication (103)
- Operations Research, Systems Engineering and Industrial Engineering (101)
- Social Media (83)
- Data Storage Systems (80)
- Medicine and Health Sciences (79)
- Education (65)
- Public Affairs, Public Policy and Public Administration (54)
- Computer and Systems Architecture (52)
- Health Information Technology (51)
- Transportation (50)
- Management Information Systems (49)
- Digital Communications and Networking (47)
- E-Commerce (44)
- Finance and Financial Management (40)
- Higher Education (30)
- Keyword
-
- Deep learning (71)
- Machine learning (59)
- Privacy (41)
- Empirical study (40)
- Deep Learning (35)
-
- Training (34)
- Security (33)
- Task analysis (32)
- Cloud computing (31)
- Data mining (31)
- Artificial intelligence (29)
- Reinforcement learning (29)
- Social media (29)
- Software engineering (29)
- Neural networks (28)
- Android (27)
- Blockchain (27)
- Feature extraction (27)
- Model Check (26)
- Access control (25)
- Large Language Models (25)
- Natural language processing (25)
- Authentication (24)
- Privacy-preserving (24)
- Servers (24)
- Anomaly detection (23)
- Information retrieval (23)
- Large language models (23)
- Visualization (23)
- Classification (22)
- Publication Year
Articles 151 - 180 of 4258
Full-Text Articles in Entire DC Network
Large Language Models As End-To-End Combinatorial Optimization Solvers, Xia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao, Yingqian Zhang
Large Language Models As End-To-End Combinatorial Optimization Solvers, Xia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao, Yingqian Zhang
Research Collection School Of Computing and Information Systems
Combinatorial optimization (CO) problems, central to decision-making scenarios like logistics and manufacturing, are traditionally solved using problem-specific algorithms requiring significant domain expertise. While large language models (LLMs) have shown promise in automating CO problem solving, existing approaches rely on intermediate steps such as code generation or solver invocation, limiting their generality and accessibility. This paper introduces a novel framework that empowers LLMs to serve as end-to-end CO solvers by directly mapping natural language problem descriptions to solutions. We propose a two-stage training strategy: supervised fine-tuning (SFT) imparts LLMs with solution generation patterns from domain-specific solvers, while a feasibility-and-optimality-aware reinforcement learning …
Agentguard: An Active Threat Discovery System For Package Confusion Using Multi-Agent Collaboration, Wei Ma, Yu Li, Zhi Chen, Ye Liu, Lingxiao Jiang, Qiang Hu, Junyi Tao
Agentguard: An Active Threat Discovery System For Package Confusion Using Multi-Agent Collaboration, Wei Ma, Yu Li, Zhi Chen, Ye Liu, Lingxiao Jiang, Qiang Hu, Junyi Tao
Research Collection School Of Computing and Information Systems
The proliferation of open-source software (OSS) has made software supply chains prime targets for attacks like Package Confusion, where adversaries publish malicious packages with names deceptively similar to legitimate ones. Existing detection methods often rely on simple lexical similarity or passive analysis of known package pairs, struggle with high false positive rates (FPR), fail to proactively identify emerging threats, and are vulnerable to adversarial evasion. To overcome these limitations, we introduce AgentGuard, a novel framework for proactive, single-input package confusion detection. AgentGuard employs a multi-agent architecture that autonomously discovers potential confusion targets using fine-tuned word embedding model to hybird semantic …
Accuracy-Enabling Differential Privacy-Preserving Truth Discovery, Man Zhang, Xinghua Li, Yinbin Miao, Bin Luo, Siqi Ma, Robert H. Deng
Accuracy-Enabling Differential Privacy-Preserving Truth Discovery, Man Zhang, Xinghua Li, Yinbin Miao, Bin Luo, Siqi Ma, Robert H. Deng
Research Collection School Of Computing and Information Systems
Perturbation-based privacy-preserving truth discovery requires the Service Provider (SP) to calculate the truthful aggregation result from perturbed data of the Data Sources (DSs), which inevitably damages the aggregation accuracy due to perturbation noise added in the data. Thus, the existing works attempt to relieve the perturbation errors by reducing noise amounts or adjusting aggregation weights of DSs. However, the former sacrifices DSs’ privacy preservation and the latter has the limited accuracy recovery performance. Aiming at it, we propose an accuracy-enabling differential privacy-preserving truth discovery consisting of an independence-guaranteed data perturbation module and a progressive-private noise elimination module. Specifically, in the …
Fl-Cdf: Collaborative Defense Framework For Backdoor Mitigation In Federated Learning, Haiyan Zhang, Xinghua Li, Yinbin Miao, Shunjie Yuan, Mengyao Zhu, Ximeng Liu, Robert H. Deng
Fl-Cdf: Collaborative Defense Framework For Backdoor Mitigation In Federated Learning, Haiyan Zhang, Xinghua Li, Yinbin Miao, Shunjie Yuan, Mengyao Zhu, Ximeng Liu, Robert H. Deng
Research Collection School Of Computing and Information Systems
Federated learning (FL) is vulnerable to backdoor attacks due to its distributed nature. Existing unilateral defense mechanisms often fail against persistent attack strategies, primarily due to their limited perspectives. To address the challenge of model misclassification on the server side caused by overlooked model similarity drift, and gradient misjudgment on the client side caused by semantic learning imbalances across classes, this paper proposes a collaborative defense framework for federated learning, termed FL-CDF. FL-CDF establishes an end-to-end defense through a bidirectional client-server collaboration mechanism. Specifically: (1) On the client side, an adversarial perturbation-based malicious neuron detection module is introduced. This module …
Mando-Llm: Heterogeneous Graph Transformers With Large Language Models For Smart Contract Vulnerability Detection, Nhat Minh Nguyen, Huu Hoang Nguyen, Long Le Thanh, Zahra Ahmadi, Thanh Nam Doan, Daoyuan Wu, Lingxiao Jiang
Mando-Llm: Heterogeneous Graph Transformers With Large Language Models For Smart Contract Vulnerability Detection, Nhat Minh Nguyen, Huu Hoang Nguyen, Long Le Thanh, Zahra Ahmadi, Thanh Nam Doan, Daoyuan Wu, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Detecting vulnerabilities in smart contracts is vital for the security and reliability of decentralized apps. To facilitate vulnerability detection, contract codes, including bug patterns, are represented as heterogeneous graphs with various nodes and edges, like control-flow and function-call graphs. However, existing graph learning techniques struggle with large, complex graphs. This paper presents MANDO-LLM, a novel framework that combines heterogeneous graph transformers (HGTs) with large language models (LLMs) for detecting vulnerabilities in smart contracts represented as heterogeneous contract graphs built upon control-flow and call graphs. MANDO-LLM uses LLMs to capture code features from control-flow and call data, customizes HGTs to learn …
Iostom: Offline Imitation Learning From Observations Via State Transition Occupancy Matching, Quang Anh Pham, Brahmanage Janaka Chathuranga Thilakarathna, Tien Mai, Akshat Kumar
Iostom: Offline Imitation Learning From Observations Via State Transition Occupancy Matching, Quang Anh Pham, Brahmanage Janaka Chathuranga Thilakarathna, Tien Mai, Akshat Kumar
Research Collection School Of Computing and Information Systems
Offline Learning from Observations (LfO) focuses on enabling agents to imitate expert behavior using datasets that contain only expert state trajectories and separate transition data with suboptimal actions. This setting is both practical and critical in real-world scenarios where direct environment interaction or access to expert action labels is costly, risky, or infeasible. Most existing LfO methods attempt to solve this problem through state or state-action occupancy matching. They typically rely on pretraining a discriminator to differentiate between expert and non-expert states, which could introduce errors and instability—especially when the discriminator is poorly trained. While recent discriminator-free methods have emerged, …
No Experts, No Problem: Avoidance Learning From Bad Demonstrations, Minh Huy Hoang, Tien Mai, Pradeep Varakantham
No Experts, No Problem: Avoidance Learning From Bad Demonstrations, Minh Huy Hoang, Tien Mai, Pradeep Varakantham
Research Collection School Of Computing and Information Systems
This paper addresses the problem of learning avoidance behavior within the context of offline imitation learning. In contrast to conventional methodologies that prioritize the replication of expert or near-expert demonstrations, our work investigates a setting where expert (or desirable) data is absent, and the objective is to learn to eschew undesirable actions by leveraging demonstrations of such behavior (i.e., learning from negative examples).To address this challenge, we propose a novel training objective grounded in the maximum entropy principle. We further characterize the fundamental properties of this objective function, reformulating the learning process as a cooperative inverse Q-learning task. Moreover, we …
Rising From Ashes: Generalized Federated Learning Via Dynamic Parameter Reset, Jiahao Wu, Ming Hu, Yanxin Yang, Xiaofei Xie, Zekai Chen, Chenyu Song, Mingsong Chen
Rising From Ashes: Generalized Federated Learning Via Dynamic Parameter Reset, Jiahao Wu, Ming Hu, Yanxin Yang, Xiaofei Xie, Zekai Chen, Chenyu Song, Mingsong Chen
Research Collection School Of Computing and Information Systems
Although Federated Learning (FL) is promising for privacy-preserving collaborative model training, it suffers from low inference performance due to heterogeneous client data. Due to heterogeneous data across clients, FL training easily learns client-specific overfitting features. Existing FL methods adopt coarsegrained averaging, which can easily cause the global model to get stuck in local optima, leading to poor generalization. Specifically, this paper presents a novel FL framework, FedPhoenix, to address this issue. It stochastically resets partial parameters in each round to destroy some features of the global model, guiding FL training to learn multiple generalized features for inference rather than specific …
Generalization Bounds For Rank‑Sparse Neural Networks, Antoine Ledent, Rodrigo Alves, Yunwen Lei
Generalization Bounds For Rank‑Sparse Neural Networks, Antoine Ledent, Rodrigo Alves, Yunwen Lei
Research Collection School Of Computing and Information Systems
It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of approximately low rank. In fact, the rank of the activations of each layer converges to a fixed value referred to as the “bottleneck rank”, which is the minimum rank required to represent the training data. This perspective is in line with the observation that regularizing linear networks (without activations) with weight decay is equivalent to minimizing the Schatten p quasi norm of the neural …
Design Principles For Customer‑Engaging Digital Service Systems: An Action Research Study, Keng Siau, Xiaofeng Chen, Xin Tan
Design Principles For Customer‑Engaging Digital Service Systems: An Action Research Study, Keng Siau, Xiaofeng Chen, Xin Tan
Research Collection School Of Computing and Information Systems
Digital services represent a business approach employed by organizations to operate in the digital environment. However, systematic development guidelines for developing quality digital service systems are lacking in the literature. The authors identified four general challenges for developing and implementing customer-engaging digital service systems (CEDSS). By employing the method of canonical action research in a digital service system project, they derived 10 design principles for developing high-quality CEDSS. They empirically evaluated the design principles in the development project and through follow-up focus group sessions. The design principles provide applicable and actionable guidelines for the development of CEDSS.
Bias Testing And Mitigation In Llm-Based Code Generation, Dong Huang, Jie M. Zhang, Qingwen Bu, Xiaofei Xie, Junjie Chen, Heming Cui
Bias Testing And Mitigation In Llm-Based Code Generation, Dong Huang, Jie M. Zhang, Qingwen Bu, Xiaofei Xie, Junjie Chen, Heming Cui
Research Collection School Of Computing and Information Systems
As the adoption of LLMs becomes more widespread in software coding ecosystems, a pressing issue has emerged: does the generated code contain social bias and unfairness, such as those related to age, gender, and race? This issue concerns the integrity, fairness, and ethical foundation of software applications that depend on the code generated by these models but are underexplored in the literature. This paper presents a novel bias testing framework that is specifically designed for code generation tasks. Based on this framework, we conduct an extensive empirical study on the biases in code generated by five widely studied LLMs (i.e., …
Genscore: Agent-Based Short-Answer Question Generation And Scoring In Software Engineering Courses, Nguyen Binh Duong Ta, Lwin Khin Shar
Genscore: Agent-Based Short-Answer Question Generation And Scoring In Software Engineering Courses, Nguyen Binh Duong Ta, Lwin Khin Shar
Research Collection School Of Computing and Information Systems
Short-answer questions are commonly used in educational assessments, as they are often viewed as a more effective way than multiple-choice questions to determine whether students have achieved the intended learning outcomes. However, manually creating appropriate questions targeting different cognitive levels such as those defined by the Bloom’s Taxonomy, and grading text answers from students are not trivial tasks for instructors. Existing work on auto-question generation and scoring in computing education typically targets coding-based questions. However, in software engineering courses, assessments can extend beyond coding to understanding of processes, DevOps methodologies, system design, etc. This work aims to address the dual …
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
When Deep Learning Meets Information Retrieval-Based Bug Localization: A Survey, Feifei Niu, Chuanyi Li, Kui Liu, Xin Xia, David Lo
Research Collection School Of Computing and Information Systems
Bug localization is a crucial aspect of software maintenance, running through the entire software lifecycle. Information retrieval-based bug localization (IRBL) identifies buggy code based on bug reports, expediting the bug resolution process for developers. Recent years have witnessed significant achievements in IRBL, propelled by the widespread adoption of deep learning (DL). To provide a comprehensive overview of the current state of the art and delve into key issues, we conduct a survey encompassing 61 IRBL studies leveraging DL. We summarize best practices in each phase of the IRBL workflow, undertake a meta-analysis of prior studies, and suggest future research directions. …
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Do Code Semantics Help? A Comprehensive Study On Execution Trace-Based Information For Code Large Language Models, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Yi Li
Research Collection School Of Computing and Information Systems
Code Large Language Models (Code LLMs) have opened a new era in programming with their impressive capabilities. However, recent research has revealed critical limitations in their ability to reason about runtime behavior and understand the actual functionality of programs, which poses significant challenges for their post-training and practical deployment. Specifically, Code LLMs encounter two principal issues: (1) a lack of proficiency in reasoning about program execution behavior, as they struggle to interpret what programs actually do during runtime, and (2) inconsistent and fragmented representation of semantic information, such as execution traces, across existing methods, which hinders their ability to generalize …
Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang
Explainable Sentiment Analysis With Deepseek-R1: Performance, Efficiency, And Few-Shot Learning, Donghao Huang, Zhaoxia Wang
Research Collection School Of Computing and Information Systems
Large language models (LLMs) have transformed sentiment analysis, yet balancing accuracy, efficiency, and explainability remains a critical challenge. This study presents the first comprehensive evaluation of DeepSeek-R1—an open-source reasoning model—against OpenAI’s GPT-4o and GPT-4o-mini. We test the full 671B model and its distilled variants, systematically documenting few-shot learning curves. Our experiments show DeepSeek-R1 achieves a 91.39% F1 score on 5-class sentiment and 99.31% accuracy on binary tasks with just 5 shots, an eightfold improvement in few-shot efficiency over GPT-4o. Architecture-specific distillation effects emerge, where a 32B Qwen2.5-based model outperforms the 70B Llama-based variant by 6.69 percentage points. While its reasoning …
Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang
Efficient Integration Of External Knowledge To Llm-Based World Models Via Retrieval-Augmented Generation And Reinforcement Learning, Chang Yang, Xinrun Wang, Qinggang Zhang, Qi Jiang, Xiao Huang
Research Collection School Of Computing and Information Systems
World models achieve remarkable success in predicting future states and planning in complex environments and Large Language Models (LLMs) serve as promising foundation to build general world models. However, their performances are usually constrained by the limited external knowledge to specific environments. Existing research attempts to enhance LLM-based world models through prompting or fine-tuning approaches, which are either requiring human knowledge or computationally extensive. Therefore, we introduce Retrieval-Augmented World Models (RAWM), a novel framework that leverages retrieval-augmented generation to efficiently integrate the external knowledge to LLM-based world models. Our main contributions are threefold: (i) We introduce a memory system and …
Sustainable Llm Inference For Edge Ai: Evaluating Quantized Llms For Energy Efficiency, Output Accuracy, And Inference Latency, Erik Johanne Husom, Arda Goknil, Merve Astekin, Lwin Khin Shar, Andre Kasen, Sagar Sen, Benedikt Andreas Mithassel, Ahmet Soylu
Sustainable Llm Inference For Edge Ai: Evaluating Quantized Llms For Energy Efficiency, Output Accuracy, And Inference Latency, Erik Johanne Husom, Arda Goknil, Merve Astekin, Lwin Khin Shar, Andre Kasen, Sagar Sen, Benedikt Andreas Mithassel, Ahmet Soylu
Research Collection School Of Computing and Information Systems
Deploying Large Language Models (LLMs) on edge devices presents significant challenges due to computational constraints, memory limitations, inference speed, and energy consumption. Model quantization has emerged as a key technique to enable efficient LLM inference by reducing model size and computational overhead. In this study, we conduct a comprehensive analysis of 28 quantized LLMs from the Ollama library, which applies by default Post-Training Quantization (PTQ) and weight-only quantization techniques, deployed on an edge device (Raspberry Pi 4 with 4GB RAM). We evaluate energy efficiency, inference performance, and output accuracy across multiple quantization levels and task types. Models are benchmarked on …
Enhancing Spatial Understanding In Mixed-Reality Presentations, Nam-Dang Vo, Van-Vinh Thai, Nam-Hoi Do, Viet-Tham Huynh, Anthony Tang, Khan-Duy Le
Enhancing Spatial Understanding In Mixed-Reality Presentations, Nam-Dang Vo, Van-Vinh Thai, Nam-Hoi Do, Viet-Tham Huynh, Anthony Tang, Khan-Duy Le
Research Collection School Of Computing and Information Systems
Mixed reality (MR) presentations often involve a presenter wearing a head-mounted display (HMD) and an audience watching via a large display, making it difficult for audiences to perceive spatial relationships between the presenter and virtual objects. We report two experiments testing three design variations: (1) scene camera placement (audience-aligned vs. opposite), (2) overlaying the presenter’s first-person view, and (3) highlighting objects in the presenter’s view. Results show that audience-aligned cameras and object highlighting improve spatial understanding, while combining third- and first-person views can further aid perception. We derive design guidelines for configuring MR presentations to better support audience comprehension.
Disc: Decentralized Identity System With Self-Sovereign Credential Aggregation, Yang Yang, Wai Keung Ching, Minming Huang, Supachate Innet, Guomin Yang, Hwee Hwa Pang, Robert H. Deng
Disc: Decentralized Identity System With Self-Sovereign Credential Aggregation, Yang Yang, Wai Keung Ching, Minming Huang, Supachate Innet, Guomin Yang, Hwee Hwa Pang, Robert H. Deng
Research Collection School Of Computing and Information Systems
The evolution of decentralized identity (DID) and self-sovereign identity (SSI) frameworks, as endorsed by W3C Verifiable Credentials (VC) and eIDAS 2.0, underscores the need for secure, efficient, and privacy-preserving credential management. However, existing credential systems often depend on centralized issuers, lack efficient aggregation mechanisms, or fail to ensure unlinkability across authentication sessions. To address these challenges, we propose DISC (Decentralized Identity System with Self-Sovereign Credential Aggregation), a novel credential system that enables multi-authority credential issuance, user-controlled credential aggregation, and unlinkable authentication. DISC allows users to aggregate credentials from multiple issuers while maintaining constant-size authentication tokens and supporting batch verification for …
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Seeing Culture: A Benchmark For Visual Reasoning And Grounding, Burak Satar, Zhixin Ma, Patrick Amadeus Irrawan, Wilfried Ariel Mulyawan, Jing Jiang, Ee-Peng Lim, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Multimodal vision-language models (VLMs) have made substantial progress in various tasks that require a combined understanding of visual and textual content, particularly in cultural understanding tasks, with the emergence of new cultural datasets. However, these datasets frequently fall short of providing cultural reasoning while underrepresenting many cultures.In this paper, we introduce the Seeing Culture Benchmark (SCB), focusing on cultural reasoning with a novel approach that requires VLMs to reason on culturally rich images in two stages: i) selecting the correct visual option with multiple-choice visual question answering (VQA), and ii) segmenting the relevant cultural artifact as evidence of reasoning. Visual …
Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen
Seeing Is Fixing: Cross-Modal Reasoning With Multimodal Llms For Visual Software Issue Fixing, Kai Huang, Jian Zhang, Xiaofei Xie, Chunyang Chen
Research Collection School Of Computing and Information Systems
Large language model (LLM)-based automated program repair (APR) techniques have shown promising results in resolving real-world github issue tasks. Existing APR systems are primarily evaluated in unimodal settings (e.g., SWE-bench), relying solely on textual issue descriptions and source code. However, these autonomous systems struggle to resolve multimodal problem scenarios (e.g., SWE-bench M) due to limitations in interpreting and leveraging visual information. In multimodal scenarios, LLMs need to rely on visual information in the graphical user interface (GUI) to understand bugs and generate fixes. To bridge this gap, we propose GUIRepair, a cross-modal reasoning approach for resolving multimodal issue scenarios by …
Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Adasteer: Your Aligned Llm Is Inherently An Adaptive Jailbreak Defender, Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Research Collection School Of Computing and Information Systems
Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. To address this, we propose AdaSteer, an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. We identify two key properties: Rejection Law (R-Law), which shows that stronger steering is needed for jailbreak inputs opposing the rejection direction, and Harmfulness Law (H-Law), which differentiates adversarial and benign inputs. AdaSteer steers input representations along both the Rejection Direction …
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Chain Of Strategy Optimization Makes Large Language Models Better Emotional Supporter, Weixiang Zhao, Xingyu Sui, Xinyang Han, Yang Deng, Yulin Hu, Jiahe Guo, Libo Qin, Qianyun Du, Shijin Wang, Yanyan Zhao, Bing Qin, Ting Liu
Research Collection School Of Computing and Information Systems
The growing emotional stress in modern society has increased the demand for Emotional Support Conversations (ESC). While Large Language Models (LLMs) show promise for ESC, they face two key challenges: (1) low strategy selection accuracy, and (2) preference bias, limiting their adaptability to users’ emotional needs. Existing supervised fine-tuning (SFT) struggles to address these issues, as it rigidly trains models on single gold-standard responses without modeling nuanced strategy trade-offs. To overcome these limitations, we propose a novel two-stage framework that optimizes strategy selection preferences at each dialogue turn. We first leverage Monte Carlo Tree Search to construct ESC-Pro, a high-quality …
Exploring Autonomous Agents: A Closer Look At Why They Fail When Completing Tasks, Ruofan Lu, Yichen Li, Yintong Huo
Exploring Autonomous Agents: A Closer Look At Why They Fail When Completing Tasks, Ruofan Lu, Yichen Li, Yintong Huo
Research Collection School Of Computing and Information Systems
Autonomous agent systems powered by Large Language Models (LLMs) have demonstrated promising capabilities in automating complex tasks. However, current evaluations largely rely on success rates without systematically analyzing the interactions, communication mechanisms, and failure causes within these systems. To bridge this gap, we present a benchmark of 34 representative programmable tasks designed to rigorously assess autonomous agents. Using this benchmark, we evaluate three popular open-source agent frameworks combined with two LLM backbones, observing a task completion rate of approximately 50%. Through in-depth failure analysis, we develop a three-tier taxonomy of failure causes aligned with task phases, highlighting planning errors, task …
Envisioning Future Interactive Web Development: Editing Webpage With Natural Language, Truong Hai Dang, Jingyu Xiao, Yintong Huo
Envisioning Future Interactive Web Development: Editing Webpage With Natural Language, Truong Hai Dang, Jingyu Xiao, Yintong Huo
Research Collection School Of Computing and Information Systems
The evolution of web applications relies on iterative code modifications, a process that is traditionally manual and time-consuming. While Large Language Models (LLMs) can generate UI code, their ability to edit existing code from new design requirements (e.g., ”center the logo”) remains a challenge. This is largely due to the absence of large-scale, high-quality tuning data to align model performance with human expectations. In this paper, we introduce a novel, automated data generation pipeline that uses LLMs to synthesize a high-quality fine-tuning dataset for web editing, named Instruct4Edit. Our approach generates diverse instructions, applies the corresponding code modifications, and performs …
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Context-Aware Hierarchical Taxonomy Generation For Scientific Papers Via Llm-Guided Multi-Aspect Clustering, Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
Research Collection School Of Computing and Information Systems
The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering. Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy. In addition, we introduce a new evaluation benchmark of 156 expert-crafted taxonomies encompassing 11.6k …
Why Stop At One Error? Benchmarking Llms As Data Science Code Debuggers For Multi-Hop And Multi-Bug Errors, Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng
Why Stop At One Error? Benchmarking Llms As Data Science Code Debuggers For Multi-Hop And Multi-Bug Errors, Zhiyu Yang, Shuo Wang, Yukun Yan, Yang Deng
Research Collection School Of Computing and Information Systems
LLMs are transforming software development, yet current code generation and code repair benchmarks mainly assess syntactic and functional correctness in simple, single-error cases. LLMs’ capabilities to autonomously find and fix runtime logical errors in complex data science code remain largely unexplored. To address this gap, we introduce DSDBench: the Data Science Debugging Benchmark, the first benchmark for systematic evaluation of LLMs on multi-hop error tracing and multi-bug detection in data science code debugging. DSDBench adapts datasets from existing data science task benchmarks, such as DABench and MatPlotBench, featuring realistic data science debugging tasks with automatically synthesized multi-hop, multi-bug code snippets. …
Generative Ai And Empirical Software Engineering: A Paradigm Shift, Christoph Treude, Margaret-Anne Storey
Generative Ai And Empirical Software Engineering: A Paradigm Shift, Christoph Treude, Margaret-Anne Storey
Research Collection School Of Computing and Information Systems
The widespread adoption of generative AI in software engineering marks a paradigm shift, offering new opportunities to design and utilize software engineering tools while influencing both developers and the artifacts they create. Traditional empirical methods in software engineering, including quantitative, qualitative, and mixed-method approaches, are well established. However, this paradigm shift introduces novel data types and redefines many concepts in the software engineering process. The roles of developers, users, agents, and researchers increasingly overlap, blurring the distinctions between these social and technical actors within the field. This paper examines how integrating AI into software engineering challenges traditional research paradigms. It …
Defects4c: Benchmarking Large Language Model Repair Capability With C/C++ Bugs, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Jiongchi Yu, Jiaolong Kong, Yi Li
Defects4c: Benchmarking Large Language Model Repair Capability With C/C++ Bugs, Jian Wang, Xiaofei Xie, Qiang Hu, Shangqing Liu, Jiongchi Yu, Jiaolong Kong, Yi Li
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) plays a critical role in enhancing the quality and reliability of software systems. While substantial progress has been made in Java-based APR, largely facilitated by benchmarks like Defects4J, there remains a significant gap in research on C/C++ program repair, despite the widespread use of C/C++ and the prevalence of associated vulnerabilities. This gap is primarily due to the lack of high-quality, open-source benchmarks tailored for C/C++. To address this issue, we introduce Defects4C, a comprehensive and executable benchmark specifically designed for C/C++ program repair. Our dataset is constructed from real-world C/C++ repositories and includes a large …
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Mmlu-Prox: A Multilingual Benchmark For Advanced Large Language Model Evaluation, Weihao Xuan, Et. Al.
Research Collection School Of Computing and Information Systems
Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack parallel questions that specifically assess cross-lingual reasoning abilities. This dual limitation makes it challenging to assess LLMs’ performance in the multilingual setting comprehensively. To fill this gap, we introduce MMLU-ProX, a comprehensive benchmark covering 29 languages, built on an English benchmark. Each language version consists of 11,829 identical questions, enabling direct cross-lingual comparisons. Additionally, to meet efficient evaluation needs, we provide a lite version containing 658 questions per language. To ensure the high quality of MMLU-ProX, we employ a rigorous development process that involves …