Open Access. Powered by Scholars. Published by Universities.®

Digital Commons Network™

Open Access. Powered by Scholars. Published by Universities.®

Research Collection School Of Computing and Information Systems

Discipline
Keyword
Publication Year

Articles 151 - 180 of 1047

Full-Text Articles in Entire DC Network

Exploring The Potential Of Large Language Models For Heterophilic Graphs, Yuxia Wu, Shujie Li, Yuan Fang, Chuan Shi May 2025

Exploring The Potential Of Large Language Models For Heterophilic Graphs, Yuxia Wu, Shujie Li, Yuan Fang, Chuan Shi

Research Collection School Of Computing and Information Systems

Large language models (LLMs) have presented significant opportunities to enhance various machine learning applications, including graph neural networks (GNNs). By leveraging the vast open-world knowledge within LLMs, we can more effectively interpret and utilize textual data to better characterize heterophilic graphs, where neighboring nodes often have different labels. However, existing approaches for heterophilic graphs overlook the rich textual data associated with nodes, which could unlock deeper insights into their heterophilic contexts. In this work, we explore the potential of LLMs for modeling heterophilic graphs and propose a novel two-stage framework: LLM-enhanced edge discriminator and LLM-guided edge reweighting. In the first …


Unlocking The Planning Capabilities Of Llms Through Maximum Diversity Fine-Tuning, Wenjun Li, Changyu Chen, Pradeep Varakantham May 2025

Unlocking The Planning Capabilities Of Llms Through Maximum Diversity Fine-Tuning, Wenjun Li, Changyu Chen, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

Large language models (LLMs) have demonstrated impressive task-solving capabilities through prompting techniques and system designs, including solving planning tasks (e.g., math proofs, basic travel planning) when sufficient data is available online and used during pre-training. However, for planning tasks with limited prior data (e.g., blocks world, advanced travel planning), the performance of LLMs, including proprietary models like GPT and Gemini, is poor. This paper investigates the impact of fine-tuning on the planning capabilities of LLMs, revealing that LLMs can achieve strong performance in planning through substantial (tens of thousands of specific examples) fine-tuning. Yet, this process incurs high economic, time, …


Reverse Modeling In Large Language Models, Sicheng Yu, Yuanchen Xu, Cunxiao Du, Yanying Zhou, Minghui Qiu, Qianru Sun, Hao Zhang, Jiawei Wu May 2025

Reverse Modeling In Large Language Models, Sicheng Yu, Yuanchen Xu, Cunxiao Du, Yanying Zhou, Minghui Qiu, Qianru Sun, Hao Zhang, Jiawei Wu

Research Collection School Of Computing and Information Systems

Humans are accustomed to reading and writing in a forward manner, and this natural bias extends to text understanding in auto-regressive large language models (LLMs). This paper investigates whether LLMs, like humans, struggle with reverse modeling, specifically with reversed text inputs. We found that publicly available pre-trained LLMs cannot understand such inputs. However, LLMs trained from scratch with both forward and reverse texts can understand them equally well during inference. Our case study shows that different-content texts result in different losses if input (to LLMs) in different directions---some get lower losses for forward while some for reverse. This leads us …


Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-Based Benchmark, Han Zhang, Zixiang Meng, Meng Luo, Hong Han, Lizi Liao, Erik Cambria, Hao Fei May 2025

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-Based Benchmark, Han Zhang, Zixiang Meng, Meng Luo, Hong Han, Lizi Liao, Erik Cambria, Hao Fei

Research Collection School Of Computing and Information Systems

Empathetic Response Generation (ERG) is one of the key tasks of the affective computing area, which aims to produce emotionally nuanced and compassionate responses to user's queries. However, existing ERG research is predominantly confined to the singleton text modality, limiting its effectiveness since human emotions are inherently conveyed through multiple modalities. To combat this, we introduce an avatar-based Multimodal ERG (MERG) task, entailing rich text, speech, and facial vision information. We first present a large-scale high-quality benchmark dataset, AvaMERG, which extends traditional text ERG by incorporating authentic human speech audio and dynamic talking-face avatar videos, encompassing a diverse range of …


Seaexam And Seabench: Benchmarking Llms With Local Multilingual Questions In Southeast Asia, Chaoqun Liu, Wenxuan Zhang, Jiahao Ying, Mahani Aljunied, Anh Tuan Luu, Lidong Bing May 2025

Seaexam And Seabench: Benchmarking Llms With Local Multilingual Questions In Southeast Asia, Chaoqun Liu, Wenxuan Zhang, Jiahao Ying, Mahani Aljunied, Anh Tuan Luu, Lidong Bing

Research Collection School Of Computing and Information Systems

This study introduces two novel benchmarks, SeaExam and SeaBench, designed to evalu ate the capabilities of Large Language Models (LLMs) in Southeast Asian (SEA) application scenarios. Unlike existing multilingual datasets primarily derived from English translations, these benchmarks are constructed based on real world scenarios from SEA regions. SeaExam draws from regional educational exams to form a comprehensive dataset that encompasses sub jects such as local history and literature. In contrast, SeaBench is crafted around multi turn, open-ended tasks that reflect daily inter actions within SEA communities. Our evalua tions demonstrate that SeaExam and SeaBench more effectively discern LLM performance on …


Frame-Voyager: Learning To Query Frames For Video Large Language Models, Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xiaolei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, Hao Zhang, Qianru Sun Apr 2025

Frame-Voyager: Learning To Query Frames For Video Large Language Models, Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xiaolei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, Hao Zhang, Qianru Sun

Research Collection School Of Computing and Information Systems

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by …


Pearl: Towards Permutation-Resilient Llms, Liang Chen, Li Shen, Yang Deng, Xiaoyan Zhao, Bin Liang, Kam-Fai Wong Apr 2025

Pearl: Towards Permutation-Resilient Llms, Liang Chen, Li Shen, Yang Deng, Xiaoyan Zhao, Bin Liang, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

The in-context learning (ICL) capability of large language models (LLMs) enables them to perform challenging tasks using provided demonstrations. However, ICL is highly sensitive to the ordering of demonstrations, leading to instability in predictions. This paper shows that this vulnerability can be exploited to design a natural attack - difficult for model providers to detect - that achieves nearly 80% success rate on LLaMA-3 by simply permuting the demonstrations. Existing mitigation methods primarily rely on post-processing and fail to enhance the model's inherent robustness to input permutations, raising concerns about safety and reliability of LLMs. To address this issue, we …


Alphaedit: Null-Space Constrained Knowledge Editing For Language Models, Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Jie Shi, Xiang Wang, Xiangnan He, Tat‑Seng Chua Apr 2025

Alphaedit: Null-Space Constrained Knowledge Editing For Language Models, Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Jie Shi, Xiang Wang, Xiangnan He, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

Large language models (LLMs) often exhibit hallucinations, producing incorrector outdated knowledge. Hence, model editing methods have emerged to enabletargeted knowledge updates. To achieve this, a prevailing paradigm is the locatingthen-editing approach, which first locates influential parameters and then edits themby introducing a perturbation. While effective, current studies have demonstrated that this perturbation inevitably disrupt the originally preserved knowledge within LLMs, especially in sequential editing scenarios. To address this, we introduce AlphaEdit, a novel solution that projects perturbation onto the null space of the preserved knowledge before applying it to the parameters. We theoretically prove that this projection ensures the output …


Llm-Enhanced Multiple Instance Learning For Joint Rumor And Stance Detection With Social Context Information, Ruichao Yang, Jing Ma, Wei Gao, Hongzhan Lin Apr 2025

Llm-Enhanced Multiple Instance Learning For Joint Rumor And Stance Detection With Social Context Information, Ruichao Yang, Jing Ma, Wei Gao, Hongzhan Lin

Research Collection School Of Computing and Information Systems

The proliferation of misinformation, such as rumors on social media, has drawn significant attention, prompting various expressions of stance among users. Although rumor detection and stance detection are distinct tasks, they can complement each other. Rumors can be identified by cross-referencing stances in related posts, and stances are influenced by the nature of the rumor. However, existing stance detection methods often require post-level stance annotations, which are costly to obtain. We propose a novel LLM-enhanced Multiple Instance Learning (MIL) approach to jointly predict post stance and claim class labels, supervised solely by claim labels, using an undirected microblog propagation model. …


Semantic Loss-Guided Data-Efficient Supervised Fine-Tuning For Safe Responses In Llms, Yuxiao Lu, Pradeep Varakantham, Pradeep Varakantham Apr 2025

Semantic Loss-Guided Data-Efficient Supervised Fine-Tuning For Safe Responses In Llms, Yuxiao Lu, Pradeep Varakantham, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, previous approaches often demand substantial human data collection or rely on the less dependable option of using another LLM to generate corrective data. In this paper, we aim to take this problem and overcome limitations of requiring significant high-quality human data. Our method requires only a small set of unsafe responses to toxic prompts, easily obtained from the unsafe LLM itself. By employing a semantic cost combined with a negative Earth Mover Distance (EMD) …


Can Llms Replace Manual Annotation Of Software Engineering Artifacts?, Toufique Ahmed, Premkumar Devanbu, Christoph Treude, Michael Pradel Apr 2025

Can Llms Replace Manual Annotation Of Software Engineering Artifacts?, Toufique Ahmed, Premkumar Devanbu, Christoph Treude, Michael Pradel

Research Collection School Of Computing and Information Systems

Experimental evaluations of software engineering innovations, e.g., tools and processes, often include human-subject studies as a component of a multi-pronged strategy to obtain greater generalizability of the findings. However, human-subject studies in our field are challenging, due to the cost and difficulty of finding and employing suitable subjects, ideally, professional programmers with varying degrees of experience. Meanwhile, large language models (LLMs) have recently started to demonstrate human-level performance in several areas. This paper explores the possibility of substituting costly human subjects with much cheaper LLM queries in evaluations of code and code-related artifacts. We study this idea by applying six …


Chatcrs: Incorporating External Knowledge And Goal Guidance For Llm-Based Conversational Recommender Systems, Chuang Li, Yang Deng, Hengchang Hu, Min-Yen Kan, Haizhou Li Apr 2025

Chatcrs: Incorporating External Knowledge And Goal Guidance For Llm-Based Conversational Recommender Systems, Chuang Li, Yang Deng, Hengchang Hu, Min-Yen Kan, Haizhou Li

Research Collection School Of Computing and Information Systems

This paper aims to efficiently enable large language models (LLMs) to use external knowledge and goal guidance in conversational recommender system (CRS) tasks. Advanced LLMs (e.g., ChatGPT) are limited in domain-specific CRS tasks for 1) generating grounded responses with recommendation-oriented knowledge, or 2) proactively leading the conversations through different dialogue goals. In this work, we first analyze those limitations through a comprehensive evaluation, showing the necessity of external knowledge and goal guidance which contribute significantly to the recommendation accuracy and language quality. In light of this finding, we propose a novel ChatCRS framework to decompose the complex CRS task into …


Agentstudio: A Toolkit For Building General Virtual Agents, Longtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang, Bo An, Shuicheng Yan Apr 2025

Agentstudio: A Toolkit For Building General Virtual Agents, Longtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang, Bo An, Shuicheng Yan

Research Collection School Of Computing and Information Systems

General virtual agents need to handle multimodal observations, master complex action spaces, and self-improve in dynamic, open-domain environments. However, existing environments are often domain-specific and require complex setups, which limits agent development and evaluation in real-world settings. As a result, current evaluations lack in-depth analyses that decompose fundamental agent capabilities. We introduce AgentStudio, a trinity of environments, tools, and benchmarks to address these issues. AgentStudio provides a lightweight, interactive environment with highly generic observation and action spaces, e.g., video observations and GUI/API actions. It integrates tools for creating online benchmark tasks, annotating GUI elements, and labeling actions in videos. Based …


Bigcodebench: Benchmarking Code Generation With Diverse Function Calls And Complex Instructions, T.Y. Zhuo, M.C. Vu, J. Chim, ..., David Lo Apr 2025

Bigcodebench: Benchmarking Code Generation With Diverse Function Calls And Complex Instructions, T.Y. Zhuo, M.C. Vu, J. Chim, ..., David Lo

Research Collection School Of Computing and Information Systems

Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks ranging from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing diverse function calls as tools to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately …


Graph Foundation Models: Concepts, Opportunities And Challenges, Jiawei Liu, Cheng Yang, Zhiyuan Lu, Junze Chen, Yibo Li, Mengmei Zhang, Ting Bai, Fang Yuan, Lichao Sun, Philip S. Yu, Chuan Shi Mar 2025

Graph Foundation Models: Concepts, Opportunities And Challenges, Jiawei Liu, Cheng Yang, Zhiyuan Lu, Junze Chen, Yibo Li, Mengmei Zhang, Ting Bai, Fang Yuan, Lichao Sun, Philip S. Yu, Chuan Shi

Research Collection School Of Computing and Information Systems

Foundation models have emerged as critical components in a variety of artificial intelligence applications, and showcase significant success in natural language processing and several other domains. Meanwhile, the field of graph machine learning is witnessing a paradigm transition from shallow methods to more sophisticated deep learning approaches. The capabilities of foundation models in generalization and adaptation motivate graph machine learning researchers to discuss the potential of developing a new graph learning paradigm. This paradigm envisions models that are pre-trained on extensive graph data and can be adapted for various graph tasks. Despite this burgeoning interest, there is a noticeable lack …


Retrieval Augmented Recipe Generation, Guoshan Liu, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang Mar 2025

Retrieval Augmented Recipe Generation, Guoshan Liu, Hailong Yin, Bin Zhu, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

The growing interest in generating recipes from food images has drawn substantial research attention in recent years. Existing works for recipe generation primarily utilize a two-stage training method—first predicting ingredients from a food image and then generating instructions from both the image and ingredients. Large Multi-modal Models (LMMs), which have achieved notable success across a variety of vision and language tasks, shed light on generating both ingredients and instructions directly from images. Nevertheless, LMMs still face the common issue of hallu- cinations during recipe generation, leading to suboptimal performance. To tackle this issue, we propose a retrieval augmented large multimodal …


A Contrastive Framework With User, Item And Review Alignment For Recommendation, Viet Hoang Dong, Yuan Fang, Hady Wirawan Lauw Mar 2025

A Contrastive Framework With User, Item And Review Alignment For Recommendation, Viet Hoang Dong, Yuan Fang, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Learning effective latent representations for users and items is the cornerstone of recommender systems. Traditional approaches rely on user-item interaction data to map users and items into a shared latent space, but the sparsity of interactions often poses challenges. While leveraging user reviews could mitigate this sparsity, existing review-aware recommendation models often exhibit two key limitations. First, they typically rely on reviews as additional features, but reviews are not universal, with many users and items lacking them. Second, such approaches do not integrate reviews into the useritem space, leading to potential divergence or inconsistency among user, item, and review representations. …


Neurovig: Integrating Event Cameras For Resource-Efficient Video Grounding, Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo Hwee Lim, Archan Misra Mar 2025

Neurovig: Integrating Event Cameras For Resource-Efficient Video Grounding, Dulanga Weerakoon, Vigneshwaran Subbaraju, Joo Hwee Lim, Archan Misra

Research Collection School Of Computing and Information Systems

Spatio-Temporal Video Grounding (STVG) - the task of identifying the target object in the field-of-view that the language instruction refers to - is a fundamental vision-language task. Current STVG approaches typically utilize feeds from an RGB camera that is assumed to be always-on and process the video frames using complex neural network pipelines. As a result they often impose prohibitive system overheads (energy latency) on pervasive devices. To address this we propose NeuroViG with two key innovations: (a) leveraging on event streams from a low-power neuromorphic event camera sensor to perform selective triggering of the more energy-hungry RGB camera for …


Explainable Neural Networks With Guarantees: A Sparse Estimation Approach, Antoine Ledent, Peng Liu Mar 2025

Explainable Neural Networks With Guarantees: A Sparse Estimation Approach, Antoine Ledent, Peng Liu

Research Collection School Of Computing and Information Systems

Balancing predictive power and interpretability has long been a challenging research area, particularly in powerful yet complex models like neural networks, where nonlinearity obstructs direct interpretation. This paper introduces a novel approach to constructing an explainable neural network that harmonizes predictiveness and explainability. Our model is designed as a linear combination of a sparse set of jointly learned features, each derived from a different trainable function applied to a single 1-dimensional input feature. Leveraging the ability to learn arbitrarily complex relationships, our neural network architecture enables automatic selection of a sparse set of important features, with the final prediction being …


Multisfl: Towards Accurate Split Federated Learning Via Multi-Model Aggregation And Knowledge Replay, Zeke Xia, Ming Hu, Dengke Yan, Ruixuan Liu, Anran Li, Xiaofei Xie, Mingsong Chen Mar 2025

Multisfl: Towards Accurate Split Federated Learning Via Multi-Model Aggregation And Knowledge Replay, Zeke Xia, Ming Hu, Dengke Yan, Ruixuan Liu, Anran Li, Xiaofei Xie, Mingsong Chen

Research Collection School Of Computing and Information Systems

Although Split Federated Learning (SFL) effectively enables knowledge sharing among resource-constrained clients, it suffers from low training performance due to the neglect of data heterogeneity and catastrophic forgetting problems. To address these issues, we propose a novel SFL approach named MultiSFL, which adopts i) an effective multimodel aggregation mechanism to alleviate gradient divergence caused by heterogeneous data and ii) a novel knowledge replay strategy to deal with the catastrophic forgetting problem. MultiSFL adopts two servers (i.e., the fed server and main server) to maintain multiple branch models for local training and an aggregated master model for knowledge sharing among branch …


Lightprof: A Lightweight Reasoning Framework For Large Language Model On Knowledge Graph, Tu Ao, Yanhua Yu, Yuling Wang, Yang Deng, Zirui Guo, Liang Pang, Pinghui Wang, Tat-Seng Chua, Xiao Zhang, Zhen Cai Mar 2025

Lightprof: A Lightweight Reasoning Framework For Large Language Model On Knowledge Graph, Tu Ao, Yanhua Yu, Yuling Wang, Yang Deng, Zirui Guo, Liang Pang, Pinghui Wang, Tat-Seng Chua, Xiao Zhang, Zhen Cai

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have impressive capabilities in text understanding and zero-shot reasoning. However, delays in knowledge updates may cause them to reason incorrectly or produce harmful results. Knowledge Graphs (KGs) provide rich and reliable contextual information for the reasoning process of LLMs by structurally organizing and connecting a wide range of entities and relations. Existing KG-based LLM reasoning methods only inject KGs’ knowledge into prompts in a textual form, ignoring its structural information. Moreover, they mostly rely on close-source models or open-source models with large parameters, which poses challenges to high resource consumption. To address this, we propose a …


Aligning Large Language Models For Faithful Integrity Against Opposing Argument, Yong Zhao, Yang Deng, See-Kiong Ng, Tat-Seng Chua Mar 2025

Aligning Large Language Models For Faithful Integrity Against Opposing Argument, Yong Zhao, Yang Deng, See-Kiong Ng, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) have demonstrated impressive capabilities in complex reasoning tasks. However, they can be easily misled by unfaithful arguments during conversations, even when their original statements are correct. To this end, we investigate the problem of maintaining faithful integrity in LLMs. This involves ensuring that LLMs adhere to their faithful statements in the face of opposing arguments and are able to correct their incorrect statements when presented with faithful arguments. In this work, we propose a novel framework, named Alignment for Faithful Integrity with Confidence Estimation (AFICE), which aims to align the LLM responses with faithful integrity. Specifically, …


Learning To Identify Seen, Unseen And Unknown In The Open World: A Practical Setting For Zero-Shot Learning, Sethupathy Parameswaran, Yuan Fang, Chandan Gautam, Savitha Ramasamy, Xiaoli Li Mar 2025

Learning To Identify Seen, Unseen And Unknown In The Open World: A Practical Setting For Zero-Shot Learning, Sethupathy Parameswaran, Yuan Fang, Chandan Gautam, Savitha Ramasamy, Xiaoli Li

Research Collection School Of Computing and Information Systems

As vision-language models advance, addressing the Zero-Shot Learning (ZSL) problem in the open world becomes increasingly crucial. Specifically, a robust model must handle three types of samples during inference: seen classes with visual and semantic information provided in training, unseen classes with only the semantic information in training, and unknown samples with no prior information from training. Existing methods either handle seen and unseen classes together (ZSL) or seen and unknown classes (known as Open-Set Recognition, OSR). However, none addresses the simultaneous handling of all three, which we term Open-Set Zero-Shot Learning (OZSL). To address this problem, we propose a …


Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang Mar 2025

Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

Automated Program Repair (APR) aims to enhance software reliability by automatically generating bug-fixing patches. Recent work has improved the state-of-the-art of APR by fine-tuning pre-trained large language models (LLMs), such as CodeT5, for APR. However, the effectiveness of fine-tuning be-comes weakened in data scarcity scenarios, and data scarcity can be a common issue in practice, limiting fine-tuning performance. To alleviate this limitation, this paper adapts prompt tuning for enhanced APR and conducts a comprehensive study to evaluate its effectiveness in data scarcity scenarios, using three LLMs of different sizes and six diverse datasets across four programming languages. Prompt tuning rewrites …


Mimic: Ai And Ar-Enhanced Multi-Modal, Immersive, Relative Instruction Comprehension, Dhanuja Wanniarachchi, Archan Misra Mar 2025

Mimic: Ai And Ar-Enhanced Multi-Modal, Immersive, Relative Instruction Comprehension, Dhanuja Wanniarachchi, Archan Misra

Research Collection School Of Computing and Information Systems

We present a multimodal instruction comprehension framework, called MImIC, that utilizes visual sensing (including LIDAR and 2D RGB sensing) & AI spatial reasoning capabilities to support more seamless and immersive interaction between humans and AI-driven situated assistive agents. MImIC's key new capability is to support disambiguation of a wider set of relative spatial references that users naturally employ while issuing spatially-situated instructions. To support enhanced visual grounding via a combination of both fully-qualified and relative attribute references, MImIC uses (a) a fine-tuned transformer-based language translation DNN to accurately convert natural verbal commands into a structured set of machine understandable constraints …


Explainable Neural Networks With Guarantee: A Sparse Estimation Approach, Antoine Ledent, Peng Liu Mar 2025

Explainable Neural Networks With Guarantee: A Sparse Estimation Approach, Antoine Ledent, Peng Liu

Research Collection School Of Computing and Information Systems

Balancing predictive power and interpretability has long been a challenging research area, particularly in powerful yet complex models like neural networks, where nonlinearity obstructs direct interpretation. This paper introduces a novel approach to constructing an explainable neural network that harmonizes predictiveness and explainability. Our model is designed as a linear combination of a sparse set of jointly learned features, each derived from a different trainable function applied to a single 1-dimensional input feature. Leveraging the ability to learn arbitrarily complex relationships, our neural network architecture enables automatic selection of a sparse set of important features, with the final prediction being …


Varium: Variational Autoencoder For Multi-Interest Representation With Inter-User Memory, Nhu Thuat Tran, Hady W. Lauw Mar 2025

Varium: Variational Autoencoder For Multi-Interest Representation With Inter-User Memory, Nhu Thuat Tran, Hady W. Lauw

Research Collection School Of Computing and Information Systems

Frameworks for discovering multiple user interest factors based on Variational AutoEncoder (VAE) has demonstrated competitive recommendation performance. However, as VAE only considers one user as input at a time, sharing across like-minded users may not be adequately facilitated. Moreover, interest sharing between users is not always available and thus, poses a challenge for VAE to explicitly model this information. To resolve this, we introduce an inter-user memory-based mechanism to unsupervisedly discover latent interest sharing between users under VAE framework. Concretely, we design a memory including an array of prototypes, each hypothetically representing a group of users sharing a particular interest. …


Simulation-Free Hierarchical Latent Policy Planning For Proactive Dialogues, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Yiheng Sun, Zerui Chen, Ming Liu, Bing Qin Mar 2025

Simulation-Free Hierarchical Latent Policy Planning For Proactive Dialogues, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Yiheng Sun, Zerui Chen, Ming Liu, Bing Qin

Research Collection School Of Computing and Information Systems

Recent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios and comprehensive policy repositories to develop such systems. However, existing approaches tend to rely on Large Language Models (LLMs) for user simulation and online learning, leading to biases that diverge from realistic scenarios and result in suboptimal efficiency. Moreover, these methods depend on manually defined, context-independent, coarse-grained policies, which not only incur high expert costs but also raise concerns regarding their completeness. In our work, we …


Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang Mar 2025

Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

In recent years, AI-based software engineering has progressed from pre-trained models to advanced agentic workflows, with Software Development Agents representing the next major leap. These agents, capable of reasoning, planning, and interacting with external environments, offer promising solutions to complex software engineering tasks. However, while much research has evaluated code generated by large language models (LLMs), comprehensive studies on agent-generated patches, particularly in real-world settings, are lacking. This study addresses that gap by evaluating 4,892 patches from 10 top-ranked agents on 500 real-world GitHub issues from SWE-Bench Verified, focusing on their impact on code quality. Our analysis shows no single …


Backdoor Token Unlearning: Exposing And Defending Backdoors In Pretrained Language Models, Peihai Jiang, Xixiang Lyu, Yige Li, Jing Ma Mar 2025

Backdoor Token Unlearning: Exposing And Defending Backdoors In Pretrained Language Models, Peihai Jiang, Xixiang Lyu, Yige Li, Jing Ma

Research Collection School Of Computing and Information Systems

Supervised fine-tuning has become the predominant method for adapting large pretrained models to downstream tasks. However, recent studies have revealed that these models are vulnerable to backdoor attacks, where even a small number of malicious samples can successfully embed backdoor triggers into the model. While most existing defense methods focus on post-training backdoor defense, efficiently defending against backdoor attacks during training phase remains largely unexplored. To address this gap, we propose a novel defense method called Backdoor Token Unlearning (BTU), which proactively detects and neutralizes trigger tokens during the training stage. Our work is based on two key findings: 1) …