Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 451 - 480 of 1897

Full-Text Articles in Artificial Intelligence and Robotics

Eduqate: Generating Adaptive Curricula Through Rmabs In Education Settings, Sidney Tio, Dexun Li, Pradeep Varakantham May 2025

Eduqate: Generating Adaptive Curricula Through Rmabs In Education Settings, Sidney Tio, Dexun Li, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

There has been significant interest in the development of personalized and adaptive educational tools that cater to a student's individual learning progress. A crucial aspect in developing such tools is in exploring how mastery can be achieved across a diverse yet related range of content in an efficient manner. While Reinforcement Learning and Multi-armed Bandits have shown promise in educational settings, existing works often assume the independence of learning content, neglecting the prevalent interdependencies between such content. In response, we introduce Education Network Restless Multi-armed Bandits (EdNetRMABs), utilizing a network to represent the relationships between interdependent arms. Subsequently, we propose …


Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-Based Benchmark, Han Zhang, Zixiang Meng, Meng Luo, Hong Han, Lizi Liao, Erik Cambria, Hao Fei May 2025

Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-Based Benchmark, Han Zhang, Zixiang Meng, Meng Luo, Hong Han, Lizi Liao, Erik Cambria, Hao Fei

Research Collection School Of Computing and Information Systems

Empathetic Response Generation (ERG) is one of the key tasks of the affective computing area, which aims to produce emotionally nuanced and compassionate responses to user's queries. However, existing ERG research is predominantly confined to the singleton text modality, limiting its effectiveness since human emotions are inherently conveyed through multiple modalities. To combat this, we introduce an avatar-based Multimodal ERG (MERG) task, entailing rich text, speech, and facial vision information. We first present a large-scale high-quality benchmark dataset, AvaMERG, which extends traditional text ERG by incorporating authentic human speech audio and dynamic talking-face avatar videos, encompassing a diverse range of …


Grounding Ai Use In Learning Science: A Conversation With Steven Miller, Steven Miller, Lieven Demeester Apr 2025

Grounding Ai Use In Learning Science: A Conversation With Steven Miller, Steven Miller, Lieven Demeester

CASTLe: Collection of Articles on Scholarship for Teaching and Learning

In this insightful interview, SMU Associate Provost (Teaching and Learning Innovation) Lieven Demeester and Professor Emeritus of Information Systems Steven Miller discuss the integration of artificial intelligence (AI) in teaching and learning, emphasising the importance of grounding AI use in the fundamentals of learning science. They explore the evolving role of education in the context of AI advancements, highlighting the need for educators to focus on the cognitive aspects of learning, such as goal-directed practice and feedback. They also address the potential of AI as a collaborative agent in group projects and the importance of maintaining accountability and quality control …


Frame-Voyager: Learning To Query Frames For Video Large Language Models, Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xiaolei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, Hao Zhang, Qianru Sun Apr 2025

Frame-Voyager: Learning To Query Frames For Video Large Language Models, Sicheng Yu, Chengkai Jin, Huanyu Wang, Zhenghao Chen, Sheng Jin, Zhongrong Zuo, Xiaolei Xu, Zhenbang Sun, Bingni Zhang, Jiawei Wu, Hao Zhang, Qianru Sun

Research Collection School Of Computing and Information Systems

Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by …


Does Chatgpt-Permitted Assessments Help Students Generate Better Answers And Learn More?, Michelle L. F. Cheong, Yun-Chen Chen Apr 2025

Does Chatgpt-Permitted Assessments Help Students Generate Better Answers And Learn More?, Michelle L. F. Cheong, Yun-Chen Chen

Research Collection School Of Computing and Information Systems

We discuss our methodology and implementation of ChatGPT-permitted assessments for a university-level spreadsheets modelling module. Through our quantitative data analysis, our students rated ChatGPT’s answers to be incorrect on average and thus will not help them generate better answers directly, representing low “Perceived usefulness” (PU), while they rated ChatGPT 3.5 with relatively high “Perceived ease of use” (PE). They gave a good “Behavioural intention” (BI) rating indicating that they were motivated to use it in future as they could still learn more about this module by using ChatGPT 3.5. We found that both PU and PE affected BI positively, with …


On Generalization Across Environments In Multi-Objective Reinforcement Learning, Jayden Jing Xiang Teoh, Pradeep Varakantham, Peter Vamplew Apr 2025

On Generalization Across Environments In Multi-Objective Reinforcement Learning, Jayden Jing Xiang Teoh, Pradeep Varakantham, Peter Vamplew

Research Collection School Of Computing and Information Systems

No abstract provided.


Pearl: Towards Permutation-Resilient Llms, Liang Chen, Li Shen, Yang Deng, Xiaoyan Zhao, Bin Liang, Kam-Fai Wong Apr 2025

Pearl: Towards Permutation-Resilient Llms, Liang Chen, Li Shen, Yang Deng, Xiaoyan Zhao, Bin Liang, Kam-Fai Wong

Research Collection School Of Computing and Information Systems

The in-context learning (ICL) capability of large language models (LLMs) enables them to perform challenging tasks using provided demonstrations. However, ICL is highly sensitive to the ordering of demonstrations, leading to instability in predictions. This paper shows that this vulnerability can be exploited to design a natural attack - difficult for model providers to detect - that achieves nearly 80% success rate on LLaMA-3 by simply permuting the demonstrations. Existing mitigation methods primarily rely on post-processing and fail to enhance the model's inherent robustness to input permutations, raising concerns about safety and reliability of LLMs. To address this issue, we …


Chatcrs: Incorporating External Knowledge And Goal Guidance For Llm-Based Conversational Recommender Systems, Chuang Li, Yang Deng, Hengchang Hu, Min-Yen Kan, Haizhou Li Apr 2025

Chatcrs: Incorporating External Knowledge And Goal Guidance For Llm-Based Conversational Recommender Systems, Chuang Li, Yang Deng, Hengchang Hu, Min-Yen Kan, Haizhou Li

Research Collection School Of Computing and Information Systems

This paper aims to efficiently enable large language models (LLMs) to use external knowledge and goal guidance in conversational recommender system (CRS) tasks. Advanced LLMs (e.g., ChatGPT) are limited in domain-specific CRS tasks for 1) generating grounded responses with recommendation-oriented knowledge, or 2) proactively leading the conversations through different dialogue goals. In this work, we first analyze those limitations through a comprehensive evaluation, showing the necessity of external knowledge and goal guidance which contribute significantly to the recommendation accuracy and language quality. In light of this finding, we propose a novel ChatCRS framework to decompose the complex CRS task into …


A Selective Vehicle Routing Problem For The Bloodmobile System, Aldy Gunawan, Samuel Alan Darmasaputra, Sy Hoang Do, Vincent F. Yu Apr 2025

A Selective Vehicle Routing Problem For The Bloodmobile System, Aldy Gunawan, Samuel Alan Darmasaputra, Sy Hoang Do, Vincent F. Yu

Research Collection School Of Computing and Information Systems

Mobile blood collection has the advantage of greater reach compared to blood drives at fixed donation sites and is preferable for individuals with limited time or means of transportation. Bloodmobiles are widely used in healthcare logistics to increase the number of donors and donation frequency and to better match blood demand with collection. Bloodmobiles are stationed at predetermined locations, while shuttles are assigned to visit these locations to collect the donated blood. This problem is formulated as the Selective Vehicle Routing Problem under the Bloodmobile System (SVRP-BM). This research extends the Selective Vehicle Routing Problem with Integrated Tours problem (SVRPwIT) …


Towards Understanding Why Fixmatch Generalizes Better Than Supervised Learning, Jingyang Li, Jiachun Pan, Vincent Tan, Kim-Chuan Toh, Pan Zhou Apr 2025

Towards Understanding Why Fixmatch Generalizes Better Than Supervised Learning, Jingyang Li, Jiachun Pan, Vincent Tan, Kim-Chuan Toh, Pan Zhou

Research Collection School Of Computing and Information Systems

Semi-supervised learning (SSL), exemplified by FixMatch (Sohn et al., 2020), has shown significant generalization advantages over supervised learning (SL), particularly in the context of deep neural networks (DNNs). However, it is still unclear, from a theoretical standpoint, why FixMatch-like SSL algorithms generalize better than SL on DNNs. In this work, we present the first theoretical justification for the enhanced test accuracy observed in FixMatch-like SSL applied to DNNs by taking convolutional neural networks (CNNs) on classification tasks as an example. Our theoretical analysis reveals that the semantic feature learning processes in FixMatch and SL are rather different. In particular, FixMatch …


Capo: Cooperative Plan Optimization For Efficient Embodied Multi-Agent Cooperation, Jie Liu, Pan Zhou, Yingjun Du, Ah-Hwee Tan, Cees Snoek, Jan-Jakob Sonke, Efstratios Gavves Apr 2025

Capo: Cooperative Plan Optimization For Efficient Embodied Multi-Agent Cooperation, Jie Liu, Pan Zhou, Yingjun Du, Ah-Hwee Tan, Cees Snoek, Jan-Jakob Sonke, Efstratios Gavves

Research Collection School Of Computing and Information Systems

In this work, we address the cooperation problem among large language model (LLM) based embodied agents, where agents must cooperate to achieve a common goal. Previous methods often execute actions extemporaneously and incoherently, without long-term strategic and cooperative planning, leading to redundant steps, failures, and even serious repercussions in complex tasks like search-and-rescue missions where discussion and cooperative plan are crucial. To solve this issue, we propose Cooperative Plan Optimization (CaPo) to enhance the cooperation efficiency of LLM-based embodied agents. Inspired by human cooperation schemes, CaPo improves cooperation efficiency with two phases: 1) meta-plan generation, and 2) progress-adaptive meta-plan and …


Configx: Modular Configuration For Evolutionary Algorithms Via Multitask Reinforcement Learning, Hongshu Guo, Zeyuan Ma, Jiacheng Chen, Yining Ma, Zhiguang Cao, Xinglin Zhang, Yue-Jiao Gong Apr 2025

Configx: Modular Configuration For Evolutionary Algorithms Via Multitask Reinforcement Learning, Hongshu Guo, Zeyuan Ma, Jiacheng Chen, Yining Ma, Zhiguang Cao, Xinglin Zhang, Yue-Jiao Gong

Research Collection School Of Computing and Information Systems

Recent advances in Meta-learning for Black-Box Optimization (MetaBBO) have shown the potential of using neural networks to dynamically configure evolutionary algorithms (EAs), enhancing their performance and adaptability across various BBO instances. However, they are often tailored to a specific EA, which limits their generalizability and necessitates retraining or redesigns for different EAs and optimization problems. To address this limitation, we introduce ConfigX, a new paradigm of the MetaBBO framework that is capable of learning a universal configuration agent (model) for boosting diverse EAs. To achieve so, our ConfigX first leverages a novel modularization system that enables the flexible combination of …


Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang Apr 2025

Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang

Research Collection School Of Computing and Information Systems

With the widespread adoption of Internet Protocol (IP) communication technology and web-based platforms, cloud manufacturing has become a significant hallmark of Industry 4.0. Integrating graph algorithms into these web-enabled environments is crucial as they facilitate the representation and analysis of complex relationships in manufacturing processes, enabling efficient decision-making and adaptability in dynamic environments. As a key scheduling problem in cloud manufacturing, the flexible job-shop scheduling problem (FJSP) finds extensive applications in real-world scenarios. However, traditional FJSP-solving methods struggle to meet the efficiency and adaptability demands of cloud manufacturing due to generalization issues and excessive computational time, while reinforcement learning-based methods …


Graph-Assisted Offline-Online Deep Reinforcement Learning For Dynamic Workflow Scheduling, Yifan Yang, Gang Chen, Hui Ma, Cong Zhang, Zhiguang Cao, Mengjie Zhang Apr 2025

Graph-Assisted Offline-Online Deep Reinforcement Learning For Dynamic Workflow Scheduling, Yifan Yang, Gang Chen, Hui Ma, Cong Zhang, Zhiguang Cao, Mengjie Zhang

Research Collection School Of Computing and Information Systems

Dynamic workflow scheduling (DWS) in cloud computing presents substantial challenges due to heterogeneous machine configurations, unpredictable workflow arrivals/patterns, and constantly evolving environments. However, existing research often assumes homogeneous setups and static conditions, limiting flexibility and adaptability in real-world scenarios. In this paper, we propose a novel Graph assisted Offline-Online Deep Reinforcement Learning (GOODRL) approach to building an effective and efficient scheduling agent for DWS. Our approach features three key innovations: (1) a task-specific graph representation and a Graph Attention Actor Network that enable the agent to dynamically assign focused tasks to heterogeneous machines while explicitly considering the future impact of …


Neural Multi-Objective Combinatorial Optimization Via Graph-Image Multimodal Fusion, Jinbiao Chen, Jiahai Wang, Zhiguang Cao, Yaoxin Wu Apr 2025

Neural Multi-Objective Combinatorial Optimization Via Graph-Image Multimodal Fusion, Jinbiao Chen, Jiahai Wang, Zhiguang Cao, Yaoxin Wu

Research Collection School Of Computing and Information Systems

Existing neural multi-objective combinatorial optimization (MOCO) methods still exhibit an optimality gap since they fail to fully exploit the intrinsic features of problem instances. A significant factor contributing to this shortfall is their reliance solely on graph-modal information. To overcome this, we propose a novel graph-image multimodal fusion (GIMF) framework that enhances neural MOCO methods by integrating graph and image information of the problem instances. Our GIMF framework comprises three key components: (1) a constructed coordinate image to better represent the spatial structure of the problem instance, (2) a problem-size adaptive resolution strategy during the image construction process to improve …


Rethinking Light Decoder-Based Solvers For Vehicle Routing Problems, Ziwei Huang, Jianan Zhou, Zhiguang Cao, Yixin Xu Apr 2025

Rethinking Light Decoder-Based Solvers For Vehicle Routing Problems, Ziwei Huang, Jianan Zhou, Zhiguang Cao, Yixin Xu

Research Collection School Of Computing and Information Systems

Light decoder-based solvers have gained popularity for solving vehicle routing problems (VRPs) due to their efficiency and ease of integration with reinforcement learning algorithms. However, they often struggle with generalization to larger problem instances or different VRP variants. This paper revisits light decoder-based approaches, analyzing the implications of their reliance on static embeddings and the inherent challenges that arise. Specifically, we demonstrate that in the light decoder paradigm, the encoder is implicitly tasked with capturing information for all potential decision scenarios during solution construction within a single set of embeddings, resulting in high information density. Furthermore, our empirical analysis reveals …


Rethinking Neural Multi-Objective Combinatorial Optimization Via Neat Weight Embedding, Jinbiao Chen, Zhiguang Cao, Jiahai Wang, Yaoxin Wu, Hanzhang Qin, Zizhen Zhang, Yue-Jiao Gong Apr 2025

Rethinking Neural Multi-Objective Combinatorial Optimization Via Neat Weight Embedding, Jinbiao Chen, Zhiguang Cao, Jiahai Wang, Yaoxin Wu, Hanzhang Qin, Zizhen Zhang, Yue-Jiao Gong

Research Collection School Of Computing and Information Systems

Recent decomposition-based neural multi-objective combinatorial optimization (MOCO) methods struggle to achieve desirable performance. Even equipped with complex learning techniques, they often suffer from significant optimality gaps in weight-specific subproblems. To address this challenge, we propose a neat weight embedding method to learn weight-specific representations, which captures weight-instance interaction for the subproblems and was overlooked by most current methods. We demonstrate the potentials of our method in two instantiations. First, we introduce a succinct addition model to learn weight-specific node embeddings, which surpassed most existing neural methods. Second, we design an enhanced conditional attention model to simultaneously learn the weight embedding …


Alphaedit: Null-Space Constrained Knowledge Editing For Language Models, Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Jie Shi, Xiang Wang, Xiangnan He, Tat‑Seng Chua Apr 2025

Alphaedit: Null-Space Constrained Knowledge Editing For Language Models, Junfeng Fang, Houcheng Jiang, Kun Wang, Yunshan Ma, Jie Shi, Xiang Wang, Xiangnan He, Tat‑Seng Chua

Research Collection School Of Computing and Information Systems

Large language models (LLMs) often exhibit hallucinations, producing incorrector outdated knowledge. Hence, model editing methods have emerged to enabletargeted knowledge updates. To achieve this, a prevailing paradigm is the locatingthen-editing approach, which first locates influential parameters and then edits themby introducing a perturbation. While effective, current studies have demonstrated that this perturbation inevitably disrupt the originally preserved knowledge within LLMs, especially in sequential editing scenarios. To address this, we introduce AlphaEdit, a novel solution that projects perturbation onto the null space of the preserved knowledge before applying it to the parameters. We theoretically prove that this projection ensures the output …


Predictive Modelling For Vessel Traffic Flow: A Comprehensive Survey From Statistics To Ai, Deshan. Chen, Chen. Huang, Tengze. Fan, Hoong Chuin Lau, Xinping. Yan Apr 2025

Predictive Modelling For Vessel Traffic Flow: A Comprehensive Survey From Statistics To Ai, Deshan. Chen, Chen. Huang, Tengze. Fan, Hoong Chuin Lau, Xinping. Yan

Research Collection School Of Computing and Information Systems

Recognizing the specific complexities of vessel traffic flow, this comprehensive survey exclusively addresses the predictive modelling in maritime transportation, tracing the evolution from conventional statistical approaches to modern artificial intelligence (AI) techniques. The survey examines a broad range of predictive targets, including vessel volume, trajectories, velocities, destinations and traffic patterns. Through bibliometric analysis utilizing Citespace, the central research themes and technological trends characterizing the vessel traffic flow prediction domain have been identified and discussed. Our analysis indicates a clear trend towards AI-based models, highlighting their increasing dominance in enhancing predictive accuracy and efficiency. Additionally, we highlight persistent challenges, such as …


Bigcodebench: Benchmarking Code Generation With Diverse Function Calls And Complex Instructions, T.Y. Zhuo, M.C. Vu, J. Chim, ..., David Lo Apr 2025

Bigcodebench: Benchmarking Code Generation With Diverse Function Calls And Complex Instructions, T.Y. Zhuo, M.C. Vu, J. Chim, ..., David Lo

Research Collection School Of Computing and Information Systems

Task automation has been greatly empowered by the recent advances in Large Language Models (LLMs) via Python code, where the tasks ranging from software engineering development to general-purpose reasoning. While current benchmarks have shown that LLMs can solve tasks using programs like human developers, the majority of their evaluations are limited to short and self-contained algorithmic tasks or standalone function calls. Solving challenging and practical tasks requires the capability of utilizing diverse function calls as tools to efficiently implement functionalities like data analysis and web development. In addition, using multiple tools to solve a task needs compositional reasoning by accurately …


The Impact Of Ai Usage On Employee Work Outcomes: The Mediating Roles Of Personal Control And Job Insecurity And The Moderating Role Of Ai Trust, Tiantian Wang Apr 2025

The Impact Of Ai Usage On Employee Work Outcomes: The Mediating Roles Of Personal Control And Job Insecurity And The Moderating Role Of Ai Trust, Tiantian Wang

Dissertations and Theses Collection (Open Access)

The widespread application of artificial intelligence (AI) technology in the workplace offers significant potential for process optimization andperformance improvement. However, the psychological mechanisms throughwhich AI usage affects employee outcomes remain underexplored. To address this gap, the present study investigated a sample of 170 employees froma media company in China, utilizing a three-wave longitudinal survey design. Specifically, this study examined how AI usage influenced employee creativity and task performance improvement through two mediatingmechanisms: the enhancement of personal control in problem-solving and the elicitation of job insecurity. Furthermore, the moderating role of trust in AI inthe relationship between AI usage and job …


Nash Bargaining Strategy In Autonomous Decision Making For Multi-Ship Collision Avoidance Based On Route Exchange, Yang Wang, Qiangsheng Ye, Hoong Chuin Lau, Tengfei Wang, Bing Wu Apr 2025

Nash Bargaining Strategy In Autonomous Decision Making For Multi-Ship Collision Avoidance Based On Route Exchange, Yang Wang, Qiangsheng Ye, Hoong Chuin Lau, Tengfei Wang, Bing Wu

Research Collection School Of Computing and Information Systems

A novel scheme is proposed for the distributed multi-ship collision avoidance (CA) problem with consideration of the autonomous, dynamic nature of the real circumstance. All the ships in the envisioned scenarios can share their decisions or intentions through route exchange, allowing them to make subsequent decisions based on the route planning in each iteration. By leveraging route exchange, the multi-ship CA problem involves iterations for negotiation, and is regarded as a staged cooperative game under conditions of complete information. The concept of closest spatio-temporal distance (CSTD) is introduced to more accurately assess collision risk between ships. A coordinated CA mechanism …


Samgpt: Text-Free Graph Foundation Model For Multi-Domain Pre-Training And Cross-Domain Adaptation, Xingtong Yu, Zechuan Gong, Chang Zhou, Yuan Fang, Hui Zhang Apr 2025

Samgpt: Text-Free Graph Foundation Model For Multi-Domain Pre-Training And Cross-Domain Adaptation, Xingtong Yu, Zechuan Gong, Chang Zhou, Yuan Fang, Hui Zhang

Research Collection School Of Computing and Information Systems

Graphs are able to model interconnected entities in many online services, supporting a wide range of applications on the Web. This raises an important question: How can we train a graph foundational model on multiple source domains and adapt to an unseen target domain? A major obstacle is that graphs from different domains often exhibit divergent characteristics. Some studies leverage large language models to align multiple domains based on textual descriptions associated with the graphs, limiting their applicability to text-attributed graphs. For text-free graphs, a few recent works attempt to align different feature distributions across domains, while generally neglecting structural …


Comadice: Offline Cooperative Multi-Agent Reinforcement Learning With Stationary Distribution Shift Regularization, The Viet Bui, Tien Mai, Hong Thanh Nguyen Apr 2025

Comadice: Offline Cooperative Multi-Agent Reinforcement Learning With Stationary Distribution Shift Regularization, The Viet Bui, Tien Mai, Hong Thanh Nguyen

Research Collection School Of Computing and Information Systems

Offline reinforcement learning (RL) has garnered significant attention for its ability to learn effective policies from pre-collected datasets without the need for further environmental interactions. While promising results have been demonstrated in single-agent settings, offline multi-agent reinforcement learning (MARL) presents additional challenges due to the large joint state-action space and the complexity of multi-agent behaviors. A key issue in offline RL is the distributional shift, which arises when the target policy being optimized deviates from the behavior policy that generated the data. This problem is exacerbated in MARL due to the interdependence between agents' local policies and the expansive joint …


Agentstudio: A Toolkit For Building General Virtual Agents, Longtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang, Bo An, Shuicheng Yan Apr 2025

Agentstudio: A Toolkit For Building General Virtual Agents, Longtao Zheng, Zhiyuan Huang, Zhenghai Xue, Xinrun Wang, Bo An, Shuicheng Yan

Research Collection School Of Computing and Information Systems

General virtual agents need to handle multimodal observations, master complex action spaces, and self-improve in dynamic, open-domain environments. However, existing environments are often domain-specific and require complex setups, which limits agent development and evaluation in real-world settings. As a result, current evaluations lack in-depth analyses that decompose fundamental agent capabilities. We introduce AgentStudio, a trinity of environments, tools, and benchmarks to address these issues. AgentStudio provides a lightweight, interactive environment with highly generic observation and action spaces, e.g., video observations and GUI/API actions. It integrates tools for creating online benchmark tasks, annotating GUI elements, and labeling actions in videos. Based …


On Minimizing Adversarial Counterfactual Error In Adversarial Reinforcement Learning, Roman Belaire, Arunesh Sinha, Pradeep Varakantham Apr 2025

On Minimizing Adversarial Counterfactual Error In Adversarial Reinforcement Learning, Roman Belaire, Arunesh Sinha, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge inherent to adversarial perturbations is that by altering the information observed by the agent, the state becomes only partially observable. Existing approaches address this by either enforcing consistent actions across nearby states or maximizing the worst-case value within adversarially perturbed observations. However, the former suffers from performance degradation when attacks succeed, while the latter tends to be overly conservative, leading to suboptimal performance in benign settings. We hypothesize that these limitations stem from their failing to account for …


Semantic Loss-Guided Data-Efficient Supervised Fine-Tuning For Safe Responses In Llms, Yuxiao Lu, Pradeep Varakantham, Pradeep Varakantham Apr 2025

Semantic Loss-Guided Data-Efficient Supervised Fine-Tuning For Safe Responses In Llms, Yuxiao Lu, Pradeep Varakantham, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

Large Language Models (LLMs) generating unsafe responses to toxic prompts is a significant issue in their applications. While various efforts aim to address this safety concern, previous approaches often demand substantial human data collection or rely on the less dependable option of using another LLM to generate corrective data. In this paper, we aim to take this problem and overcome limitations of requiring significant high-quality human data. Our method requires only a small set of unsafe responses to toxic prompts, easily obtained from the unsafe LLM itself. By employing a semantic cost combined with a negative Earth Mover Distance (EMD) …


Bootstrapping Language Models With Dpo Implicit Rewards, Changyu Chen, Zichen Liu, Chao Du, Tianyu Pang, Qian Liu, Arunesh Sinha, Pradeep Varakantham, Min Lin Apr 2025

Bootstrapping Language Models With Dpo Implicit Rewards, Changyu Chen, Zichen Liu, Chao Du, Tianyu Pang, Qian Liu, Arunesh Sinha, Pradeep Varakantham, Min Lin

Research Collection School Of Computing and Information Systems

Human alignment in large language models (LLMs) is an active area of research. A recent groundbreaking work, direct preference optimization (DPO), has greatly simplified the process from past work in reinforcement learning from human feedback (RLHF) by bypassing the reward learning stage in RLHF. DPO, after training, provides an implicit reward model. In this work, we make a novel observation that this implicit reward model can by itself be used in a bootstrapping fashion to further align the LLM. Our approach is to use the rewards from a current LLM model to construct a preference dataset, which is then used …


Why Ai’S Role In Advancing Sustainability Is Underestimated, Lipika Bhattacharya Mar 2025

Why Ai’S Role In Advancing Sustainability Is Underestimated, Lipika Bhattacharya

CCX Research

AI has quietly, but powerfully, woven itself into the fabric of our everyday lives. Yet, AI's potential impact on creating a more sustainable world is undervalued. The author examined artificial intelligence's (AI) role in advancing sustainability. She outlined how AI can be applied in various domains such as agriculture, water management, industry, urban planning, and biodiversity conservation for transformative effects.


Learning: Human Versus Machine, Sanjay Sarma Mar 2025

Learning: Human Versus Machine, Sanjay Sarma

Asian Management Insights

Outdated education paradigms must be revamped to reclaim the all-important human quality: agency.

Sanjay Sarma, CEO, President, and Dean of the Asia School of Business, Kuala Lumpur, Malaysia and the Fred Fort Flowers (1941) and Daniel Fort Flowers (1941) Professor in Mechanical Engineering at the Massachusetts Institute of Technology (MIT), shares insights on the artificial intelligence (AI)-agency revolution and how the human brain works.