Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (360)
- Engineering (296)
- Operations Research, Systems Engineering and Industrial Engineering (256)
- Business (177)
- Graphics and Human Computer Interfaces (176)
-
- Social and Behavioral Sciences (173)
- Software Engineering (137)
- Numerical Analysis and Scientific Computing (112)
- Theory and Algorithms (103)
- Public Affairs, Public Policy and Public Administration (75)
- Transportation (63)
- Programming Languages and Compilers (62)
- Information Security (43)
- Medicine and Health Sciences (43)
- Technology and Innovation (38)
- OS and Networks (37)
- Education (36)
- Law (36)
- Asian Studies (35)
- Computer Engineering (35)
- International and Area Studies (35)
- Health Information Technology (30)
- Science and Technology Law (22)
- Psychology (21)
- Library and Information Science (20)
- Finance and Financial Management (18)
- Higher Education (18)
- Keyword
-
- Artificial intelligence (98)
- Machine learning (55)
- Reinforcement learning (42)
- Deep learning (38)
- Artificial Intelligence (30)
-
- Large Language Models (30)
- Generative AI (29)
- ChatGPT (23)
- Large Language Model (23)
- Large language models (23)
- Singapore (22)
- Computer vision (19)
- Large language model (18)
- Optimization (18)
- Reinforcement Learning (18)
- Scheduling (18)
- Anomaly detection (17)
- Natural language processing (17)
- Deep reinforcement learning (16)
- Deep Learning (15)
- LLMs (15)
- Machine Learning (15)
- Vehicle routing problem (15)
- AI (14)
- Neural networks (13)
- Uncertainty (13)
- Software engineering (12)
- Graph neural networks (11)
- Metaverse (10)
- Multi-agent systems (10)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (1664)
- Dissertations and Theses Collection (Open Access) (57)
- Research Collection Lee Kong Chian School Of Business (33)
- Research Collection Yong Pung How School Of Law (31)
- Research Collection School of Social Sciences (22)
-
- Asian Management Insights (16)
- FORCE 2026 (14)
- Perspectives@SMU (11)
- Research Collection College of Integrative Studies (10)
- Research Collection Library (8)
- PhD Student’s Publications Collection (6)
- MITB Thought Leadership Series (4)
- 2024 AI for Research Week (3)
- CCX Research (3)
- LARC Research Publications (2)
- Research Collection School Of Accountancy (2)
- CASTLe: Collection of Articles on Scholarship for Teaching and Learning (1)
- Centre for AI & Data Governance (2019-2025) (1)
- Centre for Computational Law (2022-2025) (1)
- ROSA Journal Articles and Publications (1)
- Research Collection Office of Research (1)
- Research Collection School Of Economics (1)
- Research@SMU Infographics (1)
- Research@SMU: Connecting the Dots (1)
- SMU Press Releases and News (1)
- Sim Kee Boon Institute for Financial Economics (1)
- Student Publications (1)
- Publication Type
- File Type
Articles 151 - 180 of 1897
Full-Text Articles in Artificial Intelligence and Robotics
Quantum Chebyshev Transform-Based Graph Neural Networks For Financial Fraud Detection, Minrui Xu, Bingyan Guan, Bethel Hui Ting Loke, Paul R. Griffin
Quantum Chebyshev Transform-Based Graph Neural Networks For Financial Fraud Detection, Minrui Xu, Bingyan Guan, Bethel Hui Ting Loke, Paul R. Griffin
Research Collection School Of Computing and Information Systems
Financial fraud detection is a critical challenge requiring accurate identification of anomalous patterns in complex transaction networks. Graph Neural Networks (GNNs) have emerged as powerful tools for fraud detection by capturing relational structures among entities. Meanwhile, quantum computing offers new possibilities to enhance machine learning through high-dimensional Hilbert spaces and parallelism. In this paper, we propose a hybrid classical-quantum model called QCTGNN (Quantum Chebyshev Transform-based Graph Neural Network) for financial fraud detection. The QCTGNN integrates a classical graph neural network component based on Simplified Graph Convolutions (SGConv) with a quantum component that performs a Chebyshev polynomial-based transform via variational quantum …
Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He
Lagrangian Motion Fields For Long-Term Motion Generation, Yifei Yang, Zikai Huang, Chenshu Xu, Shengfeng He
Research Collection School Of Computing and Information Systems
Long-term motion generation is a challenging task that requires producing coherent and realistic sequences over extended durations. Current methods primarily rely on framewise motion representations, which capture only static spatial details and overlook temporal dynamics. This approach leads to significant redundancy across the temporal dimension, complicating the generation of effective long-term motion. To overcome these limitations, we introduce the novel concept of Lagrangian Motion Fields, specifically designed for long-term motion generation. By treating each joint as a Lagrangian particle with uniform velocity over short intervals, our approach condenses motion representations into a series of "supermotions" (analogous to superpixels). This method …
Grounding Is All You Need? Dual Temporal Grounding For Video Dialog, You Qin, Wei Ji, Xinze Lan, Hao Fei, Xun Yang, Dan Guo, Roger Zimmermann, Lizi Liao
Grounding Is All You Need? Dual Temporal Grounding For Video Dialog, You Qin, Wei Ji, Xinze Lan, Hao Fei, Xun Yang, Dan Guo, Roger Zimmermann, Lizi Liao
Research Collection School Of Computing and Information Systems
In the realm of video dialog response generation, capturing both the essence of video content and the temporal nuances of conversation history is crucial. While some approaches rely on large-scale pretrained visual-language models, often neglecting temporal dynamics, others emphasize spatial-temporal relationships within videos but demand intricate object trajectory pre-extractions and overlook dialog temporal dynamics. This paper introduces the Dual Temporal Grounding-enhanced Video Dialog model (DTGVD), designed to bridge the gap between these two approaches. DTGVD uniquely integrates the strengths of both by emphasizing dual temporal relationships. It achieves this by predicting dialog turn-specific temporal regions, selectively filtering video content, and …
Learnable Game-Theoretic Policy Optimization For Data-Centric Self-Explanation Rationalization, Yunxiao Zhao, Zhiqiang Wang, Xingtong Yu, Xiaoli Li, Jiye Liang, Ru Li
Learnable Game-Theoretic Policy Optimization For Data-Centric Self-Explanation Rationalization, Yunxiao Zhao, Zhiqiang Wang, Xingtong Yu, Xiaoli Li, Jiye Liang, Ru Li
Research Collection School Of Computing and Information Systems
Rationalization, a data-centric framework, aims to build self-explanatory models to explain the prediction outcome by generating a subset of human-intelligible pieces of the input data. It involves a cooperative game model where a generator generates the most human-intelligible parts of the input (i.e., rationales), followed by a predictor that makes predictions based on these generated rationales. Conventional rationalization methods typically impose constraints via regularization terms to calibrate or penalize undesired generation. However, these methods are suffering from a problem called mode collapse, in which the predictor produces correct predictions yet the generator consistently outputs rationales with collapsed patterns. Moreover, existing …
Mm-Attackg: A Multimodal Approach To Attack Graph Construction With Large Language Models, Yongheng Zhang, Xinyun Zhao, Yunshan Ma, Haokai Ma, Yingxiao Guan, Guozheng Yang, Yuliang Lu, Xiang Wang
Mm-Attackg: A Multimodal Approach To Attack Graph Construction With Large Language Models, Yongheng Zhang, Xinyun Zhao, Yunshan Ma, Haokai Ma, Yingxiao Guan, Guozheng Yang, Yuliang Lu, Xiang Wang
Research Collection School Of Computing and Information Systems
Cyber Threat Intelligence (CTI) parsing aims to extract key threat information from massive data, transform it into actionable intelligence, enhance threat detection and defense efficiency, including attack graph construction, intelligence fusion, and indicator extraction. Among these research topics, Attack Graph Construction (AGC) is essential for visualizing and understanding the potential attack paths of threat events from CTI reports. Existing approaches primarily construct the attack graphs purely from the textual data to reveal the logical threat relationships between entities within the attack behavioral sequence. However, they typically overlook the specific threat information inherent in visual modalities, which preserves key threat details …
Obstructive Sleep Apnea Prediction: A Comprehensive Review And Comparative Study, Thi Khanh Chi Huynh, Amonae Dabbs-Brown, Anna Jurek-Loughrey, James Mulhall, Tuan Dung Pham, Ngoc Phu Doan, Viet Hung Tran, Zichi Zhang, Xuan Hoang Nguyen, Yimeng An, Peixin Li, Phi Hung Nguyen, Thi Linh Hoang, Xinming Shi, Hans Vandierendonck, Sebastien Bailly, Jean-Louis Pépin, Thai Son Mai
Obstructive Sleep Apnea Prediction: A Comprehensive Review And Comparative Study, Thi Khanh Chi Huynh, Amonae Dabbs-Brown, Anna Jurek-Loughrey, James Mulhall, Tuan Dung Pham, Ngoc Phu Doan, Viet Hung Tran, Zichi Zhang, Xuan Hoang Nguyen, Yimeng An, Peixin Li, Phi Hung Nguyen, Thi Linh Hoang, Xinming Shi, Hans Vandierendonck, Sebastien Bailly, Jean-Louis Pépin, Thai Son Mai
Research Collection School Of Computing and Information Systems
Obstructive Sleep Apnea (OSA) is a highly prevalent sleep disorder linked to considerable public health burdens and comorbidities. However, its heterogeneous presentation and the limited accessibility of traditional diagnostic tools such as polysomnography (PSG) lead to widespread underdiagnosis. As a result, artificial intelligence (AI) approaches, including machine learning (ML) and deep learning (DL) models, have attracted attention as an alternative pathway to detection. This paper first provides a comprehensive review of AI-driven OSA diagnosis, covering different diagnosis problems, input-data types, data biases, pre-processing techniques, and model performance. We then leverage the largest clinical dataset used in OSA prediction to date, …
The Gains Do Not Make Up For The Losses: A Comprehensive Evaluation For Safety Alignment Of Large Language Models Via Machine Unlearning, Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che
The Gains Do Not Make Up For The Losses: A Comprehensive Evaluation For Safety Alignment Of Large Language Models Via Machine Unlearning, Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng, Yanyan Zhao, Bing Qin, Wanxiang Che
Research Collection School Of Computing and Information Systems
Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased, leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark MuBench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful …
Leading The Change: Staff-Driven Ai Transformation In Smu Libraries’ Collection Team, Siew Khim Lim, Fion Goh
Leading The Change: Staff-Driven Ai Transformation In Smu Libraries’ Collection Team, Siew Khim Lim, Fion Goh
Research Collection Library
Why do libraries need to use AI? It is crucial for Libraries to stay relevant in this digital age by improving efficiency, access, and user experience. By adopting AI, libraries can better manage growing digital collections, provide innovative services, and ensuring they remain essential as hubs for knowledge and learning in an AIdriven world.
Inside Out: Improving Large Model Safety, Wei Zhao
Inside Out: Improving Large Model Safety, Wei Zhao
Dissertations and Theses Collection (Open Access)
While Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) are at the frontier of current advancements in artificial intelligence, demonstrating remarkable capabilities across diverse applications, there are growing concerns about their reliability and security. LLMs remain vulnerable to adversarial attacks through carefully crafted prompts that circumvent safety mechanisms, while MLLMs face additional security challenges stemming from their multimodal nature. Despite considerable efforts in reinforcement learning from human feedback (RLHF) and supervised fine-tuning, existing safeguards have proven inadequate in addressing these critical vulnerabilities. This inadequacy stems from the fact that these models are inherently blackboxes that do not provide …
Constrained Reinforcement Learning: From Single-Agent Safety To Multi-Agent Coordination, Hao Jiang
Constrained Reinforcement Learning: From Single-Agent Safety To Multi-Agent Coordination, Hao Jiang
Dissertations and Theses Collection (Open Access)
Real-world decision-making systems such as autonomous driving and largescale ride-pooling must operate under strict safety and resource constraints. Traditional Reinforcement Learning (RL) methods, while powerful in simulation, often fail to guarantee such constraints, limiting their real-world deployment. The fundamental challenge lies in integrating constraint satisfaction with long-term reward optimization, especially when outcomes are stochastic and interdependent across multiple agents.
This dissertation advances the field of Constrained Reinforcement Learning (CRL) from both single-agent safety and multi-agent coordination perspectives. In the single-agent setting, we introduce a Reward Penalty framework that augments the state space with cumulative cost and penalizes only trajectories that …
Artificial Intelligence, Fundamental Motives, And Evolutionary Mismatch, Amy J. Lim, Jose. C. Yong, Edison Sora Tan
Artificial Intelligence, Fundamental Motives, And Evolutionary Mismatch, Amy J. Lim, Jose. C. Yong, Edison Sora Tan
Research Collection School of Social Sciences
In recent years, the intersection of artificial intelligence (AI) and psychology has garnered unprecedented attention, particularly following the advent of generative AI tools in 2022. These tools, capable of producing human-like text, images, and even deepening our understanding of cognitive processes, have not only captured the public imagination but also sparked new concerns and debates within the psychological community. While AI has been a subject of research for decades, the emergence of its generative capabilities has truly thrust AI into the spotlight. This article explores how these advancements are reshaping our understanding of human cognition and behavior, as well as …
Fluid Agency In Ai Systems: A Case For Functional Equivalence In Copyright, Patent, And Tort, Anirban Mukherjee, Hannah H. Chang
Fluid Agency In Ai Systems: A Case For Functional Equivalence In Copyright, Patent, And Tort, Anirban Mukherjee, Hannah H. Chang
Research Collection Lee Kong Chian School Of Business
Modern Artificial Intelligence (AI) systems exhibit fluid agency in multi-step workflows: lacking human-like consciousness or culpability, yet they display behavior that is (i) stochastic (probabilistic and path‑dependent), (ii) dynamic (co‑evolving with user interaction), and (iii) adaptive (able to reorient across contexts). These properties generate valuable outputs but collapse attribution, irreducibly entangling human and machine inputs. Doctrines that assume traceable provenance—authorship, inventorship, and liability—fracture under this unmappability, yielding ownership gaps and moral “crumple zones.”This Article argues that only functional equivalence stabilizes doctrine under unmappability: Where provenance is indeterminate, legal frameworks should treat human and AI contributions as equivalent for allocating rights …
Benchmarking Gaslighting Negation Attacks Against Reasoning Models, Bin Zhu, Hailong Yin, Jingjing Chen, Yu Gang Jiang
Benchmarking Gaslighting Negation Attacks Against Reasoning Models, Bin Zhu, Hailong Yin, Jingjing Chen, Yu Gang Jiang
Research Collection School Of Computing and Information Systems
Recent advances in reasoning-centric models promise improved robustness through mechanisms such as chain-of-thought prompting and test-time scaling. However, their ability to withstand gaslighting negation attacks—adversarial prompts that confidently deny correct answers—remains underexplored. In this paper, we conduct a systematic evaluation of three state-of-the-art reasoning models, i.e., OpenAI’s o4-mini, Claude-3.7-Sonnet and Gemini-2.5-Flash, across three multimodal benchmarks: MMMU, MathVista, and CharXiv. Our evaluation reveals significant accuracy drops (25–29% on average) following gaslighting negation attacks, indicating that even top-tier reasoning models struggle to preserve correct answers under manipulative user feedback. Built upon the insights of the evaluation and to further probe this vulnerability, …
Integrating Symbolic And Waveform Music Into Large Language Models, Teng Tu, Xiaohao Liu, Yunshan Ma, Ji Qi, Tat-Seng Chua
Integrating Symbolic And Waveform Music Into Large Language Models, Teng Tu, Xiaohao Liu, Yunshan Ma, Ji Qi, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Music, as a unique and integral element of human life, is characterized by its complex structures, intricate details, and the fusion of multimodal information. Recent study advance music understanding by leveraging knowledge and reasoning capabilities derived from Large Language Models (LLMs). However, they often lack compatibility and fail to fully utilize the complementary strengths of diverse representations (e.g., ABC, MIDI, Waveform). To address these limitations, we propose a unified music-language model framework, named UniMuLM, transitioning from single-representation approaches to the integration of multiple music representations for LLM. Unifying different music representation formats poses challenges such as patch integrity and boundary …
Food Recognition With Visual Language Models: Search Re-Ranking Or Retrieval-Augmented Generation?, Kian Yu Gan, Phuong Anh Nguyen, Chong-Wah Ngo
Food Recognition With Visual Language Models: Search Re-Ranking Or Retrieval-Augmented Generation?, Kian Yu Gan, Phuong Anh Nguyen, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Despite the rapid advances in Visual Language Models (VLMs), these models struggle to recognize culture-specific food items. While VLMs are effective in recognizing popular cultural dishes, their performance is suboptimal for dishes that are unique but not widely known internationally. Specifically, VLMs often generate either generic labels or hallucinated names for dishes that are localized to a particular culture. As a result, retrieval-augmented generation (RAG), which retrieves relevant recipes as references for VLMs, emerges as a promising approach. Nevertheless, recipe retrieval, which is itself imperfect, could mislead VLMs into generating inaccurate or culturally inappropriate dish names. This paper presents a …
Analysis Theories On Artificial Intelligence, Chatgpt, Data Science, And Metaverse: The Case Of Digital Medicine, Yin Yang, Xingyun Liu, Jorge Luis Cuyubamba Dominguez, Yuan Fang, Wen Xie, Bairong Shen, Keng Siau
Analysis Theories On Artificial Intelligence, Chatgpt, Data Science, And Metaverse: The Case Of Digital Medicine, Yin Yang, Xingyun Liu, Jorge Luis Cuyubamba Dominguez, Yuan Fang, Wen Xie, Bairong Shen, Keng Siau
Research Collection School Of Computing and Information Systems
Healthcare organizations are increasingly adopting digital technologies, with Artificial Intelligence (AI), Data Science, and the metaverse driving significant advancements in smart healthcare. Al facilitates personalized medicine and efficient drug development, while Data Science enables predictive analytics and big data management, enhancing patient outcomes and healthcare quality. The metaverse introduces immersive training and telemedicine platforms, revolutionizing patient engagement and healthcare research. This study conducts' a scoping review of 6,171 articles, analyzing the transformational impact of AI, ChatGPT, Data Science, and the metaverse on healthcare. It highlights the benefits and risks of these technologies, identifies research gaps in their application within the …
Editorial: Special Section On Challenges And Opportunities In Retrieval-Augmented Generation For Llms: Techniques, Trends, And Applications, Philip S. Yu, Haofen Wang, Feida Zhu
Editorial: Special Section On Challenges And Opportunities In Retrieval-Augmented Generation For Llms: Techniques, Trends, And Applications, Philip S. Yu, Haofen Wang, Feida Zhu
Research Collection School Of Computing and Information Systems
Retrieval-Augmented Generation (RAG) represents a transformative advancement for Large Language Models (LLMs) by integrating external knowledge to substantially improve accuracy and mitigate hallucinations. As a pivotal technology in the contemporary generative Artificial Intelligence (AI) landscape, RAG addresses fundamental challenges in knowledge-intensive tasks. This special issue serves as a dedicated platform to showcase these cutting-edge advancements. It features six rigorously peer-reviewed papers that present state-of-the-art research and applications in the rapidly evolving field of RAG.
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Research Collection School Of Computing and Information Systems
The rapid integration of Large Language Models (LLMs) into software engineering (SE) has revolutionized tasks from code generation to program repair, producing a massive volume of software artifacts. This surge in automated creation has exposed a critical bottleneck: the lack of scalable and reliable methods to evaluate the quality of these outputs. Human evaluation, while effective, is very costly and time-consuming. Traditional automated metrics like BLEU rely on high-quality references and struggle to capture nuanced aspects of software quality, such as readability and usefulness. In response, the LLM-as-a-Judge paradigm, which employs LLMs for automated evaluation, has emerged. This approach leverages …
Scaling Up Cooperative Multi-Agent Reinforcement Learning Through Hierarchical Heterogeneous Modular Architectures, Minghong Geng
Scaling Up Cooperative Multi-Agent Reinforcement Learning Through Hierarchical Heterogeneous Modular Architectures, Minghong Geng
Research Collection School Of Computing and Information Systems
Multi-agent reinforcement learning enables sophisticated collaborative behaviors in autonomous systems, yet fundamental scalability barriers persist: existing methods struggle to coordinate large agent populations and face challenges with extended decision-making horizons. This research develops hierarchical approaches to scale up multi-agent learning systems through two complementary directions: structural scaling for coordinating increasing numbers of agents and temporal scaling for extending decision-making horizons. This paper presents four integrated contributions: a taxonomic survey establishing hierarchical architectures as the theoretical foundation for scalable multi-agent learning systems, a benchmark for long-horizon multi-objective multi-agent reinforcement learning, a framework integrating self-organizing neural networks with multiple reinforcement learning agents …
Artem: Enhancing Large Language Model Agents With Spatial-Temporal Episodic Memory, Cassandra Hui Ming Tan, Budhitama Subagdja, Ah-Hwee Tan
Artem: Enhancing Large Language Model Agents With Spatial-Temporal Episodic Memory, Cassandra Hui Ming Tan, Budhitama Subagdja, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Current large language models (LLMs) exhibit significant deficiencies in episodic memory tasks including encoding, storing, and retrieving specific information from temporally dependent events over a long period of time. Recent approaches to handle memory tasks in LLMs, such as in-context learning, retrieval-augmented generation (RAG), and fine-tuning, may resolve the long-term retention issues, but are still inadequate to handle tasks requiring chronological awareness of the stored information. We introduce Agentic Retrieval with Temporal-Episodic Memory (ARTEM), a hybrid LLM-based agent architecture integrating LLMs with a self-organizing neural network named Spatial-Temporal Episodic Memory (STEM), designed to handle episodic memory tasks. Our approach employs …
Dystop: Dynamic Staleness Control And Topology Construction For Asynchronous Decentralized Federated Learning, Yizhou Shi, Qianpiao Ma, Yan Xu, Junlong Zhou, Ming Hu, Yunming Liao
Dystop: Dynamic Staleness Control And Topology Construction For Asynchronous Decentralized Federated Learning, Yizhou Shi, Qianpiao Ma, Yan Xu, Junlong Zhou, Ming Hu, Yunming Liao
Research Collection School Of Computing and Information Systems
Federated Learning (FL) has emerged as a potential distributed learning paradigm that enables model training on edge devices (i.e., workers) while preserving data privacy. However, its reliance on a centralized server leads to limited scalability. Decentralized federated learning (DFL) eliminates the dependency on a centralized server by enabling peer-to-peer model exchange. Existing DFL mechanisms mainly employ synchronous communication, which may result in training inefficiencies under heterogeneous and dynamic edge environments. Although a few recent asynchronous DFL (ADFL) mechanisms have been proposed to address these issues, they typically yield stale model aggregation and frequent model transmission, leading to degraded training performance …
Llamoco: Instruction Tuning Of Large Language Models For Optimization Code Generation, Zeyuan Ma, Yue-Jiao Gong, Hongshu Guo, Jiacheng Chen, Yining Ma, Zhiguang Cao
Llamoco: Instruction Tuning Of Large Language Models For Optimization Code Generation, Zeyuan Ma, Yue-Jiao Gong, Hongshu Guo, Jiacheng Chen, Yining Ma, Zhiguang Cao
Research Collection School Of Computing and Information Systems
Recently, combining the strength of large language models (LLMs) and Evolutionary Computation (EC) has shown promising results for addressing optimization problems. It typically involves either iterative next-step solution seeking or directly prompting LLMs to generate critical optimization codes. However, these methods often suffer from low computational efficiency, high sensitivity to prompt design, and a lack of domain-specific knowledge. We introduce LLaMoCo, the first instruction-tuning framework designed to adapt LLMs for solving optimization problems in a code-to-code manner. LLaMoCo features a comprehensive instruction set that includes code-style problem descriptions as input prompts and robust optimization codes from expert EC optimizers as …
Revisiting The Canonicalization For Fast And Accurate Crystal Tensor Property Prediction, Haowei Hua Hua, Jingwen Yang, Wanyu Lin, Pan Zhou
Revisiting The Canonicalization For Fast And Accurate Crystal Tensor Property Prediction, Haowei Hua Hua, Jingwen Yang, Wanyu Lin, Pan Zhou
Research Collection School Of Computing and Information Systems
Predicting the tensor properties of crystalline materials is a fundamental task in materials science. Unlike single-value property prediction, which is inherently invariant, tensor property prediction requires maintaining O(3) group tensor equivariance. Such equivariance constraint often requires specialized architecture designs to achieve effective predictions, inevitably introducing tremendous computational costs. Canonicalization, a classical technique for geometry, has recently been explored for efficient learning with symmetry. In this work, we revisit the problem of crystal tensor property prediction through the lens of canonicalization. Specifically, we demonstrate how polar decomposition, a simple yet efficient algebraic method, can serve as a form of canonicalization and …
Realign: Text-To-Motion Generation Via Step-Aware Reward-Guided Alignment, Wanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie, Pan Zhou, Hongsong Wang
Realign: Text-To-Motion Generation Via Step-Aware Reward-Guided Alignment, Wanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie, Pan Zhou, Hongsong Wang
Research Collection School Of Computing and Information Systems
Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text and motion distributions in diffusion models, which leads to semantically inconsistent or low-quality motions. To address this limitation, we propose Reward-guided sampling Alignment (ReAlign), comprising a step-aware reward model to assess alignment quality during the denoising sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a …
Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin
Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin
Research Collection School Of Computing and Information Systems
Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a …
Tempo: Training-Time Equilibration Of Modalities For Per-Sample Optimization In Multimodal Sentiment, Yi Zhao, Erik Cambria, Xiaosong E, Xianxun Zhu
Tempo: Training-Time Equilibration Of Modalities For Per-Sample Optimization In Multimodal Sentiment, Yi Zhao, Erik Cambria, Xiaosong E, Xianxun Zhu
Research Collection School Of Computing and Information Systems
Multimodal sentiment models often become over-reliant on the “easiest” modality (typically text), leading to three coupled sub-problems: (i) representation-level dominance, where weaker modalities contribute little to the fused representation; (ii) optimization-level dominance, where the strongest modality drives most gradient updates and suppresses learning in others; and (iii) robustness degradation, where audio or vision fail under noise or missing inputs at test time. We present TEMPO, a plug-and-play training framework that mitigates these issues by rebalancing learning pressure across modalities while leaving inference unchanged. For each mini-batch, TEMPO estimates relative modality strength and applies two synchronized, training-only controls: selective forward attenuation …
Airaclex: Automated Detection Of Price Oracle Manipulations Via Llm-Driven Knowledge Mining And Prompt Generation, Bo Gao, Yuan Wang, Qingsong Wei, Yong Liu, Rick Siow Mong Goh, David Lo
Airaclex: Automated Detection Of Price Oracle Manipulations Via Llm-Driven Knowledge Mining And Prompt Generation, Bo Gao, Yuan Wang, Qingsong Wei, Yong Liu, Rick Siow Mong Goh, David Lo
Research Collection School Of Computing and Information Systems
Decentralized finance (DeFi) applications depend on accurate price oracles to ensure secure and fair transactions. However, poorly integrated oracles remain susceptible to manipulation, enabling attackers to exploit smart contract logic for unfair asset valuation and financial gain. While many such vulnerabilities are only detected after deployment, smart contracts are typically immutable once deployed, making post-hoc fixes costly or infeasible. This highlights the critical need for detecting oracle manipulation risks before deployment. In this paper, we propose AiRacleX, a novel LLM-driven framework that enables pre-deployment detection of price oracle manipulation vulnerabilities by leveraging the complementary strengths of multiple large language models …
Purified Zero-Shot Sketch-Based Image Retrieval, Yang Zhou, Jingru Yang, Jin Wang, Kaixiang Huang, Guodong Lu, Shengfeng He
Purified Zero-Shot Sketch-Based Image Retrieval, Yang Zhou, Jingru Yang, Jin Wang, Kaixiang Huang, Guodong Lu, Shengfeng He
Research Collection School Of Computing and Information Systems
Sketches, as a new solution in multimedia systems that can replace natural language, are characterized by sparse visual cues such as simple strokes that differ significantly from natural images containing complex elements such as background, foreground, and texture. This misalignment poses substantial challenges for zero-shot sketch-based image retrieval (ZS-SBIR). Prior approaches match sketches to full images and tend to overlook redundant elements in natural images, leading to model distraction and semantic ambiguity. To address this issue, we introduce a distraction-agnostic framework, purified cross-domain matching (PuXIM), which operates on a straightforward principle: masking and matching. We devise a visual-cross-linguistic (VxL) sampler …
Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng
Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng
Research Collection School Of Computing and Information Systems
This work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common-place strokes, which often lead to less aesthetic outcomes. Inspired by the human painting process, i.e., observing, comparing, and drawing, we incorporate differential image analysis into a neural oil painting model, allowing the model to effectively concentrate on the incremental impact of successive brushstrokes. To operationalize this concept, we propose the Differential Query Transformer (DQ-Transformer), a new architecture that leverages differentially derived image representations enriched with positional encoding to guide the stroke …
Portfoliopilot: An Agentic Platform For Financial Portfolio Management Algorithm Development And Evaluation, Jared Chan Xu Yang, Haokai Ma, Yunshan Ma
Portfoliopilot: An Agentic Platform For Financial Portfolio Management Algorithm Development And Evaluation, Jared Chan Xu Yang, Haokai Ma, Yunshan Ma
Research Collection School Of Computing and Information Systems
Developing new portfolio-management algorithms typically demands substantial programming effort, limiting rapid experimentation and excluding finance professionals without coding skills. Current robo-advisory tools offer pre-built but rigid strategies, restricting customization and experimentation. We introduce PortfolioPilot, an open-source, agentic platform that enables users to generate bespoke portfolio through natural-language descriptions. Leveraging the Anthropic Claude API, PortfolioPilot dynamically synthesizes executable TypeScript algorithms that run in the frontend with security validation. The system integrates real-time backtesting with historical market data, classical optimization algorithms (Markowitz, LSTM, ARIMA), and interactive performance visualizations.