Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead,
2026
Singapore Management University
Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo
Research Collection School Of Computing and Information Systems
The rapid integration of Large Language Models (LLMs) into software engineering (SE) has revolutionized tasks from code generation to program repair, producing a massive volume of software artifacts. This surge in automated creation has exposed a critical bottleneck: the lack of scalable and reliable methods to evaluate the quality of these outputs. Human evaluation, while effective, is very costly and time-consuming. Traditional automated metrics like BLEU rely on high-quality references and struggle to capture nuanced aspects of software quality, such as readability and usefulness. In response, the LLM-as-a-Judge paradigm, which employs LLMs for automated evaluation, has emerged. This approach leverages …
Scaling Up Cooperative Multi-Agent Reinforcement Learning Through Hierarchical Heterogeneous Modular Architectures,
2026
Singapore Management University
Scaling Up Cooperative Multi-Agent Reinforcement Learning Through Hierarchical Heterogeneous Modular Architectures, Minghong Geng
Research Collection School Of Computing and Information Systems
Multi-agent reinforcement learning enables sophisticated collaborative behaviors in autonomous systems, yet fundamental scalability barriers persist: existing methods struggle to coordinate large agent populations and face challenges with extended decision-making horizons. This research develops hierarchical approaches to scale up multi-agent learning systems through two complementary directions: structural scaling for coordinating increasing numbers of agents and temporal scaling for extending decision-making horizons. This paper presents four integrated contributions: a taxonomic survey establishing hierarchical architectures as the theoretical foundation for scalable multi-agent learning systems, a benchmark for long-horizon multi-objective multi-agent reinforcement learning, a framework integrating self-organizing neural networks with multiple reinforcement learning agents …
Artem: Enhancing Large Language Model Agents With Spatial-Temporal Episodic Memory,
2026
Singapore Management University
Artem: Enhancing Large Language Model Agents With Spatial-Temporal Episodic Memory, Cassandra Hui Ming Tan, Budhitama Subagdja, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Current large language models (LLMs) exhibit significant deficiencies in episodic memory tasks including encoding, storing, and retrieving specific information from temporally dependent events over a long period of time. Recent approaches to handle memory tasks in LLMs, such as in-context learning, retrieval-augmented generation (RAG), and fine-tuning, may resolve the long-term retention issues, but are still inadequate to handle tasks requiring chronological awareness of the stored information. We introduce Agentic Retrieval with Temporal-Episodic Memory (ARTEM), a hybrid LLM-based agent architecture integrating LLMs with a self-organizing neural network named Spatial-Temporal Episodic Memory (STEM), designed to handle episodic memory tasks. Our approach employs …
Dystop: Dynamic Staleness Control And Topology Construction For Asynchronous Decentralized Federated Learning,
2026
Singapore Management University
Dystop: Dynamic Staleness Control And Topology Construction For Asynchronous Decentralized Federated Learning, Yizhou Shi, Qianpiao Ma, Yan Xu, Junlong Zhou, Ming Hu, Yunming Liao
Research Collection School Of Computing and Information Systems
Federated Learning (FL) has emerged as a potential distributed learning paradigm that enables model training on edge devices (i.e., workers) while preserving data privacy. However, its reliance on a centralized server leads to limited scalability. Decentralized federated learning (DFL) eliminates the dependency on a centralized server by enabling peer-to-peer model exchange. Existing DFL mechanisms mainly employ synchronous communication, which may result in training inefficiencies under heterogeneous and dynamic edge environments. Although a few recent asynchronous DFL (ADFL) mechanisms have been proposed to address these issues, they typically yield stale model aggregation and frequent model transmission, leading to degraded training performance …
Llamoco: Instruction Tuning Of Large Language Models For Optimization Code Generation,
2026
Singapore Management University
Llamoco: Instruction Tuning Of Large Language Models For Optimization Code Generation, Zeyuan Ma, Yue-Jiao Gong, Hongshu Guo, Jiacheng Chen, Yining Ma, Zhiguang Cao
Research Collection School Of Computing and Information Systems
Recently, combining the strength of large language models (LLMs) and Evolutionary Computation (EC) has shown promising results for addressing optimization problems. It typically involves either iterative next-step solution seeking or directly prompting LLMs to generate critical optimization codes. However, these methods often suffer from low computational efficiency, high sensitivity to prompt design, and a lack of domain-specific knowledge. We introduce LLaMoCo, the first instruction-tuning framework designed to adapt LLMs for solving optimization problems in a code-to-code manner. LLaMoCo features a comprehensive instruction set that includes code-style problem descriptions as input prompts and robust optimization codes from expert EC optimizers as …
Revisiting The Canonicalization For Fast And Accurate Crystal Tensor Property Prediction,
2026
Singapore Management University
Revisiting The Canonicalization For Fast And Accurate Crystal Tensor Property Prediction, Haowei Hua Hua, Jingwen Yang, Wanyu Lin, Pan Zhou
Research Collection School Of Computing and Information Systems
Predicting the tensor properties of crystalline materials is a fundamental task in materials science. Unlike single-value property prediction, which is inherently invariant, tensor property prediction requires maintaining O(3) group tensor equivariance. Such equivariance constraint often requires specialized architecture designs to achieve effective predictions, inevitably introducing tremendous computational costs. Canonicalization, a classical technique for geometry, has recently been explored for efficient learning with symmetry. In this work, we revisit the problem of crystal tensor property prediction through the lens of canonicalization. Specifically, we demonstrate how polar decomposition, a simple yet efficient algebraic method, can serve as a form of canonicalization and …
Realign: Text-To-Motion Generation Via Step-Aware Reward-Guided Alignment,
2026
Singapore Management University
Realign: Text-To-Motion Generation Via Step-Aware Reward-Guided Alignment, Wanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie, Pan Zhou, Hongsong Wang
Research Collection School Of Computing and Information Systems
Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text and motion distributions in diffusion models, which leads to semantically inconsistent or low-quality motions. To address this limitation, we propose Reward-guided sampling Alignment (ReAlign), comprising a step-aware reward model to assess alignment quality during the denoising sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a …
Reinforce Trustworthiness In Multimodal Emotional Support System,
2026
Singapore Management University
Reinforce Trustworthiness In Multimodal Emotional Support System, Huy M. Le, Dat Tien Nguyen, Ngan T. T. Vo, Tuan D. Q. Nguyen, Nguyen Le Binh, Duy Minh Ho Nguyen, Daniel Sonntag, Lizi Liao, Binh T. Nguyen
Research Collection School Of Computing and Information Systems
In today's world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources to provide empathetic, contextually relevant responses, fostering more effective interactions. However, current methods have notable limitations, often relying solely on text or converting other data types into text, or providing emotion recognition only, thus overlooking the full potential of multimodal inputs. Moreover, many studies prioritize response generation without accurately identifying critical emotional support elements or ensuring the reliability of outputs. To overcome these issues, we introduce …
Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models,
2026
Singapore Management University
Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin
Research Collection School Of Computing and Information Systems
Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a …
Tempo: Training-Time Equilibration Of Modalities For Per-Sample Optimization In Multimodal Sentiment,
2026
Singapore Management University
Tempo: Training-Time Equilibration Of Modalities For Per-Sample Optimization In Multimodal Sentiment, Yi Zhao, Erik Cambria, Xiaosong E, Xianxun Zhu
Research Collection School Of Computing and Information Systems
Multimodal sentiment models often become over-reliant on the “easiest” modality (typically text), leading to three coupled sub-problems: (i) representation-level dominance, where weaker modalities contribute little to the fused representation; (ii) optimization-level dominance, where the strongest modality drives most gradient updates and suppresses learning in others; and (iii) robustness degradation, where audio or vision fail under noise or missing inputs at test time. We present TEMPO, a plug-and-play training framework that mitigates these issues by rebalancing learning pressure across modalities while leaving inference unchanged. For each mini-batch, TEMPO estimates relative modality strength and applies two synchronized, training-only controls: selective forward attenuation …
Airaclex: Automated Detection Of Price Oracle Manipulations Via Llm-Driven Knowledge Mining And Prompt Generation,
2026
Singapore Management University
Airaclex: Automated Detection Of Price Oracle Manipulations Via Llm-Driven Knowledge Mining And Prompt Generation, Bo Gao, Yuan Wang, Qingsong Wei, Yong Liu, Rick Siow Mong Goh, David Lo
Research Collection School Of Computing and Information Systems
Decentralized finance (DeFi) applications depend on accurate price oracles to ensure secure and fair transactions. However, poorly integrated oracles remain susceptible to manipulation, enabling attackers to exploit smart contract logic for unfair asset valuation and financial gain. While many such vulnerabilities are only detected after deployment, smart contracts are typically immutable once deployed, making post-hoc fixes costly or infeasible. This highlights the critical need for detecting oracle manipulation risks before deployment. In this paper, we propose AiRacleX, a novel LLM-driven framework that enables pre-deployment detection of price oracle manipulation vulnerabilities by leveraging the complementary strengths of multiple large language models …
Purified Zero-Shot Sketch-Based Image Retrieval,
2026
Singapore Management University
Purified Zero-Shot Sketch-Based Image Retrieval, Yang Zhou, Jingru Yang, Jin Wang, Kaixiang Huang, Guodong Lu, Shengfeng He
Research Collection School Of Computing and Information Systems
Sketches, as a new solution in multimedia systems that can replace natural language, are characterized by sparse visual cues such as simple strokes that differ significantly from natural images containing complex elements such as background, foreground, and texture. This misalignment poses substantial challenges for zero-shot sketch-based image retrieval (ZS-SBIR). Prior approaches match sketches to full images and tend to overlook redundant elements in natural images, leading to model distraction and semantic ambiguity. To address this issue, we introduce a distraction-agnostic framework, purified cross-domain matching (PuXIM), which operates on a straightforward principle: masking and matching. We devise a visual-cross-linguistic (VxL) sampler …
Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting,
2026
Singapore Management University
Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng
Research Collection School Of Computing and Information Systems
This work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common-place strokes, which often lead to less aesthetic outcomes. Inspired by the human painting process, i.e., observing, comparing, and drawing, we incorporate differential image analysis into a neural oil painting model, allowing the model to effectively concentrate on the incremental impact of successive brushstrokes. To operationalize this concept, we propose the Differential Query Transformer (DQ-Transformer), a new architecture that leverages differentially derived image representations enriched with positional encoding to guide the stroke …
Potent But Stealthy: Rethink Profile Pollution Against Sequential Recommendation Via Bi-Level Constrained Reinforcement Paradigm,
2026
Singapore Management University
Potent But Stealthy: Rethink Profile Pollution Against Sequential Recommendation Via Bi-Level Constrained Reinforcement Paradigm, Jiajie Su, Zihan Nan, Yunshan Ma, Xiaobo Xia, Xiaohua Feng, Weiming Liu, Xiang Chen, Xiaolin Zheng, Chaochao Chen
Research Collection School Of Computing and Information Systems
Sequential Recommenders, which exploit dynamic user intents through interaction sequences, are vulnerable to adversarial attacks. While existing attacks primarily rely on data poisoning, they require large-scale user access or fake profiles, thus lacking practicality. In this paper, we focus on the Profile Pollution Attack that subtly contaminates partial user interactions to induce targeted mispredictions. Previous PPA methods suffer from two limitations, i.e., i) overreliance on sequence horizon impact restricts fine-grained perturbations on item transitions, and ii) holistic modifications cause detectable distribution shifts. To address these challenges, we propose a constrained reinforcement driven attack CREAT that synergizes a bi-level optimization framework …
Nondeterministic Polynomial-Time Problem Challenge: An Ever-Scaling Reasoning Benchmark For Llms,
2026
Singapore Management University
Nondeterministic Polynomial-Time Problem Challenge: An Ever-Scaling Reasoning Benchmark For Llms, Chang Yang, Ruiyu Wang, Junzhe Jiang, Qi Jiang, Qinggang Zhang, Yanchen Deng, Shuxin Li, Shuyue Hu, Bo Li, Florian T. Pokorny, Xiao Huang, Xinrun Wang
Research Collection School Of Computing and Information Systems
Reasoning is the fundamental capability of large language models (LLMs). Due to the rapid progress of LLMs, there are two main issues of current benchmarks: i) these benchmarks can be crushed in a short time (less than 1 year), and ii) these benchmarks may be easily hacked. To handle these issues, we propose the ever-scalingness for building the benchmarks which are scaling over complexity against crushing, instance against hacking and exploitation, oversight for easy verification, and coverage for real-world relevance. This paper presents Nondeterministic Polynomial-time Problem Challenge (NPPC), an ever-scaling reasoning benchmark for LLMs. Specifically, the NPPC has three main …
Portfoliopilot: An Agentic Platform For Financial Portfolio Management Algorithm Development And Evaluation,
2026
Singapore Management University
Portfoliopilot: An Agentic Platform For Financial Portfolio Management Algorithm Development And Evaluation, Jared Chan Xu Yang, Haokai Ma, Yunshan Ma
Research Collection School Of Computing and Information Systems
Developing new portfolio-management algorithms typically demands substantial programming effort, limiting rapid experimentation and excluding finance professionals without coding skills. Current robo-advisory tools offer pre-built but rigid strategies, restricting customization and experimentation. We introduce PortfolioPilot, an open-source, agentic platform that enables users to generate bespoke portfolio through natural-language descriptions. Leveraging the Anthropic Claude API, PortfolioPilot dynamically synthesizes executable TypeScript algorithms that run in the frontend with security validation. The system integrates real-time backtesting with historical market data, classical optimization algorithms (Markowitz, LSTM, ARIMA), and interactive performance visualizations.
Interpretable Multimodal Zero Shot Ecg Diagnosis Via Structured Clinical Knowledge Alignment,
2026
Singapore Management University
Interpretable Multimodal Zero Shot Ecg Diagnosis Via Structured Clinical Knowledge Alignment, Jialu Tang, Hung Manh Pham, Ignace De Lathauwer, Henk S. Schipper, Yuan Lu, Dong Ma, Aaqib Saeed
Research Collection School Of Computing and Information Systems
Electrocardiogram (ECG) interpretation is essential for cardiovascular disease diagnosis, but current automated systems often struggle with transparency and generalization to unseen conditions. To address this, we introduce ZETA, a zero-shot multimodal framework designed for interpretable ECG diagnosis aligned with clinical workflows. ZETA uniquely compares ECG signals against structured positive and negative clinical observations, which are curated through an LLM-assisted, expertvalidated process, thereby mimicking differential diagnosis. Our approach leverages a pre-trained multimodal model to align ECG and text embeddings without disease-specific fine-tuning. Empirical evaluations demonstrate ZETA’s competitive zero-shot classification performance and, importantly, provide qualitative and quantitative evidence of enhanced interpretability, grounding …
Patterns Of Llm Weaponization: A Comparative Analysis Of Exploitation Incidents Across Commercial Ai Systems,
2025
Lynn University
Patterns Of Llm Weaponization: A Comparative Analysis Of Exploitation Incidents Across Commercial Ai Systems, George Antoniou
Faculty and Staff Publications & Presentations
This comparative study examines patterns of Large Language Model (LLM) weaponization through systematic analysis of four major exploitation incidents spanning from 2023-2025. While existing research focuses on isolated incidents or theoretical vulnerabilities, this study provides one of the first comprehensive comparative frameworks analyzing exploitation patterns across state-sponsored cyber-espionage (Anthropic Claude incident), academic security research (GPT- 4 autonomous privilege escalation), social engineering platforms (SpearBot phishing framework), and underground criminal commoditization (WormGPT/FraudGPT ecosystem). Through comparative analysis across eight dimensions: Adversary sophistication, target selection, exploitation techniques, autonomy levels, detection evasion, attribution challenges, defensive gaps, and capability democratization, this research identifies critical cross-case patterns …
Quantum Readiness In Cybersecurity Education: A Framework For Preparing The Next Generation In The Post-Quantum Era,
2025
Lynn University
Quantum Readiness In Cybersecurity Education: A Framework For Preparing The Next Generation In The Post-Quantum Era, George Antoniou
Faculty and Staff Publications & Presentations
This framework addresses the critical gap between post-quantum standards and workforce readiness. Shor's algorithm demonstrates that sufficiently powerful quantum computers can break the cryptographic foundations of internet security. While the cryptography research community has developed quantum-resistant algorithms, educational institutions have not prepared students to implement these solutions. Recent surveys show fewer than half of organizations have begun planning for post-quantum cryptography (PQC) transitions (Entrust Cybersecurity Institute, 2024; U.S. Government Accountability Office, 2023; (ISC)², 2024). The NICE Framework (Newhouse, Keith, Scribner, & Witte, 2017) outlines the knowledge and skills that cybersecurity professionals should possess. The framework omits post-quantum cryptography entirely. Organizations …
From Enhancement To Substitution: A Strategic Provocation On Simulation-Based Sport,
2025
Baylor University
From Enhancement To Substitution: A Strategic Provocation On Simulation-Based Sport, Grant B. Morgan, Andreas Stamatis
Journal of Applied Sport Management
Advances in artificial intelligence, large-scale machine learning, and simulation technologies are rapidly transforming how sport is played, analyzed, and consumed. To date, most scholarly and industry discussions frame these technologies as tools that enhance embodied sport by improving performance, officiating, media production, and fan engagement. This paper extends that conversation by posing a more provocative strategic question: under what conditions might simulation move from enhancement to substitution? Focusing explicitly on sport as a business and entertainment enterprise, we argue that many of sport’s core sources of cultural and economic value—uncertainty of outcome, narrative continuity, legitimacy, and collective meaning—are structurally …
