Open Access. Powered by Scholars. Published by Universities.®

2026

Discipline
Institution
Keyword
Publication
Publication Type
File Type

Articles 931 - 960 of 967

Full-Text Articles in Artificial Intelligence and Robotics

Interpretable Machine Learning For In-Home Mild Cognitive Impairment Detection, Budhitama Subagdja, Shanthoshigaa D, Ah-Hwee Tan, Iris Rawtaer Jan 2026

Interpretable Machine Learning For In-Home Mild Cognitive Impairment Detection, Budhitama Subagdja, Shanthoshigaa D, Ah-Hwee Tan, Iris Rawtaer

Research Collection School Of Computing and Information Systems

This paper introduces a novel system for in-home cognitive health assessment using ambient sensors and a machine learning technology that can robustly detect mild cognitive impairment (MCI) despite limited available data. The learned model can explain the aspects of individuals’ daily lives led to the prediction, while reliably predicting MCI, providing more insights to healthcare workers for further clinical interventions. We developed the robust transparent machine learning model, based on fusion adaptive resonance theory (Fusion ART) neural network to learn individuals’ daily patterns of activity from continuous sensor data in terms of a suite of digital biomarkers reflecting four key …


Food Recognition With Visual Language Models: Search Re-Ranking Or Retrieval-Augmented Generation?, Kian Yu Gan, Phuong Anh Nguyen, Chong-Wah Ngo Jan 2026

Food Recognition With Visual Language Models: Search Re-Ranking Or Retrieval-Augmented Generation?, Kian Yu Gan, Phuong Anh Nguyen, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Despite the rapid advances in Visual Language Models (VLMs), these models struggle to recognize culture-specific food items. While VLMs are effective in recognizing popular cultural dishes, their performance is suboptimal for dishes that are unique but not widely known internationally. Specifically, VLMs often generate either generic labels or hallucinated names for dishes that are localized to a particular culture. As a result, retrieval-augmented generation (RAG), which retrieves relevant recipes as references for VLMs, emerges as a promising approach. Nevertheless, recipe retrieval, which is itself imperfect, could mislead VLMs into generating inaccurate or culturally inappropriate dish names. This paper presents a …


Ai Tips And Traps, Patrick Barry Jan 2026

Ai Tips And Traps, Patrick Barry

Books

Based on a series of popular courses and workshops that Professor Patrick Barry has created for students, professionals, and anyone else interested in taking a skills-based approach to artificial intelligence, this book gives you a chance to engage with important AI concepts, experiment with exploratory AI exercises, and then ultimately develop your own customized list of AI traps to try as well as AI traps to avoid.


Multimodal Representation Learning For Face Understanding: From Caption Supervision To Foundation Model Adaptation, Md Mahedi Hasan Jan 2026

Multimodal Representation Learning For Face Understanding: From Caption Supervision To Foundation Model Adaptation, Md Mahedi Hasan

Graduate Theses, Dissertations, and Problem Reports (ETD)

The rapid advancement of intelligent surveillance systems and the increasing demand for reliable biometric identification in border security, public safety, and digital forensics require robust face understanding under unconstrained conditions, including low resolution, pose variation, and occlusion. While Vision Transformer (ViT)-based foundation models have greatly improved visual representation learning, their patch-based tokenization and lack of spatial inductive bias limit their ability to capture fine-grained details in low-resolution inputs. This dissertation investigates multimodal representation learning for face understanding through natural language supervision, large-scale face-caption pre-training, and parameter-efficient foundation model adaptation. It hypothesizes that textual supervision provides complementary semantic cues that improve …


Llm-Driven Weekly Newsletter To Assess Open Source Software Project Github Health, Christian Novalski, Christopher Chavez, Ghalian Fayyadh, Kostadin Damevski Jan 2026

Llm-Driven Weekly Newsletter To Assess Open Source Software Project Github Health, Christian Novalski, Christopher Chavez, Ghalian Fayyadh, Kostadin Damevski

UROP Posters

Open Source Software (OSS) projects increasingly depend on a diverse set of contributors, including episodic participants who contribute intermittently. Episodic contributors represent a large portion of OSS communities, yet projects often struggle to retain them, leading to decreased project health and continuity. While dashboards and real-time communication tools support continuously active contributors, they often fail to serve the unique needs of episodic participants, who may struggle to remain informed and re-engage with project activity after periods of absence. In this study, we examine the effect of a weekly, email-based newsletter intervention designed to improve awareness and engagement among episodic OSS …


Artem: Enhancing Large Language Model Agents With Spatial-Temporal Episodic Memory, Cassandra Hui Ming Tan, Budhitama Subagdja, Ah-Hwee Tan Jan 2026

Artem: Enhancing Large Language Model Agents With Spatial-Temporal Episodic Memory, Cassandra Hui Ming Tan, Budhitama Subagdja, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Current large language models (LLMs) exhibit significant deficiencies in episodic memory tasks including encoding, storing, and retrieving specific information from temporally dependent events over a long period of time. Recent approaches to handle memory tasks in LLMs, such as in-context learning, retrieval-augmented generation (RAG), and fine-tuning, may resolve the long-term retention issues, but are still inadequate to handle tasks requiring chronological awareness of the stored information. We introduce Agentic Retrieval with Temporal-Episodic Memory (ARTEM), a hybrid LLM-based agent architecture integrating LLMs with a self-organizing neural network named Spatial-Temporal Episodic Memory (STEM), designed to handle episodic memory tasks. Our approach employs …


Llamoco: Instruction Tuning Of Large Language Models For Optimization Code Generation, Zeyuan Ma, Yue-Jiao Gong, Hongshu Guo, Jiacheng Chen, Yining Ma, Zhiguang Cao Jan 2026

Llamoco: Instruction Tuning Of Large Language Models For Optimization Code Generation, Zeyuan Ma, Yue-Jiao Gong, Hongshu Guo, Jiacheng Chen, Yining Ma, Zhiguang Cao

Research Collection School Of Computing and Information Systems

Recently, combining the strength of large language models (LLMs) and Evolutionary Computation (EC) has shown promising results for addressing optimization problems. It typically involves either iterative next-step solution seeking or directly prompting LLMs to generate critical optimization codes. However, these methods often suffer from low computational efficiency, high sensitivity to prompt design, and a lack of domain-specific knowledge. We introduce LLaMoCo, the first instruction-tuning framework designed to adapt LLMs for solving optimization problems in a code-to-code manner. LLaMoCo features a comprehensive instruction set that includes code-style problem descriptions as input prompts and robust optimization codes from expert EC optimizers as …


Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin Jan 2026

Cross-Modal Proxy Evolving For Ood Detection With Vision-Language Models, Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin

Research Collection School Of Computing and Information Systems

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a …


Portfoliopilot: An Agentic Platform For Financial Portfolio Management Algorithm Development And Evaluation, Jared Chan Xu Yang, Haokai Ma, Yunshan Ma Jan 2026

Portfoliopilot: An Agentic Platform For Financial Portfolio Management Algorithm Development And Evaluation, Jared Chan Xu Yang, Haokai Ma, Yunshan Ma

Research Collection School Of Computing and Information Systems

Developing new portfolio-management algorithms typically demands substantial programming effort, limiting rapid experimentation and excluding finance professionals without coding skills. Current robo-advisory tools offer pre-built but rigid strategies, restricting customization and experimentation. We introduce PortfolioPilot, an open-source, agentic platform that enables users to generate bespoke portfolio through natural-language descriptions. Leveraging the Anthropic Claude API, PortfolioPilot dynamically synthesizes executable TypeScript algorithms that run in the frontend with security validation. The system integrates real-time backtesting with historical market data, classical optimization algorithms (Markowitz, LSTM, ARIMA), and interactive performance visualizations.


Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng Jan 2026

Look, Compare And Draw: Differential Query Transformer For Automatic Oil Painting, Lingyu Liu, Yaxiong Wang, Li Zhu, Lizi Liao, Zhedong Zheng

Research Collection School Of Computing and Information Systems

This work introduces a new approach to automatic oil painting that emphasizes the creation of dynamic and expressive brushstrokes. A pivotal challenge lies in mitigating the duplicate and common-place strokes, which often lead to less aesthetic outcomes. Inspired by the human painting process, i.e., observing, comparing, and drawing, we incorporate differential image analysis into a neural oil painting model, allowing the model to effectively concentrate on the incremental impact of successive brushstrokes. To operationalize this concept, we propose the Differential Query Transformer (DQ-Transformer), a new architecture that leverages differentially derived image representations enriched with positional encoding to guide the stroke …


Dyno : Dynamic Neurosymbolic Orchestrator For Multi-Agent Systems, Ritvik Garimella, Chathurangi Shyalika, Renjith Prasad, Amit Sheth Jan 2026

Dyno : Dynamic Neurosymbolic Orchestrator For Multi-Agent Systems, Ritvik Garimella, Chathurangi Shyalika, Renjith Prasad, Amit Sheth

Publications

Large Language Model (LLM)-based multi-agent systems (LaMAS) represent an emerging paradigm for tackling complex, multi-step reasoning and decision-making problems. As these systems scale, orchestration, which is the ability to coordinate, manage, and evaluate the interactions among diverse agents, becomes central to their success. While recent orchestrators such as AgentFlow have demonstrated promise in managing communication and task delegation, they remain limited in their ability to understand task semantics, coordinate heterogeneous agent types (e.g., reactive vs. cognitive), and adaptively align outputs with human-defined goals. In this position paper, we introduce the DYNO (Dynamic Neurosymbolic Orchestrator), a system developed as part of …


Deep Learning-Based Co-Current Upward Gas-Liquid Two-Phase Flow Regime Identification In An Annular Conduit, Joshua Robert Macomber Jan 2026

Deep Learning-Based Co-Current Upward Gas-Liquid Two-Phase Flow Regime Identification In An Annular Conduit, Joshua Robert Macomber

Graduate Theses, Dissertations, and Problem Reports (ETD)

Flow regime identification in co-current upward gas-liquid flow through annular conduits remains a significant challenge in petroleum engineering, with major safety and operational implications. It is also important across industries involving the transport of multiphase fluids. Misidentifying flow regimes can introduce major operational risk, yet regime boundaries in annular gas-liquid flow are often visually complex and context dependent.

The objective of this study was to evaluate the utility of convolutional neural network (CNN) classifiers for flow regime identification. The CNN was trained using annular flow image dataset published by Texas A&M University. The dataset consists of approximately 947 RGB images …


Sparse Gradient Training For Recommender Systems, Yunke Qu, Liang Qu, Tong Chen, Xiangyu Zhao, Jianxin Li, Hongzhi Yin Jan 2026

Sparse Gradient Training For Recommender Systems, Yunke Qu, Liang Qu, Tong Chen, Xiangyu Zhao, Jianxin Li, Hongzhi Yin

Research outputs 2022 to 2026

Recommender systems are widely applied in numerous online platforms such as shopping and social media platforms. They typically utilize large embedding tables that map users and items to dense vectors of uniform sizes. As the number of users and items continues to grow, this design leads to significant memory consumption and computational inefficiencies. This challenge is particularly pronounced in scenarios such as federated learning, where model parameters are updated locally on edge devices with limited computational resources before being transmitted to a central server for aggregation. Numerous approaches have been proposed to address this issue, among which embedding pruning methods …


Realign: Text-To-Motion Generation Via Step-Aware Reward-Guided Alignment, Wanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie, Pan Zhou, Hongsong Wang Jan 2026

Realign: Text-To-Motion Generation Via Step-Aware Reward-Guided Alignment, Wanjiang Weng, Xiaofeng Tan, Junbo Wang, Guo-Sen Xie, Pan Zhou, Hongsong Wang

Research Collection School Of Computing and Information Systems

Text-to-motion generation, which synthesizes 3D human motions from text inputs, holds immense potential for applications in gaming, film, and robotics. Recently, diffusion-based methods have been shown to generate more diversity and realistic motion. However, there exists a misalignment between text and motion distributions in diffusion models, which leads to semantically inconsistent or low-quality motions. To address this limitation, we propose Reward-guided sampling Alignment (ReAlign), comprising a step-aware reward model to assess alignment quality during the denoising sampling and a reward-guided strategy that directs the diffusion process toward an optimally aligned distribution. This reward model integrates step-aware tokens and combines a …


Airaclex: Automated Detection Of Price Oracle Manipulations Via Llm-Driven Knowledge Mining And Prompt Generation, Bo Gao, Yuan Wang, Qingsong Wei, Yong Liu, Rick Siow Mong Goh, David Lo Jan 2026

Airaclex: Automated Detection Of Price Oracle Manipulations Via Llm-Driven Knowledge Mining And Prompt Generation, Bo Gao, Yuan Wang, Qingsong Wei, Yong Liu, Rick Siow Mong Goh, David Lo

Research Collection School Of Computing and Information Systems

Decentralized finance (DeFi) applications depend on accurate price oracles to ensure secure and fair transactions. However, poorly integrated oracles remain susceptible to manipulation, enabling attackers to exploit smart contract logic for unfair asset valuation and financial gain. While many such vulnerabilities are only detected after deployment, smart contracts are typically immutable once deployed, making post-hoc fixes costly or infeasible. This highlights the critical need for detecting oracle manipulation risks before deployment. In this paper, we propose AiRacleX, a novel LLM-driven framework that enables pre-deployment detection of price oracle manipulation vulnerabilities by leveraging the complementary strengths of multiple large language models …


A Knowledge-Driven, Ai-Assisted Cyber Defence Framework For Iomt Remote Patient Monitoring, Kulsoom S. Bughio, David M. Cook, Abdul M. Unar Jan 2026

A Knowledge-Driven, Ai-Assisted Cyber Defence Framework For Iomt Remote Patient Monitoring, Kulsoom S. Bughio, David M. Cook, Abdul M. Unar

Research outputs 2022 to 2026

The rapid adoption of Internet Medical Things (IoMT) technologies in remote patient monitoring has reshaped healthcare delivery by enabling continuous, real-time clinical observation outside traditional care settings. However, this shift has also expanded the cyber-attack surface across heterogeneous, resource-constrained medical devices, wireless networks, cloud services, and third-party platforms. In cyber warfare, healthcare has become an incorporated target of geopolitics, with hospitals, remote monitoring systems, and emergency health systems being used to broaden the attack surface for adversaries to exploit. Existing security approaches for IoMT environments remain largely manual, fragmented, and reactive, limiting their effectiveness in dynamically assessing vulnerabilities and supporting …


Potent But Stealthy: Rethink Profile Pollution Against Sequential Recommendation Via Bi-Level Constrained Reinforcement Paradigm, Jiajie Su, Zihan Nan, Yunshan Ma, Xiaobo Xia, Xiaohua Feng, Weiming Liu, Xiang Chen, Xiaolin Zheng, Chaochao Chen Jan 2026

Potent But Stealthy: Rethink Profile Pollution Against Sequential Recommendation Via Bi-Level Constrained Reinforcement Paradigm, Jiajie Su, Zihan Nan, Yunshan Ma, Xiaobo Xia, Xiaohua Feng, Weiming Liu, Xiang Chen, Xiaolin Zheng, Chaochao Chen

Research Collection School Of Computing and Information Systems

Sequential Recommenders, which exploit dynamic user intents through interaction sequences, are vulnerable to adversarial attacks. While existing attacks primarily rely on data poisoning, they require large-scale user access or fake profiles, thus lacking practicality. In this paper, we focus on the Profile Pollution Attack that subtly contaminates partial user interactions to induce targeted mispredictions. Previous PPA methods suffer from two limitations, i.e., i) overreliance on sequence horizon impact restricts fine-grained perturbations on item transitions, and ii) holistic modifications cause detectable distribution shifts. To address these challenges, we propose a constrained reinforcement driven attack CREAT that synergizes a bi-level optimization framework …


Nondeterministic Polynomial-Time Problem Challenge: An Ever-Scaling Reasoning Benchmark For Llms, Chang Yang, Ruiyu Wang, Junzhe Jiang, Qi Jiang, Qinggang Zhang, Yanchen Deng, Shuxin Li, Shuyue Hu, Bo Li, Florian T. Pokorny, Xiao Huang, Xinrun Wang Jan 2026

Nondeterministic Polynomial-Time Problem Challenge: An Ever-Scaling Reasoning Benchmark For Llms, Chang Yang, Ruiyu Wang, Junzhe Jiang, Qi Jiang, Qinggang Zhang, Yanchen Deng, Shuxin Li, Shuyue Hu, Bo Li, Florian T. Pokorny, Xiao Huang, Xinrun Wang

Research Collection School Of Computing and Information Systems

Reasoning is the fundamental capability of large language models (LLMs). Due to the rapid progress of LLMs, there are two main issues of current benchmarks: i) these benchmarks can be crushed in a short time (less than 1 year), and ii) these benchmarks may be easily hacked. To handle these issues, we propose the ever-scalingness for building the benchmarks which are scaling over complexity against crushing, instance against hacking and exploitation, oversight for easy verification, and coverage for real-world relevance. This paper presents Nondeterministic Polynomial-time Problem Challenge (NPPC), an ever-scaling reasoning benchmark for LLMs. Specifically, the NPPC has three main …


Interpretable Multimodal Zero Shot Ecg Diagnosis Via Structured Clinical Knowledge Alignment, Jialu Tang, Hung Manh Pham, Ignace De Lathauwer, Henk S. Schipper, Yuan Lu, Dong Ma, Aaqib Saeed Jan 2026

Interpretable Multimodal Zero Shot Ecg Diagnosis Via Structured Clinical Knowledge Alignment, Jialu Tang, Hung Manh Pham, Ignace De Lathauwer, Henk S. Schipper, Yuan Lu, Dong Ma, Aaqib Saeed

Research Collection School Of Computing and Information Systems

Electrocardiogram (ECG) interpretation is essential for cardiovascular disease diagnosis, but current automated systems often struggle with transparency and generalization to unseen conditions. To address this, we introduce ZETA, a zero-shot multimodal framework designed for interpretable ECG diagnosis aligned with clinical workflows. ZETA uniquely compares ECG signals against structured positive and negative clinical observations, which are curated through an LLM-assisted, expertvalidated process, thereby mimicking differential diagnosis. Our approach leverages a pre-trained multimodal model to align ECG and text embeddings without disease-specific fine-tuning. Empirical evaluations demonstrate ZETA’s competitive zero-shot classification performance and, importantly, provide qualitative and quantitative evidence of enhanced interpretability, grounding …


Generalization Bounds For Semi‑Supervised Matrix Completion With Distributional Side Information, Antoine Ledent, Mun Chong Soo, Minh Hieu Nong Jan 2026

Generalization Bounds For Semi‑Supervised Matrix Completion With Distributional Side Information, Antoine Ledent, Mun Chong Soo, Minh Hieu Nong

Research Collection School Of Computing and Information Systems

We study a matrix completion problem where both the ground truth R matrix and the unknown sampling distribution P over observed entries are low-rank matrices, and share a common subspace. We assume that a large amount M of unlabeled data drawn from the sampling distribution P is available, together with a small amount N of labeled data drawn from the same distribution and noisy estimates of the corresponding ground truth entries. This setting is inspired by recommender systems scenarios where the unlabeled data corresponds to ‘implicit feedback’ (consisting in interactions such as purchase, click, etc. ) and the labeled data …


Can We Read Ai’S Mind? A Quest For Transparency, Santhosh Kumar Ravindran, Estera Kot, Fiona Fui-Hoon Nah Jan 2026

Can We Read Ai’S Mind? A Quest For Transparency, Santhosh Kumar Ravindran, Estera Kot, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

Although Artificial Intelligence (AI) systems are playing an increasing role in critical domains such as healthcare, finance, and autonomous systems, their decision-making processes remain largely opaque. This paper examines the challenges of AI transparency, addressing the “black box” problem using Explainable AI (XAI) techniques such as SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME). It also examines the ethical, regulatory, and societal implications of AI opacity and proposes a Comprehensive AI Observability (CAO) Framework that integrates deep explainability, provenance tracking, and real-time monitoring to enhance AI accountability. By bridging technical solutions with governance structures, this research emphasizes the …


Actor-Critic For Continuous Action Chunks: A Reinforcement Learning Framework For Long-Horizon Robotic Manipulation With Sparse Reward, Jiarui Yang, Bin Zhu, Jingjing Chen, Yu-Gang Jiang Jan 2026

Actor-Critic For Continuous Action Chunks: A Reinforcement Learning Framework For Long-Horizon Robotic Manipulation With Sparse Reward, Jiarui Yang, Bin Zhu, Jingjing Chen, Yu-Gang Jiang

Research Collection School Of Computing and Information Systems

Existing reinforcement learning (RL) methods struggle with long-horizon robotic manipulation tasks, particularly those involving sparse rewards. While action chunking is a promising paradigm for robotic manipulation, using RL to directly learn continuous action chunks in a stable and data-efficient manner remains a critical challenge. This paper introduces AC3 (Actor-Critic for Continuous Chunks), a novel RL framework that learns to generate high-dimensional, continuous action sequences. To make this learning process stable and dataefficient, AC3 incorporates targeted stabilization mechanisms for both the actor and the critic. First, to ensure reliable policy improvement, the actor is trained with an asymmetric update rule, learning …


Ai Systems For Physicians: A Review From Socio-Technical And Human-Computer Interaction Perspectives, Wu Jiaqi Young, Fiona Fui-Hoon Nah Jan 2026

Ai Systems For Physicians: A Review From Socio-Technical And Human-Computer Interaction Perspectives, Wu Jiaqi Young, Fiona Fui-Hoon Nah

Research Collection School Of Computing and Information Systems

The adoption of artificial intelligence (AI) in healthcare is accelerating, yet successful implementations of physician-facing AI systems remain limited and uneven. This paper presents a literature review of 40 peer-reviewed studies published between November 2022 and November 2024, spanning clinical, technical, and human-computer interaction (HCI) domains. Anchored in a socio-technical perspective, the review examines our existing understanding of how technical design, user expertise, and organizational factors shape the effectiveness of AI systems in real-world clinical settings. Our analysis identifies two meta-themes: (1) context as a dynamic, multi-level influence that actively reshapes AI system behavior, and (2) trust as an emergent …


Leveraging Large Language Models For Career Mobility Analysis: A Study Of Gender, Race, And Job Change Using Us Online Resume Profiles, Palakorn Achananuparp, Ye Xu, Yao Lu, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim Jan 2026

Leveraging Large Language Models For Career Mobility Analysis: A Study Of Gender, Race, And Job Change Using Us Online Resume Profiles, Palakorn Achananuparp, Ye Xu, Yao Lu, Xavier Jayaraj Siddarth Ashok, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

We present a large-scale analysis of career mobility of college-educated U.S. workers using online resume profiles to investigate how gender, race, and job change options are associated with upward mobility. This study addresses key research questions of how the job changes affect their upward career mobility, and how the outcomes of upward career mobility differ by gender and race. We address data challenges – such as missing demographic attributes, missing wage data, and noisy occupation labels – through various data processing and Artificial Intelligence (AI) methods. In particular, we develop a large language models (LLMs) based occupation classification method known …


Thinkmatter: Panoramic-Aware Instructional Semantics For Monocular Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Hao Zhao, Bin Zhu, Qianru Sun, Xiangbo Shu Jan 2026

Thinkmatter: Panoramic-Aware Instructional Semantics For Monocular Vision-And-Language Navigation, Guangzhao Dai, Shuo Wang, Hao Zhao, Bin Zhu, Qianru Sun, Xiangbo Shu

Research Collection School Of Computing and Information Systems

Vision-and-Language Navigation in continuous environments (VLN-CE) requires an embodied robot to navigate the target destination following the natural language instruction. Most existing methods use panoramic RGB-D cameras for 360° observation of environments. However, these methods struggle in real-world applications because of the higher cost of panoramic RGB-D cameras. This paper studies a low-cost and practical VLN-CE setting, e.g., using monocular cameras of limited field of view, which means “Look Less” for visual observations and environment semantics. In this paper, we propose a ThinkMatter framework for monocular VLN-CE, where we motivate monocular robots to “Think More” by 1) generating novel views …


Memoryart: Enhancing Llms Via Multi-Memory Models With Adaptive Resonance Theory For Healthcare Agents, Renke Dai, Hebin Hu, Jiahui Zhang, Yilin Kang, Ah-Hwee Tan Jan 2026

Memoryart: Enhancing Llms Via Multi-Memory Models With Adaptive Resonance Theory For Healthcare Agents, Renke Dai, Hebin Hu, Jiahui Zhang, Yilin Kang, Ah-Hwee Tan

Research Collection School Of Computing and Information Systems

Though promising in healthcare consultation applications, large language models (LLMs) face critical limitations in retaining and utilizing long-term memory across multiturn interactions. In particular, existing memory enhancing paradigms are constrained by limited context windows and embedding-based retrieval, often failing to maintain task relevance and still suffering from memory prototype collapse in multi-turn healthcare consultation. To address these challenges, we propose a cognitively-inspired memory framework named MemoryART, which is grounded in Adaptive Resonance Theory (ART)—a cognitive and learning theory of how humans and animals adapt to dynamic environments. MemoryART employs three memory modules—working memory, episodic memory, and semantic memory to support …


Benchmarking Gaslighting Negation Attacks Against Reasoning Models, Bin Zhu, Hailong Yin, Jingjing Chen, Yu Gang Jiang Jan 2026

Benchmarking Gaslighting Negation Attacks Against Reasoning Models, Bin Zhu, Hailong Yin, Jingjing Chen, Yu Gang Jiang

Research Collection School Of Computing and Information Systems

Recent advances in reasoning-centric models promise improved robustness through mechanisms such as chain-of-thought prompting and test-time scaling. However, their ability to withstand gaslighting negation attacks—adversarial prompts that confidently deny correct answers—remains underexplored. In this paper, we conduct a systematic evaluation of three state-of-the-art reasoning models, i.e., OpenAI’s o4-mini, Claude-3.7-Sonnet and Gemini-2.5-Flash, across three multimodal benchmarks: MMMU, MathVista, and CharXiv. Our evaluation reveals significant accuracy drops (25–29% on average) following gaslighting negation attacks, indicating that even top-tier reasoning models struggle to preserve correct answers under manipulative user feedback. Built upon the insights of the evaluation and to further probe this vulnerability, …


Integrating Symbolic And Waveform Music Into Large Language Models, Teng Tu, Xiaohao Liu, Yunshan Ma, Ji Qi, Tat-Seng Chua Jan 2026

Integrating Symbolic And Waveform Music Into Large Language Models, Teng Tu, Xiaohao Liu, Yunshan Ma, Ji Qi, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Music, as a unique and integral element of human life, is characterized by its complex structures, intricate details, and the fusion of multimodal information. Recent study advance music understanding by leveraging knowledge and reasoning capabilities derived from Large Language Models (LLMs). However, they often lack compatibility and fail to fully utilize the complementary strengths of diverse representations (e.g., ABC, MIDI, Waveform). To address these limitations, we propose a unified music-language model framework, named UniMuLM, transitioning from single-representation approaches to the integration of multiple music representations for LLM. Unifying different music representation formats poses challenges such as patch integrity and boundary …


Editorial: Special Section On Challenges And Opportunities In Retrieval-Augmented Generation For Llms: Techniques, Trends, And Applications, Philip S. Yu, Haofen Wang, Feida Zhu Jan 2026

Editorial: Special Section On Challenges And Opportunities In Retrieval-Augmented Generation For Llms: Techniques, Trends, And Applications, Philip S. Yu, Haofen Wang, Feida Zhu

Research Collection School Of Computing and Information Systems

Retrieval-Augmented Generation (RAG) represents a transformative advancement for Large Language Models (LLMs) by integrating external knowledge to substantially improve accuracy and mitigate hallucinations. As a pivotal technology in the contemporary generative Artificial Intelligence (AI) landscape, RAG addresses fundamental challenges in knowledge-intensive tasks. This special issue serves as a dedicated platform to showcase these cutting-edge advancements. It features six rigorously peer-reviewed papers that present state-of-the-art research and applications in the rapidly evolving field of RAG.


Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo Jan 2026

Llm-As-A-Judge For Software Engineering: Literature Review, Vision, And The Road Ahead, Junda He, Jieke Shi, Terry Yue Zhuo, Christoph Treude, Jiamou Sun, Zhenchang Xing, Xiaoning Du, David Lo

Research Collection School Of Computing and Information Systems

The rapid integration of Large Language Models (LLMs) into software engineering (SE) has revolutionized tasks from code generation to program repair, producing a massive volume of software artifacts. This surge in automated creation has exposed a critical bottleneck: the lack of scalable and reliable methods to evaluate the quality of these outputs. Human evaluation, while effective, is very costly and time-consuming. Traditional automated metrics like BLEU rely on high-quality references and struggle to capture nuanced aspects of software quality, such as readability and usefulness. In response, the LLM-as-a-Judge paradigm, which employs LLMs for automated evaluation, has emerged. This approach leverages …