Open Access. Powered by Scholars. Published by Universities.®

Singapore Management University

Discipline
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 511 - 540 of 1897

Full-Text Articles in Artificial Intelligence and Robotics

An Aspect Performance-Aware Hypergraph Neural Network For Review-Based Recommendation, Junrui Liu, Tong Li, Di Wu, Zifang Tang, Yuan Fang, Zhen Yang Mar 2025

An Aspect Performance-Aware Hypergraph Neural Network For Review-Based Recommendation, Junrui Liu, Tong Li, Di Wu, Zifang Tang, Yuan Fang, Zhen Yang

Research Collection School Of Computing and Information Systems

Online reviews allow consumers to provide detailed feedback on various aspects of items. Existing methods utilize these aspects to model users' fine-grained preferences for specific item features through graph neural networks. We argue that the performance of items on different aspects is important for making precise recommendations, which has not been taken into account by existing approaches, due to lack of data. In this paper, we propose an aspect performance-aware hypergraph neural network (APH) for the review-based recommendation, which learns the performance of items from the conflicting sentiment polarity of user reviews. Specifically, APH comprehensively models the relationships among users, …


Marginal Benefit Driven Rl Teacher For Unsupervised Environment Design, Dexun Li, Wenjun Li, Pradeep Varakantham Mar 2025

Marginal Benefit Driven Rl Teacher For Unsupervised Environment Design, Dexun Li, Wenjun Li, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

Training generally capable agents in complex environments is a challenging task that involves identifying the “right” environments at the training stage. Recent research has highlighted the potential of the Unsupervised Environment Design framework, which generates environment instances/levels adaptively at the frontier of the agent’s capabilities using regret measures. While regret approaches have shown promise in generating feasible environments, they can produce difficult environments that are challenging for an RL agent to learn from. This is because regret represents the best-case (upper bound) learning potential and not the actual learning potential of an environment. To address this, we propose an alternative …


Simulation-Free Hierarchical Latent Policy Planning For Proactive Dialogues, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Yiheng Sun, Zerui Chen, Ming Liu, Bing Qin Mar 2025

Simulation-Free Hierarchical Latent Policy Planning For Proactive Dialogues, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Yiheng Sun, Zerui Chen, Ming Liu, Bing Qin

Research Collection School Of Computing and Information Systems

Recent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios and comprehensive policy repositories to develop such systems. However, existing approaches tend to rely on Large Language Models (LLMs) for user simulation and online learning, leading to biases that diverge from realistic scenarios and result in suboptimal efficiency. Moreover, these methods depend on manually defined, context-independent, coarse-grained policies, which not only incur high expert costs but also raise concerns regarding their completeness. In our work, we …


Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang Mar 2025

Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

Automated Program Repair (APR) aims to enhance software reliability by automatically generating bug-fixing patches. Recent work has improved the state-of-the-art of APR by fine-tuning pre-trained large language models (LLMs), such as CodeT5, for APR. However, the effectiveness of fine-tuning be-comes weakened in data scarcity scenarios, and data scarcity can be a common issue in practice, limiting fine-tuning performance. To alleviate this limitation, this paper adapts prompt tuning for enhanced APR and conducts a comprehensive study to evaluate its effectiveness in data scarcity scenarios, using three LLMs of different sizes and six diverse datasets across four programming languages. Prompt tuning rewrites …


Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang Mar 2025

Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

In recent years, AI-based software engineering has progressed from pre-trained models to advanced agentic workflows, with Software Development Agents representing the next major leap. These agents, capable of reasoning, planning, and interacting with external environments, offer promising solutions to complex software engineering tasks. However, while much research has evaluated code generated by large language models (LLMs), comprehensive studies on agent-generated patches, particularly in real-world settings, are lacking. This study addresses that gap by evaluating 4,892 patches from 10 top-ranked agents on 500 real-world GitHub issues from SWE-Bench Verified, focusing on their impact on code quality. Our analysis shows no single …


Mimic: Ai And Ar-Enhanced Multi-Modal, Immersive, Relative Instruction Comprehension, Dhanuja Wanniarachchi, Archan Misra Mar 2025

Mimic: Ai And Ar-Enhanced Multi-Modal, Immersive, Relative Instruction Comprehension, Dhanuja Wanniarachchi, Archan Misra

Research Collection School Of Computing and Information Systems

We present a multimodal instruction comprehension framework, called MImIC, that utilizes visual sensing (including LIDAR and 2D RGB sensing) & AI spatial reasoning capabilities to support more seamless and immersive interaction between humans and AI-driven situated assistive agents. MImIC's key new capability is to support disambiguation of a wider set of relative spatial references that users naturally employ while issuing spatially-situated instructions. To support enhanced visual grounding via a combination of both fully-qualified and relative attribute references, MImIC uses (a) a fine-tuned transformer-based language translation DNN to accurately convert natural verbal commands into a structured set of machine understandable constraints …


Explainable Neural Networks With Guarantee: A Sparse Estimation Approach, Antoine Ledent, Peng Liu Mar 2025

Explainable Neural Networks With Guarantee: A Sparse Estimation Approach, Antoine Ledent, Peng Liu

Research Collection School Of Computing and Information Systems

Balancing predictive power and interpretability has long been a challenging research area, particularly in powerful yet complex models like neural networks, where nonlinearity obstructs direct interpretation. This paper introduces a novel approach to constructing an explainable neural network that harmonizes predictiveness and explainability. Our model is designed as a linear combination of a sparse set of jointly learned features, each derived from a different trainable function applied to a single 1-dimensional input feature. Leveraging the ability to learn arbitrarily complex relationships, our neural network architecture enables automatic selection of a sparse set of important features, with the final prediction being …


Human-Ai Synergy In Survey Development: Implications From Large Language Models In Business And Research, Ping Fan Ke, Ka Chung Ng Feb 2025

Human-Ai Synergy In Survey Development: Implications From Large Language Models In Business And Research, Ping Fan Ke, Ka Chung Ng

Research Collection School Of Computing and Information Systems

This study examines the novel integration of Large Language Models (LLMs) into the survey development process in business and research through the development and evaluation of the Behavioral Research ASSistant (BRASS) Bot. We first analyzed the traditional scale development process to identify tasks suitable for LLM integration, including both human-in-the-loop and automated LLM data collection methods. Following this analysis, we developed the details of BRASS Bot, incorporating design principles of falsifiability and reproducibility. We then conducted a comprehensive evaluation of the BRASS Bot across a diverse set of LLMs, including GPT, Claude, Gemini, and Llama, to assess its usability, validity, …


Bridging Expert Knowledge With Deep Learning Techniques For Just-In-Time Defect Prediction, Xin Zhou, Donggyun Han, David Lo Feb 2025

Bridging Expert Knowledge With Deep Learning Techniques For Just-In-Time Defect Prediction, Xin Zhou, Donggyun Han, David Lo

Research Collection School Of Computing and Information Systems

Just-In-Time (JIT) defect prediction aims to automatically predict whether a commit is defective or not, and has been widely studied in recent years. In general, most studies can be classified into two categories: 1) simple models using traditional machine learning classifiers with hand-crafted features, and 2) complex models using deep learning techniques to automatically extract features from commit contents. Hand-crafted features used by simple models are based on expert knowledge but may not fully represent the semantic meaning of the commits. On the other hand, deep learning-based features used by complex models represent the semantic meaning of commits but may …


Proactive Conversational Ai: A Comprehensive Survey Of Advancements And Opportunities, Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, Tat-Seng Chua Feb 2025

Proactive Conversational Ai: A Comprehensive Survey Of Advancements And Opportunities, Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

Dialogue systems are designed to offer human users social support or functional services through natural language interactions. Traditional conversation research has put significant emphasis on a system's response-ability, including its capacity to understand dialogue context and generate appropriate responses. However, the key element of proactive behavior-a crucial aspect of intelligent conversations-is often overlooked in these studies. Proactivity empowers conversational agents to lead conversations towards achieving pre-defined targets or fulfilling specific goals on the system side. Proactive dialogue systems are equipped with advanced techniques to handle complex tasks, requiring strategic and motivational interactions, thus representing a significant step towards artificial general …


Loco: Low-Bit Communication Adaptor For Large-Scale Model Training, Xingyu Xie, Zhijie Lin, Kim-Chuan Toh, Pan Zhou Feb 2025

Loco: Low-Bit Communication Adaptor For Large-Scale Model Training, Xingyu Xie, Zhijie Lin, Kim-Chuan Toh, Pan Zhou

Research Collection School Of Computing and Information Systems

To efficiently train large-scale models, low-bit gradient communication compresses full-precision gradients on local GPU nodes into low-precision ones for higher gradient synchronization efficiency among GPU nodes. However, it often degrades training quality due to compression information loss. To address this, we propose the Low-bit Communication Adaptor (LoCo), which compensates gradients on local GPU nodes before compression, ensuring efficient synchronization without compromising training quality. Specifically, LoCo designs a moving average of historical compensation errors to stably estimate concurrent compression error and then adopts it to compensate for the concurrent gradient compression, yielding a less lossless compression. This mechanism allows it to …


A Causality-Aware Paradigm For Evaluating Creativity Of Multimodal Large Language Models, Zhongzhan Huang, Shanshan Zhong, Pan Zhou, Shanghua Gao, Marink Zitnik, Liang Lin Feb 2025

A Causality-Aware Paradigm For Evaluating Creativity Of Multimodal Large Language Models, Zhongzhan Huang, Shanshan Zhong, Pan Zhou, Shanghua Gao, Marink Zitnik, Liang Lin

Research Collection School Of Computing and Information Systems

Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative capabilities of LLMs is challenging due to the subjective, diverse, and data-scarce nature of creativity, especially in multimodal scenarios. In this paper, we consider the comprehensive pipeline for evaluating the creativity of multimodal LLMs, with a focus on suitable evaluation platforms and methodologies. First, we find the Oogiri game—a creativity-driven task requiring humor, associative thinking, and the ability to produce unexpected responses to text, images, or both. This game aligns well with the input-output structure of modern …


Human‑Ai And Human‑Robot Collaboration In The Age Of Generative Ai, Agentic Ai, And Artificial General Intelligence: Opportunities And Challenges, Keng Siau Feb 2025

Human‑Ai And Human‑Robot Collaboration In The Age Of Generative Ai, Agentic Ai, And Artificial General Intelligence: Opportunities And Challenges, Keng Siau

Research Collection School Of Computing and Information Systems

The advancement of Artificial Intelligence (AI) has been exponential, especially in the past few years. Most, if not all, of the AI systems we encounter and are exposed to at this point are Artificial Narrow Intelligence (ANI). ANI specializes in one area and solves problems in one area. Generative AI (GenAI) and Agentic AI (i.e., independent AI agent), at the current stage of development, are regarded as ANI. The race is currently on to develop Artificial General Intelligence (AGI). AGI refers to AI systems as smart as humans across a wide range of cognitive tasks. Recently, OpenAI’s o3 system received …


Seven Hci Grand Challenges Revisited: Five-Year Progress, Constantine Stephanidis, Gavriel Salvendy, Margherita Antona, Vincent G Duffy, Qin Gao, Waldemar Karwowski, Fiona Nah, Stavroula Ntoa, Pei-Luen Patrick Rau, Keng Siau, Jia Zhou Feb 2025

Seven Hci Grand Challenges Revisited: Five-Year Progress, Constantine Stephanidis, Gavriel Salvendy, Margherita Antona, Vincent G Duffy, Qin Gao, Waldemar Karwowski, Fiona Nah, Stavroula Ntoa, Pei-Luen Patrick Rau, Keng Siau, Jia Zhou

Research Collection School Of Computing and Information Systems

Motivated by the rapid technological advancements achieved in the last five years, and the pervasiveness of Artificial Intelligence, the paper investigates the evolving role of Human-Computer Interaction and revisits the seven grand challenges outlined in 2019: human-technology symbiosis, human-environment interactions, ethics, privacy and security, well-being, health and eudaimonia, accessibility and universal access, learning and creativity, and social organization and democracy. Through literature analysis, the paper reevaluates the status of each challenge and highlights emerging requirements. Key findings reveal the widespread impact of Artificial Intelligence across all domains and emphasize the need for improved AI transparency, alignment with human values, and …


Hotpatching On The Fly: Mitigating Drone Incidents Arising From Incorrect Configuration, Ruidong Han, Juanru Li, Zhuo Ma, David Lo, Arash Shaghaghi, Jianfeng Ma, Siqi Ma Feb 2025

Hotpatching On The Fly: Mitigating Drone Incidents Arising From Incorrect Configuration, Ruidong Han, Juanru Li, Zhuo Ma, David Lo, Arash Shaghaghi, Jianfeng Ma, Siqi Ma

Research Collection School Of Computing and Information Systems

Manufacturers offer adjustable control parameters for flight control systems to accommodate diverse environments and missions. To ensure flight safety, they also develop established boundaries, i.e., range specifications for parameter values. However, even when the configuration parameters fall within the prescribed manufacturer range, they could still lead to instability or even severe incidents like crashes, which are referred to as Range Specification Bugs. Prior research has suggested shrinking the range of parameter values to protect drones from the adverse effects of such bugs. However, narrowing the range of parameters may only reduce the probability of errors and could potentially limit the …


Ai As Your Ally: The Effects Of Ai-Assisted Venting On Negative Affect And Perceived Social Support, Meilan Hu, Xavier Cheng Wee Chua, Shu Fen Diong, K. T. A. Sandeeshwara Kasturiratna, Nadyanna M. Majeed, Andree Hartanto Feb 2025

Ai As Your Ally: The Effects Of Ai-Assisted Venting On Negative Affect And Perceived Social Support, Meilan Hu, Xavier Cheng Wee Chua, Shu Fen Diong, K. T. A. Sandeeshwara Kasturiratna, Nadyanna M. Majeed, Andree Hartanto

Research Collection School of Social Sciences

In recent years, artificial intelligence (AI) chatbots have made significant strides in generating human-like conversations. With AI's expanding capabilities in mimicking human interactions, its affordability and accessibility underscore the potential of AI chatbots to facilitate negative emotional disclosure or venting. The study's primary objective is to highlight the potential benefits of AI-assisted venting by comparing its effectiveness to venting through a traditional journaling platform in reducing negative affect and increasing perceived social support. We conducted a pre-registered within-subject experiment involving 150 participants who completed both traditional venting and AI-assisted venting conditions with counterbalancing and a wash-out period of 1-week between …


Negotiating With Gpt-4: Digital Doormat Or Skilful Counterpart?, Dorcas Quek Anderson Feb 2025

Negotiating With Gpt-4: Digital Doormat Or Skilful Counterpart?, Dorcas Quek Anderson

Research Collection Yong Pung How School Of Law

Large language models (LLMs) such as GPT-4 have been creatively harnessed in the conflict resolution arena as dialogue agents interacting with humans within negotiations, due to their capacity for in-context learning and giving human-like responses. In light of the burgeoning use of LLMs in conflict resolution training, a pilot study was conducted to ascertain the desirability of using dialogue agents built on GPT-4 in conducting simulations for students learning negotiation skills. This article discusses insights gained from the study on the reliability of LLM agents in following prompts for negotiation simulations; notable negotiation behaviour of the LLM agent; the degree …


Towards The Next Generation Of Geospatial Artificial Intelligence, Gengchen Mai, Yiqun Xie, Xiaowei Jia, Ni Lao, Jinmeng Rao, Qing Zhu, Zeping Liu, Yao-Yi Chiang, Jiao, Junfeng Feb 2025

Towards The Next Generation Of Geospatial Artificial Intelligence, Gengchen Mai, Yiqun Xie, Xiaowei Jia, Ni Lao, Jinmeng Rao, Qing Zhu, Zeping Liu, Yao-Yi Chiang, Jiao, Junfeng

Research Collection College of Integrative Studies

Geospatial Artificial Intelligence (GeoAI), as the integration of geospatial studies and AI, has become one of the fastest-developing research directions in spatial data science and geography. This rapid change in the field calls for a deeper understanding of the recent developments and envision where the field is going in the near future. In this work, we provide a quantitative analysis of the GeoAI literature from the spatial, temporal, and semantic aspects. We briefly discuss the history of AI and GeoAI by highlighting some pioneering work. Then we discuss the current landscape of GeoAI by selecting five representative subdomains including remote …


Towards Resource-Efficient Reactive And Proactive Auto-Scaling For Microservice Architectures, Hussain Ahmad, Christoph Treude, Markus Wagner, Claudia Szabo Feb 2025

Towards Resource-Efficient Reactive And Proactive Auto-Scaling For Microservice Architectures, Hussain Ahmad, Christoph Treude, Markus Wagner, Claudia Szabo

Research Collection School Of Computing and Information Systems

Microservice architectures have become increasingly popular in both academia and industry, providing enhanced agility, elasticity, and maintainability in software development and deployment. To simplify scaling operations in microservice architectures, container orchestration platforms such as Kubernetes feature Horizontal Pod Auto-scalers (HPAs) designed to adjust the resources of microservices to accommodate fluctuating workloads. However, existing HPAs are not suitable for resource-constrained environments, as they make scaling decisions based on the individual resource capacities of microservices, leading to service unavailability, resource mismanagement, and financial losses. Furthermore, the inherent delay in initializing and terminating microservice pods hinders HPAs from timely responding to workload fluctuations, …


Governing Intelligence: Singapore’S Evolving Ai Governance Framework, Jason G. Allen, Jane Loo, Jose Luna Jan 2025

Governing Intelligence: Singapore’S Evolving Ai Governance Framework, Jason G. Allen, Jane Loo, Jose Luna

Research Collection Yong Pung How School Of Law

This paper provides an outline analysis of the evolving governance framework for Artificial Intelligence (AI) in Singapore. Across the Singapore government, AI solutions are being adopted in line with Singapore’s “Smart Nation Initiative” to leverage technology to make impactful changes to the nation and the economy. In tandem, Singaporean authorities have been assiduous to release a growing number of governance documents, which we analyse together to chart the city-state’s approach to AI governance in international comparison. Characteristics of Singapore’s AI governance approach include an emphasis on consensusbuilding between stakeholders (particularly government and industry but also citizens) andvoluntary or “quasi” regulation, …


Deep Reinforcement Learning With Explicit Context Representation, Francisco Munguia-Galeano, Ah-Hwee Tan, Ze Ji Jan 2025

Deep Reinforcement Learning With Explicit Context Representation, Francisco Munguia-Galeano, Ah-Hwee Tan, Ze Ji

Research Collection School Of Computing and Information Systems

Though reinforcement learning (RL) has shown an outstanding capability for solving complex computational problems, most RL algorithms lack an explicit method that would allow learning from contextual information. On the other hand, humans often use context to identify patterns and relations among elements in the environment, along with how to avoid making wrong actions. However, what may seem like an obviously wrong decision from a human perspective could take hundreds of steps for an RL agent to learn to avoid. This article proposes a framework for discrete environments called Iota explicit context representation (IECR). The framework involves representing each state …


Lara : A Light And Anti-Overfitting Retraining Approach For Unsupervised Time Series Anomaly Detection, Feiyi Chen, Zhen Qin, Mengchu Zhou, Yingying Zhang, Shuiguang Deng, Lunting Fan, Guansong Pang, Qingsong Wen Jan 2025

Lara : A Light And Anti-Overfitting Retraining Approach For Unsupervised Time Series Anomaly Detection, Feiyi Chen, Zhen Qin, Mengchu Zhou, Yingying Zhang, Shuiguang Deng, Lunting Fan, Guansong Pang, Qingsong Wen

Research Collection School Of Computing and Information Systems

Most of current anomaly detection models assume that the normal pattern remains the same all the time. However, the normal patterns of web services can change dramatically and frequently over time. The model trained on old-distribution data becomes outdated and ineffective after such changes. Retraining the whole model whenever the pattern is changed is computationally expensive. Further, at the beginning of normal pattern changes, there is not enough observation data from the new distribution. Retraining a large neural network model with limited data is vulnerable to overfitting. Thus, we propose a Light Anti-overfitting Retraining Approach (LARA) based on deep variational …


Fuzzing Drones For Anomaly Detection: A Systematic Literature Review, Vikas Kumar Malviya, Wei Minn, Lwin Khin Shar, Lingxiao Jiang Jan 2025

Fuzzing Drones For Anomaly Detection: A Systematic Literature Review, Vikas Kumar Malviya, Wei Minn, Lwin Khin Shar, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

Drones, also referred to as Unmanned Aerial Vehicles (UAVs), are becoming popular today due to their uses in different fields and recent technological advancements which provide easy control of UAVs via mobile apps. However, UAVs may contain vulnerabilities or software bugs that cause serious safety and security concerns. For example, the communication protocol used by the UAV may contain authentication and authorization vulnerabilities, which may be exploited by attackers to gain remote access over the UAV. Drones must therefore undergo extensive testing before being released or deployed to identify and fix any software bugs or security vulnerabilities. Fuzzing is one …


Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang Jan 2025

Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang

Research Collection School Of Computing and Information Systems

Deep Neural Network-based models have demonstrated high accuracy for semantic code clone detection. However, the lack of generalization poses a threat to the trustworthiness and reliability of these models. Furthermore, the black-box nature of these models makes interpreting the model’s decisions very challenging. Currently, there is only a limited understanding of the semantic code clone detection behavior of existing models. There is a lack of transparency in understanding how a model identifies semantic code clones and the exact code components influencing its prediction. In this paper, we introduce the use of a causal interpretation framework based on the Neyman-Rubin causal …


A Review Of Chinese Sentiment Analysis: Subjects, Methods, And Trends, Zhaoxia Wang, Donghao Huang, Jingfeng Cui, Xinyue Zhang, Seng-Beng Ho, Erik Cambria Jan 2025

A Review Of Chinese Sentiment Analysis: Subjects, Methods, And Trends, Zhaoxia Wang, Donghao Huang, Jingfeng Cui, Xinyue Zhang, Seng-Beng Ho, Erik Cambria

Research Collection School Of Computing and Information Systems

Sentiment analysis has emerged as a prominent research domain within the realm of natural language processing, garnering increasing attention and a growing body of literature. While numerous literature reviews have examined sentiment analysis techniques, methods, topics and applications, there remains a gap in the literature concerning thematic trends and research methodologies in sentiment analysis, particularly in the context of Chinese text. This study addresses this gap by presenting a comprehensive survey dedicated to the progression of research subjects, methods and trends in sentiment analysis of Chinese text. Employing a framework that combines keyword co-occurrence analysis with a sophisticated community detection …


Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun Jan 2025

Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun

Research Collection School Of Computing and Information Systems

Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured by users in the free-living environment. In addition to this domain-shift problem, the real-world food datasets tend to be long-tailed distributed and some dishes of different categories exhibit subtle variations that are difficult to distinguish visually. In this paper, we present a framework empowered with large language models (LLMs) to address these challenges in food recognition. We first leverage LLMs to parse food images to generate food titles and ingredients. Then, we project the …


Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo Jan 2025

Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo

Research Collection School Of Computing and Information Systems

Cross-modal representation learning is essential for interactive text-to-video search tasks. However, the representation learning is limited by the size and quality of video-caption pairs. To improve the search accuracy, we propose to enlarge the size of available video-caption pairs by leveraging multi-model LLM on video captioning. Specifically, we use LLM to generate video captions for a large video collection (i.e., WebVid dataset) and use the generated video-caption pairs to pre-train a text-to-video search model. Additionally, we use LLM to generate fine-grained captions for test video collections to enable text-to-caption retrieval. Furthermore, we build a semantic overview of the retrieved rank …


Weakly-Supervised Semantic Segmentation With Image-Level Labels: From Traditional Models To Foundation Models, Zhaozheng Chen, Qianru Sun Jan 2025

Weakly-Supervised Semantic Segmentation With Image-Level Labels: From Traditional Models To Foundation Models, Zhaozheng Chen, Qianru Sun

Research Collection School Of Computing and Information Systems

The rapid development of deep learning has driven significant progress in image semantic segmentation—a fundamental task in computer vision. Semantic segmentation algorithms often depend on the availability of pixel-level labels (i.e., masks of objects), which are expensive, time consuming, and labor intensive. Weakly supervised semantic segmentation (WSSS) is an effective solution to avoid such labeling. It utilizes only partial or incomplete annotations and provides a cost-effective alternative to fully supervised semantic segmentation. In this article, our focus is on the WSSS with image-level labels, which is the most challenging form of WSSS. Our work has two parts. First, we conduct …


Synthesizing Multi-Person And Rare Pose Images For Human Pose Estimation, Liuqing Zhao, Zichen Tian, Zou Peng, Richang Hong, Qianru Sun Jan 2025

Synthesizing Multi-Person And Rare Pose Images For Human Pose Estimation, Liuqing Zhao, Zichen Tian, Zou Peng, Richang Hong, Qianru Sun

Research Collection School Of Computing and Information Systems

Human pose estimation (HPE) models underperform in recognizing rare poses because they suffer from data imbalance problems (i.e., there are few image samples for rare poses) in their training datasets. From a data perspective, the most intuitive solution is to synthesize data for rare poses. Specifically, the rule-based methods apply manual manipulations (such as Cutout and GridMask) to the existing data, so the limited diversity of the data constrains the model. An alternative method is to learn the underlying data distribution via deep generative models (such as ControlNet and HumanSD) and then sample “new data” from the distribution. This works …


Transforming Urban Dynamics: Harnessing Large Language Models For Smarter Mobility, Hao Xue, Ming Jin, Shirui Pan, Flora Salim, Guansong Pang Jan 2025

Transforming Urban Dynamics: Harnessing Large Language Models For Smarter Mobility, Hao Xue, Ming Jin, Shirui Pan, Flora Salim, Guansong Pang

Research Collection School Of Computing and Information Systems

Artificial intelligence (AI) has the potential to analyze mobility data and make mobility systems smarter by leveraging diverse data sources such as geospatial data, transportation logs, and real-time sensor data to optimize traffic flow, enhance public transportation systems, and support the development of autonomous vehicles. With the newly emerged generative AI paradigm, exemplified by large language models (LLMs), there is great potential to transform the current AI applications in mobility, transportation, and urban domains. This article provides an overview of recent efforts and aims to shed light on the challenges and future opportunities to facilitate the adaptation of LLMs for …