Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Databases and Information Systems (344)
- Engineering (291)
- Operations Research, Systems Engineering and Industrial Engineering (255)
- Graphics and Human Computer Interfaces (171)
- Software Engineering (129)
-
- Business (125)
- Numerical Analysis and Scientific Computing (111)
- Social and Behavioral Sciences (101)
- Theory and Algorithms (97)
- Public Affairs, Public Policy and Public Administration (66)
- Transportation (61)
- Programming Languages and Compilers (57)
- Information Security (40)
- OS and Networks (37)
- Medicine and Health Sciences (35)
- Computer Engineering (31)
- Education (24)
- Health Information Technology (24)
- Asian Studies (23)
- International and Area Studies (23)
- Finance and Financial Management (14)
- Technology and Innovation (13)
- Communication (12)
- Operations and Supply Chain Management (12)
- Social Media (12)
- Higher Education (11)
- Computer and Systems Architecture (7)
- Keyword
-
- Artificial intelligence (49)
- Machine learning (46)
- Reinforcement learning (39)
- Deep learning (37)
- Large Language Models (26)
-
- Large Language Model (22)
- Large language models (21)
- Computer vision (18)
- Optimization (18)
- Scheduling (18)
- Anomaly detection (17)
- Generative AI (17)
- Large language model (17)
- Singapore (17)
- ChatGPT (16)
- Deep reinforcement learning (16)
- Reinforcement Learning (16)
- LLMs (15)
- Vehicle routing problem (15)
- Deep Learning (14)
- Natural language processing (14)
- Artificial Intelligence (13)
- Neural networks (13)
- Uncertainty (13)
- Machine Learning (12)
- Graph neural networks (11)
- Metaverse (10)
- Multi-agent systems (10)
- Representation learning (10)
- Semantics (10)
- Publication Year
- File Type
Articles 421 - 450 of 1664
Full-Text Articles in Artificial Intelligence and Robotics
Simulation-Free Hierarchical Latent Policy Planning For Proactive Dialogues, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Yiheng Sun, Zerui Chen, Ming Liu, Bing Qin
Simulation-Free Hierarchical Latent Policy Planning For Proactive Dialogues, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Yiheng Sun, Zerui Chen, Ming Liu, Bing Qin
Research Collection School Of Computing and Information Systems
Recent advancements in proactive dialogues have garnered significant attention, particularly for more complex objectives (e.g. emotion support and persuasion). Unlike traditional task-oriented dialogues, proactive dialogues demand advanced policy planning and adaptability, requiring rich scenarios and comprehensive policy repositories to develop such systems. However, existing approaches tend to rely on Large Language Models (LLMs) for user simulation and online learning, leading to biases that diverge from realistic scenarios and result in suboptimal efficiency. Moreover, these methods depend on manually defined, context-independent, coarse-grained policies, which not only incur high expert costs but also raise concerns regarding their completeness. In our work, we …
Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang
Adapting Knowledge Prompt Tuning For Enhanced Automated Program Repair, Xuemeng Cai, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Automated Program Repair (APR) aims to enhance software reliability by automatically generating bug-fixing patches. Recent work has improved the state-of-the-art of APR by fine-tuning pre-trained large language models (LLMs), such as CodeT5, for APR. However, the effectiveness of fine-tuning be-comes weakened in data scarcity scenarios, and data scarcity can be a common issue in practice, limiting fine-tuning performance. To alleviate this limitation, this paper adapts prompt tuning for enhanced APR and conducts a comprehensive study to evaluate its effectiveness in data scarcity scenarios, using three LLMs of different sizes and six diverse datasets across four programming languages. Prompt tuning rewrites …
Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang
Evaluating Software Development Agents: Patch Patterns, Code Quality, And Issue Complexity In Real-World Github Scenarios, Zhi Chen, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
In recent years, AI-based software engineering has progressed from pre-trained models to advanced agentic workflows, with Software Development Agents representing the next major leap. These agents, capable of reasoning, planning, and interacting with external environments, offer promising solutions to complex software engineering tasks. However, while much research has evaluated code generated by large language models (LLMs), comprehensive studies on agent-generated patches, particularly in real-world settings, are lacking. This study addresses that gap by evaluating 4,892 patches from 10 top-ranked agents on 500 real-world GitHub issues from SWE-Bench Verified, focusing on their impact on code quality. Our analysis shows no single …
Drone Delivery Network Design With Uncertainties, Wenjia Zeng, Jiang Ruiwei, Hai Yang, Hai Wang
Drone Delivery Network Design With Uncertainties, Wenjia Zeng, Jiang Ruiwei, Hai Yang, Hai Wang
Research Collection School Of Computing and Information Systems
Unmanned aerial vehicles (UAVs), also called drones, are gaining popularity as an alternative delivery mode due to their faster delivery speed and reduced labor costs. Several companies, especially e-commerce giants, are conducting pilot projects that use drones to deliver fast food and groceries. In 2021, for example, Walmart partnered with Zipline in the United States to provide delivery services for areas near Walmart stores in Arkansas. In China, Meituan drone delivery services have been launched in Shenzhen and have conducted trial food delivery that cover more than 8,000 households.
Unlocking The Potential Of Black-Box Pre-Trained Gnns For Graph Few-Shot Learning, Qiannan Zhang, Shichao Pei, Yuan Fang, Xiangliang Zhang
Unlocking The Potential Of Black-Box Pre-Trained Gnns For Graph Few-Shot Learning, Qiannan Zhang, Shichao Pei, Yuan Fang, Xiangliang Zhang
Research Collection School Of Computing and Information Systems
Few-shot learning has emerged as an important problem on graphs to combat label scarcity, which can be approached by current trends in pre-trained graph neural networks (GNNs) and meta-learning. Recent efforts integrate both paradigms in a white-box setting, leaving the more realistic black-box setting largely underexplored, where the parameters and gradients in the pre-trained GNNs are inaccessible. In this paper, we study the critical problem: Leveraging black-box pre-trained GNNs for graph few-shot learning. Despite its appeal, two key issues hinder the unlocking of its potential: the inherent task gap between pre-training and downstream stages, which can introduce irrelevant knowledge and …
Learning To Identify Seen, Unseen And Unknown In The Open World: A Practical Setting For Zero-Shot Learning, Sethupathy Parameswaran, Yuan Fang, Chandan Gautam, Savitha Ramasamy, Xiaoli Li
Learning To Identify Seen, Unseen And Unknown In The Open World: A Practical Setting For Zero-Shot Learning, Sethupathy Parameswaran, Yuan Fang, Chandan Gautam, Savitha Ramasamy, Xiaoli Li
Research Collection School Of Computing and Information Systems
As vision-language models advance, addressing the Zero-Shot Learning (ZSL) problem in the open world becomes increasingly crucial. Specifically, a robust model must handle three types of samples during inference: seen classes with visual and semantic information provided in training, unseen classes with only the semantic information in training, and unknown samples with no prior information from training. Existing methods either handle seen and unseen classes together (ZSL) or seen and unknown classes (known as Open-Set Recognition, OSR). However, none addresses the simultaneous handling of all three, which we term Open-Set Zero-Shot Learning (OZSL). To address this problem, we propose a …
Dr. Tongue: Sign-Oriented Multi-Label Detection For Remote Tongue Diagnosis, Yiliang Chen, Steven S. C. Ho, Cheng Xu, Yao Jie Xie, Wing Fai Yeung, Shengfeng He, Jing Qin
Dr. Tongue: Sign-Oriented Multi-Label Detection For Remote Tongue Diagnosis, Yiliang Chen, Steven S. C. Ho, Cheng Xu, Yao Jie Xie, Wing Fai Yeung, Shengfeng He, Jing Qin
Research Collection School Of Computing and Information Systems
Tongue diagnosis is a vital tool in Western and Traditional Chinese Medicine, providing key insights into a patient's health by analyzing tongue attributes. The COVID-19 pandemic has heightened the need for accurate remote medical assessments, emphasizing the importance of precise tongue attribute recognition via telehealth. To address this, we propose a Sign-Oriented multi-label Attributes Detection framework. Our approach begins with an adaptive tongue feature extraction module that standardizes tongue images and mitigates environmental factors. This is followed by a Sign-oriented Network (SignNet) that identifies specific tongue attributes, emulating the diagnostic process of experienced practitioners and enabling comprehensive health evaluations. To …
Hotpatching On The Fly: Mitigating Drone Incidents Arising From Incorrect Configuration, Ruidong Han, Juanru Li, Zhuo Ma, David Lo, Arash Shaghaghi, Jianfeng Ma, Siqi Ma
Hotpatching On The Fly: Mitigating Drone Incidents Arising From Incorrect Configuration, Ruidong Han, Juanru Li, Zhuo Ma, David Lo, Arash Shaghaghi, Jianfeng Ma, Siqi Ma
Research Collection School Of Computing and Information Systems
Manufacturers offer adjustable control parameters for flight control systems to accommodate diverse environments and missions. To ensure flight safety, they also develop established boundaries, i.e., range specifications for parameter values. However, even when the configuration parameters fall within the prescribed manufacturer range, they could still lead to instability or even severe incidents like crashes, which are referred to as Range Specification Bugs. Prior research has suggested shrinking the range of parameter values to protect drones from the adverse effects of such bugs. However, narrowing the range of parameters may only reduce the probability of errors and could potentially limit the …
Human-Ai Synergy In Survey Development: Implications From Large Language Models In Business And Research, Ping Fan Ke, Ka Chung Ng
Human-Ai Synergy In Survey Development: Implications From Large Language Models In Business And Research, Ping Fan Ke, Ka Chung Ng
Research Collection School Of Computing and Information Systems
This study examines the novel integration of Large Language Models (LLMs) into the survey development process in business and research through the development and evaluation of the Behavioral Research ASSistant (BRASS) Bot. We first analyzed the traditional scale development process to identify tasks suitable for LLM integration, including both human-in-the-loop and automated LLM data collection methods. Following this analysis, we developed the details of BRASS Bot, incorporating design principles of falsifiability and reproducibility. We then conducted a comprehensive evaluation of the BRASS Bot across a diverse set of LLMs, including GPT, Claude, Gemini, and Llama, to assess its usability, validity, …
Bridging Expert Knowledge With Deep Learning Techniques For Just-In-Time Defect Prediction, Xin Zhou, Donggyun Han, David Lo
Bridging Expert Knowledge With Deep Learning Techniques For Just-In-Time Defect Prediction, Xin Zhou, Donggyun Han, David Lo
Research Collection School Of Computing and Information Systems
Just-In-Time (JIT) defect prediction aims to automatically predict whether a commit is defective or not, and has been widely studied in recent years. In general, most studies can be classified into two categories: 1) simple models using traditional machine learning classifiers with hand-crafted features, and 2) complex models using deep learning techniques to automatically extract features from commit contents. Hand-crafted features used by simple models are based on expert knowledge but may not fully represent the semantic meaning of the commits. On the other hand, deep learning-based features used by complex models represent the semantic meaning of commits but may …
Seven Hci Grand Challenges Revisited: Five-Year Progress, Constantine Stephanidis, Gavriel Salvendy, Margherita Antona, Vincent G Duffy, Qin Gao, Waldemar Karwowski, Fiona Nah, Stavroula Ntoa, Pei-Luen Patrick Rau, Keng Siau, Jia Zhou
Seven Hci Grand Challenges Revisited: Five-Year Progress, Constantine Stephanidis, Gavriel Salvendy, Margherita Antona, Vincent G Duffy, Qin Gao, Waldemar Karwowski, Fiona Nah, Stavroula Ntoa, Pei-Luen Patrick Rau, Keng Siau, Jia Zhou
Research Collection School Of Computing and Information Systems
Motivated by the rapid technological advancements achieved in the last five years, and the pervasiveness of Artificial Intelligence, the paper investigates the evolving role of Human-Computer Interaction and revisits the seven grand challenges outlined in 2019: human-technology symbiosis, human-environment interactions, ethics, privacy and security, well-being, health and eudaimonia, accessibility and universal access, learning and creativity, and social organization and democracy. Through literature analysis, the paper reevaluates the status of each challenge and highlights emerging requirements. Key findings reveal the widespread impact of Artificial Intelligence across all domains and emphasize the need for improved AI transparency, alignment with human values, and …
Proactive Conversational Ai: A Comprehensive Survey Of Advancements And Opportunities, Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, Tat-Seng Chua
Proactive Conversational Ai: A Comprehensive Survey Of Advancements And Opportunities, Yang Deng, Lizi Liao, Wenqiang Lei, Grace Hui Yang, Wai Lam, Tat-Seng Chua
Research Collection School Of Computing and Information Systems
Dialogue systems are designed to offer human users social support or functional services through natural language interactions. Traditional conversation research has put significant emphasis on a system's response-ability, including its capacity to understand dialogue context and generate appropriate responses. However, the key element of proactive behavior-a crucial aspect of intelligent conversations-is often overlooked in these studies. Proactivity empowers conversational agents to lead conversations towards achieving pre-defined targets or fulfilling specific goals on the system side. Proactive dialogue systems are equipped with advanced techniques to handle complex tasks, requiring strategic and motivational interactions, thus representing a significant step towards artificial general …
Loco: Low-Bit Communication Adaptor For Large-Scale Model Training, Xingyu Xie, Zhijie Lin, Kim-Chuan Toh, Pan Zhou
Loco: Low-Bit Communication Adaptor For Large-Scale Model Training, Xingyu Xie, Zhijie Lin, Kim-Chuan Toh, Pan Zhou
Research Collection School Of Computing and Information Systems
To efficiently train large-scale models, low-bit gradient communication compresses full-precision gradients on local GPU nodes into low-precision ones for higher gradient synchronization efficiency among GPU nodes. However, it often degrades training quality due to compression information loss. To address this, we propose the Low-bit Communication Adaptor (LoCo), which compensates gradients on local GPU nodes before compression, ensuring efficient synchronization without compromising training quality. Specifically, LoCo designs a moving average of historical compensation errors to stably estimate concurrent compression error and then adopts it to compensate for the concurrent gradient compression, yielding a less lossless compression. This mechanism allows it to …
A Causality-Aware Paradigm For Evaluating Creativity Of Multimodal Large Language Models, Zhongzhan Huang, Shanshan Zhong, Pan Zhou, Shanghua Gao, Marink Zitnik, Liang Lin
A Causality-Aware Paradigm For Evaluating Creativity Of Multimodal Large Language Models, Zhongzhan Huang, Shanshan Zhong, Pan Zhou, Shanghua Gao, Marink Zitnik, Liang Lin
Research Collection School Of Computing and Information Systems
Recently, numerous benchmarks have been developed to evaluate the logical reasoning abilities of large language models (LLMs). However, assessing the equally important creative capabilities of LLMs is challenging due to the subjective, diverse, and data-scarce nature of creativity, especially in multimodal scenarios. In this paper, we consider the comprehensive pipeline for evaluating the creativity of multimodal LLMs, with a focus on suitable evaluation platforms and methodologies. First, we find the Oogiri game—a creativity-driven task requiring humor, associative thinking, and the ability to produce unexpected responses to text, images, or both. This game aligns well with the input-output structure of modern …
Towards Resource-Efficient Reactive And Proactive Auto-Scaling For Microservice Architectures, Hussain Ahmad, Christoph Treude, Markus Wagner, Claudia Szabo
Towards Resource-Efficient Reactive And Proactive Auto-Scaling For Microservice Architectures, Hussain Ahmad, Christoph Treude, Markus Wagner, Claudia Szabo
Research Collection School Of Computing and Information Systems
Microservice architectures have become increasingly popular in both academia and industry, providing enhanced agility, elasticity, and maintainability in software development and deployment. To simplify scaling operations in microservice architectures, container orchestration platforms such as Kubernetes feature Horizontal Pod Auto-scalers (HPAs) designed to adjust the resources of microservices to accommodate fluctuating workloads. However, existing HPAs are not suitable for resource-constrained environments, as they make scaling decisions based on the individual resource capacities of microservices, leading to service unavailability, resource mismanagement, and financial losses. Furthermore, the inherent delay in initializing and terminating microservice pods hinders HPAs from timely responding to workload fluctuations, …
Human‑Ai And Human‑Robot Collaboration In The Age Of Generative Ai, Agentic Ai, And Artificial General Intelligence: Opportunities And Challenges, Keng Siau
Research Collection School Of Computing and Information Systems
The advancement of Artificial Intelligence (AI) has been exponential, especially in the past few years. Most, if not all, of the AI systems we encounter and are exposed to at this point are Artificial Narrow Intelligence (ANI). ANI specializes in one area and solves problems in one area. Generative AI (GenAI) and Agentic AI (i.e., independent AI agent), at the current stage of development, are regarded as ANI. The race is currently on to develop Artificial General Intelligence (AGI). AGI refers to AI systems as smart as humans across a wide range of cognitive tasks. Recently, OpenAI’s o3 system received …
Deep Reinforcement Learning With Explicit Context Representation, Francisco Munguia-Galeano, Ah-Hwee Tan, Ze Ji
Deep Reinforcement Learning With Explicit Context Representation, Francisco Munguia-Galeano, Ah-Hwee Tan, Ze Ji
Research Collection School Of Computing and Information Systems
Though reinforcement learning (RL) has shown an outstanding capability for solving complex computational problems, most RL algorithms lack an explicit method that would allow learning from contextual information. On the other hand, humans often use context to identify patterns and relations among elements in the environment, along with how to avoid making wrong actions. However, what may seem like an obviously wrong decision from a human perspective could take hundreds of steps for an RL agent to learn to avoid. This article proposes a framework for discrete environments called Iota explicit context representation (IECR). The framework involves representing each state …
Financial Named Entity Recognition: How Far Can Llm Go?, Yi-Te Lu, Yintong Huo
Financial Named Entity Recognition: How Far Can Llm Go?, Yi-Te Lu, Yintong Huo
Research Collection School Of Computing and Information Systems
The surge of large language models (LLMs) has revolutionized the extraction and analysis of crucial information from a growing volume of financial statements, announcements, and business news. Recognition for named entities to construct structured data poses a significant challenge in analyzing financial documents and is a foundational task for intelligent financial analytics. However, how effective are these generic LLMs and their performance under various prompts are yet need a better understanding. To fill in the blank, we present a systematic evaluation of state-of-the-art LLMs and prompting methods in the financial Named Entity Recognition (NER) problem. Specifically, our experimental results highlight …
Defending Federated Recommender Systems Against Untargeted Attacks: A Contribution-Aware Robust Aggregation Scheme, Ruicheng Liang, Yuanchun Jiang, Feida Zhu, Ling Cheng, Huiwen Liu
Defending Federated Recommender Systems Against Untargeted Attacks: A Contribution-Aware Robust Aggregation Scheme, Ruicheng Liang, Yuanchun Jiang, Feida Zhu, Ling Cheng, Huiwen Liu
Research Collection School Of Computing and Information Systems
Federated recommender systems (FedRSs) effectively tackle the tradeoff between recommendation accuracy and privacy preservation. However, recent studies have revealed severe vulnerabilities in FedRSs, particularly against untargeted attacks seeking to undermine their overall performance. Defense methods employed in traditional recommender systems are not applicable to FedRSs, and existing robust aggregation schemes for other federated learning-based applications have proven ineffective in FedRSs. Building on the observation that malicious clients contribute negatively to the training process, we design a novel contribution-aware robust aggregation scheme to defend FedRSs against untargeted attacks, named contribution-aware Bayesian knowledge distillation aggregation (ConDA), comprising two key components for the …
Lara : A Light And Anti-Overfitting Retraining Approach For Unsupervised Time Series Anomaly Detection, Feiyi Chen, Zhen Qin, Mengchu Zhou, Yingying Zhang, Shuiguang Deng, Lunting Fan, Guansong Pang, Qingsong Wen
Lara : A Light And Anti-Overfitting Retraining Approach For Unsupervised Time Series Anomaly Detection, Feiyi Chen, Zhen Qin, Mengchu Zhou, Yingying Zhang, Shuiguang Deng, Lunting Fan, Guansong Pang, Qingsong Wen
Research Collection School Of Computing and Information Systems
Most of current anomaly detection models assume that the normal pattern remains the same all the time. However, the normal patterns of web services can change dramatically and frequently over time. The model trained on old-distribution data becomes outdated and ineffective after such changes. Retraining the whole model whenever the pattern is changed is computationally expensive. Further, at the beginning of normal pattern changes, there is not enough observation data from the new distribution. Retraining a large neural network model with limited data is vulnerable to overfitting. Thus, we propose a Light Anti-overfitting Retraining Approach (LARA) based on deep variational …
Fuzzing Drones For Anomaly Detection: A Systematic Literature Review, Vikas Kumar Malviya, Wei Minn, Lwin Khin Shar, Lingxiao Jiang
Fuzzing Drones For Anomaly Detection: A Systematic Literature Review, Vikas Kumar Malviya, Wei Minn, Lwin Khin Shar, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Drones, also referred to as Unmanned Aerial Vehicles (UAVs), are becoming popular today due to their uses in different fields and recent technological advancements which provide easy control of UAVs via mobile apps. However, UAVs may contain vulnerabilities or software bugs that cause serious safety and security concerns. For example, the communication protocol used by the UAV may contain authentication and authorization vulnerabilities, which may be exploited by attackers to gain remote access over the UAV. Drones must therefore undergo extensive testing before being released or deployed to identify and fix any software bugs or security vulnerabilities. Fuzzing is one …
Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang
Measuring Model Alignment For Code Clone Detection Using Causal Interpretation, Shamsa Abid, Xuemeng Cai, Lingxiao Jiang
Research Collection School Of Computing and Information Systems
Deep Neural Network-based models have demonstrated high accuracy for semantic code clone detection. However, the lack of generalization poses a threat to the trustworthiness and reliability of these models. Furthermore, the black-box nature of these models makes interpreting the model’s decisions very challenging. Currently, there is only a limited understanding of the semantic code clone detection behavior of existing models. There is a lack of transparency in understanding how a model identifies semantic code clones and the exact code components influencing its prediction. In this paper, we introduce the use of a causal interpretation framework based on the Neyman-Rubin causal …
A Review Of Chinese Sentiment Analysis: Subjects, Methods, And Trends, Zhaoxia Wang, Donghao Huang, Jingfeng Cui, Xinyue Zhang, Seng-Beng Ho, Erik Cambria
A Review Of Chinese Sentiment Analysis: Subjects, Methods, And Trends, Zhaoxia Wang, Donghao Huang, Jingfeng Cui, Xinyue Zhang, Seng-Beng Ho, Erik Cambria
Research Collection School Of Computing and Information Systems
Sentiment analysis has emerged as a prominent research domain within the realm of natural language processing, garnering increasing attention and a growing body of literature. While numerous literature reviews have examined sentiment analysis techniques, methods, topics and applications, there remains a gap in the literature concerning thematic trends and research methodologies in sentiment analysis, particularly in the context of Chinese text. This study addresses this gap by presenting a comprehensive survey dedicated to the progression of research subjects, methods and trends in sentiment analysis of Chinese text. Employing a framework that combines keyword co-occurrence analysis with a sophisticated community detection …
Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun
Llms-Based Augmentation For Domain Adaptation In Long-Tailed Food Datasets, Qing Wang, Chong-Wah Ngo, Ee-Peng Lim, Qianru Sun
Research Collection School Of Computing and Information Systems
Training a model for food recognition is challenging because the training samples, which are typically crawled from the Internet, are visually different from the pictures captured by users in the free-living environment. In addition to this domain-shift problem, the real-world food datasets tend to be long-tailed distributed and some dishes of different categories exhibit subtle variations that are difficult to distinguish visually. In this paper, we present a framework empowered with large language models (LLMs) to address these challenges in food recognition. We first leverage LLMs to parse food images to generate food titles and ingredients. Then, we project the …
Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo
Interactive Video Search With Multi-Modal Llm Video Captioning, Yu-Tong Cheng, Jiaxin Wu, Zhixin Ma, Jiangshan He, Xiao-Yong Wei, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
Cross-modal representation learning is essential for interactive text-to-video search tasks. However, the representation learning is limited by the size and quality of video-caption pairs. To improve the search accuracy, we propose to enlarge the size of available video-caption pairs by leveraging multi-model LLM on video captioning. Specifically, we use LLM to generate video captions for a large video collection (i.e., WebVid dataset) and use the generated video-caption pairs to pre-train a text-to-video search model. Additionally, we use LLM to generate fine-grained captions for test video collections to enable text-to-caption retrieval. Furthermore, we build a semantic overview of the retrieved rank …
Weakly-Supervised Semantic Segmentation With Image-Level Labels: From Traditional Models To Foundation Models, Zhaozheng Chen, Qianru Sun
Weakly-Supervised Semantic Segmentation With Image-Level Labels: From Traditional Models To Foundation Models, Zhaozheng Chen, Qianru Sun
Research Collection School Of Computing and Information Systems
The rapid development of deep learning has driven significant progress in image semantic segmentation—a fundamental task in computer vision. Semantic segmentation algorithms often depend on the availability of pixel-level labels (i.e., masks of objects), which are expensive, time consuming, and labor intensive. Weakly supervised semantic segmentation (WSSS) is an effective solution to avoid such labeling. It utilizes only partial or incomplete annotations and provides a cost-effective alternative to fully supervised semantic segmentation. In this article, our focus is on the WSSS with image-level labels, which is the most challenging form of WSSS. Our work has two parts. First, we conduct …
Synthesizing Multi-Person And Rare Pose Images For Human Pose Estimation, Liuqing Zhao, Zichen Tian, Zou Peng, Richang Hong, Qianru Sun
Synthesizing Multi-Person And Rare Pose Images For Human Pose Estimation, Liuqing Zhao, Zichen Tian, Zou Peng, Richang Hong, Qianru Sun
Research Collection School Of Computing and Information Systems
Human pose estimation (HPE) models underperform in recognizing rare poses because they suffer from data imbalance problems (i.e., there are few image samples for rare poses) in their training datasets. From a data perspective, the most intuitive solution is to synthesize data for rare poses. Specifically, the rule-based methods apply manual manipulations (such as Cutout and GridMask) to the existing data, so the limited diversity of the data constrains the model. An alternative method is to learn the underlying data distribution via deep generative models (such as ControlNet and HumanSD) and then sample “new data” from the distribution. This works …
Transforming Urban Dynamics: Harnessing Large Language Models For Smarter Mobility, Hao Xue, Ming Jin, Shirui Pan, Flora Salim, Guansong Pang
Transforming Urban Dynamics: Harnessing Large Language Models For Smarter Mobility, Hao Xue, Ming Jin, Shirui Pan, Flora Salim, Guansong Pang
Research Collection School Of Computing and Information Systems
Artificial intelligence (AI) has the potential to analyze mobility data and make mobility systems smarter by leveraging diverse data sources such as geospatial data, transportation logs, and real-time sensor data to optimize traffic flow, enhance public transportation systems, and support the development of autonomous vehicles. With the newly emerged generative AI paradigm, exemplified by large language models (LLMs), there is great potential to transform the current AI applications in mobility, transportation, and urban domains. This article provides an overview of recent efforts and aims to shed light on the challenges and future opportunities to facilitate the adaptation of LLMs for …
Fedart: A Neural Model Integrating Federated Learning And Adaptive Resonance Theory, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan
Fedart: A Neural Model Integrating Federated Learning And Adaptive Resonance Theory, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Federated Learning (FL) has emerged as a promising paradigm for collaborative model training across distributed clients while preserving data privacy. However, prevailing FL approaches aggregate the clients’ local models into a global model through multi-round iterative parameter averaging. This leads to the undesirable bias of the aggregated model towards certain clients in the presence of heterogeneous data distributions among the clients. Moreover, such approaches are restricted to supervised classification tasks and do not support unsupervised clustering. To address these limitations, we propose a novel one-shot FL approach called Federated Adaptive Resonance Theory (FedART) which leverages self-organizing Adaptive Resonance Theory (ART) …
A Survey Of Multilingual Large Language Models, Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, Philip S. Yu
A Survey Of Multilingual Large Language Models, Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, Philip S. Yu
Research Collection School Of Computing and Information Systems
Multilingual large language models (MLLMs) leverage advanced large language models to process and respond to queries across multiple languages, achieving significant success in polyglot tasks. Despite these breakthroughs, a comprehensive survey summarizing existing approaches and recent developments remains absent. To this end, this paper presents a unified and thorough review of the field, highlighting recent progress and emerging trends in MLLM research. The contributions of this paper are as follows. (1) Extensive survey: to our knowledge, this is the pioneering thorough review of multilingual alignment in MLLMs. (2) Unified taxonomy: we provide a unified framework to summarize the current progress …