Open Access. Powered by Scholars. Published by Universities.®

Discipline
Institution
Keyword
Publication Year
Publication
Publication Type
File Type

Articles 2821 - 2850 of 11193

Full-Text Articles in Artificial Intelligence and Robotics

Deep Representation Learning For Time Series Forecasting, Gerald Woo Aug 2024

Deep Representation Learning For Time Series Forecasting, Gerald Woo

Dissertations and Theses Collection (Open Access)

Time series forecasting has critical applications across business and scien- tific domains, such as demand forecasting, capacity planning and management, and anomaly detection. Being able to predict the future yields immense value, allowing us to make downstream decisions with more confidence. Deep learning for time series forecasting is a burgeoning area of research, moving away from simple linear models found in classical time series analysis literature, towards more expressive, data hungry neural network architectures.

In this thesis, we develop methods leveraging deep representation learning for time series forecasting, from exploring neural network architecture designs which encode inductive biases specific to …


Causvsr: Causality Inspired Visual Sentiment Recognition, Xinyue Zhang, Zhaoxia Wang, Hailing Wang, Jing Xiang, Chunwei Wu, Guitao Cao Aug 2024

Causvsr: Causality Inspired Visual Sentiment Recognition, Xinyue Zhang, Zhaoxia Wang, Hailing Wang, Jing Xiang, Chunwei Wu, Guitao Cao

Research Collection School Of Computing and Information Systems

Visual Sentiment Recognition (VSR) is an evolving field that aims to detect emotional tendencieswithin visual content. Despite its growing significance, detecting emotions depicted in visual content,such as images, faces challenges, notably the emergence of misleading or spurious correlationsof the contextual information. In response to these challenges, we propose a causality inspired VSRapproach, called CausVSR. CausVSR is rooted in the fundamental principles of Emotional Causalitytheory, mimicking the human process from receiving emotional stimuli to deriving emotional states.CausVSR takes a deliberate stride toward conquering the VSR challenges. It harnesses the power of astructural causal model, intricately designed to encapsulate the dynamic causal …


Cross-Problem Learning For Solving Vehicle Routing Problems, Zhuoyi Lin, Yaoxin Wu, Bangjian Zhou, Zhiguang Cao, Wen Song, Yingqian Zhang, Senthilnath Jayavelu Aug 2024

Cross-Problem Learning For Solving Vehicle Routing Problems, Zhuoyi Lin, Yaoxin Wu, Bangjian Zhou, Zhiguang Cao, Wen Song, Yingqian Zhang, Senthilnath Jayavelu

Research Collection School Of Computing and Information Systems

Existing neural heuristics often train a deep architecture from scratch for each specific vehicle routing problem (VRP), ignoring the transferable knowledge across different VRP variants. This paper proposes the cross-problem learning to assist heuristics training for different downstream VRP variants. Particularly, we modularize neural architectures for complex VRPs into 1) the backbone Transformer for tackling the travelling salesman problem (TSP), and 2) the additional lightweight modules for processing problem-specific features in complex VRPs. Accordingly, we propose to pre-train the backbone Transformer for TSP, and then apply it in the process of fine-tuning the Transformer models for each target VRP variant. …


Unifying Global-Local Representations In Salient Object Detection With Transformers, Sucheng Ren, Nanxuan Zhao, Qiang Wen, Guoqiang Han, Shengfeng He Aug 2024

Unifying Global-Local Representations In Salient Object Detection With Transformers, Sucheng Ren, Nanxuan Zhao, Qiang Wen, Guoqiang Han, Shengfeng He

Research Collection School Of Computing and Information Systems

The fully convolutional network (FCN) has dominated salient object detection for a long period. However, the locality of CNN requires the model deep enough to have a global receptive field and such a deep model always leads to the loss of local details. In this paper, we introduce a new attention-based encoder, vision transformer, into salient object detection to ensure the globalization of the representations from shallow to deep layers. With the global view in very shallow layers, the transformer encoder preserves more local representations to recover the spatial details in final saliency maps. Besides, as each layer can capture …


Speaker Verification In Agent-Generated Conversations, Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Ee-Peng Lim Aug 2024

Speaker Verification In Agent-Generated Conversations, Yizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang, Ee-Peng Lim

Research Collection School Of Computing and Information Systems

The recent success of large language models (LLMs) has attracted widespread interest to develop role-playing conversational agents personalized to the characteristics and styles of different speakers to enhance their abilities to perform both general and special purpose dialogue tasks. However, the ability to personalize the generated utterances to speakers, whether conducted by human or LLM, has not been well studied. To bridge this gap, our study introduces a novel evaluation challenge: speaker verification in agent-generated conversations, which aimed to verify whether two sets of utterances originate from the same speaker. To this end, we assemble a large dataset collection encompassing …


Self-Adaptive Psro : Towards An Automatic Population-Based Game Solver, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Xiao Huang, Hau Chan, Bo An Aug 2024

Self-Adaptive Psro : Towards An Automatic Population-Based Game Solver, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Xiao Huang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Policy-Space Response Oracles (PSRO) as a general algorithmic framework has achieved state-of-the-art performance in learning equilibrium policies of two-player zero-sum games. However, the hand-crafted hyperparameter value selection in most of the existing works requires extensive domain knowledge, forming the main barrier to applying PSRO to different games. In this work, we make the first attempt to investigate the possibility of self-adaptively determining the optimal hyperparameter values in the PSRO framework. Our contributions are three-fold: (1) Using several hyperparameters, we propose a parametric PSRO that unifies the gradient descent ascent (GDA) and different PSRO variants. (2) We propose the self-adaptive PSRO …


A Multimodal Foundation Agent For Financial Trading : Tool-Augmented, Diversified, And Generalist, Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An Aug 2024

A Multimodal Foundation Agent For Financial Trading : Tool-Augmented, Diversified, And Generalist, Wentao Zhang, Lingxuan Zhao, Haochong Xia, Shuo Sun, Jiaze Sun, Molei Qin, Xinyi Li, Yuqing Zhao, Yilei Zhao, Xinyu Cai, Longtao Zheng, Xinrun Wang, Bo An

Research Collection School Of Computing and Information Systems

Financial trading is a crucial component of the markets, informed by a multimodal information landscape encompassing news, prices, and Kline charts, and encompasses diverse tasks such as quantitative trading and high-frequency trading with various assets. While advanced AI techniques like deep learning and reinforcement learning are extensively utilized in finance, their application in financial trading tasks often faces challenges due to inadequate handling of multimodal data and limited generalizability across various tasks. To address these challenges, we present FinAgent, a multimodal foundational agent with tool augmentation for financial trading. FinAgent's market intelligence module processes a diverse range of data-numerical, textual, …


Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang Aug 2024

Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang

Research Collection School Of Computing and Information Systems

High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become another appealing approach for HFT due to its terrific ability of handling high-dimensional financial data and solving sophisticated sequential decision-making problems, e.g., hierarchical reinforcement learning (HRL) has shown its promising performance on second-level HFT by training a router to select only one sub-agent from the agent pool to execute the current transaction. However, existing RL methods for HFT still have some defects: 1) standard RL-based trading agents suffer from the overfitting …


Topic Modeling On Document Networks With Dirichlet Optimal Transport Barycenter, Ce Zhang, Hady Wirawan Lauw Aug 2024

Topic Modeling On Document Networks With Dirichlet Optimal Transport Barycenter, Ce Zhang, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Text documents are often interconnected in a network structure, e.g., academic papers via citations, Web pages via hyperlinks. On the one hand, though Graph Neural Networks (GNNs) have shown promising ability to derive effective embeddings for such networked documents, they do not assume a latent topic structure and result in uninterpretable embeddings. On the other hand, topic models can infer semantically interpretable topic distributions for documents by associating each topic with a group of understandable key words. However, most topic models mainly focus on plain text within documents and fail to leverage network structure across documents. Network connectivity reveals topic …


Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An Aug 2024

Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Nash Equilibrium (NE) is the canonical solution concept of game theory, which provides an elegant tool to understand the rationalities. Though mixed strategy NE exists in any game with finite players and actions, computing NE in two- or multi-player general-sum games is PPAD-Complete. Various alternative solutions, e.g., Correlated Equilibrium (CE), and learning methods, e.g., fictitious play (FP), are proposed to approximate NE. For convenience, we call these methods as "inexact solvers", or "solvers" for short. However, the alternative solutions differ from NE and the learning methods generally fail to converge to NE. Therefore, in this work, we propose REinforcement Nash …


Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim Aug 2024

Exploring A Multimodal Fusion-Based Deep Learning Network For Detecting Facial Palsy, Heng Yim Nicole Oo, Min Hun Lee, J. H. Lim

Research Collection School Of Computing and Information Systems

Algorithmic detection of facial palsy offers the potential to improve current practices, which usually involve labor-intensive and subjective assessment by clinicians. In this paper, we present a multimodal fusion-based deep learning model that utilizes unstructured data (i.e. an image frame with facial line segments) and structured data (i.e. features of facial expressions) to detect facial palsy. We then contribute to a study to analyze the effect of different data modalities and the benefits of a multimodal fusion-based approach using videos of 21 facial palsy patients. Our experimental results show that among various data modalities (i.e. unstructured data - RGB images …


Towards Gradient-Based Time-Series Explanations Through A Spatiotemporal Attention Network, Min Hun Lee Aug 2024

Towards Gradient-Based Time-Series Explanations Through A Spatiotemporal Attention Network, Min Hun Lee

Research Collection School Of Computing and Information Systems

In this paper, we explore the feasibility of using a transformer-based, spatiotemporal attention network (STAN) for gradient-based time-series explanations. First, we trained the STAN model for video classifications using the global and local views of data and weakly supervised labels on time-series data (i.e. the type of an activity). We then leveraged a gradient-based XAI technique (e.g. saliency map) to identify salient frames of time-series data. According to the experiments using the datasets of four medically relevant activities, the STAN model demonstrated its potential to identify important frames of videos.


Certified Policy Verification And Synthesis For Mdps Under Distributional Reach-Avoidance Properties, S. Akshay, Krishnendu Chatterjee, Tobias Meggendorfer, Dorde Zikelic Aug 2024

Certified Policy Verification And Synthesis For Mdps Under Distributional Reach-Avoidance Properties, S. Akshay, Krishnendu Chatterjee, Tobias Meggendorfer, Dorde Zikelic

Research Collection School Of Computing and Information Systems

Markov Decision Processes (MDPs) are a classical model for decision making in the presence of uncertainty. Often they are viewed as state transformers with planning objectives defined with respect to paths over MDP states. An increasingly popular alternative is to view them as distribution transformers, giving rise to a sequence of probability distributions over MDP states. For instance, reachability and safety properties in modeling robot swarms or chemical reaction networks are naturally defined in terms of probability distributions over states. Verifying such distributional properties is known to be hard and often beyond the reach of classical state-based verification techniques. In …


Solving Long-Run Average Reward Robust Mdps Via Stochastic Games, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Mehrdad Karrabi, Petr Novotný, Dorde Zikelic Aug 2024

Solving Long-Run Average Reward Robust Mdps Via Stochastic Games, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Mehrdad Karrabi, Petr Novotný, Dorde Zikelic

Research Collection School Of Computing and Information Systems

Markov decision processes (MDPs) provide a standard framework for sequential decision making under uncertainty. However, MDPs do not take uncertainty in transition probabilities into account. Robust Markov decision processes (RMDPs) address this shortcoming of MDPs by assigning to each transition an uncertainty set rather than a single probability value. In this work, we consider polytopic RMDPs in which all uncertainty sets are polytopes and study the problem of solving long-run average reward polytopic RMDPs. We present a novel perspective on this problem and show that it can be reduced to solving long-run average reward turn-based stochastic games with finite state …


Task Scheduling Strategy For 3dpcp Considering Multidynamic Information Perturbation In Green Scene, Jianjia He, Jian Wu, Keng Siau Aug 2024

Task Scheduling Strategy For 3dpcp Considering Multidynamic Information Perturbation In Green Scene, Jianjia He, Jian Wu, Keng Siau

Research Collection School Of Computing and Information Systems

The 3D printing cloud platform (3DPCP) plays a pivotal role in breaking down the information silos between supply and demand, effectively reducing waste through information integration and intelligent production. However, due to the complexity of 3DPCP scheduling in green scenes and the multidynamic information perturbations, unveils problems in traditional task scheduling methods in 3DPCP. These issues manifest as incomplete considerations, subpar green performance, and weak adaptability to dynamic changes. There is an urgent need to design practical methods to realize the multidynamic information perturbations in green scenes within 3DPCP. Therefore, this article first defines the 3DPCP task scheduling problem for …


Fuel-Saving Route Planning With Data-Driven And Learning-Based Approaches: A Systematic Solution For Harbor Tugs, Shengming Wang, Xiaocai Zhang, Jing Li, Xiaoyang Wei, Hoong Chuin Lau, Bing Tian Dai, Binbin Huang Huang, Zhe Xiao, Xiuju Fu, Zheng Qin Aug 2024

Fuel-Saving Route Planning With Data-Driven And Learning-Based Approaches: A Systematic Solution For Harbor Tugs, Shengming Wang, Xiaocai Zhang, Jing Li, Xiaoyang Wei, Hoong Chuin Lau, Bing Tian Dai, Binbin Huang Huang, Zhe Xiao, Xiuju Fu, Zheng Qin

Research Collection School Of Computing and Information Systems

In recent years, there are trends toward cleaner port environments through enforcement by imposed legislation. Transit optimisation of fuel-based port service boats like harbour tugs has emerged as a critical task to reduce fuel consumption and carbon emission. In this paper, an innovative learning-based method, comprising a Reinforcement Learning (RL) model together with a fuel consumption prediction model, was proposed to formulate fuel-saving transit routes. Firstly, an ensemble model is established by combining a Long Short-Term Memory (LSTM) model with a Multilayer Perceptron (MLP) model, predicting fuel use based on tugboat movement and environment factors. Subsequently, an innovative RL based …


Optimization Of Customer Service And Driver Dispatch Areas For On-Demand Food Delivery, Jingfeng Yang, Hoong Chuin Lau, Hai Wang Aug 2024

Optimization Of Customer Service And Driver Dispatch Areas For On-Demand Food Delivery, Jingfeng Yang, Hoong Chuin Lau, Hai Wang

Research Collection School Of Computing and Information Systems

With the rapid development and popularization of mobile and wireless communication technologies, on-demand food delivery (OFD) platforms have been able to connect restaurants, customers, and drivers in real time, drastically changing dining and food delivery services. Motivated by the critical need for supply and demand management in the on-demand food delivery market, we focus on the optimization of customer service area and driver dispatch area for on-demand food delivery services. Specifically, for each restaurant, the platform needs to decide the (1) customer service area (CSA), i.e., the surrounding area within which customers can see the restaurant’s information and order food …


Enabling Sustainable Freight Forwarding Network Via Collaborative Games, Pang Jin Tan, Shih-Fen Cheng, Richard Chen Aug 2024

Enabling Sustainable Freight Forwarding Network Via Collaborative Games, Pang Jin Tan, Shih-Fen Cheng, Richard Chen

Research Collection School Of Computing and Information Systems

Freight forwarding plays a crucial role in facilitating global trade and logistics. However, as the freight forwarding market is extremely fragmented, freight forwarders often face the issue of not being able to fill the available shipping capacity. This recurrent issue motivates the creation of various freight forwarding networks that aim at exchanging capacities and demands so that the resource utilization of individual freight forwarders can be maximized. In this paper, we focus on how to design such a collaborative network based on collaborative game theory, with the Shapley value representing a fair scheme for profit sharing. Noting that the exact …


Is Aggregation The Only Choice? Federated Learning Via Layer-Wise Model Recombination, Ming Hu, Zhihao Yue, Xiaofei Xie, Cheng Chen Chen Aug 2024

Is Aggregation The Only Choice? Federated Learning Via Layer-Wise Model Recombination, Ming Hu, Zhihao Yue, Xiaofei Xie, Cheng Chen Chen

Research Collection School Of Computing and Information Systems

Although Federated Learning (FL) enables global model training Xiaofei Xie [email protected] Singapore Management University Singapore, Singapore Xian Wei [email protected] East China Normal University Shanghai, China Mingsong Chen∗ [email protected] East China Normal University Shanghai, China • Computing methodologies → Distributed artificial intelligence. across clients without compromising their raw data, due to the unevenly distributed data among clients, existing Federated Averaging (FedAvg)-based methods suffer from the problem of low inference performance. Specifically, different data distributions among clients lead to various optimization directions of local models. Aggregating local models usually results in a low-generalized global model, which performs worse on most of the …


Contrastive General Graph Matching With Adaptive Augmentation Sampling, Jianyuan Bo, Yuan Fang Aug 2024

Contrastive General Graph Matching With Adaptive Augmentation Sampling, Jianyuan Bo, Yuan Fang

Research Collection School Of Computing and Information Systems

Graph matching has important applications in pattern recognition and beyond. Current approaches predominantly adopt supervised learning, demanding extensive labeled data which can be limited or costly. Meanwhile, self-supervised learning methods for graph matching often require additional side information such as extra categorical information and input features, limiting their application to the general case. Moreover, designing the optimal graph augmentations for self-supervised graph matching presents another challenge to ensure robustness and effcacy. To address these issues, we introduce a novel Graph-centric Contrastive framework for Graph Matching (GCGM), capitalizing on a vast pool of graph augmentations for contrastive learning, yet without needing …


A Learned Generalized Geodesic Distance Function-Based Approach For Node Feature Augmentation On Graphs, Amitoz Azad, Yuan Fang Aug 2024

A Learned Generalized Geodesic Distance Function-Based Approach For Node Feature Augmentation On Graphs, Amitoz Azad, Yuan Fang

Research Collection School Of Computing and Information Systems

Geodesic distances on manifolds have numerous applications in image processing, computer graphics and computer vision. In this work, we introduce an approach called 'LGGD' (Learned Generalized Geodesic Distances). This method involves generating node features by learning a generalized geodesic distance function through a training pipeline that incorporates training data, graph topology and the node content features. The strength of this method lies in the proven robustness of the generalized geodesic distances to noise and outliers. Our contributions encompass improved performance in node classification tasks, competitive results with state-of-the-art methods on real-world graph datasets, the demonstration of the learnability of parameters …


Sibo : A Simple Booster For Parameter-Efficient Fine-Tuning, Zhihao Wen, Jie Zhang, Yuan Fang Aug 2024

Sibo : A Simple Booster For Parameter-Efficient Fine-Tuning, Zhihao Wen, Jie Zhang, Yuan Fang

Research Collection School Of Computing and Information Systems

Fine-tuning all parameters of large language models (LLMs) necessitates substantial computational power and extended time. Latest advancements in parameter-efficient fine-tuning (PEFT) techniques, such as Adapter tuning and LoRA, allow for adjustments to only a minor fraction of the parameters of these LLMs. Concurrently, it has been noted that the issue of over-smoothing diminishes the effectiveness of these Transformer-based LLMs, resulting in suboptimal performances in downstream tasks. In this paper, we present SIBO, which is a SImple BOoster to enhance PEFT, by injecting an initial residual. SIBO is straightforward and readily extensible to a range of state-of-the-art PEFT techniques to alleviate …


Heterogeneous Graph Transformer With Poly-Tokenization, Zhiyuan Lu, Yuan Fang, Cheng Yang, Chuan Shi Aug 2024

Heterogeneous Graph Transformer With Poly-Tokenization, Zhiyuan Lu, Yuan Fang, Cheng Yang, Chuan Shi

Research Collection School Of Computing and Information Systems

Graph neural networks have shown widespread success for learning on graphs, but they still face fundamental drawbacks, such as limited expressive power, over-smoothing, and over-squashing. Meanwhile, the transformer architecture offers a potential solution to these issues. However, existing graph transformers primarily cater to homogeneous graphs and are unable to model the intricate semantics of heterogeneous graphs. Moreover, unlike small molecular graphs where the entire graph can be considered as the receptive field in graph transformers, real-world heterogeneous graphs comprise a significantly larger number of nodes and cannot be entirely treated as such. Consequently, existing graph transformers struggle to capture the …


Tackling Stackelberg Network Interdiction Against A Boundedly Rational Adversary, Tien Mai, Avinandan Bose, Arunesh Sinha, Thanh Nguyen, Ayushman Kumar Singh Aug 2024

Tackling Stackelberg Network Interdiction Against A Boundedly Rational Adversary, Tien Mai, Avinandan Bose, Arunesh Sinha, Thanh Nguyen, Ayushman Kumar Singh

Research Collection School Of Computing and Information Systems

This work studies Stackelberg network interdiction games --- an important class of games in which a defender first allocates (randomized) defense resources to a set of critical nodes on a graph while an adversary chooses its path to attack these nodes accordingly. We consider a boundedly rational adversary in which the adversary's response model is based on a dynamic form of classic logit-based (quantal response) discrete choice models. The resulting optimization is non-convex and additionally, involves complex terms that sum over exponentially many paths. We tackle these computational challenges by presenting new efficient algorithms with solution guarantees. First, we present …


Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria Aug 2024

Empathyear : An Open-Source Avatar Multimodal Empathetic Chatbot, Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, Erik Cambria

Research Collection School Of Computing and Information Systems

This paper introduces EmpathyEar, a pioneering open-source, avatar-based multimodal empathetic chatbot, to fill the gap in traditional text-only empathetic response generation (ERG) systems. Leveraging the advancements of a large language model, combined with multimodal encoders and generators, EmpathyEar supports user inputs in any combination of text, sound, and vision, and produces multimodal empathetic responses, offering users, not just textual responses but also digital avatars with talking faces and synchronized speeches. A series of emotion-aware instruction-tuning is performed for comprehensive emotional understanding and generation capabilities. In this way, EmpathyEar provides users with responses that achieve a deeper emotional resonance, closely emulating …


Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin Aug 2024

Planning Like Human : A Dual-Process Framework For Dialogue Planning, Tao He, Lizi Liao, Yixin Cao, Yuanxing Liu, Ming Liu, Zerui Chen, Bing Qin

Research Collection School Of Computing and Information Systems

In proactive dialogue, the challenge lies not just in generating responses but in steering conversations toward predetermined goals, a task where Large Language Models (LLMs) typically struggle due to their reactive nature. Traditional approaches to enhance dialogue planning in LLMs, ranging from elaborate prompt engineering to the integration of policy networks, either face efficiency issues or deliver suboptimal performance. Inspired by the dual-process theory in psychology, which identifies two distinct modes of thinking—intuitive (fast) and analytical (slow), we propose the Dual-Process Dialogue Planning (DPDP) framework. DPDP embodies this theory through two complementary planning systems: an instinctive policy model for familiar …


Analyzing Temporal Complex Events With Large Language Models? A Benchmark Towards Temporal, Long Context Understanding, Zhihan Zhang, Yixin Cao, Chenchen Ye, Ma. Yunshan, Lizi Liao, Tat-Seng Chua Aug 2024

Analyzing Temporal Complex Events With Large Language Models? A Benchmark Towards Temporal, Long Context Understanding, Zhihan Zhang, Yixin Cao, Chenchen Ye, Ma. Yunshan, Lizi Liao, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

The digital landscape is rapidly evolving with an ever-increasing volume of online news, emphasizing the need for swift and precise analysis of complex events.We refer to the complex events composed of many news articles over an extended period as Temporal Complex Event (TCE). This paper proposes a novel approach using Large Language Models (LLMs) to systematically extract and analyze the event chain within TCE, characterized by their key points and timestamps. We establish a benchmark, named TCELongBench, to evaluate the proficiency of LLMs in handling temporal dynamics and understanding extensive text. This benchmark encompasses three distinct tasks - reading comprehension, …


Synergizing Large Language Models And Pre-Trained Smaller Models For Conversational Intent Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Jing Jiang Aug 2024

Synergizing Large Language Models And Pre-Trained Smaller Models For Conversational Intent Discovery, Jinggui Liang, Lizi Liao, Hao Fei, Jing Jiang

Research Collection School Of Computing and Information Systems

In Conversational Intent Discovery (CID), Small Language Models (SLMs) struggle with overfitting to familiar intents and fail to label newly discovered ones. This issue stems from their limited grasp of semantic nuances and their intrinsically discriminative framework. Therefore, we propose Synergizing Large Language Models (LLMs) with pre-trained SLMs for CID (SynCID). It harnesses the profound semantic comprehension of LLMs alongside the operational agility of SLMs. By utilizing LLMs to refine both utterances and existing intent labels, SynCID significantly enhances the semantic depth, subsequently realigning these enriched descriptors within the SLMs’ feature space to correct cluster distortion and promote robust learning …


A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua Aug 2024

A Survey On Neural Question Generation : Methods, Applications, And Prospects, Shasha Guo, Lizi Liao, Cuiping Li, Tat-Seng Chua

Research Collection School Of Computing and Information Systems

In this survey, we present a detailed examination of the advancements in Neural Question Generation (NQG), a field leveraging neural network techniques to generate relevant questions from diverse inputs like knowledge bases, texts, and images. The survey begins with an overview of NQG’s background, encompassing the task’s problem formulation, prevalent benchmark datasets, established evaluation metrics, and notable applications. It then methodically classifies NQG approaches into three predominant categories: structured NQG, which utilizes organized data sources, unstructured NQG, focusing on more loosely structured inputs like texts or visual content, and hybrid NQG, drawing on diverse input modalities. This classification is followed …


Offensive Content Detection In Online Social Platforms, Ebuka Okpala Aug 2024

Offensive Content Detection In Online Social Platforms, Ebuka Okpala

All Dissertations

Online social platforms enable users to connect with large, diverse audiences and the ability for a message or content to flow from one user to another user, user to followers, followers to user, and followers to followers. Of course, the advantages of this are apparent, and the dangers are also clearly obvious. The user-generated content could be abusive, offensive, or hateful to other users, possibly leading to adverse health effects or offline harm. As more of society's public discourse and interaction move online and these platforms grow and increase their reach, it is inherently important to protect the safety of …