Open Access. Powered by Scholars. Published by Universities.®

Deep reinforcement learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 58

Full-Text Articles in Artificial Intelligence and Robotics

Approximation And Learning-Based Algorithms For Influence Maximization In Multilayer Social Networks, Xueqin Chang, Ruize Liu, Qing Liu, Baihua Zheng, Yunjun Gao Aug 2026

Approximation And Learning-Based Algorithms For Influence Maximization In Multilayer Social Networks, Xueqin Chang, Ruize Liu, Qing Liu, Baihua Zheng, Yunjun Gao

Research Collection School Of Computing and Information Systems

Motivated by the observation that users in the real world often engage across multiple social networks simultaneously, we study the problem of influence maximization in multilayer social networks (Mlim), aiming to select a small set of nodes that maximizes the total influence spread across all layers. To this end, we introduce a hybrid propagation model that jointly captures layer-specific diffusion dynamics and probabilistic cross-layer propagation. Based on this model, we formally define the Mlim problem and establish its NP-hardness, monotonicity, and submodularity. To address the Mlim problem, we first propose a greedy baseline Mlim-Greedy, which achieves a (1-1/e) approximation. Since …


Distributed Optimization For Integrated Energy Based On Multi-Agent Reinforcement Learning, Caixia Tao, Naikun Chen, Fengyang Gao, Jiangang Zhang Feb 2026

Distributed Optimization For Integrated Energy Based On Multi-Agent Reinforcement Learning, Caixia Tao, Naikun Chen, Fengyang Gao, Jiangang Zhang

Journal of System Simulation

Abstract: To address the energy management and privacy preservation problems faced by the coordinated optimization of distributed integrated energy systems, a distributed coordinated optimization strategy based on the multi-agent proximal policy optimization algorithm was proposed. An energy management model was established under the MDP framework; the electrical and thermal heterogeneous energy characteristics were considered; a multi-region two-layer interaction mechanism was constructed. Under the framework of centralized training and decentralized execution, homomorphic encryption was utilized to avoid privacy leakage during the coordination process, while accurately quantifying individual contributions to mitigate the problem of variance explosion in multi-agent policy evaluation. In the …


Intelligent Air Combat Decision-Making Method Based On Bigru And Priority Dynamic Sampling, Zhengkun Ding, Jiaqi Liu, Junzheng Xu, Yuezhu Xu, Xingmei Wang Feb 2026

Intelligent Air Combat Decision-Making Method Based On Bigru And Priority Dynamic Sampling, Zhengkun Ding, Jiaqi Liu, Junzheng Xu, Yuezhu Xu, Xingmei Wang

Journal of System Simulation

Abstract: Current multi-agent reinforcement learning algorithms suffer from low efficiency in utilizing experience data and difficulties in setting appropriate learning rates. To address these issues, this paper proposed a BiGRU multi-agent PPO with priority sampling and dynamic learning rate. The algorithm incorporated a BiGRU network to enhance the policy network's ability to model temporal information. A priority partial sampling mechanism was introduced to improve the utilization efficiency of high-value experience data. Additionally, an improved Adam optimizer with dynamic learning rate adjustment was employed to address the challenge of learning rate configuration. Simulation experiment results demonstrate that the algorithm significantly …


Optimization Of Dynamic Weapon Target Assignment Considering Random Disturbances, Zhenzu Bai, Yizhi Hou, Zhangming He, Juhui Wei, Haiyin Zhou, Jiongqi Wang Dec 2025

Optimization Of Dynamic Weapon Target Assignment Considering Random Disturbances, Zhenzu Bai, Yizhi Hou, Zhangming He, Juhui Wei, Haiyin Zhou, Jiongqi Wang

Journal of System Simulation

Abstract: The impact of various random disturbances in the actual command and control environment of unmanned systems on problem modeling and solving of weapon target assignment was considered, and three types of uncertainty disturbance constraints were investigated. A multi-objective dynamic sensor weapon target assignment model was established. By considering the issues of model property changes caused by disturbances and insufficient robustness of the traditional single-operator solving algorithm, a multi-operator constrained multi-objective evolutionary framework based on the deep Q-network was proposed. The algorithm described the convergence, diversity, and feasibility of the population in both the objective and decision spaces. It established …


Optimization Of Service Caching And Computation Offloading In Digital Twin Cloud-Edge Networks, Jiayu Zheng, Zhuxue Mai, Zheyi Chen Nov 2025

Optimization Of Service Caching And Computation Offloading In Digital Twin Cloud-Edge Networks, Jiayu Zheng, Zhuxue Mai, Zheyi Chen

Journal of System Simulation

Abstract: In mobile edge computing (MEC), to satisfy diverse user demands by jointly optimizing service caching and computation offloading and address low-efficiency resource utilization caused by irrational resource allocation, this paper proposed a novel joint optimization of service caching and computation offloading with a convex-optimization-enabled deep reinforcement learning (JCO-CR) method. Additionally, a new model for digital twin cloud-edge networks (DTCEN) was constructed. The joint optimization of service caching and computation offloading was decoupled into two sub-problems, which were solved by an improved deep reinforcement learning method and convex optimization theory, respectively. Simulation experiments demonstrate that the proposed JCO-CR method …


Evolutionary Reinforcement Learning Based On Elite Instruction And Random Search, Jian Di, Xue Wan, Limei Jiang Nov 2025

Evolutionary Reinforcement Learning Based On Elite Instruction And Random Search, Jian Di, Xue Wan, Limei Jiang

Journal of System Simulation

Abstract: Evolutionary reinforcement learning currently suffers from low sample efficiency, a single coupling method, and poor convergence, which can affect its performance and scaling. To address this issue, an improved algorithm based on elite gradient instruction and double random search was proposed. The direction of the reinforcement strategy gradient update was corrected by introducing elite strategy gradient guidance carrying evolutionary information during reinforcement strategy training. Double stochastic search was used to replace the original evolutionary component, reducing the complexity of the algorithm while making the policy search meaningful and controllable in the parameter space. The introduction of complete replacement information …


Deep Reinforcement Learning For Solving The Stochastic E-Waste Collection Problem, Dang Viet Anh Nguyen, Aldy Gunawan, Mustafa Misir, Kwan Hui Lim, Pieter Vansteenwegen Nov 2025

Deep Reinforcement Learning For Solving The Stochastic E-Waste Collection Problem, Dang Viet Anh Nguyen, Aldy Gunawan, Mustafa Misir, Kwan Hui Lim, Pieter Vansteenwegen

Research Collection School Of Computing and Information Systems

With the growing influence of the internet and information technology, Electrical and Electronic Equipment (EEE) has become a gateway to technological innovations. However, discarded devices, also called e-waste, pose a significant threat to the environment and human health if not properly treated, disposed of, or recycled. In this study, we extend a novel model for the e-waste collection in an urban context: the Heterogeneous VRP with Multiple Time Windows and Stochastic Travel Times (HVRP-MTWSTT). We propose a solution method that employs deep reinforcement learning to guide local search heuristics (DRL-LSH). The contributions of this paper are as follows: (1) HVRP-MTWSTT …


Path Planning Of Improved Rrt Algorithm Based On Deep Reinforcement Learning, Xiuman Liang, Ziliang Liu, Zhendong Liu Oct 2025

Path Planning Of Improved Rrt Algorithm Based On Deep Reinforcement Learning, Xiuman Liang, Ziliang Liu, Zhendong Liu

Journal of System Simulation

Abstract: To address the low planning efficiency, poor safety, and limited practicability of the RRT algorithm in global path planning within complex three-dimensional environments, which fail to meet the requirements of planning the safe flight path of UAVs, an improved SAC-RRT algorithm was proposed, which fused SAC deep reinforcement learning algorithm and RRT algorithm. A target point bias strategy and a dynamic step size based on the SAC decision-making network were designed to reduce the blindness of RRT. A random point correction process was designed to optimize the position of random points based on actions from the decision network and …


Optimization Dispatch Method For High-Proportion Renewable Energy Power Systems Based On Sc-Ppo, Zhongkai Xu, Chenyang Chu, Kai Xie, Ruizhuo Zhao, Wenjun Ke Oct 2025

Optimization Dispatch Method For High-Proportion Renewable Energy Power Systems Based On Sc-Ppo, Zhongkai Xu, Chenyang Chu, Kai Xie, Ruizhuo Zhao, Wenjun Ke

Journal of System Simulation

Abstract: The high proportion of renewable energy integration brings significant challenges of randomness, multi-objective coupling, and security constraints to power systems. Traditional model-driven methods have limitations in modeling accuracy and adaptability. To address these issues, this paper proposed a safety-constrained PPO algorithm (SC-PPO). The method included three improvements. A temporal convolutional network was utilized to construct a dynamic state encoder that integrated historical operation, real-time monitoring, and prediction data to form a causal state representation. A hierarchical reward structure was designed, and an adaptive weighting mechanism based on constraint satisfaction degree was introduced to coordinate multi-objective optimization. Physical constraint projection …


Solving The Vehicle Routing Problem Based On Deep Reinforcement Learning, Ming Jiang, Tao He Sep 2025

Solving The Vehicle Routing Problem Based On Deep Reinforcement Learning, Ming Jiang, Tao He

Journal of System Simulation

Abstract: The capacitated vehicle routing problem (CVRP) is a well-known combinatorial optimization challenge recognized as NP-hard due to its significant complexity. Building upon existing research, this paper introduces a novel end-to-end deep reinforcement learning approach based on a multi-pointer Transformer to tackle the CVRP. The proposed algorithm employs an invertible residual network in the encoder to encode input features, effectively reducing memory consumption. In the decoder, a multipointer network determines the probability distribution of solutions. To further enhance the performance of CVRP solutions, the algorithm leverages the symmetry in combinatorial optimization by implementing multi-trajectory parallel processing during both training …


Robot Path Planning Based On Improved A-Ddqn Algorithm, Peilong Ni, Pengjun Mao, Ning Wang, Mengjie Yang Sep 2025

Robot Path Planning Based On Improved A-Ddqn Algorithm, Peilong Ni, Pengjun Mao, Ning Wang, Mengjie Yang

Journal of System Simulation

Abstract: An improved A-DDQN algorithm is proposed to address the challenges of reward sparsity and the inability to distinguish sample importance in traditional DQN algorithms during robot path planning. Building on the original DQN, an enhancement is made by incorporating the Double-DQN approach, which updates the predictive Q-value network based on actions selected by the Q network, rather than directly using the predicted Q-values for action selection, thereby mitigating overestimation issues. Secondly, the concept of artificial potential field (APF) is introduced to design specific rewards for each step of the robot's movement, guiding the robot and addressing the problem of …


Neuro-Ins: A Learning-Based One-Shot Node Insertion For Dynamic Routing Problems, Zhiqin Zhang, Jingfeng Yang, Zhiguang Cao, Hoong Chuin Lau Sep 2025

Neuro-Ins: A Learning-Based One-Shot Node Insertion For Dynamic Routing Problems, Zhiqin Zhang, Jingfeng Yang, Zhiguang Cao, Hoong Chuin Lau

Research Collection School Of Computing and Information Systems

The rise in instant delivery services necessitates efficient route planning in last-mile delivery scenarios, where new orders arrive dynamically and need to be integrated into existing routes. In such contexts, complete re-optimization of routes are not permitted, and node insertion to existing route sequences is the only viable option. However, many existing heuristics for node insertion, such as the Cheapest Insertion (CI) method, are myopic and often result in suboptimal solutions retrospectively. This paper presents Neuro-Ins, an initial yet novel attempt at harnessing a learning-based framework to handle the insertion of new orders for the Pickup and Delivery Problem (PDP). …


Dgl: Dynamic Global-Local Information Aggregation For Scalable Vrp Generalization With Self-Improvement Learning, Yubin Xiao, Yuesong Wu, Rui Cao, Di Wang, Zhiguang Cao, Xuan Wu, Peng Zhao, Yuanshu Li, You Zhou, Yuan Jiang Sep 2025

Dgl: Dynamic Global-Local Information Aggregation For Scalable Vrp Generalization With Self-Improvement Learning, Yubin Xiao, Yuesong Wu, Rui Cao, Di Wang, Zhiguang Cao, Xuan Wu, Peng Zhao, Yuanshu Li, You Zhou, Yuan Jiang

Research Collection School Of Computing and Information Systems

The Vehicle Routing Problem (VRP) is a critical combinatorial optimization problem with wide-reaching real-world applications, particularly in logistics, transportation. While neural network-based VRP solvers have shown impressive results on test instances similar to training data, their performance often degrades when faced with varying scales and unseen distributions, limiting their practical applicability. To overcome these limitations, we introduce DGL (Dynamic Global-Local Information Aggregation), a novel model that combines global and local information to effectively solve VRPs. DGL dynamically adjusts local node selections within a localized range, capturing local invariance across problems of different scales and distributions, thereby enhancing generalization. At the …


Preference-Based Deep Reinforcement Learning For Historical Route Estimation, Boshen Pan, Yaoxin Wu, Zhiguang Cao, Yaqing Hou, Guangyu Zou, Qiang Zhang Sep 2025

Preference-Based Deep Reinforcement Learning For Historical Route Estimation, Boshen Pan, Yaoxin Wu, Zhiguang Cao, Yaqing Hou, Guangyu Zou, Qiang Zhang

Research Collection School Of Computing and Information Systems

Recent Deep Reinforcement Learning (DRL) techniques have advanced solutions to Vehicle Routing Problems (VRPs). However, many of these methods focus exclusively on optimizing distance-oriented objectives (i.e., minimizing route length), often overlooking the implicit drivers' preferences for routes. These preferences, which are crucial in practice, are challenging to model using traditional DRL approaches. To address this gap, we propose a preference-based DRL method characterized by its reward design and optimization objective, which is specialized to learn historical route preferences. Our experiments demonstrate that the method aligns generated solutions more closely with human preferences. Moreover, it exhibits strong generalization performance across a …


Research On Policy Representation In Deep Reinforcement Learning, Zhen Chen, Zhuoyi Wu, Lin Zhang Jul 2025

Research On Policy Representation In Deep Reinforcement Learning, Zhen Chen, Zhuoyi Wu, Lin Zhang

Journal of System Simulation

Abstract: Deep reinforcement learning (DRL) has achieved remarkable success in various domains. Nevertheless, existing policy networks in DRL still face significant challenges in areas such as generalizability, multi-task adaptability, and sample efficiency. Policy representation, as a crucial research direction for enhancing DRL capabilities, aims to improve an agent's adaptability to environmental changes and novel tasks by constructing more efficient and generalizable forms of policy expression. This paper provided a concise overview of key research advances in the field of policy representation. It introduced diverse policy architectures, ranging from traditional multi-layer perceptron (MLP) -based policies to those based on pointer networks, …


Diversity Optimization For Travelling Salesman Problem Via Deep Reinforcement Learning, Qi Li, Zhiguang Cao, Yining Ma, Yaoxin Wu, Yue-Jiao Gong Jul 2025

Diversity Optimization For Travelling Salesman Problem Via Deep Reinforcement Learning, Qi Li, Zhiguang Cao, Yining Ma, Yaoxin Wu, Yue-Jiao Gong

Research Collection School Of Computing and Information Systems

Existing neural methods for the Travelling Salesman Problem (TSP) mostly aim at finding a single optimal solution. To discover diverse yet high-quality solutions for Multi-Solution TSP (MSTSP), we propose a novel deep reinforcement learning based neural solver, which is primarily featured by an encoder-decoder structured policy. Concretely, on the one hand, a Relativization Filter (RF) is designed to enhance the robustness of the encoder to affine transformations of the instances, so as to potentially improve the quality of the found solutions. On the other hand, a Multi-Attentive Adaptive Active Search (MA3S) is tailored to allow the decoders to strike a …


Collaborative Federated Learning For Robots In Heterogeneous Environments, Karlan Schneider Jun 2025

Collaborative Federated Learning For Robots In Heterogeneous Environments, Karlan Schneider

Electronic Theses and Dissertations

This research investigates the performance of Federated Averaging (FedAvg) in simulated Federated Learning (FL) scenarios with varying degrees of environmental heterogeneity among robotic agents. The study explores the impact of data heterogeneity on both the convergence of FedAvg and the fairness of learning, with regard to consistency of performance across agents. Experiments were conducted with simulated robots trained to perform a target collection task, where a subset of agents encountered an unfamiliar environment. The results demonstrate that while FedAvg exhibits resilience to the introduction of new environmental data, it struggles to ensure both convergence and fairness in heterogeneous settings. Specifically, …


A Quadrotor Trajectory Tracking Control Method Based On Deep Reinforcement Learning, Guohua Wu, Jiaheng Zeng, Dezhi Wang, Long Zheng, Wei Zou May 2025

A Quadrotor Trajectory Tracking Control Method Based On Deep Reinforcement Learning, Guohua Wu, Jiaheng Zeng, Dezhi Wang, Long Zheng, Wei Zou

Journal of System Simulation

Abstract: Traditional quadrotor controllers, constrained by fixed model equation structures, encounter challenges in addressing control errors stemming from variations in parameters and environmental disturbances. This paper proposes a deep reinforcement learning solution for the quadrotor trajectory following control problem. We present the PPO-SAG algorithm incorporated into the PPO framework, utilizing adaptive mechanisms and PID expert knowledge to enhance training convergence and stability. Target functions incorporating distance constraint penalties and entropy policies are designed in alignment with the characteristics of the given problem. We also devise innovative disturbance-adaptive structures and trajectory feature selection mechanisms to augment control error information and extract …


Multi-Uav Reconnaissance Mission Planning Via Deep Reinforcement Learning With Simulated Annealing, Mingfeng Fan, Huan Liu, Guohua Wu, Aldy Gunawan, Guillaume Sartoretti Mar 2025

Multi-Uav Reconnaissance Mission Planning Via Deep Reinforcement Learning With Simulated Annealing, Mingfeng Fan, Huan Liu, Guohua Wu, Aldy Gunawan, Guillaume Sartoretti

Research Collection School Of Computing and Information Systems

Unmanned aerial vehicles (UAVs) are widely used in reconnaissance missions due to their autonomy and flexibility. Efficient mission planning for multiple UAVs is crucial for tasks such as traffic monitoring and data collection. However, existing approaches to multi-UAV reconnaissance mission planning problem (MURMPP) often struggle with high computational demands, leading to suboptimal solutions. To overcome this challenge, we introduce a divide-and-conquer framework that splits the problem into two phases: target allocation and UAV routing, effectively reducing computational complexity. Specifically, we propose a hybrid method, SA-NNO-DRL, which combines the nearest neighbor optima-based deep reinforcement learning (NNO-DRL) approach with simulated annealing (SA). …


Decision Modeling And Solution Based On Game Adversarial Complex Systems, Jiachen Jiang, Zhengxuan Jia, Zhao Xu, Tingyu Lin, Pengpeng Zhao, Yiming Ou Jan 2025

Decision Modeling And Solution Based On Game Adversarial Complex Systems, Jiachen Jiang, Zhengxuan Jia, Zhao Xu, Tingyu Lin, Pengpeng Zhao, Yiming Ou

Journal of System Simulation

Abstract: In view of the complex situation of the current game which will be large-scale, high-intensity, not omniscient, and strong confrontation, and in response to the lack of flexibility and long iteration cycles in traditional game decision-making, the model of the unmanned complex game system is built according to the background of the unmanned red and blue game. Based on deep reinforcement learning technology, intelligent decision-making algorithms are studied in the background of unmanned red and blue games. With the help of deep neural networks and Bellman's optimal principle, the search of the huge solution space is more efficient, and …


Qos-Aware Link Adaptation For Beyond 5g Networks: A Deep Reinforcement Learning Approach, Ali Parsa, Neda Moghim, Sachin Shetty Jan 2025

Qos-Aware Link Adaptation For Beyond 5g Networks: A Deep Reinforcement Learning Approach, Ali Parsa, Neda Moghim, Sachin Shetty

Center for Secure and Intelligent Critical Systems (CSICS) Publications

Modern wireless communication systems face increasingly complex challenges due to rapidly changing channel conditions and the growing diversity of application-specific Quality of Service (QoS) requirements. Traditional link adaptation mechanisms primarily aim to maximize throughput and often lack the flexibility to support emerging applications, such as Extended Reality (XR) and Virtual Reality (VR), which demand simultaneous guarantees for high data rates, ultra low latency, and high reliability. These stringent and multidimensional QoS needs call for more intelligent and adaptive solutions. In this paper, we propose QDRLLA (QoS-aware Deep Reinforcement Learning-based Link Adaptation), a novel framework that employs deep reinforcement learning to …


Adversarial Generative Flow Network For Solving Vehicle Routing Problems, Ni Zhang, Jingfeng Yang, Zhiguang Cao, Xu Chi Jan 2025

Adversarial Generative Flow Network For Solving Vehicle Routing Problems, Ni Zhang, Jingfeng Yang, Zhiguang Cao, Xu Chi

Research Collection School Of Computing and Information Systems

Recent research into solving vehicle routing problems (VRPs) has gained significant traction, particularly through the application of deep (reinforcement) learning for end-to-end solution construction. However, many current construction-based neural solvers predominantly utilize Transformer architectures, which can face scalability challenges and struggle to produce diverse solutions. To address these limitations, we introduce a novel framework beyond Transformer-based approaches, i.e., Adversarial Generative Flow Networks (AGFN). This framework integrates the generative flow network (GFlowNet)-a probabilistic model inherently adept at generating diverse solutions (routes)-with a complementary model for discriminating (or evaluating) the solutions. These models are trained alternately in an adversarial manner to improve …


Path Planning Of Desert Robot Based On Deep Reinforcement Learning, Ming Li, Wangzhong Ye, Jiehua Yan Dec 2024

Path Planning Of Desert Robot Based On Deep Reinforcement Learning, Ming Li, Wangzhong Ye, Jiehua Yan

Journal of System Simulation

Abstract: Due to the complexity and variability of the desert environment, the key to the high-efficient of mobile robot is how to avoid obstacles and plan its path. To solve the problems of poor search efficiency and slow convergence of deep reinforcement learning algorithm in complex environment, an improved deep reinforcement learning path planning algorithm is proposed. The exploration factor is improved and dynamically adjusted according to the convergence degree of the algorithm, so that the exploration factor dynamically decreases with the increase of the understanding degree of the agent to the environment, thus speeding up the convergence speed of …


Intersections Of Living And Machine Agencies: Art-Based Models Of Adaptive Conversation With The More-Than-Human World, Carlos Castellanos Nov 2024

Intersections Of Living And Machine Agencies: Art-Based Models Of Adaptive Conversation With The More-Than-Human World, Carlos Castellanos

Tradition Innovations in Arts, Design, and Media Higher Education

Today’s AI systems are not built to have reciprocal interplay with their environments and thus they demonstrate little interest in emergence, adaptation or developing mutually productive relationships with the natural world. Is a different kind of AI possible? How can artists contribute to its development? In this essay, I will discuss ways in which the arts might help guide us towards a new kind of AI, built upon adaptive conversation (i.e. shared construction of meaning) with nature. I will discuss how can work with AI while also challenging its prevailing ontology, and even suggest alternative ontologies and epistemologies. I will …


A Method For Key Node Identification In Operational Target System Based On War Gaming, Yongfu Zhang, Yang Liu, He Yuan Nov 2024

A Method For Key Node Identification In Operational Target System Based On War Gaming, Yongfu Zhang, Yang Liu, He Yuan

Journal of System Simulation

Abstract: The identification of key nodes in an operational target system is an important basis for combat command decision-making. Due to the lack of experimental verification of key node identification in the current operational target system in a campaign-level dynamic confrontation environment, a complex network model of operational target system with large-scale entities and complex interaction relationship was constructed by taking integrated air defense network as an example, with the help of the data derived from the large joint war gaming; the characteristics of wargame data were considered, and the value characteristics of combat targets and network structure characteristics were …


Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun Jul 2024

Learning Topological Representations With Bidirectional Graph Attention Network For Solving Job Shop Scheduling Problem, Cong Zhang, Zhiguang Cao, Yaoxin Wu, Wen Song, Jing Sun

Research Collection School Of Computing and Information Systems

Existing learning-based methods for solving job shop scheduling problems (JSSP) usually use off-the-shelf GNN models tailored to undirected graphs and neglect the rich and meaningful topological structures of disjunctive graphs (DGs). This paper proposes the topology-aware bidirectional graph attention network (TBGAT), a novel GNN architecture based on the attention mechanism, to embed the DG for solving JSSP in a local search framework. Specifically, TBGAT embeds the DG from a forward and a backward view, respectively, where the messages are propagated by following the different topologies of the views and aggregated via graph attention. Then, we propose a novel operator based …


Collaborative Deep Reinforcement Learning For Solving Multi-Objective Vehicle Routing Problems, Yaoxin Wu, Mingfeng Fan, Zhiguang Cao, Ruobin Gao, Yaqing Hou, Guillaume Sartoretti May 2024

Collaborative Deep Reinforcement Learning For Solving Multi-Objective Vehicle Routing Problems, Yaoxin Wu, Mingfeng Fan, Zhiguang Cao, Ruobin Gao, Yaqing Hou, Guillaume Sartoretti

Research Collection School Of Computing and Information Systems

Existing deep reinforcement learning (DRL) methods for multi-objective vehicle routing problems (MOVRPs) typically decompose an MOVRP into subproblems with respective preferences and then train policies to solve corresponding subproblems. However, such a paradigm is still less effective in tackling the intricate interactions among subproblems, thus holding back the quality of the Pareto solutions. To counteract this limitation, we introduce a collaborative deep reinforcement learning method. We first propose a preference-based attention network (PAN) that allows the DRL agents to reason out solutions to subproblems in parallel, where a shared encoder learns the instance embedding and a decoder is tailored for …


Online Control Of Adaptive Large Neighborhood Search Using Deep Reinforcement Learning, Reijnen Reijnen, Yingqian Zhang, Hoong Chuin Lau, Zaharah Bukhsh May 2024

Online Control Of Adaptive Large Neighborhood Search Using Deep Reinforcement Learning, Reijnen Reijnen, Yingqian Zhang, Hoong Chuin Lau, Zaharah Bukhsh

Research Collection School Of Computing and Information Systems

The Adaptive Large Neighborhood Search (ALNS) algorithm has shown considerable success in solving combinatorial optimization problems (COPs). Nonetheless, the performance of ALNS relies on the proper configuration of its selection and acceptance parameters, which is known to be a complex and resource-intensive task. To address this, we introduce a Deep Reinforcement Learning (DRL) based approach called DR-ALNS that selects operators, adjusts parameters, and controls the acceptance criterion throughout the search. The proposed method aims to learn, based on the state of the search, to configure ALNS for the next iteration to yield more effective solutions for the given optimization problem. …


Intelligent Optimization Of Coal Terminal Unloading Scheduling Based On Improved D3qn Algorithm, Baoxin Qin, Yuxiao Zhang, Sirui Wu, Weichong Cao, Zhan Li Mar 2024

Intelligent Optimization Of Coal Terminal Unloading Scheduling Based On Improved D3qn Algorithm, Baoxin Qin, Yuxiao Zhang, Sirui Wu, Weichong Cao, Zhan Li

Journal of System Simulation

Abstract: Intelligent decision scheduling can improve the operation efficiency of large ports, which is one of the important research directions for the implementation of artificial intelligence technology in the smart port scenario. This article studies the intelligent unloading scheduling tasks of coal terminals and abstracts them as a Markov sequence decision problem. A deep reinforcement learning model for this problem is established, and an improved D3QN algorithm is proposed to realize intelligent optimization of unloading scheduling decisions by considering the characteristics of high action space dimension and sparse feasible action in the model. The simulation results show that for the …


Age Of Sensing Empowered Holographic Isac Framework For Nextg Wireless Networks: A Vae And Drl Approach, Apurba Adhikary, Avi Deb Raha, Yu Qiao, Md. Shirajum Munir, Monishanker Halder, Choong Seon Hong Jan 2024

Age Of Sensing Empowered Holographic Isac Framework For Nextg Wireless Networks: A Vae And Drl Approach, Apurba Adhikary, Avi Deb Raha, Yu Qiao, Md. Shirajum Munir, Monishanker Halder, Choong Seon Hong

School of Cybersecurity Faculty Publications

This paper proposes an artificial intelligence (AI) framework that leverages integrated sensing and communication (ISAC), aided by the age of sensing (AoS) to ensure the timely location updates of the users for a holographic MIMO (HMIMO)- enabled wireless network. The AI-driven framework guarantees optimal power allocation for efficient beamforming by activating the minimal number of grids from the HMIMO base station. An optimization problem is formulated to maximize the sensing utility function, aiming to maximize the signal-to-interference-plus-noise ratio (SINR) of the received signal, beam-pattern gains to improve the sensing SINR of reflected echo signals and maximizing the evidence lower bound …