Open Access. Powered by Scholars. Published by Universities.®

Systems Science Commons

Open Access. Powered by Scholars. Published by Universities.®

DRL

Publication Year

Articles 1 - 19 of 19

Full-Text Articles in Systems Science

Trajectory-Aware Dynamic Offloading For High-Mobility Vehicular Edge Systems, Haoran Xu, Zhiqin Huang, Youwu Hu, Zheyi Chen Jul 2026

Trajectory-Aware Dynamic Offloading For High-Mobility Vehicular Edge Systems, Haoran Xu, Zhiqin Huang, Youwu Hu, Zheyi Chen

Journal of System Simulation

Abstract: To address the problems of service interruptions, task failures, and resource waste caused by high mobility of intelligent vehicles (IVs) in vehicular edge computing (VEC), this paper proposes a trajectory-aware dynamic offloading (TADO) framework for high-mobility vehicular edge systems. A lightweight T-pattern trajectory-aware algorithm is designed to efficiently predict the next-hop road side unit (RSU) by mining spatio-temporal patterns from historical trajectories of vehicles, offering a forward-looking reference for offloading decisions. A joint optimization model is constructed, and a service interruption risk factor driven by trajectory prediction is introduced. An improved DRL method is developed. It takes the predicted …


Resource-Efficient Continuous Learning Framework For Edge Real-Time Video Analytics, Shuxia Wu, Junjie Zhang, Delong Chen, Zheyi Chen Feb 2026

Resource-Efficient Continuous Learning Framework For Edge Real-Time Video Analytics, Shuxia Wu, Junjie Zhang, Delong Chen, Zheyi Chen

Journal of System Simulation

Abstract: By deploying lightweight models at the network edge, edge systems can provide services of real-time video analytics. However, due to the data drift caused by the discrepancy between model training and actual deployment, it is challenging to construct lightweight models that match real-world environments. To address this challenge, a resource-efficient continuous learning framework for edge real-time video analytics (CL4VA) was proposed. A region of interest-granularity predictor for accuracy degradation was introduced to efficiently select key samples from real-time video streams. A two-layer mixed sample pool was constructed to adaptively trigger the model's continuous learning and avoid the issue of …


Strike Strategy Planning Method Of Unmanned Ground Vehicles Based On Improved Ppo Algorithm, Bingkun Wang, Yue Wang, Mei Yang, Pengnian Zhang, Bohao Fan, Jie Tang Feb 2026

Strike Strategy Planning Method Of Unmanned Ground Vehicles Based On Improved Ppo Algorithm, Bingkun Wang, Yue Wang, Mei Yang, Pengnian Zhang, Bohao Fan, Jie Tang

Journal of System Simulation

Abstract: An improved PPO algorithm based on the hybrid action space and gated recurrent unit (GRU) is proposed to address the limitations of predefined strike rules in maximizing the hitting accuracy of unmanned ground vehicles and the difficult coupling and optimization of continuous motion planning and discrete strike decision-making. The environmental model and target model are built for the process of unmanned ground vehicles' strike missions, coupled with a three-layer model for unmanned ground vehicles that fuses kinematic constraints, situational awareness, and dynamic decision-making. Two distinct policy networks are employed, including the continuous motion planning network for path planning, and …


Knowledge Closed-Loop Driving-Based Intelligent Game Confrontation Simulation, Quan Liu, Yu Wang, Linyue Liu, Hao Chen, Jian Huang Feb 2026

Knowledge Closed-Loop Driving-Based Intelligent Game Confrontation Simulation, Quan Liu, Yu Wang, Linyue Liu, Hao Chen, Jian Huang

Journal of System Simulation

Abstract: For human-machine intelligence integration and collaborative intelligence enhancement, a “knowledge-model-data-knowledge” closed-loop paradigm for combat simulation is proposed to guide the design of a DRL-based game confrontation simulation architecture. By building a combat priori knowledge-guided DRL agent model, mining and analyzing the time series data of agent interactions generated during simulations, and extracting combat posterior knowledge that expands the cognition boundaries of commanders, the knowledge closed-loop driving mechanism for intelligent combat simulations is achieved. The experimental results indicate that the proposed mechanism can effectively endow the combat simulation system with intelligence growth capabilities, providing valuable reference for the deepening …


A Drl⁃Based Approach For Distributed Equipment Nodes Selection, Ziyi Wang, Kai Zhang, Dianwei Qian, Yuzhen Liu Jun 2025

A Drl⁃Based Approach For Distributed Equipment Nodes Selection, Ziyi Wang, Kai Zhang, Dianwei Qian, Yuzhen Liu

Journal of System Simulation

Abstract: Aiming at the problem of insufficient solution speed and poor generalization of traditional algorithms in large-scale scenarios, this paper intelligently solves the large-scale distributed equipment system preference problem based on deep reinforcement learning. According to the characteristics of distributed equipment system combat, using the complex network to its graph form modeling, and based on the attention mechanism to the equipment between the connecting edge relationship for the characterization, in order to build a distributed equipment system digital simulation environment. Simulation results show that compared with the genetic evolutionary algorithm, the obtained model has obvious advantages in terms of solution …


Trajectory Planning Of Quadruped Robot Over Obstacle With Single Leg Based On Deep Reinforcement Learning, Min Li, Sen Zhang, Xiangguang Zeng, Gang Wang, Tongwei Zhang, Dijie Xie, Wenzhe Ren, Tao Zhang Apr 2025

Trajectory Planning Of Quadruped Robot Over Obstacle With Single Leg Based On Deep Reinforcement Learning, Min Li, Sen Zhang, Xiangguang Zeng, Gang Wang, Tongwei Zhang, Dijie Xie, Wenzhe Ren, Tao Zhang

Journal of System Simulation

Abstract: Aiming at the problems of joint vibration and high energy consumption of quadruped robot in the process of walking over obstacles, a foot trajectory planning method of quadruped robot based on deep reinforcement learning SAC algorithm is proposed. Based on robot kinematics and Monte Carlo method, the motion space of the single-legged foot of quadruped robot is analyzed. A compound seventhdegree polynomial trajectory of the quadruped robot is planned. The SAC algorithm is used to train and obtain the low energy consumption obstacle crossing strategy of four-legged robot under different obstacle environment. The simulation results show that the compound …


Uav Path Planning Based On Improved Deep Deterministic Policy Gradients, Sen Zhang, Qiangqiang Dai Apr 2025

Uav Path Planning Based On Improved Deep Deterministic Policy Gradients, Sen Zhang, Qiangqiang Dai

Journal of System Simulation

Abstract: Aiming at the problems of poor convergence and invalid exploration when UAVs perform path planning in complex environments, an improved deep deterministic policy gradient(DDPG) algorithm is proposed. Using a dual experience pooling mechanism to store success and failure experiences separately, the algorithm is able to use the success experience to strengthen the strategy optimization and learn from the failure experience to avoid the wrong path; an APF method is introduced to add a bootstrap term to the planning, which is combined with the exploration of noisy actions in a randomized sampling process to dynamically integrate the selected actions; multi-objective …


Research On Pedestrian Avoidance Strategy For Agv Based On Deep Reinforcement Learning, He Wang, Jianing Xu, Guangyu Yan Mar 2025

Research On Pedestrian Avoidance Strategy For Agv Based On Deep Reinforcement Learning, He Wang, Jianing Xu, Guangyu Yan

Journal of System Simulation

Abstract: To ensure the safety and comfort of pedestrians during Automated Guided Vehicle (AGV) obstacle avoidance in smart factory environments, a deep reinforcement learning-based end-to-end obstacle avoidance method is proposed. The YOLOv8 module is introduced to extract pedestrian pose information, and a visual-based state space is designed. A reinforcement learning mechanism is formulated based on personal space theory, penalizing AGV behaviors such as entering pedestrian comfort space and collisions. A virtual simulation system is constructed, utilizing PPO algorithm along with LSTM network layer for obstacle avoidance strategy training and simulation experiments. Simulation results indicate that this obstacle avoidance strategy, under …


Reinforcement Learning Modeling Of Missile Penetration Decision Based On Combat Simulation, Bin Zhang, Yonglin Lei, Qun Li, Yuan Gao, Yong Chen, Jiajun Zhu, Chenlong Bao Mar 2025

Reinforcement Learning Modeling Of Missile Penetration Decision Based On Combat Simulation, Bin Zhang, Yonglin Lei, Qun Li, Yuan Gao, Yong Chen, Jiajun Zhu, Chenlong Bao

Journal of System Simulation

Abstract: Penetration capability is a primary measure of missile systems. In response to the shortcomings of traditional knowledge-based decision-making methods that are difficult to adaptively evolve, an intelligent penetration decision-making based on combat simulation and DRL is proposed. A missile intelligent decision-making training environment is constructed based on the WESS system. Taking missile maneuver penetration decision-making as an example, a maneuver penetration decisionmaking network model is designed and trained based on the SAC-discrete algorithm and the test of intelligence is conducted. Experimental results show that the intelligent decision model derived from machine learning has a better combat outcome than traditional …


Intelligent Service Migration Towards Mec-Based Iov Systems, Sijin Huang, Jia Wen, Zheyi Chen Feb 2025

Intelligent Service Migration Towards Mec-Based Iov Systems, Sijin Huang, Jia Wen, Zheyi Chen

Journal of System Simulation

Abstract: To address the problem of QoS degradation during the vehicle movement, a novel service migration via convex-optimization-enabled deep reinforcement learning (SeMiR) method is proposed. The optimization problem is decomposed into two sub-problems and solved separately. For the service migration sub-problem, an improved deep reinforcement learning based service migration method is designed to explore the optimal migration policy. For the resource allocation sub-problem, a convex optimization based resource allocation method is developed to derive the optimal resource allocation for each MEC server under the given migration decisions, thereby improving the performance of service migration. Experimental results show that the SeMiR …


Research On The Target Allocation Method For Air Defense And Anti-Missile Defense Of Naval Ships, Shuaidi Fei, Changlong Cai, Fei Liu, Minghui Chen, Xiaoming Liu Feb 2025

Research On The Target Allocation Method For Air Defense And Anti-Missile Defense Of Naval Ships, Shuaidi Fei, Changlong Cai, Fei Liu, Minghui Chen, Xiaoming Liu

Journal of System Simulation

Abstract: To solve the problems of multiple types of state information and correlation of time-series state information encountered in the dynamic weapon target assignment problem, a dynamic weapon target assignment method based on an improved deep reinforcement learning algorithm is proposed. A multiinput assignment model of target missile-interceptor unit, interceptor unit, and defense unit under multiwave target and multi-phase is constructed. A multi-input state space is designed, and a Markov decision process is established in conjunction with the problem model. A feature extraction network combining multi-input information processing and gated recurrent network is designed, which improves the ability to extract …


Edge Surveillance Task Offloading And Resource Allocation Algorithm Based On Drl, Chao Li, Jiabao Li, Caichang Ding, Zhiwei Ye, Fangwei Zuo Sep 2024

Edge Surveillance Task Offloading And Resource Allocation Algorithm Based On Drl, Chao Li, Jiabao Li, Caichang Ding, Zhiwei Ye, Fangwei Zuo

Journal of System Simulation

Abstract: For the resource limitation of intensive surveillance tasks in edge computing, a surveillance task offloading and resource allocation algorithm based on DRL is proposed. With the optimization objectives of surveillance task delay and recognition accuracy, the joint decision objective optimization solution of task offloading, wireless channel allocation, and image compression rate was modeled as a Markov decision process. To address the problem of slow and unstable algorithm convergence due to the high volatility of training samples caused by the dynamic nature of wireless channels and the randomness of surveillance tasks, an attention mechanism is used to jointly encode channel …


Construction Of A Virtual Interactive System For Orchards Based On Digital Twin, Hongjun Wang, Junqiang Lin, Xiangjun Zou, Po Zhang, Mingxuan Zhou, Weirui Zou, Yunchao Tang, Lufeng Luo Jun 2024

Construction Of A Virtual Interactive System For Orchards Based On Digital Twin, Hongjun Wang, Junqiang Lin, Xiangjun Zou, Po Zhang, Mingxuan Zhou, Weirui Zou, Yunchao Tang, Lufeng Luo

Journal of System Simulation

Abstract: Aiming at the low visibility, poor real-time, weak adaptability and single interaction mode in orchard planting management system, a six-dimensional model of orchard digital twin system for planting management process is proposed. The system model construction theory and technology system is discussed from four aspects, entity modeling of management elements, dynamic modeling of management process, simulation modeling of management system and optimization modeling of management strategy. Based on the six-dimensional model, supported by the theory and technology system, the virtual interactive system architecture of the orchard based on the digital twin is designed, and the key technologies of the …


Simulation Of Robotic Peg-In-Hole Assembly Strategy Based On Drl, Zilu Zhu, Yongkui Liu, Lin Zhang, Lihui Wang, Tingyu Lin Jun 2024

Simulation Of Robotic Peg-In-Hole Assembly Strategy Based On Drl, Zilu Zhu, Yongkui Liu, Lin Zhang, Lihui Wang, Tingyu Lin

Journal of System Simulation

Abstract: Aiming at the existing peg-in-hole assembly method problems of dependence on accurate contact state models, difficulties in data acquisition, low sampling efficiency, and poor security, a simulation research method for robot peg-in-hole assembly strategy based on DRL is proposed. A simulation environment of robot peg-in-hole assembly based on ROS-Gazebo is built, and a method of gravity compensation for force/torque sensor based on a least square method is proposed. The reinforcement learning paradigm is employed to model the robot peg-in-hole assembly, and a method based on soft actor-critic(SAC) algorithm is proposed. The communication mechanism between the simulation environment and the …


Gradient-Based Deep Reinforcement Learning Interpretation Methods, Yuan Wang, Lin Xu, Xiaoze Gong, Yongliang Zhang, Yongli Wang May 2024

Gradient-Based Deep Reinforcement Learning Interpretation Methods, Yuan Wang, Lin Xu, Xiaoze Gong, Yongliang Zhang, Yongli Wang

Journal of System Simulation

Abstract: The learning process and working mechanism of deep reinforcement learning methods such as DQN are not transparent, and their decision basis and reliability cannot be perceived, which makes the decisions made by the model highly questionable and greatly limits the application scenarios of deep reinforcement learning. To explain the decision-making mechanism of intelligent agents, this paper proposes a gradient based saliency map generation algorithm SMGG. It uses the gradient information of feature maps generated by high-level convolutional layers to calculate the importance of different feature maps. With the known structure and internal parameters of the model, starting from the …


Efficiency Optimization Method For Data Sampling In Power Grid Topology Scheduling Simulation, Yingying Zhao, Pusen Dong, Tianchen Zhu, Fan Li, Yun Su, Zhenying Tai, Qingyun Sun, Hang Fan Feb 2024

Efficiency Optimization Method For Data Sampling In Power Grid Topology Scheduling Simulation, Yingying Zhao, Pusen Dong, Tianchen Zhu, Fan Li, Yun Su, Zhenying Tai, Qingyun Sun, Hang Fan

Journal of System Simulation

Abstract: To address the large simulation computational workload and low simulation speed caused by the scale and complexity of the new power system, a simulation acceleration method for topology scheduling based on the distributed and quantization mechanisms is proposed. The parallelization of topology scheduling models is used to increase the scale of data simulation sampling in unit time. The introduced quantization operators accelerate the computation speed of the topology scheduling model operators, reduces the time cost of the every single simulation. Case studies confirm the effectiveness of the topology simulation acceleration, in which the available transfer capacity of the simulated …


Research On Motion Planning Of Hexapod Robot Based On Drl And Free Gait, Xinpeng Wang, Huiqiao Fu, Guizhou Deng, Kaiqiang Tang, Chunlin Chen, Canghao Liu Feb 2024

Research On Motion Planning Of Hexapod Robot Based On Drl And Free Gait, Xinpeng Wang, Huiqiao Fu, Guizhou Deng, Kaiqiang Tang, Chunlin Chen, Canghao Liu

Journal of System Simulation

Abstract: To improve the passability and the motion performance of the hexapod robot in the unstructured environment, a multi-contact motion planning algorithm based on DRL and free gait planner is proposed. Firstly, the free gait planner obtains the reachable footholds under the target state and outputs the optimal gait sequence. The center of mass motion policy of the hexapod robot in the randomly generated plum blossom pile environment is obtained by using deep reinforcement learning training. To ensure the reachability between adjacent states of the robot in motion, the state transition feasibility model is used to judge the state transition …


Flipper Control Method For Tracked Robot Based On Deep Reinforcement Learning, Hainan Pan, Bailiang Chen, Kaihong Huang, Junkai Ren, Chuang Cheng, Huimin Lu, Hui Zhang Feb 2024

Flipper Control Method For Tracked Robot Based On Deep Reinforcement Learning, Hainan Pan, Bailiang Chen, Kaihong Huang, Junkai Ren, Chuang Cheng, Huimin Lu, Hui Zhang

Journal of System Simulation

Abstract: Tracked robots with flippers have certain terrain adaptation capabilities. To improve the intelligent operation level of robots in complex environments, it is significant to realize the flipper autonomously control. Combining the expert experience in obstacle crossing and optimization indicators, Markov decision process(MDP) modeling of the robot's flipper control problem is carried out and a simulation training environment based on physics simulation engine Pymunk is built. A deep reinforcement learning control algorithm based on dueling double DQN(D3QN) network is proposed for controlling the flippers. With terrain information and robot state as the input and the four flippers' angle as the …


Research On Multi-Aircraft Air Combat Behavior Modeling Based On Hierarchical Intelligent Modeling Methods, Yukun Wang, Ze Wang, Liwei Dong, Ni Li Oct 2023

Research On Multi-Aircraft Air Combat Behavior Modeling Based On Hierarchical Intelligent Modeling Methods, Yukun Wang, Ze Wang, Liwei Dong, Ni Li

Journal of System Simulation

Abstract: In response to the problem of the difficulty of decision-making in the game of force under the constraints of high-dimensional state-space in multi-machine air combat confrontation scenarios, a force intelligent agent decision-making generation strategy based on deep reinforcement learning is adopted. The developing situational cognition and reward feedback generation algorithms for force intelligentgame are proposed, a behavior modeling hierarchical framework based on hybrid intelligence modeling method is constructed, which solve the technical difficulty of sparse reward in the reinforcement learning process. It provides an feasible reinforcement learning training method that can solve the large-scale, multi-model, and multi-element air combat …