Open Access. Powered by Scholars. Published by Universities.®

Systems Science Commons

Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 30 of 34

Full-Text Articles in Systems Science

Simulation Study Of Elevator Group Control Scheduling Based On Real-Time Occupancy Perception, Yiyong Han, Shiyu Wang, Yang Ye, Zhen Zhang, Fengque Pei, Minghai Yuan Jul 2026

Simulation Study Of Elevator Group Control Scheduling Based On Real-Time Occupancy Perception, Yiyong Han, Shiyu Wang, Yang Ye, Zhen Zhang, Fengque Pei, Minghai Yuan

Journal of System Simulation

Abstract: To address the conflict between car space allocation and peak passenger flow response efficiency in elevator group control scheduling, a multi-objective scheduling method based on proximal policy optimization (PPO) with real-time occupancy perception was proposed. A simulation environment considering car capacity constraints was constructed, and a reward-penalty mechanism with average passenger waiting time, system energy consumption, and car congestion as optimization objectives was designed. Based on this, state and action spaces were defined to form a PPO-based scheduling framework; a simulation platform integrating traffic flow visualization, policy scheduling, and performance evaluation was developed. Simulation results show that this method …


Research On Output Feedback Control Based On Reinforcement Learning For Overhead Crane, Minghui Li, Daoxiang Gao Jun 2026

Research On Output Feedback Control Based On Reinforcement Learning For Overhead Crane, Minghui Li, Daoxiang Gao

Journal of System Simulation

An output feedback control algorithm is designed based on reinforcement learning for the optimal control problem of overhead crane system. A high gain observer (HGO) is designed using output data to estimate the unmeasurable states of the overhead crane system. Based on the estimated states from the high-gain observer, a policy iteration (PI) method is designed with integral reinforcement learning, which uses Critic and Actor neural networks to approximate the optimal value function and control strategy, and adjusts the neural network weights in real time through online adaptive algorithms. According to the Lyapunov stability theory, the uniform ultimate boundedness of …


Improved Nsga-Ii For Dual-Resource Flexible Job Shop Scheduling Considering Worker Load, Guohui Zhang, Yuan Ren, Changjun Wu, Xiaofei Kou Jun 2026

Improved Nsga-Ii For Dual-Resource Flexible Job Shop Scheduling Considering Worker Load, Guohui Zhang, Yuan Ren, Changjun Wu, Xiaofei Kou

Journal of System Simulation

For the dual-resource-constrained flexible job shop scheduling problem considering worker load, an evolutionary algorithm integrating reinforcement learning was proposed. A three-stage encoding conforming to the problem characteristics was designed, and three initialization methods were combined to improve the population quality; a left-insertion decoding method based on worker load was designed to ensure that the completion time of the operation is less than the maximum processable time of the worker on the current day; two neighborhood structures based on the critical path were constructed to enhance the local exploration ability of the population; reinforcement learning was integrated to enable the …


Review On Optimization Of Simulation Modeling Strategies For Spacecraft Orbit Avoidance, Guozheng Li, Rui Wang, Shichao Fan, Xintong Cai, Xinyue Zhai Apr 2026

Review On Optimization Of Simulation Modeling Strategies For Spacecraft Orbit Avoidance, Guozheng Li, Rui Wang, Shichao Fan, Xintong Cai, Xinyue Zhai

Journal of System Simulation

Abstract: The number of on-orbit spacecraft increases exponentially; the space environment becomes more complex, and the collision risk of on-orbit spacecraft increases significantly. On-orbit safety is thus severely threatened, posing higher requirements for orbit avoidance methods. The costs and risks of space activities are extremely high, making simulation an effective method to solve complex problems of orbit avoidance. The modeling, solution, and simulation methods for the two core issues of spacecraft orbit avoidance, "collision avoidance" and "pursuit-evasion games", were systematically reviewed, and the existing shortcomings were analyzed. The applications of technologies such as deep reinforcement learning in promoting orbit avoidance …


Robot Path Planning By Reinforcement Learning Based On Sac3q-Hdm, Dequan Li, Wan Xiong Mar 2026

Robot Path Planning By Reinforcement Learning Based On Sac3q-Hdm, Dequan Li, Wan Xiong

Journal of System Simulation

Abstract: To address the issues of overestimated and underestimated biases, low sample utilization rate, and the inability to balance exploration and exploitation in reinforcement learning for path planning, an improved SAC method was proposed. The size balance of entropy was explored and utilized through adaptive temperature coefficient adjustment; on the basis of the SAC framework, a triple Critic architecture was introduced to dynamically weight and fuse the minimum and average values through Qvalue uncertainty, balancing overestimated and underestimated biases. A mixed dynamic sampling experience replay buffer was designed; experience data was partitioned based on reward thresholds; sampling ratios were dynamically …


Reinforcement Learning Based Method For Uav Team Orienteering Optimization Under Multi-Constraint Condition, Can Yang, Kai Chen, Feng Zhu Feb 2026

Reinforcement Learning Based Method For Uav Team Orienteering Optimization Under Multi-Constraint Condition, Can Yang, Kai Chen, Feng Zhu

Journal of System Simulation

Abstract: Traditional optimization methods struggle with efficiency, while reinforcement learning approaches often yield low solution quality and high training costs. In response, this paper proposes an attention mechanism-based reinforcement learning method. A dynamic attention strategy network with multi-information fusion is designed to improve solution quality. A visibility-graph approach is employed to simplify threat zone constraints and speed up convergence, and a decoding sequence reordering mechanism is introduced for further performance optimization of the solution. The simulation results show that the method generates high-quality solutions within milliseconds, achieving total rewards that approach or even surpass those obtained by traditional solvers …


Auv Path Planning Based On Behavior Cloning And Improved Dqn In Partially Unknown Environments, Lijing Xing, Min Li, Xiangguang Zeng, Ping Zhang, Bei Peng Nov 2025

Auv Path Planning Based On Behavior Cloning And Improved Dqn In Partially Unknown Environments, Lijing Xing, Min Li, Xiangguang Zeng, Ping Zhang, Bei Peng

Journal of System Simulation

Abstract: To address the problems of large randomness and slow convergence of the DQN dynamic path planning algorithm for a single autonomous underwater vehicle (AUV) in a partially unknown environment, a path planning method combining behavior cloning with A* algorithm and DQN (BA_DQN) was proposed. Based on the known environmental information, an improved A* algorithm incorporating ocean current resistance was proposed to guide DQN, thereby reducing the randomness of the DQN algorithm. By considering the complexity of the marine environment, the sampling probability was improved again after expanding the positive experience pool to enhance the training success rate. To address …


Modeling And Simulation Of Traffic Signal Control Based On Mlp With Improved Gcn-Td3, Deqi Huang, Yating Tu, Zhenhua Zhang, Xin Guo Oct 2025

Modeling And Simulation Of Traffic Signal Control Based On Mlp With Improved Gcn-Td3, Deqi Huang, Yating Tu, Zhenhua Zhang, Xin Guo

Journal of System Simulation

Abstract: To address the issues of uneven traffic flow at urban intersections, limited road capacity, and the poor coordination of existing traffic signal control algorithms, a traffic signal control algorithm based on graph convolutional reinforcement learning was proposed. By utilizing a multilayer perceptron, the dynamic features of vehicles and phase information at the controlled intersection and its neighboring intersections were extracted. A graph convolutional neural network was then employed to aggregate these vehicle dynamic features into potential features representing regional traffic. The control strategy was derived through multiple iterations of an improved twin delayed deep deterministic policy gradient (TD3) algorithm. …


Signal Timing Optimization Via Reinforcement Learning With Traffic Flow Prediction, Ming Xu, Jinye Li, Dongyu Zuo, Jing Zhang Apr 2025

Signal Timing Optimization Via Reinforcement Learning With Traffic Flow Prediction, Ming Xu, Jinye Li, Dongyu Zuo, Jing Zhang

Journal of System Simulation

Abstract: In response to the existing reinforcement learning-based traffic signal control methods that do not consider the changing trends in traffic flow, leading to congestion and inability to adapt to complex and variable road conditions, we propose a traffic signal timing optimization reinforcement learning method based on flow prediction. A phase timing amplitude control model is introduced. This model analyzes the spatiotemporal characteristics of historical traffic data to predict the flow for the next time slot and calculates a reasonable range for phase timing based on the prediction results. The H-PPO algorithm is employed to control the signal phase while …


Research On Digital Twin Simulation Method Of Industrial Robot Integrated With Reinforcement Learning, Tianyue Miao, Lu Wang, Jiaxiao He, Nenggang Xie Dec 2024

Research On Digital Twin Simulation Method Of Industrial Robot Integrated With Reinforcement Learning, Tianyue Miao, Lu Wang, Jiaxiao He, Nenggang Xie

Journal of System Simulation

Abstract: In response to the lack of comprehensive functionality and limited application scenarios in the current field of industrial robot digital twin systems, which results in low versatility, a method for constructing a digital twin system for industrial robots with high versatility is proposed. A four-dimensional system architecture for the digital twin is designed, and the components and functions of the four-dimensional system are analyzed, based on the system level planning of the four-dimensional system, the concept of integrating reinforcement learning into the virtual replacement of real concept is defined. By constructing a multi-attribute virtual model and using TCP communication …


Research On Scheduling Strategies Simulation For Building Air-Conditioning Systems Based On Transfer Imitation Learning, Qiaochu Wang, Yan Ding, Chuanzhi Liang, Haozheng Zhang, Chen Huang Dec 2024

Research On Scheduling Strategies Simulation For Building Air-Conditioning Systems Based On Transfer Imitation Learning, Qiaochu Wang, Yan Ding, Chuanzhi Liang, Haozheng Zhang, Chen Huang

Journal of System Simulation

Abstract: To solve the problem of unstable performance and inefficient training process of low-quality data conditions at the initial stage of online deployment of air conditioner scheduling, we propose a migration-imitation learning-based air conditioning scheduling strategy simulation method. Reinforcement learning methods are used to generate building operation strategies. A standard building simulation model serves as the source domain, upon which migration learning is applied. An imitation learning loss function is incorporated into the intelligent loss function to enhance algorithm performance. The results indicate that, compared with the non-use of migration learning, the proposed method can improve the operational efficiency by …


End-To-End Motion Planning Of Unmanned Vehicles Based On Multimodal Deep Reinforcement Learning, Kaiyuan Ding, Askar Hamdulla, Bin Zhu, Eksan Firkat, Zhengtang Ma Nov 2024

End-To-End Motion Planning Of Unmanned Vehicles Based On Multimodal Deep Reinforcement Learning, Kaiyuan Ding, Askar Hamdulla, Bin Zhu, Eksan Firkat, Zhengtang Ma

Journal of System Simulation

Abstract: Since the agent cannot sense the surrounding environment and cannot successfully avoid obstacles, reinforcement learning fails to be generalized to robot motion planning in difficult terrain. Therefore, a solution based on multimodal deep reinforcement learning, which learns to blend proprioceptive states with high-dimensional depth sensor inputs, is proposed for the motion planning of unmanned vehicles. To be specific, proprioceptive states offer contact measurement for immediate reaction, and the unmanned vehicle can learn and forecast environmental changes with its attached visual sensors, proactively navigating around obstacles and uneven terrains numerous time steps ahead. TransProAct (transformer-based proactive action), a unique end-to-end …


Research On Learnable Wargame Agent Driven By Battle Scheme, Yifeng Sun, Zhi Li, Jiang Wu, Yubin Wang Jul 2024

Research On Learnable Wargame Agent Driven By Battle Scheme, Yifeng Sun, Zhi Li, Jiang Wu, Yubin Wang

Journal of System Simulation

Abstract: To enable the agent to cope with complex battle scenarios and objectives in wargame, a learnable wargame agent architecture driven by a battle scheme is proposed. By analyzing the "attachment characteristics" and "loose coupling characteristics" of the agent to wargame system, the learnable requirements of the agent are obtained. In the design of the agent framework, battle schemes are used to reduce the learning range of the agent. The finite state machine corresponds to the knowledge of the operational phase in the battle scheme, and the decision-making space of the agent is determined according to the framework of the …


Task Analysis Methods Based On Deep Reinforcement Learning, Xue Gong, Pengfei Peng, Li Rong, Yalian Zheng, Jun Jiang Jul 2024

Task Analysis Methods Based On Deep Reinforcement Learning, Xue Gong, Pengfei Peng, Li Rong, Yalian Zheng, Jun Jiang

Journal of System Simulation

Abstract: In response to the high coupling of task interaction and many influencing factors in task analysis, a task analysis method based on sequence decoupling and deep reinforcement learning (DRL) is proposed, which can achieve task decomposition and task sequence reconstruction under complex constraints. The method designs an environment for deep reinforcement learning based on task information interaction, while improving the SumTree algorithm based on the difference between the loss functions of the target network and the evaluation network, achieving the priority evaluation among tasks. The activation function operation mechanism is introduced into the deep reinforcement learning network, followed by …


Real-Time Scheduling Method For Dynamic Flexible Job Shop Scheduling, Quan Jiang, Jingxuan Wei Jul 2024

Real-Time Scheduling Method For Dynamic Flexible Job Shop Scheduling, Quan Jiang, Jingxuan Wei

Journal of System Simulation

Abstract: A multi-objective dynamic flexible job shop scheduling problem model with machine breakdown and random jobs arrival is constructed to address the interference of dynamic events in manufacturing processing on the scheduling scheme, and a real-time scheduling method with multiobjective proximal policy optimization (MPPO) algorithm is proposed. The MPPO algorithm trains two agents, routing agent (RA) and sequencing agent (SA), for real-time scheduling and real-time processing of dynamic events. It employs a linear combination of weight vectors and reward vectors as reward signals and stores the agents' parameters for each weight vector to optimize multiple objectives. The required state information, …


Research And Development Of Simulation Training Platform For Multi-Agent Collaborative Decision-Making, Cheng Cheng, Zhijie Chen, Ziming Guo, Ni Li Dec 2023

Research And Development Of Simulation Training Platform For Multi-Agent Collaborative Decision-Making, Cheng Cheng, Zhijie Chen, Ziming Guo, Ni Li

Journal of System Simulation

Abstract: Reinforcement learning simulation platform can be an interactive and training environment for reinforcement learning. In order to make the simulation platform compatible with the multi-agent reinforcement learning algorithms and meet the needs of simulation in military field, the similar processes in multi-agent reinforcement learning algorithms are refined and a unified interface is designed to embed and verify different types of deep reinforcement learning algorithms on the simulation platform and to optimize the back-end service of the simulation platform to accelerate the training process of the algorithm model. The experimental results show that, by unifing the interface, the simulation platform …


Intercell Dynamic Scheduling Method Based On Deep Reinforcement Learning, Jing Ni, Mengke Ma Nov 2023

Intercell Dynamic Scheduling Method Based On Deep Reinforcement Learning, Jing Ni, Mengke Ma

Journal of System Simulation

Abstract: In order to solve the intercell scheduling problem of dynamic arrival of machining tasks and realize adaptive scheduling in the complex and changeable environment of the intelligent factory, a scheduling method based on a deep Q network is proposed. A complex network with cells as nodes and workpiece intercell machining path as directed edges is constructed, and the degree value is introduced to define the state space with intercell scheduling characteristics. A compound scheduling rule composed of a workpiece layer, unit layer, and machine layer is designed, and hierarchical optimization makes the scheduling scheme more global. Since double deep …


Imitative Generation Of Optimal Guidance Law Based On Reinforcement Learning, Zhengxuan Jia, Tingyu Lin, Yingying Xiao, Guoqiang Shi, Hao Wang, Bi Zeng, Yiming Ou, Pengpeng Zhao Nov 2023

Imitative Generation Of Optimal Guidance Law Based On Reinforcement Learning, Zhengxuan Jia, Tingyu Lin, Yingying Xiao, Guoqiang Shi, Hao Wang, Bi Zeng, Yiming Ou, Pengpeng Zhao

Journal of System Simulation

Abstract: Under the background of high-speed maneuvering target interception, an optimal guidance law generation method for head-on interception independent of target acceleration estimation is proposed based on deep reinforcement learning. In addition, its effectiveness is verified through simulation experiments. As the simulation results suggest, the proposed method successfully achieves head-on interception of high-speed maneuvering targets in 3D space and largely reduces the requirement for target estimation with strong uncertainty, and it is more applicable than the optimal control method.


Uav-Enabled Task Offloading Strategy For Vehicular Edge Computing Networks, Feng Hu, Haiyang Gu, Jun Lin Nov 2023

Uav-Enabled Task Offloading Strategy For Vehicular Edge Computing Networks, Feng Hu, Haiyang Gu, Jun Lin

Journal of System Simulation

Abstract: As intelligent vehicles are equipped with more and more sensors, the explosive growth of sensor data is generated, which brings severe challenges to vehicular communication and computing. In addition, the modern road presents a three-dimensional structure, and the system architecture of traditional vehicular networks cannot guarantee full coverage and seamless computing. A task offloading strategy for UAV-assisted and 6G-enabled (Sixth Generation) vehicular edge computing networks is proposed. Furthermore, a flexible and intelligent vehicular edge computing mode is composed by vehicles and UAVs, which provide three-dimensional edge computing services for delay-sensitive and computation-intensive vehicular tasks, and ensure timely processing and …


Aircraft Assignment Method For Optimal Utilization Of Maintenance Intervals, Runxia Guo, Yifu Wang Sep 2023

Aircraft Assignment Method For Optimal Utilization Of Maintenance Intervals, Runxia Guo, Yifu Wang

Journal of System Simulation

Abstract: The aircraft assignment problem is studied from a maintenance assurance perspective. In order to ensure its continuous airworthiness, civil aircraft are required to perform maintenance tasks, i. e., scheduled inspections, at specified intervals. The scheduled inspection interval is usually controlled by the number of flight cycles (FC), flight hours (FH), or flight days (FD), whichever comes first. In order to make balanced use of the inspection interval, an aircraft assignment model for a given fleet size is developed to optimize the maintenance interval utilization, and it is solved by a reinforcement learning algorithm to minimize the variance of the …


Multi-Agent Cooperative Combat Simulation In Naval Battlefield With Reinforcement Learning, Ding Shi, Xuefeng Yan, Lina Gong, Jingxuan Zhang, Donghai Guan, Mingqiang Wei Apr 2023

Multi-Agent Cooperative Combat Simulation In Naval Battlefield With Reinforcement Learning, Ding Shi, Xuefeng Yan, Lina Gong, Jingxuan Zhang, Donghai Guan, Mingqiang Wei

Journal of System Simulation

Abstract: Due to the rapidly-changed situations of future naval battlefields, it is urgent to realize the high-quality combat simulation in naval battlefields based on artificial intelligence to comprehensively optimize and improve the combat effectiveness of our army and defeat the enemy. The collaboration of combat units is the key point and how to realize the balanced decision-making among multiple agents is the first task. Based on decoupling priority experience replay mechanism and attention mechanism, a multi-agent reinforcement learning-based cooperative combat simulation (MARL-CCSA) network is proposed. Based on the expert experience, a multi-scale reward function is designed, on which a naval …


Research On Unmanned Swarm Combat System Adaptive Evolution Model Simulation, Zhiqiang Li, Yuanlong Li, Laixiang Yin, Xiangping Ma Apr 2023

Research On Unmanned Swarm Combat System Adaptive Evolution Model Simulation, Zhiqiang Li, Yuanlong Li, Laixiang Yin, Xiangping Ma

Journal of System Simulation

Abstract: Aiming at the fact that the intelligent unmanned swarm combat system is mainly composed of large-scale combat individuals with limited behavioral capabilities and has limited ability to adapt to the changes of battlefield environment and combat opponents, a learning evolution method combining genetic algorithm and reinforcement learning is proposed to construct an individual-based unmanned bee colony combat system evolution model. To improve the adaptive evolution efficiency of bee colony combat system, an improved genetic algorithm is proposed to improve the learning and evolution speed of bee colony individuals by using individual-specific mutation optimization strategy. Simulation experiment on …


Dqn-Based Joint Scheduling Method Of Heterogeneous Tt&C Resources, Naiyang Xue, Dan Ding, Yutong Jia, Zhiqiang Wang, Yuan Liu Feb 2023

Dqn-Based Joint Scheduling Method Of Heterogeneous Tt&C Resources, Naiyang Xue, Dan Ding, Yutong Jia, Zhiqiang Wang, Yuan Liu

Journal of System Simulation

Abstract: Joint scheduling of heterogeneous TT&C resources as research object, a deep Q network (DQN) algorithm based on reinforcement learning is proposed. The characteristics of the joint scheduling problem of heterogeneous TT&C resources being fully analyzied and mathematical language being used to describe the constraints affecting the solution, a resource joint scheduling model is established. From the perspective of applying reinforcement learning, two neural networks with the same structure and the action selection strategies based onεgreedy algorithm are respectively designed after Markov decision process description, and DQN solution framework is established. The simulation results show that DQN-based heterogeneous …


Reinforcement-Learning-Based Adaptive Tracking Control For A Space Continuum Robot Based On Reinforcement Learning, Da Jiang, Zhiqin Cai, Zhongzhen Liu, Haijun Peng, Zhigang Wu Oct 2022

Reinforcement-Learning-Based Adaptive Tracking Control For A Space Continuum Robot Based On Reinforcement Learning, Da Jiang, Zhiqin Cai, Zhongzhen Liu, Haijun Peng, Zhigang Wu

Journal of System Simulation

Abstract: Aiming at the tracking control for three-arm space continuum robot in space active debris removal manipulation, an adaptive sliding mode control algorithm based on deep reinforcement learning is proposed. Through BP network, a data-driven dynamic model is developed as the predictive model to guide the reinforcement learning to adjust the sliding mode controller's parameters online, and finally realize a real-time tracking control. Simulation results show that the proposed data-driven predictive model can accurately predict the robot's dynamic characteristics with the relative error within ±1% to random trajectories. Compared with the fixed-parameter sliding mode controller, the proposed adaptive controller …


Application Of Improved Q Learning Algorithm In Job Shop Scheduling Problem, Yejian Zhao, Yanhong Wang, Jun Zhang, Hongxia Yu, Zhongda Tian Jun 2022

Application Of Improved Q Learning Algorithm In Job Shop Scheduling Problem, Yejian Zhao, Yanhong Wang, Jun Zhang, Hongxia Yu, Zhongda Tian

Journal of System Simulation

Abstract: Aiming at the job shop scheduling in a dynamic environment, a dynamic scheduling algorithm based on an improved Q learning algorithm and dispatching rules is proposed. The state space of the dynamic scheduling algorithm is described with the concept of "the urgency of remaining tasks" and a reward function with the purpose of "the higher the slack, the higher the penalty" is disigned. In view of the problem that the greedy strategy will select the sub-optimal actions in the later stage of learning, the traditional Q learning algorithm is improved by introducing an action selection strategy based on the …


Research On The Construction Method Of Simulation Evaluation Index Of Operation Effectiveness Operation Concept Traction, Ziwei Zhang, Liang Li, Zhiming Dong, Yifei Wang, Li Duan Mar 2022

Research On The Construction Method Of Simulation Evaluation Index Of Operation Effectiveness Operation Concept Traction, Ziwei Zhang, Liang Li, Zhiming Dong, Yifei Wang, Li Duan

Journal of System Simulation

Abstract: Agents are difficult to be directly modeled and simulated due to the complexity of their own interaction and learning behaviors. Aiming at the common problems in the discrete simulation of the agent, the event transfer mechanism of the discrete event system specification (DEVS) atomic model is applied to express the interaction and learning of an agent. Through the interaction mode of the agent, the transfer control of multi-state external events, the port connection mode, as well as the introduction of reinforcement learning event transfer representation, a discrete simulation construction method of the agent based on the DEVS atomic model …


Dqn-Based Path Planning Method And Simulation For Submarine And Warship In Naval Battlefield, Xiaodong Huang, Haitao Yuan, Bi Jing, Liu Tao Oct 2021

Dqn-Based Path Planning Method And Simulation For Submarine And Warship In Naval Battlefield, Xiaodong Huang, Haitao Yuan, Bi Jing, Liu Tao

Journal of System Simulation

Abstract: To realize multi-agent intelligent planning and target tracking in complex naval battlefield environment, the work focuses on agents (submarine or warship), and proposes a simulation method based on reinforcement learning algorithm called Deep Q Network (DQN). Two neural networks with the same structure and different parameters are designed to update real and predicted Q values for the convergence of value functions. An ε-greedy algorithm is proposed to design an action selection mechanism, and a reward function is designed for the naval battlefield environment to increase the update velocity and generalization ability of Learning with Experience Replay (LER). Simulation results …


Research On Experimental Method Of Joint Operation Simulation Based On Human-Machine Hybrid Intelligence, Ma Jun, Jingyu Yang, Wu Xi Oct 2021

Research On Experimental Method Of Joint Operation Simulation Based On Human-Machine Hybrid Intelligence, Ma Jun, Jingyu Yang, Wu Xi

Journal of System Simulation

Abstract: In view of the difficulties that the joint operation simulation experiment methods are mainly for guiding equipment evaluation and demonstration, which is difficult to effectively support the research of operation problems, a joint operation simulation experiment method based on human-machine hybrid intelligence is proposed. The classification, generation and accumulation process of the knowledge in joint operation simulation experiment are clarified. Through the detailed descriptions of experimental interaction process, experimental operation process, experimental driving mode, simulation operation mode, supporting system structure, etc., a joint operation simulation experiment framework based on man-machine hybrid intelligence is constructed. It provides a new method …


Study On Next-Generation Strategic Wargame System, Wu Xi, Xianglin Meng, Jingyu Yang Sep 2021

Study On Next-Generation Strategic Wargame System, Wu Xi, Xianglin Meng, Jingyu Yang

Journal of System Simulation

Abstract: Strategic wargame is an important support to the strategic decision. The research status and challenges of the strategic wargame are analyzed, and the influence of big data and artificial intelligence technology on the strategic wargame system is studied. The prospects and key technologies of the next-generation strategic wargame system are studied, including the construction of event association graph for strategic topics, generation of strategic decision sparse samples based on generative adversarial nets, gaming strategy learning of human-in-loop hybrid enhancement, and public opinion dissemination modeling technology based on social network. The development trend of the strategic wargame is proposed.


Self-Learning-Based Multiple Spacecraft Evasion Decision Making Simulation Under Sparse Reward Condition, Zhao Yu, Jifeng Guo, Yan Peng, Chengchao Bai Aug 2021

Self-Learning-Based Multiple Spacecraft Evasion Decision Making Simulation Under Sparse Reward Condition, Zhao Yu, Jifeng Guo, Yan Peng, Chengchao Bai

Journal of System Simulation

Abstract: In order to improve the ability of spacecraft formation to evade multiple interceptors, aiming at the low success rate of traditional procedural maneuver evasion, a multi-agent cooperative autonomous decision-making algorithm, which is based on deep reinforcement learning method, is proposed. Based on the actor-critic architecture, a multi-agent reinforcement learning algorithm is designed, in which a weighted linear fitting method is proposed to solve the reliability allocation problem of the self-learning system. To solve the sparse reward problem in task scenario, a sparse reward reinforcement learning method based on inverse value method is proposed. According to the task scenario, …