Open Access. Powered by Scholars. Published by Universities.®

Reinforcement learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 31 - 60 of 140

Full-Text Articles in Artificial Intelligence and Robotics

Efficient Task Scheduling In Cloud Infrastructures Using Dynamic Score-Based Allocation And Deep Q-Learning, Shadman Sakib Jan 2025

Efficient Task Scheduling In Cloud Infrastructures Using Dynamic Score-Based Allocation And Deep Q-Learning, Shadman Sakib

Graduate Theses/Dissertations

Cloud computing has grown rapidly in recent years, mainly due to the sharp increase in data transferred over the internet. This growth makes task scheduling a key and challenging part of cloud systems, as it helps distribute user requests across servers to minimize response time, prevent overloading, and ensure smooth user experience. This thesis proposes two novel approaches for dynamic task scheduling in cloud environments. First, a novel Score-Based Dynamic Load Balancing (SBDLB) strategy is developed, which leverages system parameters to allocate tasks efficiently across virtual machines (VMs) in data centers. SBDLB ensures balanced workload distribution by continuously evaluating VM …


Harnessing The Power Of Gradient-Based Simulations For Multi-Objective Optimization In Particle Accelerators, Kishansingh Rajput, Malachi Schram, Auralee Edelen, Jonathan Colen, Armen Kasparian, Ryan Roussel, Adam Carpenter, He Zhang, Jay Benesch Jan 2025

Harnessing The Power Of Gradient-Based Simulations For Multi-Objective Optimization In Particle Accelerators, Kishansingh Rajput, Malachi Schram, Auralee Edelen, Jonathan Colen, Armen Kasparian, Ryan Roussel, Adam Carpenter, He Zhang, Jay Benesch

Data Science Faculty Publications

Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The …


Research On Digital Twin Simulation Method Of Industrial Robot Integrated With Reinforcement Learning, Tianyue Miao, Lu Wang, Jiaxiao He, Nenggang Xie Dec 2024

Research On Digital Twin Simulation Method Of Industrial Robot Integrated With Reinforcement Learning, Tianyue Miao, Lu Wang, Jiaxiao He, Nenggang Xie

Journal of System Simulation

Abstract: In response to the lack of comprehensive functionality and limited application scenarios in the current field of industrial robot digital twin systems, which results in low versatility, a method for constructing a digital twin system for industrial robots with high versatility is proposed. A four-dimensional system architecture for the digital twin is designed, and the components and functions of the four-dimensional system are analyzed, based on the system level planning of the four-dimensional system, the concept of integrating reinforcement learning into the virtual replacement of real concept is defined. By constructing a multi-attribute virtual model and using TCP communication …


Research On Scheduling Strategies Simulation For Building Air-Conditioning Systems Based On Transfer Imitation Learning, Qiaochu Wang, Yan Ding, Chuanzhi Liang, Haozheng Zhang, Chen Huang Dec 2024

Research On Scheduling Strategies Simulation For Building Air-Conditioning Systems Based On Transfer Imitation Learning, Qiaochu Wang, Yan Ding, Chuanzhi Liang, Haozheng Zhang, Chen Huang

Journal of System Simulation

Abstract: To solve the problem of unstable performance and inefficient training process of low-quality data conditions at the initial stage of online deployment of air conditioner scheduling, we propose a migration-imitation learning-based air conditioning scheduling strategy simulation method. Reinforcement learning methods are used to generate building operation strategies. A standard building simulation model serves as the source domain, upon which migration learning is applied. An imitation learning loss function is incorporated into the intelligent loss function to enhance algorithm performance. The results indicate that, compared with the non-use of migration learning, the proposed method can improve the operational efficiency by …


Q-Learning In Starclash, Hanani Pankaj Dec 2024

Q-Learning In Starclash, Hanani Pankaj

2024 Fall Honors Capstone Projects - Archive

Developers create video games using Artificial Intelligence (AI) agents to provide a challenging opponent in a single-player game. However, studies show that when Reinforcement Learning (RL) agents are used, they outperform the AI agents. This project sought to test how RL agents would perform in StarClash, a video game without RL agents, using Q-Learning. This was done by creating two Q-Learning agents: a Simple agent and an Advanced (more complex) agent. These two agents were tested against each other and a Random AI agent. As expected, the Advanced agent did better than the Simple agent but only performed slightly better, …


Sprinql : Sub-Optimal Demonstrations Driven Offline Imitation Learning, Minh Huy Hoang, Tien Mai, Pradeep Varakantham Dec 2024

Sprinql : Sub-Optimal Demonstrations Driven Offline Imitation Learning, Minh Huy Hoang, Tien Mai, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

We focus on offline imitation learning (IL), which aims to mimic an expert's behavior using demonstrations without any interaction with the environment. One of the main challenges in offline IL is the limited support of expert demonstrations, which typically cover only a small fraction of the state-action space. While it may not be feasible to obtain numerous expert demonstrations, it is often possible to gather a larger set of sub-optimal demonstrations. For example, in treatment optimization problems, there are varying levels of doctor treatments available for different chronic conditions. These range from treatment specialists and experienced general practitioners to less …


Reinforcement Learning Based Online Request Scheduling Framework For Workload-Adaptive Edge Deep Learning Inference, Xinrui Tan, Hongjia Li, Xiaofei Xie, Lu Guo, Nirwan Ansari, Xueqing Huang, Liming Wang, Zhen Xu, Yang Liu Dec 2024

Reinforcement Learning Based Online Request Scheduling Framework For Workload-Adaptive Edge Deep Learning Inference, Xinrui Tan, Hongjia Li, Xiaofei Xie, Lu Guo, Nirwan Ansari, Xueqing Huang, Liming Wang, Zhen Xu, Yang Liu

Research Collection School Of Computing and Information Systems

The recent advances of deep learning in various mobile and Internet-of-Things applications, coupled with the emergence of edge computing, have led to a strong trend of performing deep learning inference on the edge servers located physically close to the end devices. This trend presents the challenge of how to meet the quality-of-service requirements of inference tasks at the resource-constrained network edge, especially under variable or even bursty inference workloads. Solutions to this challenge have not yet been reported in the related literature. In the present paper, we tackle this challenge by means of workload-adaptive inference request scheduling: in different workload …


End-To-End Motion Planning Of Unmanned Vehicles Based On Multimodal Deep Reinforcement Learning, Kaiyuan Ding, Askar Hamdulla, Bin Zhu, Eksan Firkat, Zhengtang Ma Nov 2024

End-To-End Motion Planning Of Unmanned Vehicles Based On Multimodal Deep Reinforcement Learning, Kaiyuan Ding, Askar Hamdulla, Bin Zhu, Eksan Firkat, Zhengtang Ma

Journal of System Simulation

Abstract: Since the agent cannot sense the surrounding environment and cannot successfully avoid obstacles, reinforcement learning fails to be generalized to robot motion planning in difficult terrain. Therefore, a solution based on multimodal deep reinforcement learning, which learns to blend proprioceptive states with high-dimensional depth sensor inputs, is proposed for the motion planning of unmanned vehicles. To be specific, proprioceptive states offer contact measurement for immediate reaction, and the unmanned vehicle can learn and forecast environmental changes with its attached visual sensors, proactively navigating around obstacles and uneven terrains numerous time steps ahead. TransProAct (transformer-based proactive action), a unique end-to-end …


Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang Aug 2024

Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang

Research Collection School Of Computing and Information Systems

High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become another appealing approach for HFT due to its terrific ability of handling high-dimensional financial data and solving sophisticated sequential decision-making problems, e.g., hierarchical reinforcement learning (HRL) has shown its promising performance on second-level HFT by training a router to select only one sub-agent from the agent pool to execute the current transaction. However, existing RL methods for HFT still have some defects: 1) standard RL-based trading agents suffer from the overfitting …


Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An Aug 2024

Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Nash Equilibrium (NE) is the canonical solution concept of game theory, which provides an elegant tool to understand the rationalities. Though mixed strategy NE exists in any game with finite players and actions, computing NE in two- or multi-player general-sum games is PPAD-Complete. Various alternative solutions, e.g., Correlated Equilibrium (CE), and learning methods, e.g., fictitious play (FP), are proposed to approximate NE. For convenience, we call these methods as "inexact solvers", or "solvers" for short. However, the alternative solutions differ from NE and the learning methods generally fail to converge to NE. Therefore, in this work, we propose REinforcement Nash …


Research On Learnable Wargame Agent Driven By Battle Scheme, Yifeng Sun, Zhi Li, Jiang Wu, Yubin Wang Jul 2024

Research On Learnable Wargame Agent Driven By Battle Scheme, Yifeng Sun, Zhi Li, Jiang Wu, Yubin Wang

Journal of System Simulation

Abstract: To enable the agent to cope with complex battle scenarios and objectives in wargame, a learnable wargame agent architecture driven by a battle scheme is proposed. By analyzing the "attachment characteristics" and "loose coupling characteristics" of the agent to wargame system, the learnable requirements of the agent are obtained. In the design of the agent framework, battle schemes are used to reduce the learning range of the agent. The finite state machine corresponds to the knowledge of the operational phase in the battle scheme, and the decision-making space of the agent is determined according to the framework of the …


Task Analysis Methods Based On Deep Reinforcement Learning, Xue Gong, Pengfei Peng, Li Rong, Yalian Zheng, Jun Jiang Jul 2024

Task Analysis Methods Based On Deep Reinforcement Learning, Xue Gong, Pengfei Peng, Li Rong, Yalian Zheng, Jun Jiang

Journal of System Simulation

Abstract: In response to the high coupling of task interaction and many influencing factors in task analysis, a task analysis method based on sequence decoupling and deep reinforcement learning (DRL) is proposed, which can achieve task decomposition and task sequence reconstruction under complex constraints. The method designs an environment for deep reinforcement learning based on task information interaction, while improving the SumTree algorithm based on the difference between the loss functions of the target network and the evaluation network, achieving the priority evaluation among tasks. The activation function operation mechanism is introduced into the deep reinforcement learning network, followed by …


Real-Time Scheduling Method For Dynamic Flexible Job Shop Scheduling, Quan Jiang, Jingxuan Wei Jul 2024

Real-Time Scheduling Method For Dynamic Flexible Job Shop Scheduling, Quan Jiang, Jingxuan Wei

Journal of System Simulation

Abstract: A multi-objective dynamic flexible job shop scheduling problem model with machine breakdown and random jobs arrival is constructed to address the interference of dynamic events in manufacturing processing on the scheduling scheme, and a real-time scheduling method with multiobjective proximal policy optimization (MPPO) algorithm is proposed. The MPPO algorithm trains two agents, routing agent (RA) and sequencing agent (SA), for real-time scheduling and real-time processing of dynamic events. It employs a linear combination of weight vectors and reward vectors as reward signals and stores the agents' parameters for each weight vector to optimize multiple objectives. The required state information, …


Design Of Long-Distance Entanglement Distribution Protocols For Quantum Networks, Stav Haldar Jul 2024

Design Of Long-Distance Entanglement Distribution Protocols For Quantum Networks, Stav Haldar

LSU Doctoral Dissertations

Future quantum technologies such as quantum communication, quantum sensing, and distributed quantum computation, will rely on networks of shared entanglement between spatially separated nodes. Distributing entanglement between these nodes, especially over long distances, currently remains a challenge, due to limitations resulting from the fragility of quantum systems, such as photon losses, non-ideal measurements, and quantum memories with short coherence times. In the absence of full-scale fault-tolerant quantum error correction, which can in principle overcome these limitations, we should understand the extent to which we can circumvent these limitations. In this work, we provide improved protocols and policies for entanglement distribution …


Configurable Mirror Descent : Towards A Unification Of Decision Making, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Hau Chan, Bo An Jul 2024

Configurable Mirror Descent : Towards A Unification Of Decision Making, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Decision-making problems, categorized as single-agent, e.g., Atari, cooperative multi-agent, e.g., Hanabi, competitive multi-agent, e.g., Hold’em poker, and mixed cooperative and competitive, e.g., football, are ubiquitous in the real world. Although various methods have been proposed to address the specific decision-making categories, these methods typically evolve independently and cannot generalize to other categories. Therefore, a fundamental question for decision-making is: Can we develop a single algorithm to tackle ALL categories of decision-making problems? There are several main challenges to address this question: i) different decision-making categories involve different numbers of agents and different relationships between agents, ii) different categories have different …


Augmenting Decision With Hypothesis In Reinforcement Learning, Minh Quang Nguyen, Hady Wirawan Lauw Jul 2024

Augmenting Decision With Hypothesis In Reinforcement Learning, Minh Quang Nguyen, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Value-based reinforcement learning is the current State-Of-The-Art due to high sampling efficiency. However, our study shows it suffers from low exploitation in early training period and bias sensitiveness. To address these issues, we propose to augment the decision-making process with hypothesis, a weak form of environment description. Our approach relies on prompting the learning agent with accurate hypotheses, and designing a ready-to-adapt policy through incremental learning. We propose the ALH algorithm, showing detailed analyses on a typical learning scheme and a diverse set of Mujoco benchmarks. Our algorithm produces a significant improvement over value-based learning algorithms and other strong baselines. …


Reinforcement Learning With Maskable Stock Representation For Portfolio Management In Customizable Stock Pools, Wentao Zhang, Yilei Zhao, Shuo Sun, Jie Ying, Yonggang Xie, Zitao Song, Xinrun Wang, Bo An May 2024

Reinforcement Learning With Maskable Stock Representation For Portfolio Management In Customizable Stock Pools, Wentao Zhang, Yilei Zhao, Shuo Sun, Jie Ying, Yonggang Xie, Zitao Song, Xinrun Wang, Bo An

Research Collection School Of Computing and Information Systems

Portfolio management (PM) is a fundamental financial trading task, which explores the optimal periodical reallocation of capitals into different stocks to pursue long-term profits. Reinforcement learning (RL) has recently shown its potential to train profitable agents for PM through interacting with financial markets. However, existing work mostly focuses on fixed stock pools, which is inconsistent with investors’ practical demand. Specifically, the target stock pool of different investors varies dramatically due to their discrepancy on market states and individual investors may temporally adjust stocks they desire to trade (e.g., adding one popular stocks), which lead to customizable stock pools (CSPs). Existing …


Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs Mar 2024

Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs

Theses and Dissertations

Leveraging the Advanced Framework for Simulation, Integration, and Modeling (AFSIM) we investigate the use of reinforcement learning (RL) techniques for imbuing AUCAV agents with high-quality behaviors for the within-visual-range air combat maneuvering problem (ACMP). We formulate the 2v2 WVR ACMP as a Markov decision process wherein friendly AUCAVs are equipped with DEW capabilities and operate with 6 degrees of freedom. We utilize the Double Deep Q-Network RL algorithm, which centrally trains two friendly AUCAVs and employ a phased learning approach, initially exposing the AUCAVs to a dense reward environment for early training, followed by a sparse reward environment to encourage …


A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae Mar 2024

A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae

Theses and Dissertations

A growing demand exists for interpretable artificial intelligence models, leading to extensive research efforts to enhance the explainability and transparency of policies generated by reinforcement learning (RL) methods. This research develops random forest-based RL algorithms as a logical progression in this academic pursuit. The algorithms are evaluated using three standard benchmark environments from OpenAI gym — CartPole, MountainCar, and LunarLander — and compared to implementations of the Deep Q-learning Network (DQN) and Double DQN (DDQN) algorithms for various metrics, including performance, robustness, efficiency, and interpretability. The random forest-based algorithms exhibit superior performance to both neural network-based algorithms in two out …


Reward Penalties On Augmented States For Solving Richly Constrained Rl Effectively, Jiang Hao, Tien Mai, Pradeep Varakanthan, Minh Huy Hoang Mar 2024

Reward Penalties On Augmented States For Solving Richly Constrained Rl Effectively, Jiang Hao, Tien Mai, Pradeep Varakanthan, Minh Huy Hoang

Research Collection School Of Computing and Information Systems

Constrained Reinforcement Learning employs trajectory-based cost constraints (such as expected cost, Value at Risk, or Conditional VaR cost) to compute safe policies. The challenge lies in handling these constraints effectively while optimizing expected reward. Existing methods convert such trajectory-based constraints into local cost constraints, but they rely on cost estimates, leading to either aggressive or conservative solutions with regards to cost. We propose an unconstrained formulation that employs reward penalties over states augmented with costs to compute safe policies. Unlike standard primal-dual methods, our approach penalizes only infeasible trajectories through state augmentation. This ensures that increasing the penalty parameter always …


De Novo Drug Design Using Transformer-Based Machine Translation And Reinforcement Learning Of An Adaptive Monte Carlo Tree Search, Dony Ang, Cyril Rakovski, Hagop S. Atamian Jan 2024

De Novo Drug Design Using Transformer-Based Machine Translation And Reinforcement Learning Of An Adaptive Monte Carlo Tree Search, Dony Ang, Cyril Rakovski, Hagop S. Atamian

Biology, Chemistry, and Environmental Sciences Faculty Articles and Research

The discovery of novel therapeutic compounds through de novo drug design represents a critical challenge in the field of pharmaceutical research. Traditional drug discovery approaches are often resource intensive and time consuming, leading researchers to explore innovative methods that harness the power of deep learning and reinforcement learning techniques. Here, we introduce a novel drug design approach called drugAI that leverages the Encoder–Decoder Transformer architecture in tandem with Reinforcement Learning via a Monte Carlo Tree Search (RL-MCTS) to expedite the process of drug discovery while ensuring the production of valid small molecules with drug-like characteristics and strong binding affinities towards …


Reinforcement Learning For Optimal Kicking Actions In Humanoid Robotics: Advancing Robotic Autonomy And Versatility, Suresh Dodda, Sathish Kumar Chintala, Sukender Reddy Mallreddy, Sharath Chandra Macha, Yashwanth Vasa, Sapan Bharadwaj Bonala, Navin Kamuni, Sujatha Alla Jan 2024

Reinforcement Learning For Optimal Kicking Actions In Humanoid Robotics: Advancing Robotic Autonomy And Versatility, Suresh Dodda, Sathish Kumar Chintala, Sukender Reddy Mallreddy, Sharath Chandra Macha, Yashwanth Vasa, Sapan Bharadwaj Bonala, Navin Kamuni, Sujatha Alla

Engineering Management & Systems Engineering Faculty Publications

Acquiring the necessary skills to perform a work effectively and efficiently requires a significant investment of time and computing power. Previous applications of Reinforcement Learning (RL) for action optimization in humanoid robotics have shown how promising this technology is for moving robotics towards true autonomy and versatility. Therefore, this study offers the first use of RL to create an entirely optimal kicking action for the Alderbaran Nao robot. Kicking motions that were steady, precise, quick, and able to kick farther than any existing RoboCup squad were generated by optimizing for a multi-objective reward function. We demonstrate that the ideal kicking …


Reinforcement Learning: Applying Low Discrepancy Action Selection To Deep Deterministic Policy Gradient, Aleksandr Svishchev Jan 2024

Reinforcement Learning: Applying Low Discrepancy Action Selection To Deep Deterministic Policy Gradient, Aleksandr Svishchev

College of Graduate Studies: Theses & Dissertations

Reinforcement learning (RL) is a subfield of machine learning concerned with agents learning to behave optimally by interacting with an environment. One of the most important topics in RL is how the agent should explore, that is, how to choose actions in order to rate their impact on long-term reward. For example, a simple baseline strategy might be uniformly random action selection. This thesis investigates the heuristic idea that agents will learn faster if they explore by factoring the environment’s state into their decision and intentionally choose actions which are as different as possible from what they have previously observed. …


Advancing Household Robotics: Deep Interactive Reinforcement Learning For Efficient Training And Enhanced Performance, Arpita Soni, Sujatha Alla, Suresh Dodda, Hemanth Volikatla Jan 2024

Advancing Household Robotics: Deep Interactive Reinforcement Learning For Efficient Training And Enhanced Performance, Arpita Soni, Sujatha Alla, Suresh Dodda, Hemanth Volikatla

Engineering Management & Systems Engineering Faculty Publications

The market for domestic robots—made to perform household chore, is growing as these robots relieve people of everyday responsibilities. Domestic robots are generally welcomed for their role in easing human labour, in contrast to industrial robots, which are frequently criticised for displacing human workers. But before these robots can carry out domestic chores, they need to become proficient in a number of minor activities, such as recognizing their surroundings, making decisions, and picking up on human behaviours. Reinforcement learning, or RL, has emerged as a key robotics technology that enables robots to interact with their environment and learn how to …


Enhancing Heart Disease Prediction With Reinforcement Learning And Data Augmentation, Gayathri R., Sangeetha S. K. B., Sandeep Kumar Mathivanan, Hariharan Rajadurai, Benjula Anbu Malar Mb, Saurav Mallik, Hong Qin Jan 2024

Enhancing Heart Disease Prediction With Reinforcement Learning And Data Augmentation, Gayathri R., Sangeetha S. K. B., Sandeep Kumar Mathivanan, Hariharan Rajadurai, Benjula Anbu Malar Mb, Saurav Mallik, Hong Qin

Computer Science Faculty Publications

The study presents a novel method to improve the prediction accuracy of cardiac disease by combining data augmentation techniques with reinforcement learning. The complex nature of cardiac data frequently presents challenges for traditional machine learning models, which results in subpar performance. In response, our fusion methodology improves predictive capabilities by augmenting data and utilizing reinforcement learning's skill at sequential decision-making. Our method predicts cardiac disease with an astounding 94 % accuracy rate, which is an outstanding result. This significant improvement outperforms existing techniques and shows a deeper comprehension of intricate data relationships. The amalgamation of reinforcement learning and data augmentation …


A Memory Efficient Deep Recurrent Q-Learning Approach For Autonomous Wildfire Surveillance, Jeremy A. Cantor Jan 2024

A Memory Efficient Deep Recurrent Q-Learning Approach For Autonomous Wildfire Surveillance, Jeremy A. Cantor

UNF Graduate Theses and Dissertations

Previous literature demonstrates that autonomous UAVs (unmanned aerial vehicles) have the po- tential to be utilized for wildfire surveillance. This advanced technology empowers firefighters by providing them with critical information, thereby facilitating more informed decision-making processes. This thesis applies deep Q-learning techniques to the problem of control policy design under the objective that the UAVs collectively identify the maximum number of locations that are under fire, assuming the UAVs can share their observations. The prohibitively large state space underlying the control policy motivates a neural network approximation, but prior work used only convolutional layers to extract spatial fire information from …


Research And Development Of Simulation Training Platform For Multi-Agent Collaborative Decision-Making, Cheng Cheng, Zhijie Chen, Ziming Guo, Ni Li Dec 2023

Research And Development Of Simulation Training Platform For Multi-Agent Collaborative Decision-Making, Cheng Cheng, Zhijie Chen, Ziming Guo, Ni Li

Journal of System Simulation

Abstract: Reinforcement learning simulation platform can be an interactive and training environment for reinforcement learning. In order to make the simulation platform compatible with the multi-agent reinforcement learning algorithms and meet the needs of simulation in military field, the similar processes in multi-agent reinforcement learning algorithms are refined and a unified interface is designed to embed and verify different types of deep reinforcement learning algorithms on the simulation platform and to optimize the back-end service of the simulation platform to accelerate the training process of the algorithm model. The experimental results show that, by unifing the interface, the simulation platform …


Intercell Dynamic Scheduling Method Based On Deep Reinforcement Learning, Jing Ni, Mengke Ma Nov 2023

Intercell Dynamic Scheduling Method Based On Deep Reinforcement Learning, Jing Ni, Mengke Ma

Journal of System Simulation

Abstract: In order to solve the intercell scheduling problem of dynamic arrival of machining tasks and realize adaptive scheduling in the complex and changeable environment of the intelligent factory, a scheduling method based on a deep Q network is proposed. A complex network with cells as nodes and workpiece intercell machining path as directed edges is constructed, and the degree value is introduced to define the state space with intercell scheduling characteristics. A compound scheduling rule composed of a workpiece layer, unit layer, and machine layer is designed, and hierarchical optimization makes the scheduling scheme more global. Since double deep …


Imitative Generation Of Optimal Guidance Law Based On Reinforcement Learning, Zhengxuan Jia, Tingyu Lin, Yingying Xiao, Guoqiang Shi, Hao Wang, Bi Zeng, Yiming Ou, Pengpeng Zhao Nov 2023

Imitative Generation Of Optimal Guidance Law Based On Reinforcement Learning, Zhengxuan Jia, Tingyu Lin, Yingying Xiao, Guoqiang Shi, Hao Wang, Bi Zeng, Yiming Ou, Pengpeng Zhao

Journal of System Simulation

Abstract: Under the background of high-speed maneuvering target interception, an optimal guidance law generation method for head-on interception independent of target acceleration estimation is proposed based on deep reinforcement learning. In addition, its effectiveness is verified through simulation experiments. As the simulation results suggest, the proposed method successfully achieves head-on interception of high-speed maneuvering targets in 3D space and largely reduces the requirement for target estimation with strong uncertainty, and it is more applicable than the optimal control method.


Uav-Enabled Task Offloading Strategy For Vehicular Edge Computing Networks, Feng Hu, Haiyang Gu, Jun Lin Nov 2023

Uav-Enabled Task Offloading Strategy For Vehicular Edge Computing Networks, Feng Hu, Haiyang Gu, Jun Lin

Journal of System Simulation

Abstract: As intelligent vehicles are equipped with more and more sensors, the explosive growth of sensor data is generated, which brings severe challenges to vehicular communication and computing. In addition, the modern road presents a three-dimensional structure, and the system architecture of traditional vehicular networks cannot guarantee full coverage and seamless computing. A task offloading strategy for UAV-assisted and 6G-enabled (Sixth Generation) vehicular edge computing networks is proposed. Furthermore, a flexible and intelligent vehicular edge computing mode is composed by vehicles and UAVs, which provide three-dimensional edge computing services for delay-sensitive and computation-intensive vehicular tasks, and ensure timely processing and …