Open Access. Powered by Scholars. Published by Universities.®

Computer Sciences Commons

Open Access. Powered by Scholars. Published by Universities.®

Reinforcement learning

Discipline
Institution
Publication Year
Publication
Publication Type
File Type

Articles 61 - 90 of 259

Full-Text Articles in Computer Sciences

Real-Time Scheduling Method For Dynamic Flexible Job Shop Scheduling, Quan Jiang, Jingxuan Wei Jul 2024

Real-Time Scheduling Method For Dynamic Flexible Job Shop Scheduling, Quan Jiang, Jingxuan Wei

Journal of System Simulation

Abstract: A multi-objective dynamic flexible job shop scheduling problem model with machine breakdown and random jobs arrival is constructed to address the interference of dynamic events in manufacturing processing on the scheduling scheme, and a real-time scheduling method with multiobjective proximal policy optimization (MPPO) algorithm is proposed. The MPPO algorithm trains two agents, routing agent (RA) and sequencing agent (SA), for real-time scheduling and real-time processing of dynamic events. It employs a linear combination of weight vectors and reward vectors as reward signals and stores the agents' parameters for each weight vector to optimize multiple objectives. The required state information, …


Design Of Long-Distance Entanglement Distribution Protocols For Quantum Networks, Stav Haldar Jul 2024

Design Of Long-Distance Entanglement Distribution Protocols For Quantum Networks, Stav Haldar

LSU Doctoral Dissertations

Future quantum technologies such as quantum communication, quantum sensing, and distributed quantum computation, will rely on networks of shared entanglement between spatially separated nodes. Distributing entanglement between these nodes, especially over long distances, currently remains a challenge, due to limitations resulting from the fragility of quantum systems, such as photon losses, non-ideal measurements, and quantum memories with short coherence times. In the absence of full-scale fault-tolerant quantum error correction, which can in principle overcome these limitations, we should understand the extent to which we can circumvent these limitations. In this work, we provide improved protocols and policies for entanglement distribution …


Configurable Mirror Descent : Towards A Unification Of Decision Making, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Hau Chan, Bo An Jul 2024

Configurable Mirror Descent : Towards A Unification Of Decision Making, Pengdeng Li, Shuxin Li, Chang Yang, Xinrun Wang, Hau Chan, Bo An

Research Collection School Of Computing and Information Systems

Decision-making problems, categorized as single-agent, e.g., Atari, cooperative multi-agent, e.g., Hanabi, competitive multi-agent, e.g., Hold’em poker, and mixed cooperative and competitive, e.g., football, are ubiquitous in the real world. Although various methods have been proposed to address the specific decision-making categories, these methods typically evolve independently and cannot generalize to other categories. Therefore, a fundamental question for decision-making is: Can we develop a single algorithm to tackle ALL categories of decision-making problems? There are several main challenges to address this question: i) different decision-making categories involve different numbers of agents and different relationships between agents, ii) different categories have different …


Augmenting Decision With Hypothesis In Reinforcement Learning, Minh Quang Nguyen, Hady Wirawan Lauw Jul 2024

Augmenting Decision With Hypothesis In Reinforcement Learning, Minh Quang Nguyen, Hady Wirawan Lauw

Research Collection School Of Computing and Information Systems

Value-based reinforcement learning is the current State-Of-The-Art due to high sampling efficiency. However, our study shows it suffers from low exploitation in early training period and bias sensitiveness. To address these issues, we propose to augment the decision-making process with hypothesis, a weak form of environment description. Our approach relies on prompting the learning agent with accurate hypotheses, and designing a ready-to-adapt policy through incremental learning. We propose the ALH algorithm, showing detailed analyses on a typical learning scheme and a diverse set of Mujoco benchmarks. Our algorithm produces a significant improvement over value-based learning algorithms and other strong baselines. …


Reinforcement Learning With Maskable Stock Representation For Portfolio Management In Customizable Stock Pools, Wentao Zhang, Yilei Zhao, Shuo Sun, Jie Ying, Yonggang Xie, Zitao Song, Xinrun Wang, Bo An May 2024

Reinforcement Learning With Maskable Stock Representation For Portfolio Management In Customizable Stock Pools, Wentao Zhang, Yilei Zhao, Shuo Sun, Jie Ying, Yonggang Xie, Zitao Song, Xinrun Wang, Bo An

Research Collection School Of Computing and Information Systems

Portfolio management (PM) is a fundamental financial trading task, which explores the optimal periodical reallocation of capitals into different stocks to pursue long-term profits. Reinforcement learning (RL) has recently shown its potential to train profitable agents for PM through interacting with financial markets. However, existing work mostly focuses on fixed stock pools, which is inconsistent with investors’ practical demand. Specifically, the target stock pool of different investors varies dramatically due to their discrepancy on market states and individual investors may temporally adjust stocks they desire to trade (e.g., adding one popular stocks), which lead to customizable stock pools (CSPs). Existing …


Scaling Up Cooperative Multi-Agent Reinforcement Learning Systems, Minghong Geng May 2024

Scaling Up Cooperative Multi-Agent Reinforcement Learning Systems, Minghong Geng

Research Collection School Of Computing and Information Systems

Cooperative multi-agent reinforcement learning methods aim to learn effective collaborative behaviours of multiple agents performing complex tasks. However, existing MARL methods are commonly proposed for fairly small-scale multi-agent benchmark problems, wherein both the number of agents and the length of the time horizons are typically restricted. My initial work investigates hierarchical controls of multi-agent systems, where a unified overarching framework coordinates multiple smaller multi-agent subsystems, tackling complex, long-horizon tasks that involve multiple objectives. Addressing another critical need in the field, my research introduces a comprehensive benchmark for evaluating MARL methods in long-horizon, multi-agent, and multi-objective scenarios. This benchmark aims to …


Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs Mar 2024

Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs

Theses and Dissertations

Leveraging the Advanced Framework for Simulation, Integration, and Modeling (AFSIM) we investigate the use of reinforcement learning (RL) techniques for imbuing AUCAV agents with high-quality behaviors for the within-visual-range air combat maneuvering problem (ACMP). We formulate the 2v2 WVR ACMP as a Markov decision process wherein friendly AUCAVs are equipped with DEW capabilities and operate with 6 degrees of freedom. We utilize the Double Deep Q-Network RL algorithm, which centrally trains two friendly AUCAVs and employ a phased learning approach, initially exposing the AUCAVs to a dense reward environment for early training, followed by a sparse reward environment to encourage …


A Reinforcement Learning Approach To The 2v2 Beyond Visual Range Air Combat Maneuvering Problem, Jacob J. Pike Mar 2024

A Reinforcement Learning Approach To The 2v2 Beyond Visual Range Air Combat Maneuvering Problem, Jacob J. Pike

Theses and Dissertations

This research examines a 2v2 air combat maneuvering problem (ACMP) in a Beyond Visual Range (BVR) environment. A discrete-time, infinite-horizon Markov Decision Process (MDP) model represents the BVR-ACMP, seeking to determine high-quality policies for a pair of autonomous aircraft to execute tactical maneuvers and firing decisions. The Advanced Framework for Simulation, Integration, and Modeling (AFSIM) characterizes the complex six-degree of freedom (6-DOF) aircraft operations, encompassing kinematics, sensors, and weapons. Given the high dimensionality and continuous nature of the state and decision variables, a deep reinforcement learning (RL) solution approach is adopted wherein the value function is approximated via a Neural …


A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae Mar 2024

A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae

Theses and Dissertations

A growing demand exists for interpretable artificial intelligence models, leading to extensive research efforts to enhance the explainability and transparency of policies generated by reinforcement learning (RL) methods. This research develops random forest-based RL algorithms as a logical progression in this academic pursuit. The algorithms are evaluated using three standard benchmark environments from OpenAI gym — CartPole, MountainCar, and LunarLander — and compared to implementations of the Deep Q-learning Network (DQN) and Double DQN (DDQN) algorithms for various metrics, including performance, robustness, efficiency, and interpretability. The random forest-based algorithms exhibit superior performance to both neural network-based algorithms in two out …


Reward Penalties On Augmented States For Solving Richly Constrained Rl Effectively, Jiang Hao, Tien Mai, Pradeep Varakanthan, Minh Huy Hoang Mar 2024

Reward Penalties On Augmented States For Solving Richly Constrained Rl Effectively, Jiang Hao, Tien Mai, Pradeep Varakanthan, Minh Huy Hoang

Research Collection School Of Computing and Information Systems

Constrained Reinforcement Learning employs trajectory-based cost constraints (such as expected cost, Value at Risk, or Conditional VaR cost) to compute safe policies. The challenge lies in handling these constraints effectively while optimizing expected reward. Existing methods convert such trajectory-based constraints into local cost constraints, but they rely on cost estimates, leading to either aggressive or conservative solutions with regards to cost. We propose an unconstrained formulation that employs reward penalties over states augmented with costs to compute safe policies. Unlike standard primal-dual methods, our approach penalizes only infeasible trajectories through state augmentation. This ensures that increasing the penalty parameter always …


Decentralized Multimedia Data Sharing In Iov: A Learning-Based Equilibrium Of Supply And Demand, Jiani Fan, Minrui Xu, Jiale Guo, Lwin Khin Shar, Jiawen Kang, Dusit Niyato, Kwok-Yan Lam Mar 2024

Decentralized Multimedia Data Sharing In Iov: A Learning-Based Equilibrium Of Supply And Demand, Jiani Fan, Minrui Xu, Jiale Guo, Lwin Khin Shar, Jiawen Kang, Dusit Niyato, Kwok-Yan Lam

Research Collection School Of Computing and Information Systems

The Internet of Vehicles (IoV) has great potential to transform transportation systems by enhancing road safety, reducing traffic congestion, and improving user experience through onboard infotainment applications. Decentralized data sharing can improve security, privacy, reliability, and facilitate infotainment data sharing in IoVs. However, decentralized data sharing may not achieve the expected efficiency if there are IoV users who only want to consume the shared data but are not willing to contribute their own data to the community, resulting in incomplete information observed by other vehicles and infrastructure, which can introduce additional transmission latency. Therefore, in this paper, by modeling the …


De Novo Drug Design Using Transformer-Based Machine Translation And Reinforcement Learning Of An Adaptive Monte Carlo Tree Search, Dony Ang, Cyril Rakovski, Hagop S. Atamian Jan 2024

De Novo Drug Design Using Transformer-Based Machine Translation And Reinforcement Learning Of An Adaptive Monte Carlo Tree Search, Dony Ang, Cyril Rakovski, Hagop S. Atamian

Biology, Chemistry, and Environmental Sciences Faculty Articles and Research

The discovery of novel therapeutic compounds through de novo drug design represents a critical challenge in the field of pharmaceutical research. Traditional drug discovery approaches are often resource intensive and time consuming, leading researchers to explore innovative methods that harness the power of deep learning and reinforcement learning techniques. Here, we introduce a novel drug design approach called drugAI that leverages the Encoder–Decoder Transformer architecture in tandem with Reinforcement Learning via a Monte Carlo Tree Search (RL-MCTS) to expedite the process of drug discovery while ensuring the production of valid small molecules with drug-like characteristics and strong binding affinities towards …


Freyr⁺: Harvesting Idle Resources In Serverless Computing Via Deep Reinforcement Learning, Hanfei Yu, Hao Wang, Jian Li, Xu Yuan, Seung Jong Park Jan 2024

Freyr⁺: Harvesting Idle Resources In Serverless Computing Via Deep Reinforcement Learning, Hanfei Yu, Hao Wang, Jian Li, Xu Yuan, Seung Jong Park

Computer Science Faculty Research & Creative Works

Serverless computing has revolutionized online service development and deployment with ease-to-use operations, auto-scaling, fine-grained resource allocation, and pay-as-you-go pricing. However, a gap remains in configuring serverless functions - the actual resource consumption may vary due to function types, dependencies, and input data sizes, thus mismatching the static resource configuration by users. Dynamic resource consumption against static configuration may lead to either poor function execution performance or low utilization. This paper proposes Freyr+, a novel resource manager (RM) that dynamically harvests idle resources from over-provisioned functions to accelerate under-provisioned functions for serverless platforms. Freyr+ monitors each function's resource utilization in real-time …


Reinforcement Learning For Optimal Kicking Actions In Humanoid Robotics: Advancing Robotic Autonomy And Versatility, Suresh Dodda, Sathish Kumar Chintala, Sukender Reddy Mallreddy, Sharath Chandra Macha, Yashwanth Vasa, Sapan Bharadwaj Bonala, Navin Kamuni, Sujatha Alla Jan 2024

Reinforcement Learning For Optimal Kicking Actions In Humanoid Robotics: Advancing Robotic Autonomy And Versatility, Suresh Dodda, Sathish Kumar Chintala, Sukender Reddy Mallreddy, Sharath Chandra Macha, Yashwanth Vasa, Sapan Bharadwaj Bonala, Navin Kamuni, Sujatha Alla

Engineering Management & Systems Engineering Faculty Publications

Acquiring the necessary skills to perform a work effectively and efficiently requires a significant investment of time and computing power. Previous applications of Reinforcement Learning (RL) for action optimization in humanoid robotics have shown how promising this technology is for moving robotics towards true autonomy and versatility. Therefore, this study offers the first use of RL to create an entirely optimal kicking action for the Alderbaran Nao robot. Kicking motions that were steady, precise, quick, and able to kick farther than any existing RoboCup squad were generated by optimizing for a multi-objective reward function. We demonstrate that the ideal kicking …


Energy Consumption Optimization Of Uav-Assisted Traffic Monitoring Scheme With Tiny Reinforcement Learning, Xiangjie Kong, Chenhao Ni, Gaohui Duan, Guojiang Shen, Yao Yang, Sajal K. Das Jan 2024

Energy Consumption Optimization Of Uav-Assisted Traffic Monitoring Scheme With Tiny Reinforcement Learning, Xiangjie Kong, Chenhao Ni, Gaohui Duan, Guojiang Shen, Yao Yang, Sajal K. Das

Computer Science Faculty Research & Creative Works

Unmanned Aerial Vehicles (UAVs) can capture pictures of road conditions in all directions and from different angles by carrying high-definition cameras, which helps gather relevant road data more effectively. However, due to their limited energy capacity, drones face challenges in performing related tasks for an extended period. Therefore, a crucial concern is how to plan the path of UAVs and minimize energy consumption. To address this problem, we propose a multi-agent deep deterministic policy gradient based (MADDPG) algorithm for UAV path planning (MAUP). Considering the energy consumption and memory usage of MAUP, we have conducted optimizations to reduce consumption on …


Reinforcement Learning: Applying Low Discrepancy Action Selection To Deep Deterministic Policy Gradient, Aleksandr Svishchev Jan 2024

Reinforcement Learning: Applying Low Discrepancy Action Selection To Deep Deterministic Policy Gradient, Aleksandr Svishchev

College of Graduate Studies: Theses & Dissertations

Reinforcement learning (RL) is a subfield of machine learning concerned with agents learning to behave optimally by interacting with an environment. One of the most important topics in RL is how the agent should explore, that is, how to choose actions in order to rate their impact on long-term reward. For example, a simple baseline strategy might be uniformly random action selection. This thesis investigates the heuristic idea that agents will learn faster if they explore by factoring the environment’s state into their decision and intentionally choose actions which are as different as possible from what they have previously observed. …


Reinforcement Learning-Based Constrained Optimal Control Of Strict-Feedback Nonlinear Systems: Application To Autonomous Underwater Vehicles, Behzad Farzanegan, S. Jagannathan Jan 2024

Reinforcement Learning-Based Constrained Optimal Control Of Strict-Feedback Nonlinear Systems: Application To Autonomous Underwater Vehicles, Behzad Farzanegan, S. Jagannathan

Electrical and Computer Engineering Faculty Research & Creative Works

This paper addresses a constrained neural network (NN)-based optimal tracking scheme for a class of uncertain nonlinear discrete-time systems in strict-feedback form by using a control barrier function (CBF). First, a modified barrier-type cost function is introduced for each subsystem, guiding the actual system trajectory toward the safe set or desired trajectory while avoiding unwanted sets. To address the tracking problem, an augmented system is employed to convert the time-varying optimal tracking to a time-invariant optimal regulation. Then, an actor-critic framework is employed with the backstepping technique to obtain both virtual and actual optimal control policies for each subsystem to …


A New Cache Replacement Policy In Named Data Network Based On Fib Table Information, Mehran Hosseinzadeh, Neda Moghim, Samira Taheri, Nasrin Gholami Jan 2024

A New Cache Replacement Policy In Named Data Network Based On Fib Table Information, Mehran Hosseinzadeh, Neda Moghim, Samira Taheri, Nasrin Gholami

VMASC Publications

Named Data Network (NDN) is proposed for the Internet as an information-centric architecture. Content storing in the router’s cache plays a significant role in NDN. When a router’s cache becomes full, a cache replacement policy determines which content should be discarded for the new content storage. This paper proposes a new cache replacement policy called Discard of Fast Retrievable Content (DFRC). In DFRC, the retrieval time of the content is evaluated using the FIB table information, and the content with less retrieval time receives more discard priority. An impact weight is also used to involve both the grade of retrieval …


Modeling Coupled Driving Behavior During Lane Change: A Multi-Agent Transformer Reinforcement Learning Approach, Hongyu Guo, Mehdi Keyvan-Ekbatani, Kun Xie Jan 2024

Modeling Coupled Driving Behavior During Lane Change: A Multi-Agent Transformer Reinforcement Learning Approach, Hongyu Guo, Mehdi Keyvan-Ekbatani, Kun Xie

Civil & Environmental Engineering Faculty Publications

In a lane change (LC) scenario, the lane change vehicle interacts with surrounding vehicles. The interactions not only affect their driving behaviors but also influence the traffic flow. This study aims to model the coupled behavior of the lane changer and the follower in the target lane during LC. Large-scale real-world connected vehicle (CV) data from the Safety Pilot Model Deployment (SPMD) program are used to extract LCs and study vehicle interactions. A multi-agent Transformer-based deep deterministic policy gradient (MA-TDDPG) method is proposed to model the coupled behaviors during LC. The multi-agent framework can handle the multiple agents’ behaviors with …


Lifelong Learning-Based Optimal Trajectory Tracking Control Of Constrained Nonlinear Affine Systems Using Deep Neural Networks, Irfan Ganie, Sarangapani Jagannathan Jan 2024

Lifelong Learning-Based Optimal Trajectory Tracking Control Of Constrained Nonlinear Affine Systems Using Deep Neural Networks, Irfan Ganie, Sarangapani Jagannathan

Electrical and Computer Engineering Faculty Research & Creative Works

This article presents a novel lifelong integral reinforcement learning (LIRL)-based optimal trajectory tracking scheme using the multilayer (MNN) or deep neural network (Deep NN) for the uncertain nonlinear continuous-time (CT) affine systems subject to state constraints. A critic MNN, which approximates the value function, and a second NN identifier are together used to generate the optimal control policies. The weights of the critic MNN are tuned online using a novel singular value decomposition (SVD)-based method, which can be extended to MNN with the N-hidden layers. Moreover, an online lifelong learning (LL) scheme is incorporated with the critic MNN to mitigate …


Learning From The Past: Using Peer Data To Improve Course Recommendations In Personalized Education, Colton Walker, Sahra Sedigh Sarvestani, Ali R. Hurson Jan 2024

Learning From The Past: Using Peer Data To Improve Course Recommendations In Personalized Education, Colton Walker, Sahra Sedigh Sarvestani, Ali R. Hurson

Electrical and Computer Engineering Faculty Research & Creative Works

This research introduces a recommendation system designed to enhance student success by intelligently personalizing the semester schedules and graduation path based on the student's performance, interests, and background; and inspired by the academic journeys of similar students who have successfully graduated in the past. The proposed recommender system leverages a combination of Markov decision processes, Q-Learning, and collaborative filtering techniques to identify graduation paths with a higher likelihood of success for the student. The proposed model is versatile and generic and can be adapted to various disciplines if sufficient past historical data is available. The proposed model has been prototyped …


Advancing Household Robotics: Deep Interactive Reinforcement Learning For Efficient Training And Enhanced Performance, Arpita Soni, Sujatha Alla, Suresh Dodda, Hemanth Volikatla Jan 2024

Advancing Household Robotics: Deep Interactive Reinforcement Learning For Efficient Training And Enhanced Performance, Arpita Soni, Sujatha Alla, Suresh Dodda, Hemanth Volikatla

Engineering Management & Systems Engineering Faculty Publications

The market for domestic robots—made to perform household chore, is growing as these robots relieve people of everyday responsibilities. Domestic robots are generally welcomed for their role in easing human labour, in contrast to industrial robots, which are frequently criticised for displacing human workers. But before these robots can carry out domestic chores, they need to become proficient in a number of minor activities, such as recognizing their surroundings, making decisions, and picking up on human behaviours. Reinforcement learning, or RL, has emerged as a key robotics technology that enables robots to interact with their environment and learn how to …


A Memory Efficient Deep Recurrent Q-Learning Approach For Autonomous Wildfire Surveillance, Jeremy A. Cantor Jan 2024

A Memory Efficient Deep Recurrent Q-Learning Approach For Autonomous Wildfire Surveillance, Jeremy A. Cantor

UNF Graduate Theses and Dissertations

Previous literature demonstrates that autonomous UAVs (unmanned aerial vehicles) have the po- tential to be utilized for wildfire surveillance. This advanced technology empowers firefighters by providing them with critical information, thereby facilitating more informed decision-making processes. This thesis applies deep Q-learning techniques to the problem of control policy design under the objective that the UAVs collectively identify the maximum number of locations that are under fire, assuming the UAVs can share their observations. The prohibitively large state space underlying the control policy motivates a neural network approximation, but prior work used only convolutional layers to extract spatial fire information from …


Enhancing Heart Disease Prediction With Reinforcement Learning And Data Augmentation, Gayathri R., Sangeetha S. K. B., Sandeep Kumar Mathivanan, Hariharan Rajadurai, Benjula Anbu Malar Mb, Saurav Mallik, Hong Qin Jan 2024

Enhancing Heart Disease Prediction With Reinforcement Learning And Data Augmentation, Gayathri R., Sangeetha S. K. B., Sandeep Kumar Mathivanan, Hariharan Rajadurai, Benjula Anbu Malar Mb, Saurav Mallik, Hong Qin

Computer Science Faculty Publications

The study presents a novel method to improve the prediction accuracy of cardiac disease by combining data augmentation techniques with reinforcement learning. The complex nature of cardiac data frequently presents challenges for traditional machine learning models, which results in subpar performance. In response, our fusion methodology improves predictive capabilities by augmenting data and utilizing reinforcement learning's skill at sequential decision-making. Our method predicts cardiac disease with an astounding 94 % accuracy rate, which is an outstanding result. This significant improvement outperforms existing techniques and shows a deeper comprehension of intricate data relationships. The amalgamation of reinforcement learning and data augmentation …


Research And Development Of Simulation Training Platform For Multi-Agent Collaborative Decision-Making, Cheng Cheng, Zhijie Chen, Ziming Guo, Ni Li Dec 2023

Research And Development Of Simulation Training Platform For Multi-Agent Collaborative Decision-Making, Cheng Cheng, Zhijie Chen, Ziming Guo, Ni Li

Journal of System Simulation

Abstract: Reinforcement learning simulation platform can be an interactive and training environment for reinforcement learning. In order to make the simulation platform compatible with the multi-agent reinforcement learning algorithms and meet the needs of simulation in military field, the similar processes in multi-agent reinforcement learning algorithms are refined and a unified interface is designed to embed and verify different types of deep reinforcement learning algorithms on the simulation platform and to optimize the back-end service of the simulation platform to accelerate the training process of the algorithm model. The experimental results show that, by unifing the interface, the simulation platform …


Neural Airport Ground Handling, Yaoxin Wu, Jianan Zhou, Yunwen Xia, Xianli Zhang, Zhiguang Cao, Jie Zhang Dec 2023

Neural Airport Ground Handling, Yaoxin Wu, Jianan Zhou, Yunwen Xia, Xianli Zhang, Zhiguang Cao, Jie Zhang

Research Collection School Of Computing and Information Systems

Airport ground handling (AGH) offers necessary operations to flights during their turnarounds and is of great importance to the efficiency of airport management and the economics of aviation. Such a problem involves the interplay among the operations that leads to NP-hard problems with complex constraints. Hence, existing methods for AGH are usually designed with massive domain knowledge but still fail to yield high-quality solutions efficiently. In this paper, we aim to enhance the solution quality and computation efficiency for solving AGH. Particularly, we first model AGH as a multiple-fleet vehicle routing problem (VRP) with miscellaneous constraints including precedence, time windows, …


Intercell Dynamic Scheduling Method Based On Deep Reinforcement Learning, Jing Ni, Mengke Ma Nov 2023

Intercell Dynamic Scheduling Method Based On Deep Reinforcement Learning, Jing Ni, Mengke Ma

Journal of System Simulation

Abstract: In order to solve the intercell scheduling problem of dynamic arrival of machining tasks and realize adaptive scheduling in the complex and changeable environment of the intelligent factory, a scheduling method based on a deep Q network is proposed. A complex network with cells as nodes and workpiece intercell machining path as directed edges is constructed, and the degree value is introduced to define the state space with intercell scheduling characteristics. A compound scheduling rule composed of a workpiece layer, unit layer, and machine layer is designed, and hierarchical optimization makes the scheduling scheme more global. Since double deep …


Imitative Generation Of Optimal Guidance Law Based On Reinforcement Learning, Zhengxuan Jia, Tingyu Lin, Yingying Xiao, Guoqiang Shi, Hao Wang, Bi Zeng, Yiming Ou, Pengpeng Zhao Nov 2023

Imitative Generation Of Optimal Guidance Law Based On Reinforcement Learning, Zhengxuan Jia, Tingyu Lin, Yingying Xiao, Guoqiang Shi, Hao Wang, Bi Zeng, Yiming Ou, Pengpeng Zhao

Journal of System Simulation

Abstract: Under the background of high-speed maneuvering target interception, an optimal guidance law generation method for head-on interception independent of target acceleration estimation is proposed based on deep reinforcement learning. In addition, its effectiveness is verified through simulation experiments. As the simulation results suggest, the proposed method successfully achieves head-on interception of high-speed maneuvering targets in 3D space and largely reduces the requirement for target estimation with strong uncertainty, and it is more applicable than the optimal control method.


Uav-Enabled Task Offloading Strategy For Vehicular Edge Computing Networks, Feng Hu, Haiyang Gu, Jun Lin Nov 2023

Uav-Enabled Task Offloading Strategy For Vehicular Edge Computing Networks, Feng Hu, Haiyang Gu, Jun Lin

Journal of System Simulation

Abstract: As intelligent vehicles are equipped with more and more sensors, the explosive growth of sensor data is generated, which brings severe challenges to vehicular communication and computing. In addition, the modern road presents a three-dimensional structure, and the system architecture of traditional vehicular networks cannot guarantee full coverage and seamless computing. A task offloading strategy for UAV-assisted and 6G-enabled (Sixth Generation) vehicular edge computing networks is proposed. Furthermore, a flexible and intelligent vehicular edge computing mode is composed by vehicles and UAVs, which provide three-dimensional edge computing services for delay-sensitive and computation-intensive vehicular tasks, and ensure timely processing and …


Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He Nov 2023

Cgt-Gan: Clip-Guided Text Gan For Image Captioning, Jiarui Yu, Haoran Li, Yanbin Hao, Bin Zhu, Tong Xu, Xiangnan He

Research Collection School Of Computing and Information Systems

The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated image-caption pairs. Recent advanced CLIP-based image captioning without human annotations follows a text-only training paradigm, i.e., reconstructing text from shared embedding space. Nevertheless, these approaches are limited by the training/inference gap or huge storage requirements for text embeddings. Given that it is trivial to obtain images in the real world, we propose CLIP-guided text GAN (CgT-GAN), which incorporates images into the training process to enable the model to "see" real visual modality. Particularly, we use adversarial training to teach CgT-GAN to mimic …