Open Access. Powered by Scholars. Published by Universities.®

Reinforcement learning

Discipline
Institution
Publication Year
Publication
Publication Type

Articles 1 - 30 of 140

Full-Text Articles in Artificial Intelligence and Robotics

Keyframe Selection From Motion Capture Data With Dual-Agent Reinforcement Learning, Kun Hu, Wang, Clinton Mo, Mingyang Ma, Shaohui Mei, Zebin Chen, Zhiyong Wang Nov 2026

Keyframe Selection From Motion Capture Data With Dual-Agent Reinforcement Learning, Kun Hu, Wang, Clinton Mo, Mingyang Ma, Shaohui Mei, Zebin Chen, Zhiyong Wang

Research outputs 2022 to 2026

Animation production workflows centered around motion capture techniques require animators to edit motions based on a set of keyframes. However, most existing keyframe selection methods are optimization-based, which suffer from the issues of flexibility and efficiency. In this paper, a novel deep reinforcement learning method with dual agents are proposed for unsupervised keyframe selection. First, an S-Agent and an R-Agent evaluate the actions of selection and refinement, respectively. A deep spatio-temporal network, namely graph keyframe evaluation network (GKEN), is proposed for the agents. Then, an animation specified reward is devised based on reconstruction, which fulfills three important properties of the …


A Statistical Mechanics Approach To Reinforcement Learning, Jacob Adamczyk Aug 2026

A Statistical Mechanics Approach To Reinforcement Learning, Jacob Adamczyk

Graduate Doctoral Dissertations

Reinforcement learning (RL), the study of optimal decision-making over long timescales in stochastic systems, has recently seen remarkable advances due in large part to the efforts of the deep learning community. RL has witnessed great success in solving problems in video games, robotics, biological control, and language modeling. However, a unified statistical mechanics framework to understand and develop the corresponding algorithms is lacking. To address this issue, we begin by showing that the reinforcement learning problem can be formulated and solved using the tools of statistical mechanics. Drawing on physical principles of free energy minimization and invariance, we address important …


Assessing Flaws In Captcha Security Through Progress In Ai, Jaydon Stanislowski Jul 2026

Assessing Flaws In Captcha Security Through Progress In Ai, Jaydon Stanislowski

Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal

Protecting the internet from the threat of malicious bot activity is an important problem as AI tools become more powerful and commonplace over time. To that end, security measures are employed across websites in the form of CAPTCHAs, short challenges designed to identify and block fake web traffic. Yet, they become less effective over time as AI becomes more powerful, and thus more capable of solving them. This paper examines recent research on the threat to CAPTCHA security posed by current AI models and how this security can be reinforced over time, focusing primarily on Google’s reCAPTCHA v3.


Simulation Study Of Elevator Group Control Scheduling Based On Real-Time Occupancy Perception, Yiyong Han, Shiyu Wang, Yang Ye, Zhen Zhang, Fengque Pei, Minghai Yuan Jul 2026

Simulation Study Of Elevator Group Control Scheduling Based On Real-Time Occupancy Perception, Yiyong Han, Shiyu Wang, Yang Ye, Zhen Zhang, Fengque Pei, Minghai Yuan

Journal of System Simulation

Abstract: To address the conflict between car space allocation and peak passenger flow response efficiency in elevator group control scheduling, a multi-objective scheduling method based on proximal policy optimization (PPO) with real-time occupancy perception was proposed. A simulation environment considering car capacity constraints was constructed, and a reward-penalty mechanism with average passenger waiting time, system energy consumption, and car congestion as optimization objectives was designed. Based on this, state and action spaces were defined to form a PPO-based scheduling framework; a simulation platform integrating traffic flow visualization, policy scheduling, and performance evaluation was developed. Simulation results show that this method …


Research On Output Feedback Control Based On Reinforcement Learning For Overhead Crane, Minghui Li, Daoxiang Gao Jun 2026

Research On Output Feedback Control Based On Reinforcement Learning For Overhead Crane, Minghui Li, Daoxiang Gao

Journal of System Simulation

An output feedback control algorithm is designed based on reinforcement learning for the optimal control problem of overhead crane system. A high gain observer (HGO) is designed using output data to estimate the unmeasurable states of the overhead crane system. Based on the estimated states from the high-gain observer, a policy iteration (PI) method is designed with integral reinforcement learning, which uses Critic and Actor neural networks to approximate the optimal value function and control strategy, and adjusts the neural network weights in real time through online adaptive algorithms. According to the Lyapunov stability theory, the uniform ultimate boundedness of …


Improved Nsga-Ii For Dual-Resource Flexible Job Shop Scheduling Considering Worker Load, Guohui Zhang, Yuan Ren, Changjun Wu, Xiaofei Kou Jun 2026

Improved Nsga-Ii For Dual-Resource Flexible Job Shop Scheduling Considering Worker Load, Guohui Zhang, Yuan Ren, Changjun Wu, Xiaofei Kou

Journal of System Simulation

For the dual-resource-constrained flexible job shop scheduling problem considering worker load, an evolutionary algorithm integrating reinforcement learning was proposed. A three-stage encoding conforming to the problem characteristics was designed, and three initialization methods were combined to improve the population quality; a left-insertion decoding method based on worker load was designed to ensure that the completion time of the operation is less than the maximum processable time of the worker on the current day; two neighborhood structures based on the critical path were constructed to enhance the local exploration ability of the population; reinforcement learning was integrated to enable the …


Shaping Emergent Competitive And Cooperative Behaviors In Multi-Agent General-Sum Games, Ethan F. Erickson May 2026

Shaping Emergent Competitive And Cooperative Behaviors In Multi-Agent General-Sum Games, Ethan F. Erickson

Honors Projects

Reinforcement learning (RL) algorithms can train agents to solve problems in environments using complex behaviors that are not explicitly programmed, known as emergent behaviors. The goal of our research is to investigate how different RL reward values influence the emergence of competitive and cooperative behaviors in games with teams of multiple agents. Specifically, we focus on general-sum games, in which the sum of gains and losses of each team may be non-zero, allowing situations for agents to mutually benefit or mutually fail. Using Unity’s ML-Agents Toolkit to train agents with RL self-play in bounded 2D environments, we identify high-level behaviors …


Review On Optimization Of Simulation Modeling Strategies For Spacecraft Orbit Avoidance, Guozheng Li, Rui Wang, Shichao Fan, Xintong Cai, Xinyue Zhai Apr 2026

Review On Optimization Of Simulation Modeling Strategies For Spacecraft Orbit Avoidance, Guozheng Li, Rui Wang, Shichao Fan, Xintong Cai, Xinyue Zhai

Journal of System Simulation

Abstract: The number of on-orbit spacecraft increases exponentially; the space environment becomes more complex, and the collision risk of on-orbit spacecraft increases significantly. On-orbit safety is thus severely threatened, posing higher requirements for orbit avoidance methods. The costs and risks of space activities are extremely high, making simulation an effective method to solve complex problems of orbit avoidance. The modeling, solution, and simulation methods for the two core issues of spacecraft orbit avoidance, "collision avoidance" and "pursuit-evasion games", were systematically reviewed, and the existing shortcomings were analyzed. The applications of technologies such as deep reinforcement learning in promoting orbit avoidance …


Robot Path Planning By Reinforcement Learning Based On Sac3q-Hdm, Dequan Li, Wan Xiong Mar 2026

Robot Path Planning By Reinforcement Learning Based On Sac3q-Hdm, Dequan Li, Wan Xiong

Journal of System Simulation

Abstract: To address the issues of overestimated and underestimated biases, low sample utilization rate, and the inability to balance exploration and exploitation in reinforcement learning for path planning, an improved SAC method was proposed. The size balance of entropy was explored and utilized through adaptive temperature coefficient adjustment; on the basis of the SAC framework, a triple Critic architecture was introduced to dynamically weight and fuse the minimum and average values through Qvalue uncertainty, balancing overestimated and underestimated biases. A mixed dynamic sampling experience replay buffer was designed; experience data was partitioned based on reward thresholds; sampling ratios were dynamically …


Learning To Search For Vehicle Routing With Multiple Time Windows, Kuan Xu, Zhiguang Cao, Chenlong Zheng, Lindong Liu Mar 2026

Learning To Search For Vehicle Routing With Multiple Time Windows, Kuan Xu, Zhiguang Cao, Chenlong Zheng, Lindong Liu

Research Collection School Of Computing and Information Systems

In this study, we propose a reinforcement learning-based adaptive variable neighborhood search (RL-AVNS) method designed for effectively solving the Vehicle Routing Problem with Multiple Time Windows (VRPMTW). Unlike traditional adaptive approaches that rely solely on historical operator performance, our method integrates a reinforcement learning framework to dynamically select neighborhood operators based on real-time solution states and learned experience. We introduce a fitness metric that quantifies customers’ temporal flexibility to improve the shaking phase, and employ a transformer-based neural policy network to intelligently guide operator selection during the local search. Extensive computational experiments are conducted on realistic scenarios derived from the …


Reinforcement Learning Based Method For Uav Team Orienteering Optimization Under Multi-Constraint Condition, Can Yang, Kai Chen, Feng Zhu Feb 2026

Reinforcement Learning Based Method For Uav Team Orienteering Optimization Under Multi-Constraint Condition, Can Yang, Kai Chen, Feng Zhu

Journal of System Simulation

Abstract: Traditional optimization methods struggle with efficiency, while reinforcement learning approaches often yield low solution quality and high training costs. In response, this paper proposes an attention mechanism-based reinforcement learning method. A dynamic attention strategy network with multi-information fusion is designed to improve solution quality. A visibility-graph approach is employed to simplify threat zone constraints and speed up convergence, and a decoding sequence reordering mechanism is introduced for further performance optimization of the solution. The simulation results show that the method generates high-quality solutions within milliseconds, achieving total rewards that approach or even surpass those obtained by traditional solvers …


Aura: An Ai-Powered Multimodal Prototype For Adaptive Apraxia Of Speech Therapy And Communication Support, Omotayo Omoyemi, Rachel K. Johnson Jan 2026

Aura: An Ai-Powered Multimodal Prototype For Adaptive Apraxia Of Speech Therapy And Communication Support, Omotayo Omoyemi, Rachel K. Johnson

Speech-Language Pathology Faculty Publications

Apraxia of Speech (AOS) is a motor speech disorder that significantly limits communication and requires intensive, long-term therapy. Access to consistent treatment is often constrained by shortages of speech-language pathologists, high costs, and limited opportunities for continuous monitoring outside clinical settings. Recent advances in Artificial Intelligence (AI) provide new opportunities to support scalable and personalized speech therapy.

This paper presents AURA (Adaptive Understanding and Relearning Assistant for Apraxia), a multimodal AI framework designed to support speech therapy, progress monitoring, and communication for individuals with AOS. The system integrates speech analysis, machine learning–based error detection, reinforcement learning for adaptive therapy, and …


Design And Analysis Of Modern Quantum Neural Network Architectures For Intelligent Systems, Lakshmi Chandrakanth Kasireddy, Prabhakara Rao Kapula, Dineshkumar Rajendran, Neha Bharani, Srikanth Pulipeti, Islombek Khushvaktov Jan 2026

Design And Analysis Of Modern Quantum Neural Network Architectures For Intelligent Systems, Lakshmi Chandrakanth Kasireddy, Prabhakara Rao Kapula, Dineshkumar Rajendran, Neha Bharani, Srikanth Pulipeti, Islombek Khushvaktov

Computer Science Faculty Publications

Quantum neural networks (QNNs) offer a principled pathway for integrating quantum computation with machine learning through superposition- and entanglement-based representations. This chapter proposes an architecture-aware design and evaluation framework for modern QNNs, emphasizing robustness and system feasibility alongside predictive performance. Multiple architectures variational QNNs, quantum convolutional neural networks, tensor-network hybrids, and fully quantum models—are assessed under a unified protocol. Experimental analysis shows that the proposed architecture-search–guided QNN achieves 91.8% classification accuracy and an F1-score of 0.914, outperforming fixed-template variational QNNs by approximately 5.6 percentage points. Under depolarizing noise with probability p = 0.10, the proposed model retains 85.3% accuracy, whereas …


ℵ-Ipomdp: Mitigating Deception In A Cognitive Hierarchy With Off-Policy Counterfactual Anomaly Detection, Nitay Alon, Joseph M. Barnby, Stefan Sarkadi, Lion Schulz, Jeffrey S. Rosenschein, Peter Dayan Jan 2026

ℵ-Ipomdp: Mitigating Deception In A Cognitive Hierarchy With Off-Policy Counterfactual Anomaly Detection, Nitay Alon, Joseph M. Barnby, Stefan Sarkadi, Lion Schulz, Jeffrey S. Rosenschein, Peter Dayan

Research outputs 2022 to 2026

Social agents with finitely nested opponent models are vulnerable to manipulation by agents with deeper recursive capabilities. This imbalance, rooted in logic and the theory of recursive modelling frameworks, cannot be solved directly. We propose a computational framework called ℵ-IPOMDP, which augments the Bayesian inference of model-based RL agents with an anomaly detection algorithm and an out-of-belief policy. Our mechanism allows agents to realize that they are being deceived, even if they cannot understand how, and to deter opponents via a credible threat. We test this framework in both a mixed-motive and a zero-sum game. Our results demonstrate the ℵ-mechanism’s …


Explainable Physics-Based Constraints On Reinforcement Learning For Accelerator Optimization, Jonathan Colen, Malachi Schram, Kishansingh Rajput, Armen Kasparian Jan 2026

Explainable Physics-Based Constraints On Reinforcement Learning For Accelerator Optimization, Jonathan Colen, Malachi Schram, Kishansingh Rajput, Armen Kasparian

Data Science Faculty Publications

We present a reinforcement learning (RL) framework for optimizing particle accelerator experiments that builds explainable physics-based constraints on agent behavior. The goal is to increase transparency and trust by letting users verify that the agent’s decision-making process incorporates suitable physics. Our algorithm uses a learnable surrogate function for physical observables, such as energy, and uses them to fine-tune how actions are chosen. This surrogate can be represented by a neural network or by an interpretable sparse dictionary model. We test our algorithm on a range of particle accelerator optimization environments designed to emulate the Continuous Electron Beam Accelerator Facility at …


Large Language Models As End-To-End Combinatorial Optimization Solvers, Xia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao, Yingqian Zhang Dec 2025

Large Language Models As End-To-End Combinatorial Optimization Solvers, Xia Jiang, Yaoxin Wu, Minshuo Li, Zhiguang Cao, Yingqian Zhang

Research Collection School Of Computing and Information Systems

Combinatorial optimization (CO) problems, central to decision-making scenarios like logistics and manufacturing, are traditionally solved using problem-specific algorithms requiring significant domain expertise. While large language models (LLMs) have shown promise in automating CO problem solving, existing approaches rely on intermediate steps such as code generation or solver invocation, limiting their generality and accessibility. This paper introduces a novel framework that empowers LLMs to serve as end-to-end CO solvers by directly mapping natural language problem descriptions to solutions. We propose a two-stage training strategy: supervised fine-tuning (SFT) imparts LLMs with solution generation patterns from domain-specific solvers, while a feasibility-and-optimality-aware reinforcement learning …


Multi-Task Vehicle Routing Solver Via Mixture Of Specialized Experts Under State-Decomposable Mdp, Yuxin Pan, Zhiguang Cao, Chengyang Gu, Liu Liu, Peilin Zhao, Yize Chen, Fangzhen Lin Dec 2025

Multi-Task Vehicle Routing Solver Via Mixture Of Specialized Experts Under State-Decomposable Mdp, Yuxin Pan, Zhiguang Cao, Chengyang Gu, Liu Liu, Peilin Zhao, Yize Chen, Fangzhen Lin

Research Collection School Of Computing and Information Systems

Existing neural methods for multi-task vehicle routing problems (VRPs) typically learn unified solvers to handle multiple constraints simultaneously. However, they often underutilize the compositional structure of VRP variants, each derivable from a common set of basis VRP variants. This critical oversight causes unified solvers to miss out the potential benefits of basis solvers, each specialized for a basis VRP variant. To overcome this limitation, we propose a framework that enables unified solvers to perceive the shared-component nature across VRP variants by proactively reusing basis solvers, while mitigating the exponential growth of trained neural solvers. Specifically, we introduce a State-Decomposable MDP …


Auv Path Planning Based On Behavior Cloning And Improved Dqn In Partially Unknown Environments, Lijing Xing, Min Li, Xiangguang Zeng, Ping Zhang, Bei Peng Nov 2025

Auv Path Planning Based On Behavior Cloning And Improved Dqn In Partially Unknown Environments, Lijing Xing, Min Li, Xiangguang Zeng, Ping Zhang, Bei Peng

Journal of System Simulation

Abstract: To address the problems of large randomness and slow convergence of the DQN dynamic path planning algorithm for a single autonomous underwater vehicle (AUV) in a partially unknown environment, a path planning method combining behavior cloning with A* algorithm and DQN (BA_DQN) was proposed. Based on the known environmental information, an improved A* algorithm incorporating ocean current resistance was proposed to guide DQN, thereby reducing the randomness of the DQN algorithm. By considering the complexity of the marine environment, the sampling probability was improved again after expanding the positive experience pool to enhance the training success rate. To address …


Modeling And Simulation Of Traffic Signal Control Based On Mlp With Improved Gcn-Td3, Deqi Huang, Yating Tu, Zhenhua Zhang, Xin Guo Oct 2025

Modeling And Simulation Of Traffic Signal Control Based On Mlp With Improved Gcn-Td3, Deqi Huang, Yating Tu, Zhenhua Zhang, Xin Guo

Journal of System Simulation

Abstract: To address the issues of uneven traffic flow at urban intersections, limited road capacity, and the poor coordination of existing traffic signal control algorithms, a traffic signal control algorithm based on graph convolutional reinforcement learning was proposed. By utilizing a multilayer perceptron, the dynamic features of vehicles and phase information at the controlled intersection and its neighboring intersections were extracted. A graph convolutional neural network was then employed to aggregate these vehicle dynamic features into potential features representing regional traffic. The control strategy was derived through multiple iterations of an improved twin delayed deep deterministic policy gradient (TD3) algorithm. …


Robust Ai Solutions For Financial Markets Through Generative Modeling, Dynamic Graph Learning, And Reinforcement-Based Portfolio Optimization, Jingyi Gu Aug 2025

Robust Ai Solutions For Financial Markets Through Generative Modeling, Dynamic Graph Learning, And Reinforcement-Based Portfolio Optimization, Jingyi Gu

Dissertations

Financial markets are inherently uncertain and dynamic, driven by complex factors such as macroeconomic signals, investor sentiment, and evolving inter-asset relationships. While machine learning has advanced financial modeling, existing approaches often fall short in addressing the real-world intricacies of finance. This dissertation confronts two critical challenges, human-driven stochasticity and risk-intensive decision-making under real-world trading constraints, while seizing a pivotal opportunity, the structural dynamics of evolving financial systems. These elements are foundational to advancing robust and practical financial intelligence.

To this end, this dissertation develops a unified framework for robust financial modeling and decision-making. The framework is architected as a progressive, …


Rl4co: An Extensive Reinforcement Learning For Combinatorial Optimization Benchmark, Federico Berto, Et. Al Aug 2025

Rl4co: An Extensive Reinforcement Learning For Combinatorial Optimization Benchmark, Federico Berto, Et. Al

Research Collection School Of Computing and Information Systems

Combinatorial optimization (CO) is fundamental to several real-world applications, from logistics and scheduling to hardware design and resource allocation. Deep reinforcement learning (RL) has recently shown significant benefits in solving CO problems, reducing reliance on domain expertise and improving computational efficiency. However, the absence of a unified benchmarking framework leads to inconsistent evaluations, limits reproducibility, and increases engineering overhead, raising barriers to adoption for new researchers. To address these challenges, we introduce RL4CO, a unified and extensive benchmark with in-depth library coverage of 27 CO problem environments and 23 state-of-the-art baselines. Built on efficient software libraries and best practices in …


Interactive Generative Modeling: A Pathway For Improved Simulation And Decision Making, Changyu Chen Jun 2025

Interactive Generative Modeling: A Pathway For Improved Simulation And Decision Making, Changyu Chen

Dissertations and Theses Collection (Open Access)

This dissertation presents Interactive Generative Modeling (IGM), a unified perspective that integrates interactive paradigm and generative modeling to advance the development of general-purpose intelligent systems. IGM is motivated by the observation that while reinforcement learning (RL) has mastered a wide range of complex simulated tasks, it struggles to generalize in high-dimensional, open-ended tasks. In contrast, generative models excel in such settings due to their expressivity and their ability to serve as powerful priors (e.g., LLMs pretrained on massive corpora). By bridging these two paradigms, IGM offers a promising path forward.

The first direction explored in this dissertation is IGM for …


Machine Learning And Optimization For Intelligent Decision-Making, Elson Cibaku May 2025

Machine Learning And Optimization For Intelligent Decision-Making, Elson Cibaku

Dissertations

This dissertation presents a series of innovative machine learning and optimization model designs that address complex operational challenges across logistics and power systems. By integrating advanced neural architectures with robust optimization techniques, the work delivers scalable solutions designed to improve efficiency, reliability, and decision-making in dynamic and real-world environments. The first study introduces a two-stage approach to effective vaccine distribution. This framework tackles the capacitated vehicle routing problem by combining adaptive clustering techniques with reinforcement learning and a simulated annealing pickup policy. Through extensive computational experiments, the approach demonstrates substantial improvements in routing efficiency, reducing both computational time and logistical …


Model-Based Reinforcement Learning And Deep Learning For Power Converter Circuit Design Automation, Shaoze Fan May 2025

Model-Based Reinforcement Learning And Deep Learning For Power Converter Circuit Design Automation, Shaoze Fan

Dissertations

This dissertation presents a comprehensive automated framework for power converter design, leveraging reinforcement learning (RL) and graph-transformer networks (GTN) to address critical inefficiencies in traditional manual topology optimization. Motivated by the combinatorial increase of circuit design spaces and the computational cost of iterative simulations, this work develops a robust framework for generating energy-efficient topologies requiring rapid and reliable circuit design.

The framework integrates three key components: (1) an upper-confidence-bound-tree-based (UCT-based) RL model for circuit topology space exploration, (2) parallelized UCT algorithms to accelerate exploration processes, (3) a Graph-Transformer-based Network enabling fast circuit performance evaluation. Experimental validation demonstrates the whole framework …


Satellite Reorientation Using Reinforcement Learning Under Unknown Attitude Failure, Matthew Willoughby May 2025

Satellite Reorientation Using Reinforcement Learning Under Unknown Attitude Failure, Matthew Willoughby

Doctoral Dissertations and Master's Theses

This study presents a reinforcement learning (RL) approach for reestablishing communication with deep-space satellites under unknown attitude determination and control system (ADCS) failures. When traditional fault-tolerant control methods cannot restore signal, the proposed RL controller acts as a last-resort measure by autonomously reorienting the satellite’s antenna toward Earth while charging the battery via solar panels. A generic reward function, designed for the RL-based method, enables the controller to adapt to diverse failure scenarios, including severe actuator noise, misalignment, and complete actuator failure. Simulations are conducted in the Basilisk environment and trained with the tonic framework and demonstrate ranging capabilities of …


Controlling A Mobile Inverted Pendulum And Optimizing Leaning Angle To Apply Force Using Reinforcement Learning, Aryan Mediratta May 2025

Controlling A Mobile Inverted Pendulum And Optimizing Leaning Angle To Apply Force Using Reinforcement Learning, Aryan Mediratta

2025 Spring Honors Capstone Projects - Archive

Reinforcement Learning is a Machine Learning paradigm that involves simulating learning through rewards and penalties in intelligent systems. This technique is often employed in robotics when traditional control methods are insufficient or when human intuition does not provide a good solution on how to control robot systems, This project involves training a Segway-style Mobile Inverted Pendulum (MIP) robot to balance and push a box forward. The BeagleBone Blue board is used that includes a built-in Inertial Measurement Unit (IMU) and encoder ports. These sensors enable the system to measure its current state. The goal is to find the optimal leaning …


Signal Timing Optimization Via Reinforcement Learning With Traffic Flow Prediction, Ming Xu, Jinye Li, Dongyu Zuo, Jing Zhang Apr 2025

Signal Timing Optimization Via Reinforcement Learning With Traffic Flow Prediction, Ming Xu, Jinye Li, Dongyu Zuo, Jing Zhang

Journal of System Simulation

Abstract: In response to the existing reinforcement learning-based traffic signal control methods that do not consider the changing trends in traffic flow, leading to congestion and inability to adapt to complex and variable road conditions, we propose a traffic signal timing optimization reinforcement learning method based on flow prediction. A phase timing amplitude control model is introduced. This model analyzes the spatiotemporal characteristics of historical traffic data to predict the flow for the next time slot and calculates a reasonable range for phase timing based on the prediction results. The H-PPO algorithm is employed to control the signal phase while …


Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang Apr 2025

Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang

Research Collection School Of Computing and Information Systems

With the widespread adoption of Internet Protocol (IP) communication technology and web-based platforms, cloud manufacturing has become a significant hallmark of Industry 4.0. Integrating graph algorithms into these web-enabled environments is crucial as they facilitate the representation and analysis of complex relationships in manufacturing processes, enabling efficient decision-making and adaptability in dynamic environments. As a key scheduling problem in cloud manufacturing, the flexible job-shop scheduling problem (FJSP) finds extensive applications in real-world scenarios. However, traditional FJSP-solving methods struggle to meet the efficiency and adaptability demands of cloud manufacturing due to generalization issues and excessive computational time, while reinforcement learning-based methods …


Graph-Assisted Offline-Online Deep Reinforcement Learning For Dynamic Workflow Scheduling, Yifan Yang, Gang Chen, Hui Ma, Cong Zhang, Zhiguang Cao, Mengjie Zhang Apr 2025

Graph-Assisted Offline-Online Deep Reinforcement Learning For Dynamic Workflow Scheduling, Yifan Yang, Gang Chen, Hui Ma, Cong Zhang, Zhiguang Cao, Mengjie Zhang

Research Collection School Of Computing and Information Systems

Dynamic workflow scheduling (DWS) in cloud computing presents substantial challenges due to heterogeneous machine configurations, unpredictable workflow arrivals/patterns, and constantly evolving environments. However, existing research often assumes homogeneous setups and static conditions, limiting flexibility and adaptability in real-world scenarios. In this paper, we propose a novel Graph assisted Offline-Online Deep Reinforcement Learning (GOODRL) approach to building an effective and efficient scheduling agent for DWS. Our approach features three key innovations: (1) a task-specific graph representation and a Graph Attention Actor Network that enable the agent to dynamically assign focused tasks to heterogeneous machines while explicitly considering the future impact of …


On Minimizing Adversarial Counterfactual Error In Adversarial Reinforcement Learning, Roman Belaire, Arunesh Sinha, Pradeep Varakantham Apr 2025

On Minimizing Adversarial Counterfactual Error In Adversarial Reinforcement Learning, Roman Belaire, Arunesh Sinha, Pradeep Varakantham

Research Collection School Of Computing and Information Systems

Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge inherent to adversarial perturbations is that by altering the information observed by the agent, the state becomes only partially observable. Existing approaches address this by either enforcing consistent actions across nearby states or maximizing the worst-case value within adversarially perturbed observations. However, the former suffers from performance degradation when attacks succeed, while the latter tends to be overly conservative, leading to suboptimal performance in benign settings. We hypothesize that these limitations stem from their failing to account for …