Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (140)
- Engineering (102)
- Computer Engineering (57)
- Operations Research, Systems Engineering and Industrial Engineering (53)
- Numerical Analysis and Scientific Computing (47)
-
- Systems Science (34)
- Electrical and Computer Engineering (29)
- Databases and Information Systems (27)
- Theory and Algorithms (23)
- Social and Behavioral Sciences (14)
- OS and Networks (11)
- Public Affairs, Public Policy and Public Administration (9)
- Software Engineering (9)
- Transportation (9)
- Graphics and Human Computer Interfaces (7)
- Medicine and Health Sciences (7)
- Applied Mathematics (5)
- Business (5)
- Operational Research (5)
- Other Computer Sciences (5)
- Computer and Systems Architecture (4)
- Data Science (4)
- Information Security (4)
- Physics (4)
- Aerospace Engineering (3)
- Chemistry (3)
- Controls and Control Theory (3)
- Programming Languages and Compilers (3)
- Institution
-
- Singapore Management University (72)
- China Simulation Federation (34)
- Missouri University of Science and Technology (16)
- Brigham Young University (12)
- Air Force Institute of Technology (9)
-
- Old Dominion University (9)
- University of Texas at Arlington (9)
- TÜBİTAK (8)
- MBZUAI (7)
- New Jersey Institute of Technology (5)
- Portland State University (5)
- University of Texas Rio Grande Valley (5)
- Edith Cowan University (4)
- San Jose State University (4)
- University of Denver (4)
- Utah State University (4)
- Zayed University (3)
- Bucknell University (2)
- Chapman University (2)
- Georgia Southern University (2)
- Missouri State University (2)
- Nova Southeastern University (2)
- Technological University Dublin (2)
- University of Kentucky (2)
- University of Minnesota Morris Digital Well (2)
- University of Nebraska - Lincoln (2)
- University of Nevada, Las Vegas (2)
- University of South Carolina (2)
- Western Michigan University (2)
- California State University, San Bernardino (1)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (69)
- Journal of System Simulation (34)
- Theses and Dissertations (16)
- Electrical and Computer Engineering Faculty Research & Creative Works (12)
- Faculty Publications (9)
-
- Turkish Journal of Electrical Engineering and Computer Sciences (8)
- Dissertations (7)
- Machine Learning Faculty Publications (7)
- Electronic Theses and Dissertations (6)
- Computer Science and Engineering Dissertations - Archive (5)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (4)
- Computer Science Faculty Publications (4)
- Computer Science Faculty Research & Creative Works (4)
- Dissertations and Theses (4)
- Master's Projects (4)
- All Works (3)
- Dissertations and Theses Collection (Open Access) (3)
- Research outputs 2022 to 2026 (3)
- Articles (2)
- CCAC Theses and Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science and Engineering Theses - Archive (2)
- Data Science Faculty Publications (2)
- Electrical & Computer Engineering Faculty Research (2)
- Engineering Management & Systems Engineering Faculty Publications (2)
- Graduate Theses/Dissertations (2)
- Journal of Undergraduate Research (2)
- Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal (2)
- School of Computing: Dissertations, Theses, and Student Research (2)
- Theses and Dissertations--Computer Science (2)
- Publication Type
Articles 31 - 60 of 259
Full-Text Articles in Computer Sciences
Interactive Generative Modeling: A Pathway For Improved Simulation And Decision Making, Changyu Chen
Interactive Generative Modeling: A Pathway For Improved Simulation And Decision Making, Changyu Chen
Dissertations and Theses Collection (Open Access)
This dissertation presents Interactive Generative Modeling (IGM), a unified perspective that integrates interactive paradigm and generative modeling to advance the development of general-purpose intelligent systems. IGM is motivated by the observation that while reinforcement learning (RL) has mastered a wide range of complex simulated tasks, it struggles to generalize in high-dimensional, open-ended tasks. In contrast, generative models excel in such settings due to their expressivity and their ability to serve as powerful priors (e.g., LLMs pretrained on massive corpora). By bridging these two paradigms, IGM offers a promising path forward.
The first direction explored in this dissertation is IGM for …
Machine Learning And Optimization For Intelligent Decision-Making, Elson Cibaku
Machine Learning And Optimization For Intelligent Decision-Making, Elson Cibaku
Dissertations
This dissertation presents a series of innovative machine learning and optimization model designs that address complex operational challenges across logistics and power systems. By integrating advanced neural architectures with robust optimization techniques, the work delivers scalable solutions designed to improve efficiency, reliability, and decision-making in dynamic and real-world environments. The first study introduces a two-stage approach to effective vaccine distribution. This framework tackles the capacitated vehicle routing problem by combining adaptive clustering techniques with reinforcement learning and a simulated annealing pickup policy. Through extensive computational experiments, the approach demonstrates substantial improvements in routing efficiency, reducing both computational time and logistical …
Model-Based Reinforcement Learning And Deep Learning For Power Converter Circuit Design Automation, Shaoze Fan
Model-Based Reinforcement Learning And Deep Learning For Power Converter Circuit Design Automation, Shaoze Fan
Dissertations
This dissertation presents a comprehensive automated framework for power converter design, leveraging reinforcement learning (RL) and graph-transformer networks (GTN) to address critical inefficiencies in traditional manual topology optimization. Motivated by the combinatorial increase of circuit design spaces and the computational cost of iterative simulations, this work develops a robust framework for generating energy-efficient topologies requiring rapid and reliable circuit design.
The framework integrates three key components: (1) an upper-confidence-bound-tree-based (UCT-based) RL model for circuit topology space exploration, (2) parallelized UCT algorithms to accelerate exploration processes, (3) a Graph-Transformer-based Network enabling fast circuit performance evaluation. Experimental validation demonstrates the whole framework …
Satellite Reorientation Using Reinforcement Learning Under Unknown Attitude Failure, Matthew Willoughby
Satellite Reorientation Using Reinforcement Learning Under Unknown Attitude Failure, Matthew Willoughby
Doctoral Dissertations and Master's Theses
This study presents a reinforcement learning (RL) approach for reestablishing communication with deep-space satellites under unknown attitude determination and control system (ADCS) failures. When traditional fault-tolerant control methods cannot restore signal, the proposed RL controller acts as a last-resort measure by autonomously reorienting the satellite’s antenna toward Earth while charging the battery via solar panels. A generic reward function, designed for the RL-based method, enables the controller to adapt to diverse failure scenarios, including severe actuator noise, misalignment, and complete actuator failure. Simulations are conducted in the Basilisk environment and trained with the tonic framework and demonstrate ranging capabilities of …
Controlling A Mobile Inverted Pendulum And Optimizing Leaning Angle To Apply Force Using Reinforcement Learning, Aryan Mediratta
Controlling A Mobile Inverted Pendulum And Optimizing Leaning Angle To Apply Force Using Reinforcement Learning, Aryan Mediratta
2025 Spring Honors Capstone Projects - Archive
Reinforcement Learning is a Machine Learning paradigm that involves simulating learning through rewards and penalties in intelligent systems. This technique is often employed in robotics when traditional control methods are insufficient or when human intuition does not provide a good solution on how to control robot systems, This project involves training a Segway-style Mobile Inverted Pendulum (MIP) robot to balance and push a box forward. The BeagleBone Blue board is used that includes a built-in Inertial Measurement Unit (IMU) and encoder ports. These sensors enable the system to measure its current state. The goal is to find the optimal leaning …
Generalizable Skill Learning In Robotic Agents Using Transformer Models, Erik Enriquez
Generalizable Skill Learning In Robotic Agents Using Transformer Models, Erik Enriquez
Theses and Dissertations
This work explores the application of Transformer models to robotic skill learning, aiming to enhance generalization across various physical tasks and environments with continuous control. Despite their success in other domains, our experiments reveal that the utility of Transformers in robotics heavily depends on pretraining strategies. Specifically, Transformers pretrained on reinforcement learning tasks generalized effectively, while those trained with task-agnostic masking strategies did not. These findings challenge assumptions about the universality of Transformer-based methods and underscore the importance of domain-aligned pretraining for developing versatile robotic agents.
Real-Time Rectifying Flight Control Misconfiguration Using Intelligent Agent, Ruidong Han, Shangzhi Xu, Juanru Li, Elisa Bertino, David Lo, Jianfeng Ma, Siqi Ma
Real-Time Rectifying Flight Control Misconfiguration Using Intelligent Agent, Ruidong Han, Shangzhi Xu, Juanru Li, Elisa Bertino, David Lo, Jianfeng Ma, Siqi Ma
Research Collection School Of Computing and Information Systems
Configurations are supported by most flight control systems, allowing users to control a flying drone adapted to complexities such as environmental changes or mission alterations. Such an advanced functionality also introduces a significant problem—misconfiguration settings. It may cause drone instability, threaten drone safety, and potentially lead to substantial financial loss. However, detecting and rectifying misconfigurations across different flight control systems is challenging because (1) (mis)configuration-related code snippets might be syntactically correct and thus hard to identify through traditional code analysis; (2) the response to each configuration varies under different flying scenarios.In this article, we propose and implement a novel rectification …
Signal Timing Optimization Via Reinforcement Learning With Traffic Flow Prediction, Ming Xu, Jinye Li, Dongyu Zuo, Jing Zhang
Signal Timing Optimization Via Reinforcement Learning With Traffic Flow Prediction, Ming Xu, Jinye Li, Dongyu Zuo, Jing Zhang
Journal of System Simulation
Abstract: In response to the existing reinforcement learning-based traffic signal control methods that do not consider the changing trends in traffic flow, leading to congestion and inability to adapt to complex and variable road conditions, we propose a traffic signal timing optimization reinforcement learning method based on flow prediction. A phase timing amplitude control model is introduced. This model analyzes the spatiotemporal characteristics of historical traffic data to predict the flow for the next time slot and calculates a reasonable range for phase timing based on the prediction results. The H-PPO algorithm is employed to control the signal phase while …
Graph-Assisted Offline-Online Deep Reinforcement Learning For Dynamic Workflow Scheduling, Yifan Yang, Gang Chen, Hui Ma, Cong Zhang, Zhiguang Cao, Mengjie Zhang
Graph-Assisted Offline-Online Deep Reinforcement Learning For Dynamic Workflow Scheduling, Yifan Yang, Gang Chen, Hui Ma, Cong Zhang, Zhiguang Cao, Mengjie Zhang
Research Collection School Of Computing and Information Systems
Dynamic workflow scheduling (DWS) in cloud computing presents substantial challenges due to heterogeneous machine configurations, unpredictable workflow arrivals/patterns, and constantly evolving environments. However, existing research often assumes homogeneous setups and static conditions, limiting flexibility and adaptability in real-world scenarios. In this paper, we propose a novel Graph assisted Offline-Online Deep Reinforcement Learning (GOODRL) approach to building an effective and efficient scheduling agent for DWS. Our approach features three key innovations: (1) a task-specific graph representation and a Graph Attention Actor Network that enable the agent to dynamically assign focused tasks to heterogeneous machines while explicitly considering the future impact of …
On Minimizing Adversarial Counterfactual Error In Adversarial Reinforcement Learning, Roman Belaire, Arunesh Sinha, Pradeep Varakantham
On Minimizing Adversarial Counterfactual Error In Adversarial Reinforcement Learning, Roman Belaire, Arunesh Sinha, Pradeep Varakantham
Research Collection School Of Computing and Information Systems
Deep Reinforcement Learning (DRL) policies are highly susceptible to adversarial noise in observations, which poses significant risks in safety-critical scenarios. The challenge inherent to adversarial perturbations is that by altering the information observed by the agent, the state becomes only partially observable. Existing approaches address this by either enforcing consistent actions across nearby states or maximizing the worst-case value within adversarially perturbed observations. However, the former suffers from performance degradation when attacks succeed, while the latter tends to be overly conservative, leading to suboptimal performance in benign settings. We hypothesize that these limitations stem from their failing to account for …
Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang
Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang
Research Collection School Of Computing and Information Systems
With the widespread adoption of Internet Protocol (IP) communication technology and web-based platforms, cloud manufacturing has become a significant hallmark of Industry 4.0. Integrating graph algorithms into these web-enabled environments is crucial as they facilitate the representation and analysis of complex relationships in manufacturing processes, enabling efficient decision-making and adaptability in dynamic environments. As a key scheduling problem in cloud manufacturing, the flexible job-shop scheduling problem (FJSP) finds extensive applications in real-world scenarios. However, traditional FJSP-solving methods struggle to meet the efficiency and adaptability demands of cloud manufacturing due to generalization issues and excessive computational time, while reinforcement learning-based methods …
Exploring & Exploiting High-Order Graph Structure For Sparse Knowledge Graph Completion, Tao He, Ming Liu, Yixin Cao, Zekun Wang, Zihao Zheng, Bing Qin
Exploring & Exploiting High-Order Graph Structure For Sparse Knowledge Graph Completion, Tao He, Ming Liu, Yixin Cao, Zekun Wang, Zihao Zheng, Bing Qin
Research Collection School Of Computing and Information Systems
Sparse Knowledge Graph (KG) scenarios pose a challenge for previous Knowledge Graph Completion (KGC) methods, that is, the completion performance decreases rapidly with the increase of graph sparsity. This problem is also exacerbated because of the widespread existence of sparse KGs in practical applications. To alleviate this challenge, we present a novel framework, LR-GCN, that is able to automatically capture valuable long-range dependency among entities to supplement insufficient structure features and distill logical reasoning knowledge for sparse KGC. The proposed approach comprises two main components: a GNN-based predictor and a reasoning path distiller. The reasoning path distiller explores high-order graph …
Balancing Anarchy And Efficiency: Partial Team Formations And Learning In Potential Games, Muhammed Sayin
Balancing Anarchy And Efficiency: Partial Team Formations And Learning In Potential Games, Muhammed Sayin
Turkish Journal of Electrical Engineering and Computer Sciences
Non-cooperative multi-agent learning, focusing on individual rationality (anarchy), often falls short in achieving system-wide efficiency in potential games, a class of games with applications in decentralized control and optimization. On the other hand, cooperative approaches prioritize system efficiency but often via global coordination, which could be impractical, e.g., for large-scale and less controlled environments. To address this dilemma, we propose a novel framework that introduces partial team formations, allowing team members with shared objectives to coordinate their actions while maintaining team-wise rationality for improved system-wide efficiency without the burden of global coordination. We model such interactions as a multi-team game …
Improved Optimal Tracking Of Uncertain Nonlinear Discrete-Time Systems Using Experience Replay, Maxwell Geiger, Sarangapani Jagannathan
Improved Optimal Tracking Of Uncertain Nonlinear Discrete-Time Systems Using Experience Replay, Maxwell Geiger, Sarangapani Jagannathan
Electrical and Computer Engineering Faculty Research & Creative Works
This paper addresses the infinite horizon optimal tracking control problem for partially uncertain control-affine nonlinear discrete-time (DT) systems, where the control input dynamics are known. Multi-layer critic and actor neural networks (MNNs) are utilized for online estimation of the infinite horizon value function and optimal control input. The NN weights are tuned online using a direct temporal difference error (TDE)-driven learning approach, which modifies the singular values of the gradient with respect to the NN weights to accelerate their convergence. The critic NN uses a novel experience replay technique to improve sample efficiency without introducing biased TDEs and guarantee the …
Reinforcement Learning-Based Nonlinear Optimal Discrete-Time Control Of Power Systems, Vijay Kumar Singh, Behzad Farzanegan, S. Jagannathan
Reinforcement Learning-Based Nonlinear Optimal Discrete-Time Control Of Power Systems, Vijay Kumar Singh, Behzad Farzanegan, S. Jagannathan
Electrical and Computer Engineering Faculty Research & Creative Works
This paper presents a partially model-free adaptive optimal tracking control method for power systems, specifically targeting a synchronous generator connected through a reactive transmission line. By integrating the tracking error dynamics with reference trajectory dynamics, an augmented system is created. A discounted performance function is introduced to address the nonlinear tracking problem optimally. Unlike traditional methods that compute feedforward and feedback terms separately, the proposed approach calculates both simultaneously by minimizing the discounted performance function. The discrete-time tracking Bellman and Hamilton-Jacobi-Bellman (HJB) equations are derived, and a reinforcement learning (RL)-based technique is employed to solve the optimal policy online without …
Anti-Jamming Attack Mixed Strategy For Formation Tracking Control Via Game-Theoretical Reinforcement Learning, Lei Xue, Bei Ma, Yongbao Wu, Jian Liu, Chaoxu Mu, Donald C. Wunsch
Anti-Jamming Attack Mixed Strategy For Formation Tracking Control Via Game-Theoretical Reinforcement Learning, Lei Xue, Bei Ma, Yongbao Wu, Jian Liu, Chaoxu Mu, Donald C. Wunsch
Electrical and Computer Engineering Faculty Research & Creative Works
Communication plays a role in multi-UAV to perform formation tracking missions. In complex environments, UAV communication is often subject to jamming attacks, affecting the formation process. Therefore, studying the formation tracking control problem in jamming attacks is of great significance. Typically, the actions of the UAV consist of two fundamental modules: mobility strategy and communication strategy. In this paper, we design an anti-jamming attack mixed strategy for formation tracking control of the multi-UAV system. In practical scenarios, multi-UAV systems not only require the accomplishment of formation maneuvers but also necessitate effective mitigation of jamming attacks caused by other UAVs. Therefore, …
Harnessing The Power Of Gradient-Based Simulations For Multi-Objective Optimization In Particle Accelerators, Kishansingh Rajput, Malachi Schram, Auralee Edelen, Jonathan Colen, Armen Kasparian, Ryan Roussel, Adam Carpenter, He Zhang, Jay Benesch
Harnessing The Power Of Gradient-Based Simulations For Multi-Objective Optimization In Particle Accelerators, Kishansingh Rajput, Malachi Schram, Auralee Edelen, Jonathan Colen, Armen Kasparian, Ryan Roussel, Adam Carpenter, He Zhang, Jay Benesch
Data Science Faculty Publications
Particle accelerator operation requires simultaneous optimization of multiple objectives. Multi-objective optimization (MOO) is particularly challenging due to trade-offs between the objectives. Evolutionary algorithms, such as genetic algorithms (GAs), have been leveraged for many optimization problems, however, they do not apply to complex control problems by design. This paper demonstrates the power of differentiability for solving MOO problems in particle accelerators using a deep differentiable reinforcement learning (DDRL) algorithm. We compare the DDRL algorithm with model-free reinforcement learning (MFRL), GA, and Bayesian optimization (BO) for simultaneous optimization of heat load and trip rates in the continuous electron beam accelerator facility. The …
Efficient Task Scheduling In Cloud Infrastructures Using Dynamic Score-Based Allocation And Deep Q-Learning, Shadman Sakib
Efficient Task Scheduling In Cloud Infrastructures Using Dynamic Score-Based Allocation And Deep Q-Learning, Shadman Sakib
Graduate Theses/Dissertations
Cloud computing has grown rapidly in recent years, mainly due to the sharp increase in data transferred over the internet. This growth makes task scheduling a key and challenging part of cloud systems, as it helps distribute user requests across servers to minimize response time, prevent overloading, and ensure smooth user experience. This thesis proposes two novel approaches for dynamic task scheduling in cloud environments. First, a novel Score-Based Dynamic Load Balancing (SBDLB) strategy is developed, which leverages system parameters to allocate tasks efficiently across virtual machines (VMs) in data centers. SBDLB ensures balanced workload distribution by continuously evaluating VM …
Dataset Generation For Routing Policy Study In Ad Hoc Wireless Networks, Vishnu Vishnu Priya
Dataset Generation For Routing Policy Study In Ad Hoc Wireless Networks, Vishnu Vishnu Priya
Browse all Theses and Dissertations
Ad Hoc wireless networks, with their decentralized architecture and dynamic topology, present challenges in reliable and energy-efficient routing. While machine learning (ML) and reinforcement learning (RL) offer promising solutions, progress is limited by the lack of realistic, high-fidelity datasets. This research introduces a simulation-based framework for generating four diverse datasets representing combinations of node mobility (mobile vs. static) and spatial distribution (random vs. clustered). Each dataset captures critical metrics such as Signal-to-Interference-plus-Noise Ratio (SINR), bottleneck rate, and power consumption across multi-hop paths. A lookahead-based greedy routing algorithm with scenario-aware power control is implemented to emulate practical behavior. Supervised ML models, …
Research On Digital Twin Simulation Method Of Industrial Robot Integrated With Reinforcement Learning, Tianyue Miao, Lu Wang, Jiaxiao He, Nenggang Xie
Research On Digital Twin Simulation Method Of Industrial Robot Integrated With Reinforcement Learning, Tianyue Miao, Lu Wang, Jiaxiao He, Nenggang Xie
Journal of System Simulation
Abstract: In response to the lack of comprehensive functionality and limited application scenarios in the current field of industrial robot digital twin systems, which results in low versatility, a method for constructing a digital twin system for industrial robots with high versatility is proposed. A four-dimensional system architecture for the digital twin is designed, and the components and functions of the four-dimensional system are analyzed, based on the system level planning of the four-dimensional system, the concept of integrating reinforcement learning into the virtual replacement of real concept is defined. By constructing a multi-attribute virtual model and using TCP communication …
Research On Scheduling Strategies Simulation For Building Air-Conditioning Systems Based On Transfer Imitation Learning, Qiaochu Wang, Yan Ding, Chuanzhi Liang, Haozheng Zhang, Chen Huang
Research On Scheduling Strategies Simulation For Building Air-Conditioning Systems Based On Transfer Imitation Learning, Qiaochu Wang, Yan Ding, Chuanzhi Liang, Haozheng Zhang, Chen Huang
Journal of System Simulation
Abstract: To solve the problem of unstable performance and inefficient training process of low-quality data conditions at the initial stage of online deployment of air conditioner scheduling, we propose a migration-imitation learning-based air conditioning scheduling strategy simulation method. Reinforcement learning methods are used to generate building operation strategies. A standard building simulation model serves as the source domain, upon which migration learning is applied. An imitation learning loss function is incorporated into the intelligent loss function to enhance algorithm performance. The results indicate that, compared with the non-use of migration learning, the proposed method can improve the operational efficiency by …
Q-Learning In Starclash, Hanani Pankaj
Q-Learning In Starclash, Hanani Pankaj
2024 Fall Honors Capstone Projects - Archive
Developers create video games using Artificial Intelligence (AI) agents to provide a challenging opponent in a single-player game. However, studies show that when Reinforcement Learning (RL) agents are used, they outperform the AI agents. This project sought to test how RL agents would perform in StarClash, a video game without RL agents, using Q-Learning. This was done by creating two Q-Learning agents: a Simple agent and an Advanced (more complex) agent. These two agents were tested against each other and a Random AI agent. As expected, the Advanced agent did better than the Simple agent but only performed slightly better, …
Proof Of Concept: Simulating Drone Tracking In A Border Security Context, Jose Ruben Espinoza
Proof Of Concept: Simulating Drone Tracking In A Border Security Context, Jose Ruben Espinoza
Theses and Dissertations
Worldwide availability of drone technology has risen to unprecedented levels within the past century due to its commercial availability. While there has been various positive applications of such technology, it has additionally found usage within security critical contexts. Specifically, there have been reports of illegal drug smuggling along the Mexico-United States border in which quadrocopter based drones have been used. Within our research we aim to showcase, as a proof of concept, that autonomous drone technology can be leveraged within a defensive approach via the usage of reinforcement learning and object detection for security critical contexts. To promote the importance …
Reinforcement Learning Based Online Request Scheduling Framework For Workload-Adaptive Edge Deep Learning Inference, Xinrui Tan, Hongjia Li, Xiaofei Xie, Lu Guo, Nirwan Ansari, Xueqing Huang, Liming Wang, Zhen Xu, Yang Liu
Reinforcement Learning Based Online Request Scheduling Framework For Workload-Adaptive Edge Deep Learning Inference, Xinrui Tan, Hongjia Li, Xiaofei Xie, Lu Guo, Nirwan Ansari, Xueqing Huang, Liming Wang, Zhen Xu, Yang Liu
Research Collection School Of Computing and Information Systems
The recent advances of deep learning in various mobile and Internet-of-Things applications, coupled with the emergence of edge computing, have led to a strong trend of performing deep learning inference on the edge servers located physically close to the end devices. This trend presents the challenge of how to meet the quality-of-service requirements of inference tasks at the resource-constrained network edge, especially under variable or even bursty inference workloads. Solutions to this challenge have not yet been reported in the related literature. In the present paper, we tackle this challenge by means of workload-adaptive inference request scheduling: in different workload …
Sprinql : Sub-Optimal Demonstrations Driven Offline Imitation Learning, Minh Huy Hoang, Tien Mai, Pradeep Varakantham
Sprinql : Sub-Optimal Demonstrations Driven Offline Imitation Learning, Minh Huy Hoang, Tien Mai, Pradeep Varakantham
Research Collection School Of Computing and Information Systems
We focus on offline imitation learning (IL), which aims to mimic an expert's behavior using demonstrations without any interaction with the environment. One of the main challenges in offline IL is the limited support of expert demonstrations, which typically cover only a small fraction of the state-action space. While it may not be feasible to obtain numerous expert demonstrations, it is often possible to gather a larger set of sub-optimal demonstrations. For example, in treatment optimization problems, there are varying levels of doctor treatments available for different chronic conditions. These range from treatment specialists and experienced general practitioners to less …
End-To-End Motion Planning Of Unmanned Vehicles Based On Multimodal Deep Reinforcement Learning, Kaiyuan Ding, Askar Hamdulla, Bin Zhu, Eksan Firkat, Zhengtang Ma
End-To-End Motion Planning Of Unmanned Vehicles Based On Multimodal Deep Reinforcement Learning, Kaiyuan Ding, Askar Hamdulla, Bin Zhu, Eksan Firkat, Zhengtang Ma
Journal of System Simulation
Abstract: Since the agent cannot sense the surrounding environment and cannot successfully avoid obstacles, reinforcement learning fails to be generalized to robot motion planning in difficult terrain. Therefore, a solution based on multimodal deep reinforcement learning, which learns to blend proprioceptive states with high-dimensional depth sensor inputs, is proposed for the motion planning of unmanned vehicles. To be specific, proprioceptive states offer contact measurement for immediate reaction, and the unmanned vehicle can learn and forecast environmental changes with its attached visual sensors, proactively navigating around obstacles and uneven terrains numerous time steps ahead. TransProAct (transformer-based proactive action), a unique end-to-end …
Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An
Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An
Research Collection School Of Computing and Information Systems
Nash Equilibrium (NE) is the canonical solution concept of game theory, which provides an elegant tool to understand the rationalities. Though mixed strategy NE exists in any game with finite players and actions, computing NE in two- or multi-player general-sum games is PPAD-Complete. Various alternative solutions, e.g., Correlated Equilibrium (CE), and learning methods, e.g., fictitious play (FP), are proposed to approximate NE. For convenience, we call these methods as "inexact solvers", or "solvers" for short. However, the alternative solutions differ from NE and the learning methods generally fail to converge to NE. Therefore, in this work, we propose REinforcement Nash …
Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang
Macrohft : Memory Augmented Context-Aware Reinforcement Learning On High Frequency Trading, Chuqiao Zong, Chaojie Wang, Molei Qin, Lei Feng, Xinrun Wang, Xinrun Wang
Research Collection School Of Computing and Information Systems
High-frequency trading (HFT) that executes algorithmic trading in short time scales, has recently occupied the majority of cryptocurrency market. Besides traditional quantitative trading methods, reinforcement learning (RL) has become another appealing approach for HFT due to its terrific ability of handling high-dimensional financial data and solving sophisticated sequential decision-making problems, e.g., hierarchical reinforcement learning (HRL) has shown its promising performance on second-level HFT by training a router to select only one sub-agent from the agent pool to execute the current transaction. However, existing RL methods for HFT still have some defects: 1) standard RL-based trading agents suffer from the overfitting …
Research On Learnable Wargame Agent Driven By Battle Scheme, Yifeng Sun, Zhi Li, Jiang Wu, Yubin Wang
Research On Learnable Wargame Agent Driven By Battle Scheme, Yifeng Sun, Zhi Li, Jiang Wu, Yubin Wang
Journal of System Simulation
Abstract: To enable the agent to cope with complex battle scenarios and objectives in wargame, a learnable wargame agent architecture driven by a battle scheme is proposed. By analyzing the "attachment characteristics" and "loose coupling characteristics" of the agent to wargame system, the learnable requirements of the agent are obtained. In the design of the agent framework, battle schemes are used to reduce the learning range of the agent. The finite state machine corresponds to the knowledge of the operational phase in the battle scheme, and the decision-making space of the agent is determined according to the framework of the …
Task Analysis Methods Based On Deep Reinforcement Learning, Xue Gong, Pengfei Peng, Li Rong, Yalian Zheng, Jun Jiang
Task Analysis Methods Based On Deep Reinforcement Learning, Xue Gong, Pengfei Peng, Li Rong, Yalian Zheng, Jun Jiang
Journal of System Simulation
Abstract: In response to the high coupling of task interaction and many influencing factors in task analysis, a task analysis method based on sequence decoupling and deep reinforcement learning (DRL) is proposed, which can achieve task decomposition and task sequence reconstruction under complex constraints. The method designs an environment for deep reinforcement learning based on task information interaction, while improving the SumTree algorithm based on the difference between the loss functions of the target network and the evaluation network, achieving the priority evaluation among tasks. The activation function operation mechanism is introduced into the deep reinforcement learning network, followed by …