Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Artificial Intelligence and Robotics (140)
- Engineering (102)
- Computer Engineering (57)
- Operations Research, Systems Engineering and Industrial Engineering (53)
- Numerical Analysis and Scientific Computing (47)
-
- Systems Science (34)
- Electrical and Computer Engineering (29)
- Databases and Information Systems (27)
- Theory and Algorithms (23)
- Social and Behavioral Sciences (14)
- OS and Networks (11)
- Public Affairs, Public Policy and Public Administration (9)
- Software Engineering (9)
- Transportation (9)
- Graphics and Human Computer Interfaces (7)
- Medicine and Health Sciences (7)
- Applied Mathematics (5)
- Business (5)
- Operational Research (5)
- Other Computer Sciences (5)
- Computer and Systems Architecture (4)
- Data Science (4)
- Information Security (4)
- Physics (4)
- Aerospace Engineering (3)
- Chemistry (3)
- Controls and Control Theory (3)
- Programming Languages and Compilers (3)
- Institution
-
- Singapore Management University (72)
- China Simulation Federation (34)
- Missouri University of Science and Technology (16)
- Brigham Young University (12)
- Air Force Institute of Technology (9)
-
- Old Dominion University (9)
- University of Texas at Arlington (9)
- TÜBİTAK (8)
- MBZUAI (7)
- New Jersey Institute of Technology (5)
- Portland State University (5)
- University of Texas Rio Grande Valley (5)
- Edith Cowan University (4)
- San Jose State University (4)
- University of Denver (4)
- Utah State University (4)
- Zayed University (3)
- Bucknell University (2)
- Chapman University (2)
- Georgia Southern University (2)
- Missouri State University (2)
- Nova Southeastern University (2)
- Technological University Dublin (2)
- University of Kentucky (2)
- University of Minnesota Morris Digital Well (2)
- University of Nebraska - Lincoln (2)
- University of Nevada, Las Vegas (2)
- University of South Carolina (2)
- Western Michigan University (2)
- California State University, San Bernardino (1)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (69)
- Journal of System Simulation (34)
- Theses and Dissertations (16)
- Electrical and Computer Engineering Faculty Research & Creative Works (12)
- Faculty Publications (9)
-
- Turkish Journal of Electrical Engineering and Computer Sciences (8)
- Dissertations (7)
- Machine Learning Faculty Publications (7)
- Electronic Theses and Dissertations (6)
- Computer Science and Engineering Dissertations - Archive (5)
- All Graduate Theses and Dissertations, Spring 1920 to Summer 2023 (4)
- Computer Science Faculty Publications (4)
- Computer Science Faculty Research & Creative Works (4)
- Dissertations and Theses (4)
- Master's Projects (4)
- All Works (3)
- Dissertations and Theses Collection (Open Access) (3)
- Research outputs 2022 to 2026 (3)
- Articles (2)
- CCAC Theses and Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science and Engineering Theses - Archive (2)
- Data Science Faculty Publications (2)
- Electrical & Computer Engineering Faculty Research (2)
- Engineering Management & Systems Engineering Faculty Publications (2)
- Graduate Theses/Dissertations (2)
- Journal of Undergraduate Research (2)
- Scholarly Horizons: University of Minnesota, Morris Undergraduate Journal (2)
- School of Computing: Dissertations, Theses, and Student Research (2)
- Theses and Dissertations--Computer Science (2)
- Publication Type
Articles 91 - 120 of 259
Full-Text Articles in Computer Sciences
Aircraft Assignment Method For Optimal Utilization Of Maintenance Intervals, Runxia Guo, Yifu Wang
Aircraft Assignment Method For Optimal Utilization Of Maintenance Intervals, Runxia Guo, Yifu Wang
Journal of System Simulation
Abstract: The aircraft assignment problem is studied from a maintenance assurance perspective. In order to ensure its continuous airworthiness, civil aircraft are required to perform maintenance tasks, i. e., scheduled inspections, at specified intervals. The scheduled inspection interval is usually controlled by the number of flight cycles (FC), flight hours (FH), or flight days (FD), whichever comes first. In order to make balanced use of the inspection interval, an aircraft assignment model for a given fleet size is developed to optimize the maintenance interval utilization, and it is solved by a reinforcement learning algorithm to minimize the variance of the …
Asynchronous Fdrl-Based Low-Latency Computation Offloading For Integrated Terrestrial And Non-Terrestrial Power Iot, Sifeng Li, Sunxuan Zhang, Zhao Wang, Zhenyu Zhou, Xiaoyan Wang, Shahid Mumtaz, Mohsen Guizani, Valerio Frascolla
Asynchronous Fdrl-Based Low-Latency Computation Offloading For Integrated Terrestrial And Non-Terrestrial Power Iot, Sifeng Li, Sunxuan Zhang, Zhao Wang, Zhenyu Zhou, Xiaoyan Wang, Shahid Mumtaz, Mohsen Guizani, Valerio Frascolla
Machine Learning Faculty Publications
Integrated terrestrial and non-terrestrial power internet of things (IPIoT) has emerged as a paradigm shift to three-dimensional vertical communication networks for power systems in the 6G era. Computation offloading plays key roles in enabling real-time data processing and analysis for electric services. However, computation offloading in IPIoT still faces challenges of coupling between task offloading and computation resource allocation, resource heterogeneity and dynamics, and degraded model training caused by electromagnetic interference (EMI). In this article, we propose an asynchronous federated deep reinforcement learning (AFDRL)-based computation offloading framework for IPIoT, where models are uploaded asynchronously for federated averaging to relieve network …
Reinforcement Learning Approach To Stochastic Vehicle Routing Problem With Correlated Demands, Zangir Iklassov, Ikboljon Sobirov, Ruben Solozabal, Martin Takac
Reinforcement Learning Approach To Stochastic Vehicle Routing Problem With Correlated Demands, Zangir Iklassov, Ikboljon Sobirov, Ruben Solozabal, Martin Takac
Machine Learning Faculty Publications
We present a novel end-to-end framework for solving the Vehicle Routing Problem with stochastic demands (VRPSD) using Reinforcement Learning (RL). Our formulation incorporates the correlation between stochastic demands through other observable stochastic variables, thereby offering an experimental demonstration of the theoretical premise that non-i.i.d. stochastic demands provide opportunities for improved routing solutions. Our approach bridges the gap in the application of RL to VRPSD and consists of a parameterized stochastic policy optimized using a policy gradient algorithm to generate a sequence of actions that form the solution. Our model outperforms previous state-of-the-art metaheuristics and demonstrates robustness to changes in the …
On-Line Environment Adaptation For User Performance Optimization, Subharag Sarkar
On-Line Environment Adaptation For User Performance Optimization, Subharag Sarkar
Computer Science and Engineering Dissertations - Archive
In today’s fast-paced and globally connected world, businesses are creating products with more significance to user personalization and customization. This has amplified the importance of capturing and learning user preferences as more information from users can lead to the designing and development of products that will improve user engagement and performance. Numerous algorithms based on collaborative filtering and recommender systems have been used to learn user preferences, but almost all of them require big datasets to train on. This creates a dependency on collecting more and more user information which might lead to ethical considerations and privacy concerns. To solve …
Transferable Curricula Through Difficulty Conditioned Generators, Sidney Tio, Pradeep Varakantham
Transferable Curricula Through Difficulty Conditioned Generators, Sidney Tio, Pradeep Varakantham
Research Collection School Of Computing and Information Systems
Advancements in reinforcement learning (RL) have demonstrated superhuman performance in complex tasks such as Starcraft, Go, Chess etc. However, knowledge transfer from Artificial "Experts" to humans remain a significant challenge. A promising avenue for such transfer would be the use of curricula. Recent methods in curricula generation focuses on training RL agents efficiently, yet such methods rely on surrogate measures to track student progress, and are not suited for training robots in the real world (or more ambitiously humans). In this paper, we introduce a method named Parameterized Environment Response Model (PERM) that shows promising results in training RL agents …
Reinforcement Learning For Sequential Decision Making With Constraints, Jiajing Ling
Reinforcement Learning For Sequential Decision Making With Constraints, Jiajing Ling
Dissertations and Theses Collection (Open Access)
Reinforcement learning is a widely used approach to tackle problems in sequential decision making where an agent learns from rewards or penalties. However, in decision-making problems that involve safety or limited resources, the agent's exploration is often limited by constraints. To model such problems, constrained Markov decision processes and constrained decentralized partially observable Markov decision processes have been proposed for single-agent and multi-agent settings, respectively. A significant challenge in solving constrained Dec-POMDP is determining the contribution of each agent to the primary objective and constraint violations. To address this issue, we propose a fictitious play-based method that uses Lagrangian Relaxation …
Multi-View Hypergraph Contrastive Policy Learning For Conversational Recommendation, Sen Zhao, Wei Wei, Xian-Ling Mao, Shuai: Yang Zhu, Zujie Wen, Dangyang Chen, Feida Zhu, Feida Zhu
Multi-View Hypergraph Contrastive Policy Learning For Conversational Recommendation, Sen Zhao, Wei Wei, Xian-Ling Mao, Shuai: Yang Zhu, Zujie Wen, Dangyang Chen, Feida Zhu, Feida Zhu
Research Collection School Of Computing and Information Systems
Conversational recommendation systems (CRS) aim to interactively acquire user preferences and accordingly recommend items to users. Accurately learning the dynamic user preferences is of crucial importance for CRS. Previous works learn the user preferences with pairwise relations from the interactive conversation and item knowledge, while largely ignoring the fact that factors for a relationship in CRS are multiplex. Specifically, the user likes/dislikes the items that satisfy some attributes (Like/Dislike view). Moreover social influence is another important factor that affects user preference towards the item (Social view), while is largely ignored by previous works in CRS. The user preferences from these …
Imitation Improvement Learning For Large-Scale Capacitated Vehicle Routing Problems, The Viet Bui, Tien Mai
Imitation Improvement Learning For Large-Scale Capacitated Vehicle Routing Problems, The Viet Bui, Tien Mai
Research Collection School Of Computing and Information Systems
Recent works using deep reinforcement learning (RL) to solve routing problems such as the capacitated vehicle routing problem (CVRP) have focused on improvement learning-based methods, which involve improving a given solution until it becomes near-optimal. Although adequate solutions can be achieved for small problem instances, their efficiency degrades for large-scale ones. In this work, we propose a newimprovement learning-based framework based on imitation learning where classical heuristics serve as experts to encourage the policy model to mimic and produce similar or better solutions. Moreover, to improve scalability, we propose Clockwise Clustering, a novel augmented framework for decomposing large-scale CVRP into …
An Investigation Into Machine Learning Techniques For Designing Dynamic Difficulty Agents In Real-Time Games, Ryan Adare Dunagan
An Investigation Into Machine Learning Techniques For Designing Dynamic Difficulty Agents In Real-Time Games, Ryan Adare Dunagan
Electronic Theses and Dissertations
Video games are an incredibly popular pastime enjoyed by people of all ages world wide. Many different kinds of games exist, but most games feature some elements of the player overcoming some challenge, usually through gameplay. These challenges are insurmountable for some people and may turn them off to video games as a pastime. Games can be made more accessible to players of little skill and/or experience through the use of Dynamic Difficulty Adjustment (DDA) systems that adjust the difficulty of the game in response to the player’s performance. This research seeks to establish the effectiveness of machine learning techniques …
Dynamic Police Patrol Scheduling With Multi-Agent Reinforcement Learning, Songhan Wong, Waldy Joe, Hoong Chuin Lau
Dynamic Police Patrol Scheduling With Multi-Agent Reinforcement Learning, Songhan Wong, Waldy Joe, Hoong Chuin Lau
Research Collection School Of Computing and Information Systems
Effective police patrol scheduling is essential in projecting police presence and ensuring readiness in responding to unexpected events in urban environments. However, scheduling patrols can be a challenging task as it requires balancing between two conflicting objectives namely projecting presence (proactive patrol) and incident response (reactive patrol). This task is made even more challenging with the fact that patrol schedules do not remain static as occurrences of dynamic incidents can disrupt the existing schedules. In this paper, we propose a solution to this problem using Multi-Agent Reinforcement Learning (MARL) to address the Dynamic Bi-objective Police Patrol Dispatching and Rescheduling Problem …
Sim-To-Real Reinforcement Learning Framework For Autonomous Aerial Leaf Sampling, Ashraful Islam
Sim-To-Real Reinforcement Learning Framework For Autonomous Aerial Leaf Sampling, Ashraful Islam
School of Computing: Dissertations, Theses, and Student Research
Using unmanned aerial systems (UAS) for leaf sampling is contributing to a better understanding of the influence of climate change on plant species, and the dynamics of forest ecology by studying hard-to-reach tree canopies. Currently, multiple skilled operators are required for UAS maneuvering and using the leaf sampling tool. This often limits sampling to only the canopy top or periphery. Sim-to-real reinforcement learning (RL) can be leveraged to tackle challenges in the autonomous operation of aerial leaf sampling in the changing environment of a tree canopy. However, trans- ferring an RL controller that is learned in simulation to real UAS …
Neural Network Architecture Optimization Using Reinforcement Learning, Raghav Vadhera
Neural Network Architecture Optimization Using Reinforcement Learning, Raghav Vadhera
Computer Science and Engineering Dissertations - Archive
Deep learning has emerged as an increasingly valuable tool, employed across a myriad of applications. However, the intricacies of deep learning systems, stemming from their sensitivity to specific network architectures, have rendered them challenging for non-experts to harness, thus highlighting the need for automatic network architecture optimization. Prior research predominantly optimizes a network for a single problem through architecture search, necessitating extensive training of various architectures during optimization.\\ To tackle this issue and unlock the potential for transferability across tasks, this dissertation presents a groundbreaking approach that employs Reinforcement Learning to develop a network optimization policy based on an abstract …
Detecting Complex Cyber Attacks Using Decoys With Online Reinforcement Learning, Marcus Gutierrez
Detecting Complex Cyber Attacks Using Decoys With Online Reinforcement Learning, Marcus Gutierrez
Open Access Theses & Dissertations
Most vulnerabilities discovered in cybersecurity can be associated with their own singular piece of software. I investigate complex vulnerabilities, which may require multiple software to be present. These complex vulnerabilities represent 16.6% of all documented vulnerabilities and are more dangerous on average than their simple vulnerability counterparts. In addition to this, because they often require multiple pieces of software to be present, they are harder to identify overall as specific combinations are needed for the vulnerability to appear.
I consider the motivating scenario where an attacker is repeatedly deploying exploits that use complex vulnerabilities into an Airport Wi-Fi. The network …
Reinforced Adaptation Network For Partial Domain Adaptation, Keyu Wu, Min Wu, Zhenghua Chen, Ruibing Jin, Wei Cui, Zhiguang Cao, Xiaoli Li
Reinforced Adaptation Network For Partial Domain Adaptation, Keyu Wu, Min Wu, Zhenghua Chen, Ruibing Jin, Wei Cui, Zhiguang Cao, Xiaoli Li
Research Collection School Of Computing and Information Systems
Domain adaptation enables generalized learning in new environments by transferring knowledge from label-rich source domains to label-scarce target domains. As a more realistic extension, partial domain adaptation (PDA) relaxes the assumption of fully shared label space, and instead deals with the scenario where the target label space is a subset of the source label space. In this paper, we propose a Reinforced Adaptation Network (RAN) to address the challenging PDA problem. Specifically, a deep reinforcement learning model is proposed to learn source data selection policies. Meanwhile, a domain adaptation model is presented to simultaneously determine rewards and learn domain-invariant feature …
Multi-Agent Cooperative Combat Simulation In Naval Battlefield With Reinforcement Learning, Ding Shi, Xuefeng Yan, Lina Gong, Jingxuan Zhang, Donghai Guan, Mingqiang Wei
Multi-Agent Cooperative Combat Simulation In Naval Battlefield With Reinforcement Learning, Ding Shi, Xuefeng Yan, Lina Gong, Jingxuan Zhang, Donghai Guan, Mingqiang Wei
Journal of System Simulation
Abstract: Due to the rapidly-changed situations of future naval battlefields, it is urgent to realize the high-quality combat simulation in naval battlefields based on artificial intelligence to comprehensively optimize and improve the combat effectiveness of our army and defeat the enemy. The collaboration of combat units is the key point and how to realize the balanced decision-making among multiple agents is the first task. Based on decoupling priority experience replay mechanism and attention mechanism, a multi-agent reinforcement learning-based cooperative combat simulation (MARL-CCSA) network is proposed. Based on the expert experience, a multi-scale reward function is designed, on which a naval …
Research On Unmanned Swarm Combat System Adaptive Evolution Model Simulation, Zhiqiang Li, Yuanlong Li, Laixiang Yin, Xiangping Ma
Research On Unmanned Swarm Combat System Adaptive Evolution Model Simulation, Zhiqiang Li, Yuanlong Li, Laixiang Yin, Xiangping Ma
Journal of System Simulation
Abstract: Aiming at the fact that the intelligent unmanned swarm combat system is mainly composed of large-scale combat individuals with limited behavioral capabilities and has limited ability to adapt to the changes of battlefield environment and combat opponents, a learning evolution method combining genetic algorithm and reinforcement learning is proposed to construct an individual-based unmanned bee colony combat system evolution model. To improve the adaptive evolution efficiency of bee colony combat system, an improved genetic algorithm is proposed to improve the learning and evolution speed of bee colony individuals by using individual-specific mutation optimization strategy. Simulation experiment on …
A Reinforcement Learning Approach To A Beyond Visual Range Air Combat Maneuvering Problem, Caleb A. Taylor
A Reinforcement Learning Approach To A Beyond Visual Range Air Combat Maneuvering Problem, Caleb A. Taylor
Theses and Dissertations
A one-versus-one air combat maneuvering problem is considered wherein a friendly autonomous aircraft must engage and defeat an adversary autonomous aircraft in a beyond visual range environment. The Advanced Framework for Simulation, Integration, and Modeling (AFSIM) is leveraged to model the complex and interdependent operations of aircraft, sensors, and weapons utilized in beyond visual range air combat. We formulate a Markov decision process to obtain high-quality decision policies wherein our autonomous aircraft makes maneuvering and missile firing decisions. We utilize a reinforcement learning solution procedure that implements a linear value function approximation to represent state-decision pairs due to the high …
Dqn-Based Joint Scheduling Method Of Heterogeneous Tt&C Resources, Naiyang Xue, Dan Ding, Yutong Jia, Zhiqiang Wang, Yuan Liu
Dqn-Based Joint Scheduling Method Of Heterogeneous Tt&C Resources, Naiyang Xue, Dan Ding, Yutong Jia, Zhiqiang Wang, Yuan Liu
Journal of System Simulation
Abstract: Joint scheduling of heterogeneous TT&C resources as research object, a deep Q network (DQN) algorithm based on reinforcement learning is proposed. The characteristics of the joint scheduling problem of heterogeneous TT&C resources being fully analyzied and mathematical language being used to describe the constraints affecting the solution, a resource joint scheduling model is established. From the perspective of applying reinforcement learning, two neural networks with the same structure and the action selection strategies based onεgreedy algorithm are respectively designed after Markov decision process description, and DQN solution framework is established. The simulation results show that DQN-based heterogeneous …
Constrained Reinforcement Learning In Hard Exploration Problems, Pankayaraj Pathmanathan, Pradeep Varakantham
Constrained Reinforcement Learning In Hard Exploration Problems, Pankayaraj Pathmanathan, Pradeep Varakantham
Research Collection School Of Computing and Information Systems
One approach to guaranteeing safety in Reinforcement Learning is through cost constraints that are imposed on trajectories. Recent works in constrained RL have developed methods that ensure constraints can be enforced even at learning time while maximizing the overall value of the policy. Unfortunately, as demonstrated in our experimental results, such approaches do not perform well on complex multi-level tasks, with longer episode lengths or sparse rewards. To that end, wepropose a scalable hierarchical approach for constrained RL problems that employs backward cost value functions in the context of task hierarchy and a novel intrinsic reward function in lower levels …
Learning Feature Embedding Refiner For Solving Vehicle Routing Problems, Jingwen Li, Yining Ma, Zhiguang Cao, Yaoxin Wu, Wen Song, Jie Zhang, Yeow Meng Chee
Learning Feature Embedding Refiner For Solving Vehicle Routing Problems, Jingwen Li, Yining Ma, Zhiguang Cao, Yaoxin Wu, Wen Song, Jie Zhang, Yeow Meng Chee
Research Collection School Of Computing and Information Systems
While the encoder–decoder structure is widely used in the recent neural construction methods for learning to solve vehicle routing problems (VRPs), they are less effective in searching solutions due to deterministic feature embeddings and deterministic probability distributions. In this article, we propose the feature embedding refiner (FER) with a novel and generic encoder–refiner–decoder structure to boost the existing encoder–decoder structured deep models. It is model-agnostic that the encoder and the decoder can be from any pretrained neural construction method. Regarding the introduced refiner network, we design its architecture by combining the standard gated recurrent units (GRU) cell with two new …
Malbot-Drl: Malware Botnet Detection Using Deep Reinforcement Learning In Iot Networks, Mohammad Al-Fawa'reh, Jumana Abu-Khalaf, Patryk Szewczyk, James J. Kang
Malbot-Drl: Malware Botnet Detection Using Deep Reinforcement Learning In Iot Networks, Mohammad Al-Fawa'reh, Jumana Abu-Khalaf, Patryk Szewczyk, James J. Kang
Research outputs 2022 to 2026
In the dynamic landscape of cyber threats, multi-stage malware botnets have surfaced as significant threats of concern. These sophisticated threats can exploit Internet of Things (IoT) devices to undertake an array of cyberattacks, ranging from basic infections to complex operations such as phishing, cryptojacking, and distributed denial of service (DDoS) attacks. Existing machine learning solutions are often constrained by their limited generalizability across various datasets and their inability to adapt to the mutable patterns of malware attacks in real world environments, a challenge known as model drift. This limitation highlights the pressing need for adaptive Intrusion Detection Systems (IDS), capable …
Optimizing Constraint Selection In A Design Verification Environment For Efficient Coverage Closure, Vanessa Cooper
Optimizing Constraint Selection In A Design Verification Environment For Efficient Coverage Closure, Vanessa Cooper
CCAC Theses and Dissertations
No abstract provided.
Continual Optimal Adaptive Tracking Of Uncertain Nonlinear Continuous-Time Systems Using Multilayer Neural Networks, Irfan Ganie, S. (Sarangapani) Jagannathan
Continual Optimal Adaptive Tracking Of Uncertain Nonlinear Continuous-Time Systems Using Multilayer Neural Networks, Irfan Ganie, S. (Sarangapani) Jagannathan
Electrical and Computer Engineering Faculty Research & Creative Works
This study provides a lifelong integral reinforcement learning (LIRL)-based optimal tracking scheme for uncertain nonlinear continuous-time (CT) systems using multilayer neural network (MNN). In this LIRL framework, the optimal control policies are generated by using both the critic neural network (NN) weights and single-layer NN identifier. The critic MNN weight tuning is accomplished using an improved singular value decomposition (SVD) of its activation function gradient. The NN identifier, on the other hand, provides the control coefficient matrix for computing the control policies. An online weight velocity attenuation (WVA)-based consolidation scheme is proposed wherein the significance of weights is derived by …
The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson
The Basil Technique: Bias Adaptive Statistical Inference Learning Agents For Learning From Human Feedback, Jonathan Indigo Watson
Theses and Dissertations--Computer Science
We introduce a novel approach for learning behaviors using human-provided feedback that is subject to systematic bias. Our method, known as BASIL, models the feedback signal as a combination of a heuristic evaluation of an action's utility and a probabilistically-drawn bias value, characterized by unknown parameters. We present both the general framework for our technique and specific algorithms for biases drawn from a normal distribution. We evaluate our approach across various environments and tasks, comparing it to interactive and non-interactive machine learning methods, including deep learning techniques, using human trainers and a synthetic oracle with feedback distorted to varying degrees. …
Navigating Classic Atari Games With Deep Learning, Ayan Abhiranya Singh
Navigating Classic Atari Games With Deep Learning, Ayan Abhiranya Singh
Master's Projects
Games for the Atari 2600 console provide great environments for testing reinforcement learning algorithms. In reinforcement learning algorithms, an agent typically learns about its environment via the delivery of periodic rewards. Deep Q-Learning, a variant of Q-Learning, utilizes neural networks which train a Q-function to predict the highest future reward given an input state and action. Deep Q-learning has shown great results in training agents to play Atari 2600 games like Space Invaders and Breakout. However, Deep Q-Learning has historically struggled with learning how to play games with greater emphasis on exploration and delayed rewards, like Ms. PacMan. In this …
Moment-Based Reinforcement Learning For Ensemble Control, Yao-Chi Yu, Vignesh Narayanan, Jr-Shin Li
Moment-Based Reinforcement Learning For Ensemble Control, Yao-Chi Yu, Vignesh Narayanan, Jr-Shin Li
Publications
Problems involving controlling the collective behavior of a population of structurally similar dynamical systems, the so-called ensemble control, arise in diverse emerging applications and pose a grand challenge in systems science and control engineering. Owing to the severely under-actuated nature and the difficulty of placing large-scale sensor networks, ensemble systems are limited to being actuated and monitored at the population level. Moreover, mathematical models describing the dynamics of ensemble systems are often elusive. Therefore, it is essential to design broadcast controls that excite the entire population in such a way that the heterogeneity in system dynamics are robustly compensated. In …
Reinforcement Learning Enhanced Pichunter For Interactive Search, Zhixin Ma, Jiaxin Wu, Weixiong Loo, Chong-Wah Ngo
Reinforcement Learning Enhanced Pichunter For Interactive Search, Zhixin Ma, Jiaxin Wu, Weixiong Loo, Chong-Wah Ngo
Research Collection School Of Computing and Information Systems
With the tremendous increase in video data size, search performance could be impacted significantly. Specifically, in an interactive system, a real-time system allows a user to browse, search and refine a query. Without a speedy system quickly, the main ingredient to engage a user to stay focused, an interactive system becomes less effective even with a sophisticated deep learning system. This paper addresses this challenge by leveraging approximate search, Bayesian inference, and reinforcement learning. For approximate search, we apply a hierarchical navigable small world, which is an efficient approximate nearest neighbor search algorithm. To quickly prune the search scope, we …
Intelligent Adaptive Gossip-Based Broadcast Protocol For Uav-Mec Using Multi-Agent Deep Reinforcement Learning, Zen Ren, Xinghua Li, Yinbin Miao, Zhuowen Li, Zihao Wang, Mengyao Zhu, Ximeng Liu, Deng, Robert H.
Intelligent Adaptive Gossip-Based Broadcast Protocol For Uav-Mec Using Multi-Agent Deep Reinforcement Learning, Zen Ren, Xinghua Li, Yinbin Miao, Zhuowen Li, Zihao Wang, Mengyao Zhu, Ximeng Liu, Deng, Robert H.
Research Collection School Of Computing and Information Systems
UAV-assisted mobile edge computing (UAV-MEC) has been proposed to offer computing resources for smart devices and user equipment. UAV cluster aided MEC rather than one UAV-aided MEC as edge pool is the newest edge computing architecture. Unfortunately, the data packet exchange during edge computing within the UAV cluster hasn't received enough attention. UAVs need to collaborate for the wide implementation of MEC, relying on the gossip-based broadcast protocol. However, gossip has the problem of long propagation delay, where the forwarding probability and neighbors are two factors that are difficult to balance. The existing works improve gossip from only one factor, …
End-To-End Hierarchical Reinforcement Learning With Integrated Subgoal Discovery, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan, Chai Quek
End-To-End Hierarchical Reinforcement Learning With Integrated Subgoal Discovery, Shubham Pateria, Budhitama Subagdja, Ah-Hwee Tan, Chai Quek
Research Collection School Of Computing and Information Systems
Hierarchical reinforcement learning (HRL) is a promising approach to perform long-horizon goal-reaching tasks by decomposing the goals into subgoals. In a holistic HRL paradigm, an agent must autonomously discover such subgoals and also learn a hierarchy of policies that uses them to reach the goals. Recently introduced end-to-end HRL methods accomplish this by using the higher-level policy in the hierarchy to directly search the useful subgoals in a continuous subgoal space. However, learning such a policy may be challenging when the subgoal space is large. We propose integrated discovery of salient subgoals (LIDOSS), an end-to-end HRL method with an integrated …
Reinforcement-Learning-Based Adaptive Tracking Control For A Space Continuum Robot Based On Reinforcement Learning, Da Jiang, Zhiqin Cai, Zhongzhen Liu, Haijun Peng, Zhigang Wu
Reinforcement-Learning-Based Adaptive Tracking Control For A Space Continuum Robot Based On Reinforcement Learning, Da Jiang, Zhiqin Cai, Zhongzhen Liu, Haijun Peng, Zhigang Wu
Journal of System Simulation
Abstract: Aiming at the tracking control for three-arm space continuum robot in space active debris removal manipulation, an adaptive sliding mode control algorithm based on deep reinforcement learning is proposed. Through BP network, a data-driven dynamic model is developed as the predictive model to guide the reinforcement learning to adjust the sliding mode controller's parameters online, and finally realize a real-time tracking control. Simulation results show that the proposed data-driven predictive model can accurately predict the robot's dynamic characteristics with the relative error within ±1% to random trajectories. Compared with the fixed-parameter sliding mode controller, the proposed adaptive controller …