Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Physical Sciences and Mathematics (7)
- Computer Sciences (5)
- Artificial Intelligence and Robotics (4)
- Aviation (3)
- Business (1)
-
- Business Intelligence (1)
- Civil and Environmental Engineering (1)
- Computer Engineering (1)
- Data Science (1)
- Design of Experiments and Sample Surveys (1)
- E-Commerce (1)
- Industrial Engineering (1)
- Management Information Systems (1)
- Management Sciences and Quantitative Methods (1)
- Operations and Supply Chain Management (1)
- Statistics and Probability (1)
- Transportation Engineering (1)
- Institution
- Publication
- Publication Type
Articles 1 - 12 of 12
Full-Text Articles in Operational Research
Learning To Dogfight: Proximal Policy Optimization Vs. Double Deep Q Network For 2v2 Air Combat With Directed Energy Weapons In Afsim, Caden W. Wilson
Learning To Dogfight: Proximal Policy Optimization Vs. Double Deep Q Network For 2v2 Air Combat With Directed Energy Weapons In Afsim, Caden W. Wilson
Theses and Dissertations
This research utilizes reinforcement learning (RL) to train two blue agents each imbued with a directed energy weapon (DEW) in a 2v2 within visual range air combat maneuvering problem. A phased solution approach is employed to repeatedly tune and train several RL algorithm implementations: Proximal Policy Optimization (PPO) and Double Deep Q Network (DDQN). Phase I of training includes reward shaping for basic flight elements such as altitude, airspeed, and target proximity. Phase II of training builds off policies developed in Phase I, but rewards emphasize winning the aerial engagement by any means necessary. DDQN significantly outperforms PPO in Phase …
Reinforcement Learning For Aeromedical Evacuation In Nonstationary Combat Environments, Micah J. Kartchner
Reinforcement Learning For Aeromedical Evacuation In Nonstationary Combat Environments, Micah J. Kartchner
Theses and Dissertations
This research formulates the medical evacuation (MEDEVAC) dispatching problem as a sequential decision process and investigates the application of reinforcement learning under nonstationary conditions. We model the dynamic arrival rate of MEDEVAC requests using a nonstationary Hawkes process and design a Double Deep Q-Network algorithm that incorporates belief states to anticipate future requests. Through computational experimentation, we analyze the impact of belief formulation on decision quality and system performance. Results indicate that policies incorporating belief states significantly outperform myopic dispatching policies, reducing urgent casualty wait times by up to 49.68% and increasing on-time evacuations by up to 21.91%.
Proximal Policy Optimization Applied To The Beyond Visual Range Air Combat Maneuvering Problem, Daniel B. Joseph
Proximal Policy Optimization Applied To The Beyond Visual Range Air Combat Maneuvering Problem, Daniel B. Joseph
Theses and Dissertations
Artificial intelligence (AI) grows ever-more important in warfighting. Emerging technologies allow for the use of AI to control aircraft and weapons systems. This research investigates the application of reinforcement learning (RL) through the Proximal Policy Optimization (PPO) algorithm to a two-versus-two (2v2) beyond-visual-range (BVR) air combat maneuvering problem (ACMP). Implemented in the Advanced Framework for Simulation, Integration, and Modeling (AFSIM), the methodology frames the engagement as a Markov decision process, wherein an autonomous RL agent learns continuous control decisions—throttle, pitch, roll, and yaw—under a cooperative communication scheme. A multi-phase curriculum-learning approach facilitates the progressive acquisition of flight stability, weapon deployment, …
Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs
Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs
Theses and Dissertations
Leveraging the Advanced Framework for Simulation, Integration, and Modeling (AFSIM) we investigate the use of reinforcement learning (RL) techniques for imbuing AUCAV agents with high-quality behaviors for the within-visual-range air combat maneuvering problem (ACMP). We formulate the 2v2 WVR ACMP as a Markov decision process wherein friendly AUCAVs are equipped with DEW capabilities and operate with 6 degrees of freedom. We utilize the Double Deep Q-Network RL algorithm, which centrally trains two friendly AUCAVs and employ a phased learning approach, initially exposing the AUCAVs to a dense reward environment for early training, followed by a sparse reward environment to encourage …
A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae
A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae
Theses and Dissertations
A growing demand exists for interpretable artificial intelligence models, leading to extensive research efforts to enhance the explainability and transparency of policies generated by reinforcement learning (RL) methods. This research develops random forest-based RL algorithms as a logical progression in this academic pursuit. The algorithms are evaluated using three standard benchmark environments from OpenAI gym — CartPole, MountainCar, and LunarLander — and compared to implementations of the Deep Q-learning Network (DQN) and Double DQN (DDQN) algorithms for various metrics, including performance, robustness, efficiency, and interpretability. The random forest-based algorithms exhibit superior performance to both neural network-based algorithms in two out …
A Reinforcement Learning Approach To The 2v2 Beyond Visual Range Air Combat Maneuvering Problem, Jacob J. Pike
A Reinforcement Learning Approach To The 2v2 Beyond Visual Range Air Combat Maneuvering Problem, Jacob J. Pike
Theses and Dissertations
This research examines a 2v2 air combat maneuvering problem (ACMP) in a Beyond Visual Range (BVR) environment. A discrete-time, infinite-horizon Markov Decision Process (MDP) model represents the BVR-ACMP, seeking to determine high-quality policies for a pair of autonomous aircraft to execute tactical maneuvers and firing decisions. The Advanced Framework for Simulation, Integration, and Modeling (AFSIM) characterizes the complex six-degree of freedom (6-DOF) aircraft operations, encompassing kinematics, sensors, and weapons. Given the high dimensionality and continuous nature of the state and decision variables, a deep reinforcement learning (RL) solution approach is adopted wherein the value function is approximated via a Neural …
Improving Deep Reinforcement Learning Methodology For Autonomous Defense And Escort Of Military High-Value Assets, Joseph Liles Iv
Improving Deep Reinforcement Learning Methodology For Autonomous Defense And Escort Of Military High-Value Assets, Joseph Liles Iv
Theses and Dissertations
This dissertation explores the application of machine learning to the control of autonomous unmanned combat aerial vehicles (AUCAVs). In particular, this research applies deep reinforcement learning methodologies to a defensive air combat scenario wherein a fleet of AUCAVs protects a military high-value asset (HVA). A collection of air battle management scenarios along with an original simulation environment and a set of designed computational experiments support the approximation of high-quality decision policies by employing Markov decision processes, approximate dynamic programming algorithms, and deep neural networks for value function approximation.
Multiagent Routing Problem With Dynamic Target Arrivals Solved Via Approximate Dynamic Programming, Andrew E. Mogan
Multiagent Routing Problem With Dynamic Target Arrivals Solved Via Approximate Dynamic Programming, Andrew E. Mogan
Theses and Dissertations
This research formulates and solves the multiagent routing problem with dynamic target arrivals (MRP-DTA), a stochastic system wherein a team of autonomous unmanned aerial vehicles (AUAVs) executes a strike coordination and reconnaissance (SCAR) mission against a notional adversary. Dynamic target arrivals that occur during the mission present the team of AUAVs with a sequential decision-making process which we model via a Markov Decision Process (MDP). To combat the curse of dimensionality, we construct and implement a hybrid approximate dynamic programming (ADP) algorithmic framework that employs a parametric cost function approximation (CFA) which augments a direct lookahead (DLA) model via a …
Team Air Combat Using Model-Based Reinforcement Learning, David A. Mottice
Team Air Combat Using Model-Based Reinforcement Learning, David A. Mottice
Theses and Dissertations
We formulate the first generalized air combat maneuvering problem (ACMP), called the MvN ACMP, wherein M friendly AUCAVs engage against N enemy AUCAVs, developing a Markov decision process (MDP) model to control the team of M Blue AUCAVs. The MDP model leverages a 5-degree-of-freedom aircraft state transition model and formulates a directed energy weapon capability. Instead, a model-based reinforcement learning approach is adopted wherein an approximate policy iteration algorithmic strategy is implemented to attain high-quality approximate policies relative to a high performing benchmark policy. The ADP algorithm utilizes a multi-layer neural network for the value function approximation regression mechanism. One-versus-one …
High-Density Parking For Autonomous Vehicles., Parag J. Siddique
High-Density Parking For Autonomous Vehicles., Parag J. Siddique
Electronic Theses and Dissertations
In a common parking lot, much of the space is devoted to lanes. Lanes must not be blocked for one simple reason: a blocked car might need to leave before the car that blocks it. However, the advent of autonomous vehicles gives us an opportunity to overcome this constraint, and to achieve a higher storage capacity of cars. Taking advantage of self-parking and intelligent communication systems of autonomous vehicles, we propose puzzle-based parking, a high-density design for a parking lot. We introduce a novel method of vehicle parking, which leads to maximum parking density. We then propose a heuristic method …
Assessment Of Adaptability Of A Supply Chain Trading Agent’S Strategy: Evolutionary Game Theory Approach, Yoon Sang Lee, Riyaz T. Sikora
Assessment Of Adaptability Of A Supply Chain Trading Agent’S Strategy: Evolutionary Game Theory Approach, Yoon Sang Lee, Riyaz T. Sikora
Journal of International Technology and Information Management
With the increase in the complexity of supply chain management, the use of intelligent agents for automated trading has gained popularity (Collins, Arunachalam, B, et al. 2006). The performance of supply-chain agents depends on not just the market environment (supply and demand patterns) but also on what types of other agents they are competing with. For designers of such agents it is important to ascertain that their agents are robust and can adapt to changing market and competitive environments. However, to date there has not been any work done that assesses the adaptability of a trading agent’s strategy in the …
Coping With The Curse Of Dimensionality By Combining Linear Programming And Reinforcement Learning, Scott H. Burton
Coping With The Curse Of Dimensionality By Combining Linear Programming And Reinforcement Learning, Scott H. Burton
All Graduate Theses and Dissertations, Spring 1920 to Summer 2023
Reinforcement learning techniques offer a very powerful method of finding solutions in unpredictable problem environments where human supervision is not possible. However, in many real world situations, the state space needed to represent the solutions becomes so large that using these methods becomes infeasible. Often the vast majority of these states are not valuable in finding the optimal solution. This work introduces a novel method of using linear programming to identify and represent the small area of the state space that is most likely to lead to a near-optimal solution, significantly reducing the memory requirements and time needed to arrive …