Open Access. Powered by Scholars. Published by Universities.®

Articles 1 - 12 of 12

Full-Text Articles in Operational Research

Learning To Dogfight: Proximal Policy Optimization Vs. Double Deep Q Network For 2v2 Air Combat With Directed Energy Weapons In Afsim, Caden W. Wilson Mar 2025

Learning To Dogfight: Proximal Policy Optimization Vs. Double Deep Q Network For 2v2 Air Combat With Directed Energy Weapons In Afsim, Caden W. Wilson

Theses and Dissertations

This research utilizes reinforcement learning (RL) to train two blue agents each imbued with a directed energy weapon (DEW) in a 2v2 within visual range air combat maneuvering problem. A phased solution approach is employed to repeatedly tune and train several RL algorithm implementations: Proximal Policy Optimization (PPO) and Double Deep Q Network (DDQN). Phase I of training includes reward shaping for basic flight elements such as altitude, airspeed, and target proximity. Phase II of training builds off policies developed in Phase I, but rewards emphasize winning the aerial engagement by any means necessary. DDQN significantly outperforms PPO in Phase …


Reinforcement Learning For Aeromedical Evacuation In Nonstationary Combat Environments, Micah J. Kartchner Mar 2025

Reinforcement Learning For Aeromedical Evacuation In Nonstationary Combat Environments, Micah J. Kartchner

Theses and Dissertations

This research formulates the medical evacuation (MEDEVAC) dispatching problem as a sequential decision process and investigates the application of reinforcement learning under nonstationary conditions. We model the dynamic arrival rate of MEDEVAC requests using a nonstationary Hawkes process and design a Double Deep Q-Network algorithm that incorporates belief states to anticipate future requests. Through computational experimentation, we analyze the impact of belief formulation on decision quality and system performance. Results indicate that policies incorporating belief states significantly outperform myopic dispatching policies, reducing urgent casualty wait times by up to 49.68% and increasing on-time evacuations by up to 21.91%.


Proximal Policy Optimization Applied To The Beyond Visual Range Air Combat Maneuvering Problem, Daniel B. Joseph Mar 2025

Proximal Policy Optimization Applied To The Beyond Visual Range Air Combat Maneuvering Problem, Daniel B. Joseph

Theses and Dissertations

Artificial intelligence (AI) grows ever-more important in warfighting. Emerging technologies allow for the use of AI to control aircraft and weapons systems. This research investigates the application of reinforcement learning (RL) through the Proximal Policy Optimization (PPO) algorithm to a two-versus-two (2v2) beyond-visual-range (BVR) air combat maneuvering problem (ACMP). Implemented in the Advanced Framework for Simulation, Integration, and Modeling (AFSIM), the methodology frames the engagement as a Markov decision process, wherein an autonomous RL agent learns continuous control decisions—throttle, pitch, roll, and yaw—under a cooperative communication scheme. A multi-phase curriculum-learning approach facilitates the progressive acquisition of flight stability, weapon deployment, …


Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs Mar 2024

Reinforcement Learning For Team Based Air Combat Maneuvering Decisions With Directed Energy Weaponry, Joshua D. Combs

Theses and Dissertations

Leveraging the Advanced Framework for Simulation, Integration, and Modeling (AFSIM) we investigate the use of reinforcement learning (RL) techniques for imbuing AUCAV agents with high-quality behaviors for the within-visual-range air combat maneuvering problem (ACMP). We formulate the 2v2 WVR ACMP as a Markov decision process wherein friendly AUCAVs are equipped with DEW capabilities and operate with 6 degrees of freedom. We utilize the Double Deep Q-Network RL algorithm, which centrally trains two friendly AUCAVs and employ a phased learning approach, initially exposing the AUCAVs to a dense reward environment for early training, followed by a sparse reward environment to encourage …


A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae Mar 2024

A Random Forest-Based Q-Learning Algorithm: Toward Interpretable Artificial Intelligence, Victor R. Rae

Theses and Dissertations

A growing demand exists for interpretable artificial intelligence models, leading to extensive research efforts to enhance the explainability and transparency of policies generated by reinforcement learning (RL) methods. This research develops random forest-based RL algorithms as a logical progression in this academic pursuit. The algorithms are evaluated using three standard benchmark environments from OpenAI gym — CartPole, MountainCar, and LunarLander — and compared to implementations of the Deep Q-learning Network (DQN) and Double DQN (DDQN) algorithms for various metrics, including performance, robustness, efficiency, and interpretability. The random forest-based algorithms exhibit superior performance to both neural network-based algorithms in two out …


A Reinforcement Learning Approach To The 2v2 Beyond Visual Range Air Combat Maneuvering Problem, Jacob J. Pike Mar 2024

A Reinforcement Learning Approach To The 2v2 Beyond Visual Range Air Combat Maneuvering Problem, Jacob J. Pike

Theses and Dissertations

This research examines a 2v2 air combat maneuvering problem (ACMP) in a Beyond Visual Range (BVR) environment. A discrete-time, infinite-horizon Markov Decision Process (MDP) model represents the BVR-ACMP, seeking to determine high-quality policies for a pair of autonomous aircraft to execute tactical maneuvers and firing decisions. The Advanced Framework for Simulation, Integration, and Modeling (AFSIM) characterizes the complex six-degree of freedom (6-DOF) aircraft operations, encompassing kinematics, sensors, and weapons. Given the high dimensionality and continuous nature of the state and decision variables, a deep reinforcement learning (RL) solution approach is adopted wherein the value function is approximated via a Neural …


Improving Deep Reinforcement Learning Methodology For Autonomous Defense And Escort Of Military High-Value Assets, Joseph Liles Iv Sep 2023

Improving Deep Reinforcement Learning Methodology For Autonomous Defense And Escort Of Military High-Value Assets, Joseph Liles Iv

Theses and Dissertations

This dissertation explores the application of machine learning to the control of autonomous unmanned combat aerial vehicles (AUCAVs). In particular, this research applies deep reinforcement learning methodologies to a defensive air combat scenario wherein a fleet of AUCAVs protects a military high-value asset (HVA). A collection of air battle management scenarios along with an original simulation environment and a set of designed computational experiments support the approximation of high-quality decision policies by employing Markov decision processes, approximate dynamic programming algorithms, and deep neural networks for value function approximation.


Multiagent Routing Problem With Dynamic Target Arrivals Solved Via Approximate Dynamic Programming, Andrew E. Mogan Mar 2022

Multiagent Routing Problem With Dynamic Target Arrivals Solved Via Approximate Dynamic Programming, Andrew E. Mogan

Theses and Dissertations

This research formulates and solves the multiagent routing problem with dynamic target arrivals (MRP-DTA), a stochastic system wherein a team of autonomous unmanned aerial vehicles (AUAVs) executes a strike coordination and reconnaissance (SCAR) mission against a notional adversary. Dynamic target arrivals that occur during the mission present the team of AUAVs with a sequential decision-making process which we model via a Markov Decision Process (MDP). To combat the curse of dimensionality, we construct and implement a hybrid approximate dynamic programming (ADP) algorithmic framework that employs a parametric cost function approximation (CFA) which augments a direct lookahead (DLA) model via a …


Team Air Combat Using Model-Based Reinforcement Learning, David A. Mottice Mar 2022

Team Air Combat Using Model-Based Reinforcement Learning, David A. Mottice

Theses and Dissertations

We formulate the first generalized air combat maneuvering problem (ACMP), called the MvN ACMP, wherein M friendly AUCAVs engage against N enemy AUCAVs, developing a Markov decision process (MDP) model to control the team of M Blue AUCAVs. The MDP model leverages a 5-degree-of-freedom aircraft state transition model and formulates a directed energy weapon capability. Instead, a model-based reinforcement learning approach is adopted wherein an approximate policy iteration algorithmic strategy is implemented to attain high-quality approximate policies relative to a high performing benchmark policy. The ADP algorithm utilizes a multi-layer neural network for the value function approximation regression mechanism. One-versus-one …


High-Density Parking For Autonomous Vehicles., Parag J. Siddique Aug 2021

High-Density Parking For Autonomous Vehicles., Parag J. Siddique

Electronic Theses and Dissertations

In a common parking lot, much of the space is devoted to lanes. Lanes must not be blocked for one simple reason: a blocked car might need to leave before the car that blocks it. However, the advent of autonomous vehicles gives us an opportunity to overcome this constraint, and to achieve a higher storage capacity of cars. Taking advantage of self-parking and intelligent communication systems of autonomous vehicles, we propose puzzle-based parking, a high-density design for a parking lot. We introduce a novel method of vehicle parking, which leads to maximum parking density. We then propose a heuristic method …


Assessment Of Adaptability Of A Supply Chain Trading Agent’S Strategy: Evolutionary Game Theory Approach, Yoon Sang Lee, Riyaz T. Sikora Jan 2019

Assessment Of Adaptability Of A Supply Chain Trading Agent’S Strategy: Evolutionary Game Theory Approach, Yoon Sang Lee, Riyaz T. Sikora

Journal of International Technology and Information Management

With the increase in the complexity of supply chain management, the use of intelligent agents for automated trading has gained popularity (Collins, Arunachalam, B, et al. 2006). The performance of supply-chain agents depends on not just the market environment (supply and demand patterns) but also on what types of other agents they are competing with. For designers of such agents it is important to ascertain that their agents are robust and can adapt to changing market and competitive environments. However, to date there has not been any work done that assesses the adaptability of a trading agent’s strategy in the …


Coping With The Curse Of Dimensionality By Combining Linear Programming And Reinforcement Learning, Scott H. Burton May 2010

Coping With The Curse Of Dimensionality By Combining Linear Programming And Reinforcement Learning, Scott H. Burton

All Graduate Theses and Dissertations, Spring 1920 to Summer 2023

Reinforcement learning techniques offer a very powerful method of finding solutions in unpredictable problem environments where human supervision is not possible. However, in many real world situations, the state space needed to represent the solutions becomes so large that using these methods becomes infeasible. Often the vast majority of these states are not valuable in finding the optimal solution. This work introduces a novel method of using linear programming to identify and represent the small area of the state space that is most likely to lead to a near-optimal solution, significantly reducing the memory requirements and time needed to arrive …