Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (20)
- Operations Research, Systems Engineering and Industrial Engineering (9)
- Data Science (7)
- Social and Behavioral Sciences (7)
- Theory and Algorithms (7)
-
- Databases and Information Systems (6)
- Computer Engineering (5)
- Electrical and Computer Engineering (4)
- Life Sciences (4)
- Mathematics (4)
- Numerical Analysis and Scientific Computing (4)
- Statistics and Probability (4)
- Aerospace Engineering (3)
- Bioinformatics (3)
- Robotics (3)
- Applied Mathematics (2)
- Applied Statistics (2)
- Behavioral Economics (2)
- Controls and Control Theory (2)
- Economics (2)
- Graphics and Human Computer Interfaces (2)
- Operational Research (2)
- Other Computer Sciences (2)
- Power and Energy (2)
- Public Affairs, Public Policy and Public Administration (2)
- Systems Architecture (2)
- Systems Engineering (2)
- Institution
-
- Singapore Management University (18)
- San Jose State University (6)
- Missouri University of Science and Technology (4)
- University of Kentucky (4)
- Clemson University (3)
-
- University of South Florida (3)
- Dartmouth College (2)
- Georgia Southern University (2)
- University of Arkansas, Fayetteville (2)
- University of New Mexico (2)
- American University in Cairo (1)
- California Polytechnic State University, San Luis Obispo (1)
- Chapman University (1)
- Claremont Colleges (1)
- Embry-Riddle Aeronautical University (1)
- New Jersey Institute of Technology (1)
- Northern Illinois University (1)
- Old Dominion University (1)
- University of Connecticut (1)
- University of Nebraska at Omaha (1)
- University of New Hampshire (1)
- University of South Carolina (1)
- University of Texas Rio Grande Valley (1)
- University of Texas at Arlington (1)
- University of Texas at El Paso (1)
- Utah State University (1)
- Western University (1)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (16)
- Master's Projects (6)
- Masters Theses (3)
- Theses and Dissertations--Computer Science (3)
- USF Tampa Graduate Theses and Dissertations (3)
-
- All Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science ETDs (2)
- Dissertations and Theses Collection (Open Access) (2)
- Graduate Theses and Dissertations (2)
- Theses and Dissertations (2)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- All Theses (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Computer Science Senior Theses (1)
- Computer Science and Engineering Theses - Archive (1)
- Dartmouth College Ph.D Dissertations (1)
- Dissertations (1)
- Doctoral Dissertations (1)
- Doctoral Dissertations and Master's Theses (1)
- Electrical and Computer Engineering Publications (1)
- Graduate Research Theses & Dissertations (1)
- Graduate Student Government Association Research Conference (1)
- Honors Scholar Theses (1)
- Honors Theses and Capstones (1)
- Master's Theses (1)
- Open Access Theses & Dissertations (1)
- Pitzer Senior Theses (1)
- Publications (1)
- Theses and Dissertations--Electrical and Computer Engineering (1)
- Publication Type
Articles 31 - 60 of 63
Full-Text Articles in Artificial Intelligence and Robotics
Reinforcement Learning Approach To Coordinate Real-World Multi-Agent Dynamic Routing And Scheduling, Joe Waldy
Reinforcement Learning Approach To Coordinate Real-World Multi-Agent Dynamic Routing And Scheduling, Joe Waldy
Dissertations and Theses Collection (Open Access)
In this thesis, we study new variants of routing and scheduling problems motivated by real-world problems from the urban logistics and law enforcement domains. In particular, we focus on two key aspects: dynamic and multi-agent. While routing problems such as the Vehicle Routing Problem (VRP) is well-studied in the Operations Research (OR) community, we know that in real-world route planning today, initially-planned route plans and schedules may be disrupted by dynamically-occurring events. In addition, routing and scheduling plans cannot be done in silos due to the presence of other agents which may be independent and self-interested. These requirements create …
Adaptive Multi-Scale Place Cell Representations And Replay For Spatial Navigation And Learning In Autonomous Robots, Pablo Scleidorovich
Adaptive Multi-Scale Place Cell Representations And Replay For Spatial Navigation And Learning In Autonomous Robots, Pablo Scleidorovich
USF Tampa Graduate Theses and Dissertations
Place cells are one of the most widely studied neurons thought to play a vital role in spatial cognition. Extensive studies show that their activity in the rodent hippocampus is highly correlated with the animal’s spatial location, forming “place fields” of smaller sizes near the dorsal pole and larger sizes near the ventral pole. Despite advances, it is yet unclear how this multi-scale representation enables navigation in complex environments.
In this dissertation, we analyze the place cell representation from a computational point of view, evaluating how multi-scale place fields impact navigation in large and cluttered environments. The objectives are to …
Symplectically Integrated Symbolic Regression Of Hamiltonian Dynamical Systems, Daniel Dipietro
Symplectically Integrated Symbolic Regression Of Hamiltonian Dynamical Systems, Daniel Dipietro
Computer Science Senior Theses
Here we present Symplectically Integrated Symbolic Regression (SISR), a novel technique for learning physical governing equations from data. SISR employs a deep symbolic regression approach, using a multi-layer LSTMRNN with mutation to probabilistically sample Hamiltonian symbolic expressions. Using symplectic neural networks, we develop a model-agnostic approach for extracting meaningful physical priors from the data that can be imposed on-the-fly into the RNN output, limiting its search space. Hamiltonians generated by the RNN are optimized and assessed using a fourth-order symplectic integration scheme; prediction performance is used to train the LSTM-RNN to generate increasingly better functions via a risk-seeking policy gradients …
Reinforcement Learning Approach To Solve Dynamic Bi-Objective Police Patrol Dispatching And Rescheduling Problem, Waldy Joe, Hoong Chuin Lau, Jonathan Pan
Reinforcement Learning Approach To Solve Dynamic Bi-Objective Police Patrol Dispatching And Rescheduling Problem, Waldy Joe, Hoong Chuin Lau, Jonathan Pan
Research Collection School Of Computing and Information Systems
Police patrol aims to fulfill two main objectives namely to project presence and to respond to incidents in a timely manner. Incidents happen dynamically and can disrupt the initially-planned patrol schedules. The key decisions to be made will be which patrol agent to be dispatched to respond to an incident and subsequently how to adapt the patrol schedules in response to such dynamically-occurring incidents whilst still fulfilling both objectives; which sometimes can be conflicting. In this paper, we define this real-world problem as a Dynamic Bi-Objective Police Patrol Dispatching and Rescheduling Problem and propose a solution approach that combines Deep …
Hierarchical Value Decomposition For Effective On-Demand Ride Pooling, Hao Jiang, Pradeep Varakantham
Hierarchical Value Decomposition For Effective On-Demand Ride Pooling, Hao Jiang, Pradeep Varakantham
Research Collection School Of Computing and Information Systems
On-demand ride-pooling (e.g., UberPool, GrabShare) services focus on serving multiple different customer requests using each vehicle, i.e., an empty or partially filled vehicle can be assigned requests from different passengers with different origins and destinations. On the other hand, in Taxi on Demand (ToD) services (e.g., UberX), one vehicle is assigned to only one request at a time. On-demand ride pooling is not only beneficial to customers (lower cost), drivers (higher revenue per trip) and aggregation companies (higher revenue), but is also of crucial importance to the environment as it reduces the number of vehicles required on the roads. Since …
Analyzing Decision-Making In Robot Soccer For Attacking Behaviors, Justin Rodney
Analyzing Decision-Making In Robot Soccer For Attacking Behaviors, Justin Rodney
USF Tampa Graduate Theses and Dissertations
In robotics soccer, decision-making is critical to the performance of a team’s SoftwareSystem. The University of South Florida’s (USF) RoboBulls team implements behavior for the robots by using traditional methods such as analytical geometry to path plan and determine whether an action should be taken. In recent works, Machine Learning (ML) and Reinforcement Learning (RL) techniques have been used to calculate the probability of success for a pass or goal, and even train models for performing low-level skills such as traveling towards a ball and shooting it towards the goal[1, 2]. Open-source frameworks have been created for training Reinforcement Learning …
Iseeq: Information Seeking Question Generation Using Dynamic Meta-Information Retrieval And Knowledge Graphs, Manas Gaur, Kalpa Gunaratna, Vijay Srinivasan, Hongxia Jin
Iseeq: Information Seeking Question Generation Using Dynamic Meta-Information Retrieval And Knowledge Graphs, Manas Gaur, Kalpa Gunaratna, Vijay Srinivasan, Hongxia Jin
Publications
Conversational Information Seeking (CIS) is a relatively new research area within conversational AI that attempts to seek information from end-users in order to understand and satisfy users’ needs. If realized, such a system has far-reaching benefits in the real world; for example, a CIS system can assist clinicians in pre-screening or triaging patients in healthcare. A key open sub-problem in CIS that remains unaddressed in the literature is generating Information Seeking Questions (ISQs) based on a short initial query from the end user. To address this open problem, we propose Information SEEking Question generator (ISEEQ), a novel approach for generating …
Reinforcement Learning: Low Discrepancy Action Selection For Continuous States And Actions, Jedidiah Lindborg
Reinforcement Learning: Low Discrepancy Action Selection For Continuous States And Actions, Jedidiah Lindborg
College of Graduate Studies: Theses & Dissertations
In reinforcement learning the process of selecting an action during the exploration or exploitation stage is difficult to optimize. The purpose of this thesis is to create an action selection process for an agent by employing a low discrepancy action selection (LDAS) method. This should allow the agent to quickly determine the utility of its actions by prioritizing actions that are dissimilar to ones that it has already picked. In this way the learning process should be faster for the agent and result in more optimal policies.
Whole File Chunk Based Deduplication Using Reinforcement Learning, Xincheng Yuan
Whole File Chunk Based Deduplication Using Reinforcement Learning, Xincheng Yuan
Master's Projects
Deduplication is the process of removing replicated data content from storage facilities like online databases, cloud datastore, local file systems, etc., which is commonly performed as part of data preprocessing to eliminate redundant data that requires unnecessary storage spaces and computing power. Deduplication is even more specifically essential for file backup systems since duplicated files will presumably consume more storage space, especially with a short backup period like daily [8]. A common technique in this field involves splitting files into chunks whose hashes can be compared using data structures or techniques like clustering. In this project we explore the possibility …
Burst-Induced Multi-Armed Bandit For Learning Recommendation, Rodrigo Alves, Antoine Ledent, Marius Kloft
Burst-Induced Multi-Armed Bandit For Learning Recommendation, Rodrigo Alves, Antoine Ledent, Marius Kloft
Research Collection School Of Computing and Information Systems
In this paper, we introduce a non-stationary and context-free Multi-Armed Bandit (MAB) problem and a novel algorithm (which we refer to as BMAB) to solve it. The problem is context-free in the sense that no side information about users or items is needed. We work in a continuous-time setting where each timestamp corresponds to a visit by a user and a corresponding decision regarding recommendation. The main novelty is that we model the reward distribution as a consequence of variations in the intensity of the activity, and thereby we assist the exploration/exploitation dilemma by exploring the temporal dynamics of the …
Reinforcement Learning Algorithms: An Overview And Classification, Fadi Almahamid, Katarina Grolinger
Reinforcement Learning Algorithms: An Overview And Classification, Fadi Almahamid, Katarina Grolinger
Electrical and Computer Engineering Publications
The desire to make applications and machines more intelligent and the aspiration to enable their operation without human interaction have been driving innovations in neural networks, deep learning, and other machine learning techniques. Although reinforcement learning has been primarily used in video games, recent advancements and the development of diverse and powerful reinforcement algorithms have enabled the reinforcement learning community to move from playing video games to solving complex real-life problems in autonomous systems such as self-driving cars, delivery drones, and automated robotics. Understanding the environment of an application and the algorithms’ limitations plays a vital role in selecting the …
Toward Deep Supervised Anomaly Detection: Reinforcement Learning From Partially Labeled Anomaly Data, Guansong Pang, Anton Van Den Hengel, Chunhua Shen, Longbing Cao
Toward Deep Supervised Anomaly Detection: Reinforcement Learning From Partially Labeled Anomaly Data, Guansong Pang, Anton Van Den Hengel, Chunhua Shen, Longbing Cao
Research Collection School Of Computing and Information Systems
We consider the problem of anomaly detection with a small set of partially labeled anomaly examples and a large-scale unlabeled dataset. This is a common scenario in many important applications. Existing related methods either exclusively fit the limited anomaly examples that typically do not span the entire set of anomalies, or proceed with unsupervised learning from the unlabeled data. We propose here instead a deep reinforcement learning-based approach that enables an end-to-end optimization of the detection of both labeled and unlabeled anomalies. This approach learns the known abnormality by automatically interacting with an anomalybiased simulation environment, while continuously extending the …
Scheduling Allocation And Inventory Replenishment Problems Under Uncertainty: Applications In Managing Electric Vehicle And Drone Battery Swap Stations, Amin Asadi
Graduate Theses and Dissertations
In this dissertation, motivated by electric vehicle (EV) and drone application growth, we propose novel optimization problems and solution techniques for managing the operations at EV and drone battery swap stations. In Chapter 2, we introduce a novel class of stochastic scheduling allocation and inventory replenishment problems (SAIRP), which determines the recharging, discharging, and replacement decisions at a swap station over time to maximize the expected total profit. We use Markov Decision Process (MDP) to model SAIRPs facing uncertain demands, varying costs, and battery degradation. Considering battery degradation is crucial as it relaxes the assumption that charging/discharging batteries do not …
Optimization And Machine Learning Methods For Solving Combinatorial Problems In Urban Transportation, Aigerim Bogyrbayeva
Optimization And Machine Learning Methods For Solving Combinatorial Problems In Urban Transportation, Aigerim Bogyrbayeva
USF Tampa Graduate Theses and Dissertations
This dissertation investigates three applications of emerging technologies for urban trans- portation. In the first chapter, we design a new market for fractional ownership of au- tonomous vehicles (AVs), in which an AV is co-leased by a group of individuals. We present a practical iterative auction based on the combinatorial clock auction to match the interested customers together and determine their payments. In designing such an auction, we con- sider continuous-time items (time slots) which are defined by bidders, and naturally exploit driverless mobility of AVs to form co-leasing groups. To relieve the computational burdens of both bidders and the …
Reinforcement Learning For Realistic Robotic Training: A Survey, Andres Jaramillo
Reinforcement Learning For Realistic Robotic Training: A Survey, Andres Jaramillo
Honors Scholar Theses
Reinforcement learning is a widely popular topic that has resulted in a plethora of
research papers and interest from academia and industry. When applied with robotics,
the field has showed some promising signs that robots can achieve levels of complex
cognitive abilities rivaling humans, but the goal of creating sapient robots is far from
a reality due to many challenges involved with training robots in a real world setting.
This paper will provide a survey regarding the keys towards realistic robotic training by
detailing the challenges and overviewing the reinforcement learning solutions involved
in getting a robot to think like …
Deep Learning For Multi-Tissue Cancer Classification Of Gene Expressions, Tarek Khorshed
Deep Learning For Multi-Tissue Cancer Classification Of Gene Expressions, Tarek Khorshed
Theses and Dissertations
We contribute in saving the lives of cancer patients through early detection and diagnosis, since one of the major challenges in cancer treatment is that patients are diagnosed at very late stages when appropriate medical interventions become less effective and full curative treatment is no longer achievable. Cancer classification using gene expressions is extremely challenging given the complexity and high dimensionality of the data. Current classification methods typically rely on samples collected from a single tissue type and perform a prerequisite of gene feature selection to avoid processing the full set of genes. These methods fall short in taking advantage …
Neural Network Supervised And Reinforcement Learning For Neurological, Diagnostic, And Modeling Problems, Donald Wunsch Iii
Neural Network Supervised And Reinforcement Learning For Neurological, Diagnostic, And Modeling Problems, Donald Wunsch Iii
Masters Theses
“As the medical world becomes increasingly intertwined with the tech sphere, machine learning on medical datasets and mathematical models becomes an attractive application. This research looks at the predictive capabilities of neural networks and other machine learning algorithms, and assesses the validity of several feature selection strategies to reduce the negative effects of high dataset dimensionality. Our results indicate that several feature selection methods can maintain high validation and test accuracy on classification tasks, with neural networks performing best, for both single class and multi-class classification applications. This research also evaluates a proof-of-concept application of a deep-Q-learning network (DQN) to …
Virtual Robot Locomotion On Variable Terrain With Adversarial Reinforcement Learning, Phong Nguyen
Virtual Robot Locomotion On Variable Terrain With Adversarial Reinforcement Learning, Phong Nguyen
Master's Projects
Reinforcement Learning (RL) is a machine learning technique where an agent learns to perform a complex action by going through a repeated process of trial and error to maximize a well-defined reward function. This form of learning has found applications in robot locomotion where it has been used to teach robots to traverse complex terrain. While RL algorithms may work well in training robot locomotion, they tend to not generalize well when the agent is brought into an environment that it has never encountered before. Possible solutions from the literature include training a destabilizing adversary alongside the locomotive learning agent. …
A Comprehensive And Modular Robotic Control Framework For Model-Less Control Law Development Using Reinforcement Learning For Soft Robotics, Charles Sullivan
A Comprehensive And Modular Robotic Control Framework For Model-Less Control Law Development Using Reinforcement Learning For Soft Robotics, Charles Sullivan
Open Access Theses & Dissertations
Soft robotics is a growing field in robotics research. Heavily inspired by biological systems, these robots are made of softer, non-linear, materials such as elastomers and are actuated using several novel methods, from fluidic actuation channels to shape changing materials such as electro-active polymers. Highly non-linear materials make modeling difficult, and sensors are still an area of active research. These issues have rendered typical control and modeling techniques often inadequate for soft robotics. Reinforcement learning is a branch of machine learning that focuses on model-less control by mapping states to actions that maximize a specific reward signal. Reinforcement learning has …
Landing Throttleable Hybrid Rockets With Hierarchical Reinforcement Learning In A Simulated Environment, Francesco Alessandro Stefano Mikulis-Borsoi
Landing Throttleable Hybrid Rockets With Hierarchical Reinforcement Learning In A Simulated Environment, Francesco Alessandro Stefano Mikulis-Borsoi
Honors Theses and Capstones
In this paper, I develop a hierarchical Markov Decision Process (MDP) structure for completing the task of vertical rocket landing. I start by covering the background of this problem, and formally defining its constraints. In order to reduce mistakes while formulating different MDPs, I define and develop the criteria for a standardized MDP definition format. I then decompose the problem into several sub-problems of vertical landing, namely velocity control and vertical stability control. By exploiting MDP coupling and symmetrical properties, I am able to significantly reduce the size of the state space compared to a unified MDP formulation. This paper …
A Comparative Analysis Of Reinforcement Learning Applied To Task-Space Reaching With A Robotic Manipulator With And Without Gravity Compensation, Jonathan Fugal
A Comparative Analysis Of Reinforcement Learning Applied To Task-Space Reaching With A Robotic Manipulator With And Without Gravity Compensation, Jonathan Fugal
Theses and Dissertations--Electrical and Computer Engineering
Advances in computing power in recent years have facilitated developments in autonomous robotic systems. These robotic systems can be used in prosthetic limbs, wearhouse packaging and sorting, assembly line production, as well as many other applications. Designing these autonomous systems typically requires robotic system and world models (for classical control based strategies) or time consuming and computationally expensive training (for learning based strategies). Often these requirements are difficult to fulfill. There are ways to combine classical control and learning based strategies that can mitigate both requirements. One of these ways is to use a gravity compensated torque control with reinforcement …
End-To-End Deep Reinforcement Learning For Multi-Agent Collaborative Exploration, Zichen Chen, Budhitama Subagdja, Ah-Hwee Tan
End-To-End Deep Reinforcement Learning For Multi-Agent Collaborative Exploration, Zichen Chen, Budhitama Subagdja, Ah-Hwee Tan
Research Collection School Of Computing and Information Systems
Exploring an unknown environment by multiple autonomous robots is a major challenge in robotics domains. As multiple robots are assigned to explore different locations, they may interfere each other making the overall tasks less efficient. In this paper, we present a new model called CNN-based Multi-agent Proximal Policy Optimization (CMAPPO) to multi-agent exploration wherein the agents learn the effective strategy to allocate and explore the environment using a new deep reinforcement learning architecture. The model combines convolutional neural network to process multi-channel visual inputs, curriculum-based learning, and PPO algorithm for motivation based reinforcement learning. Evaluations show that the proposed method …
Learning To Play The Trading Game, Neeraj Kulkarni
Learning To Play The Trading Game, Neeraj Kulkarni
Master's Projects
Can we train a stock trading bot that can take decisions in high-entropy envi- ronments like stock markets to generate profits based on some optimal policy? Can we further extend this learning for any general trading problem? Quantitative Al- gorithms are responsible for more than 75% of the stock trading around the world. Creating a stock market prediction model is comparatively easy. But creating a prof- itable prediction model is still considered as a challenging task in the field of machine learning and deep learning due to the unpredictability of the financial markets. Us- ing biologically inspired computing techniques of …
Ai Dining Suggestion App, Bao Pham
Ai Dining Suggestion App, Bao Pham
Master's Projects
Trying to decide what to eat can sometimes be challenging and time-consuming for people. Google and Yelp have large scale data sets of restaurant information as well as Application Program Interfaces (APIs) for using them. This restaurant data includes time, price range, traffic, temperature, etc. The goal of this project is to build an app that eases the process of finding a restaurant to eat. This app has a Tinder-like user friendly User Interface (UI) design to change the common way that lists of restaurants are presented to users on mobile apps. It also uses the help of Artificial Intelligence …
Sky Surveys Scheduling Using Reinforcement Learning, Andres Felipe Alba Hernandez
Sky Surveys Scheduling Using Reinforcement Learning, Andres Felipe Alba Hernandez
Graduate Research Theses & Dissertations
Modern cosmic sky surveys (e.g., CMB S4, DES, LSST) collect a complex diversity of astronomical objects. Each of class of objects presents different requirements for observation time and sensitivity. For determining the best sequence of exposures for mapping the sky systematically, conventional scheduling methods do not optimize the use of survey time and resources. Dynamic sky survey scheduling is an NP-hard problem that has been therefore treated primarily with heuristic methods. We present an alternative scheduling method based on reinforcement learning (RL) that aims to optimize the use of telescope resources for scheduling sky surveys.
We present an exploration of …
Virtual Robot Climbing Using Reinforcement Learning, Ujjawal Garg
Virtual Robot Climbing Using Reinforcement Learning, Ujjawal Garg
Master's Projects
Reinforcement Learning (RL) is a field of Artificial Intelligence that has gained a lot of attention in recent years. In this project, RL research was used to design and train an agent to climb and navigate through an environment with slopes. We compared and evaluated the performance of two state-of-the-art reinforcement learning algorithms for locomotion related tasks, Deep Deterministic Policy Gradients (DDPG) and Trust Region Policy Optimisation (TRPO). We observed that, on an average, training with TRPO was three times faster than DDPG, and also much more stable for the locomotion control tasks that we experimented. We conducted experiments and …
Influencing Exploration In Actor-Critic Reinforcement Learning Algorithms, Andrew R. Gough
Influencing Exploration In Actor-Critic Reinforcement Learning Algorithms, Andrew R. Gough
Master's Theses
Reinforcement Learning (RL) is a subset of machine learning primarily concerned with goal-directed learning and optimal decision making. RL agents learn based on a reward signal discovered from trial and error in complex, uncertain environments with the goal of maximizing positive reward signals. RL approaches need to scale up as they are applied to more complex environments with extremely large state spaces. Inefficient exploration methods cannot sufficiently explore complex environments in a reasonable amount of time, and optimal policies will be unrealized resulting in RL agents failing to solve an environment.
This thesis proposes a novel variant of the Actor-Advantage …
Improving Asynchronous Advantage Actor Critic With A More Intelligent Exploration Strategy, James B. Holliday
Improving Asynchronous Advantage Actor Critic With A More Intelligent Exploration Strategy, James B. Holliday
Graduate Theses and Dissertations
We propose a simple and efficient modification to the Asynchronous Advantage Actor Critic (A3C)
algorithm that improves training. In 2016 Google’s DeepMind set a new standard for state-of-theart
reinforcement learning performance with the introduction of the A3C algorithm. The goal of
this research is to show that A3C can be improved by the use of a new novel exploration strategy we
call “Follow then Forage Exploration” (FFE). FFE forces the agents to follow the best known path
at the beginning of a training episode and then later in the episode the agent is forced to “forage”
and explores randomly. In …
A New Reinforcement Learning Algorithm With Fixed Exploration For Semi-Markov Decision Processes, Angelo Michael Encapera
A New Reinforcement Learning Algorithm With Fixed Exploration For Semi-Markov Decision Processes, Angelo Michael Encapera
Masters Theses
"Artificial intelligence or machine learning techniques are currently being widely applied for solving problems within the field of data analytics. This work presents and demonstrates the use of a new machine learning algorithm for solving semi-Markov decision processes (SMDPs). SMDPs are encountered in the domain of Reinforcement Learning to solve control problems in discrete-event systems. The new algorithm developed here is called iSMART, an acronym for imaging Semi-Markov Average Reward Technique. The algorithm uses a constant exploration rate, unlike its precursor R-SMART, which required exploration decay. The major difference between R-SMART and iSMART is that the latter uses, in addition …
A Bounded Actor-Critic Algorithm For Reinforcement Learning, Ryan Jacob Lawhead
A Bounded Actor-Critic Algorithm For Reinforcement Learning, Ryan Jacob Lawhead
Masters Theses
"This thesis presents a new actor-critic algorithm from the domain of reinforcement learning to solve Markov and semi-Markov decision processes (or problems) in the field of airline revenue management (ARM). The ARM problem is one of control optimization in which a decision-maker must accept or reject a customer based on a requested fare. This thesis focuses on the so-called single-leg version of the ARM problem, which can be cast as a semi-Markov decision process (SMDP). Large-scale Markov decision processes (MDPs) and SMDPs suffer from the curses of dimensionality and modeling, making it difficult to create the transition probability matrices (TPMs) …