Open Access. Powered by Scholars. Published by Universities.®
Artificial Intelligence and Robotics Commons™
Open Access. Powered by Scholars. Published by Universities.®
- Discipline
-
- Engineering (20)
- Operations Research, Systems Engineering and Industrial Engineering (9)
- Data Science (7)
- Social and Behavioral Sciences (7)
- Theory and Algorithms (7)
-
- Databases and Information Systems (6)
- Computer Engineering (5)
- Electrical and Computer Engineering (4)
- Life Sciences (4)
- Mathematics (4)
- Numerical Analysis and Scientific Computing (4)
- Statistics and Probability (4)
- Aerospace Engineering (3)
- Bioinformatics (3)
- Robotics (3)
- Applied Mathematics (2)
- Applied Statistics (2)
- Behavioral Economics (2)
- Controls and Control Theory (2)
- Economics (2)
- Graphics and Human Computer Interfaces (2)
- Operational Research (2)
- Other Computer Sciences (2)
- Power and Energy (2)
- Public Affairs, Public Policy and Public Administration (2)
- Systems Architecture (2)
- Systems Engineering (2)
- Institution
-
- Singapore Management University (18)
- San Jose State University (6)
- Missouri University of Science and Technology (4)
- University of Kentucky (4)
- Clemson University (3)
-
- University of South Florida (3)
- Dartmouth College (2)
- Georgia Southern University (2)
- University of Arkansas, Fayetteville (2)
- University of New Mexico (2)
- American University in Cairo (1)
- California Polytechnic State University, San Luis Obispo (1)
- Chapman University (1)
- Claremont Colleges (1)
- Embry-Riddle Aeronautical University (1)
- New Jersey Institute of Technology (1)
- Northern Illinois University (1)
- Old Dominion University (1)
- University of Connecticut (1)
- University of Nebraska at Omaha (1)
- University of New Hampshire (1)
- University of South Carolina (1)
- University of Texas Rio Grande Valley (1)
- University of Texas at Arlington (1)
- University of Texas at El Paso (1)
- Utah State University (1)
- Western University (1)
- Publication Year
- Publication
-
- Research Collection School Of Computing and Information Systems (16)
- Master's Projects (6)
- Masters Theses (3)
- Theses and Dissertations--Computer Science (3)
- USF Tampa Graduate Theses and Dissertations (3)
-
- All Dissertations (2)
- College of Graduate Studies: Theses & Dissertations (2)
- Computer Science ETDs (2)
- Dissertations and Theses Collection (Open Access) (2)
- Graduate Theses and Dissertations (2)
- Theses and Dissertations (2)
- All Graduate Theses and Dissertations, Fall 2023 to Present (1)
- All Theses (1)
- Computational and Data Sciences (PhD) Dissertations (1)
- Computer Science Senior Theses (1)
- Computer Science and Engineering Theses - Archive (1)
- Dartmouth College Ph.D Dissertations (1)
- Dissertations (1)
- Doctoral Dissertations (1)
- Doctoral Dissertations and Master's Theses (1)
- Electrical and Computer Engineering Publications (1)
- Graduate Research Theses & Dissertations (1)
- Graduate Student Government Association Research Conference (1)
- Honors Scholar Theses (1)
- Honors Theses and Capstones (1)
- Master's Theses (1)
- Open Access Theses & Dissertations (1)
- Pitzer Senior Theses (1)
- Publications (1)
- Theses and Dissertations--Electrical and Computer Engineering (1)
- Publication Type
Articles 1 - 30 of 63
Full-Text Articles in Artificial Intelligence and Robotics
Mapping Homogeneous Configuration States For Learning Based Motion Planners, Yazied Hasan
Mapping Homogeneous Configuration States For Learning Based Motion Planners, Yazied Hasan
Computer Science ETDs
Reinforcement learning (RL) excels at solving complex tasks, but training times can become prohibitively large for challenging motion-planning problems. Methods that address this cost often require additional training or tuning, counteracting the goal of reducing training time. A more effective approach is to exploit inherent task equivalences: many elements of the state space, dynamics, or structure are functionally interchangeable, enabling simplification or knowledge reuse. We present learning solutions that leverage these equivalences to enhance the RL process. First, we leverage the symmetry of homogeneous multi-agent teams to simplify the task to a single strategy. Second, we map correspondences between distinct …
Sequential Robustness In Adversarial Reinforcement Learning, Roman Lok-Ming Belaire
Sequential Robustness In Adversarial Reinforcement Learning, Roman Lok-Ming Belaire
Dissertations and Theses Collection (Open Access)
My goal is to build autonomous systems that expand the reach of human capability in challenging domains such as undersea and space exploration, disaster response, and large-scale infrastructure. In everyday settings, these systems will increasingly appear in safety-critical applications such as autonomous driving, robotics, and industrial manufacturing. A central requirement for these systems is the ability to operate reliably under uncertainty, particularly when the environment behaves in unanticipated ways.
The robust handling of unforeseen environment dynamics is therefore a technical cornerstone of autonomous decision-making; Adversarial attacks provide a useful and principled lens through which to study this problem. Adversarial \textit{robustness}, …
Teaching Machines To Deter: Exploring Strategic Deterrence In Ai Models, Will Taylor
Teaching Machines To Deter: Exploring Strategic Deterrence In Ai Models, Will Taylor
Theses/Capstones/Creative Projects
This capstone project investigates whether deterrence can emerge as a meaningful strategy within a zero-sum stochastic game using multi-agent reinforcement learning (MARL). After outlining core concepts in game theory and deterrence, the study models a simplified deterrence environment in which two minimax-Q agents repeatedly interact under uncertainty and adversarial incentives. The agents learn from rewards shaped by escalation costs, unilateral vulnerability, and the stabilizing benefits of restraint. Results show that both agents consistently converge toward a conservative, status-quo strategy, overwhelmingly selecting the Maintain action while avoiding both escalation and restraint in most scenarios. This behavior reflects the risk-averse logic of …
Enhancing Action And Ingredient Modeling For Semantically Grounded Recipe Generation, Guoshan Liu, Bin Zhu, Yian Li, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang
Enhancing Action And Ingredient Modeling For Semantically Grounded Recipe Generation, Guoshan Liu, Bin Zhu, Yian Li, Jingjing Chen, Chong-Wah Ngo, Yu-Gang Jiang
Research Collection School Of Computing and Information Systems
Recent advances in Multimodal Large Language Models (MLMMs) have enabled recipe generation from food images, yet outputs often contain semantically incorrect actions or ingredients despite high lexical scores (e.g., BLEU, ROUGE). To address this gap, we propose a semantically grounded framework that predicts and validates actions and ingredients as internal context for instruction generation. Our two-stage pipeline combines supervised fine-tuning (SFT) with reinforcement fine-tuning (RFT): SFT builds foundational accuracy using an Action-Reasoning dataset and ingredient corpus, while RFT employs frequency-aware rewards to improve long-tail action prediction and ingredient generalization. A Semantic Confidence Scoring and Rectification (SCSR) module further filters and …
From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios
From Attention To Reasoning: Beyond Accuracy In Multimodal Ai, Wayner Barrios
Dartmouth College Ph.D Dissertations
Multimodal large language models have achieved impressive performance on vision-language benchmarks by integrating visual encoders with large language models. Yet a critical gap persists between benchmark accuracy and genuine multimodal understanding: current evaluation frameworks assess performance by final answers alone, rewarding confident predictions while leaving systematic reasoning failures undetected.
This thesis addresses this gap through a unified framework that progresses from understanding to reasoning, using video as the most comprehensive multimodal testbed. Video inherently combines vision, audio, and language with temporal dynamics and massive token redundancy; techniques developed for video's comprehensive challenges transfer naturally to simpler multimodal tasks.
On understanding …
Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang
Real-Time Motion-Controllable Autoregressive Video Diffusion, Kesen Zhao, Jiaxin Shi, Beier Zhu, Junbao Zhou, Xiaolong Shen, Yuan Zhou, Qianru Sun, Hanwang Zhang
Research Collection School Of Computing and Information Systems
Real-time motion-controllable video generation remains challenging due to the inherent latency of bidirectional diffusion models and the lack of effective autoregressive (AR) approaches. Existing AR video diffusion models are limited to simple control signals or text-to-video generation, and often suffer from quality degradation and motion artifacts in few-step generation. To address these challenges, we propose AR-Drag, the first RL-enhanced few-step AR video diffusion model for real-time image-to-video generation with diverse motion control. We first fine-tune a base I2V model to support basic motion control, then further improve it via reinforcement learning with a trajectory-based reward model. Our design preserves the …
Multi-Agent Robotaxi Dispatch Coordination In A Real-World Simulation – Optimizing Rider Assignment, Rebalancing, And Charging Using Battery-Dependent Rewards And Welfare Maximization, Paden Thompson
All Graduate Theses and Dissertations, Fall 2023 to Present
We propose an approach to coordinate a robotaxi fleet for an autonomous ride-hail service. This is a service similar to a traditional ride-hailing service (Uber, Lyft), where customers request a ride and are then picked up in a car and dropped off in a new location; except, driverless vehicles called robotaxis are used to transport the customers.
Our approach teaches helpful coordination strategies to a robotaxi fleet while taking into account the individual battery level of the robotaxis. Each robotaxi acts as an individual agent in our simulation and can choose to pick up a rider, reposition to a new …
Visual-Enhanced Multimodal Framework For Flexible Job Shop Scheduling Problem, Peng Zhao, Zhiguang Cao, Di Wang, Wen Song, Wei Pang, You Zhou, Yuan Jiang
Visual-Enhanced Multimodal Framework For Flexible Job Shop Scheduling Problem, Peng Zhao, Zhiguang Cao, Di Wang, Wen Song, Wei Pang, You Zhou, Yuan Jiang
Research Collection School Of Computing and Information Systems
Multimodal models leverage complementary information across modalities to enrich feature representations. While visual information shows potential in representing structure for some combinatorial optimization problems (COPs), its application to complex scheduling like the Flexible Job Shop Scheduling Problem (FJSP) remains underexplored. Current learning-based FJSP solvers predominantly rely on handcrafted state features. This dependence can lead to inconsistencies and may not fully capture the problem's intricate dynamics. Crucially, these methods overlook visual modalities. Visual representations offer a distinct advantage by inherently capturing the global topological structure and complex resource interactions within the FJSP state. Unlike localized handcrafted features, this holistic, structural view …
Reinforcement Learning, Modeling Markets, And Professional Basketball Free Agency, Jacob Cohn
Reinforcement Learning, Modeling Markets, And Professional Basketball Free Agency, Jacob Cohn
Computational and Data Sciences (PhD) Dissertations
This dissertation presents a reinforcement learning-based approach to modeling and optimizing decision-making in professional basketball free agency and related economic environments. A Markov Decision Process (MDP) framework is introduced to capture the strategic interactions of NBA teams bidding for free agents under budgetary and roster constraints. To address computational scalability challenges, a reinforcement learning (RL) environment is developed, leveraging Proximal Policy Optimization (PPO) to approximate optimal policies for team decision-making.
Empirical results demonstrate that the RL agent successfully learns strategic bidding behavior that aligns with dynamic programming benchmarks in simplified settings while scaling effectively to larger, intractable environments. The study …
Tamos: Task-Aware Multi-Agent Orchestrator System, Joshit Mohanty, Sandeep Kumar Nayak, Sumit Lahiri
Tamos: Task-Aware Multi-Agent Orchestrator System, Joshit Mohanty, Sandeep Kumar Nayak, Sumit Lahiri
Graduate Student Government Association Research Conference
Large language models (LLMs) are increasingly at the core of multi-agent systems (MAS). However, the high resource demand, error propagation, and lack of adaptive evaluation mechanisms pose significant challenges in deploying these agentic solutions at scale. To address these concerns, this research proposes a Task-Aware Multi-Agent Orchestrator System designed to refine the agentic framework, categorizing tasks autonomously, assigning specialized evaluation datasets, and balancing token usage against functional effectiveness. This approach underscores robust data management, including AsyncHow, Mosaic AI, and Synthetic Preference Optimization (PO) corpora. Each dataset targets specific dimensions of agent performance, such as dynamic task decomposition and tool integration …
Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang
Dual Operation Aggregation Graph Neural Networks For Solving Flexible Job-Shop Scheduling Problem With Reinforcement Learning, Peng Zhao, You Zhou, Di Wang, Zhiguang Cao, Yubin Xiao, Xuan Wu, Yuanshu Li, Hongjia Liu, Wei Du, Yuan Jiang, Liupu Wang
Research Collection School Of Computing and Information Systems
With the widespread adoption of Internet Protocol (IP) communication technology and web-based platforms, cloud manufacturing has become a significant hallmark of Industry 4.0. Integrating graph algorithms into these web-enabled environments is crucial as they facilitate the representation and analysis of complex relationships in manufacturing processes, enabling efficient decision-making and adaptability in dynamic environments. As a key scheduling problem in cloud manufacturing, the flexible job-shop scheduling problem (FJSP) finds extensive applications in real-world scenarios. However, traditional FJSP-solving methods struggle to meet the efficiency and adaptability demands of cloud manufacturing due to generalization issues and excessive computational time, while reinforcement learning-based methods …
A Deep Reinforcement Learning Framework For Sequential Art Creation, Asmin Pothula
A Deep Reinforcement Learning Framework For Sequential Art Creation, Asmin Pothula
Computer Science and Engineering Theses - Archive
Most computational art systems rely on generative models that produce a complete artwork in a single pass, without capturing the gradual, decision-driven process through which human artists construct visual pieces. Prior research in sequential, stroke-based image generation, including differentiable neural painters and model-based reinforcement learning agents, has explored step-by-step creation, but these systems typically aim to reconstruct the input image within the same visual representation space, closely matching brushstrokes, textures, or colors to the target. In contrast, this thesis investigates sequential art creation in a different artistic representation, where the final artwork does not share the same visual form as …
Optimizing Decision-Making In A Cerebral Palsy Model Using Reinforcement Learning, Richard Ampah
Optimizing Decision-Making In A Cerebral Palsy Model Using Reinforcement Learning, Richard Ampah
Pitzer Senior Theses
This study presents an original interdisciplinary investigation into how reinforcement learning (RL) can model motor and cognitive defects and potentially improve motor and cognitive functions in individuals with cerebral palsy (CP), a non-progressive neurological disorder that impairs movement and adaptability. Integrating computational neuroscience and machine learning, the research applies policy gradient methods and Markov Decision Processes (MDPs) to simulate adaptive learning in agents with and without CP-related constraints.
The central aim is to compare the cumulative rewards of optimal policies, derived from value iteration, and human-like learning policies using the REINFORCE algorithm, both with and without the Bellman baseline. The …
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
Machine Learning Methods For Intrusion Detection And Response In Network Security, Ayomide Oyemaja
College of Graduate Studies: Theses & Dissertations
Intrusion Detection Systems (IDS) play a crucial role in computer network security by identifying malicious activities and potential cyberattacks. This thesis combines machine learning and cybersecurity by applying Reinforcement Learning (RL) in intrusion detection and response using the NSL-KDD dataset.
We designed and implemented a Q-learning framework where an agent learns to classify network traffic over time by interacting with the environment and receiving rewards based on detection accuracy. We also look at the importance of feature selection and classification techniques and how effective they are in improving model performance, reducing the complexity of computation, and producing more desirable results. …
Safety Through Feedback In Constrained Rl, Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
Safety Through Feedback In Constrained Rl, Shashank Reddy Chirra, Pradeep Varakantham, Praveen Paruchuri
Research Collection School Of Computing and Information Systems
In safety-critical RL settings, the inclusion of an additional cost function is often favoured over the arduous task of modifying the reward function to ensure the agent's safe behaviour. However, designing or evaluating such a cost function can be prohibitively expensive. For instance, in the domain of self-driving, designing a cost function that encompasses all unsafe behaviours (e.g., aggressive lane changes, risky overtakes) is inherently complex, it must also consider all the actors present in the scene making it expensive to evaluate. In such scenarios, the cost function can be learned from feedback collected offline in between training rounds. This …
Personalized Driving Using Inverse Reinforcement Learning, Rodrigo J. Gonzalez Salinas
Personalized Driving Using Inverse Reinforcement Learning, Rodrigo J. Gonzalez Salinas
Theses and Dissertations
This thesis introduces an autonomous driving controller designed to replicate individual driving behaviors based on a provided demonstration. The controller employs Inverse Reinforcement Learning (IRL) to formulate the reward function associated with the provided demonstration. IRL is implemented through a dual-feedback loop system. The inner loop utilizes Q-learning, a model-free reinforcement learning technique, to optimize the Hamilton-Jacobi-Bellman (HJB) equation and derive an appropriate control solution. The outer loop leverages this derived control solution to generate parameters for the reward function, which are subsequently integrated into the HJB equation. The ultimate control policy is deduced from the final reward function obtained …
Reinforcement Learning For Strategic Airport Slot Scheduling: Analysis Of State Observations And Reward Designs, Anh Nguyen-Duy, Duc-Thinh Pham, Jian-Yi Lye, Nguyen Binh Duong Ta
Reinforcement Learning For Strategic Airport Slot Scheduling: Analysis Of State Observations And Reward Designs, Anh Nguyen-Duy, Duc-Thinh Pham, Jian-Yi Lye, Nguyen Binh Duong Ta
Research Collection School Of Computing and Information Systems
Due to the NP-hard nature, the strategic airport slot scheduling problem is calling for exploring sub-optimal approaches, such as heuristics and learning-based approaches. Moreover, the continuous increase in air traffic demand requires approaches that can work well in new scenarios. While heuristics rely on a fixed set of rules, which limits the ability to explore new solutions, Reinforcement Learning offers a versatile framework to automate the search and generalize to unseen scenarios. Finding a suitable state observation and reward structure design is essential in using Reinforcement Learning. In this paper, we investigate the impact of providing the Reinforcement Learning agent …
Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An
Reinforcement Nash Equilibrium Solver, Xinrun Wang, Chang Yang, Shuxin Li, Pengdeng Li, Xiao Huang, Hau Chan, Bo An
Research Collection School Of Computing and Information Systems
Nash Equilibrium (NE) is the canonical solution concept of game theory, which provides an elegant tool to understand the rationalities. Computing NE in two- or multi-player general-sum games is PPAD-Complete. Therefore, in this work, we propose REinforcement Nash Equilibrium Solver (RENES), which trains a single policy to modify the games with different sizes and applies the solvers on the modified games where the obtained solution is evaluated on the original games. Specifically, our contributions are threefold. i) We represent the games as ��-rank response graphs and leverage graph neural network (GNN) to handle the games with different sizes as inputs; …
Earnhft: Efficient Hierarchical Reinforcement Learning For High Frequency Trading, Molei Qin, Shuo Sun, Wentao Zhang, Haochong Xia, Xinrun Wang, Bo An
Earnhft: Efficient Hierarchical Reinforcement Learning For High Frequency Trading, Molei Qin, Shuo Sun, Wentao Zhang, Haochong Xia, Xinrun Wang, Bo An
Research Collection School Of Computing and Information Systems
High-frequency trading (HFT) is using computer algorithms to make trading decisions in short time scales (e.g., second-level), which is widely used in the Cryptocurrency (Crypto) market, (e.g., Bitcoin). Reinforcement learning (RL) in financial research has shown stellar performance on many quantitative trading tasks. However, most methods focus on low-frequency trading, e.g., day-level, which cannot be directly applied to HFT because of two challenges. First, RL for HFT involves dealing with extremely long trajectories (e.g., 2.4 million steps per month), which is hard to optimize and evaluate. Second, the dramatic price fluctuations and market trend changes of Crypto make existing algorithms …
Transition-Informed Reinforcement Learning For Large-Scale Stackelberg Mean-Field Games., Pengdeng Li, Runsheng Yu, Xinrun Wang, Bo An
Transition-Informed Reinforcement Learning For Large-Scale Stackelberg Mean-Field Games., Pengdeng Li, Runsheng Yu, Xinrun Wang, Bo An
Research Collection School Of Computing and Information Systems
Many real-world scenarios including fleet management and Ad auctions can be modeled as Stackelberg mean-field games (SMFGs) where a leader aims to incentivize a large number of homogeneous self-interested followers to maximize her utility. Existing works focus on cases with a small number of heterogeneous followers, e.g., 5-10, and suffer from scalability issue when the number of followers increases. There are three major challenges in solving large-scale SMFGs: i) classical methods based on solving differential equations fail as they require exact dynamics parameters, ii) learning by interacting with environment is data-inefficient, and iii) complex interaction between the leader and followers …
Finding Hierarchies To Improve Learning In Hierarchical Reinforcement Learning, Roy Mobley
Finding Hierarchies To Improve Learning In Hierarchical Reinforcement Learning, Roy Mobley
Theses and Dissertations--Computer Science
Reinforcement Learning (RL) is an approach to allowing computer agents to try and learn how to solve problems by learning what actions are best to take in a given situation. RL is effective for learning what to do in an environment, but as the problem grows larger, the amount of information needed grows exponentially, making RL less effective on complex problems. A big challenge, often called the curse of dimensionality, is that the number of states and possible number of actions in an environment can grow too large to sufficiently test every possible combination of state and action. One method …
Smart Applications And Resource Management In Internet Of Things, Zeinab Akhavan
Smart Applications And Resource Management In Internet Of Things, Zeinab Akhavan
Computer Science ETDs
Internet of Things (IoT) technologies are currently the principal solutions driving smart cities. These new technologies such as Cyber Physical Systems, 5G and data analytic have emerged to address various cities' infrastructure issues ranging from transportation and energy management to healthcare systems. An IoT setting primarily consists of a wide range of users and devices as a massive network interacting with different layers of the city infrastructure resulting in generating sheer volume of data to enable smart city services. The goal of smart city services is to create value for the entire ecosystem, whether this is health, education, transportation, energy, …
Online Aircraft System Identification Using A Novel Parameter Informed Reinforcement Learning Method, Nathan Schaff
Online Aircraft System Identification Using A Novel Parameter Informed Reinforcement Learning Method, Nathan Schaff
Doctoral Dissertations and Master's Theses
This thesis presents the development and analysis of a novel method for training reinforcement learning neural networks for online aircraft system identification of multiple similar linear systems, such as all fixed wing aircraft. This approach, termed Parameter Informed Reinforcement Learning (PIRL), dictates that reinforcement learning neural networks should be trained using input and output trajectory/history data as is convention; however, the PIRL method also includes any known and relevant aircraft parameters, such as airspeed, altitude, center of gravity location and/or others. Through this, the PIRL Agent is better suited to identify novel/test-set aircraft.
First, the PIRL method is applied to …
Quantifying Balance: Computational And Learning Frameworks For The Characterization Of Balance In Bipedal Systems, Kubra Akbas
Quantifying Balance: Computational And Learning Frameworks For The Characterization Of Balance In Bipedal Systems, Kubra Akbas
Dissertations
In clinical practice and general healthcare settings, the lack of reliable and objective balance and stability assessment metrics hinders the tracking of patient performance progression during rehabilitation; the assessment of bipedal balance plays a crucial role in understanding stability and falls in humans and other bipeds, while providing clinicians important information regarding rehabilitation outcomes. Bipedal balance has often been examined through kinematic or kinetic quantities, such as the Zero Moment Point and Center of Pressure; however, analyzing balance specifically through the body's Center of Mass (COM) state offers a holistic and easily comprehensible view of balance and stability.
Building upon …
Motion Synthesis And Control For Autonomous Agents Using Generative Models And Reinforcement Learning, Pei Xu
All Dissertations
Imitating and predicting human motions have wide applications in both graphics and robotics, from developing realistic models of human movement and behavior in immersive virtual worlds and games to improving autonomous navigation for service agents deployed in the real world. Traditional approaches for motion imitation and prediction typically rely on pre-defined rules to model agent behaviors or use reinforcement learning with manually designed reward functions. Despite impressive results, such approaches cannot effectively capture the diversity of motor behaviors and the decision making capabilities of human beings. Furthermore, manually designing a model or reward function to explicitly describe human motion characteristics …
Imitating Opponent To Win: Adversarial Policy Imitation Learning In Two-Player Competitive Games, The Viet Bui, Tien Mai, Thanh H. Nguyen
Imitating Opponent To Win: Adversarial Policy Imitation Learning In Two-Player Competitive Games, The Viet Bui, Tien Mai, Thanh H. Nguyen
Research Collection School Of Computing and Information Systems
Recent research on vulnerabilities of deep reinforcement learning (RL) has shown that adversarial policies adopted by an adversary agent can influence a target RL agent (victim agent) to perform poorly in a multi-agent environment. In existing studies, adversarial policies are directly trained based on experiences of interacting with the victim agent. There is a key shortcoming of this approach --- knowledge derived from historical interactions may not be properly generalized to unexplored policy regions of the victim agent, making the trained adversarial policy significantly less effective. In this work, we design a new effective adversarial policy learning algorithm that overcomes …
Gaslight: Attacking Hard-Label Black-Box Classifiers Via Deep Reinforcement Learning, Rajat Sethi
Gaslight: Attacking Hard-Label Black-Box Classifiers Via Deep Reinforcement Learning, Rajat Sethi
All Theses
Through artificial intelligence, algorithms can classify arrays of data, such as images or videos, into a predefined set of categories. With enough labeled data, a classifier can analyze an input’s components and calculate confidence scores for each category. However, machine learning relies heavily on approximation, which allows attackers to exploit classifiers by providing adversarial
examples. Specifically, attackers can modify their input so that the victim classifier cannot correctly label it, while a human observer would be unable to notice the difference.
This thesis proposes Gaslight, a system that uses deep reinforcement learning to generate adversarial examples against a victim classifier. …
Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian
Practical Ai Value Alignment Using Stories, Md Sultan Al Nahian
Theses and Dissertations--Computer Science
As more machine learning agents interact with humans, it is increasingly a prospect that an agent trained to perform a task optimally - using only a measure of task performance as feedback--can violate societal norms for acceptable behavior or cause harm. Consequently, it becomes necessary to prioritize task performance and ensure that AI actions do not have detrimental effects. Value alignment is a property of intelligent agents, wherein they solely pursue goals and activities that are non-harmful and beneficial to humans. Current approaches to value alignment largely depend on imitation learning or learning from demonstration methods. However, the dynamic nature …
Peer-To-Peer Energy Trading In Smart Residential Environment With User Behavioral Modeling, Ashutosh Timilsina
Peer-To-Peer Energy Trading In Smart Residential Environment With User Behavioral Modeling, Ashutosh Timilsina
Theses and Dissertations--Computer Science
Electric power systems are transforming from a centralized unidirectional market to a decentralized open market. With this shift, the end-users have the possibility to actively participate in local energy exchanges, with or without the involvement of the main grid. Rapidly reducing prices for Renewable Energy Technologies (RETs), supported by their ease of installation and operation, with the facilitation of Electric Vehicles (EV) and Smart Grid (SG) technologies to make bidirectional flow of energy possible, has contributed to this changing landscape in the distribution side of the traditional power grid.
Trading energy among users in a decentralized fashion has been referred …
Airport Assignment For Emergency Aircraft Using Reinforcement Learning, Saketh Kamatham
Airport Assignment For Emergency Aircraft Using Reinforcement Learning, Saketh Kamatham
Master's Projects
The volume of air traffic is increasing exponentially every day. The Air Traffic Control (ATC) at the airport has to handle aircraft runway assignments for landing and takeoff and airspace maintenance by directing passing aircraft through the airspace safely. If any aircraft is facing a technical issue or problem and is in a state of emergency, it requires expedited landing to respond to that emergency. The ATC gives this aircraft priority to landing and assistance. This process is very strenuous as the ATC has to deal with multiple aspects along with the emergency aircraft. It is the duty of the …