Multi-agent Off-policy Actor-critic Reinforcement Learning For Partially Observable Environments
2024 Β· Ainur Zhaikhan, Ali H. Sayed
Abstract
This study proposes the use of a social learning method to estimate a global state within a multi-agent off-policy actor-critic algorithm for reinforcement learning (RL) operating in a partially observable environment. We assume that the network of agents operates in a fully-decentralized manner, possessing the capability to exchange variables with their immediate neighbors. The proposed design methodology is supported by an analysis demonstrating that the difference between final outcomes, obtained when the global state is fully observed versus estimated through the social learning method, is \(\epsilon\)-bounded when an appropriate number of iterations of social learning updates are implemented. Unlike many existing dec-POMDP-based RL approaches, the proposed algorithm is suitable for model-free multi-agent reinforcement learning as it does not require knowledge of a transition model. Furthermore, experimental results illustrate the efficacy of the algorithm and demonstrate its super
Authors
(none)
Tags
Stats
Related papers
- Optimal Decision-making In Mixed-agent Partially Observable Stochastic Environments Via Reinforcement Learning (2019)0.00
- Actor-critic Policy Optimization In Partially Observable Multiagent Environments (2018)0.00
- Reinforcement Learning Under Partial Observability Guided By Learned Environment Models (2022)6.34
- Provably Efficient Reinforcement Learning In Partially Observable Dynamical Systems (2022)0.00
- Belief States For Cooperative Multi-agent Reinforcement Learning Under Partial Observability (2025)0.00
- Unbiased Asymmetric Reinforcement Learning Under Partial Observability (2021)2.26
- Deep Decentralized Multi-task Multi-agent Reinforcement Learning Under Partial Observability (2017)0.00
- Centralized Model And Exploration Policy For Multi-agent RL (2021)0.00