Beyond Conservatism: Diffusion Policies In Offline Multi-agent Reinforcement Learning
2023 Β· Zhuoran Li, Ling Pan, Longbo Huang
Abstract
We present a novel Diffusion Offline Multi-agent Model (DOM2) for offline Multi-Agent Reinforcement Learning (MARL). Different from existing algorithms that rely mainly on conservatism in policy design, DOM2 enhances policy expressiveness and diversity based on diffusion. Specifically, we incorporate a diffusion model into the policy network and propose a trajectory-based data-augmentation scheme in training. These key ingredients make our algorithm more robust to environment changes and achieve significant improvements in performance, generalization and data-efficiency. Our extensive experimental results demonstrate that DOM2 outperforms existing state-of-the-art methods in multi-agent particle and multi-agent MuJoCo environments, and generalizes significantly better in shifted environments thanks to its high expressiveness and diversity. Furthermore, DOM2 shows superior data efficiency and can achieve state-of-the-art performance with \(20+\) times less data compared to existing algo
Authors
(none)
Tags
Stats
Related papers
- Diffusion Models For Offline Multi-agent Reinforcement Learning With Safety Constraints (2024)0.00
- Madiff: Offline Multi-agent Learning With Diffusion Models (2023)2.26
- Diffusion Policies With Value-conditional Optimization For Offline Reinforcement Learning (2025)0.00
- Preferred-action-optimized Diffusion Policies For Offline Reinforcement Learning (2024)0.00
- Diffpogan: Diffusion Policies With Generative Adversarial Networks For Offline Reinforcement Learning (2024)0.00
- Long-horizon Rollout Via Dynamics Diffusion For Offline Reinforcement Learning (2024)1.81
- Policy Representation Via Diffusion Probability Model For Reinforcement Learning (2023)0.00
- Plan Better Amid Conservatism: Offline Multi-agent Reinforcement Learning With Actor Rectification (2021)0.00