Abstract
A challenging problem in multi-agent reinforcement learning (MARL) is to ensure that the policy converges quickly and is effective with limited computing resources. This paper extends the second-order optimization to MARL using Kronecker-factored approximate curvature (K-FAC) to approximate the natural gradient update. And it solves the challenge of training policy networks in MARL which requires a lot of time and computing costs. We propose a Heterogeneous-agent Trust Region algorithm using K-FAC (HAKTR). Further more, we endow HAKTR with monotonic performance improvement based on the multi-agent advantage decomposition theorem. Our algorithm is evaluated on continuous tasks in the MuJoCo environment. The experimental results show that HAKTR can achieve higher rewards with less computing costs compared to the baselines such as HATRPO and HAPPO. Moreover, HAKTR has good scalability regarding the number of agents and can be applied to large-scale networks.
Author supplied keywords
Cite
CITATION STYLE
Yu, J., Wu, F., & Zhao, J. (2022). Trust Region Method Using K-FAC in Multi-Agent Reinforcement Learning. In ACM International Conference Proceeding Series. Association for Computing Machinery. https://doi.org/10.1145/3579654.3579702
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.