Trust Region Method Using K-FAC in Multi-Agent Reinforcement Learning

0Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.
Get full text

Abstract

A challenging problem in multi-agent reinforcement learning (MARL) is to ensure that the policy converges quickly and is effective with limited computing resources. This paper extends the second-order optimization to MARL using Kronecker-factored approximate curvature (K-FAC) to approximate the natural gradient update. And it solves the challenge of training policy networks in MARL which requires a lot of time and computing costs. We propose a Heterogeneous-agent Trust Region algorithm using K-FAC (HAKTR). Further more, we endow HAKTR with monotonic performance improvement based on the multi-agent advantage decomposition theorem. Our algorithm is evaluated on continuous tasks in the MuJoCo environment. The experimental results show that HAKTR can achieve higher rewards with less computing costs compared to the baselines such as HATRPO and HAPPO. Moreover, HAKTR has good scalability regarding the number of agents and can be applied to large-scale networks.

Cite

CITATION STYLE

APA

Yu, J., Wu, F., & Zhao, J. (2022). Trust Region Method Using K-FAC in Multi-Agent Reinforcement Learning. In ACM International Conference Proceeding Series. Association for Computing Machinery. https://doi.org/10.1145/3579654.3579702

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free