Multi-Agent Reinforcement Learning With Deep Networks for Diverse Q-Vectors

4Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

In multi-agent reinforcement learning (MARL) tasks, the state-action value, commonly referred to as the (Formula presented.) -value, can vary among agents because of their individual rewards, resulting in a (Formula presented.) -vector. Determining an optimal policy is challenging, as it involves more than just maximizing a single (Formula presented.) -value. Various optimal policies, such as a Nash equilibrium, have been studied in this context. Algorithms like Nash Q-learning and Nash Actor-Critic have shown effectiveness in these scenarios. This paper extends this research by proposing a deep Q-networks algorithm capable of learning various (Formula presented.) -vectors using Max, Nash, and Maximin strategies. We validate the effectiveness of our approach in a dual-arm robotic environment, a representative human cyber-physical systems (HCPS) scenario, where two robotic arms collaborate to lift a pot or hand over a hammer to each other. This setting highlights how incorporating MARL into HCPS can address real-world complexities such as physical constraints, communication overhead, and dynamic interactions among multiple agents.

Cite

CITATION STYLE

APA

Luo, Z., Chen, Z., Liu, S., & Welsh, J. (2025, January 1). Multi-Agent Reinforcement Learning With Deep Networks for Diverse Q-Vectors. Electronics Letters. John Wiley and Sons Inc. https://doi.org/10.1049/ell2.70342

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free