Abstract
In multi-agent reinforcement learning (MARL) tasks, the state-action value, commonly referred to as the (Formula presented.) -value, can vary among agents because of their individual rewards, resulting in a (Formula presented.) -vector. Determining an optimal policy is challenging, as it involves more than just maximizing a single (Formula presented.) -value. Various optimal policies, such as a Nash equilibrium, have been studied in this context. Algorithms like Nash Q-learning and Nash Actor-Critic have shown effectiveness in these scenarios. This paper extends this research by proposing a deep Q-networks algorithm capable of learning various (Formula presented.) -vectors using Max, Nash, and Maximin strategies. We validate the effectiveness of our approach in a dual-arm robotic environment, a representative human cyber-physical systems (HCPS) scenario, where two robotic arms collaborate to lift a pot or hand over a hammer to each other. This setting highlights how incorporating MARL into HCPS can address real-world complexities such as physical constraints, communication overhead, and dynamic interactions among multiple agents.
Author supplied keywords
Cite
CITATION STYLE
Luo, Z., Chen, Z., Liu, S., & Welsh, J. (2025, January 1). Multi-Agent Reinforcement Learning With Deep Networks for Diverse Q-Vectors. Electronics Letters. John Wiley and Sons Inc. https://doi.org/10.1049/ell2.70342
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.