MALight: A Deep Reinforcement Learning Traffic Light Control Algorithm with Pressure and Attentive Experience Replay

1Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

This study proposes a new algorithm MALight based on multi-step deep Q network (DQN) and attentive experience replay (AER). Multi-step DQN samples multiple consecutive experiences within a time step, combines them into a longterm sample, and uses them to update the Q network to reduce the bias caused by inaccurate Q value estimation, which could accelerate the convergence of Q network. During training, we adopted the concept of AER to prioritize learning experiences close to the current state to enable the agent to learn better strategies. Finally, we conducted simulation experiments in the city traffic simulator CityFlow using both synthetic and real-world traffic flow datasets. The evaluation results show that MALight can accelerate the convergence speed of the network, effectively improve the traffic capacity of intersections, and optimize the average travel time of intersections.

Cite

CITATION STYLE

APA

Kong, Y., Li, Y., & Hsia, C. H. (2024). MALight: A Deep Reinforcement Learning Traffic Light Control Algorithm with Pressure and Attentive Experience Replay. Journal of Internet Technology, 25(7), 955–962. https://doi.org/10.70003/160792642024122507001

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free