Abstract
This study proposes a new algorithm MALight based on multi-step deep Q network (DQN) and attentive experience replay (AER). Multi-step DQN samples multiple consecutive experiences within a time step, combines them into a longterm sample, and uses them to update the Q network to reduce the bias caused by inaccurate Q value estimation, which could accelerate the convergence of Q network. During training, we adopted the concept of AER to prioritize learning experiences close to the current state to enable the agent to learn better strategies. Finally, we conducted simulation experiments in the city traffic simulator CityFlow using both synthetic and real-world traffic flow datasets. The evaluation results show that MALight can accelerate the convergence speed of the network, effectively improve the traffic capacity of intersections, and optimize the average travel time of intersections.
Author supplied keywords
Cite
CITATION STYLE
Kong, Y., Li, Y., & Hsia, C. H. (2024). MALight: A Deep Reinforcement Learning Traffic Light Control Algorithm with Pressure and Attentive Experience Replay. Journal of Internet Technology, 25(7), 955–962. https://doi.org/10.70003/160792642024122507001
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.