Phase Re-service in Reinforcement Learning Traffic Signal Control

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zhang, Zhiyao, Gunter, George, Quinones-Grueiro, Marcos, Zhang, Yuhang, Barbour, William, Biswas, Gautam, Work, Daniel
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910551619141632
author Zhang, Zhiyao
Gunter, George
Quinones-Grueiro, Marcos
Zhang, Yuhang
Barbour, William
Biswas, Gautam
Work, Daniel
author_facet Zhang, Zhiyao
Gunter, George
Quinones-Grueiro, Marcos
Zhang, Yuhang
Barbour, William
Biswas, Gautam
Work, Daniel
contents This article proposes a novel approach to traffic signal control that combines phase re-service with reinforcement learning (RL). The RL agent directly determines the duration of the next phase in a pre-defined sequence. Before the RL agent's decision is executed, we use the shock wave theory to estimate queue expansion at the designated movement allowed for re-service and decide if phase re-service is necessary. If necessary, a temporary phase re-service is inserted before the next regular phase. We formulate the RL problem as a semi-Markov decision process (SMDP) and solve it with proximal policy optimization (PPO). We conducted a series of experiments that showed significant improvements thanks to the introduction of phase re-service. Vehicle delays are reduced by up to 29.95% of the average and up to 59.21% of the standard deviation. The number of stops is reduced by 26.05% on average with 45.77% less standard deviation.
format Preprint
id arxiv_https___arxiv_org_abs_2407_14775
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Phase Re-service in Reinforcement Learning Traffic Signal Control
Zhang, Zhiyao
Gunter, George
Quinones-Grueiro, Marcos
Zhang, Yuhang
Barbour, William
Biswas, Gautam
Work, Daniel
Systems and Control
This article proposes a novel approach to traffic signal control that combines phase re-service with reinforcement learning (RL). The RL agent directly determines the duration of the next phase in a pre-defined sequence. Before the RL agent's decision is executed, we use the shock wave theory to estimate queue expansion at the designated movement allowed for re-service and decide if phase re-service is necessary. If necessary, a temporary phase re-service is inserted before the next regular phase. We formulate the RL problem as a semi-Markov decision process (SMDP) and solve it with proximal policy optimization (PPO). We conducted a series of experiments that showed significant improvements thanks to the introduction of phase re-service. Vehicle delays are reduced by up to 29.95% of the average and up to 59.21% of the standard deviation. The number of stops is reduced by 26.05% on average with 45.77% less standard deviation.
title Phase Re-service in Reinforcement Learning Traffic Signal Control
topic Systems and Control
url https://arxiv.org/abs/2407.14775