Phase Re-service in Reinforcement Learning Traffic Signal Control
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866910551619141632 |
|---|---|
| author | Zhang, Zhiyao Gunter, George Quinones-Grueiro, Marcos Zhang, Yuhang Barbour, William Biswas, Gautam Work, Daniel |
| author_facet | Zhang, Zhiyao Gunter, George Quinones-Grueiro, Marcos Zhang, Yuhang Barbour, William Biswas, Gautam Work, Daniel |
| contents | This article proposes a novel approach to traffic signal control that combines phase re-service with reinforcement learning (RL). The RL agent directly determines the duration of the next phase in a pre-defined sequence. Before the RL agent's decision is executed, we use the shock wave theory to estimate queue expansion at the designated movement allowed for re-service and decide if phase re-service is necessary. If necessary, a temporary phase re-service is inserted before the next regular phase. We formulate the RL problem as a semi-Markov decision process (SMDP) and solve it with proximal policy optimization (PPO). We conducted a series of experiments that showed significant improvements thanks to the introduction of phase re-service. Vehicle delays are reduced by up to 29.95% of the average and up to 59.21% of the standard deviation. The number of stops is reduced by 26.05% on average with 45.77% less standard deviation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_14775 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Phase Re-service in Reinforcement Learning Traffic Signal Control Zhang, Zhiyao Gunter, George Quinones-Grueiro, Marcos Zhang, Yuhang Barbour, William Biswas, Gautam Work, Daniel Systems and Control This article proposes a novel approach to traffic signal control that combines phase re-service with reinforcement learning (RL). The RL agent directly determines the duration of the next phase in a pre-defined sequence. Before the RL agent's decision is executed, we use the shock wave theory to estimate queue expansion at the designated movement allowed for re-service and decide if phase re-service is necessary. If necessary, a temporary phase re-service is inserted before the next regular phase. We formulate the RL problem as a semi-Markov decision process (SMDP) and solve it with proximal policy optimization (PPO). We conducted a series of experiments that showed significant improvements thanks to the introduction of phase re-service. Vehicle delays are reduced by up to 29.95% of the average and up to 59.21% of the standard deviation. The number of stops is reduced by 26.05% on average with 45.77% less standard deviation. |
| title | Phase Re-service in Reinforcement Learning Traffic Signal Control |
| topic | Systems and Control |
| url | https://arxiv.org/abs/2407.14775 |