GuideLight: "Industrial Solution" Guidance for More Practical Traffic Signal Control Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jiang, Haoyuan, Xiong, Xuantang, Li, Ziyue, Mao, Hangyu, Sui, Guanghu, Ruan, Jingqing, Cheng, Yuheng, Wei, Hua, Ketter, Wolfgang, Zhao, Rui
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910527845826560
author Jiang, Haoyuan
Xiong, Xuantang
Li, Ziyue
Mao, Hangyu
Sui, Guanghu
Ruan, Jingqing
Cheng, Yuheng
Wei, Hua
Ketter, Wolfgang
Zhao, Rui
author_facet Jiang, Haoyuan
Xiong, Xuantang
Li, Ziyue
Mao, Hangyu
Sui, Guanghu
Ruan, Jingqing
Cheng, Yuheng
Wei, Hua
Ketter, Wolfgang
Zhao, Rui
contents Currently, traffic signal control (TSC) methods based on reinforcement learning (RL) have proven superior to traditional methods. However, most RL methods face difficulties when applied in the real world due to three factors: input, output, and the cycle-flow relation. The industry's observable input is much more limited than simulation-based RL methods. For real-world solutions, only flow can be reliably collected, whereas common RL methods need more. For the output action, most RL methods focus on acyclic control, which real-world signal controllers do not support. Most importantly, industry standards require a consistent cycle-flow relationship: non-decreasing and different response strategies for low, medium, and high-level flows, which is ignored by the RL methods. To narrow the gap between RL methods and industry standards, we innovatively propose to use industry solutions to guide the RL agent. Specifically, we design behavior cloning and curriculum learning to guide the agent to mimic and meet industry requirements and, at the same time, leverage the power of exploration and exploitation in RL for better performance. We theoretically prove that such guidance can largely decrease the sample complexity to polynomials in the horizon when searching for an optimal policy. Our rigid experiments show that our method has good cycle-flow relation and superior performance.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10811
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GuideLight: "Industrial Solution" Guidance for More Practical Traffic Signal Control Agents
Jiang, Haoyuan
Xiong, Xuantang
Li, Ziyue
Mao, Hangyu
Sui, Guanghu
Ruan, Jingqing
Cheng, Yuheng
Wei, Hua
Ketter, Wolfgang
Zhao, Rui
Multiagent Systems
Artificial Intelligence
Machine Learning
Currently, traffic signal control (TSC) methods based on reinforcement learning (RL) have proven superior to traditional methods. However, most RL methods face difficulties when applied in the real world due to three factors: input, output, and the cycle-flow relation. The industry's observable input is much more limited than simulation-based RL methods. For real-world solutions, only flow can be reliably collected, whereas common RL methods need more. For the output action, most RL methods focus on acyclic control, which real-world signal controllers do not support. Most importantly, industry standards require a consistent cycle-flow relationship: non-decreasing and different response strategies for low, medium, and high-level flows, which is ignored by the RL methods. To narrow the gap between RL methods and industry standards, we innovatively propose to use industry solutions to guide the RL agent. Specifically, we design behavior cloning and curriculum learning to guide the agent to mimic and meet industry requirements and, at the same time, leverage the power of exploration and exploitation in RL for better performance. We theoretically prove that such guidance can largely decrease the sample complexity to polynomials in the horizon when searching for an optimal policy. Our rigid experiments show that our method has good cycle-flow relation and superior performance.
title GuideLight: "Industrial Solution" Guidance for More Practical Traffic Signal Control Agents
topic Multiagent Systems
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2407.10811