WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Runsheng, Lin, Hubert, Jeon, Wonseok, Feng, Hao, Zou, Yuliang, Sun, Liting, Gorman, John, Tolstaya, Ekaterina, Tang, Sarah, White, Brandyn, Sapp, Ben, Tan, Mingxing, Hwang, Jyh-Jing, Anguelov, Dragomir
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908649966796800
author Xu, Runsheng
Lin, Hubert
Jeon, Wonseok
Feng, Hao
Zou, Yuliang
Sun, Liting
Gorman, John
Tolstaya, Ekaterina
Tang, Sarah
White, Brandyn
Sapp, Ben
Tan, Mingxing
Hwang, Jyh-Jing
Anguelov, Dragomir
author_facet Xu, Runsheng
Lin, Hubert
Jeon, Wonseok
Feng, Hao
Zou, Yuliang
Sun, Liting
Gorman, John
Tolstaya, Ekaterina
Tang, Sarah
White, Brandyn
Sapp, Ben
Tan, Mingxing
Hwang, Jyh-Jing
Anguelov, Dragomir
contents Vision-based end-to-end (E2E) driving has garnered significant interest in the research community due to its scalability and synergy with multimodal large language models (MLLMs). However, current E2E driving benchmarks primarily feature nominal scenarios, failing to adequately test the true potential of these systems. Furthermore, existing open-loop evaluation metrics often fall short in capturing the multi-modal nature of driving or effectively evaluating performance in long-tail scenarios. To address these gaps, we introduce the Waymo Open Dataset for End-to-End Driving (WOD-E2E). WOD-E2E contains 4,021 driving segments (approximately 12 hours), specifically curated for challenging long-tail scenarios that that are rare in daily life with an occurring frequency of less than 0.03%. Concretely, each segment in WOD-E2E includes the high-level routing information, ego states, and 360-degree camera views from 8 surrounding cameras. To evaluate the E2E driving performance on these long-tail situations, we propose a novel open-loop evaluation metric: Rater Feedback Score (RFS). Unlike conventional metrics that measure the distance between predicted way points and the logs, RFS measures how closely the predicted trajectory matches rater-annotated trajectory preference labels. We have released rater preference labels for all WOD-E2E validation set segments, while the held out test set labels have been used for the 2025 WOD-E2E Challenge. Through our work, we aim to foster state of the art research into generalizable, robust, and safe end-to-end autonomous driving agents capable of handling complex real-world situations.
format Preprint
id arxiv_https___arxiv_org_abs_2510_26125
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
Xu, Runsheng
Lin, Hubert
Jeon, Wonseok
Feng, Hao
Zou, Yuliang
Sun, Liting
Gorman, John
Tolstaya, Ekaterina
Tang, Sarah
White, Brandyn
Sapp, Ben
Tan, Mingxing
Hwang, Jyh-Jing
Anguelov, Dragomir
Computer Vision and Pattern Recognition
Artificial Intelligence
Vision-based end-to-end (E2E) driving has garnered significant interest in the research community due to its scalability and synergy with multimodal large language models (MLLMs). However, current E2E driving benchmarks primarily feature nominal scenarios, failing to adequately test the true potential of these systems. Furthermore, existing open-loop evaluation metrics often fall short in capturing the multi-modal nature of driving or effectively evaluating performance in long-tail scenarios. To address these gaps, we introduce the Waymo Open Dataset for End-to-End Driving (WOD-E2E). WOD-E2E contains 4,021 driving segments (approximately 12 hours), specifically curated for challenging long-tail scenarios that that are rare in daily life with an occurring frequency of less than 0.03%. Concretely, each segment in WOD-E2E includes the high-level routing information, ego states, and 360-degree camera views from 8 surrounding cameras. To evaluate the E2E driving performance on these long-tail situations, we propose a novel open-loop evaluation metric: Rater Feedback Score (RFS). Unlike conventional metrics that measure the distance between predicted way points and the logs, RFS measures how closely the predicted trajectory matches rater-annotated trajectory preference labels. We have released rater preference labels for all WOD-E2E validation set segments, while the held out test set labels have been used for the 2025 WOD-E2E Challenge. Through our work, we aim to foster state of the art research into generalizable, robust, and safe end-to-end autonomous driving agents capable of handling complex real-world situations.
title WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.26125