Valeo4Cast: A Modular Approach to End-to-End Forecasting

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xu, Yihong, Zablocki, Éloi, Boulch, Alexandre, Puy, Gilles, Chen, Mickael, Bartoccioni, Florent, Samet, Nermin, Siméoni, Oriane, Gidaris, Spyros, Vu, Tuan-Hung, Bursuc, Andrei, Valle, Eduardo, Marlet, Renaud, Cord, Matthieu
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916411403665408
author Xu, Yihong
Zablocki, Éloi
Boulch, Alexandre
Puy, Gilles
Chen, Mickael
Bartoccioni, Florent
Samet, Nermin
Siméoni, Oriane
Gidaris, Spyros
Vu, Tuan-Hung
Bursuc, Andrei
Valle, Eduardo
Marlet, Renaud
Cord, Matthieu
author_facet Xu, Yihong
Zablocki, Éloi
Boulch, Alexandre
Puy, Gilles
Chen, Mickael
Bartoccioni, Florent
Samet, Nermin
Siméoni, Oriane
Gidaris, Spyros
Vu, Tuan-Hung
Bursuc, Andrei
Valle, Eduardo
Marlet, Renaud
Cord, Matthieu
contents Motion forecasting is crucial in autonomous driving systems to anticipate the future trajectories of surrounding agents such as pedestrians, vehicles, and traffic signals. In end-to-end forecasting, the model must jointly detect and track from sensor data (cameras or LiDARs) the past trajectories of the different elements of the scene and predict their future locations. We depart from the current trend of tackling this task via end-to-end training from perception to forecasting, and instead use a modular approach. We individually build and train detection, tracking and forecasting modules. We then only use consecutive finetuning steps to integrate the modules better and alleviate compounding errors. We conduct an in-depth study on the finetuning strategies and it reveals that our simple yet effective approach significantly improves performance on the end-to-end forecasting benchmark. Consequently, our solution ranks first in the Argoverse 2 End-to-end Forecasting Challenge, with 63.82 mAPf. We surpass forecasting results by +17.1 points over last year's winner and by +13.3 points over this year's runner-up. This remarkable performance in forecasting can be explained by our modular paradigm, which integrates finetuning strategies and significantly outperforms the end-to-end-trained counterparts. The code, model weights and results are made available https://github.com/valeoai/valeo4cast.
format Preprint
id arxiv_https___arxiv_org_abs_2406_08113
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Valeo4Cast: A Modular Approach to End-to-End Forecasting
Xu, Yihong
Zablocki, Éloi
Boulch, Alexandre
Puy, Gilles
Chen, Mickael
Bartoccioni, Florent
Samet, Nermin
Siméoni, Oriane
Gidaris, Spyros
Vu, Tuan-Hung
Bursuc, Andrei
Valle, Eduardo
Marlet, Renaud
Cord, Matthieu
Computer Vision and Pattern Recognition
Robotics
Motion forecasting is crucial in autonomous driving systems to anticipate the future trajectories of surrounding agents such as pedestrians, vehicles, and traffic signals. In end-to-end forecasting, the model must jointly detect and track from sensor data (cameras or LiDARs) the past trajectories of the different elements of the scene and predict their future locations. We depart from the current trend of tackling this task via end-to-end training from perception to forecasting, and instead use a modular approach. We individually build and train detection, tracking and forecasting modules. We then only use consecutive finetuning steps to integrate the modules better and alleviate compounding errors. We conduct an in-depth study on the finetuning strategies and it reveals that our simple yet effective approach significantly improves performance on the end-to-end forecasting benchmark. Consequently, our solution ranks first in the Argoverse 2 End-to-end Forecasting Challenge, with 63.82 mAPf. We surpass forecasting results by +17.1 points over last year's winner and by +13.3 points over this year's runner-up. This remarkable performance in forecasting can be explained by our modular paradigm, which integrates finetuning strategies and significantly outperforms the end-to-end-trained counterparts. The code, model weights and results are made available https://github.com/valeoai/valeo4cast.
title Valeo4Cast: A Modular Approach to End-to-End Forecasting
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2406.08113