TopoDiffuser: A Diffusion-Based Multimodal Trajectory Prediction Model with Topometric Maps

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Zehui, Wang, Junhui, Shi, Yongliang, Gao, Chao, Zhou, Guyue
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916875036786688
author Xu, Zehui
Wang, Junhui
Shi, Yongliang
Gao, Chao
Zhou, Guyue
author_facet Xu, Zehui
Wang, Junhui
Shi, Yongliang
Gao, Chao
Zhou, Guyue
contents This paper introduces TopoDiffuser, a diffusion-based framework for multimodal trajectory prediction that incorporates topometric maps to generate accurate, diverse, and road-compliant future motion forecasts. By embedding structural cues from topometric maps into the denoising process of a conditional diffusion model, the proposed approach enables trajectory generation that naturally adheres to road geometry without relying on explicit constraints. A multimodal conditioning encoder fuses LiDAR observations, historical motion, and route information into a unified bird's-eye-view (BEV) representation. Extensive experiments on the KITTI benchmark demonstrate that TopoDiffuser outperforms state-of-the-art methods, while maintaining strong geometric consistency. Ablation studies further validate the contribution of each input modality, as well as the impact of denoising steps and the number of trajectory samples. To support future research, we publicly release our code at https://github.com/EI-Nav/TopoDiffuser.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00303
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TopoDiffuser: A Diffusion-Based Multimodal Trajectory Prediction Model with Topometric Maps
Xu, Zehui
Wang, Junhui
Shi, Yongliang
Gao, Chao
Zhou, Guyue
Robotics
This paper introduces TopoDiffuser, a diffusion-based framework for multimodal trajectory prediction that incorporates topometric maps to generate accurate, diverse, and road-compliant future motion forecasts. By embedding structural cues from topometric maps into the denoising process of a conditional diffusion model, the proposed approach enables trajectory generation that naturally adheres to road geometry without relying on explicit constraints. A multimodal conditioning encoder fuses LiDAR observations, historical motion, and route information into a unified bird's-eye-view (BEV) representation. Extensive experiments on the KITTI benchmark demonstrate that TopoDiffuser outperforms state-of-the-art methods, while maintaining strong geometric consistency. Ablation studies further validate the contribution of each input modality, as well as the impact of denoising steps and the number of trajectory samples. To support future research, we publicly release our code at https://github.com/EI-Nav/TopoDiffuser.
title TopoDiffuser: A Diffusion-Based Multimodal Trajectory Prediction Model with Topometric Maps
topic Robotics
url https://arxiv.org/abs/2508.00303