LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Karimov, Mirlan, Spasojevic, Teodora, Braun, Markus, Wiederer, Julian, Belagiannis, Vasileios, Pollefeys, Marc
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866908816402022400
author Karimov, Mirlan
Spasojevic, Teodora
Braun, Markus
Wiederer, Julian
Belagiannis, Vasileios
Pollefeys, Marc
author_facet Karimov, Mirlan
Spasojevic, Teodora
Braun, Markus
Wiederer, Julian
Belagiannis, Vasileios
Pollefeys, Marc
contents Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model towards temporally consistent generation of dynamic objects, limiting their utility as scalable and generalizable data engines. In this work, we propose Localized Semantic Alignment (LSA), a simple yet effective framework for fine-tuning pre-trained video generation models. LSA enhances temporal consistency by aligning semantic features between ground-truth and generated video clips. Specifically, we compare the output of an off-the-shelf feature extraction model between the ground-truth and generated video clips localized around dynamic objects inducing a semantic feature consistency loss. We fine-tune the base model by combining this loss with the standard diffusion loss. The model fine-tuned for a single epoch with our novel loss outperforms the baselines in common video generation evaluation metrics. To further test the temporal consistency in generated videos we adapt two additional metrics from object detection task, namely mAP and mIoU. Extensive experiments on nuScenes and KITTI datasets show the effectiveness of our approach in enhancing temporal consistency in video generation without the need for external control signals during inference and any computational overheads.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05966
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
Karimov, Mirlan
Spasojevic, Teodora
Braun, Markus
Wiederer, Julian
Belagiannis, Vasileios
Pollefeys, Marc
Computer Vision and Pattern Recognition
Artificial Intelligence
Controllable video generation has emerged as a versatile tool for autonomous driving, enabling realistic synthesis of traffic scenarios. However, existing methods depend on control signals at inference time to guide the generative model towards temporally consistent generation of dynamic objects, limiting their utility as scalable and generalizable data engines. In this work, we propose Localized Semantic Alignment (LSA), a simple yet effective framework for fine-tuning pre-trained video generation models. LSA enhances temporal consistency by aligning semantic features between ground-truth and generated video clips. Specifically, we compare the output of an off-the-shelf feature extraction model between the ground-truth and generated video clips localized around dynamic objects inducing a semantic feature consistency loss. We fine-tune the base model by combining this loss with the standard diffusion loss. The model fine-tuned for a single epoch with our novel loss outperforms the baselines in common video generation evaluation metrics. To further test the temporal consistency in generated videos we adapt two additional metrics from object detection task, namely mAP and mIoU. Extensive experiments on nuScenes and KITTI datasets show the effectiveness of our approach in enhancing temporal consistency in video generation without the need for external control signals during inference and any computational overheads.
title LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.05966