DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wang, Zhechao, Zeng, Yiming, Ma, Lufan, Fu, Zeqing, Bai, Chen, Lin, Ziyao, Lu, Cheng
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910033673977856
author Wang, Zhechao
Zeng, Yiming
Ma, Lufan
Fu, Zeqing
Bai, Chen
Lin, Ziyao
Lu, Cheng
author_facet Wang, Zhechao
Zeng, Yiming
Ma, Lufan
Fu, Zeqing
Bai, Chen
Lin, Ziyao
Lu, Cheng
contents Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes as geometric conditions in diffusion models for conditional scene generation. However, implicit inter-condition dependency causes generation failures when control conditions change independently. Additionally, these methods suffer from insufficient details in both semantic and structural aspects. Specifically, brief and view-invariant captions restrict semantic contexts, resulting in weak background modeling. Meanwhile, the standard denoising loss with uniform spatial weighting neglects foreground structural details, causing visual distortions and blurriness. To address these challenges, we propose DrivePTS, which incorporates three key innovations. Firstly, our framework adopts a progressive learning strategy to mitigate inter-dependency between geometric conditions, reinforced by an explicit mutual information constraint. Secondly, a Vision-Language Model is utilized to generate multi-view hierarchical descriptions across six semantic aspects, providing fine-grained textual guidance. Thirdly, a frequency-guided structure loss is introduced to strengthen the model's sensitivity to high-frequency elements, improving foreground structural fidelity. Extensive experiments demonstrate that our DrivePTS achieves state-of-the-art fidelity and controllability in generating diverse driving scenes. Notably, DrivePTS successfully generates rare scenes where prior methods fail, highlighting its strong generalization ability.
format Preprint
id arxiv_https___arxiv_org_abs_2602_22549
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation
Wang, Zhechao
Zeng, Yiming
Ma, Lufan
Fu, Zeqing
Bai, Chen
Lin, Ziyao
Lu, Cheng
Computer Vision and Pattern Recognition
Artificial Intelligence
Synthesis of diverse driving scenes serves as a crucial data augmentation technique for validating the robustness and generalizability of autonomous driving systems. Current methods aggregate high-definition (HD) maps and 3D bounding boxes as geometric conditions in diffusion models for conditional scene generation. However, implicit inter-condition dependency causes generation failures when control conditions change independently. Additionally, these methods suffer from insufficient details in both semantic and structural aspects. Specifically, brief and view-invariant captions restrict semantic contexts, resulting in weak background modeling. Meanwhile, the standard denoising loss with uniform spatial weighting neglects foreground structural details, causing visual distortions and blurriness. To address these challenges, we propose DrivePTS, which incorporates three key innovations. Firstly, our framework adopts a progressive learning strategy to mitigate inter-dependency between geometric conditions, reinforced by an explicit mutual information constraint. Secondly, a Vision-Language Model is utilized to generate multi-view hierarchical descriptions across six semantic aspects, providing fine-grained textual guidance. Thirdly, a frequency-guided structure loss is introduced to strengthen the model's sensitivity to high-frequency elements, improving foreground structural fidelity. Extensive experiments demonstrate that our DrivePTS achieves state-of-the-art fidelity and controllability in generating diverse driving scenes. Notably, DrivePTS successfully generates rare scenes where prior methods fail, highlighting its strong generalization ability.
title DrivePTS: A Progressive Learning Framework with Textual and Structural Enhancement for Driving Scene Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2602.22549