Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fu, Danzhen, Hu, Jiagao, Zhou, Daiguo, Wang, Fei, Wang, Zepeng, Liao, Wenhua
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911086844837888
author Fu, Danzhen
Hu, Jiagao
Zhou, Daiguo
Wang, Fei
Wang, Zepeng
Liao, Wenhua
author_facet Fu, Danzhen
Hu, Jiagao
Zhou, Daiguo
Wang, Fei
Wang, Zepeng
Liao, Wenhua
contents Pedestrian detection models in autonomous driving systems often lack robustness due to insufficient representation of dangerous pedestrian scenarios in training datasets. To address this limitation, we present a novel framework for controllable pedestrian video editing in multi-view driving scenarios by integrating video inpainting and human motion control techniques. Our approach begins by identifying pedestrian regions of interest across multiple camera views, expanding detection bounding boxes with a fixed ratio, and resizing and stitching these regions into a unified canvas while preserving cross-view spatial relationships. A binary mask is then applied to designate the editable area, within which pedestrian editing is guided by pose sequence control conditions. This enables flexible editing functionalities, including pedestrian insertion, replacement, and removal. Extensive experiments demonstrate that our framework achieves high-quality pedestrian editing with strong visual realism, spatiotemporal coherence, and cross-view consistency. These results establish the proposed method as a robust and versatile solution for multi-view pedestrian video generation, with broad potential for applications in data augmentation and scenario simulation in autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2508_00299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
Fu, Danzhen
Hu, Jiagao
Zhou, Daiguo
Wang, Fei
Wang, Zepeng
Liao, Wenhua
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Pedestrian detection models in autonomous driving systems often lack robustness due to insufficient representation of dangerous pedestrian scenarios in training datasets. To address this limitation, we present a novel framework for controllable pedestrian video editing in multi-view driving scenarios by integrating video inpainting and human motion control techniques. Our approach begins by identifying pedestrian regions of interest across multiple camera views, expanding detection bounding boxes with a fixed ratio, and resizing and stitching these regions into a unified canvas while preserving cross-view spatial relationships. A binary mask is then applied to designate the editable area, within which pedestrian editing is guided by pose sequence control conditions. This enables flexible editing functionalities, including pedestrian insertion, replacement, and removal. Extensive experiments demonstrate that our framework achieves high-quality pedestrian editing with strong visual realism, spatiotemporal coherence, and cross-view consistency. These results establish the proposed method as a robust and versatile solution for multi-view pedestrian video generation, with broad potential for applications in data augmentation and scenario simulation in autonomous driving.
title Controllable Pedestrian Video Editing for Multi-View Driving Scenarios via Motion Sequence
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2508.00299