Generative Spatiotemporal Data Augmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Jinfan, Luo, Lixin, Eum, Sungmin, Kwon, Heesung, Park, Jeong Joon
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915673292144640
author Zhou, Jinfan
Luo, Lixin
Eum, Sungmin
Kwon, Heesung
Park, Jeong Joon
author_facet Zhou, Jinfan
Luo, Lixin
Eum, Sungmin
Kwon, Heesung
Park, Jeong Joon
contents We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometric transforms or appearance perturbations, our method leverages off-the-shelf video diffusion models to generate realistic 3D spatial and temporal variations from a given image dataset. Incorporating these synthesized video clips as supplemental training data yields consistent performance gains in low-data settings, such as UAV-captured imagery where annotations are scarce. Beyond empirical improvements, we provide practical guidelines for (i) choosing an appropriate spatiotemporal generative setup, (ii) transferring annotations to synthetic frames, and (iii) addressing disocclusion - regions newly revealed and unlabeled in generated views. Experiments on COCO subsets and UAV-captured datasets show that, when applied judiciously, spatiotemporal augmentation broadens the data distribution along axes underrepresented by traditional and prior generative methods, offering an effective lever for improving model performance in data-scarce regimes.
format Preprint
id arxiv_https___arxiv_org_abs_2512_12508
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generative Spatiotemporal Data Augmentation
Zhou, Jinfan
Luo, Lixin
Eum, Sungmin
Kwon, Heesung
Park, Jeong Joon
Computer Vision and Pattern Recognition
Machine Learning
We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometric transforms or appearance perturbations, our method leverages off-the-shelf video diffusion models to generate realistic 3D spatial and temporal variations from a given image dataset. Incorporating these synthesized video clips as supplemental training data yields consistent performance gains in low-data settings, such as UAV-captured imagery where annotations are scarce. Beyond empirical improvements, we provide practical guidelines for (i) choosing an appropriate spatiotemporal generative setup, (ii) transferring annotations to synthetic frames, and (iii) addressing disocclusion - regions newly revealed and unlabeled in generated views. Experiments on COCO subsets and UAV-captured datasets show that, when applied judiciously, spatiotemporal augmentation broadens the data distribution along axes underrepresented by traditional and prior generative methods, offering an effective lever for improving model performance in data-scarce regimes.
title Generative Spatiotemporal Data Augmentation
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2512.12508