Turbo4DGen: Ultra-Fast Acceleration for 4D Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Man, Yuanbin, Huang, Ying, Ren, Zhile, Yin, Miao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910088594194432
author Man, Yuanbin
Huang, Ying
Ren, Zhile
Yin, Miao
author_facet Man, Yuanbin
Huang, Ying
Ren, Zhile
Yin, Miao
contents 4D generation, or dynamic 3D content generation, integrates spatial, temporal, and view dimensions to model realistic dynamic scenes, playing a foundational role in advancing world models and physical AI. However, maintaining long-chain consistency across both frames and viewpoints through the unique spatio-camera-motion (SCM) attention mechanism introduces substantial computational and memory overhead, often leading to out-of-memory (OOM) failures and prohibitive generation times. To address these challenges, we propose Turbo4DGen, an ultra-fast acceleration framework for diffusion-based multi-view 4D content generation. Turbo4DGen introduces a spatiotemporal cache mechanism that persistently reuses intermediate attention across denoising steps, combined with dynamically semantic-aware attention pruning and an adaptive SCM chain bypass scheduler, to drastically reduce redundant SCM attention computation. Our experimental results show that Turbo4DGen achieves an average 9.7$\times$ speedup without quality degradation on the ObjaverseDy and Consistent4D datasets. To the best of our knowledge, Turbo4DGen is the first dedicated acceleration framework for 4D generation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29572
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
Man, Yuanbin
Huang, Ying
Ren, Zhile
Yin, Miao
Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
4D generation, or dynamic 3D content generation, integrates spatial, temporal, and view dimensions to model realistic dynamic scenes, playing a foundational role in advancing world models and physical AI. However, maintaining long-chain consistency across both frames and viewpoints through the unique spatio-camera-motion (SCM) attention mechanism introduces substantial computational and memory overhead, often leading to out-of-memory (OOM) failures and prohibitive generation times. To address these challenges, we propose Turbo4DGen, an ultra-fast acceleration framework for diffusion-based multi-view 4D content generation. Turbo4DGen introduces a spatiotemporal cache mechanism that persistently reuses intermediate attention across denoising steps, combined with dynamically semantic-aware attention pruning and an adaptive SCM chain bypass scheduler, to drastically reduce redundant SCM attention computation. Our experimental results show that Turbo4DGen achieves an average 9.7$\times$ speedup without quality degradation on the ObjaverseDy and Consistent4D datasets. To the best of our knowledge, Turbo4DGen is the first dedicated acceleration framework for 4D generation.
title Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
topic Graphics
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.29572