Saved in:
Bibliographic Details
Main Authors: You, Xin, Zhang, Minghui, Zhang, Hanxiao, Yang, Jie, Navab, Nassir
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.17333
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918069571420160
author You, Xin
Zhang, Minghui
Zhang, Hanxiao
Yang, Jie
Navab, Nassir
author_facet You, Xin
Zhang, Minghui
Zhang, Hanxiao
Yang, Jie
Navab, Nassir
contents Temporal modeling on regular respiration-induced motions is crucial to image-guided clinical applications. Existing methods cannot simulate temporal motions unless high-dose imaging scans including starting and ending frames exist simultaneously. However, in the preoperative data acquisition stage, the slight movement of patients may result in dynamic backgrounds between the first and last frames in a respiratory period. This additional deviation can hardly be removed by image registration, thus affecting the temporal modeling. To address that limitation, we pioneeringly simulate the regular motion process via the image-to-video (I2V) synthesis framework, which animates with the first frame to forecast future frames of a given length. Besides, to promote the temporal consistency of animated videos, we devise the Temporal Differential Diffusion Model to generate temporal differential fields, which measure the relative differential representations between adjacent frames. The prompt attention layer is devised for fine-grained differential fields, and the field augmented layer is adopted to better interact these fields with the I2V framework, promoting more accurate temporal variation of synthesized videos. Extensive results on ACDC cardiac and 4D Lung datasets reveal that our approach simulates 4D videos along the intrinsic motion trajectory, rivaling other competitive methods on perceptual similarity and temporal consistency. Codes will be available soon.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17333
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Temporal Differential Fields for 4D Motion Modeling via Image-to-Video Synthesis
You, Xin
Zhang, Minghui
Zhang, Hanxiao
Yang, Jie
Navab, Nassir
Computer Vision and Pattern Recognition
Temporal modeling on regular respiration-induced motions is crucial to image-guided clinical applications. Existing methods cannot simulate temporal motions unless high-dose imaging scans including starting and ending frames exist simultaneously. However, in the preoperative data acquisition stage, the slight movement of patients may result in dynamic backgrounds between the first and last frames in a respiratory period. This additional deviation can hardly be removed by image registration, thus affecting the temporal modeling. To address that limitation, we pioneeringly simulate the regular motion process via the image-to-video (I2V) synthesis framework, which animates with the first frame to forecast future frames of a given length. Besides, to promote the temporal consistency of animated videos, we devise the Temporal Differential Diffusion Model to generate temporal differential fields, which measure the relative differential representations between adjacent frames. The prompt attention layer is devised for fine-grained differential fields, and the field augmented layer is adopted to better interact these fields with the I2V framework, promoting more accurate temporal variation of synthesized videos. Extensive results on ACDC cardiac and 4D Lung datasets reveal that our approach simulates 4D videos along the intrinsic motion trajectory, rivaling other competitive methods on perceptual similarity and temporal consistency. Codes will be available soon.
title Temporal Differential Fields for 4D Motion Modeling via Image-to-Video Synthesis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.17333