Unsupervised Cardiac Video Translation Via Motion Feature Guided Diffusion Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Deb, Swakshar, Wu, Nian, Epstein, Frederick H., Zhang, Miaomiao
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912468028096512
author Deb, Swakshar
Wu, Nian
Epstein, Frederick H.
Zhang, Miaomiao
author_facet Deb, Swakshar
Wu, Nian
Epstein, Frederick H.
Zhang, Miaomiao
contents This paper presents a novel motion feature guided diffusion model for unpaired video-to-video translation (MFD-V2V), designed to synthesize dynamic, high-contrast cine cardiac magnetic resonance (CMR) from lower-contrast, artifact-prone displacement encoding with stimulated echoes (DENSE) CMR sequences. To achieve this, we first introduce a Latent Temporal Multi-Attention (LTMA) registration network that effectively learns more accurate and consistent cardiac motions from cine CMR image videos. A multi-level motion feature guided diffusion model, equipped with a specialized Spatio-Temporal Motion Encoder (STME) to extract fine-grained motion conditioning, is then developed to improve synthesis quality and fidelity. We evaluate our method, MFD-V2V, on a comprehensive cardiac dataset, demonstrating superior performance over the state-of-the-art in both quantitative metrics and qualitative assessments. Furthermore, we show the benefits of our synthesized cine CMRs improving downstream clinical and analytical tasks, underscoring the broader impact of our approach. Our code is publicly available at https://github.com/SwaksharDeb/MFD-V2V.
format Preprint
id arxiv_https___arxiv_org_abs_2507_02003
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Cardiac Video Translation Via Motion Feature Guided Diffusion Model
Deb, Swakshar
Wu, Nian
Epstein, Frederick H.
Zhang, Miaomiao
Image and Video Processing
This paper presents a novel motion feature guided diffusion model for unpaired video-to-video translation (MFD-V2V), designed to synthesize dynamic, high-contrast cine cardiac magnetic resonance (CMR) from lower-contrast, artifact-prone displacement encoding with stimulated echoes (DENSE) CMR sequences. To achieve this, we first introduce a Latent Temporal Multi-Attention (LTMA) registration network that effectively learns more accurate and consistent cardiac motions from cine CMR image videos. A multi-level motion feature guided diffusion model, equipped with a specialized Spatio-Temporal Motion Encoder (STME) to extract fine-grained motion conditioning, is then developed to improve synthesis quality and fidelity. We evaluate our method, MFD-V2V, on a comprehensive cardiac dataset, demonstrating superior performance over the state-of-the-art in both quantitative metrics and qualitative assessments. Furthermore, we show the benefits of our synthesized cine CMRs improving downstream clinical and analytical tasks, underscoring the broader impact of our approach. Our code is publicly available at https://github.com/SwaksharDeb/MFD-V2V.
title Unsupervised Cardiac Video Translation Via Motion Feature Guided Diffusion Model
topic Image and Video Processing
url https://arxiv.org/abs/2507.02003