VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Seyfarth, Marvin, Dar, Salman Ul Hassan, Frisch, Yannik, Wild, Philipp, Frey, Norbert, André, Florian, Engelhardt, Sandy
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915892425654272
author Seyfarth, Marvin
Dar, Salman Ul Hassan
Frisch, Yannik
Wild, Philipp
Frey, Norbert
André, Florian
Engelhardt, Sandy
author_facet Seyfarth, Marvin
Dar, Salman Ul Hassan
Frisch, Yannik
Wild, Philipp
Frey, Norbert
André, Florian
Engelhardt, Sandy
contents Diffusion models have become a leading approach for high-fidelity medical image synthesis. However, most existing methods for 3D medical image generation rely on convolutional U-Net backbones within latent diffusion frameworks. While effective, these architectures impose strong locality biases and limited receptive fields, which may constrain scalability, global context integration, and flexible conditioning. In this work, we introduce VolDiT, the first purely transformer-based 3D Diffusion Transformer for volumetric medical image synthesis. Our approach extends diffusion transformers to native 3D data through volumetric patch embeddings and global self-attention operating directly over 3D tokens. To enable structured control, we propose a timestep-gated control adapter that maps segmentation masks into learnable control tokens that modulate transformer layers during denoising. This token-level conditioning mechanism allows precise spatial guidance while preserving the modeling advantages of transformer architectures. We evaluate our model on high-resolution 3D medical image synthesis tasks and compare it to state-of-the-art 3D latent diffusion models based on U-Nets. Results demonstrate improved global coherence, superior generative fidelity, and enhanced controllability. Our findings suggest that fully transformerbased diffusion models provide a flexible foundation for volumetric medical image synthesis. The code and models trained on public data are available at https://github.com/Cardio-AI/voldit.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25181
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers
Seyfarth, Marvin
Dar, Salman Ul Hassan
Frisch, Yannik
Wild, Philipp
Frey, Norbert
André, Florian
Engelhardt, Sandy
Computer Vision and Pattern Recognition
Diffusion models have become a leading approach for high-fidelity medical image synthesis. However, most existing methods for 3D medical image generation rely on convolutional U-Net backbones within latent diffusion frameworks. While effective, these architectures impose strong locality biases and limited receptive fields, which may constrain scalability, global context integration, and flexible conditioning. In this work, we introduce VolDiT, the first purely transformer-based 3D Diffusion Transformer for volumetric medical image synthesis. Our approach extends diffusion transformers to native 3D data through volumetric patch embeddings and global self-attention operating directly over 3D tokens. To enable structured control, we propose a timestep-gated control adapter that maps segmentation masks into learnable control tokens that modulate transformer layers during denoising. This token-level conditioning mechanism allows precise spatial guidance while preserving the modeling advantages of transformer architectures. We evaluate our model on high-resolution 3D medical image synthesis tasks and compare it to state-of-the-art 3D latent diffusion models based on U-Nets. Results demonstrate improved global coherence, superior generative fidelity, and enhanced controllability. Our findings suggest that fully transformerbased diffusion models provide a flexible foundation for volumetric medical image synthesis. The code and models trained on public data are available at https://github.com/Cardio-AI/voldit.
title VolDiT: Controllable Volumetric Medical Image Synthesis with Diffusion Transformers
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.25181