Efficient Image-to-Image Schrödinger Bridge for CT Field of View Extension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhenhao, Ni, Song, Yang, Long, Yin, Xiaojie, Yu, Haijun, Wang, Jiazhou, Han, Hongbin, Hu, Weigang, Huang, Yixing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915922486231040
author Li, Zhenhao
Ni, Song
Yang, Long
Yin, Xiaojie
Yu, Haijun
Wang, Jiazhou
Han, Hongbin
Hu, Weigang
Huang, Yixing
author_facet Li, Zhenhao
Ni, Song
Yang, Long
Yin, Xiaojie
Yu, Haijun
Wang, Jiazhou
Han, Hongbin
Hu, Weigang
Huang, Yixing
contents Computed tomography (CT) is a cornerstone imaging modality for non-invasive, high-resolution visualization of internal anatomical structures. However, when the scanned object exceeds the scanner's field of view (FOV), projection data are truncated, resulting in incomplete reconstructions and pronounced artifacts near FOV boundaries. Conventional reconstruction algorithms struggle to recover accurate anatomy from such data, limiting clinical reliability. Deep learning approaches have been explored for FOV extension, with diffusion generative models representing the latest advances in image synthesis. Yet, conventional diffusion models are computationally demanding and slow at inference due to their iterative sampling process. To address these limitations, we propose an efficient CT FOV extension framework based on the image-to-image Schrödinger Bridge (I$^2$SB) diffusion model. Unlike traditional diffusion models that synthesize images from pure Gaussian noise, I$^2$SB learns a direct stochastic mapping between paired limited-FOV and extended-FOV images. This direct correspondence yields a more interpretable and traceable generative process, enhancing anatomical consistency and structural fidelity in reconstructions. I$^2$SB achieves superior quantitative performance, with root-mean-square error (RMSE) values of 49.8 HU on simulated noisy data and 152.0 HU on real data, outperforming state-of-the-art diffusion models such as conditional denoising diffusion probabilistic models (cDDPM) and patch-based diffusion methods. Moreover, its one-step inference enables reconstruction in just 0.19 s per 2D slice, representing over a 700-fold speedup compared to cDDPM (135 s) and surpassing DiffusionGAN (0.58 s), the second fastest. This combination of accuracy and efficiency indicates that I$^2$SB has potential for real-time or clinical deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11211
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Image-to-Image Schrödinger Bridge for CT Field of View Extension
Li, Zhenhao
Ni, Song
Yang, Long
Yin, Xiaojie
Yu, Haijun
Wang, Jiazhou
Han, Hongbin
Hu, Weigang
Huang, Yixing
Image and Video Processing
Computer Vision and Pattern Recognition
Computed tomography (CT) is a cornerstone imaging modality for non-invasive, high-resolution visualization of internal anatomical structures. However, when the scanned object exceeds the scanner's field of view (FOV), projection data are truncated, resulting in incomplete reconstructions and pronounced artifacts near FOV boundaries. Conventional reconstruction algorithms struggle to recover accurate anatomy from such data, limiting clinical reliability. Deep learning approaches have been explored for FOV extension, with diffusion generative models representing the latest advances in image synthesis. Yet, conventional diffusion models are computationally demanding and slow at inference due to their iterative sampling process. To address these limitations, we propose an efficient CT FOV extension framework based on the image-to-image Schrödinger Bridge (I$^2$SB) diffusion model. Unlike traditional diffusion models that synthesize images from pure Gaussian noise, I$^2$SB learns a direct stochastic mapping between paired limited-FOV and extended-FOV images. This direct correspondence yields a more interpretable and traceable generative process, enhancing anatomical consistency and structural fidelity in reconstructions. I$^2$SB achieves superior quantitative performance, with root-mean-square error (RMSE) values of 49.8 HU on simulated noisy data and 152.0 HU on real data, outperforming state-of-the-art diffusion models such as conditional denoising diffusion probabilistic models (cDDPM) and patch-based diffusion methods. Moreover, its one-step inference enables reconstruction in just 0.19 s per 2D slice, representing over a 700-fold speedup compared to cDDPM (135 s) and surpassing DiffusionGAN (0.58 s), the second fastest. This combination of accuracy and efficiency indicates that I$^2$SB has potential for real-time or clinical deployment.
title Efficient Image-to-Image Schrödinger Bridge for CT Field of View Extension
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.11211