A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jena, Rohit, Zope, Vedant, Chaudhari, Pratik, Gee, James C.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908566283091968
author Jena, Rohit
Zope, Vedant
Chaudhari, Pratik
Gee, James C.
author_facet Jena, Rohit
Zope, Vedant
Chaudhari, Pratik
Gee, James C.
contents In this work, we propose FFDP, a set of IO-aware non-GEMM fused kernels supplemented with a distributed framework for image registration at unprecedented scales. Image registration is an inverse problem fundamental to biomedical and life sciences, but algorithms have not scaled in tandem with image acquisition capabilities. Our framework complements existing model parallelism techniques proposed for large-scale transformer training by optimizing non-GEMM bottlenecks and enabling convolution-aware tensor sharding. We demonstrate unprecedented capabilities by performing multimodal registration of a 100 micron ex-vivo human brain MRI volume at native resolution - an inverse problem more than 570x larger than a standard clinical datum in about a minute using only 8 A6000 GPUs. FFDP accelerates existing state-of-the-art optimization and deep learning registration pipelines by upto 6 - 7x while reducing peak memory consumption by 20 - 59%. Comparative analysis on a 250 micron dataset shows that FFDP can fit upto 64x larger problems than existing SOTA on a single GPU, and highlights both the performance and efficiency gains of FFDP compared to SOTA image registration methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25044
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
Jena, Rohit
Zope, Vedant
Chaudhari, Pratik
Gee, James C.
Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
In this work, we propose FFDP, a set of IO-aware non-GEMM fused kernels supplemented with a distributed framework for image registration at unprecedented scales. Image registration is an inverse problem fundamental to biomedical and life sciences, but algorithms have not scaled in tandem with image acquisition capabilities. Our framework complements existing model parallelism techniques proposed for large-scale transformer training by optimizing non-GEMM bottlenecks and enabling convolution-aware tensor sharding. We demonstrate unprecedented capabilities by performing multimodal registration of a 100 micron ex-vivo human brain MRI volume at native resolution - an inverse problem more than 570x larger than a standard clinical datum in about a minute using only 8 A6000 GPUs. FFDP accelerates existing state-of-the-art optimization and deep learning registration pipelines by upto 6 - 7x while reducing peak memory consumption by 20 - 59%. Comparative analysis on a 250 micron dataset shows that FFDP can fit upto 64x larger problems than existing SOTA on a single GPU, and highlights both the performance and efficiency gains of FFDP compared to SOTA image registration methods.
title A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
topic Computer Vision and Pattern Recognition
Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2509.25044