Vision-Language Models for Automated 3D PET/CT Report Generation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Jiao, Wenpei, Shang, Kun, Li, Hui, Yan, Ke, Zhang, Jiajin, Yang, Guangjie, Guo, Lijuan, Wan, Yan, Yang, Xing, Jin, Dakai, Xie, Zhaoheng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918217614622720
author Jiao, Wenpei
Shang, Kun
Li, Hui
Yan, Ke
Zhang, Jiajin
Yang, Guangjie
Guo, Lijuan
Wan, Yan
Yang, Xing
Jin, Dakai
Xie, Zhaoheng
author_facet Jiao, Wenpei
Shang, Kun
Li, Hui
Yan, Ke
Zhang, Jiajin
Yang, Guangjie
Guo, Lijuan
Wan, Yan
Yang, Xing
Jin, Dakai
Xie, Zhaoheng
contents Positron emission tomography/computed tomography (PET/CT) is essential in oncology, yet the rapid expansion of scanners has outpaced the availability of trained specialists, making automated PET/CT report generation (PETRG) increasingly important for reducing clinical workload. Compared with structural imaging (e.g., X-ray, CT, and MRI), functional PET poses distinct challenges: metabolic patterns vary with tracer physiology, and whole-body 3D contextual information is required rather than local-region interpretation. To advance PETRG, we propose PETRG-3D, an end-to-end 3D dual-branch framework that separately encodes PET and CT volumes and incorporates style-adaptive prompts to mitigate inter-hospital variability in reporting practices. We construct PETRG-Lym, a multi-center lymphoma dataset collected from four hospitals (824 reports w/ 245,509 paired PET/CT slices), and construct AutoPET-RG-Lym, a publicly accessible PETRG benchmark derived from open imaging data but equipped with new expert-written, clinically validated reports (135 cases). To assess clinical utility, we introduce PETRG-Score, a lymphoma-specific evaluation protocol that jointly measures metabolic and structural findings across curated anatomical regions. Experiments show that PETRG-3D substantially outperforms existing methods on both natural language metrics (e.g., +31.49\% ROUGE-L) and clinical efficacy metrics (e.g., +8.18\% PET-All), highlighting the benefits of volumetric dual-modality modeling and style-aware prompting. Overall, this work establishes a foundation for future PET/CT-specific models emphasizing disease-aware reasoning and clinically reliable evaluation. Codes, models, and AutoPET-RG-Lym will be released.
format Preprint
id arxiv_https___arxiv_org_abs_2511_20145
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Vision-Language Models for Automated 3D PET/CT Report Generation
Jiao, Wenpei
Shang, Kun
Li, Hui
Yan, Ke
Zhang, Jiajin
Yang, Guangjie
Guo, Lijuan
Wan, Yan
Yang, Xing
Jin, Dakai
Xie, Zhaoheng
Computer Vision and Pattern Recognition
Positron emission tomography/computed tomography (PET/CT) is essential in oncology, yet the rapid expansion of scanners has outpaced the availability of trained specialists, making automated PET/CT report generation (PETRG) increasingly important for reducing clinical workload. Compared with structural imaging (e.g., X-ray, CT, and MRI), functional PET poses distinct challenges: metabolic patterns vary with tracer physiology, and whole-body 3D contextual information is required rather than local-region interpretation. To advance PETRG, we propose PETRG-3D, an end-to-end 3D dual-branch framework that separately encodes PET and CT volumes and incorporates style-adaptive prompts to mitigate inter-hospital variability in reporting practices. We construct PETRG-Lym, a multi-center lymphoma dataset collected from four hospitals (824 reports w/ 245,509 paired PET/CT slices), and construct AutoPET-RG-Lym, a publicly accessible PETRG benchmark derived from open imaging data but equipped with new expert-written, clinically validated reports (135 cases). To assess clinical utility, we introduce PETRG-Score, a lymphoma-specific evaluation protocol that jointly measures metabolic and structural findings across curated anatomical regions. Experiments show that PETRG-3D substantially outperforms existing methods on both natural language metrics (e.g., +31.49\% ROUGE-L) and clinical efficacy metrics (e.g., +8.18\% PET-All), highlighting the benefits of volumetric dual-modality modeling and style-aware prompting. Overall, this work establishes a foundation for future PET/CT-specific models emphasizing disease-aware reasoning and clinically reliable evaluation. Codes, models, and AutoPET-RG-Lym will be released.
title Vision-Language Models for Automated 3D PET/CT Report Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.20145