DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yuhe, Zhang, Kun, Ma, Haoran, Yan, Rui, Li, Yingtai, Wang, Rongsheng, Zhou, Shaohua Kevin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917351626113024
author Tian, Yuhe
Zhang, Kun
Ma, Haoran
Yan, Rui
Li, Yingtai
Wang, Rongsheng
Zhou, Shaohua Kevin
author_facet Tian, Yuhe
Zhang, Kun
Ma, Haoran
Yan, Rui
Li, Yingtai
Wang, Rongsheng
Zhou, Shaohua Kevin
contents While large language models (LLMs) have advanced CT report generation, existing methods typically encode 3D volumes holistically, failing to distinguish informative cues from redundant anatomical background. Inspired by radiological cognitive subtraction, we propose Differential Visual Prompting (DiffVP), which conditions report generation on explicit, high-level semantic scan-to-reference differences rather than solely on absolute visual features. DiffVP employs a hierarchical difference extractor to capture complementary global and local semantic discrepancies into a shared latent space, along with a difference-to-prompt generator that transforms these signals into learnable visual prefix tokens for LLM conditioning. These difference prompts serve as structured conditioning signals that implicitly suppress invariant anatomy while amplifying diagnostically relevant visual evidence, thereby facilitating accurate report generation without explicit lesion localization. On two large-scale benchmarks, DiffVP consistently outperforms prior methods, improving the average BLEU-1-4 by +10.98 and +4.36, respectively, and further boosts clinical efficacy on RadGenome-ChestCT (F1 score 0.421). All codes will be released at https://github.com/ArielTYH/DiffVP/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_17718
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation
Tian, Yuhe
Zhang, Kun
Ma, Haoran
Yan, Rui
Li, Yingtai
Wang, Rongsheng
Zhou, Shaohua Kevin
Computer Vision and Pattern Recognition
While large language models (LLMs) have advanced CT report generation, existing methods typically encode 3D volumes holistically, failing to distinguish informative cues from redundant anatomical background. Inspired by radiological cognitive subtraction, we propose Differential Visual Prompting (DiffVP), which conditions report generation on explicit, high-level semantic scan-to-reference differences rather than solely on absolute visual features. DiffVP employs a hierarchical difference extractor to capture complementary global and local semantic discrepancies into a shared latent space, along with a difference-to-prompt generator that transforms these signals into learnable visual prefix tokens for LLM conditioning. These difference prompts serve as structured conditioning signals that implicitly suppress invariant anatomy while amplifying diagnostically relevant visual evidence, thereby facilitating accurate report generation without explicit lesion localization. On two large-scale benchmarks, DiffVP consistently outperforms prior methods, improving the average BLEU-1-4 by +10.98 and +4.36, respectively, and further boosts clinical efficacy on RadGenome-ChestCT (F1 score 0.421). All codes will be released at https://github.com/ArielTYH/DiffVP/.
title DiffVP: Differential Visual Semantic Prompting for LLM-Based CT Report Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.17718