Saved in:
Bibliographic Details
Main Authors: Deng, Xiwei, He, Xianchun, Bao, Jianfeng, Zhou, Yudan, Cai, Shuhui, Cai, Congbo, Chen, Zhong
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.18309
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913912607211520
author Deng, Xiwei
He, Xianchun
Bao, Jianfeng
Zhou, Yudan
Cai, Shuhui
Cai, Congbo
Chen, Zhong
author_facet Deng, Xiwei
He, Xianchun
Bao, Jianfeng
Zhou, Yudan
Cai, Shuhui
Cai, Congbo
Chen, Zhong
contents CT report generation (CTRG) aims to automatically generate diagnostic reports for 3D volumes, relieving clinicians' workload and improving patient care. Despite clinical value, existing works fail to effectively incorporate diagnostic information from multiple anatomical views and lack related clinical expertise essential for accurate and reliable diagnosis. To resolve these limitations, we propose a novel Multi-view perception Knowledge-enhanced TansfoRmer (MvKeTR) to mimic the diagnostic workflow of clinicians. Just as radiologists first examine CT scans from multiple planes, a Multi-View Perception Aggregator (MVPA) with view-aware attention is proposed to synthesize diagnostic information from multiple anatomical views effectively. Then, inspired by how radiologists further refer to relevant clinical records to guide diagnostic decision-making, a Cross-Modal Knowledge Enhancer (CMKE) is devised to retrieve the most similar reports based on the query volume to incorporate domain knowledge into the diagnosis procedure. Furthermore, instead of traditional MLPs, we employ Kolmogorov-Arnold Networks (KANs) as the fundamental building blocks of both modules, which exhibit superior parameter efficiency and reduced spectral bias to better capture high-frequency components critical for CT interpretation while mitigating overfitting. Extensive experiments on the public CTRG-Chest-548 K dataset demonstrate that our method outpaces prior state-of-the-art (SOTA) models across almost all metrics. The code is available at https://github.com/xiweideng/MvKeTR.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18309
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement
Deng, Xiwei
He, Xianchun
Bao, Jianfeng
Zhou, Yudan
Cai, Shuhui
Cai, Congbo
Chen, Zhong
Computer Vision and Pattern Recognition
Artificial Intelligence
CT report generation (CTRG) aims to automatically generate diagnostic reports for 3D volumes, relieving clinicians' workload and improving patient care. Despite clinical value, existing works fail to effectively incorporate diagnostic information from multiple anatomical views and lack related clinical expertise essential for accurate and reliable diagnosis. To resolve these limitations, we propose a novel Multi-view perception Knowledge-enhanced TansfoRmer (MvKeTR) to mimic the diagnostic workflow of clinicians. Just as radiologists first examine CT scans from multiple planes, a Multi-View Perception Aggregator (MVPA) with view-aware attention is proposed to synthesize diagnostic information from multiple anatomical views effectively. Then, inspired by how radiologists further refer to relevant clinical records to guide diagnostic decision-making, a Cross-Modal Knowledge Enhancer (CMKE) is devised to retrieve the most similar reports based on the query volume to incorporate domain knowledge into the diagnosis procedure. Furthermore, instead of traditional MLPs, we employ Kolmogorov-Arnold Networks (KANs) as the fundamental building blocks of both modules, which exhibit superior parameter efficiency and reduced spectral bias to better capture high-frequency components critical for CT interpretation while mitigating overfitting. Extensive experiments on the public CTRG-Chest-548 K dataset demonstrate that our method outpaces prior state-of-the-art (SOTA) models across almost all metrics. The code is available at https://github.com/xiweideng/MvKeTR.
title MvKeTR: Chest CT Report Generation with Multi-View Perception and Knowledge Enhancement
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.18309