FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pham, Trong Thang, Ho, Ngoc-Vuong, Bui, Nhat-Tan, Phan, Thinh, Brijesh, Patel, Adjeroh, Donald, Doretto, Gianfranco, Nguyen, Anh, Wu, Carol C., Nguyen, Hien, Le, Ngan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909401054445568
author Pham, Trong Thang
Ho, Ngoc-Vuong
Bui, Nhat-Tan
Phan, Thinh
Brijesh, Patel
Adjeroh, Donald
Doretto, Gianfranco
Nguyen, Anh
Wu, Carol C.
Nguyen, Hien
Le, Ngan
author_facet Pham, Trong Thang
Ho, Ngoc-Vuong
Bui, Nhat-Tan
Phan, Thinh
Brijesh, Patel
Adjeroh, Donald
Doretto, Gianfranco
Nguyen, Anh
Wu, Carol C.
Nguyen, Hien
Le, Ngan
contents Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems. Despite the growth of diverse datasets and methods focusing on report generation, there remains a notable gap in how closely these models' generated reports align with the interpretations of real radiologists. In this study, we tackle this challenge by initially introducing Fine-Grained CXR (FG-CXR) dataset, which provides fine-grained paired information between the captions generated by radiologists and the corresponding gaze attention heatmaps for each anatomy. Unlike existing datasets that include a raw sequence of gaze alongside a report, with significant misalignment between gaze location and report content, our FG-CXR dataset offers a more grained alignment between gaze attention and diagnosis transcript. Furthermore, our analysis reveals that simply applying black-box image captioning methods to generate reports cannot adequately explain which information in CXR is utilized and how long needs to attend to accurately generate reports. Consequently, we propose a novel explainable radiologist's attention generator network (Gen-XAI) that mimics the diagnosis process of radiologists, explicitly constraining its output to closely align with both radiologist's gaze attention and transcript. Finally, we perform extensive experiments to illustrate the effectiveness of our method. Our datasets and checkpoint is available at https://github.com/UARK-AICV/FG-CXR.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15413
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
Pham, Trong Thang
Ho, Ngoc-Vuong
Bui, Nhat-Tan
Phan, Thinh
Brijesh, Patel
Adjeroh, Donald
Doretto, Gianfranco
Nguyen, Anh
Wu, Carol C.
Nguyen, Hien
Le, Ngan
Computer Vision and Pattern Recognition
Artificial Intelligence
Developing an interpretable system for generating reports in chest X-ray (CXR) analysis is becoming increasingly crucial in Computer-aided Diagnosis (CAD) systems, enabling radiologists to comprehend the decisions made by these systems. Despite the growth of diverse datasets and methods focusing on report generation, there remains a notable gap in how closely these models' generated reports align with the interpretations of real radiologists. In this study, we tackle this challenge by initially introducing Fine-Grained CXR (FG-CXR) dataset, which provides fine-grained paired information between the captions generated by radiologists and the corresponding gaze attention heatmaps for each anatomy. Unlike existing datasets that include a raw sequence of gaze alongside a report, with significant misalignment between gaze location and report content, our FG-CXR dataset offers a more grained alignment between gaze attention and diagnosis transcript. Furthermore, our analysis reveals that simply applying black-box image captioning methods to generate reports cannot adequately explain which information in CXR is utilized and how long needs to attend to accurately generate reports. Consequently, we propose a novel explainable radiologist's attention generator network (Gen-XAI) that mimics the diagnosis process of radiologists, explicitly constraining its output to closely align with both radiologist's gaze attention and transcript. Finally, we perform extensive experiments to illustrate the effectiveness of our method. Our datasets and checkpoint is available at https://github.com/UARK-AICV/FG-CXR.
title FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2411.15413