Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shu, Shelley Zixin, Luo, Haozhe, Poellinger, Alexander, Reyes, Mauricio
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917013333475328
author Shu, Shelley Zixin
Luo, Haozhe
Poellinger, Alexander
Reyes, Mauricio
author_facet Shu, Shelley Zixin
Luo, Haozhe
Poellinger, Alexander
Reyes, Mauricio
contents Transformer-based deep learning models have demonstrated exceptional performance in medical imaging by leveraging attention mechanisms for feature representation and interpretability. However, these models are prone to learning spurious correlations, leading to biases and limited generalization. While human-AI attention alignment can mitigate these issues, it often depends on costly manual supervision. In this work, we propose a Hybrid Explanation-Guided Learning (H-EGL) framework that combines self-supervised and human-guided constraints to enhance attention alignment and improve generalization. The self-supervised component of H-EGL leverages class-distinctive attention without relying on restrictive priors, promoting robustness and flexibility. We validate our approach on chest X-ray classification using the Vision Transformer (ViT), where H-EGL outperforms two state-of-the-art Explanation-Guided Learning (EGL) methods, demonstrating superior classification accuracy and generalization capability. Additionally, it produces attention maps that are better aligned with human expertise.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12704
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis
Shu, Shelley Zixin
Luo, Haozhe
Poellinger, Alexander
Reyes, Mauricio
Computer Vision and Pattern Recognition
Artificial Intelligence
Transformer-based deep learning models have demonstrated exceptional performance in medical imaging by leveraging attention mechanisms for feature representation and interpretability. However, these models are prone to learning spurious correlations, leading to biases and limited generalization. While human-AI attention alignment can mitigate these issues, it often depends on costly manual supervision. In this work, we propose a Hybrid Explanation-Guided Learning (H-EGL) framework that combines self-supervised and human-guided constraints to enhance attention alignment and improve generalization. The self-supervised component of H-EGL leverages class-distinctive attention without relying on restrictive priors, promoting robustness and flexibility. We validate our approach on chest X-ray classification using the Vision Transformer (ViT), where H-EGL outperforms two state-of-the-art Explanation-Guided Learning (EGL) methods, demonstrating superior classification accuracy and generalization capability. Additionally, it produces attention maps that are better aligned with human expertise.
title Hybrid Explanation-Guided Learning for Transformer-Based Chest X-Ray Diagnosis
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2510.12704