Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Onari, Mohsen Abbaspour, Magister, Lucie Charlotte, Wu, Yaoxin, Lupi, Amalia, Creazzo, Dario, Tordin, Mattia, Di Donatantonio, Luigi, Quaia, Emilio, Zhang, Chao, Grau, Isel, Nobile, Marco S., Zhang, Yingqian, Liò, Pietro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908476803907584
author Onari, Mohsen Abbaspour
Magister, Lucie Charlotte
Wu, Yaoxin
Lupi, Amalia
Creazzo, Dario
Tordin, Mattia
Di Donatantonio, Luigi
Quaia, Emilio
Zhang, Chao
Grau, Isel
Nobile, Marco S.
Zhang, Yingqian
Liò, Pietro
author_facet Onari, Mohsen Abbaspour
Magister, Lucie Charlotte
Wu, Yaoxin
Lupi, Amalia
Creazzo, Dario
Tordin, Mattia
Di Donatantonio, Luigi
Quaia, Emilio
Zhang, Chao
Grau, Isel
Nobile, Marco S.
Zhang, Yingqian
Liò, Pietro
contents Distal myopathy represents a genetically heterogeneous group of skeletal muscle disorders with broad clinical manifestations, posing diagnostic challenges in radiology. To address this, we propose a novel multimodal attention-aware fusion architecture that combines features extracted from two distinct deep learning models, one capturing global contextual information and the other focusing on local details, representing complementary aspects of the input data. Uniquely, our approach integrates these features through an attention gate mechanism, enhancing both predictive performance and interpretability. Our method achieves a high classification accuracy on the BUSI benchmark and a proprietary distal myopathy dataset, while also generating clinically relevant saliency maps that support transparent decision-making in medical diagnosis. We rigorously evaluated interpretability through (1) functionally grounded metrics, coherence scoring against reference masks and incremental deletion analysis, and (2) application-grounded validation with seven expert radiologists. While our fusion strategy boosts predictive performance relative to single-stream and alternative fusion strategies, both quantitative and qualitative evaluations reveal persistent gaps in anatomical specificity and clinical usefulness of the interpretability. These findings highlight the need for richer, context-aware interpretability methods and human-in-the-loop feedback to meet clinicians' expectations in real-world diagnostic settings.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01316
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust
Onari, Mohsen Abbaspour
Magister, Lucie Charlotte
Wu, Yaoxin
Lupi, Amalia
Creazzo, Dario
Tordin, Mattia
Di Donatantonio, Luigi
Quaia, Emilio
Zhang, Chao
Grau, Isel
Nobile, Marco S.
Zhang, Yingqian
Liò, Pietro
Computer Vision and Pattern Recognition
Human-Computer Interaction
Distal myopathy represents a genetically heterogeneous group of skeletal muscle disorders with broad clinical manifestations, posing diagnostic challenges in radiology. To address this, we propose a novel multimodal attention-aware fusion architecture that combines features extracted from two distinct deep learning models, one capturing global contextual information and the other focusing on local details, representing complementary aspects of the input data. Uniquely, our approach integrates these features through an attention gate mechanism, enhancing both predictive performance and interpretability. Our method achieves a high classification accuracy on the BUSI benchmark and a proprietary distal myopathy dataset, while also generating clinically relevant saliency maps that support transparent decision-making in medical diagnosis. We rigorously evaluated interpretability through (1) functionally grounded metrics, coherence scoring against reference masks and incremental deletion analysis, and (2) application-grounded validation with seven expert radiologists. While our fusion strategy boosts predictive performance relative to single-stream and alternative fusion strategies, both quantitative and qualitative evaluations reveal persistent gaps in anatomical specificity and clinical usefulness of the interpretability. These findings highlight the need for richer, context-aware interpretability methods and human-in-the-loop feedback to meet clinicians' expectations in real-world diagnostic settings.
title Multimodal Attention-Aware Fusion for Diagnosing Distal Myopathy: Evaluating Model Interpretability and Clinician Trust
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
url https://arxiv.org/abs/2508.01316