Bridging the gap between Performance and Interpretability: An Explainable Disentangled Multimodal Framework for Cancer Survival Prediction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Eijpe, Aniek, Lakbir, Soufyan, Cesur, Melis Erdal, Oliveira, Sara P., Chatzimparmpas, Angelos, Abeln, Sanne, Silva, Wilson
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918366814404608
author Eijpe, Aniek
Lakbir, Soufyan
Cesur, Melis Erdal
Oliveira, Sara P.
Chatzimparmpas, Angelos
Abeln, Sanne
Silva, Wilson
author_facet Eijpe, Aniek
Lakbir, Soufyan
Cesur, Melis Erdal
Oliveira, Sara P.
Chatzimparmpas, Angelos
Abeln, Sanne
Silva, Wilson
contents While multimodal survival prediction models are increasingly more accurate, their complexity often reduces interpretability, limiting insight into how different data sources influence predictions. To address this, we introduce DIMAFx, an explainable multimodal framework for cancer survival prediction that produces disentangled, interpretable modality-specific and modality-shared representations from histopathology whole-slide images and transcriptomics data. Across multiple cancer cohorts, DIMAFx achieves state-of-the-art performance and improved representation disentanglement. Leveraging its interpretable design and SHapley Additive exPlanations, DIMAFx systematically reveals key multimodal interactions and the biological information encoded in the disentangled representations. In breast cancer survival prediction, the most predictive features contain modality-shared information, including one capturing solid tumor morphology contextualized primarily by late estrogen response, where higher-grade morphology aligned with pathway upregulation and increased risk, consistent with known breast cancer biology. Key modality-specific features capture microenvironmental signals from interacting adipose and stromal morphologies. These results show that multimodal models can overcome the traditional trade-off between performance and explainability, supporting their application in precision medicine.
format Preprint
id arxiv_https___arxiv_org_abs_2603_02162
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bridging the gap between Performance and Interpretability: An Explainable Disentangled Multimodal Framework for Cancer Survival Prediction
Eijpe, Aniek
Lakbir, Soufyan
Cesur, Melis Erdal
Oliveira, Sara P.
Chatzimparmpas, Angelos
Abeln, Sanne
Silva, Wilson
Computer Vision and Pattern Recognition
While multimodal survival prediction models are increasingly more accurate, their complexity often reduces interpretability, limiting insight into how different data sources influence predictions. To address this, we introduce DIMAFx, an explainable multimodal framework for cancer survival prediction that produces disentangled, interpretable modality-specific and modality-shared representations from histopathology whole-slide images and transcriptomics data. Across multiple cancer cohorts, DIMAFx achieves state-of-the-art performance and improved representation disentanglement. Leveraging its interpretable design and SHapley Additive exPlanations, DIMAFx systematically reveals key multimodal interactions and the biological information encoded in the disentangled representations. In breast cancer survival prediction, the most predictive features contain modality-shared information, including one capturing solid tumor morphology contextualized primarily by late estrogen response, where higher-grade morphology aligned with pathway upregulation and increased risk, consistent with known breast cancer biology. Key modality-specific features capture microenvironmental signals from interacting adipose and stromal morphologies. These results show that multimodal models can overcome the traditional trade-off between performance and explainability, supporting their application in precision medicine.
title Bridging the gap between Performance and Interpretability: An Explainable Disentangled Multimodal Framework for Cancer Survival Prediction
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.02162