A Step towards Interpretable Multimodal AI Models with MultiFIX

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Malafaia, Mafalda, Schlender, Thalea, Alderliesten, Tanja, Bosman, Peter A. N.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908367260221440
author Malafaia, Mafalda
Schlender, Thalea
Alderliesten, Tanja
Bosman, Peter A. N.
author_facet Malafaia, Mafalda
Schlender, Thalea
Alderliesten, Tanja
Bosman, Peter A. N.
contents Real-world problems are often dependent on multiple data modalities, making multimodal fusion essential for leveraging diverse information sources. In high-stakes domains, such as in healthcare, understanding how each modality contributes to the prediction is critical to ensure trustworthy and interpretable AI models. We present MultiFIX, an interpretability-driven multimodal data fusion pipeline that explicitly engineers distinct features from different modalities and combines them to make the final prediction. Initially, only deep learning components are used to train a model from data. The black-box (deep learning) components are subsequently either explained using post-hoc methods such as Grad-CAM for images or fully replaced by interpretable blocks, namely symbolic expressions for tabular data, resulting in an explainable model. We study the use of MultiFIX using several training strategies for feature extraction and predictive modeling. Besides highlighting strengths and weaknesses of MultiFIX, experiments on a variety of synthetic datasets with varying degrees of interaction between modalities demonstrate that MultiFIX can generate multimodal models that can be used to accurately explain both the extracted features and their integration without compromising predictive performance.
format Preprint
id arxiv_https___arxiv_org_abs_2505_11262
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Step towards Interpretable Multimodal AI Models with MultiFIX
Malafaia, Mafalda
Schlender, Thalea
Alderliesten, Tanja
Bosman, Peter A. N.
Neural and Evolutionary Computing
Real-world problems are often dependent on multiple data modalities, making multimodal fusion essential for leveraging diverse information sources. In high-stakes domains, such as in healthcare, understanding how each modality contributes to the prediction is critical to ensure trustworthy and interpretable AI models. We present MultiFIX, an interpretability-driven multimodal data fusion pipeline that explicitly engineers distinct features from different modalities and combines them to make the final prediction. Initially, only deep learning components are used to train a model from data. The black-box (deep learning) components are subsequently either explained using post-hoc methods such as Grad-CAM for images or fully replaced by interpretable blocks, namely symbolic expressions for tabular data, resulting in an explainable model. We study the use of MultiFIX using several training strategies for feature extraction and predictive modeling. Besides highlighting strengths and weaknesses of MultiFIX, experiments on a variety of synthetic datasets with varying degrees of interaction between modalities demonstrate that MultiFIX can generate multimodal models that can be used to accurately explain both the extracted features and their integration without compromising predictive performance.
title A Step towards Interpretable Multimodal AI Models with MultiFIX
topic Neural and Evolutionary Computing
url https://arxiv.org/abs/2505.11262