A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Shilin, An, Wenbin, Tian, Feng, Nan, Fang, Liu, Qidong, Liu, Jun, Shah, Nazaraf, Chen, Ping
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915069859725312
author Sun, Shilin
An, Wenbin
Tian, Feng
Nan, Fang
Liu, Qidong
Liu, Jun
Shah, Nazaraf
Chen, Ping
author_facet Sun, Shilin
An, Wenbin
Tian, Feng
Nan, Fang
Liu, Qidong
Liu, Jun
Shah, Nazaraf
Chen, Ping
contents Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challenges in interpreting the "black-box" nature of AI models. To address these concerns, eXplainable AI (XAI) has emerged with a focus on transparency and interpretability to enhance human understanding and trust in AI decision-making processes. In the context of multimodal data fusion and complex reasoning scenarios, the proposal of Multimodal eXplainable AI (MXAI) integrates multiple modalities for prediction and explanation tasks. Meanwhile, the advent of Large Language Models (LLMs) has led to remarkable breakthroughs in natural language processing, yet their complexity has further exacerbated the issue of MXAI. To gain key insights into the development of MXAI methods and provide crucial guidance for building more transparent, fair, and trustworthy AI systems, we review the MXAI methods from a historical perspective and categorize them across four eras: traditional machine learning, deep learning, discriminative foundation models, and generative LLMs. We also review evaluation metrics and datasets used in MXAI research, concluding with a discussion of future challenges and directions. A project related to this review has been created at https://github.com/ShilinSun/mxai_review.
format Preprint
id arxiv_https___arxiv_org_abs_2412_14056
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
Sun, Shilin
An, Wenbin
Tian, Feng
Nan, Fang
Liu, Qidong
Liu, Jun
Shah, Nazaraf
Chen, Ping
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
Artificial intelligence (AI) has rapidly developed through advancements in computational power and the growth of massive datasets. However, this progress has also heightened challenges in interpreting the "black-box" nature of AI models. To address these concerns, eXplainable AI (XAI) has emerged with a focus on transparency and interpretability to enhance human understanding and trust in AI decision-making processes. In the context of multimodal data fusion and complex reasoning scenarios, the proposal of Multimodal eXplainable AI (MXAI) integrates multiple modalities for prediction and explanation tasks. Meanwhile, the advent of Large Language Models (LLMs) has led to remarkable breakthroughs in natural language processing, yet their complexity has further exacerbated the issue of MXAI. To gain key insights into the development of MXAI methods and provide crucial guidance for building more transparent, fair, and trustworthy AI systems, we review the MXAI methods from a historical perspective and categorize them across four eras: traditional machine learning, deep learning, discriminative foundation models, and generative LLMs. We also review evaluation metrics and datasets used in MXAI research, concluding with a discussion of future challenges and directions. A project related to this review has been created at https://github.com/ShilinSun/mxai_review.
title A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Multimedia
url https://arxiv.org/abs/2412.14056