A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Shilin, An, Wenbin, Tian, Feng, Nan, Fang, Liu, Qidong, Liu, Jun, Shah, Nazaraf, Chen, Ping |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
von: An, Wenbin, et al.
Veröffentlicht: (2025)
von: An, Wenbin, et al.
Veröffentlicht: (2025)
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
von: Wang, Bing, et al.
Veröffentlicht: (2025)
von: Wang, Bing, et al.
Veröffentlicht: (2025)
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
von: An, Wenbin, et al.
Veröffentlicht: (2024)
von: An, Wenbin, et al.
Veröffentlicht: (2024)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
von: Song, Xiao, et al.
Veröffentlicht: (2023)
von: Song, Xiao, et al.
Veröffentlicht: (2023)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
von: Jiang, Jingjing, et al.
Veröffentlicht: (2025)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
von: Cai, Zhongang, et al.
Veröffentlicht: (2025)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
von: Wang, Kangsheng, et al.
Veröffentlicht: (2025)
von: Wang, Kangsheng, et al.
Veröffentlicht: (2025)
A New Hybrid Intelligent Approach for Multimodal Detection of Suspected Disinformation on TikTok
von: Guerrero-Sosa, Jared D. T., et al.
Veröffentlicht: (2025)
von: Guerrero-Sosa, Jared D. T., et al.
Veröffentlicht: (2025)
Foundations of Multisensory Artificial Intelligence
von: Liang, Paul Pu
Veröffentlicht: (2024)
von: Liang, Paul Pu
Veröffentlicht: (2024)
Advancing Grounded Multimodal Named Entity Recognition via LLM-Based Reformulation and Box-Based Segmentation
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
von: Li, Jinyuan, et al.
Veröffentlicht: (2024)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
von: Qi, Peng, et al.
Veröffentlicht: (2024)
von: Qi, Peng, et al.
Veröffentlicht: (2024)
Hierarchical Aligned Multimodal Learning for NER on Tweet Posts
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
von: Liu, Peipei, et al.
Veröffentlicht: (2023)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
Magnifier Prompt: Tackling Multimodal Hallucination via Extremely Simple Instructions
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
von: Fu, Yuhan, et al.
Veröffentlicht: (2024)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
von: Ku, Max, et al.
Veröffentlicht: (2025)
von: Ku, Max, et al.
Veröffentlicht: (2025)
LLMs Meet Multimodal Generation and Editing: A Survey
von: He, Yingqing, et al.
Veröffentlicht: (2024)
von: He, Yingqing, et al.
Veröffentlicht: (2024)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
von: Wang, Bing, et al.
Veröffentlicht: (2024)
von: Wang, Bing, et al.
Veröffentlicht: (2024)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
von: Bin, Yi, et al.
Veröffentlicht: (2024)
von: Bin, Yi, et al.
Veröffentlicht: (2024)
DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
von: Maji, Arijit, et al.
Veröffentlicht: (2025)
VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations
von: Galarnyk, Michael, et al.
Veröffentlicht: (2025)
von: Galarnyk, Michael, et al.
Veröffentlicht: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
von: Gan, Chengguang, et al.
Veröffentlicht: (2025)
Towards Robust Multimodal Sentiment Analysis with Incomplete Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
von: Masumura, Ryo, et al.
Veröffentlicht: (2025)
von: Masumura, Ryo, et al.
Veröffentlicht: (2025)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
von: Lv, Zheqi, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
von: Wang, Bing, et al.
Veröffentlicht: (2025)
von: Wang, Bing, et al.
Veröffentlicht: (2025)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
von: Ghaleb, Esam, et al.
Veröffentlicht: (2025)
von: Ghaleb, Esam, et al.
Veröffentlicht: (2025)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
von: Wu, Jiaying, et al.
Veröffentlicht: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
von: Chow, Wei, et al.
Veröffentlicht: (2025)
von: Chow, Wei, et al.
Veröffentlicht: (2025)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
Can We Edit Multimodal Large Language Models?
von: Cheng, Siyuan, et al.
Veröffentlicht: (2023)
von: Cheng, Siyuan, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
von: An, Wenbin, et al.
Veröffentlicht: (2025) -
Remember Past, Anticipate Future: Learning Continual Multimodal Misinformation Detectors
von: Wang, Bing, et al.
Veröffentlicht: (2025) -
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
von: An, Wenbin, et al.
Veröffentlicht: (2024) -
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026) -
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
von: Song, Xiao, et al.
Veröffentlicht: (2023)