Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
Fuente:
arXiv
Saved in:
| Main Authors: | An, Wenbin, Nie, Jiahao, Wu, Yaqiang, Tian, Feng, Lu, Shijian, Zheng, Qinghua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
by: An, Wenbin, et al.
Published: (2024)
by: An, Wenbin, et al.
Published: (2024)
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024)
by: He, Yingqing, et al.
Published: (2024)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025)
by: Lai, Zhengzhao, et al.
Published: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
by: Zhao, Zhixian, et al.
Published: (2026)
by: Zhao, Zhixian, et al.
Published: (2026)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
by: Bin, Yi, et al.
Published: (2024)
by: Bin, Yi, et al.
Published: (2024)
RaTA-Tool: Retrieval-based Tool Selection with Multimodal Large Language Models
by: Mattioli, Gabriele, et al.
Published: (2026)
by: Mattioli, Gabriele, et al.
Published: (2026)
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
by: Sun, Shilin, et al.
Published: (2024)
by: Sun, Shilin, et al.
Published: (2024)
The Revolution of Multimodal Large Language Models: A Survey
by: Caffagni, Davide, et al.
Published: (2024)
by: Caffagni, Davide, et al.
Published: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Rethinking Radiology Report Generation via Causal Inspired Counterfactual Augmentation
by: Song, Xiao, et al.
Published: (2023)
by: Song, Xiao, et al.
Published: (2023)
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
by: Wu, Jiaying, et al.
Published: (2025)
by: Wu, Jiaying, et al.
Published: (2025)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
by: Wang, Wenxuan, et al.
Published: (2025)
by: Wang, Wenxuan, et al.
Published: (2025)
Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective
by: Wang, Bing, et al.
Published: (2025)
by: Wang, Bing, et al.
Published: (2025)
Hierarchical Adaptive Expert for Multimodal Sentiment Analysis
by: Qin, Jiahao, et al.
Published: (2025)
by: Qin, Jiahao, et al.
Published: (2025)
Noise-Tolerant Learning for Audio-Visual Action Recognition
by: Han, Haochen, et al.
Published: (2022)
by: Han, Haochen, et al.
Published: (2022)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
MIPS at SemEval-2024 Task 3: Multimodal Emotion-Cause Pair Extraction in Conversations with Multimodal Language Models
by: Cheng, Zebang, et al.
Published: (2024)
by: Cheng, Zebang, et al.
Published: (2024)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
by: Bai, Tianyi, et al.
Published: (2024)
by: Bai, Tianyi, et al.
Published: (2024)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
by: Jiang, Ruixiang, et al.
Published: (2025)
by: Jiang, Ruixiang, et al.
Published: (2025)
Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
by: Jiang, Jingjing, et al.
Published: (2025)
by: Jiang, Jingjing, et al.
Published: (2025)
Harmfully Manipulated Images Matter in Multimodal Misinformation Detection
by: Wang, Bing, et al.
Published: (2024)
by: Wang, Bing, et al.
Published: (2024)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
by: Lv, Zheqi, et al.
Published: (2025)
by: Lv, Zheqi, et al.
Published: (2025)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
by: Zhang, Xueqiao, et al.
Published: (2025)
by: Zhang, Xueqiao, et al.
Published: (2025)
Audio-visual training for improved grounding in video-text LLMs
by: Sagare, Shivprasad, et al.
Published: (2024)
by: Sagare, Shivprasad, et al.
Published: (2024)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
by: Cai, Zhongang, et al.
Published: (2025)
by: Cai, Zhongang, et al.
Published: (2025)
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
by: Compagnoni, Alberto, et al.
Published: (2025)
by: Compagnoni, Alberto, et al.
Published: (2025)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
by: Caffagni, Davide, et al.
Published: (2024)
by: Caffagni, Davide, et al.
Published: (2024)
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
by: Wang, Kangsheng, et al.
Published: (2025)
by: Wang, Kangsheng, et al.
Published: (2025)
M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction
by: Gan, Chengguang, et al.
Published: (2025)
by: Gan, Chengguang, et al.
Published: (2025)
Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice
by: Westermann, Hannes, et al.
Published: (2024)
by: Westermann, Hannes, et al.
Published: (2024)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025)
by: Masumura, Ryo, et al.
Published: (2025)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
by: Song, Shezheng, et al.
Published: (2023)
by: Song, Shezheng, et al.
Published: (2023)
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue
by: Ghaleb, Esam, et al.
Published: (2025)
by: Ghaleb, Esam, et al.
Published: (2025)
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
by: Cocchi, Federico, et al.
Published: (2024)
by: Cocchi, Federico, et al.
Published: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
by: Jiang, Chaoya, et al.
Published: (2024)
by: Jiang, Chaoya, et al.
Published: (2024)
Similar Items
-
Knowledge Acquisition Disentanglement for Knowledge-based Visual Question Answering with Large Language Models
by: An, Wenbin, et al.
Published: (2024) -
LLMs Meet Multimodal Generation and Editing: A Survey
by: He, Yingqing, et al.
Published: (2024) -
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026) -
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
by: Lai, Zhengzhao, et al.
Published: (2025) -
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
by: Zhao, Zhixian, et al.
Published: (2026)