Explaining Multi-modal Large Language Models by Analyzing their Vision Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Giulivi, Loris, Boracchi, Giacomo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet
by: Giulivi, Loris, et al.
Published: (2024)
by: Giulivi, Loris, et al.
Published: (2024)
SE3D: A Framework For Saliency Method Evaluation In 3D Imaging
by: Wiśniewski, Mariusz, et al.
Published: (2024)
by: Wiśniewski, Mariusz, et al.
Published: (2024)
An expert-driven data generation pipeline for histological images
by: Basla, Roberto, et al.
Published: (2024)
by: Basla, Roberto, et al.
Published: (2024)
MultiLink: Multi-class Structure Recovery via Agglomerative Clustering and Model Selection
by: Magri, Luca, et al.
Published: (2025)
by: Magri, Luca, et al.
Published: (2025)
Convolutional Set Transformer
by: Chinello, Federico, et al.
Published: (2025)
by: Chinello, Federico, et al.
Published: (2025)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
Sparks of Explainability: Recent Advancements in Explaining Large Vision Models
by: Fel, Thomas
Published: (2025)
by: Fel, Thomas
Published: (2025)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
by: Liao, Pan, et al.
Published: (2026)
by: Liao, Pan, et al.
Published: (2026)
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
by: Li, Nanxi, et al.
Published: (2026)
by: Li, Nanxi, et al.
Published: (2026)
DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models
by: Zhang, Yudong, et al.
Published: (2024)
by: Zhang, Yudong, et al.
Published: (2024)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025)
by: He, Hulingxiao, et al.
Published: (2025)
Temporal-consistent CAMs for Weakly Supervised Video Segmentation in Waste Sorting
by: Marelli, Andrea, et al.
Published: (2025)
by: Marelli, Andrea, et al.
Published: (2025)
PIF: Anomaly detection via preference embedding
by: Leveni, Filippo, et al.
Published: (2025)
by: Leveni, Filippo, et al.
Published: (2025)
Hashing for Structure-based Anomaly Detection
by: Leveni, Filippo, et al.
Published: (2025)
by: Leveni, Filippo, et al.
Published: (2025)
Preference Isolation Forest for Structure-based Anomaly Detection
by: Leveni, Filippo, et al.
Published: (2025)
by: Leveni, Filippo, et al.
Published: (2025)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
by: Ohi, Masanari, et al.
Published: (2024)
by: Ohi, Masanari, et al.
Published: (2024)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
MVP-Bench: Can Large Vision--Language Models Conduct Multi-level Visual Perception Like Humans?
by: Li, Guanzhen, et al.
Published: (2024)
by: Li, Guanzhen, et al.
Published: (2024)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
by: Li, Yanwei, et al.
Published: (2024)
by: Li, Yanwei, et al.
Published: (2024)
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
by: Chen, Haokun, et al.
Published: (2025)
by: Chen, Haokun, et al.
Published: (2025)
MultiEmo-Bench: Multi-label Visual Emotion Analysis for Multi-modal Large Language Models
by: Chen, Tianwei, et al.
Published: (2026)
by: Chen, Tianwei, et al.
Published: (2026)
Adaptation of Multi-modal Representation Models for Multi-task Surgical Computer Vision
by: Walimbe, Soham, et al.
Published: (2025)
by: Walimbe, Soham, et al.
Published: (2025)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
Advancing Visual Large Language Model for Multi-granular Versatile Perception
by: Xiang, Wentao, et al.
Published: (2025)
by: Xiang, Wentao, et al.
Published: (2025)
ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
by: Choi, Sangbum, et al.
Published: (2025)
by: Choi, Sangbum, et al.
Published: (2025)
Visual Hallucinations of Multi-modal Large Language Models
by: Huang, Wen, et al.
Published: (2024)
by: Huang, Wen, et al.
Published: (2024)
Revealing Multi-View Hallucination in Large Vision-Language Models
by: Park, Wooje, et al.
Published: (2026)
by: Park, Wooje, et al.
Published: (2026)
A Deep Learning Pipeline for Solid Waste Detection in Remote Sensing Images
by: Gibellini, Federico, et al.
Published: (2025)
by: Gibellini, Federico, et al.
Published: (2025)
Learning Emergent Modular Representations in Multi-modality Medical Vision Foundation Models
by: He, Yuting, et al.
Published: (2026)
by: He, Yuting, et al.
Published: (2026)
Explaining multimodal LLMs via intra-modal token interactions
by: Liang, Jiawei, et al.
Published: (2025)
by: Liang, Jiawei, et al.
Published: (2025)
AirCache: Activating Inter-modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model Inference
by: Huang, Kai, et al.
Published: (2025)
by: Huang, Kai, et al.
Published: (2025)
Towards Cross-modal Backward-compatible Representation Learning for Vision-Language Models
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
FakeShield: Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models
by: Xu, Zhipei, et al.
Published: (2024)
by: Xu, Zhipei, et al.
Published: (2024)
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models
by: Mou, Tingshu, et al.
Published: (2026)
by: Mou, Tingshu, et al.
Published: (2026)
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability
by: Zhao, Fei, et al.
Published: (2024)
by: Zhao, Fei, et al.
Published: (2024)
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
by: Liu, Chonghan, et al.
Published: (2025)
by: Liu, Chonghan, et al.
Published: (2025)
Think Hierarchically, Act Dynamically: Hierarchical Multi-modal Fusion and Reasoning for Vision-and-Language Navigation
by: Yue, Junrong, et al.
Published: (2025)
by: Yue, Junrong, et al.
Published: (2025)
Medical Large Vision Language Models with Multi-Image Visual Ability
by: Yang, Xikai, et al.
Published: (2025)
by: Yang, Xikai, et al.
Published: (2025)
PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
by: Sun, Fengyuan, et al.
Published: (2025)
by: Sun, Fengyuan, et al.
Published: (2025)
Why Only Text: Empowering Vision-and-Language Navigation with Multi-modal Prompts
by: Hong, Haodong, et al.
Published: (2024)
by: Hong, Haodong, et al.
Published: (2024)
Similar Items
-
Concept Visualization: Explaining the CLIP Multi-modal Embedding Using WordNet
by: Giulivi, Loris, et al.
Published: (2024) -
SE3D: A Framework For Saliency Method Evaluation In 3D Imaging
by: Wiśniewski, Mariusz, et al.
Published: (2024) -
An expert-driven data generation pipeline for histological images
by: Basla, Roberto, et al.
Published: (2024) -
MultiLink: Multi-class Structure Recovery via Agglomerative Clustering and Model Selection
by: Magri, Luca, et al.
Published: (2025) -
Convolutional Set Transformer
by: Chinello, Federico, et al.
Published: (2025)