Gespeichert in:
| Hauptverfasser: | Zhao, Yi, Zhang, Yilin, Xiang, Rong, Li, Jing, Li, Hillming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2402.01735 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
von: Zhao, Yi, et al.
Veröffentlicht: (2026)
von: Zhao, Yi, et al.
Veröffentlicht: (2026)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
Large Multimodal Agents: A Survey
von: Xie, Junlin, et al.
Veröffentlicht: (2024)
von: Xie, Junlin, et al.
Veröffentlicht: (2024)
A Survey on Agentic Multimodal Large Language Models
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
An Examination of the Compositionality of Large Generative Vision-Language Models
von: Ma, Teli, et al.
Veröffentlicht: (2023)
von: Ma, Teli, et al.
Veröffentlicht: (2023)
CoViPAL: Layer-wise Contextualized Visual Token Pruning for Large Vision-Language Models
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
von: Tang, Zicong, et al.
Veröffentlicht: (2025)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
von: Song, Shezheng, et al.
Veröffentlicht: (2023)
A Survey on Evaluation of Multimodal Large Language Models
von: Huang, Jiaxing, et al.
Veröffentlicht: (2024)
von: Huang, Jiaxing, et al.
Veröffentlicht: (2024)
AI for Service: Proactive Assistance with AI Glasses
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
InsTALL: Context-aware Instructional Task Assistance with Multi-modal Large Language Models
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
von: Nguyen, Pha, et al.
Veröffentlicht: (2025)
A Survey on Multimodal Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
von: Yin, Shukang, et al.
Veröffentlicht: (2023)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
von: Gao, Silin, et al.
Veröffentlicht: (2025)
von: Gao, Silin, et al.
Veröffentlicht: (2025)
Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
von: Lee, Yi-Lun, et al.
Veröffentlicht: (2024)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
von: Chen, Qiguang, et al.
Veröffentlicht: (2026)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education
von: Wang, Junling, et al.
Veröffentlicht: (2026)
von: Wang, Junling, et al.
Veröffentlicht: (2026)
Speed Always Wins: A Survey on Efficient Architectures for Large Language Models
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
von: Sun, Weigao, et al.
Veröffentlicht: (2025)
SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
von: Hong, Jixiang, et al.
Veröffentlicht: (2025)
VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
A Survey of Multimodal Large Language Model from A Data-centric Perspective
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
von: Bai, Tianyi, et al.
Veröffentlicht: (2024)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
von: Lin, Xiao, et al.
Veröffentlicht: (2025)
Pedestrian Attribute Recognition: A New Benchmark Dataset and A Large Language Model Augmented Framework
von: Jin, Jiandong, et al.
Veröffentlicht: (2024)
von: Jin, Jiandong, et al.
Veröffentlicht: (2024)
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
von: Jia, Mengdi, et al.
Veröffentlicht: (2025)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Robust Multimodal Large Language Models Against Modality Conflict
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2025)
von: Zhang, Zongmeng, et al.
Veröffentlicht: (2025)
HyperLLaVA: Dynamic Visual and Language Expert Tuning for Multimodal Large Language Models
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
von: Zhang, Wenqiao, et al.
Veröffentlicht: (2024)
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
AviationLMM: A Large Multimodal Foundation Model for Civil Aviation
von: Li, Wenbin, et al.
Veröffentlicht: (2026)
von: Li, Wenbin, et al.
Veröffentlicht: (2026)
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
von: Zou, Chengke, et al.
Veröffentlicht: (2024)
From Pixels to Insights: A Survey on Automatic Chart Understanding in the Era of Large Foundation Models
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2024)
von: Huang, Kung-Hsiang, et al.
Veröffentlicht: (2024)
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhao, Haoyu, et al.
Veröffentlicht: (2024)
A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
von: Li, Zongxia, et al.
Veröffentlicht: (2025)
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
von: Ma, Xingjun, et al.
Veröffentlicht: (2025)
von: Ma, Xingjun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation
von: Zhao, Yi, et al.
Veröffentlicht: (2026) -
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024) -
Large Multimodal Agents: A Survey
von: Xie, Junlin, et al.
Veröffentlicht: (2024) -
A Survey on Agentic Multimodal Large Language Models
von: Yao, Huanjin, et al.
Veröffentlicht: (2025) -
An Examination of the Compositionality of Large Generative Vision-Language Models
von: Ma, Teli, et al.
Veröffentlicht: (2023)