MMInA: Benchmarking Multihop Multimodal Internet Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tian, Shulin, Zhang, Ziniu, Chen, Liangyu, Liu, Ziwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
von: Zhang, Fan, et al.
Veröffentlicht: (2024)
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
von: Li, Shilong, et al.
Veröffentlicht: (2025)
von: Li, Shilong, et al.
Veröffentlicht: (2025)
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026)
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026)
Large Multimodal Agents: A Survey
von: Xie, Junlin, et al.
Veröffentlicht: (2024)
von: Xie, Junlin, et al.
Veröffentlicht: (2024)
HippoCamp: Benchmarking Contextual Agents on Personal Computers
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
von: Yang, Zhe, et al.
Veröffentlicht: (2026)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
von: Xu, Yichen, et al.
Veröffentlicht: (2025)
RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
von: Chen, Haoyu, et al.
Veröffentlicht: (2024)
Benchmarking and Analyzing Generative Data for Visual Recognition
von: Li, Bo, et al.
Veröffentlicht: (2023)
von: Li, Bo, et al.
Veröffentlicht: (2023)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark
von: Zhang, Ge, et al.
Veröffentlicht: (2024)
von: Zhang, Ge, et al.
Veröffentlicht: (2024)
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation
von: Huang, Jinsheng, et al.
Veröffentlicht: (2024)
von: Huang, Jinsheng, et al.
Veröffentlicht: (2024)
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
Multi-Dimensional Insights: Benchmarking Real-World Personalization in Large Multimodal Models
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
von: Zhang, YiFan, et al.
Veröffentlicht: (2024)
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
von: Yue, Xiang, et al.
Veröffentlicht: (2023)
von: Yue, Xiang, et al.
Veröffentlicht: (2023)
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
Benchmarking Egocentric Clinical Intent Understanding Capability for Medical Multimodal Large Language Models
von: Liu, Shaonan, et al.
Veröffentlicht: (2026)
von: Liu, Shaonan, et al.
Veröffentlicht: (2026)
FM2DS: Few-Shot Multimodal Multihop Data Synthesis with Knowledge Distillation for Question Answering
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
VisualAgentBench: Towards Large Multimodal Models as Visual Foundation Agents
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
von: Liu, Xiao, et al.
Veröffentlicht: (2024)
A Multimodal Automated Interpretability Agent
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024)
von: Shaham, Tamar Rott, et al.
Veröffentlicht: (2024)
GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented Understanding
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
von: Gu, Zishan, et al.
Veröffentlicht: (2024)
MLLM-CompBench: A Comparative Reasoning Benchmark for Multimodal LLMs
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
von: Song, Ziyang, et al.
Veröffentlicht: (2025)
von: Song, Ziyang, et al.
Veröffentlicht: (2025)
A Survey on Benchmarks of Multimodal Large Language Models
von: Li, Jian, et al.
Veröffentlicht: (2024)
von: Li, Jian, et al.
Veröffentlicht: (2024)
MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
von: Fu, Chaoyou, et al.
Veröffentlicht: (2024)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
von: Chen, Xinyi, et al.
Veröffentlicht: (2023)
von: Chen, Xinyi, et al.
Veröffentlicht: (2023)
DART: Leveraging Multi-Agent Disagreement for Tool Recruitment in Multimodal Reasoning
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
von: Sivakumaran, Nithin, et al.
Veröffentlicht: (2025)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
von: Fang, Ye, et al.
Veröffentlicht: (2024)
von: Fang, Ye, et al.
Veröffentlicht: (2024)
MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
von: Jiang, Yulun, et al.
Veröffentlicht: (2025)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
von: Saxena, Rohit, et al.
Veröffentlicht: (2025)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
von: Patil, Vaidehi, et al.
Veröffentlicht: (2025)
From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping
von: Wu, Yu, et al.
Veröffentlicht: (2026)
von: Wu, Yu, et al.
Veröffentlicht: (2026)
FileGram: Grounding Agent Personalization in File-System Behavioral Traces
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
von: Liu, Shuai, et al.
Veröffentlicht: (2026)
ChartM$^3$: Benchmarking Chart Editing with Multimodal Instructions
von: Yang, Donglu, et al.
Veröffentlicht: (2025)
von: Yang, Donglu, et al.
Veröffentlicht: (2025)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
von: Yu, Shoubin, et al.
Veröffentlicht: (2026)
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
von: Zhang, Fan, et al.
Veröffentlicht: (2024) -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
von: Li, Shilong, et al.
Veröffentlicht: (2025) -
PerceptionComp: A Video Benchmark for Complex Perception-Centric Reasoning
von: Li, Shaoxuan, et al.
Veröffentlicht: (2026) -
Large Multimodal Agents: A Survey
von: Xie, Junlin, et al.
Veröffentlicht: (2024) -
HippoCamp: Benchmarking Contextual Agents on Personal Computers
von: Yang, Zhe, et al.
Veröffentlicht: (2026)