Gespeichert in:
| Hauptverfasser: | Kou, Qian, Shi, Xiaofeng, Li, Yulin, Qiu, Xiaosong, Wang, Xinyang, Zhou, Hua, Dongxing, Cao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2605.30794 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
von: Pal, Aniket, et al.
Veröffentlicht: (2025)
von: Pal, Aniket, et al.
Veröffentlicht: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024)
von: Li, Keliang, et al.
Veröffentlicht: (2024)
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
von: Amirloo, Elmira, et al.
Veröffentlicht: (2024)
von: Amirloo, Elmira, et al.
Veröffentlicht: (2024)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
von: Baek, Jeonghun, et al.
Veröffentlicht: (2025)
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
von: Zeng, Tong, et al.
Veröffentlicht: (2025)
von: Zeng, Tong, et al.
Veröffentlicht: (2025)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
von: Wang, Youze, et al.
Veröffentlicht: (2025)
von: Wang, Youze, et al.
Veröffentlicht: (2025)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
von: Xie, Wulin, et al.
Veröffentlicht: (2025)
von: Xie, Wulin, et al.
Veröffentlicht: (2025)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
von: Yang, Sicheng, et al.
Veröffentlicht: (2026)
Existence Is Chaos: Enhancing 3D Human Motion Prediction with Uncertainty Consideration
von: Wang, Zhihao, et al.
Veröffentlicht: (2024)
von: Wang, Zhihao, et al.
Veröffentlicht: (2024)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
von: Ge, Jinchao, et al.
Veröffentlicht: (2025)
von: Ge, Jinchao, et al.
Veröffentlicht: (2025)
AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
von: Gauba, Aruna, et al.
Veröffentlicht: (2025)
von: Gauba, Aruna, et al.
Veröffentlicht: (2025)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
von: Lin, Weifeng, et al.
Veröffentlicht: (2024)
von: Lin, Weifeng, et al.
Veröffentlicht: (2024)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
von: Kurz, Paul Jonas, et al.
Veröffentlicht: (2026)
von: Kurz, Paul Jonas, et al.
Veröffentlicht: (2026)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
von: Wang, Chunwei, et al.
Veröffentlicht: (2024)
von: Wang, Chunwei, et al.
Veröffentlicht: (2024)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
von: Kahl, Kim-Celine, et al.
Veröffentlicht: (2024)
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
von: Ning, Shan, et al.
Veröffentlicht: (2026)
von: Ning, Shan, et al.
Veröffentlicht: (2026)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
von: Hou, Wenjun, et al.
Veröffentlicht: (2024)
von: Hou, Wenjun, et al.
Veröffentlicht: (2024)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
von: Shi, Mengqi, et al.
Veröffentlicht: (2026)
von: Shi, Mengqi, et al.
Veröffentlicht: (2026)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
von: Zhang, Chengyi, et al.
Veröffentlicht: (2026)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
MechVerse: Evaluating Physical Motion Consistency in Video Generation Models
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
von: Su, Xin, et al.
Veröffentlicht: (2024)
von: Su, Xin, et al.
Veröffentlicht: (2024)
Embodied Scene Understanding for Vision Language Models via MetaVQA
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
von: Wang, Weizhen, et al.
Veröffentlicht: (2025)
Is ChatGPT-5 Ready for Mammogram VQA?
von: Li, Qiang, et al.
Veröffentlicht: (2025)
von: Li, Qiang, et al.
Veröffentlicht: (2025)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
von: Zhou, Xunchu, et al.
Veröffentlicht: (2024)
von: Zhou, Xunchu, et al.
Veröffentlicht: (2024)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
von: Xia, Peng, et al.
Veröffentlicht: (2024)
von: Xia, Peng, et al.
Veröffentlicht: (2024)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
von: Hu, Yutao, et al.
Veröffentlicht: (2024)
von: Hu, Yutao, et al.
Veröffentlicht: (2024)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
von: Xia, Haotian, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
von: Zhou, Hao, et al.
Veröffentlicht: (2024)
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
von: Luo, Yulin, et al.
Veröffentlicht: (2025)
von: Luo, Yulin, et al.
Veröffentlicht: (2025)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
von: Zhang, Yulin, et al.
Veröffentlicht: (2026)
von: Zhang, Yulin, et al.
Veröffentlicht: (2026)
WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring
von: Habibpour, Mobin, et al.
Veröffentlicht: (2026)
von: Habibpour, Mobin, et al.
Veröffentlicht: (2026)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
von: Ma, Jiatong, et al.
Veröffentlicht: (2026)
von: Ma, Jiatong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
von: Pal, Aniket, et al.
Veröffentlicht: (2025) -
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
von: Rostamkhani, Mohammadmostafa, et al.
Veröffentlicht: (2024) -
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024) -
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
von: Liu, Bo, et al.
Veröffentlicht: (2025) -
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
von: Zhang, Weichen, et al.
Veröffentlicht: (2025)