MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Kou, Qian, Shi, Xiaofeng, Li, Yulin, Qiu, Xiaosong, Wang, Xinyang, Zhou, Hua, Dongxing, Cao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025)
by: Pal, Aniket, et al.
Published: (2025)
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024)
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024)
by: Li, Keliang, et al.
Published: (2024)
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
by: Liu, Bo, et al.
Published: (2025)
by: Liu, Bo, et al.
Published: (2025)
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)
by: Zhang, Weichen, et al.
Published: (2025)
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
by: Amirloo, Elmira, et al.
Published: (2024)
by: Amirloo, Elmira, et al.
Published: (2024)
Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
by: Wang, Youze, et al.
Published: (2025)
by: Wang, Youze, et al.
Published: (2025)
Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
by: Zeng, Tong, et al.
Published: (2025)
by: Zeng, Tong, et al.
Published: (2025)
MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
by: Baek, Jeonghun, et al.
Published: (2025)
by: Baek, Jeonghun, et al.
Published: (2025)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
by: Xie, Wulin, et al.
Published: (2025)
by: Xie, Wulin, et al.
Published: (2025)
AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
by: Gauba, Aruna, et al.
Published: (2025)
by: Gauba, Aruna, et al.
Published: (2025)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
by: Ge, Jinchao, et al.
Published: (2025)
by: Ge, Jinchao, et al.
Published: (2025)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)
by: Kurz, Paul Jonas, et al.
Published: (2026)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
by: Wang, Chunwei, et al.
Published: (2024)
by: Wang, Chunwei, et al.
Published: (2024)
Existence Is Chaos: Enhancing 3D Human Motion Prediction with Uncertainty Consideration
by: Wang, Zhihao, et al.
Published: (2024)
by: Wang, Zhihao, et al.
Published: (2024)
Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
by: Lin, Weifeng, et al.
Published: (2024)
by: Lin, Weifeng, et al.
Published: (2024)
Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum
by: Ning, Shan, et al.
Published: (2026)
by: Ning, Shan, et al.
Published: (2026)
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding
by: Shi, Mengqi, et al.
Published: (2026)
by: Shi, Mengqi, et al.
Published: (2026)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering
by: Zhang, Chengyi, et al.
Published: (2026)
by: Zhang, Chengyi, et al.
Published: (2026)
SURE-VQA: Systematic Understanding of Robustness Evaluation in Medical VQA Tasks
by: Kahl, Kim-Celine, et al.
Published: (2024)
by: Kahl, Kim-Celine, et al.
Published: (2024)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
by: Zhou, Xingcheng, et al.
Published: (2026)
by: Zhou, Xingcheng, et al.
Published: (2026)
MechVerse: Evaluating Physical Motion Consistency in Video Generation Models
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
by: Wang, Xucheng, et al.
Published: (2026)
by: Wang, Xucheng, et al.
Published: (2026)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
by: Su, Xin, et al.
Published: (2024)
by: Su, Xin, et al.
Published: (2024)
Embodied Scene Understanding for Vision Language Models via MetaVQA
by: Wang, Weizhen, et al.
Published: (2025)
by: Wang, Weizhen, et al.
Published: (2025)
Visual Robustness Benchmark for Visual Question Answering (VQA)
by: Ishmam, Md Farhan, et al.
Published: (2024)
by: Ishmam, Md Farhan, et al.
Published: (2024)
SPORTU: A Comprehensive Sports Understanding Benchmark for Multimodal Large Language Models
by: Xia, Haotian, et al.
Published: (2024)
by: Xia, Haotian, et al.
Published: (2024)
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving
by: Zhou, Hao, et al.
Published: (2024)
by: Zhou, Hao, et al.
Published: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
by: Ouyang, Kun, et al.
Published: (2024)
by: Ouyang, Kun, et al.
Published: (2024)
OmniMedVQA: A New Large-Scale Comprehensive Evaluation Benchmark for Medical LVLM
by: Hu, Yutao, et al.
Published: (2024)
by: Hu, Yutao, et al.
Published: (2024)
Light-VQA+: A Video Quality Assessment Model for Exposure Correction with Vision-Language Guidance
by: Zhou, Xunchu, et al.
Published: (2024)
by: Zhou, Xunchu, et al.
Published: (2024)
Is ChatGPT-5 Ready for Mammogram VQA?
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
by: Luo, Yulin, et al.
Published: (2025)
by: Luo, Yulin, et al.
Published: (2025)
M$^3$-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
by: Ma, Jiatong, et al.
Published: (2026)
by: Ma, Jiatong, et al.
Published: (2026)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
by: Zhang, Yulin, et al.
Published: (2026)
by: Zhang, Yulin, et al.
Published: (2026)
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
by: Lin, Kevin Qinghong, et al.
Published: (2025)
by: Lin, Kevin Qinghong, et al.
Published: (2025)
WildFireVQA: A Large-Scale Radiometric Thermal VQA Benchmark for Aerial Wildfire Monitoring
by: Habibpour, Mobin, et al.
Published: (2026)
by: Habibpour, Mobin, et al.
Published: (2026)
Similar Items
-
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
by: Pal, Aniket, et al.
Published: (2025) -
Illusory VQA: Benchmarking and Enhancing Multimodal Models on Visual Illusions
by: Rostamkhani, Mohammadmostafa, et al.
Published: (2024) -
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
by: Li, Keliang, et al.
Published: (2024) -
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
by: Liu, Bo, et al.
Published: (2025) -
Open3D-VQA: A Benchmark for Comprehensive Spatial Reasoning with Multimodal Large Language Model in Open Space
by: Zhang, Weichen, et al.
Published: (2025)