AECV-Bench: Benchmarking Multimodal Models on Architectural and Engineering Drawings Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kondratenko, Aleksei, Birhane, Mussie, Hsain, Houssame E., Maciocci, Guido |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction
von: Mankodiya, Harsh, et al.
Veröffentlicht: (2026)
von: Mankodiya, Harsh, et al.
Veröffentlicht: (2026)
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
von: Kou, Qian, et al.
Veröffentlicht: (2026)
von: Kou, Qian, et al.
Veröffentlicht: (2026)
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
von: Chang, Yifan, et al.
Veröffentlicht: (2025)
von: Chang, Yifan, et al.
Veröffentlicht: (2025)
CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
von: Zhu, Qingqing, et al.
Veröffentlicht: (2026)
von: Zhu, Qingqing, et al.
Veröffentlicht: (2026)
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)
PetroBench: A Benchmark for Large Language Models in Petroleum Engineering
von: Wang, Xiang, et al.
Veröffentlicht: (2026)
von: Wang, Xiang, et al.
Veröffentlicht: (2026)
ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
von: Wang, Yuhang, et al.
Veröffentlicht: (2026)
SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
von: Xu, Peiran, et al.
Veröffentlicht: (2025)
UmniBench: Unified Understand and Generation Model Oriented Omni-dimensional Benchmark
von: Liu, Kai, et al.
Veröffentlicht: (2025)
von: Liu, Kai, et al.
Veröffentlicht: (2025)
MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling
von: Li, Zhong, et al.
Veröffentlicht: (2026)
von: Li, Zhong, et al.
Veröffentlicht: (2026)
IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
von: Fu, Deqing, et al.
Veröffentlicht: (2024)
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
von: Sukthanker, Rhea Sanjay, et al.
Veröffentlicht: (2024)
von: Sukthanker, Rhea Sanjay, et al.
Veröffentlicht: (2024)
Adversarial Vulnerability Transcends Computational Paradigms: Feature Engineering Provides No Defense Against Neural Adversarial Transfer
von: Hsain, Achraf, et al.
Veröffentlicht: (2026)
von: Hsain, Achraf, et al.
Veröffentlicht: (2026)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
von: Cai, Mu, et al.
Veröffentlicht: (2024)
von: Cai, Mu, et al.
Veröffentlicht: (2024)
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
von: Doris, Anna C., et al.
Veröffentlicht: (2024)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
von: Wang, Zeyu, et al.
Veröffentlicht: (2026)
MMR-Bench: A Comprehensive Benchmark for Multimodal LLM Routing
von: Ma, Haoxuan, et al.
Veröffentlicht: (2026)
von: Ma, Haoxuan, et al.
Veröffentlicht: (2026)
EngiBench: A Benchmark for Evaluating Large Language Models on Engineering Problem Solving
von: Zhou, Xiyuan, et al.
Veröffentlicht: (2025)
von: Zhou, Xiyuan, et al.
Veröffentlicht: (2025)
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
von: Liu, Yunqi, et al.
Veröffentlicht: (2026)
von: Liu, Yunqi, et al.
Veröffentlicht: (2026)
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
von: Hu, He, et al.
Veröffentlicht: (2025)
von: Hu, He, et al.
Veröffentlicht: (2025)
RxnBench: A Multimodal Benchmark for Evaluating Large Language Models on Chemical Reaction Understanding from Scientific Literature
von: Li, Hanzheng, et al.
Veröffentlicht: (2025)
von: Li, Hanzheng, et al.
Veröffentlicht: (2025)
ALE-Bench: A Benchmark for Long-Horizon Objective-Driven Algorithm Engineering
von: Imajuku, Yuki, et al.
Veröffentlicht: (2025)
von: Imajuku, Yuki, et al.
Veröffentlicht: (2025)
FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning
von: Mao, Mingyang, et al.
Veröffentlicht: (2026)
von: Mao, Mingyang, et al.
Veröffentlicht: (2026)
MU-Bench: A Multitask Multimodal Benchmark for Machine Unlearning
von: Cheng, Jiali, et al.
Veröffentlicht: (2024)
von: Cheng, Jiali, et al.
Veröffentlicht: (2024)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
WXImpactBench: A Disruptive Weather Impact Understanding Benchmark for Evaluating Large Language Models
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
von: Yu, Yongan, et al.
Veröffentlicht: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
KnowMe-Bench: Benchmarking Person Understanding for Lifelong Digital Companions
von: Wu, Tingyu, et al.
Veröffentlicht: (2026)
von: Wu, Tingyu, et al.
Veröffentlicht: (2026)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
von: Zhu, Hengchuan, et al.
Veröffentlicht: (2025)
AttackSeqBench: Benchmarking the Capabilities of LLMs for Attack Sequences Understanding
von: Ma, Haokai, et al.
Veröffentlicht: (2025)
von: Ma, Haokai, et al.
Veröffentlicht: (2025)
MPCI-Bench: A Benchmark for Multimodal Pairwise Contextual Integrity Evaluation of Language Model Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2026)
von: Wang, Shouju, et al.
Veröffentlicht: (2026)
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
von: Pan, Yaning, et al.
Veröffentlicht: (2025)
von: Pan, Yaning, et al.
Veröffentlicht: (2025)
DrafterBench: Benchmarking Large Language Models for Tasks Automation in Civil Engineering
von: Li, Yinsheng, et al.
Veröffentlicht: (2025)
von: Li, Yinsheng, et al.
Veröffentlicht: (2025)
LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
AgriBench: A Hierarchical Agriculture Benchmark for Multimodal Large Language Models
von: Zhou, Yutong, et al.
Veröffentlicht: (2024)
von: Zhou, Yutong, et al.
Veröffentlicht: (2024)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
von: Deng, Andong, et al.
Veröffentlicht: (2025)
von: Deng, Andong, et al.
Veröffentlicht: (2025)
EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection
von: Xu, Ancheng, et al.
Veröffentlicht: (2025)
von: Xu, Ancheng, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
Do Multimodal Language Models Really Understand Direction? A Benchmark for Compass Direction Reasoning
von: Yin, Hang, et al.
Veröffentlicht: (2024)
von: Yin, Hang, et al.
Veröffentlicht: (2024)
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
von: Qiu, Jielin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction
von: Mankodiya, Harsh, et al.
Veröffentlicht: (2026) -
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
von: Kou, Qian, et al.
Veröffentlicht: (2026) -
SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
von: Chang, Yifan, et al.
Veröffentlicht: (2025) -
CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
von: Zhu, Qingqing, et al.
Veröffentlicht: (2026) -
II-Bench: An Image Implication Understanding Benchmark for Multimodal Large Language Models
von: Liu, Ziqiang, et al.
Veröffentlicht: (2024)