Gespeichert in:
| Hauptverfasser: | Heap, Thomas, Aitchison, Laurence, Cahill, Emma, Rodriguez, Adriana Casado |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.18540 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
APML: Adaptive Probabilistic Matching Loss for Robust 3D Point Cloud Reconstruction
von: Sharifipour, Sasan, et al.
Veröffentlicht: (2025)
von: Sharifipour, Sasan, et al.
Veröffentlicht: (2025)
PushupBench: Your VLM is not good at counting pushups
von: Li, Shengzhi, et al.
Veröffentlicht: (2026)
von: Li, Shengzhi, et al.
Veröffentlicht: (2026)
Video-Bench: Human-Aligned Video Generation Benchmark
von: Han, Hui, et al.
Veröffentlicht: (2025)
von: Han, Hui, et al.
Veröffentlicht: (2025)
VEU-Bench: Towards Comprehensive Understanding of Video Editing
von: Li, Bozheng, et al.
Veröffentlicht: (2025)
von: Li, Bozheng, et al.
Veröffentlicht: (2025)
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)
Seeing the Big Picture: Evaluating Multimodal LLMs' Ability to Interpret and Grade Handwritten Student Work
von: Henkel, Owen, et al.
Veröffentlicht: (2025)
von: Henkel, Owen, et al.
Veröffentlicht: (2025)
JourneyBench: A Challenging One-Stop Vision-Language Understanding Benchmark of Generated Images
von: Wang, Zhecan, et al.
Veröffentlicht: (2024)
von: Wang, Zhecan, et al.
Veröffentlicht: (2024)
AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
von: Li, Danrui, et al.
Veröffentlicht: (2026)
von: Li, Danrui, et al.
Veröffentlicht: (2026)
HY3D-Bench: Generation of 3D Assets
von: Hunyuan3D, Team, et al.
Veröffentlicht: (2026)
von: Hunyuan3D, Team, et al.
Veröffentlicht: (2026)
A-Bench: Are LMMs Masters at Evaluating AI-generated Images?
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
von: Zhang, Zicheng, et al.
Veröffentlicht: (2024)
TurtleBench: A Visual Programming Benchmark in Turtle Geometry
von: Rismanchian, Sina, et al.
Veröffentlicht: (2024)
von: Rismanchian, Sina, et al.
Veröffentlicht: (2024)
μ-Bench: A Vision-Language Benchmark for Microscopy Understanding
von: Lozano, Alejandro, et al.
Veröffentlicht: (2024)
von: Lozano, Alejandro, et al.
Veröffentlicht: (2024)
ViLCo-Bench: VIdeo Language COntinual learning Benchmark
von: Tang, Tianqi, et al.
Veröffentlicht: (2024)
von: Tang, Tianqi, et al.
Veröffentlicht: (2024)
LocateBench: Evaluating the Locating Ability of Vision Language Models
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
VideoGameBench: Can Vision-Language Models complete popular video games?
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
von: Zhang, Alex L., et al.
Veröffentlicht: (2025)
VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing
von: Luo, Zhiming, et al.
Veröffentlicht: (2026)
von: Luo, Zhiming, et al.
Veröffentlicht: (2026)
Omni IIE Bench: Benchmarking the Practical Capabilities of Image Editing Models
von: Yang, Yujia, et al.
Veröffentlicht: (2026)
von: Yang, Yujia, et al.
Veröffentlicht: (2026)
CT-Bench: A Benchmark for Multimodal Lesion Understanding in Computed Tomography
von: Zhu, Qingqing, et al.
Veröffentlicht: (2026)
von: Zhu, Qingqing, et al.
Veröffentlicht: (2026)
SeqBench: Benchmarking Sequential Narrative Generation in Text-to-Video Models
von: Tang, Zhengxu, et al.
Veröffentlicht: (2025)
von: Tang, Zhengxu, et al.
Veröffentlicht: (2025)
Waste-Bench: A Comprehensive Benchmark for Evaluating VLLMs in Cluttered Environments
von: Ali, Muhammad, et al.
Veröffentlicht: (2025)
von: Ali, Muhammad, et al.
Veröffentlicht: (2025)
EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
von: Yuan, Botai, et al.
Veröffentlicht: (2025)
StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
von: Lin, Junming, et al.
Veröffentlicht: (2024)
von: Lin, Junming, et al.
Veröffentlicht: (2024)
SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
von: Kong, Fanqi, et al.
Veröffentlicht: (2025)
von: Kong, Fanqi, et al.
Veröffentlicht: (2025)
Hydra-Bench: A Benchmark for Multi-Modal Leaf Wetness Sensing
von: Liu, Yimeng, et al.
Veröffentlicht: (2025)
von: Liu, Yimeng, et al.
Veröffentlicht: (2025)
AVTrustBench: Assessing and Enhancing Reliability and Robustness in Audio-Visual LLMs
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
von: Chowdhury, Sanjoy, et al.
Veröffentlicht: (2025)
WorldModelBench: Judging Video Generation Models As World Models
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
PoseBench: Benchmarking the Robustness of Pose Estimation Models under Corruptions
von: Ma, Sihan, et al.
Veröffentlicht: (2024)
von: Ma, Sihan, et al.
Veröffentlicht: (2024)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
GlazyBench: A Benchmark for Ceramic Glaze Property Prediction and Image Generation
von: Zhai, Ziyu, et al.
Veröffentlicht: (2026)
von: Zhai, Ziyu, et al.
Veröffentlicht: (2026)
VisBrowse-Bench: Benchmarking Visual-Native Search for Multimodal Browsing Agents
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
von: Zhang, Zhengbo, et al.
Veröffentlicht: (2026)
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
von: Jia, Zi-Yi, et al.
Veröffentlicht: (2026)
von: Jia, Zi-Yi, et al.
Veröffentlicht: (2026)
LithoBench: Benchmarking Large Multimodal Models for Remote-Sensing Lithology Interpretation
von: Wang, Jun, et al.
Veröffentlicht: (2026)
von: Wang, Jun, et al.
Veröffentlicht: (2026)
MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
von: Chen, Yifan, et al.
Veröffentlicht: (2026)
von: Chen, Yifan, et al.
Veröffentlicht: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language Models
von: Saxena, Rohit, et al.
Veröffentlicht: (2026)
von: Saxena, Rohit, et al.
Veröffentlicht: (2026)
DO-Bench: An Attributable Benchmark for Diagnosing Object Hallucination in Vision-Language Models
von: Wang, JiYang, et al.
Veröffentlicht: (2026)
von: Wang, JiYang, et al.
Veröffentlicht: (2026)
SurgBench: A Unified Large-Scale Benchmark for Surgical Video Analysis
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
von: Wei, Jianhui, et al.
Veröffentlicht: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
GEO-Bench-2: From Performance to Capability, Rethinking Evaluation in Geospatial AI
von: Simumba, Naomi, et al.
Veröffentlicht: (2025)
von: Simumba, Naomi, et al.
Veröffentlicht: (2025)
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
von: Gao, Tianyi, et al.
Veröffentlicht: (2025)
von: Gao, Tianyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
APML: Adaptive Probabilistic Matching Loss for Robust 3D Point Cloud Reconstruction
von: Sharifipour, Sasan, et al.
Veröffentlicht: (2025) -
PushupBench: Your VLM is not good at counting pushups
von: Li, Shengzhi, et al.
Veröffentlicht: (2026) -
Video-Bench: Human-Aligned Video Generation Benchmark
von: Han, Hui, et al.
Veröffentlicht: (2025) -
VEU-Bench: Towards Comprehensive Understanding of Video Editing
von: Li, Bozheng, et al.
Veröffentlicht: (2025) -
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
von: Ouyang, Kun, et al.
Veröffentlicht: (2024)