GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhu, Xiaorong, Jia, Ziheng, Wang, Jiarui, Zhao, Xiangyu, Duan, Haodong, Min, Xiongkuo, Wang, Jia, Zhang, Zicheng, Zhai, Guangtao
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913975369728000
author Zhu, Xiaorong
Jia, Ziheng
Wang, Jiarui
Zhao, Xiangyu
Duan, Haodong
Min, Xiongkuo
Wang, Jia
Zhang, Zicheng
Zhai, Guangtao
author_facet Zhu, Xiaorong
Jia, Ziheng
Wang, Jiarui
Zhao, Xiangyu
Duan, Haodong
Min, Xiongkuo
Wang, Jia
Zhang, Zicheng
Zhai, Guangtao
contents The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained physical principles especially in geometric optics, remains underexplored. To address this gap, we introduce GOBench, the first benchmark to systematically evaluate MLLMs' ability across two tasks: 1) Generating Optically Authentic Imagery and 2) Understanding Underlying Optical Phenomena. We curates high-quality prompts of geometric optical scenarios and use MLLMs to construct GOBench-Gen-1k dataset.We then organize subjective experiments to assess the generated imagery based on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, revealing MLLMs' generation flaws that violate optical principles. For the understanding task, we apply crafted evaluation instructions to test optical understanding ability of eleven prominent MLLMs. The experimental results demonstrate that current models face significant challenges in both optical generation and understanding. The top-performing generative model, GPT-4o-Image, cannot perfectly complete all generation tasks, and the best-performing MLLM model, Gemini-2.5Pro, attains a mere 37.35\% accuracy in optical understanding. Database and codes are publicly available at https://github.com/aiben-ch/GOBench.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00991
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
Zhu, Xiaorong
Jia, Ziheng
Wang, Jiarui
Zhao, Xiangyu
Duan, Haodong
Min, Xiongkuo
Wang, Jia
Zhang, Zicheng
Zhai, Guangtao
Computer Vision and Pattern Recognition
The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained physical principles especially in geometric optics, remains underexplored. To address this gap, we introduce GOBench, the first benchmark to systematically evaluate MLLMs' ability across two tasks: 1) Generating Optically Authentic Imagery and 2) Understanding Underlying Optical Phenomena. We curates high-quality prompts of geometric optical scenarios and use MLLMs to construct GOBench-Gen-1k dataset.We then organize subjective experiments to assess the generated imagery based on Optical Authenticity, Aesthetic Quality, and Instruction Fidelity, revealing MLLMs' generation flaws that violate optical principles. For the understanding task, we apply crafted evaluation instructions to test optical understanding ability of eleven prominent MLLMs. The experimental results demonstrate that current models face significant challenges in both optical generation and understanding. The top-performing generative model, GPT-4o-Image, cannot perfectly complete all generation tasks, and the best-performing MLLM model, Gemini-2.5Pro, attains a mere 37.35\% accuracy in optical understanding. Database and codes are publicly available at https://github.com/aiben-ch/GOBench.
title GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.00991