Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909900099026944 |
|---|---|
| author | Xiong, Shengwu. Zou, Tianyu. Wang, Cong. Li, Xuelong |
| author_facet | Xiong, Shengwu. Zou, Tianyu. Wang, Cong. Li, Xuelong |
| contents | Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristically-grouped tasks have vague cognitive targets, overlapping abilities, redundant indicators, and weak diagnostic power. We therefore propose a structural-equation-modeling-aligned framework that quantifies internal validity, dimensional separability, and component contributions, and introduce a Piaget-inspired capability hierarchy that stratifies MLLM abilities into Perception, Memory, and Reasoning. Reorganizing existing tasks under this theory, we build the GOLD benchmark, whose experiments show superior interpretability, lower indicator redundancy, and clearer cognitive consistency than prior benchmarks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21572 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling Xiong, Shengwu. Zou, Tianyu. Wang, Cong. Li, Xuelong Computation and Language Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristically-grouped tasks have vague cognitive targets, overlapping abilities, redundant indicators, and weak diagnostic power. We therefore propose a structural-equation-modeling-aligned framework that quantifies internal validity, dimensional separability, and component contributions, and introduce a Piaget-inspired capability hierarchy that stratifies MLLM abilities into Perception, Memory, and Reasoning. Reorganizing existing tasks under this theory, we build the GOLD benchmark, whose experiments show superior interpretability, lower indicator redundancy, and clearer cognitive consistency than prior benchmarks. |
| title | Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2506.21572 |