Saved in:
Bibliographic Details
Main Authors: Xiong, Shengwu., Zou, Tianyu., Wang, Cong., Li, Xuelong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2506.21572
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909900099026944
author Xiong, Shengwu.
Zou, Tianyu.
Wang, Cong.
Li, Xuelong
author_facet Xiong, Shengwu.
Zou, Tianyu.
Wang, Cong.
Li, Xuelong
contents Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristically-grouped tasks have vague cognitive targets, overlapping abilities, redundant indicators, and weak diagnostic power. We therefore propose a structural-equation-modeling-aligned framework that quantifies internal validity, dimensional separability, and component contributions, and introduce a Piaget-inspired capability hierarchy that stratifies MLLM abilities into Perception, Memory, and Reasoning. Reorganizing existing tasks under this theory, we build the GOLD benchmark, whose experiments show superior interpretability, lower indicator redundancy, and clearer cognitive consistency than prior benchmarks.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21572
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
Xiong, Shengwu.
Zou, Tianyu.
Wang, Cong.
Li, Xuelong
Computation and Language
Evaluating multimodal large language models (MLLMs) is fundamentally challenged by the absence of structured, interpretable, and theoretically grounded benchmarks; current heuristically-grouped tasks have vague cognitive targets, overlapping abilities, redundant indicators, and weak diagnostic power. We therefore propose a structural-equation-modeling-aligned framework that quantifies internal validity, dimensional separability, and component contributions, and introduce a Piaget-inspired capability hierarchy that stratifies MLLM abilities into Perception, Memory, and Reasoning. Reorganizing existing tasks under this theory, we build the GOLD benchmark, whose experiments show superior interpretability, lower indicator redundancy, and clearer cognitive consistency than prior benchmarks.
title Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
topic Computation and Language
url https://arxiv.org/abs/2506.21572