Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ruiz, Tomas, Agustoslu, Tanalp, Schwemmer, Carsten
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915877419483136
author Ruiz, Tomas
Agustoslu, Tanalp
Schwemmer, Carsten
author_facet Ruiz, Tomas
Agustoslu, Tanalp
Schwemmer, Carsten
contents Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) development. We address this gap by introducing an evaluation protocol for multimodal large language model (MLLM) benchmarking that explicitly accounts for two conditions: (1) human label agreement and (2) disagreement. We apply this protocol to two state-of-the-art MLLM families (Gemma 3, Qwen 2.5 VL) using non-aggregated human annotations from a social media content classification dataset. Across tasks, we find that larger models tend to perform best on high-agreement subsets, yet often underperform medium-sized models when human disagreement is high, indicating that parameter count alone does not determine sensitivity to ambiguity and subjectivity. These results show that benchmarks based solely on consensus labels can overstate model capabilities in such domains and that incorporating human label variation yields more realistic and robust assessments of MLLMs in content moderation pipelines.
format Preprint
id arxiv_https___arxiv_org_abs_2603_19744
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
Ruiz, Tomas
Agustoslu, Tanalp
Schwemmer, Carsten
Computation and Language
Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) development. We address this gap by introducing an evaluation protocol for multimodal large language model (MLLM) benchmarking that explicitly accounts for two conditions: (1) human label agreement and (2) disagreement. We apply this protocol to two state-of-the-art MLLM families (Gemma 3, Qwen 2.5 VL) using non-aggregated human annotations from a social media content classification dataset. Across tasks, we find that larger models tend to perform best on high-agreement subsets, yet often underperform medium-sized models when human disagreement is high, indicating that parameter count alone does not determine sensitivity to ambiguity and subjectivity. These results show that benchmarks based solely on consensus labels can overstate model capabilities in such domains and that incorporating human label variation yields more realistic and robust assessments of MLLMs in content moderation pipelines.
title Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
topic Computation and Language
url https://arxiv.org/abs/2603.19744