Benchmarking Vision Language Models on German Factual Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peinl, René, Tischler, Vincent
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916811997446144
author Peinl, René
Tischler, Vincent
author_facet Peinl, René
Tischler, Vincent
contents Similar to LLMs, the development of vision language models is mainly driven by English datasets and models trained in English and Chinese language, whereas support for other languages, even those considered high-resource languages such as German, remains significantly weaker. In this work we present an analysis of open-weight VLMs on factual knowledge in the German and English language. We disentangle the image-related aspects from the textual ones by analyzing accu-racy with jury-as-a-judge in both prompt languages and images from German and international contexts. We found that for celebrities and sights, VLMs struggle because they are lacking visual cognition of German image contents. For animals and plants, the tested models can often correctly identify the image contents ac-cording to the scientific name or English common name but fail in German lan-guage. Cars and supermarket products were identified equally well in English and German images across both prompt languages.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11108
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Benchmarking Vision Language Models on German Factual Data
Peinl, René
Tischler, Vincent
Computation and Language
68T45 (Primary), 68T07 (Secondary), 68T10 (Secondary)
I.4.0
Similar to LLMs, the development of vision language models is mainly driven by English datasets and models trained in English and Chinese language, whereas support for other languages, even those considered high-resource languages such as German, remains significantly weaker. In this work we present an analysis of open-weight VLMs on factual knowledge in the German and English language. We disentangle the image-related aspects from the textual ones by analyzing accu-racy with jury-as-a-judge in both prompt languages and images from German and international contexts. We found that for celebrities and sights, VLMs struggle because they are lacking visual cognition of German image contents. For animals and plants, the tested models can often correctly identify the image contents ac-cording to the scientific name or English common name but fail in German lan-guage. Cars and supermarket products were identified equally well in English and German images across both prompt languages.
title Benchmarking Vision Language Models on German Factual Data
topic Computation and Language
68T45 (Primary), 68T07 (Secondary), 68T10 (Secondary)
I.4.0
url https://arxiv.org/abs/2504.11108