MIEB: Massive Image Embedding Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Chenghao, Chung, Isaac, Kerboua, Imene, Stirling, Jamie, Zhang, Xin, Kardos, Márton, Solomatin, Roman, Moubayed, Noura Al, Enevoldsen, Kenneth, Muennighoff, Niklas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910911151734784
author Xiao, Chenghao
Chung, Isaac
Kerboua, Imene
Stirling, Jamie
Zhang, Xin
Kardos, Márton
Solomatin, Roman
Moubayed, Noura Al
Enevoldsen, Kenneth
Muennighoff, Niklas
author_facet Xiao, Chenghao
Chung, Isaac
Kerboua, Imene
Stirling, Jamie
Zhang, Xin
Kardos, Márton
Solomatin, Roman
Moubayed, Noura Al
Enevoldsen, Kenneth
Muennighoff, Niklas
contents Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10471
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MIEB: Massive Image Embedding Benchmark
Xiao, Chenghao
Chung, Isaac
Kerboua, Imene
Stirling, Jamie
Zhang, Xin
Kardos, Márton
Solomatin, Roman
Moubayed, Noura Al
Enevoldsen, Kenneth
Muennighoff, Niklas
Computer Vision and Pattern Recognition
Computation and Language
Image representations are often evaluated through disjointed, task-specific protocols, leading to a fragmented understanding of model capabilities. For instance, it is unclear whether an image embedding model adept at clustering images is equally good at retrieving relevant images given a piece of text. We introduce the Massive Image Embedding Benchmark (MIEB) to evaluate the performance of image and image-text embedding models across the broadest spectrum to date. MIEB spans 38 languages across 130 individual tasks, which we group into 8 high-level categories. We benchmark 50 models across our benchmark, finding that no single method dominates across all task categories. We reveal hidden capabilities in advanced vision models such as their accurate visual representation of texts, and their yet limited capabilities in interleaved encodings and matching images and texts in the presence of confounders. We also show that the performance of vision encoders on MIEB correlates highly with their performance when used in multimodal large language models. Our code, dataset, and leaderboard are publicly available at https://github.com/embeddings-benchmark/mteb.
title MIEB: Massive Image Embedding Benchmark
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2504.10471