BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915470120058880 |
|---|---|
| author | Santos, João Guilherme Alves Bonás, Giovana Kerche Almeida, Thales Sales |
| author_facet | Santos, João Guilherme Alves Bonás, Giovana Kerche Almeida, Thales Sales |
| contents | With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We present an updated version of the BLUEX dataset, now including 2024-2025 exams and automatically generated image captions using state-of-the-art models, enhancing its relevance for data contamination studies in LLM pretraining. Captioning strategies increase accessibility to text-only models by more than 40%, producing 1,422 usable questions, more than doubling the number in the original BLUEX. We evaluated commercial and open-source LLMs and their ability to leverage visual context through captions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_21294 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning Santos, João Guilherme Alves Bonás, Giovana Kerche Almeida, Thales Sales Computation and Language Artificial Intelligence With the growing capabilities of Large Language Models (LLMs), there is an increasing need for robust evaluation methods, especially in multilingual and non-English contexts. We present an updated version of the BLUEX dataset, now including 2024-2025 exams and automatically generated image captions using state-of-the-art models, enhancing its relevance for data contamination studies in LLM pretraining. Captioning strategies increase accessibility to text-only models by more than 40%, producing 1,422 usable questions, more than doubling the number in the original BLUEX. We evaluated commercial and open-source LLMs and their ability to leverage visual context through captions. |
| title | BLUEX Revisited: Enhancing Benchmark Coverage with Automatic Captioning |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2508.21294 |