Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Martynov, Nikita, Mordasheva, Anastasia, Gorbetskiy, Dmitriy, Astafurov, Danil, Isaeva, Ulyana, Basyrova, Elina, Skachkov, Sergey, Berestova, Victoria, Ivanov, Nikolay, Zanina, Valeriia, Fenogenova, Alena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911295744245760
author Martynov, Nikita
Mordasheva, Anastasia
Gorbetskiy, Dmitriy
Astafurov, Danil
Isaeva, Ulyana
Basyrova, Elina
Skachkov, Sergey
Berestova, Victoria
Ivanov, Nikolay
Zanina, Valeriia
Fenogenova, Alena
author_facet Martynov, Nikita
Mordasheva, Anastasia
Gorbetskiy, Dmitriy
Astafurov, Danil
Isaeva, Ulyana
Basyrova, Elina
Skachkov, Sergey
Berestova, Victoria
Ivanov, Nikolay
Zanina, Valeriia
Fenogenova, Alena
contents We introduce POLLUX, a comprehensive open-source benchmark designed to evaluate the generative capabilities of large language models (LLMs) in Russian. Our main contribution is a novel evaluation methodology that enhances the interpretability of LLM assessment. For each task type, we define a set of detailed criteria and develop a scoring protocol where models evaluate responses and provide justifications for their ratings. This enables transparent, criteria-driven evaluation beyond traditional resource-consuming, side-by-side human comparisons. POLLUX includes a detailed, fine-grained taxonomy of 35 task types covering diverse generative domains such as code generation, creative writing, and practical assistant use cases, totaling 2,100 manually crafted and professionally authored prompts. Each task is categorized by difficulty (easy/medium/hard), with experts constructing the dataset entirely from scratch. We also release a family of LLM-as-a-Judge (7B and 32B) evaluators trained for nuanced assessment of generative outputs. This approach provides scalable, interpretable evaluation and annotation tools for model development, effectively replacing costly and less precise human judgments.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24616
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
Martynov, Nikita
Mordasheva, Anastasia
Gorbetskiy, Dmitriy
Astafurov, Danil
Isaeva, Ulyana
Basyrova, Elina
Skachkov, Sergey
Berestova, Victoria
Ivanov, Nikolay
Zanina, Valeriia
Fenogenova, Alena
Computation and Language
Artificial Intelligence
We introduce POLLUX, a comprehensive open-source benchmark designed to evaluate the generative capabilities of large language models (LLMs) in Russian. Our main contribution is a novel evaluation methodology that enhances the interpretability of LLM assessment. For each task type, we define a set of detailed criteria and develop a scoring protocol where models evaluate responses and provide justifications for their ratings. This enables transparent, criteria-driven evaluation beyond traditional resource-consuming, side-by-side human comparisons. POLLUX includes a detailed, fine-grained taxonomy of 35 task types covering diverse generative domains such as code generation, creative writing, and practical assistant use cases, totaling 2,100 manually crafted and professionally authored prompts. Each task is categorized by difficulty (easy/medium/hard), with experts constructing the dataset entirely from scratch. We also release a family of LLM-as-a-Judge (7B and 32B) evaluators trained for nuanced assessment of generative outputs. This approach provides scalable, interpretable evaluation and annotation tools for model development, effectively replacing costly and less precise human judgments.
title Eye of Judgement: Dissecting the Evaluation of Russian-speaking LLMs with POLLUX
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.24616