Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Suominen, Osma, Inkinen, Juho, Lehtinen, Mona
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916911878504448
author Suominen, Osma
Inkinen, Juho
Lehtinen, Mona
author_facet Suominen, Osma
Inkinen, Juho
Lehtinen, Mona
contents This paper presents the Annif system in the LLMs4Subjects shared task (Subtask 2) at GermEval-2025. The task required creating subject predictions for bibliographic records using large language models, with a special focus on computational efficiency. Our system, based on the Annif automated subject indexing toolkit, refines our previous system from the first LLMs4Subjects shared task, which produced excellent results. We further improved the system by using many small and efficient language models for translation and synthetic data generation and by using LLMs for ranking candidate subjects. Our system ranked 1st in the overall quantitative evaluation of and 1st in the qualitative evaluation of Subtask 2.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15877
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs
Suominen, Osma
Inkinen, Juho
Lehtinen, Mona
Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
I.2.7
This paper presents the Annif system in the LLMs4Subjects shared task (Subtask 2) at GermEval-2025. The task required creating subject predictions for bibliographic records using large language models, with a special focus on computational efficiency. Our system, based on the Annif automated subject indexing toolkit, refines our previous system from the first LLMs4Subjects shared task, which produced excellent results. We further improved the system by using many small and efficient language models for translation and synthetic data generation and by using LLMs for ranking candidate subjects. Our system ranked 1st in the overall quantitative evaluation of and 1st in the qualitative evaluation of Subtask 2.
title Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs
topic Computation and Language
Artificial Intelligence
Information Retrieval
Machine Learning
I.2.7
url https://arxiv.org/abs/2508.15877