Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Acharya, Arkadeep, Murthy, Rudra, Kumar, Vishwajeet, Sen, Jaydeep
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914915534503936
author Acharya, Arkadeep
Murthy, Rudra
Kumar, Vishwajeet
Sen, Jaydeep
author_facet Acharya, Arkadeep
Murthy, Rudra
Kumar, Vishwajeet
Sen, Jaydeep
contents Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, there is a lack of comprehensive benchmark for evaluating retrieval models in Hindi. To address this gap, we introduce the Hindi version of the BEIR benchmark, which includes a subset of English BEIR datasets translated to Hindi, existing Hindi retrieval datasets, and synthetically created datasets for retrieval. The benchmark is comprised of $15$ datasets spanning across $8$ distinct tasks. We evaluate state-of-the-art multilingual retrieval models on this benchmark to identify task and domain-specific challenges and their impact on retrieval performance. By releasing this benchmark and a set of relevant baselines, we enable researchers to understand the limitations and capabilities of current Hindi retrieval models, promoting advancements in this critical area. The datasets from Hindi-BEIR are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2408_09437
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
Acharya, Arkadeep
Murthy, Rudra
Kumar, Vishwajeet
Sen, Jaydeep
Information Retrieval
Computation and Language
Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, there is a lack of comprehensive benchmark for evaluating retrieval models in Hindi. To address this gap, we introduce the Hindi version of the BEIR benchmark, which includes a subset of English BEIR datasets translated to Hindi, existing Hindi retrieval datasets, and synthetically created datasets for retrieval. The benchmark is comprised of $15$ datasets spanning across $8$ distinct tasks. We evaluate state-of-the-art multilingual retrieval models on this benchmark to identify task and domain-specific challenges and their impact on retrieval performance. By releasing this benchmark and a set of relevant baselines, we enable researchers to understand the limitations and capabilities of current Hindi retrieval models, promoting advancements in this critical area. The datasets from Hindi-BEIR are publicly available.
title Hindi-BEIR : A Large Scale Retrieval Benchmark in Hindi
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2408.09437