ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Joshi, Neha, Gogoi, Pamir, Mirza, Aasim, Jansari, Aayush, Yadavalli, Aditya, Pandey, Ayushi, Shukla, Arunima, Sudharsan, Deepthi, Bali, Kalika, Seshadri, Vivek
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914175819710464
author Joshi, Neha
Gogoi, Pamir
Mirza, Aasim
Jansari, Aayush
Yadavalli, Aditya
Pandey, Ayushi
Shukla, Arunima
Sudharsan, Deepthi
Bali, Kalika
Seshadri, Vivek
author_facet Joshi, Neha
Gogoi, Pamir
Mirza, Aasim
Jansari, Aayush
Yadavalli, Aditya
Pandey, Ayushi
Shukla, Arunima
Sudharsan, Deepthi
Bali, Kalika
Seshadri, Vivek
contents We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered languages. These recipes, rich in linguistic and cultural nuance, were collected using a mobile interface designed for contributors with low digital literacy. Endangered Language Recipes (ELR)-1000 -- captures not only culinary practices but also the socio-cultural context embedded in indigenous food traditions. We evaluate the performance of several state-of-the-art large language models (LLMs) on translating these recipes into English and find the following: despite the models' capabilities, they struggle with low-resource, culturally-specific language. However, we observe that providing targeted context -- including background information about the languages, translation examples, and guidelines for cultural preservation -- leads to significant improvements in translation quality. Our results underscore the need for benchmarks that cater to underrepresented languages and domains to advance equitable and culturally-aware language technologies. As part of this work, we release the ELR-1000 dataset to the NLP community, hoping it motivates the development of language technologies for endangered languages.
format Preprint
id arxiv_https___arxiv_org_abs_2512_01077
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
Joshi, Neha
Gogoi, Pamir
Mirza, Aasim
Jansari, Aayush
Yadavalli, Aditya
Pandey, Ayushi
Shukla, Arunima
Sudharsan, Deepthi
Bali, Kalika
Seshadri, Vivek
Computation and Language
Human-Computer Interaction
We present a culturally-grounded multimodal dataset of 1,060 traditional recipes crowdsourced from rural communities across remote regions of Eastern India, spanning 10 endangered languages. These recipes, rich in linguistic and cultural nuance, were collected using a mobile interface designed for contributors with low digital literacy. Endangered Language Recipes (ELR)-1000 -- captures not only culinary practices but also the socio-cultural context embedded in indigenous food traditions. We evaluate the performance of several state-of-the-art large language models (LLMs) on translating these recipes into English and find the following: despite the models' capabilities, they struggle with low-resource, culturally-specific language. However, we observe that providing targeted context -- including background information about the languages, translation examples, and guidelines for cultural preservation -- leads to significant improvements in translation quality. Our results underscore the need for benchmarks that cater to underrepresented languages and domains to advance equitable and culturally-aware language technologies. As part of this work, we release the ELR-1000 dataset to the NLP community, hoping it motivates the development of language technologies for endangered languages.
title ELR-1000: A Community-Generated Dataset for Endangered Indic Indigenous Languages
topic Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2512.01077