Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Litschko, Robert, Kraus, Oliver, Blaschke, Verena, Plank, Barbara
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916552017707008
author Litschko, Robert
Kraus, Oliver
Blaschke, Verena
Plank, Barbara
author_facet Litschko, Robert
Kraus, Oliver
Blaschke, Verena
Plank, Barbara
contents A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limited attention. Dialect retrieval poses unique challenges due to the limited availability of resources to train retrieval models and the high variability in non-standardized languages. We study these challenges on the example of German dialects and introduce the first German dialect retrieval dataset, dubbed WikiDIR, which consists of seven German dialects extracted from Wikipedia. Using WikiDIR, we demonstrate the weakness of lexical methods in dealing with high lexical variation in dialects. We further show that commonly used zero-shot cross-lingual transfer approach with multilingual encoders do not transfer well to extremely low-resource setups, motivating the need for resource-lean and dialect-specific retrieval models. We finally demonstrate that (document) translation is an effective way to reduce the dialect gap in CDIR.
format Preprint
id arxiv_https___arxiv_org_abs_2412_12806
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages
Litschko, Robert
Kraus, Oliver
Blaschke, Verena
Plank, Barbara
Computation and Language
Information Retrieval
A large amount of local and culture-specific knowledge (e.g., people, traditions, food) can only be found in documents written in dialects. While there has been extensive research conducted on cross-lingual information retrieval (CLIR), the field of cross-dialect retrieval (CDIR) has received limited attention. Dialect retrieval poses unique challenges due to the limited availability of resources to train retrieval models and the high variability in non-standardized languages. We study these challenges on the example of German dialects and introduce the first German dialect retrieval dataset, dubbed WikiDIR, which consists of seven German dialects extracted from Wikipedia. Using WikiDIR, we demonstrate the weakness of lexical methods in dealing with high lexical variation in dialects. We further show that commonly used zero-shot cross-lingual transfer approach with multilingual encoders do not transfer well to extremely low-resource setups, motivating the need for resource-lean and dialect-specific retrieval models. We finally demonstrate that (document) translation is an effective way to reduce the dialect gap in CDIR.
title Cross-Dialect Information Retrieval: Information Access in Low-Resource and High-Variance Languages
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2412.12806