MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916437692514304 |
|---|---|
| author | Sun, Weiwei Shi, Zhengliang Wu, Jiulong Yan, Lingyong Ma, Xinyu Liu, Yiding Cao, Min Yin, Dawei Ren, Zhaochun |
| author_facet | Sun, Weiwei Shi, Zhengliang Wu, Jiulong Yan, Lingyong Ma, Xinyu Liu, Yiding Cao, Min Yin, Dawei Ren, Zhaochun |
| contents | Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_10127 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | MAIR: A Massive Benchmark for Evaluating Instructed Retrieval Sun, Weiwei Shi, Zhengliang Wu, Jiulong Yan, Lingyong Ma, Xinyu Liu, Yiding Cao, Min Yin, Dawei Ren, Zhaochun Information Retrieval Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair. |
| title | MAIR: A Massive Benchmark for Evaluating Instructed Retrieval |
| topic | Information Retrieval |
| url | https://arxiv.org/abs/2410.10127 |