MAIR: A Massive Benchmark for Evaluating Instructed Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Weiwei, Shi, Zhengliang, Wu, Jiulong, Yan, Lingyong, Ma, Xinyu, Liu, Yiding, Cao, Min, Yin, Dawei, Ren, Zhaochun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916437692514304
author Sun, Weiwei
Shi, Zhengliang
Wu, Jiulong
Yan, Lingyong
Ma, Xinyu
Liu, Yiding
Cao, Min
Yin, Dawei
Ren, Zhaochun
author_facet Sun, Weiwei
Shi, Zhengliang
Wu, Jiulong
Yan, Lingyong
Ma, Xinyu
Liu, Yiding
Cao, Min
Yin, Dawei
Ren, Zhaochun
contents Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair.
format Preprint
id arxiv_https___arxiv_org_abs_2410_10127
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
Sun, Weiwei
Shi, Zhengliang
Wu, Jiulong
Yan, Lingyong
Ma, Xinyu
Liu, Yiding
Cao, Min
Yin, Dawei
Ren, Zhaochun
Information Retrieval
Recent information retrieval (IR) models are pre-trained and instruction-tuned on massive datasets and tasks, enabling them to perform well on a wide range of tasks and potentially generalize to unseen tasks with instructions. However, existing IR benchmarks focus on a limited scope of tasks, making them insufficient for evaluating the latest IR models. In this paper, we propose MAIR (Massive Instructed Retrieval Benchmark), a heterogeneous IR benchmark that includes 126 distinct IR tasks across 6 domains, collected from existing datasets. We benchmark state-of-the-art instruction-tuned text embedding models and re-ranking models. Our experiments reveal that instruction-tuned models generally achieve superior performance compared to non-instruction-tuned models on MAIR. Additionally, our results suggest that current instruction-tuned text embedding models and re-ranking models still lack effectiveness in specific long-tail tasks. MAIR is publicly available at https://github.com/sunnweiwei/Mair.
title MAIR: A Massive Benchmark for Evaluating Instructed Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2410.10127