Overview of the TREC 2023 deep learning track

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Craswell, Nick, Mitra, Bhaskar, Yilmaz, Emine, Rahmani, Hossein A., Campos, Daniel, Lin, Jimmy, Voorhees, Ellen M., Soboroff, Ian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916840112914432
author Craswell, Nick
Mitra, Bhaskar
Yilmaz, Emine
Rahmani, Hossein A.
Campos, Daniel
Lin, Jimmy
Voorhees, Ellen M.
Soboroff, Ian
author_facet Craswell, Nick
Mitra, Bhaskar
Yilmaz, Emine
Rahmani, Hossein A.
Campos, Daniel
Lin, Jimmy
Voorhees, Ellen M.
Soboroff, Ian
contents This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels available for both passage and document ranking tasks. We mostly repeated last year's design, to get another matching test set, based on the larger, cleaner, less-biased v2 passage and document set, with passage ranking as primary and document ranking as a secondary task (using labels inferred from passage). As we did last year, we sample from MS MARCO queries that were completely held out, unused in corpus construction, unlike the test queries in the first three years. This approach yields a more difficult test with more headroom for improvement. Alongside the usual MS MARCO (human) queries from MS MARCO, this year we generated synthetic queries using a fine-tuned T5 model and using a GPT-4 prompt. The new headline result this year is that runs using Large Language Model (LLM) prompting in some way outperformed runs that use the "nnlm" approach, which was the best approach in the previous four years. Since this is the last year of the track, future iterations of prompt-based ranking can happen in other tracks. Human relevance assessments were applied to all query types, not just human MS MARCO queries. Evaluation using synthetic queries gave similar results to human queries, with system ordering agreement of $τ=0.8487$. However, human effort was needed to select a subset of the synthetic queries that were usable. We did not see clear evidence of bias, where runs using GPT-4 were favored when evaluated using synthetic GPT-4 queries, or where runs using T5 were favored when evaluated on synthetic T5 queries.
format Preprint
id arxiv_https___arxiv_org_abs_2507_08890
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Overview of the TREC 2023 deep learning track
Craswell, Nick
Mitra, Bhaskar
Yilmaz, Emine
Rahmani, Hossein A.
Campos, Daniel
Lin, Jimmy
Voorhees, Ellen M.
Soboroff, Ian
Information Retrieval
Artificial Intelligence
Computation and Language
This is the fifth year of the TREC Deep Learning track. As in previous years, we leverage the MS MARCO datasets that made hundreds of thousands of human-annotated training labels available for both passage and document ranking tasks. We mostly repeated last year's design, to get another matching test set, based on the larger, cleaner, less-biased v2 passage and document set, with passage ranking as primary and document ranking as a secondary task (using labels inferred from passage). As we did last year, we sample from MS MARCO queries that were completely held out, unused in corpus construction, unlike the test queries in the first three years. This approach yields a more difficult test with more headroom for improvement. Alongside the usual MS MARCO (human) queries from MS MARCO, this year we generated synthetic queries using a fine-tuned T5 model and using a GPT-4 prompt. The new headline result this year is that runs using Large Language Model (LLM) prompting in some way outperformed runs that use the "nnlm" approach, which was the best approach in the previous four years. Since this is the last year of the track, future iterations of prompt-based ranking can happen in other tracks. Human relevance assessments were applied to all query types, not just human MS MARCO queries. Evaluation using synthetic queries gave similar results to human queries, with system ordering agreement of $τ=0.8487$. However, human effort was needed to select a subset of the synthetic queries that were usable. We did not see clear evidence of bias, where runs using GPT-4 were favored when evaluated using synthetic GPT-4 queries, or where runs using T5 were favored when evaluated on synthetic T5 queries.
title Overview of the TREC 2023 deep learning track
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2507.08890