Reproducing NevIR: Negation in Neural Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Elsen, Coen van den, Barkhof, Francien, Nijdam, Thijmen, Lupart, Simon, Aliannejadi, Mohammad
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908348651143168
author Elsen, Coen van den
Barkhof, Francien
Nijdam, Thijmen
Lupart, Simon
Aliannejadi, Mohammad
author_facet Elsen, Coen van den
Barkhof, Francien
Nijdam, Thijmen
Lupart, Simon
Aliannejadi, Mohammad
contents Negation is a fundamental aspect of human communication, yet it remains a challenge for Language Models (LMs) in Information Retrieval (IR). Despite the heavy reliance of modern neural IR systems on LMs, little attention has been given to their handling of negation. In this study, we reproduce and extend the findings of NevIR, a benchmark study that revealed most IR models perform at or below the level of random ranking when dealing with negation. We replicate NevIR's original experiments and evaluate newly developed state-of-the-art IR models. Our findings show that a recently emerging category-listwise Large Language Model (LLM) re-rankers-outperforms other models but still underperforms human performance. Additionally, we leverage ExcluIR, a benchmark dataset designed for exclusionary queries with extensive negation, to assess the generalisability of negation understanding. Our findings suggest that fine-tuning on one dataset does not reliably improve performance on the other, indicating notable differences in their data distributions. Furthermore, we observe that only cross-encoders and listwise LLM re-rankers achieve reasonable performance across both negation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Reproducing NevIR: Negation in Neural Information Retrieval
Elsen, Coen van den
Barkhof, Francien
Nijdam, Thijmen
Lupart, Simon
Aliannejadi, Mohammad
Information Retrieval
Negation is a fundamental aspect of human communication, yet it remains a challenge for Language Models (LMs) in Information Retrieval (IR). Despite the heavy reliance of modern neural IR systems on LMs, little attention has been given to their handling of negation. In this study, we reproduce and extend the findings of NevIR, a benchmark study that revealed most IR models perform at or below the level of random ranking when dealing with negation. We replicate NevIR's original experiments and evaluate newly developed state-of-the-art IR models. Our findings show that a recently emerging category-listwise Large Language Model (LLM) re-rankers-outperforms other models but still underperforms human performance. Additionally, we leverage ExcluIR, a benchmark dataset designed for exclusionary queries with extensive negation, to assess the generalisability of negation understanding. Our findings suggest that fine-tuning on one dataset does not reliably improve performance on the other, indicating notable differences in their data distributions. Furthermore, we observe that only cross-encoders and listwise LLM re-rankers achieve reasonable performance across both negation tasks.
title Reproducing NevIR: Negation in Neural Information Retrieval
topic Information Retrieval
url https://arxiv.org/abs/2502.13506