UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Schimanski, Tobias, Lewandowski, Stefanie, Woerle, Christian, Reichenau, Nicola, Huryn, Yauheni, Leippold, Markus
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917430841835520
author Schimanski, Tobias
Lewandowski, Stefanie
Woerle, Christian
Reichenau, Nicola
Huryn, Yauheni
Leippold, Markus
author_facet Schimanski, Tobias
Lewandowski, Stefanie
Woerle, Christian
Reichenau, Nicola
Huryn, Yauheni
Leippold, Markus
contents Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevance is dominated by aspects of similarity in texts, leaving unobserved whether the text is truly useful for addressing the query. For instance, when answering whether Paris is larger than Berlin, texts about Paris being in France are relevant (lexical/semantic similarity), but not useful. In this paper, we introduce UsefulBench, a domain-specific dataset curated by three professional analysts labeling whether a text is connected to a query (relevance) or holds practical value in responding to it (usefulness). We show that classic similarity-based information retrieval aligns more strongly with relevance. While LLM-based systems can counteract this bias, we find that domain-specific problems require a high degree of expertise, which current LLMs do not fully incorporate. We explore approaches to (partially) overcome this challenge. However, UsefulBench presents a dataset challenge for targeted information retrieval systems.
format Preprint
id arxiv_https___arxiv_org_abs_2604_15827
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval
Schimanski, Tobias
Lewandowski, Stefanie
Woerle, Christian
Reichenau, Nicola
Huryn, Yauheni
Leippold, Markus
Information Retrieval
Computation and Language
Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevance is dominated by aspects of similarity in texts, leaving unobserved whether the text is truly useful for addressing the query. For instance, when answering whether Paris is larger than Berlin, texts about Paris being in France are relevant (lexical/semantic similarity), but not useful. In this paper, we introduce UsefulBench, a domain-specific dataset curated by three professional analysts labeling whether a text is connected to a query (relevance) or holds practical value in responding to it (usefulness). We show that classic similarity-based information retrieval aligns more strongly with relevance. While LLM-based systems can counteract this bias, we find that domain-specific problems require a high degree of expertise, which current LLMs do not fully incorporate. We explore approaches to (partially) overcome this challenge. However, UsefulBench presents a dataset challenge for targeted information retrieval systems.
title UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval
topic Information Retrieval
Computation and Language
url https://arxiv.org/abs/2604.15827