Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Chaturvedi, Sarthak, Acharya, Anurag, Meyur, Rounak, Hayashi, Koby, Munikoti, Sai, Horawalavithana, Sameera
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911025777868800
author Chaturvedi, Sarthak
Acharya, Anurag
Meyur, Rounak
Hayashi, Koby
Munikoti, Sai
Horawalavithana, Sameera
author_facet Chaturvedi, Sarthak
Acharya, Anurag
Meyur, Rounak
Hayashi, Koby
Munikoti, Sai
Horawalavithana, Sameera
contents Evaluation benchmark characteristics may distort the true benefits of domain adaptation in retrieval models. This creates misleading assessments that influence deployment decisions in specialized domains. We show that two benchmarks with drastically different features such as topic diversity, boundary overlap, and semantic complexity can influence the perceived benefits of fine-tuning. Using environmental regulatory document retrieval as a case study, we fine-tune ColBERTv2 model on Environmental Impact Statements (EIS) from federal agencies. We evaluate these models across two benchmarks with different semantic structures. Our findings reveal that identical domain adaptation approaches show very different perceived benefits depending on evaluation methodology. On one benchmark, with clearly separated topic boundaries, domain adaptation shows small improvements (maximum 0.61% NDCG gain). However, on the other benchmark with overlapping semantic structures, the same models demonstrate large improvements (up to 2.22% NDCG gain), a 3.6-fold difference in the performance benefit. We compare these benchmarks through topic diversity metrics, finding that the higher-performing benchmark shows 11% higher average cosine distances between contexts and 23% lower silhouette scores, directly contributing to the observed performance difference. These results demonstrate that benchmark selection strongly determines assessments of retrieval system effectiveness in specialized domains. Evaluation frameworks with well-separated topics regularly underestimate domain adaptation benefits, while those with overlapping semantic boundaries reveal improvements that better reflect real-world regulatory document complexity. Our findings have important implications for developing and deploying AI systems for interdisciplinary domains that integrate multiple topics.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21581
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
Chaturvedi, Sarthak
Acharya, Anurag
Meyur, Rounak
Hayashi, Koby
Munikoti, Sai
Horawalavithana, Sameera
Information Retrieval
Artificial Intelligence
Machine Learning
Evaluation benchmark characteristics may distort the true benefits of domain adaptation in retrieval models. This creates misleading assessments that influence deployment decisions in specialized domains. We show that two benchmarks with drastically different features such as topic diversity, boundary overlap, and semantic complexity can influence the perceived benefits of fine-tuning. Using environmental regulatory document retrieval as a case study, we fine-tune ColBERTv2 model on Environmental Impact Statements (EIS) from federal agencies. We evaluate these models across two benchmarks with different semantic structures. Our findings reveal that identical domain adaptation approaches show very different perceived benefits depending on evaluation methodology. On one benchmark, with clearly separated topic boundaries, domain adaptation shows small improvements (maximum 0.61% NDCG gain). However, on the other benchmark with overlapping semantic structures, the same models demonstrate large improvements (up to 2.22% NDCG gain), a 3.6-fold difference in the performance benefit. We compare these benchmarks through topic diversity metrics, finding that the higher-performing benchmark shows 11% higher average cosine distances between contexts and 23% lower silhouette scores, directly contributing to the observed performance difference. These results demonstrate that benchmark selection strongly determines assessments of retrieval system effectiveness in specialized domains. Evaluation frameworks with well-separated topics regularly underestimate domain adaptation benefits, while those with overlapping semantic boundaries reveal improvements that better reflect real-world regulatory document complexity. Our findings have important implications for developing and deploying AI systems for interdisciplinary domains that integrate multiple topics.
title Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains
topic Information Retrieval
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.21581