IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information Retrieval
Fuente:
arXiv
Guardado en:
| Autores principales: | Song, Tingyu, Gan, Guo, Shang, Mingsheng, Zhao, Yilun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
por: Zhao, Yilun, et al.
Publicado: (2026)
por: Zhao, Yilun, et al.
Publicado: (2026)
FollowTable: A Benchmark for Instruction-Following Table Retrieval
por: Jin, Rihui, et al.
Publicado: (2026)
por: Jin, Rihui, et al.
Publicado: (2026)
LimRank: Less is More for Reasoning-Intensive Information Reranking
por: Song, Tingyu, et al.
Publicado: (2025)
por: Song, Tingyu, et al.
Publicado: (2025)
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
por: Song, Tingyu, et al.
Publicado: (2026)
por: Song, Tingyu, et al.
Publicado: (2026)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
por: Wang, Shuting, et al.
Publicado: (2024)
por: Wang, Shuting, et al.
Publicado: (2024)
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
por: Weller, Orion, et al.
Publicado: (2024)
por: Weller, Orion, et al.
Publicado: (2024)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
por: Hu, Tiansheng, et al.
Publicado: (2026)
por: Hu, Tiansheng, et al.
Publicado: (2026)
mFollowIR: a Multilingual Benchmark for Instruction Following in Retrieval
por: Weller, Orion, et al.
Publicado: (2025)
por: Weller, Orion, et al.
Publicado: (2025)
Towards Better Instruction Following Retrieval Models
por: Zhuang, Yuchen, et al.
Publicado: (2025)
por: Zhuang, Yuchen, et al.
Publicado: (2025)
CoIR: A Comprehensive Benchmark for Code Information Retrieval Models
por: Li, Xiangyang, et al.
Publicado: (2024)
por: Li, Xiangyang, et al.
Publicado: (2024)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
por: Dai, Sunhao, et al.
Publicado: (2024)
por: Dai, Sunhao, et al.
Publicado: (2024)
Building Russian Benchmark for Evaluation of Information Retrieval Models
por: Kovalev, Grigory, et al.
Publicado: (2025)
por: Kovalev, Grigory, et al.
Publicado: (2025)
MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
por: Zhang, Siyue, et al.
Publicado: (2025)
por: Zhang, Siyue, et al.
Publicado: (2025)
Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain
por: Guo, Jing, et al.
Publicado: (2025)
por: Guo, Jing, et al.
Publicado: (2025)
LogEval: A Comprehensive Benchmark Suite for Large Language Models In Log Analysis
por: Cui, Tianyu, et al.
Publicado: (2024)
por: Cui, Tianyu, et al.
Publicado: (2024)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
por: Chen, Jianlyu, et al.
Publicado: (2024)
por: Chen, Jianlyu, et al.
Publicado: (2024)
Toward General Instruction-Following Alignment for Retrieval-Augmented Generation
por: Dong, Guanting, et al.
Publicado: (2024)
por: Dong, Guanting, et al.
Publicado: (2024)
JUÁ -- A Benchmark for Information Retrieval in Brazilian Legal Text Collections
por: Pereira, Jayr, et al.
Publicado: (2026)
por: Pereira, Jayr, et al.
Publicado: (2026)
On Synthetic Data Strategies for Domain-Specific Generative Retrieval
por: Wen, Haoyang, et al.
Publicado: (2025)
por: Wen, Haoyang, et al.
Publicado: (2025)
Evaluating the Robustness of Retrieval-Augmented Generation to Adversarial Evidence in the Health Domain
por: Amirshahi, Shakiba, et al.
Publicado: (2025)
por: Amirshahi, Shakiba, et al.
Publicado: (2025)
PosIR: Position-Aware Heterogeneous Information Retrieval Benchmark
por: Zeng, Ziyang, et al.
Publicado: (2026)
por: Zeng, Ziyang, et al.
Publicado: (2026)
Evaluating Generative Ad Hoc Information Retrieval
por: Gienapp, Lukas, et al.
Publicado: (2023)
por: Gienapp, Lukas, et al.
Publicado: (2023)
RVR: Retrieve-Verify-Retrieve for Comprehensive Question Answering
por: Qian, Deniz, et al.
Publicado: (2026)
por: Qian, Deniz, et al.
Publicado: (2026)
Wikipedia-based Datasets in Russian Information Retrieval Benchmark RusBEIR
por: Kovalev, Grigory, et al.
Publicado: (2025)
por: Kovalev, Grigory, et al.
Publicado: (2025)
A Comprehensive Taxonomy of Negation for NLP and Neural Retrievers
por: Petcu, Roxana, et al.
Publicado: (2025)
por: Petcu, Roxana, et al.
Publicado: (2025)
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
por: Xi, Yunjia, et al.
Publicado: (2025)
por: Xi, Yunjia, et al.
Publicado: (2025)
Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation
por: Abdallah, Abdelrahman, et al.
Publicado: (2025)
por: Abdallah, Abdelrahman, et al.
Publicado: (2025)
A Survey of Reasoning-Intensive Retrieval: Progress and Challenges
por: Wei, Yiyang, et al.
Publicado: (2026)
por: Wei, Yiyang, et al.
Publicado: (2026)
Benchmarking Information Retrieval Models on Complex Retrieval Tasks
por: Killingback, Julian, et al.
Publicado: (2025)
por: Killingback, Julian, et al.
Publicado: (2025)
REAR: A Relevance-Aware Retrieval-Augmented Framework for Open-Domain Question Answering
por: Wang, Yuhao, et al.
Publicado: (2024)
por: Wang, Yuhao, et al.
Publicado: (2024)
Evaluating LLM Abilities to Understand Tabular Electronic Health Records: A Comprehensive Study of Patient Data Extraction and Retrieval
por: Lovon, Jesus, et al.
Publicado: (2025)
por: Lovon, Jesus, et al.
Publicado: (2025)
STELLA: Self-Reflective Terminology-Aware Framework for Building an Aerospace Information Retrieval Benchmark
por: Kim, Bongmin
Publicado: (2026)
por: Kim, Bongmin
Publicado: (2026)
Anveshana: A New Benchmark Dataset for Cross-Lingual Information Retrieval On English Queries and Sanskrit Documents
por: Jagadeeshan, Manoj Balaji, et al.
Publicado: (2025)
por: Jagadeeshan, Manoj Balaji, et al.
Publicado: (2025)
MIRB: Mathematical Information Retrieval Benchmark
por: Ju, Haocheng, et al.
Publicado: (2025)
por: Ju, Haocheng, et al.
Publicado: (2025)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
por: Cheng, Yiruo, et al.
Publicado: (2024)
por: Cheng, Yiruo, et al.
Publicado: (2024)
RAR-b: Reasoning as Retrieval Benchmark
por: Xiao, Chenghao, et al.
Publicado: (2024)
por: Xiao, Chenghao, et al.
Publicado: (2024)
DAPR: A Benchmark on Document-Aware Passage Retrieval
por: Wang, Kexin, et al.
Publicado: (2023)
por: Wang, Kexin, et al.
Publicado: (2023)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
por: Rau, David, et al.
Publicado: (2024)
por: Rau, David, et al.
Publicado: (2024)
UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers
por: Min, Dehai, et al.
Publicado: (2024)
por: Min, Dehai, et al.
Publicado: (2024)
Plan-and-Refine: Diverse and Comprehensive Retrieval-Augmented Generation
por: Salemi, Alireza, et al.
Publicado: (2025)
por: Salemi, Alireza, et al.
Publicado: (2025)
Ejemplares similares
-
Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
por: Zhao, Yilun, et al.
Publicado: (2026) -
FollowTable: A Benchmark for Instruction-Following Table Retrieval
por: Jin, Rihui, et al.
Publicado: (2026) -
LimRank: Less is More for Reasoning-Intensive Information Reranking
por: Song, Tingyu, et al.
Publicado: (2025) -
Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
por: Song, Tingyu, et al.
Publicado: (2026) -
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
por: Wang, Shuting, et al.
Publicado: (2024)