RAGCap-Bench: Benchmarking Capabilities of LLMs in Agentic Retrieval Augmented Generation Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Jingru, Zhang, Chen, Liu, Stephen Y., Li, Haizhou |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
by: Thakur, Nandan, et al.
Published: (2024)
by: Thakur, Nandan, et al.
Published: (2024)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025)
by: Chen, Haotian, et al.
Published: (2025)
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
by: Kuo, Tzu-Lin, et al.
Published: (2024)
by: Kuo, Tzu-Lin, et al.
Published: (2024)
DentalBench: Benchmarking and Advancing LLMs Capability for Bilingual Dentistry Understanding
by: Zhu, Hengchuan, et al.
Published: (2025)
by: Zhu, Hengchuan, et al.
Published: (2025)
LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generation
by: Itai, Koki, et al.
Published: (2026)
by: Itai, Koki, et al.
Published: (2026)
BenchBench: Benchmarking Automated Benchmark Generation
by: Zheng, Yandan, et al.
Published: (2026)
by: Zheng, Yandan, et al.
Published: (2026)
Retrieval-Augmented Dialogue Knowledge Aggregation for Expressive Conversational Speech Synthesis
by: Liu, Rui, et al.
Published: (2025)
by: Liu, Rui, et al.
Published: (2025)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
by: Singh, Harman, et al.
Published: (2024)
by: Singh, Harman, et al.
Published: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
Thought-Retriever: Don't Just Retrieve Raw Data, Retrieve Thoughts for Memory-Augmented Agentic Systems
by: Feng, Tao, et al.
Published: (2026)
by: Feng, Tao, et al.
Published: (2026)
Legal-DC: Benchmarking Retrieval-Augmented Generation for Legal Documents
by: Li, Yaocong, et al.
Published: (2026)
by: Li, Yaocong, et al.
Published: (2026)
GraphSearch: An Agentic Deep Searching Workflow for Graph Retrieval-Augmented Generation
by: Yang, Cehao, et al.
Published: (2025)
by: Yang, Cehao, et al.
Published: (2025)
ReZG: Retrieval-Augmented Zero-Shot Counter Narrative Generation for Hate Speech
by: Jiang, Shuyu, et al.
Published: (2023)
by: Jiang, Shuyu, et al.
Published: (2023)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
by: Lim, Hyeonseok, et al.
Published: (2024)
by: Lim, Hyeonseok, et al.
Published: (2024)
CRAB: A Benchmark for Evaluating Curation of Retrieval-Augmented LLMs in Biomedicine
by: Zhong, Hanmeng, et al.
Published: (2025)
by: Zhong, Hanmeng, et al.
Published: (2025)
TritonBench: Benchmarking Large Language Model Capabilities for Generating Triton Operators
by: Li, Jianling, et al.
Published: (2025)
by: Li, Jianling, et al.
Published: (2025)
VoiceBench: Benchmarking LLM-Based Voice Assistants
by: Chen, Yiming, et al.
Published: (2024)
by: Chen, Yiming, et al.
Published: (2024)
Graph Retrieval-Augmented Generation: A Survey
by: Peng, Boci, et al.
Published: (2024)
by: Peng, Boci, et al.
Published: (2024)
LLMs are Biased Evaluators But Not Biased for Retrieval Augmented Generation
by: Chen, Yen-Shan, et al.
Published: (2024)
by: Chen, Yen-Shan, et al.
Published: (2024)
Benchmarking Retrieval-Augmented Generation for Medicine
by: Xiong, Guangzhi, et al.
Published: (2024)
by: Xiong, Guangzhi, et al.
Published: (2024)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
by: Friel, Robert, et al.
Published: (2024)
by: Friel, Robert, et al.
Published: (2024)
PerfCodeBench: Benchmarking LLMs for System-Level High-Performance Code Optimization
by: Jing, Huihao, et al.
Published: (2026)
by: Jing, Huihao, et al.
Published: (2026)
Agentic Retrieval-Augmented Generation: A Survey on Agentic RAG
by: Singh, Aditi, et al.
Published: (2025)
by: Singh, Aditi, et al.
Published: (2025)
Does RAG Introduce Unfairness in LLMs? Evaluating Fairness in Retrieval-Augmented Generation Systems
by: Wu, Xuyang, et al.
Published: (2024)
by: Wu, Xuyang, et al.
Published: (2024)
A-RAG: Scaling Agentic Retrieval-Augmented Generation via Hierarchical Retrieval Interfaces
by: Du, Mingxuan, et al.
Published: (2026)
by: Du, Mingxuan, et al.
Published: (2026)
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation
by: Zhou, Li, et al.
Published: (2025)
by: Zhou, Li, et al.
Published: (2025)
RadFabric: Agentic AI System with Reasoning Capability for Radiology
by: Chen, Wenting, et al.
Published: (2025)
by: Chen, Wenting, et al.
Published: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
by: Chen, Jianlyu, et al.
Published: (2024)
by: Chen, Jianlyu, et al.
Published: (2024)
WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives
by: Yu, Yongan, et al.
Published: (2025)
by: Yu, Yongan, et al.
Published: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
RiddleBench: A New Generative Reasoning Benchmark for LLMs
by: Halder, Deepon, et al.
Published: (2025)
by: Halder, Deepon, et al.
Published: (2025)
MRAG: Benchmarking Retrieval-Augmented Generation for Bio-medicine
by: Li, Liz, et al.
Published: (2026)
by: Li, Liz, et al.
Published: (2026)
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing
by: Li, Kuan, et al.
Published: (2025)
by: Li, Kuan, et al.
Published: (2025)
GameDevBench: Evaluating Agentic Capabilities Through Game Development
by: Chi, Wayne, et al.
Published: (2026)
by: Chi, Wayne, et al.
Published: (2026)
CodeRAG-Bench: Can Retrieval Augment Code Generation?
by: Wang, Zora Zhiruo, et al.
Published: (2024)
by: Wang, Zora Zhiruo, et al.
Published: (2024)
Dialogue Benchmark Generation from Knowledge Graphs with Cost-Effective Retrieval-Augmented LLMs
by: Omar, Reham, et al.
Published: (2025)
by: Omar, Reham, et al.
Published: (2025)
Enhancing Large Language Models (LLMs) for Telecommunications using Knowledge Graphs and Retrieval-Augmented Generation
by: Yuan, Dun, et al.
Published: (2025)
by: Yuan, Dun, et al.
Published: (2025)
Similar Items
-
MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
by: Thakur, Nandan, et al.
Published: (2024) -
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
by: Chen, Haotian, et al.
Published: (2025) -
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
by: Xi, Yunjia, et al.
Published: (2025) -
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025) -
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
by: Kuo, Tzu-Lin, et al.
Published: (2024)