MIRAGE-Bench: Automatic Multilingual Benchmark Arena for Retrieval-Augmented Generation Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Thakur, Nandan, Kazi, Suleman, Luo, Ge, Lin, Jimmy, Ahmad, Amin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
di: Kuissi, Nathan, et al.
Pubblicazione: (2026)
di: Kuissi, Nathan, et al.
Pubblicazione: (2026)
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
di: Park, Chanhee, et al.
Pubblicazione: (2025)
di: Park, Chanhee, et al.
Pubblicazione: (2025)
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
di: Thakur, Nandan, et al.
Pubblicazione: (2023)
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
Ragnarök: A Reusable RAG Framework and Baselines for TREC 2024 Retrieval-Augmented Generation Track
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
di: Pradeep, Ronak, et al.
Pubblicazione: (2024)
ORBIT: Scalable and Verifiable Data Generation for Search Agents on a Tight Budget
di: Thakur, Nandan, et al.
Pubblicazione: (2026)
di: Thakur, Nandan, et al.
Pubblicazione: (2026)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
di: Bao, Forrest Sheng, et al.
Pubblicazione: (2024)
di: Bao, Forrest Sheng, et al.
Pubblicazione: (2024)
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation
di: Lim, Hyeonseok, et al.
Pubblicazione: (2024)
di: Lim, Hyeonseok, et al.
Pubblicazione: (2024)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
di: Chen, Haotian, et al.
Pubblicazione: (2025)
di: Chen, Haotian, et al.
Pubblicazione: (2025)
MEMERAG: A Multilingual End-to-End Meta-Evaluation Benchmark for Retrieval Augmented Generation
di: Blandón, María Andrea Cruz, et al.
Pubblicazione: (2025)
di: Blandón, María Andrea Cruz, et al.
Pubblicazione: (2025)
RAGBench: Explainable Benchmark for Retrieval-Augmented Generation Systems
di: Friel, Robert, et al.
Pubblicazione: (2024)
di: Friel, Robert, et al.
Pubblicazione: (2024)
Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Task
di: Ranaldi, Leonardo, et al.
Pubblicazione: (2025)
di: Ranaldi, Leonardo, et al.
Pubblicazione: (2025)
On the Consistency of Multilingual Context Utilization in Retrieval-Augmented Generation
di: Qi, Jirui, et al.
Pubblicazione: (2025)
di: Qi, Jirui, et al.
Pubblicazione: (2025)
Support Evaluation for the TREC 2024 RAG Track: Comparing Human versus LLM Judges
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
di: Thakur, Nandan, et al.
Pubblicazione: (2025)
WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives
di: Yu, Yongan, et al.
Pubblicazione: (2025)
di: Yu, Yongan, et al.
Pubblicazione: (2025)
Benchmarking Retrieval-Augmented Generation for Medicine
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
di: Xiong, Guangzhi, et al.
Pubblicazione: (2024)
A Survey on Retrieval-Augmented Text Generation for Large Language Models
di: Huang, Yizheng, et al.
Pubblicazione: (2024)
di: Huang, Yizheng, et al.
Pubblicazione: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
di: Jin, Zhuoran, et al.
Pubblicazione: (2024)
di: Jin, Zhuoran, et al.
Pubblicazione: (2024)
HoH: A Dynamic Benchmark for Evaluating the Impact of Outdated Information on Retrieval-Augmented Generation
di: Ouyang, Jie, et al.
Pubblicazione: (2025)
di: Ouyang, Jie, et al.
Pubblicazione: (2025)
Towards Global Retrieval Augmented Generation: A Benchmark for Corpus-Level Reasoning
di: Luo, Qi, et al.
Pubblicazione: (2025)
di: Luo, Qi, et al.
Pubblicazione: (2025)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
di: Fu, Jia, et al.
Pubblicazione: (2024)
di: Fu, Jia, et al.
Pubblicazione: (2024)
MRAG: Benchmarking Retrieval-Augmented Generation for Bio-medicine
di: Li, Liz, et al.
Pubblicazione: (2026)
di: Li, Liz, et al.
Pubblicazione: (2026)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
di: Katsis, Yannis, et al.
Pubblicazione: (2025)
di: Katsis, Yannis, et al.
Pubblicazione: (2025)
CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analytics
di: Agarwal, Parth, et al.
Pubblicazione: (2025)
di: Agarwal, Parth, et al.
Pubblicazione: (2025)
From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
di: Li, Tianle, et al.
Pubblicazione: (2024)
di: Li, Tianle, et al.
Pubblicazione: (2024)
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation
di: Mao, Qianren, et al.
Pubblicazione: (2024)
di: Mao, Qianren, et al.
Pubblicazione: (2024)
Benchmarking Retrieval-Augmented Generation for Chemistry
di: Zhong, Xianrui, et al.
Pubblicazione: (2025)
di: Zhong, Xianrui, et al.
Pubblicazione: (2025)
Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs
di: Conia, Simone, et al.
Pubblicazione: (2024)
di: Conia, Simone, et al.
Pubblicazione: (2024)
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
di: Xiao, Yilin, et al.
Pubblicazione: (2025)
di: Xiao, Yilin, et al.
Pubblicazione: (2025)
Retrieval-Augmented Generation Systems for Intellectual Property via Synthetic Multi-Angle Fine-tuning
di: Ren, Runtao, et al.
Pubblicazione: (2025)
di: Ren, Runtao, et al.
Pubblicazione: (2025)
RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
di: Han, Rujun, et al.
Pubblicazione: (2024)
di: Han, Rujun, et al.
Pubblicazione: (2024)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
di: Liu, Chaoqun, et al.
Pubblicazione: (2025)
di: Liu, Chaoqun, et al.
Pubblicazione: (2025)
ArenaBencher: Automatic Benchmark Evolution via Multi-Model Competitive Evaluation
di: Liu, Qin, et al.
Pubblicazione: (2025)
di: Liu, Qin, et al.
Pubblicazione: (2025)
MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language Models
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering
di: Sultana, Nusrat, et al.
Pubblicazione: (2026)
di: Sultana, Nusrat, et al.
Pubblicazione: (2026)
Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language
di: Singh, Jaskaranjeet, et al.
Pubblicazione: (2025)
di: Singh, Jaskaranjeet, et al.
Pubblicazione: (2025)
Towards Automated Smart Contract Generation: Evaluation, Benchmarking, and Retrieval-Augmented Repair
di: Chen, Zaoyu, et al.
Pubblicazione: (2025)
di: Chen, Zaoyu, et al.
Pubblicazione: (2025)
MASLegalBench: Benchmarking Multi-Agent Systems in Deductive Legal Reasoning
di: Jing, Huihao, et al.
Pubblicazione: (2025)
di: Jing, Huihao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
di: Kuissi, Nathan, et al.
Pubblicazione: (2026) -
MIRAGE: A Metric-Intensive Benchmark for Retrieval-Augmented Generation Evaluation
di: Park, Chanhee, et al.
Pubblicazione: (2025) -
Leveraging LLMs for Synthesizing Training Data Across Many Languages in Multilingual Dense Retrieval
di: Thakur, Nandan, et al.
Pubblicazione: (2023) -
FreshStack: Building Realistic Benchmarks for Evaluating Retrieval on Technical Documents
di: Thakur, Nandan, et al.
Pubblicazione: (2025) -
Hard Negatives, Hard Lessons: Revisiting Training Data Quality for Robust Information Retrieval with LLMs
di: Thakur, Nandan, et al.
Pubblicazione: (2025)