MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Katsis, Yannis, Rosenthal, Sara, Fadnis, Kshitij, Gunasekara, Chulaka, Lee, Young-Suk, Popa, Lucian, Shah, Vraj, Zhu, Huaiyu, Contractor, Danish, Danilevsky, Marina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
by: Rosenthal, Sara, et al.
Published: (2026)
by: Rosenthal, Sara, et al.
Published: (2026)
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
by: Fadnis, Kshitij, et al.
Published: (2025)
by: Fadnis, Kshitij, et al.
Published: (2025)
A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks
by: Rosenthal, Sara, et al.
Published: (2025)
by: Rosenthal, Sara, et al.
Published: (2025)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
by: Danilevsky, Marina, et al.
Published: (2025)
by: Danilevsky, Marina, et al.
Published: (2025)
InspectorRAGet: An Introspection Platform for RAG Evaluation
by: Fadnis, Kshitij, et al.
Published: (2024)
by: Fadnis, Kshitij, et al.
Published: (2024)
Multi-Document Grounded Multi-Turn Synthetic Dialog Generation
by: Lee, Young-Suk, et al.
Published: (2024)
by: Lee, Young-Suk, et al.
Published: (2024)
Reducing the Scope of Language Models
by: Yunis, David, et al.
Published: (2024)
by: Yunis, David, et al.
Published: (2024)
Activated LoRA: Fine-tuned LLMs for Intrinsics
by: Greenewald, Kristjan, et al.
Published: (2025)
by: Greenewald, Kristjan, et al.
Published: (2025)
Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
by: Crouse, Maxwell, et al.
Published: (2026)
by: Crouse, Maxwell, et al.
Published: (2026)
A Survey of the State of Explainable AI for Natural Language Processing
by: Danilevsky, Marina, et al.
Published: (2020)
by: Danilevsky, Marina, et al.
Published: (2020)
Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
by: Venkateswaran, Praveen, et al.
Published: (2025)
by: Venkateswaran, Praveen, et al.
Published: (2025)
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
by: Li, Haitao, et al.
Published: (2025)
by: Li, Haitao, et al.
Published: (2025)
Training with Pseudo-Code for Instruction Following
by: Kumar, Prince, et al.
Published: (2025)
by: Kumar, Prince, et al.
Published: (2025)
KCIF: Knowledge-Conditioned Instruction Following
by: Murthy, Rudra, et al.
Published: (2024)
by: Murthy, Rudra, et al.
Published: (2024)
DELIFT: Data Efficient Language model Instruction Fine Tuning
by: Agarwal, Ishika, et al.
Published: (2024)
by: Agarwal, Ishika, et al.
Published: (2024)
New Tools are Needed for Tracking Adherence to AI Model Behavioral Use Clauses
by: McDuff, Daniel, et al.
Published: (2025)
by: McDuff, Daniel, et al.
Published: (2025)
Leveraging Large Language Models to Enhance Domain Expert Inclusion in Data Science Workflows
by: Shih, Jasmine Y., et al.
Published: (2024)
by: Shih, Jasmine Y., et al.
Published: (2024)
The sinusoidal valley: a recipe for high peaks in the scalar and induced tensor spectra
by: Katsis, Aris
Published: (2025)
by: Katsis, Aris
Published: (2025)
T-Retriever: Tree-based Hierarchical Retrieval Augmented Generation for Textual Graphs
by: Wei, Chunyu, et al.
Published: (2026)
by: Wei, Chunyu, et al.
Published: (2026)
eDCF: Estimating Intrinsic Dimension using Local Connectivity
by: Gupta, Dhruv, et al.
Published: (2025)
by: Gupta, Dhruv, et al.
Published: (2025)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
by: Cheng, Yiruo, et al.
Published: (2024)
by: Cheng, Yiruo, et al.
Published: (2024)
Towards Retrieval Augmented Generation over Large Video Libraries
by: Tevissen, Yannis, et al.
Published: (2024)
by: Tevissen, Yannis, et al.
Published: (2024)
ConTReGen: Context-driven Tree-structured Retrieval for Open-domain Long-form Text Generation
by: Roy, Kashob Kumar, et al.
Published: (2024)
by: Roy, Kashob Kumar, et al.
Published: (2024)
Rome Constipation Symptoms Augmented by Painful Defecation Predicts Specific Subtypes of Refractory Constipation
by: Christian Lambiase, et al.
Published: (2025)
by: Christian Lambiase, et al.
Published: (2025)
Persistent and Conversational Multi-Method Explainability for Trustworthy Financial AI
by: Makridis, Georgios, et al.
Published: (2026)
by: Makridis, Georgios, et al.
Published: (2026)
Personalized Turn-Level User Conversation Satisfaction Benchmark
by: Wang, Zhefan, et al.
Published: (2026)
by: Wang, Zhefan, et al.
Published: (2026)
ZT-SDN: An ML-powered Zero-Trust Architecture for Software-Defined Networks
by: Katsis, Charalampos, et al.
Published: (2024)
by: Katsis, Charalampos, et al.
Published: (2024)
Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance
by: Nagy, Akos, et al.
Published: (2025)
by: Nagy, Akos, et al.
Published: (2025)
Conversation AI Dialog for Medicare powered by Finetuning and Retrieval Augmented Generation
by: Agrawal, Atharva Mangeshkumar, et al.
Published: (2025)
by: Agrawal, Atharva Mangeshkumar, et al.
Published: (2025)
The 3rd Place Solution of CCIR CUP 2025: A Framework for Retrieval-Augmented Generation in Multi-Turn Legal Conversation
by: Li, Da, et al.
Published: (2025)
by: Li, Da, et al.
Published: (2025)
Composed Image Retrieval for Training-Free Domain Conversion
by: Efthymiadis, Nikos, et al.
Published: (2024)
by: Efthymiadis, Nikos, et al.
Published: (2024)
Generative AI in Higher Education: Evidence from an Elite College
by: Contractor, Zara, et al.
Published: (2025)
by: Contractor, Zara, et al.
Published: (2025)
Thurston construction mapping classes with minimal dilatation
by: Contractor, Maryam, et al.
Published: (2024)
by: Contractor, Maryam, et al.
Published: (2024)
MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
by: Lee, Young-Jun, et al.
Published: (2025)
by: Lee, Young-Jun, et al.
Published: (2025)
LARA: Linguistic-Adaptive Retrieval-Augmentation for Multi-Turn Intent Classification
by: Liu, Junhua, et al.
Published: (2024)
by: Liu, Junhua, et al.
Published: (2024)
A rare case of Fanconi anemia with Mitomycin C sensitivity: A pediatrics case report
by: Vraj Bhatt, et al.
Published: (2024)
by: Vraj Bhatt, et al.
Published: (2024)
Retrieval Augmented Conversational Recommendation with Reinforcement Learning
by: Yue, Zhenrui, et al.
Published: (2026)
by: Yue, Zhenrui, et al.
Published: (2026)
Adaptive Retrieval-Augmented Generation for Conversational Systems
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
Reasoning-Augmented Conversation for Multi-Turn Jailbreak Attacks on Large Language Models
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling
by: Elder, Benjamin, et al.
Published: (2025)
by: Elder, Benjamin, et al.
Published: (2025)
Similar Items
-
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
by: Rosenthal, Sara, et al.
Published: (2026) -
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
by: Fadnis, Kshitij, et al.
Published: (2025) -
A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks
by: Rosenthal, Sara, et al.
Published: (2025) -
A Library of LLM Intrinsics for Retrieval-Augmented Generation
by: Danilevsky, Marina, et al.
Published: (2025) -
InspectorRAGet: An Introspection Platform for RAG Evaluation
by: Fadnis, Kshitij, et al.
Published: (2024)