MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
Fuente:
arXiv
Saved in:
| Main Authors: | Rosenthal, Sara, Katsis, Yannis, Shah, Vraj, He, Lihong, Popa, Lucian, Danilevsky, Marina |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
by: Katsis, Yannis, et al.
Published: (2025)
by: Katsis, Yannis, et al.
Published: (2025)
A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks
by: Rosenthal, Sara, et al.
Published: (2025)
by: Rosenthal, Sara, et al.
Published: (2025)
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
by: Fadnis, Kshitij, et al.
Published: (2025)
by: Fadnis, Kshitij, et al.
Published: (2025)
InspectorRAGet: An Introspection Platform for RAG Evaluation
by: Fadnis, Kshitij, et al.
Published: (2024)
by: Fadnis, Kshitij, et al.
Published: (2024)
A Survey of the State of Explainable AI for Natural Language Processing
by: Danilevsky, Marina, et al.
Published: (2020)
by: Danilevsky, Marina, et al.
Published: (2020)
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
by: Li, Haitao, et al.
Published: (2025)
by: Li, Haitao, et al.
Published: (2025)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
by: Sirdeshmukh, Ved, et al.
Published: (2025)
by: Sirdeshmukh, Ved, et al.
Published: (2025)
A Library of LLM Intrinsics for Retrieval-Augmented Generation
by: Danilevsky, Marina, et al.
Published: (2025)
by: Danilevsky, Marina, et al.
Published: (2025)
DELIFT: Data Efficient Language model Instruction Fine Tuning
by: Agarwal, Ishika, et al.
Published: (2024)
by: Agarwal, Ishika, et al.
Published: (2024)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
H-RAG at SemEval-2026 Task 8: Hierarchical Parent-Child Retrieval for Multi-Turn RAG Conversations
by: Elchafei, Passant, et al.
Published: (2026)
by: Elchafei, Passant, et al.
Published: (2026)
Rethinking Stateful Tool Use in Multi-Turn Dialogues: Benchmarks and Challenges
by: Wang, Hongru, et al.
Published: (2025)
by: Wang, Hongru, et al.
Published: (2025)
CLAPNQ: Cohesive Long-form Answers from Passages in Natural Questions for RAG systems
by: Rosenthal, Sara, et al.
Published: (2024)
by: Rosenthal, Sara, et al.
Published: (2024)
Personalized Turn-Level User Conversation Satisfaction Benchmark
by: Wang, Zhefan, et al.
Published: (2026)
by: Wang, Zhefan, et al.
Published: (2026)
Eliciting Behaviors in Multi-Turn Conversations
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
LLMs Get Lost In Multi-Turn Conversation
by: Laban, Philippe, et al.
Published: (2025)
by: Laban, Philippe, et al.
Published: (2025)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
by: He, Yun, et al.
Published: (2024)
by: He, Yun, et al.
Published: (2024)
InfoQuest: Evaluating Multi-Turn Dialogue Agents for Open-Ended Conversations with Hidden Context
by: de Oliveira, Bryan L. M., et al.
Published: (2025)
by: de Oliveira, Bryan L. M., et al.
Published: (2025)
ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
by: Luo, Sichun, et al.
Published: (2025)
by: Luo, Sichun, et al.
Published: (2025)
MTP: A Dataset for Multi-Modal Turning Points in Casual Conversations
by: Ho, Gia-Bao Dinh, et al.
Published: (2024)
by: Ho, Gia-Bao Dinh, et al.
Published: (2024)
TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models
by: Zhang, Yiran, et al.
Published: (2025)
by: Zhang, Yiran, et al.
Published: (2025)
ORAssistant: A Custom RAG-based Conversational Assistant for OpenROAD
by: Kaintura, Aviral, et al.
Published: (2024)
by: Kaintura, Aviral, et al.
Published: (2024)
MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
by: Epstein, Elliot L., et al.
Published: (2024)
by: Epstein, Elliot L., et al.
Published: (2024)
SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following
by: Canaverde, Beatriz, et al.
Published: (2026)
by: Canaverde, Beatriz, et al.
Published: (2026)
Boiling the Frog: A Multi-Turn Benchmark for Agentic Safety
by: Bisconti, Piercosma, et al.
Published: (2026)
by: Bisconti, Piercosma, et al.
Published: (2026)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
by: Pakhomov, Egor, et al.
Published: (2025)
by: Pakhomov, Egor, et al.
Published: (2025)
C-MTCSD: A Chinese Multi-Turn Conversational Stance Detection Dataset
by: Niu, Fuqiang, et al.
Published: (2025)
by: Niu, Fuqiang, et al.
Published: (2025)
Proactive Guidance of Multi-Turn Conversation in Industrial Search
by: Li, Xiaoyu, et al.
Published: (2025)
by: Li, Xiaoyu, et al.
Published: (2025)
NaturalTurn: A Method to Segment Speech into Psychologically Meaningful Conversational Turns
by: Cooney, Gus, et al.
Published: (2024)
by: Cooney, Gus, et al.
Published: (2024)
Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap
by: Chen, Tianlang, et al.
Published: (2026)
by: Chen, Tianlang, et al.
Published: (2026)
MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
by: Singh, Jyotika, et al.
Published: (2026)
by: Singh, Jyotika, et al.
Published: (2026)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
by: Juneja, Prerna, et al.
Published: (2026)
by: Juneja, Prerna, et al.
Published: (2026)
Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
by: Feng, Xueyang, et al.
Published: (2025)
by: Feng, Xueyang, et al.
Published: (2025)
Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
by: Ramnath, Sahana, et al.
Published: (2025)
by: Ramnath, Sahana, et al.
Published: (2025)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
by: Liu, Junyu, et al.
Published: (2026)
by: Liu, Junyu, et al.
Published: (2026)
CRAG -- Comprehensive RAG Benchmark
by: Yang, Xiao, et al.
Published: (2024)
by: Yang, Xiao, et al.
Published: (2024)
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
Quantifying Conversational Reliability of Large Language Models under Multi-Turn Interaction
by: Myung, Jiyoon
Published: (2026)
by: Myung, Jiyoon
Published: (2026)
Similar Items
-
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
by: Katsis, Yannis, et al.
Published: (2025) -
A Longitudinal Study on Different Annotator Feedback Loops in Complex RAG Tasks
by: Rosenthal, Sara, et al.
Published: (2025) -
RAGAPHENE: A RAG Annotation Platform with Human Enhancements and Edits
by: Fadnis, Kshitij, et al.
Published: (2025) -
InspectorRAGet: An Introspection Platform for RAG Evaluation
by: Fadnis, Kshitij, et al.
Published: (2024) -
A Survey of the State of Explainable AI for Natural Language Processing
by: Danilevsky, Marina, et al.
Published: (2020)