Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yuyang, Kerbusch, Philip J. M., Pruim, Raimon H. R., Käfer, Tobias |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA
by: Alushi, Klejda, et al.
Published: (2026)
by: Alushi, Klejda, et al.
Published: (2026)
From Transcripts to AI Agents: Knowledge Extraction, RAG Integration, and Robust Evaluation of Conversational AI Assistants
by: Pachtrachai, Krittin, et al.
Published: (2026)
by: Pachtrachai, Krittin, et al.
Published: (2026)
Mentalic Net: Development of RAG-based Conversational AI and Evaluation Framework for Mental Health Support
by: Dutta, Anandi, et al.
Published: (2025)
by: Dutta, Anandi, et al.
Published: (2025)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching
by: Ding, Yuyang, et al.
Published: (2024)
by: Ding, Yuyang, et al.
Published: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
by: Athanasiou, Dimosthenis, et al.
Published: (2026)
Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline
by: Zhong, Philip, et al.
Published: (2026)
by: Zhong, Philip, et al.
Published: (2026)
Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
by: Saxena, Yash, et al.
Published: (2025)
by: Saxena, Yash, et al.
Published: (2025)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
by: Gao, Joshua, et al.
Published: (2025)
by: Gao, Joshua, et al.
Published: (2025)
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
by: Hasan, Md. Toufique, et al.
Published: (2026)
by: Hasan, Md. Toufique, et al.
Published: (2026)
Foundation Metrics for Evaluating Effectiveness of Healthcare Conversations Powered by Generative AI
by: Abbasian, Mahyar, et al.
Published: (2023)
by: Abbasian, Mahyar, et al.
Published: (2023)
Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy
by: DeMarco, Michael R.
Published: (2026)
by: DeMarco, Michael R.
Published: (2026)
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables
by: Geng, Xuzhao, et al.
Published: (2025)
by: Geng, Xuzhao, et al.
Published: (2025)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
by: Singh, Ishneet Sukhvinder, et al.
Published: (2024)
ORAssistant: A Custom RAG-based Conversational Assistant for OpenROAD
by: Kaintura, Aviral, et al.
Published: (2024)
by: Kaintura, Aviral, et al.
Published: (2024)
Dynamic Contexts for Generating Suggestion Questions in RAG Based Conversational Systems
by: Tayal, Anuja, et al.
Published: (2024)
by: Tayal, Anuja, et al.
Published: (2024)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
by: Li, Kun, et al.
Published: (2025)
by: Li, Kun, et al.
Published: (2025)
Long Context vs. RAG for LLMs: An Evaluation and Revisits
by: Li, Xinze, et al.
Published: (2024)
by: Li, Xinze, et al.
Published: (2024)
Domain-Specific Data Generation Framework for RAG Adaptation
by: Tian, Chris Xing, et al.
Published: (2025)
by: Tian, Chris Xing, et al.
Published: (2025)
Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs
by: Anuyah, Sydney, et al.
Published: (2026)
by: Anuyah, Sydney, et al.
Published: (2026)
Evaluation of RAG Metrics for Question Answering in the Telecom Domain
by: Roychowdhury, Sujoy, et al.
Published: (2024)
by: Roychowdhury, Sujoy, et al.
Published: (2024)
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
by: Zhao, Suifeng, et al.
Published: (2025)
by: Zhao, Suifeng, et al.
Published: (2025)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
by: Li, Yuyang, et al.
Published: (2026)
by: Li, Yuyang, et al.
Published: (2026)
QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent Debate
by: Zhao, Jihao, et al.
Published: (2026)
by: Zhao, Jihao, et al.
Published: (2026)
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
by: Atf, Zahra, et al.
Published: (2025)
by: Atf, Zahra, et al.
Published: (2025)
Reinforcement Learning for Optimizing RAG for Domain Chatbots
by: Kulkarni, Mandar, et al.
Published: (2024)
by: Kulkarni, Mandar, et al.
Published: (2024)
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
by: Xiao, Yilin, et al.
Published: (2025)
by: Xiao, Yilin, et al.
Published: (2025)
RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
by: Han, Rujun, et al.
Published: (2024)
by: Han, Rujun, et al.
Published: (2024)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
by: Pakhomov, Egor, et al.
Published: (2025)
by: Pakhomov, Egor, et al.
Published: (2025)
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
by: Rosenthal, Sara, et al.
Published: (2026)
by: Rosenthal, Sara, et al.
Published: (2026)
Open-Domain Text Evaluation via Contrastive Distribution Methods
by: Lu, Sidi, et al.
Published: (2023)
by: Lu, Sidi, et al.
Published: (2023)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
by: Juneja, Prerna, et al.
Published: (2026)
by: Juneja, Prerna, et al.
Published: (2026)
RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
by: Kim, Kiseung, et al.
Published: (2024)
by: Kim, Kiseung, et al.
Published: (2024)
Context Volume Drives Performance: Tackling Domain Shift in Extremely Low-Resource Translation via RAG
by: Setiawan, David Samuel, et al.
Published: (2026)
by: Setiawan, David Samuel, et al.
Published: (2026)
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI
by: Schirmer, Miriam, et al.
Published: (2024)
by: Schirmer, Miriam, et al.
Published: (2024)
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
by: Zhou, Xinyun, et al.
Published: (2025)
by: Zhou, Xinyun, et al.
Published: (2025)
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
BSharedRAG: Backbone Shared Retrieval-Augmented Generation for the E-commerce Domain
by: Guan, Kaisi, et al.
Published: (2024)
by: Guan, Kaisi, et al.
Published: (2024)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
by: Ridder, Fabian, et al.
Published: (2024)
by: Ridder, Fabian, et al.
Published: (2024)
Similar Items
-
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA
by: Alushi, Klejda, et al.
Published: (2026) -
From Transcripts to AI Agents: Knowledge Extraction, RAG Integration, and Robust Evaluation of Conversational AI Assistants
by: Pachtrachai, Krittin, et al.
Published: (2026) -
Mentalic Net: Development of RAG-based Conversational AI and Evaluation Framework for Mental Health Support
by: Dutta, Anandi, et al.
Published: (2025) -
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
by: Wang, Shuting, et al.
Published: (2024) -
Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching
by: Ding, Yuyang, et al.
Published: (2024)