Evaluating the Performance of RAG Methods for Conversational AI in the Airport Domain
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yuyang, Kerbusch, Philip J. M., Pruim, Raimon H. R., Käfer, Tobias |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA
von: Alushi, Klejda, et al.
Veröffentlicht: (2026)
von: Alushi, Klejda, et al.
Veröffentlicht: (2026)
From Transcripts to AI Agents: Knowledge Extraction, RAG Integration, and Robust Evaluation of Conversational AI Assistants
von: Pachtrachai, Krittin, et al.
Veröffentlicht: (2026)
von: Pachtrachai, Krittin, et al.
Veröffentlicht: (2026)
Mentalic Net: Development of RAG-based Conversational AI and Evaluation Framework for Mental Health Support
von: Dutta, Anandi, et al.
Veröffentlicht: (2025)
von: Dutta, Anandi, et al.
Veröffentlicht: (2025)
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching
von: Ding, Yuyang, et al.
Veröffentlicht: (2024)
von: Ding, Yuyang, et al.
Veröffentlicht: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
von: Wang, Shuting, et al.
Veröffentlicht: (2024)
AILS-NTUA at SemEval-2026 Task 8: Evaluating Multi-Turn RAG Conversations
von: Athanasiou, Dimosthenis, et al.
Veröffentlicht: (2026)
von: Athanasiou, Dimosthenis, et al.
Veröffentlicht: (2026)
Evaluating AI Meeting Summaries with a Reusable Cross-Domain Pipeline
von: Zhong, Philip, et al.
Veröffentlicht: (2026)
von: Zhong, Philip, et al.
Veröffentlicht: (2026)
Ranking Free RAG: Replacing Re-ranking with Selection in RAG for Sensitive Domains
von: Saxena, Yash, et al.
Veröffentlicht: (2025)
von: Saxena, Yash, et al.
Veröffentlicht: (2025)
RAGalyst: Automated Human-Aligned Agentic Evaluation for Domain-Specific RAG
von: Gao, Joshua, et al.
Veröffentlicht: (2025)
von: Gao, Joshua, et al.
Veröffentlicht: (2025)
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
von: Hasan, Md. Toufique, et al.
Veröffentlicht: (2026)
von: Hasan, Md. Toufique, et al.
Veröffentlicht: (2026)
Foundation Metrics for Evaluating Effectiveness of Healthcare Conversations Powered by Generative AI
von: Abbasian, Mahyar, et al.
Veröffentlicht: (2023)
von: Abbasian, Mahyar, et al.
Veröffentlicht: (2023)
Evaluating Factual Density in Multi-Source RAG: A Study in Medical AI Accuracy
von: DeMarco, Michael R.
Veröffentlicht: (2026)
von: DeMarco, Michael R.
Veröffentlicht: (2026)
Enhancing RAG with Active Learning on Conversation Records: Reject Incapables and Answer Capables
von: Geng, Xuzhao, et al.
Veröffentlicht: (2025)
von: Geng, Xuzhao, et al.
Veröffentlicht: (2025)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
von: Singh, Ishneet Sukhvinder, et al.
Veröffentlicht: (2024)
von: Singh, Ishneet Sukhvinder, et al.
Veröffentlicht: (2024)
ORAssistant: A Custom RAG-based Conversational Assistant for OpenROAD
von: Kaintura, Aviral, et al.
Veröffentlicht: (2024)
von: Kaintura, Aviral, et al.
Veröffentlicht: (2024)
Dynamic Contexts for Generating Suggestion Questions in RAG Based Conversational Systems
von: Tayal, Anuja, et al.
Veröffentlicht: (2024)
von: Tayal, Anuja, et al.
Veröffentlicht: (2024)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
von: Li, Kun, et al.
Veröffentlicht: (2025)
von: Li, Kun, et al.
Veröffentlicht: (2025)
Long Context vs. RAG for LLMs: An Evaluation and Revisits
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
Domain-Specific Data Generation Framework for RAG Adaptation
von: Tian, Chris Xing, et al.
Veröffentlicht: (2025)
von: Tian, Chris Xing, et al.
Veröffentlicht: (2025)
Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs
von: Anuyah, Sydney, et al.
Veröffentlicht: (2026)
von: Anuyah, Sydney, et al.
Veröffentlicht: (2026)
Evaluation of RAG Metrics for Question Answering in the Telecom Domain
von: Roychowdhury, Sujoy, et al.
Veröffentlicht: (2024)
von: Roychowdhury, Sujoy, et al.
Veröffentlicht: (2024)
FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
von: Zhao, Suifeng, et al.
Veröffentlicht: (2025)
von: Zhao, Suifeng, et al.
Veröffentlicht: (2025)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
QChunker: Learning Question-Aware Text Chunking for Domain RAG via Multi-Agent Debate
von: Zhao, Jihao, et al.
Veröffentlicht: (2026)
von: Zhao, Jihao, et al.
Veröffentlicht: (2026)
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
von: Atf, Zahra, et al.
Veröffentlicht: (2025)
von: Atf, Zahra, et al.
Veröffentlicht: (2025)
Reinforcement Learning for Optimizing RAG for Domain Chatbots
von: Kulkarni, Mandar, et al.
Veröffentlicht: (2024)
von: Kulkarni, Mandar, et al.
Veröffentlicht: (2024)
GraphRAG-Bench: Challenging Domain-Specific Reasoning for Evaluating Graph Retrieval-Augmented Generation
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
von: Xiao, Yilin, et al.
Veröffentlicht: (2025)
RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
von: Han, Rujun, et al.
Veröffentlicht: (2024)
von: Han, Rujun, et al.
Veröffentlicht: (2024)
Convomem Benchmark: Why Your First 150 Conversations Don't Need RAG
von: Pakhomov, Egor, et al.
Veröffentlicht: (2025)
von: Pakhomov, Egor, et al.
Veröffentlicht: (2025)
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
von: Rosenthal, Sara, et al.
Veröffentlicht: (2026)
von: Rosenthal, Sara, et al.
Veröffentlicht: (2026)
Open-Domain Text Evaluation via Contrastive Distribution Methods
von: Lu, Sidi, et al.
Veröffentlicht: (2023)
von: Lu, Sidi, et al.
Veröffentlicht: (2023)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
von: Juneja, Prerna, et al.
Veröffentlicht: (2026)
von: Juneja, Prerna, et al.
Veröffentlicht: (2026)
RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
von: Kim, Kiseung, et al.
Veröffentlicht: (2024)
von: Kim, Kiseung, et al.
Veröffentlicht: (2024)
Context Volume Drives Performance: Tackling Domain Shift in Extremely Low-Resource Translation via RAG
von: Setiawan, David Samuel, et al.
Veröffentlicht: (2026)
von: Setiawan, David Samuel, et al.
Veröffentlicht: (2026)
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI
von: Schirmer, Miriam, et al.
Veröffentlicht: (2024)
von: Schirmer, Miriam, et al.
Veröffentlicht: (2024)
EmoRAG: Evaluating RAG Robustness to Symbolic Perturbations
von: Zhou, Xinyun, et al.
Veröffentlicht: (2025)
von: Zhou, Xinyun, et al.
Veröffentlicht: (2025)
RAFT: Adapting Language Model to Domain Specific RAG
von: Zhang, Tianjun, et al.
Veröffentlicht: (2024)
von: Zhang, Tianjun, et al.
Veröffentlicht: (2024)
BSharedRAG: Backbone Shared Retrieval-Augmented Generation for the E-commerce Domain
von: Guan, Kaisi, et al.
Veröffentlicht: (2024)
von: Guan, Kaisi, et al.
Veröffentlicht: (2024)
The HalluRAG Dataset: Detecting Closed-Domain Hallucinations in RAG Applications Using an LLM's Internal States
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)
von: Ridder, Fabian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA
von: Alushi, Klejda, et al.
Veröffentlicht: (2026) -
From Transcripts to AI Agents: Knowledge Extraction, RAG Integration, and Robust Evaluation of Conversational AI Assistants
von: Pachtrachai, Krittin, et al.
Veröffentlicht: (2026) -
Mentalic Net: Development of RAG-based Conversational AI and Evaluation Framework for Mental Health Support
von: Dutta, Anandi, et al.
Veröffentlicht: (2025) -
OmniEval: An Omnidirectional and Automatic RAG Evaluation Benchmark in Financial Domain
von: Wang, Shuting, et al.
Veröffentlicht: (2024) -
Boosting Large Language Models with Socratic Method for Conversational Mathematics Teaching
von: Ding, Yuyang, et al.
Veröffentlicht: (2024)