Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
Fuente:
arXiv
Guardado en:
| Autores principales: | Afzal, Anum, Kowsik, Alexander, Fani, Rajna, Matthes, Florian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data
por: Afzal, Anum, et al.
Publicado: (2024)
por: Afzal, Anum, et al.
Publicado: (2024)
On the Influence of Context Size and Model Choice in Retrieval-Augmented Generation Systems
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
por: Afzal, Anum, et al.
Publicado: (2026)
por: Afzal, Anum, et al.
Publicado: (2026)
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
por: Afzal, Anum, et al.
Publicado: (2025)
por: Afzal, Anum, et al.
Publicado: (2025)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
por: Afzal, Anum, et al.
Publicado: (2025)
por: Afzal, Anum, et al.
Publicado: (2025)
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
por: Afzal, Anum, et al.
Publicado: (2025)
por: Afzal, Anum, et al.
Publicado: (2025)
AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
por: Afzal, Anum, et al.
Publicado: (2024)
por: Afzal, Anum, et al.
Publicado: (2024)
Improving Health Question Answering with Reliable and Time-Aware Evidence Retrieval
por: Vladika, Juraj, et al.
Publicado: (2024)
por: Vladika, Juraj, et al.
Publicado: (2024)
Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
por: Chiang, Wei-Lin, et al.
Publicado: (2024)
por: Chiang, Wei-Lin, et al.
Publicado: (2024)
Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering
por: Han, Rujun, et al.
Publicado: (2024)
por: Han, Rujun, et al.
Publicado: (2024)
Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
por: Afzal, Anum, et al.
Publicado: (2025)
por: Afzal, Anum, et al.
Publicado: (2025)
Cross-lingual Editing in Multilingual Language Models
por: Beniwal, Himanshu, et al.
Publicado: (2024)
por: Beniwal, Himanshu, et al.
Publicado: (2024)
Improving Reliability and Explainability of Medical Question Answering through Atomic Fact Checking in Retrieval-Augmented LLMs
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA
por: Bonomo, Tommaso, et al.
Publicado: (2025)
por: Bonomo, Tommaso, et al.
Publicado: (2025)
MedBioLM: Optimizing Medical and Biological QA with Fine-Tuned Large Language Models and Retrieval-Augmented Generation
por: Kim, Seonok
Publicado: (2025)
por: Kim, Seonok
Publicado: (2025)
ORCHID: Orchestrated Retrieval-Augmented Classification with Human-in-the-Loop Intelligent Decision-Making for High-Risk Property
por: Mahbub, Maria, et al.
Publicado: (2025)
por: Mahbub, Maria, et al.
Publicado: (2025)
SuRe: Summarizing Retrievals using Answer Candidates for Open-domain QA of LLMs
por: Kim, Jaehyung, et al.
Publicado: (2024)
por: Kim, Jaehyung, et al.
Publicado: (2024)
Leveraging Retrieval-Augmented Generation for Culturally Inclusive Hakka Chatbots: Design Insights and User Perceptions
por: Chang, Chen-Chi, et al.
Publicado: (2024)
por: Chang, Chen-Chi, et al.
Publicado: (2024)
HealthFC: Verifying Health Claims with Evidence-Based Medical Fact-Checking
por: Vladika, Juraj, et al.
Publicado: (2023)
por: Vladika, Juraj, et al.
Publicado: (2023)
Step-by-Step Fact Verification System for Medical Claims with Explainable Reasoning
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
Comparing Knowledge Sources for Open-Domain Scientific Claim Verification
por: Vladika, Juraj, et al.
Publicado: (2024)
por: Vladika, Juraj, et al.
Publicado: (2024)
MedSEBA: Synthesizing Evidence-Based Answers Grounded in Evolving Medical Literature
por: Vladika, Juraj, et al.
Publicado: (2025)
por: Vladika, Juraj, et al.
Publicado: (2025)
SciRerankBench: Benchmarking Rerankers Towards Scientific Retrieval-Augmented Generated LLMs
por: Chen, Haotian, et al.
Publicado: (2025)
por: Chen, Haotian, et al.
Publicado: (2025)
CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation
por: Zhao, Kaiwen, et al.
Publicado: (2025)
por: Zhao, Kaiwen, et al.
Publicado: (2025)
RE-RAG: Improving Open-Domain QA Performance and Interpretability with Relevance Estimator in Retrieval-Augmented Generation
por: Kim, Kiseung, et al.
Publicado: (2024)
por: Kim, Kiseung, et al.
Publicado: (2024)
RAG-BioQA: A Retrieval-Augmented Generation Framework for Long-Form Biomedical Question Answering
por: Panchumarthi, Lovely Yeswanth, et al.
Publicado: (2025)
por: Panchumarthi, Lovely Yeswanth, et al.
Publicado: (2025)
Knowledge Graph-Driven Retrieval-Augmented Generation: Integrating Deepseek-R1 with Weaviate for Advanced Chatbot Applications
por: Lecu, Alexandru, et al.
Publicado: (2025)
por: Lecu, Alexandru, et al.
Publicado: (2025)
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
por: Xiao, Yunpeng, et al.
Publicado: (2025)
por: Xiao, Yunpeng, et al.
Publicado: (2025)
Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA
por: Ou, Justice, et al.
Publicado: (2025)
por: Ou, Justice, et al.
Publicado: (2025)
DO-RAG: A Domain-Specific QA Framework Using Knowledge Graph-Enhanced Retrieval-Augmented Generation
por: Opoku, David Osei, et al.
Publicado: (2025)
por: Opoku, David Osei, et al.
Publicado: (2025)
MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
por: Kim, Seonok
Publicado: (2025)
por: Kim, Seonok
Publicado: (2025)
MuTSE: A Human-in-the-Loop Multi-use Text Simplification Evaluator
por: Roscan, Rares-Alexandru, et al.
Publicado: (2026)
por: Roscan, Rares-Alexandru, et al.
Publicado: (2026)
ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation
por: Gao, Jing, et al.
Publicado: (2025)
por: Gao, Jing, et al.
Publicado: (2025)
PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs
por: Huang, Tianyi, et al.
Publicado: (2026)
por: Huang, Tianyi, et al.
Publicado: (2026)
Evaluation of Retrieval-Augmented Generation: A Survey
por: Yu, Hao, et al.
Publicado: (2024)
por: Yu, Hao, et al.
Publicado: (2024)
HALO: Hallucination Analysis and Learning Optimization to Empower LLMs with Retrieval-Augmented Context for Guided Clinical Decision Making
por: Anjum, Sumera, et al.
Publicado: (2024)
por: Anjum, Sumera, et al.
Publicado: (2024)
NOVI : Chatbot System for University Novice with BERT and LLMs
por: Nam, Yoonji, et al.
Publicado: (2024)
por: Nam, Yoonji, et al.
Publicado: (2024)
Toward Faithful Retrieval-Augmented Generation with Sparse Autoencoders
por: Xiong, Guangzhi, et al.
Publicado: (2025)
por: Xiong, Guangzhi, et al.
Publicado: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
por: Wei, Jianhui, et al.
Publicado: (2025)
por: Wei, Jianhui, et al.
Publicado: (2025)
Ejemplares similares
-
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data
por: Afzal, Anum, et al.
Publicado: (2024) -
On the Influence of Context Size and Model Choice in Retrieval-Augmented Generation Systems
por: Vladika, Juraj, et al.
Publicado: (2025) -
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
por: Afzal, Anum, et al.
Publicado: (2026) -
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
por: Afzal, Anum, et al.
Publicado: (2025) -
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
por: Afzal, Anum, et al.
Publicado: (2025)