MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Yuelyu, Kwak, Min Gu, Zhang, Hang, Wu, Xizhi, Li, Chenyu, Wang, Yanshan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909986629615616
author Ji, Yuelyu
Kwak, Min Gu
Zhang, Hang
Wu, Xizhi
Li, Chenyu
Wang, Yanshan
author_facet Ji, Yuelyu
Kwak, Min Gu
Zhang, Hang
Wu, Xizhi
Li, Chenyu
Wang, Yanshan
contents Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications. We introduce MedRAGChecker, a claim-level verification and diagnostic framework for biomedical RAG. Given a question, retrieved evidence, and a generated answer, MedRAGChecker decomposes the answer into atomic claims and estimates claim support by combining evidence-grounded natural language inference (NLI) with biomedical knowledge-graph (KG) consistency signals. Aggregating claim decisions yields answer-level diagnostics that help disentangle retrieval and generation failures, including faithfulness, under-evidence, contradiction, and safety-critical error rates. To enable scalable evaluation, we distill the pipeline into compact biomedical models and use an ensemble verifier with class-specific reliability weighting. Experiments on four biomedical QA benchmarks show that MedRAGChecker reliably flags unsupported and contradicted claims and reveals distinct risk profiles across generators, particularly on safety-critical biomedical relations.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06519
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
Ji, Yuelyu
Kwak, Min Gu
Zhang, Hang
Wu, Xizhi
Li, Chenyu
Wang, Yanshan
Computation and Language
Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications. We introduce MedRAGChecker, a claim-level verification and diagnostic framework for biomedical RAG. Given a question, retrieved evidence, and a generated answer, MedRAGChecker decomposes the answer into atomic claims and estimates claim support by combining evidence-grounded natural language inference (NLI) with biomedical knowledge-graph (KG) consistency signals. Aggregating claim decisions yields answer-level diagnostics that help disentangle retrieval and generation failures, including faithfulness, under-evidence, contradiction, and safety-critical error rates. To enable scalable evaluation, we distill the pipeline into compact biomedical models and use an ensemble verifier with class-specific reliability weighting. Experiments on four biomedical QA benchmarks show that MedRAGChecker reliably flags unsupported and contradicted claims and reveals distinct risk profiles across generators, particularly on safety-critical biomedical relations.
title MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
topic Computation and Language
url https://arxiv.org/abs/2601.06519