MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909986629615616 |
|---|---|
| author | Ji, Yuelyu Kwak, Min Gu Zhang, Hang Wu, Xizhi Li, Chenyu Wang, Yanshan |
| author_facet | Ji, Yuelyu Kwak, Min Gu Zhang, Hang Wu, Xizhi Li, Chenyu Wang, Yanshan |
| contents | Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications.
We introduce MedRAGChecker, a claim-level verification and diagnostic framework for biomedical RAG.
Given a question, retrieved evidence, and a generated answer, MedRAGChecker decomposes the answer into atomic claims and estimates claim support by combining evidence-grounded natural language inference (NLI) with biomedical knowledge-graph (KG) consistency signals.
Aggregating claim decisions yields answer-level diagnostics that help disentangle retrieval and generation failures, including faithfulness, under-evidence, contradiction, and safety-critical error rates.
To enable scalable evaluation, we distill the pipeline into compact biomedical models and use an ensemble verifier with class-specific reliability weighting.
Experiments on four biomedical QA benchmarks show that MedRAGChecker reliably flags unsupported and contradicted claims and reveals distinct risk profiles across generators, particularly on safety-critical biomedical relations. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2601_06519 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation Ji, Yuelyu Kwak, Min Gu Zhang, Hang Wu, Xizhi Li, Chenyu Wang, Yanshan Computation and Language Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications. We introduce MedRAGChecker, a claim-level verification and diagnostic framework for biomedical RAG. Given a question, retrieved evidence, and a generated answer, MedRAGChecker decomposes the answer into atomic claims and estimates claim support by combining evidence-grounded natural language inference (NLI) with biomedical knowledge-graph (KG) consistency signals. Aggregating claim decisions yields answer-level diagnostics that help disentangle retrieval and generation failures, including faithfulness, under-evidence, contradiction, and safety-critical error rates. To enable scalable evaluation, we distill the pipeline into compact biomedical models and use an ensemble verifier with class-specific reliability weighting. Experiments on four biomedical QA benchmarks show that MedRAGChecker reliably flags unsupported and contradicted claims and reveals distinct risk profiles across generators, particularly on safety-critical biomedical relations. |
| title | MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation |
| topic | Computation and Language |
| url | https://arxiv.org/abs/2601.06519 |