Contradiction Detection in RAG Systems: Evaluating LLMs as Context Validators for Improved Information Consistency
Fuente:
arXiv
Saved in:
| Main Authors: | Gokul, Vignesh, Tenneti, Srikanth, Nakkiran, Alwarappan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction
by: Maheshwari, Harsh, et al.
Published: (2025)
by: Maheshwari, Harsh, et al.
Published: (2025)
Beyond Factual Grounding: The Case for Opinion-Aware Retrieval-Augmented Generation
by: Agrawal, Aditya, et al.
Published: (2026)
by: Agrawal, Aditya, et al.
Published: (2026)
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025)
by: Kumar, Adarsh, et al.
Published: (2025)
Confidence Improves Self-Consistency in LLMs
by: Taubenfeld, Amir, et al.
Published: (2025)
by: Taubenfeld, Amir, et al.
Published: (2025)
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
by: Parihar, Shweta, et al.
Published: (2026)
by: Parihar, Shweta, et al.
Published: (2026)
HealthContradict: Evaluating Biomedical Knowledge Conflicts in Language Models
by: Zhang, Boya, et al.
Published: (2025)
by: Zhang, Boya, et al.
Published: (2025)
Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
by: Jin, Bowen, et al.
Published: (2024)
by: Jin, Bowen, et al.
Published: (2024)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs
by: Li, Yangning, et al.
Published: (2025)
by: Li, Yangning, et al.
Published: (2025)
Face4RAG: Factual Consistency Evaluation for Retrieval Augmented Generation in Chinese
by: Xu, Yunqi, et al.
Published: (2024)
by: Xu, Yunqi, et al.
Published: (2024)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
by: Hou, Yufang, et al.
Published: (2024)
by: Hou, Yufang, et al.
Published: (2024)
Evaluating Role-Consistency in LLMs for Counselor Training
by: Rudolph, Eric, et al.
Published: (2026)
by: Rudolph, Eric, et al.
Published: (2026)
Evidence-backed Fact Checking using RAG and Few-Shot In-Context Learning with LLMs
by: Singhal, Ronit, et al.
Published: (2024)
by: Singhal, Ronit, et al.
Published: (2024)
Evaluating the Sensitivity of LLMs to Prior Context
by: Hankache, Robert, et al.
Published: (2025)
by: Hankache, Robert, et al.
Published: (2025)
FB-RAG: Improving RAG with Forward and Backward Lookup
by: Chawla, Kushal, et al.
Published: (2025)
by: Chawla, Kushal, et al.
Published: (2025)
Assessing the Quality of AI-Generated Clinical Notes: A Validated Evaluation of a Large Language Model Scribe
by: Palm, Erin, et al.
Published: (2025)
by: Palm, Erin, et al.
Published: (2025)
Knowledge-Graph Based RAG System Evaluation Framework
by: Dong, Sicheng, et al.
Published: (2025)
by: Dong, Sicheng, et al.
Published: (2025)
AXCEL: Automated eXplainable Consistency Evaluation using LLMs
by: Sreekar, P Aditya, et al.
Published: (2024)
by: Sreekar, P Aditya, et al.
Published: (2024)
LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs -- No Silver Bullet for LC or RAG Routing
by: Li, Kuan, et al.
Published: (2025)
by: Li, Kuan, et al.
Published: (2025)
Should We Fine-Tune or RAG? Evaluating Different Techniques to Adapt LLMs for Dialogue
by: Alghisi, Simone, et al.
Published: (2024)
by: Alghisi, Simone, et al.
Published: (2024)
Enhancing RAG Efficiency with Adaptive Context Compression
by: Guo, Shuyu, et al.
Published: (2025)
by: Guo, Shuyu, et al.
Published: (2025)
RAG Makes Guardrails Unsafe? Investigating Robustness of Guardrails under RAG-style Contexts
by: She, Yining, et al.
Published: (2025)
by: She, Yining, et al.
Published: (2025)
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
Systematic Evaluation of Long-Context LLMs on Financial Concepts
by: Gupta, Lavanya, et al.
Published: (2024)
by: Gupta, Lavanya, et al.
Published: (2024)
SteLLA: A Structured Grading System Using LLMs with RAG
by: Qiu, Hefei, et al.
Published: (2025)
by: Qiu, Hefei, et al.
Published: (2025)
CTRL-RAG: Contrastive Likelihood Reward Based Reinforcement Learning for Context-Faithful RAG Models
by: Tan, Zhehao, et al.
Published: (2026)
by: Tan, Zhehao, et al.
Published: (2026)
SFR-RAG: Towards Contextually Faithful LLMs
by: Nguyen, Xuan-Phi, et al.
Published: (2024)
by: Nguyen, Xuan-Phi, et al.
Published: (2024)
Are LLMs Enough for Hyperpartisan, Fake, Polarized and Harmful Content Detection? Evaluating In-Context Learning vs. Fine-Tuning
by: Maggini, Michele Joshua, et al.
Published: (2025)
by: Maggini, Michele Joshua, et al.
Published: (2025)
Improving Labeling Consistency with Detailed Constitutional Definitions and AI-Driven Evaluation
by: Berlin, Konstantin, et al.
Published: (2026)
by: Berlin, Konstantin, et al.
Published: (2026)
From Amateur to Master: Infusing Knowledge into LLMs via Automated Curriculum Learning
by: Neema, Nishit, et al.
Published: (2025)
by: Neema, Nishit, et al.
Published: (2025)
Are LLMs Truly Multilingual? Exploring Zero-Shot Multilingual Capability of LLMs for Information Retrieval: An Italian Healthcare Use Case
by: Kembu, Vignesh Kumar, et al.
Published: (2025)
by: Kembu, Vignesh Kumar, et al.
Published: (2025)
On the Overscaling Curse of Parallel Thinking: System Efficacy Contradicts Sample Efficiency
by: Wang, Yiming, et al.
Published: (2026)
by: Wang, Yiming, et al.
Published: (2026)
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
by: Lissak, Shir, et al.
Published: (2024)
by: Lissak, Shir, et al.
Published: (2024)
Probabilistic distances-based hallucination detection in LLMs with RAG
by: Oblovatny, Rodion, et al.
Published: (2025)
by: Oblovatny, Rodion, et al.
Published: (2025)
ECoRAG: Evidentiality-guided Compression for Long Context RAG
by: Jeong, Yeonseok, et al.
Published: (2025)
by: Jeong, Yeonseok, et al.
Published: (2025)
Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systems
by: Chhabra, Mukul, et al.
Published: (2026)
by: Chhabra, Mukul, et al.
Published: (2026)
StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization
by: Li, Zhuoqun, et al.
Published: (2024)
by: Li, Zhuoqun, et al.
Published: (2024)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
by: Ma, Huanhuan, et al.
Published: (2025)
by: Ma, Huanhuan, et al.
Published: (2025)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
by: Cui, Hao, et al.
Published: (2025)
by: Cui, Hao, et al.
Published: (2025)
Supervised Fine-Tuning or In-Context Learning? Evaluating LLMs for Clinical NER
by: Baroian, Andrei
Published: (2025)
by: Baroian, Andrei
Published: (2025)
Similar Items
-
CiteFix: Enhancing RAG Accuracy Through Post-Processing Citation Correction
by: Maheshwari, Harsh, et al.
Published: (2025) -
Beyond Factual Grounding: The Case for Opinion-Aware Retrieval-Augmented Generation
by: Agrawal, Aditya, et al.
Published: (2026) -
Improving the Reliability of LLMs: Combining CoT, RAG, Self-Consistency, and Self-Verification
by: Kumar, Adarsh, et al.
Published: (2025) -
Confidence Improves Self-Consistency in LLMs
by: Taubenfeld, Amir, et al.
Published: (2025) -
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
by: Parihar, Shweta, et al.
Published: (2026)