Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Shuliang, Li, Xinze, Liu, Zhenghao, Yan, Yukun, Yang, Cheng, Zeng, Zheni, Liu, Zhiyuan, Sun, Maosong, Yu, Ge |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Structured Knowledge Representation through Contextual Pages for Retrieval-Augmented Generation
von: Li, Xinze, et al.
Veröffentlicht: (2026)
von: Li, Xinze, et al.
Veröffentlicht: (2026)
RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughts
von: Wu, Mingyan, et al.
Veröffentlicht: (2025)
von: Wu, Mingyan, et al.
Veröffentlicht: (2025)
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
NaviRAG: Towards Active Knowledge Navigation for Retrieval-Augmented Generation
von: Dai, Jihao, et al.
Veröffentlicht: (2026)
von: Dai, Jihao, et al.
Veröffentlicht: (2026)
PersLLM: A Personified Training Approach for Large Language Models
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)
Finding What Matters: Anchoring Context Knowledge with Evolving Indices for Iterative Retrieval
von: Wu, Mingyan, et al.
Veröffentlicht: (2026)
von: Wu, Mingyan, et al.
Veröffentlicht: (2026)
Scientific Knowledge-driven Decoding Constraints Improving the Reliability of LLMs
von: Ma, Maotian, et al.
Veröffentlicht: (2026)
von: Ma, Maotian, et al.
Veröffentlicht: (2026)
KG-Infused RAG: Augmenting Corpus-Based RAG with External Knowledge Graphs
von: Wu, Dingjun, et al.
Veröffentlicht: (2025)
von: Wu, Dingjun, et al.
Veröffentlicht: (2025)
Enhancing Long-Chain Reasoning Distillation through Error-Aware Self-Reflection
von: Wu, Zhuoyang, et al.
Veröffentlicht: (2025)
von: Wu, Zhuoyang, et al.
Veröffentlicht: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
ClueAnchor: Clue-Anchored Knowledge Reasoning Exploration and Optimization for Retrieval-Augmented Generation
von: Chen, Hao, et al.
Veröffentlicht: (2025)
von: Chen, Hao, et al.
Veröffentlicht: (2025)
ThinkNote: Enhancing Knowledge Integration and Utilization of Large Language Models via Constructivist Cognition Modeling
von: Xu, Zhipeng, et al.
Veröffentlicht: (2024)
von: Xu, Zhipeng, et al.
Veröffentlicht: (2024)
KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG
von: Li, Yongjian, et al.
Veröffentlicht: (2025)
von: Li, Yongjian, et al.
Veröffentlicht: (2025)
Building A Coding Assistant via the Retrieval-Augmented Language Model
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
von: Huang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Huang, Pengcheng, et al.
Veröffentlicht: (2025)
MetaMem: Evolving Meta-Memory for Knowledge Utilization through Self-Reflective Symbolic Optimization
von: Xin, Haidong, et al.
Veröffentlicht: (2026)
von: Xin, Haidong, et al.
Veröffentlicht: (2026)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
von: Yang, Bo, et al.
Veröffentlicht: (2026)
von: Yang, Bo, et al.
Veröffentlicht: (2026)
JudgeLRM: Large Reasoning Models as a Judge
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
von: Chen, Nuo, et al.
Veröffentlicht: (2025)
KBAlign: Efficient Self Adaptation on Specific Knowledge Bases
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)
Mitigating Judgment Preference Bias in Large Language Models through Group-Based Polling
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
von: Thakur, Aman Singh, et al.
Veröffentlicht: (2024)
DeepNote: Note-Centric Deep Retrieval-Augmented Generation
von: Wang, Ruobing, et al.
Veröffentlicht: (2024)
von: Wang, Ruobing, et al.
Veröffentlicht: (2024)
Improve LLM-as-a-Judge Ability as a General Ability
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
FedJudge: Federated Legal Large Language Model
von: Yue, Linan, et al.
Veröffentlicht: (2023)
von: Yue, Linan, et al.
Veröffentlicht: (2023)
MR. Judge: Multimodal Reasoner as a Judge
von: Pi, Renjie, et al.
Veröffentlicht: (2025)
von: Pi, Renjie, et al.
Veröffentlicht: (2025)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
von: Liu, Yixin, et al.
Veröffentlicht: (2025)
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
von: Zhao, Yuwei, et al.
Veröffentlicht: (2024)
von: Zhao, Yuwei, et al.
Veröffentlicht: (2024)
JudgeLM: Fine-tuned Large Language Models are Scalable Judges
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2023)
Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QA
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
von: Belmadani, Ikram, et al.
Veröffentlicht: (2026)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
von: Cao, Maosong, et al.
Veröffentlicht: (2024)
RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
von: Tang, Zhenwei, et al.
Veröffentlicht: (2026)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
von: Laskar, Md Tahmid Rahman, et al.
Veröffentlicht: (2025)
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
von: Liu, Zhenghao, et al.
Veröffentlicht: (2026)
von: Liu, Zhenghao, et al.
Veröffentlicht: (2026)
Say More with Less: Understanding Prompt Learning Behaviors through Gist Compression
von: Li, Xinze, et al.
Veröffentlicht: (2024)
von: Li, Xinze, et al.
Veröffentlicht: (2024)
CodeJudge: Evaluating Code Generation with Large Language Models
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
von: Tong, Weixi, et al.
Veröffentlicht: (2024)
CyclicJudge: Mitigating Judge Bias Efficiently in LLM-based Evaluation
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
von: Zhu, Ziyi, et al.
Veröffentlicht: (2026)
ConPET: Continual Parameter-Efficient Tuning for Large Language Models
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
von: Song, Chenyang, et al.
Veröffentlicht: (2023)
Judge Anything: MLLM as a Judge Across Any Modality
von: Pu, Shu, et al.
Veröffentlicht: (2025)
von: Pu, Shu, et al.
Veröffentlicht: (2025)
Agent-as-a-Judge
von: You, Runyang, et al.
Veröffentlicht: (2026)
von: You, Runyang, et al.
Veröffentlicht: (2026)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Structured Knowledge Representation through Contextual Pages for Retrieval-Augmented Generation
von: Li, Xinze, et al.
Veröffentlicht: (2026) -
RankCoT: Refining Knowledge for Retrieval-Augmented Generation through Ranking Chain-of-Thoughts
von: Wu, Mingyan, et al.
Veröffentlicht: (2025) -
RAG-DDR: Optimizing Retrieval-Augmented Generation Using Differentiable Data Rewards
von: Li, Xinze, et al.
Veröffentlicht: (2024) -
NaviRAG: Towards Active Knowledge Navigation for Retrieval-Augmented Generation
von: Dai, Jihao, et al.
Veröffentlicht: (2026) -
PersLLM: A Personified Training Approach for Large Language Models
von: Zeng, Zheni, et al.
Veröffentlicht: (2024)