JudgeAgent: Beyond Static Benchmarks for Knowledge-Driven and Dynamic LLM Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shi, Zhichao, Jiang, Xuhui, Xu, Chengjin, Yao, Cangli, Ma, Shengjia, Shen, Yinghan, Li, Zixuan, Guo, Jian, Wang, Yuanzhuo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Practical Entity Alignment Method Design: Insights from New Highly Heterogeneous Knowledge Graph Datasets
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
Unlocking the Power of Large Language Models for Entity Alignment
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
von: Gu, Jiawei, et al.
Veröffentlicht: (2024)
On the Evolution of Knowledge Graphs: A Survey and Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023)
A Survey on Large Language Model Hallucination via a Creativity Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)
DataArc-SynData-Toolkit: A Unified Closed-Loop Framework for Multi-Path, Multimodal, and Multilingual Data Synthesis
von: Shi, Zhichao, et al.
Veröffentlicht: (2026)
von: Shi, Zhichao, et al.
Veröffentlicht: (2026)
Synthesize-on-Graph: Knowledgeable Synthetic Data Generation for Continue Pre-training of Large Language Models
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
von: Ma, Shengjie, et al.
Veröffentlicht: (2025)
The Evaluation Game: Beyond Static LLM Benchmarking
von: Wang, Paul, et al.
Veröffentlicht: (2026)
von: Wang, Paul, et al.
Veröffentlicht: (2026)
Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
von: Wu, Xiaojun, et al.
Veröffentlicht: (2025)
von: Wu, Xiaojun, et al.
Veröffentlicht: (2025)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
von: Lee, Dongryeol, et al.
Veröffentlicht: (2026)
von: Lee, Dongryeol, et al.
Veröffentlicht: (2026)
SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
von: Ma, Peixian, et al.
Veröffentlicht: (2025)
von: Ma, Peixian, et al.
Veröffentlicht: (2025)
Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation
von: Ma, Shengjie, et al.
Veröffentlicht: (2024)
von: Ma, Shengjie, et al.
Veröffentlicht: (2024)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Retrieval, Reasoning, Re-ranking: A Context-Enriched Framework for Knowledge Graph Completion
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
Context-aware Inductive Knowledge Graph Completion with Latent Type Constraints and Subgraph Reasoning
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
von: Li, Muzhi, et al.
Veröffentlicht: (2024)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
von: Yoa, Seungdong, et al.
Veröffentlicht: (2026)
von: Yoa, Seungdong, et al.
Veröffentlicht: (2026)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Context Graph
von: Xu, Chengjin, et al.
Veröffentlicht: (2024)
von: Xu, Chengjin, et al.
Veröffentlicht: (2024)
RuleRAG: Rule-Guided Retrieval-Augmented Generation with Language Models for Question Answering
von: Chen, Zhongwu, et al.
Veröffentlicht: (2024)
von: Chen, Zhongwu, et al.
Veröffentlicht: (2024)
Continual Pretraining on Encrypted Synthetic Data for Privacy-Preserving LLMs
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2024)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2024)
Beyond the Illusion of Consensus: From Surface Heuristics to Knowledge-Grounded Evaluation in LLM-as-a-Judge
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Beyond Static Summarization: Proactive Memory Extraction for LLM Agents
von: Yang, Chengyuan, et al.
Veröffentlicht: (2026)
von: Yang, Chengyuan, et al.
Veröffentlicht: (2026)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
von: Yin, Sheng, et al.
Veröffentlicht: (2024)
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
von: Jiang, Hongchao, et al.
Veröffentlicht: (2025)
Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024)
von: Szymanski, Annalisa, et al.
Veröffentlicht: (2024)
AJ-Bench: Benchmarking Agent-as-a-Judge for Environment-Aware Evaluation
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
von: Shi, Wentao, et al.
Veröffentlicht: (2026)
R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
von: Yuan, Tongxin, et al.
Veröffentlicht: (2024)
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)
von: Chen, Jiaju, et al.
Veröffentlicht: (2025)
Select2Reason: Efficient Instruction-Tuning Data Selection for Long-CoT Reasoning
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
von: Yang, Cehao, et al.
Veröffentlicht: (2025)
Conflicts Make Large Reasoning Models Vulnerable to Attacks
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
von: Liu, Honghao, et al.
Veröffentlicht: (2026)
Beyond Consensus: Mitigating the Agreeableness Bias in LLM Judge Evaluations
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025)
von: Jain, Suryaansh, et al.
Veröffentlicht: (2025)
LLM4DESIGN: An Automated Multi-Modal System for Architectural and Environmental Design
von: Chen, Ran, et al.
Veröffentlicht: (2024)
von: Chen, Ran, et al.
Veröffentlicht: (2024)
Automatically Benchmarking LLM Code Agents through Agent-Driven Annotation and Evaluation
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
von: Fu, Lingyue, et al.
Veröffentlicht: (2025)
Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
von: Shen, Yiyang, et al.
Veröffentlicht: (2026)
Disentanglement Beyond Static vs. Dynamic: A Benchmark and Evaluation Framework for Multi-Factor Sequential Representations
von: Barami, Tal, et al.
Veröffentlicht: (2025)
von: Barami, Tal, et al.
Veröffentlicht: (2025)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
von: Chen, Junjie, et al.
Veröffentlicht: (2026)
Agent-as-a-Judge: Evaluate Agents with Agents
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
von: Zhuge, Mingchen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Toward Practical Entity Alignment Method Design: Insights from New Highly Heterogeneous Knowledge Graph Datasets
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023) -
Unlocking the Power of Large Language Models for Entity Alignment
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024) -
A Survey on LLM-as-a-Judge
von: Gu, Jiawei, et al.
Veröffentlicht: (2024) -
On the Evolution of Knowledge Graphs: A Survey and Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2023) -
A Survey on Large Language Model Hallucination via a Creativity Perspective
von: Jiang, Xuhui, et al.
Veröffentlicht: (2024)