Curse of Knowledge: When Complex Evaluation Context Benefits yet Biases LLM Judges
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Weiyuan, Wang, Xintao, Yuan, Siyu, Xu, Rui, Chen, Jiangjie, Dong, Qingqing, Xiao, Yanghua, Yang, Deqing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base
di: Yuan, Siyu, et al.
Pubblicazione: (2023)
di: Yuan, Siyu, et al.
Pubblicazione: (2023)
ARIA: Training Language Agents with Intention-Driven Reward Aggregation
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
di: Yang, Ruihan, et al.
Pubblicazione: (2025)
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
di: Yuan, Xinfeng, et al.
Pubblicazione: (2024)
di: Yuan, Xinfeng, et al.
Pubblicazione: (2024)
SelfGoal: Your Language Agents Already Know How to Achieve High-level Goals
di: Yang, Ruihan, et al.
Pubblicazione: (2024)
di: Yang, Ruihan, et al.
Pubblicazione: (2024)
Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data
di: Ran, Yiting, et al.
Pubblicazione: (2024)
di: Ran, Yiting, et al.
Pubblicazione: (2024)
CoSER: A Comprehensive Literary Dataset and Framework for Training and Evaluating LLM Role-Playing and Persona Simulation
di: Wang, Xintao, et al.
Pubblicazione: (2025)
di: Wang, Xintao, et al.
Pubblicazione: (2025)
Past Meets Present: Creating Historical Analogy with Large Language Models
di: Li, Nianqi, et al.
Pubblicazione: (2024)
di: Li, Nianqi, et al.
Pubblicazione: (2024)
EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
di: Yuan, Siyu, et al.
Pubblicazione: (2024)
di: Yuan, Siyu, et al.
Pubblicazione: (2024)
Enhancing Language Agent Strategic Reasoning through Self-Play in Adversarial Games
di: Zhang, Yikai, et al.
Pubblicazione: (2025)
di: Zhang, Yikai, et al.
Pubblicazione: (2025)
TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological Interviews
di: Wang, Xintao, et al.
Pubblicazione: (2023)
di: Wang, Xintao, et al.
Pubblicazione: (2023)
SurveyAgent: A Conversational System for Personalized and Efficient Research Survey
di: Wang, Xintao, et al.
Pubblicazione: (2024)
di: Wang, Xintao, et al.
Pubblicazione: (2024)
BookWorld: From Novels to Interactive Agent Societies for Creative Story Generation
di: Ran, Yiting, et al.
Pubblicazione: (2025)
di: Ran, Yiting, et al.
Pubblicazione: (2025)
What Makes an Ideal Quote? Recommending "Unexpected yet Rational" Quotations via Novelty
di: Zhang, Bowei, et al.
Pubblicazione: (2025)
di: Zhang, Bowei, et al.
Pubblicazione: (2025)
Can LLMs Learn to Map the World from Local Descriptions?
di: Xia, Sirui, et al.
Pubblicazione: (2025)
di: Xia, Sirui, et al.
Pubblicazione: (2025)
HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patterns
di: Wang, Xintao, et al.
Pubblicazione: (2026)
di: Wang, Xintao, et al.
Pubblicazione: (2026)
Light Up the Shadows: Enhance Long-Tailed Entity Grounding with Concept-Guided Vision-Language Models
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
di: Zhang, Yikai, et al.
Pubblicazione: (2024)
DEEPER Insight into Your User: Directed Persona Refinement for Dynamic Persona Modeling
di: Chen, Aili, et al.
Pubblicazione: (2025)
di: Chen, Aili, et al.
Pubblicazione: (2025)
Chain-of-Knowledge: Integrating Knowledge Reasoning into Large Language Models by Learning from Knowledge Graphs
di: Zhang, Yifei, et al.
Pubblicazione: (2024)
di: Zhang, Yifei, et al.
Pubblicazione: (2024)
Put Your Money Where Your Mouth Is: Evaluating Strategic Planning and Execution of LLM Agents in an Auction Arena
di: Chen, Jiangjie, et al.
Pubblicazione: (2023)
di: Chen, Jiangjie, et al.
Pubblicazione: (2023)
Revealing the Barriers of Language Agents in Planning
di: Xie, Jian, et al.
Pubblicazione: (2024)
di: Xie, Jian, et al.
Pubblicazione: (2024)
Character is Destiny: Can Role-Playing Language Agents Make Persona-Driven Decisions?
di: Xu, Rui, et al.
Pubblicazione: (2024)
di: Xu, Rui, et al.
Pubblicazione: (2024)
From AI Assistant to AI Scientist: Autonomous Discovery of LLM-RL Algorithms with LLM Agents
di: Xia, Sirui, et al.
Pubblicazione: (2026)
di: Xia, Sirui, et al.
Pubblicazione: (2026)
"A good pun is its own reword": Can Large Language Models Understand Puns?
di: Xu, Zhijun, et al.
Pubblicazione: (2024)
di: Xu, Zhijun, et al.
Pubblicazione: (2024)
Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
di: Moon, Jiwon, et al.
Pubblicazione: (2025)
PowerAttention: Exponentially Scaling of Receptive Fields for Effective Sparse Attention
di: Chen, Lida, et al.
Pubblicazione: (2025)
di: Chen, Lida, et al.
Pubblicazione: (2025)
GumbelSoft: Diversified Language Model Watermarking via the GumbelMax-trick
di: Fu, Jiayi, et al.
Pubblicazione: (2024)
di: Fu, Jiayi, et al.
Pubblicazione: (2024)
Improving Recall of Large Language Models: A Model Collaboration Approach for Relational Triple Extraction
di: Ding, Zepeng, et al.
Pubblicazione: (2024)
di: Ding, Zepeng, et al.
Pubblicazione: (2024)
TravelAgent: An AI Assistant for Personalized Travel Planning
di: Chen, Aili, et al.
Pubblicazione: (2024)
di: Chen, Aili, et al.
Pubblicazione: (2024)
Reasoning Model Is Superior LLM-Judge, Yet Suffers from Biases
di: Huang, Hui, et al.
Pubblicazione: (2026)
di: Huang, Hui, et al.
Pubblicazione: (2026)
Negation Triplet Extraction with Syntactic Dependency and Semantic Consistency
di: Shi, Yuchen, et al.
Pubblicazione: (2024)
di: Shi, Yuchen, et al.
Pubblicazione: (2024)
How Easily do Irrelevant Inputs Skew the Responses of Large Language Models?
di: Wu, Siye, et al.
Pubblicazione: (2024)
di: Wu, Siye, et al.
Pubblicazione: (2024)
Implicit Reasoning in Transformers is Reasoning through Shortcuts
di: Lin, Tianhe, et al.
Pubblicazione: (2025)
di: Lin, Tianhe, et al.
Pubblicazione: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
di: Ye, Jiayi, et al.
Pubblicazione: (2024)
Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
di: Lee, Dongryeol, et al.
Pubblicazione: (2026)
Exploiting Duality in Open Information Extraction with Predicate Prompt
di: Chen, Zhen, et al.
Pubblicazione: (2024)
di: Chen, Zhen, et al.
Pubblicazione: (2024)
FairJudge: An Adaptive, Debiased, and Consistent LLM-as-a-Judge
di: Yang, Bo, et al.
Pubblicazione: (2026)
di: Yang, Bo, et al.
Pubblicazione: (2026)
Boosting Scientific Concepts Understanding: Can Analogy from Teacher Models Empower Student Models?
di: Yuan, Siyu, et al.
Pubblicazione: (2024)
di: Yuan, Siyu, et al.
Pubblicazione: (2024)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
di: Wang, Bing, et al.
Pubblicazione: (2026)
di: Wang, Bing, et al.
Pubblicazione: (2026)
Documenti analoghi
-
ANALOGYKB: Unlocking Analogical Reasoning of Language Models with A Million-scale Knowledge Base
di: Yuan, Siyu, et al.
Pubblicazione: (2023) -
ARIA: Training Language Agents with Intention-Driven Reward Aggregation
di: Yang, Ruihan, et al.
Pubblicazione: (2025) -
Evaluating Character Understanding of Large Language Models via Character Profiling from Fictional Works
di: Yuan, Xinfeng, et al.
Pubblicazione: (2024) -
SelfGoal: Your Language Agents Already Know How to Achieve High-level Goals
di: Yang, Ruihan, et al.
Pubblicazione: (2024) -
Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data
di: Ran, Yiting, et al.
Pubblicazione: (2024)