Context Over Content: Exposing Evaluation Faking in Automated Judges
Fuente:
arXiv
Saved in:
| Main Authors: | Gupta, Manan, Nair, Inderjeet, Wang, Lu, Kumar, Dhruv |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
by: Gupta, Manan, et al.
Published: (2026)
by: Gupta, Manan, et al.
Published: (2026)
Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models
by: Nair, Inderjeet, et al.
Published: (2026)
by: Nair, Inderjeet, et al.
Published: (2026)
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
by: Nair, Inderjeet, et al.
Published: (2024)
by: Nair, Inderjeet, et al.
Published: (2024)
MIDGARD: Self-Consistency Using Minimum Description Length for Structured Commonsense Reasoning
by: Nair, Inderjeet, et al.
Published: (2024)
by: Nair, Inderjeet, et al.
Published: (2024)
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths
by: Nair, Inderjeet, et al.
Published: (2025)
by: Nair, Inderjeet, et al.
Published: (2025)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
by: Xu, Austin, et al.
Published: (2025)
by: Xu, Austin, et al.
Published: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
On Evaluating LLM Alignment by Evaluating LLMs as Judges
by: Liu, Yixin, et al.
Published: (2025)
by: Liu, Yixin, et al.
Published: (2025)
Large Visual-Language Models Are Also Good Classifiers: A Study of In-Context Multimodal Fake News Detection
by: Jiang, Ye, et al.
Published: (2024)
by: Jiang, Ye, et al.
Published: (2024)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
by: Singh, Joykirat, et al.
Published: (2024)
by: Singh, Joykirat, et al.
Published: (2024)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
ALAS: Autonomous Learning Agent for Self-Updating Language Models
by: Atreja, Dhruv
Published: (2025)
by: Atreja, Dhruv
Published: (2025)
Becoming Experienced Judges: Selective Test-Time Learning for Evaluators
by: Jwa, Seungyeon, et al.
Published: (2025)
by: Jwa, Seungyeon, et al.
Published: (2025)
Towards Compute-Optimal Many-Shot In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2025)
by: Golchin, Shahriar, et al.
Published: (2025)
Cross-Modal Augmentation for Few-Shot Multimodal Fake News Detection
by: Jiang, Ye, et al.
Published: (2024)
by: Jiang, Ye, et al.
Published: (2024)
FaKnow: A Unified Library for Fake News Detection
by: Zhu, Yiyuan, et al.
Published: (2024)
by: Zhu, Yiyuan, et al.
Published: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
by: Hong, Yihan, et al.
Published: (2026)
by: Hong, Yihan, et al.
Published: (2026)
Beyond LLM-as-a-Judge: Deterministic Metrics for Multilingual Generative Text Evaluation
by: Alam, Firoj, et al.
Published: (2026)
by: Alam, Firoj, et al.
Published: (2026)
RLHF: A comprehensive Survey for Cultural, Multimodal and Low Latency Alignment Methods
by: Sharma, Raghav, et al.
Published: (2025)
by: Sharma, Raghav, et al.
Published: (2025)
Revisiting In-Context Learning with Long Context Language Models
by: Baek, Jinheon, et al.
Published: (2024)
by: Baek, Jinheon, et al.
Published: (2024)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
by: Li, Zhuochun, et al.
Published: (2026)
by: Li, Zhuochun, et al.
Published: (2026)
Multi-view autoencoders for Fake News Detection
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
by: Pereira, Ingryd V. S. T., et al.
Published: (2025)
Context Tuning for In-Context Optimization
by: Lu, Jack, et al.
Published: (2025)
by: Lu, Jack, et al.
Published: (2025)
CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation
by: Muhamed, Aashiq
Published: (2025)
by: Muhamed, Aashiq
Published: (2025)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
The Compliance Paradox: Semantic-Instruction Decoupling in Automated Academic Code Evaluation
by: Sahoo, Devanshu, et al.
Published: (2026)
by: Sahoo, Devanshu, et al.
Published: (2026)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
by: Jain, Ojas, et al.
Published: (2026)
by: Jain, Ojas, et al.
Published: (2026)
Fake News Detection After LLM Laundering: Measurement and Explanation
by: Das, Rupak Kumar, et al.
Published: (2025)
by: Das, Rupak Kumar, et al.
Published: (2025)
ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
by: Wang, Zhilin, et al.
Published: (2025)
by: Wang, Zhilin, et al.
Published: (2025)
Balanced Accuracy: The Right Metric for Evaluating LLM Judges -- Explained through Youden's J statistic
by: Collot, Stephane, et al.
Published: (2025)
by: Collot, Stephane, et al.
Published: (2025)
JAF: Judge Agent Forest
by: Garg, Sahil, et al.
Published: (2026)
by: Garg, Sahil, et al.
Published: (2026)
Exposing Limitations of Language Model Agents in Sequential-Task Compositions on the Web
by: Furuta, Hiroki, et al.
Published: (2023)
by: Furuta, Hiroki, et al.
Published: (2023)
The Perfect Blend: Redefining RLHF with Mixture of Judges
by: Xu, Tengyu, et al.
Published: (2024)
by: Xu, Tengyu, et al.
Published: (2024)
Exposing propaganda: an analysis of stylistic cues comparing human annotations and machine classification
by: Faye, Géraud, et al.
Published: (2024)
by: Faye, Géraud, et al.
Published: (2024)
Neighborhood-Order Learning Graph Attention Network for Fake News Detection
by: Lakzaei, Batool, et al.
Published: (2025)
by: Lakzaei, Batool, et al.
Published: (2025)
HSFN: Hierarchical Selection for Fake News Detection building Heterogeneous Ensemble
by: Coutinho, Sara B., et al.
Published: (2025)
by: Coutinho, Sara B., et al.
Published: (2025)
Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
by: Liu, Yixin, et al.
Published: (2026)
by: Liu, Yixin, et al.
Published: (2026)
Dialogue Without Limits: Constant-Sized KV Caches for Extended Responses in LLMs
by: Ghadia, Ravi, et al.
Published: (2025)
by: Ghadia, Ravi, et al.
Published: (2025)
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
by: Nie, Allen, et al.
Published: (2024)
by: Nie, Allen, et al.
Published: (2024)
Similar Items
-
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
by: Gupta, Manan, et al.
Published: (2026) -
Latent Phase-Shift Rollback: Inference-Time Error Correction via Residual Stream Monitoring and KV-Cache Steering
by: Gupta, Manan, et al.
Published: (2026) -
Value-Conflict Diagnostics Reveal Widespread Alignment Faking in Language Models
by: Nair, Inderjeet, et al.
Published: (2026) -
Closing the Loop: Learning to Generate Writing Feedback via Language Model Simulated Student Revisions
by: Nair, Inderjeet, et al.
Published: (2024) -
MIDGARD: Self-Consistency Using Minimum Description Length for Structured Commonsense Reasoning
by: Nair, Inderjeet, et al.
Published: (2024)