Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Liu, Yixin, Yu, Yue, Su, DiJia, Wang, Sid, Wang, Xuewei, Jiang, Song, Liu, Bo, Cohan, Arman, Tian, Yuandong, Chen, Zhengxing
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914388125941760
author Liu, Yixin
Yu, Yue
Su, DiJia
Wang, Sid
Wang, Xuewei
Jiang, Song
Liu, Bo
Cohan, Arman
Tian, Yuandong
Chen, Zhengxing
author_facet Liu, Yixin
Yu, Yue
Su, DiJia
Wang, Sid
Wang, Xuewei
Jiang, Song
Liu, Bo
Cohan, Arman
Tian, Yuandong
Chen, Zhengxing
contents Reasoning LLMs-as-Judges, which can benefit from inference-time scaling, provide a promising path for extending the success of reasoning models to non-verifiable domains where the output correctness/quality cannot be directly checked. However, while reasoning judges have shown better performance on static evaluation benchmarks, their effectiveness in actual policy training has not been systematically examined. Therefore, we conduct a rigorous study to investigate the actual impact of non-reasoning and reasoning judges in reinforcement-learning-based LLM alignment. Our controlled synthetic setting, where a "gold-standard" judge (gpt-oss-120b) provides preference annotations to train smaller judges, reveals key differences between non-reasoning and reasoning judges: non-reasoning judges lead to reward hacking easily, while reasoning judges can lead to policies that achieve strong performance when evaluated by the gold-standard judge. Interestingly, we find that the reasoning-judge-trained policies achieve such strong performance by learning to generate highly effective adversarial outputs that can also score well on popular benchmarks such as Arena-Hard by deceiving other LLM-judges. Combined with our further analysis, our study highlights both important findings and room for improvements for applying (reasoning) LLM-judges in non-verifiable LLM post-training.
format Preprint
id arxiv_https___arxiv_org_abs_2603_12246
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
Liu, Yixin
Yu, Yue
Su, DiJia
Wang, Sid
Wang, Xuewei
Jiang, Song
Liu, Bo
Cohan, Arman
Tian, Yuandong
Chen, Zhengxing
Artificial Intelligence
Computation and Language
Machine Learning
Reasoning LLMs-as-Judges, which can benefit from inference-time scaling, provide a promising path for extending the success of reasoning models to non-verifiable domains where the output correctness/quality cannot be directly checked. However, while reasoning judges have shown better performance on static evaluation benchmarks, their effectiveness in actual policy training has not been systematically examined. Therefore, we conduct a rigorous study to investigate the actual impact of non-reasoning and reasoning judges in reinforcement-learning-based LLM alignment. Our controlled synthetic setting, where a "gold-standard" judge (gpt-oss-120b) provides preference annotations to train smaller judges, reveals key differences between non-reasoning and reasoning judges: non-reasoning judges lead to reward hacking easily, while reasoning judges can lead to policies that achieve strong performance when evaluated by the gold-standard judge. Interestingly, we find that the reasoning-judge-trained policies achieve such strong performance by learning to generate highly effective adversarial outputs that can also score well on popular benchmarks such as Arena-Hard by deceiving other LLM-judges. Combined with our further analysis, our study highlights both important findings and room for improvements for applying (reasoning) LLM-judges in non-verifiable LLM post-training.
title Examining Reasoning LLMs-as-Judges in Non-Verifiable LLM Post-Training
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.12246