AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Miao, Yongliang, Liang, Yangyang, Du, Mengnan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912819998359552
author Miao, Yongliang
Liang, Yangyang
Du, Mengnan
author_facet Miao, Yongliang
Liang, Yangyang
Du, Mengnan
contents Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strategy to condense sequences into scalar scores. This paradigm, however, suffers from two key limitations: a static inductive bias that misaligns with task-dependent preference signals, and a representational mismatch, as the backbone is optimized for generation rather than fine-grained discrimination. To address this, we propose AdaJudge, a unified framework that jointly adapts representation and aggregation. AdaJudge first refines backbone representations into a discrimination-oriented space via gated refinement blocks. It then replaces the static readout with an adaptive multi-view pooling module that dynamically routes and combines evidence. Extensive experiments on RM-Bench and JudgeBench show that AdaJudge outperforms strong off-the-shelf reward models and traditional pooling baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2601_08097
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
Miao, Yongliang
Liang, Yangyang
Du, Mengnan
Computation and Language
Machine Learning
Reward modeling is essential for aligning large language models with human preferences, yet predominant architectures rely on a static pooling strategy to condense sequences into scalar scores. This paradigm, however, suffers from two key limitations: a static inductive bias that misaligns with task-dependent preference signals, and a representational mismatch, as the backbone is optimized for generation rather than fine-grained discrimination. To address this, we propose AdaJudge, a unified framework that jointly adapts representation and aggregation. AdaJudge first refines backbone representations into a discrimination-oriented space via gated refinement blocks. It then replaces the static readout with an adaptive multi-view pooling module that dynamically routes and combines evidence. Extensive experiments on RM-Bench and JudgeBench show that AdaJudge outperforms strong off-the-shelf reward models and traditional pooling baselines.
title AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2601.08097