S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Shaoning, Yu, Jiachen, Wang, Zongqi, Yang, Xuewei, Gu, Tianle, Yang, Yujiu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improve LLM-as-a-Judge Ability as a General Ability
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
von: Yu, Jiachen, et al.
Veröffentlicht: (2025)
Reward Modeling from Natural Language Human Feedback
von: Wang, Zongqi, et al.
Veröffentlicht: (2026)
von: Wang, Zongqi, et al.
Veröffentlicht: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
SCAN: Structured Capability Assessment and Navigation for LLMs
von: Wang, Zongqi, et al.
Veröffentlicht: (2025)
von: Wang, Zongqi, et al.
Veröffentlicht: (2025)
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Robust and Minimally Invasive Watermarking for EaaS
von: Wang, Zongqi, et al.
Veröffentlicht: (2024)
von: Wang, Zongqi, et al.
Veröffentlicht: (2024)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
von: Liu, Tao, et al.
Veröffentlicht: (2026)
von: Liu, Tao, et al.
Veröffentlicht: (2026)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
von: Gu, Tianle, et al.
Veröffentlicht: (2025)
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
von: Yu, Jiachen, et al.
Veröffentlicht: (2026)
von: Yu, Jiachen, et al.
Veröffentlicht: (2026)
MorphMark: Flexible Adaptive Watermarking for Large Language Models
von: Wang, Zongqi, et al.
Veröffentlicht: (2025)
von: Wang, Zongqi, et al.
Veröffentlicht: (2025)
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
von: Jia, Ruipeng, et al.
Veröffentlicht: (2025)
von: Jia, Ruipeng, et al.
Veröffentlicht: (2025)
From Sparse Decisions to Dense Reasoning: A Multi-attribute Trajectory Paradigm for Multimodal Moderation
von: Gu, Tianle, et al.
Veröffentlicht: (2026)
von: Gu, Tianle, et al.
Veröffentlicht: (2026)
Distributional Clarity: The Hidden Driver of RL-Friendliness in Large Language Models
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
von: Sun, Shaoning, et al.
Veröffentlicht: (2026)
Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
von: Liu, Jinliang, et al.
Veröffentlicht: (2025)
von: Liu, Jinliang, et al.
Veröffentlicht: (2025)
Adaptive Segment-level Reward: Bridging the Gap Between Action and Reward Space in Alignment
von: Li, Yanshi, et al.
Veröffentlicht: (2024)
von: Li, Yanshi, et al.
Veröffentlicht: (2024)
Speculative Reward Model Boosts Decision Making Ability of LLMs Cost-Effectively
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
von: Xiao, Zeguan, et al.
Veröffentlicht: (2025)
MEOW: MEMOry Supervised LLM Unlearning Via Inverted Facts
von: Gu, Tianle, et al.
Veröffentlicht: (2024)
von: Gu, Tianle, et al.
Veröffentlicht: (2024)
P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling
von: Zhang, Pinyi, et al.
Veröffentlicht: (2026)
von: Zhang, Pinyi, et al.
Veröffentlicht: (2026)
Assessing the Ability of Neural TTS Systems to Model Consonant-Induced F0 Perturbation
von: Yang, Tianle, et al.
Veröffentlicht: (2026)
von: Yang, Tianle, et al.
Veröffentlicht: (2026)
Solving Math Word Problems via Cooperative Reasoning induced Language Models
von: Zhu, Xinyu, et al.
Veröffentlicht: (2022)
von: Zhu, Xinyu, et al.
Veröffentlicht: (2022)
Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Chain of History: Learning and Forecasting with LLMs for Temporal Knowledge Graph Completion
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
von: Luo, Ruilin, et al.
Veröffentlicht: (2024)
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
von: Zhou, Xin, et al.
Veröffentlicht: (2025)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
von: Ko, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Ko, Hyunwoo, et al.
Veröffentlicht: (2025)
A Novel Self-Evolution Framework for Large Language Models
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
von: Sun, Haoran, et al.
Veröffentlicht: (2025)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
von: Shi, Chufan, et al.
Veröffentlicht: (2024)
ReviewAgents: Bridging the Gap Between Human and AI-Generated Paper Reviews
von: Gao, Xian, et al.
Veröffentlicht: (2025)
von: Gao, Xian, et al.
Veröffentlicht: (2025)
ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2026)
IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
von: Wen, Bosi, et al.
Veröffentlicht: (2026)
AdaJudge: Adaptive Multi-Perspective Judging for Reward Modeling
von: Miao, Yongliang, et al.
Veröffentlicht: (2026)
von: Miao, Yongliang, et al.
Veröffentlicht: (2026)
R-Align: Enhancing Generative Reward Models through Rationale-Centric Meta-Judging
von: Lai, Yanlin, et al.
Veröffentlicht: (2026)
von: Lai, Yanlin, et al.
Veröffentlicht: (2026)
From Sub-Ability Diagnosis to Human-Aligned Generation: Bridging the Gap for Text Length Control via MARKERGEN
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
von: Yuan, Peiwen, et al.
Veröffentlicht: (2025)
Distilling Bayesian Belief States into Language Models for Auditable Negotiation
von: Cui, Zongqi, et al.
Veröffentlicht: (2026)
von: Cui, Zongqi, et al.
Veröffentlicht: (2026)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
GRAM: A Generative Foundation Reward Model for Reward Generalization
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
von: Wang, Chenglong, et al.
Veröffentlicht: (2025)
Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning
von: Yang, Zhaorui, et al.
Veröffentlicht: (2024)
von: Yang, Zhaorui, et al.
Veröffentlicht: (2024)
SongSage: A Large Musical Language Model with Lyric Generative Pre-training
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
von: Guo, Jiani, et al.
Veröffentlicht: (2026)
Bridging the Semantic Gap: Contrastive Rewards for Multilingual Text-to-SQL with GRPO
von: Kattamuri, Ashish, et al.
Veröffentlicht: (2025)
von: Kattamuri, Ashish, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Improve LLM-as-a-Judge Ability as a General Ability
von: Yu, Jiachen, et al.
Veröffentlicht: (2025) -
Reward Modeling from Natural Language Human Feedback
von: Wang, Zongqi, et al.
Veröffentlicht: (2026) -
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026) -
SCAN: Structured Capability Assessment and Navigation for LLMs
von: Wang, Zongqi, et al.
Veröffentlicht: (2025) -
Invisible Entropy: Towards Safe and Efficient Low-Entropy LLM Watermarking
von: Gu, Tianle, et al.
Veröffentlicht: (2025)