CE-RM: A Pointwise Generative Reward Model Optimized via Two-Stage Rollout and Unified Criteria
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Xinyu, He, Yancheng, Wang, Weixun, Feng, Tao, Lin, Li, Liu, Jiashun, Su, Wenbo, Zheng, Bo, Wan, Xiaojun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
by: Liu, Jiashun, et al.
Published: (2025)
by: Liu, Jiashun, et al.
Published: (2025)
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024)
by: Li, Shilong, et al.
Published: (2024)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025)
by: Liu, Zihe, et al.
Published: (2025)
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Think-J: Learning to Think for Generative LLM-as-a-Judge
by: Huang, Hui, et al.
Published: (2025)
by: Huang, Hui, et al.
Published: (2025)
Can Large Language Models Detect Errors in Long Chain-of-Thought Reasoning?
by: He, Yancheng, et al.
Published: (2025)
by: He, Yancheng, et al.
Published: (2025)
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling
by: Jian, Ai, et al.
Published: (2025)
by: Jian, Ai, et al.
Published: (2025)
ProgCo: Program Helps Self-Correction of Large Language Models
by: Song, Xiaoshuai, et al.
Published: (2025)
by: Song, Xiaoshuai, et al.
Published: (2025)
Are LLM-based Evaluators Confusing NLG Quality Criteria?
by: Hu, Xinyu, et al.
Published: (2024)
by: Hu, Xinyu, et al.
Published: (2024)
Complementary Reinforcement Learning
by: Muhtar, Dilxat, et al.
Published: (2026)
by: Muhtar, Dilxat, et al.
Published: (2026)
Deconstructing Long Chain-of-Thought: A Structured Reasoning Optimization Framework for Long CoT Distillation
by: Luo, Yijia, et al.
Published: (2025)
by: Luo, Yijia, et al.
Published: (2025)
Token Preference Optimization with Self-Calibrated Visual-Anchored Rewards for Hallucination Mitigation
by: Gu, Jihao, et al.
Published: (2024)
by: Gu, Jihao, et al.
Published: (2024)
ToolRM: Towards Agentic Tool-Use Reward Modeling
by: Li, Renhao, et al.
Published: (2025)
by: Li, Renhao, et al.
Published: (2025)
IRPM: Intergroup Relative Preference Modeling for Pointwise Generative Reward Models
by: Song, Haonan, et al.
Published: (2026)
by: Song, Haonan, et al.
Published: (2026)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
AIR: Complex Instruction Generation via Automatic Iterative Refinement
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
by: Lu, Xiaodong, et al.
Published: (2026)
by: Lu, Xiaodong, et al.
Published: (2026)
InfoRM: Mitigating Reward Hacking in RLHF via Information-Theoretic Reward Modeling
by: Miao, Yuchun, et al.
Published: (2024)
by: Miao, Yuchun, et al.
Published: (2024)
Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
by: Xing, Shangyu, et al.
Published: (2025)
by: Xing, Shangyu, et al.
Published: (2025)
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
by: Gao, Mingqi, et al.
Published: (2024)
by: Gao, Mingqi, et al.
Published: (2024)
Optimistic Model Rollouts for Pessimistic Offline Policy Optimization
by: Zhai, Yuanzhao, et al.
Published: (2024)
by: Zhai, Yuanzhao, et al.
Published: (2024)
Quantile Reward Policy Optimization: Alignment with Pointwise Regression and Exact Partition Functions
by: Matrenok, Simon, et al.
Published: (2025)
by: Matrenok, Simon, et al.
Published: (2025)
SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMs
by: Zhang, Zhenliang, et al.
Published: (2025)
by: Zhang, Zhenliang, et al.
Published: (2025)
CFunModel: A "Funny" Language Model Capable of Chinese Humor Generation and Processing
by: Yu, Zhenghan, et al.
Published: (2025)
by: Yu, Zhenghan, et al.
Published: (2025)
SMART-RAG: Selection using Determinantal Matrices for Augmented Retrieval
by: Li, Jiatao, et al.
Published: (2024)
by: Li, Jiatao, et al.
Published: (2024)
Analysis of Two-Stage Rollout Designs with Clustering for Causal Inference under Network Interference
by: Cortez-Rodriguez, Mayleen, et al.
Published: (2024)
by: Cortez-Rodriguez, Mayleen, et al.
Published: (2024)
Proof-RM: A Scalable and Generalizable Reward Model for Math Proof
by: Yang, Haotong, et al.
Published: (2026)
by: Yang, Haotong, et al.
Published: (2026)
Weighted-Reward Preference Optimization for Implicit Model Fusion
by: Yang, Ziyi, et al.
Published: (2024)
by: Yang, Ziyi, et al.
Published: (2024)
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
by: He, Yancheng, et al.
Published: (2024)
by: He, Yancheng, et al.
Published: (2024)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
AgentRM: Enhancing Agent Generalization with Reward Modeling
by: Xia, Yu, et al.
Published: (2025)
by: Xia, Yu, et al.
Published: (2025)
DR$^2$Seg: Decomposed Two-Stage Rollouts for Efficient Reasoning Segmentation in Multimodal Large Language Models
by: He, Yulin, et al.
Published: (2026)
by: He, Yulin, et al.
Published: (2026)
Beyond Pointwise Scores: Decomposed Criteria-Based Evaluation of LLM Responses
by: Yu, Fangyi, et al.
Published: (2025)
by: Yu, Fangyi, et al.
Published: (2025)
RM-R1: Reward Modeling as Reasoning
by: Chen, Xiusi, et al.
Published: (2025)
by: Chen, Xiusi, et al.
Published: (2025)
WiS Platform: Enhancing Evaluation of LLM-Based Multi-Agent Systems Through Game-Based Analysis
by: Hu, Chengwei, et al.
Published: (2024)
by: Hu, Chengwei, et al.
Published: (2024)
A Unified Approach to Two Pointwise Ergodic Theorems: Double Recurrence and Return Times
by: Krause, Ben
Published: (2025)
by: Krause, Ben
Published: (2025)
UCS: A Unified Approach to Cell Segmentation for Subcellular Spatial Transcriptomics
by: Yuheng Chen, et al.
Published: (2025)
by: Yuheng Chen, et al.
Published: (2025)
RecRM-Bench: Benchmarking Multidimensional Reward Modeling for Agentic Recommender Systems
by: Zeng, Wenwen, et al.
Published: (2026)
by: Zeng, Wenwen, et al.
Published: (2026)
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Similar Items
-
Asymmetric Proximal Policy Optimization: mini-critics boost LLM reasoning
by: Liu, Jiashun, et al.
Published: (2025) -
2D-DPO: Scaling Direct Preference Optimization with 2-Dimensional Supervision
by: Li, Shilong, et al.
Published: (2024) -
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025) -
Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
by: Liu, Zihe, et al.
Published: (2025) -
Attention Illuminates LLM Reasoning: The Preplan-and-Anchor Rhythm Enables Fine-Grained Policy Optimization
by: Li, Yang, et al.
Published: (2025)