Beyond Length Scaling: Synergizing Breadth and Depth for Generative Reward Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Qiyuan, Wang, Yufei, Wu, Tianhe, Xu, Can, Sun, Qingfeng, Zheng, Kai, Liu, Xue, Ma, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RubricBench: Aligning Model-Generated Rubrics with Human Standards
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026)
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024)
Towards a Unified Paradigm: Integrating Recommendation Systems as a New Language in Large Models
von: Zheng, Kai, et al.
Veröffentlicht: (2024)
von: Zheng, Kai, et al.
Veröffentlicht: (2024)
Zero Reinforcement Learning Towards General Domains
von: Zeng, Yuyuan, et al.
Veröffentlicht: (2025)
von: Zeng, Yuyuan, et al.
Veröffentlicht: (2025)
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
von: Luo, Haipeng, et al.
Veröffentlicht: (2025)
OffSeeker: Online Reinforcement Learning Is Not All You Need for Deep Research Agents
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
von: Zhou, Yuhang, et al.
Veröffentlicht: (2026)
In-Depth and In-Breadth: Pre-training Multimodal Language Models Customized for Comprehensive Chart Understanding
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
von: Fan, Wan-Cyuan, et al.
Veröffentlicht: (2025)
Bias Fitting to Mitigate Length Bias of Reward Model in RLHF
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
von: Zhao, Kangwen, et al.
Veröffentlicht: (2025)
Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
von: Polonuer, Joaquín, et al.
Veröffentlicht: (2026)
von: Polonuer, Joaquín, et al.
Veröffentlicht: (2026)
A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2025)
Scaling Autonomous Agents via Automatic Reward Modeling And Planning
von: Chen, Zhenfang, et al.
Veröffentlicht: (2025)
von: Chen, Zhenfang, et al.
Veröffentlicht: (2025)
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
Expected Runtime Comparisons Between Breadth-First Search and Constant-Depth Restarting Random Walks
von: Platnick, Daniel, et al.
Veröffentlicht: (2024)
von: Platnick, Daniel, et al.
Veröffentlicht: (2024)
Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
von: Wang, Chaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Chaoqi, et al.
Veröffentlicht: (2025)
Collaborative Performance Prediction for Large Language Models
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2024)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
Inference-Time Scaling for Generalist Reward Modeling
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
von: Liu, Zijun, et al.
Veröffentlicht: (2025)
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
von: Chen, Xuetian, et al.
Veröffentlicht: (2025)
von: Chen, Xuetian, et al.
Veröffentlicht: (2025)
Compute Allocation in Evolutionary Search: From Depth-Breadth to Multi-Armed Bandits
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
von: Xing, Sixue, et al.
Veröffentlicht: (2026)
Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
von: Li, Yanhao, et al.
Veröffentlicht: (2025)
von: Li, Yanhao, et al.
Veröffentlicht: (2025)
Beyond Surface Judgments: Human-Grounded Risk Evaluation of LLM-Generated Disinformation
von: Xu, Zonghuan, et al.
Veröffentlicht: (2026)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2026)
Mamba Modulation: On the Length Generalization of Mamba
von: Lu, Peng, et al.
Veröffentlicht: (2025)
von: Lu, Peng, et al.
Veröffentlicht: (2025)
In-Context Reward Adaptation for Robust Preference Modeling
von: Sun, Zhenyu, et al.
Veröffentlicht: (2026)
von: Sun, Zhenyu, et al.
Veröffentlicht: (2026)
Dual Engines of Thoughts: A Depth-Breadth Integration Framework for Open-Ended Analysis
von: Yu, Fei-Hsuan, et al.
Veröffentlicht: (2025)
von: Yu, Fei-Hsuan, et al.
Veröffentlicht: (2025)
Beyond Parameters: Exploring Virtual Logic Depth for Scaling Laws
von: Zhu, Ruike, et al.
Veröffentlicht: (2025)
von: Zhu, Ruike, et al.
Veröffentlicht: (2025)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
Confidence as a Reward: Transforming LLMs into Reward Models
von: Du, He, et al.
Veröffentlicht: (2025)
von: Du, He, et al.
Veröffentlicht: (2025)
FormalRewardBench: A Benchmark for Formal Theorem Proving Reward Models
von: Uluşan, Zeynel A., et al.
Veröffentlicht: (2026)
von: Uluşan, Zeynel A., et al.
Veröffentlicht: (2026)
Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling
von: Liu, Yuchen, et al.
Veröffentlicht: (2026)
von: Liu, Yuchen, et al.
Veröffentlicht: (2026)
From Static to Interactive: Authoring Interactive Visualizations via Natural Language
von: Liu, Can, et al.
Veröffentlicht: (2026)
von: Liu, Can, et al.
Veröffentlicht: (2026)
WizardLM: Empowering large pre-trained language models to follow complex instructions
von: Xu, Can, et al.
Veröffentlicht: (2023)
von: Xu, Can, et al.
Veröffentlicht: (2023)
SpeakRL: Synergizing Reasoning, Speaking, and Acting in Language Models with Reinforcement Learning
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
von: Acikgoz, Emre Can, et al.
Veröffentlicht: (2025)
Structural Reward Model: Enhancing Interpretability, Efficiency, and Scalability in Reward Modeling
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyu, et al.
Veröffentlicht: (2025)
CoLD: Counterfactually-Guided Length Debiasing for Process Reward Models in Mathematical Reasoning
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
von: Zheng, Congmin, et al.
Veröffentlicht: (2025)
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
von: Hu, Mengkang, et al.
Veröffentlicht: (2024)
Learning What Matters: Dynamic Dimension Selection and Aggregation for Interpretable Vision-Language Reward Modeling
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
von: Chen, Qiyuan, et al.
Veröffentlicht: (2026)
SemiReward: A General Reward Model for Semi-supervised Learning
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
von: Li, Siyuan, et al.
Veröffentlicht: (2023)
Beyond Sparse Rewards: Enhancing Reinforcement Learning with Language Model Critique in Text Generation
von: Cao, Meng, et al.
Veröffentlicht: (2024)
von: Cao, Meng, et al.
Veröffentlicht: (2024)
From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation
von: Guo, Zhihan, et al.
Veröffentlicht: (2025)
von: Guo, Zhihan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
RubricBench: Aligning Model-Generated Rubrics with Human Standards
von: Zhang, Qiyuan, et al.
Veröffentlicht: (2026) -
Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion
von: Chen, Jiuhai, et al.
Veröffentlicht: (2024) -
Towards a Unified Paradigm: Integrating Recommendation Systems as a New Language in Large Models
von: Zheng, Kai, et al.
Veröffentlicht: (2024) -
Zero Reinforcement Learning Towards General Domains
von: Zeng, Yuyuan, et al.
Veröffentlicht: (2025) -
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)