Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Fuente:
arXiv
Guardado en:
| Autores principales: | Su, Yi, Yu, Dian, Song, Linfeng, Li, Juntao, Mi, Haitao, Tu, Zhaopeng, Zhang, Min, Yu, Dong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
por: Liu, Xiaoyuan, et al.
Publicado: (2025)
Teaching LLMs to Refine with Tools
por: Yu, Dian, et al.
Publicado: (2024)
por: Yu, Dian, et al.
Publicado: (2024)
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
por: Wang, Ante, et al.
Publicado: (2025)
por: Wang, Ante, et al.
Publicado: (2025)
SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
por: Ding, Yuyang, et al.
Publicado: (2025)
por: Ding, Yuyang, et al.
Publicado: (2025)
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
por: Yu, Dian, et al.
Publicado: (2024)
por: Yu, Dian, et al.
Publicado: (2024)
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
por: Wang, Yue, et al.
Publicado: (2025)
por: Wang, Yue, et al.
Publicado: (2025)
Every Question Has Its Own Value: Reinforcement Learning with Explicit Human Values
por: Yu, Dian, et al.
Publicado: (2025)
por: Yu, Dian, et al.
Publicado: (2025)
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
por: He, Zhiwei, et al.
Publicado: (2025)
por: He, Zhiwei, et al.
Publicado: (2025)
Verified Critical Step Optimization for LLM Agents
por: Li, Mukai, et al.
Publicado: (2026)
por: Li, Mukai, et al.
Publicado: (2026)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
por: Dai, Runpeng, et al.
Publicado: (2025)
por: Dai, Runpeng, et al.
Publicado: (2025)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
por: Tian, Ye, et al.
Publicado: (2024)
por: Tian, Ye, et al.
Publicado: (2024)
AgentV-RL: Scaling Reward Modeling with Agentic Verifier
por: Zhang, Jiazheng, et al.
Publicado: (2026)
por: Zhang, Jiazheng, et al.
Publicado: (2026)
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
por: Liang, Zhenwen, et al.
Publicado: (2025)
por: Liang, Zhenwen, et al.
Publicado: (2025)
LiteSearch: Efficacious Tree Search for LLM
por: Wang, Ante, et al.
Publicado: (2024)
por: Wang, Ante, et al.
Publicado: (2024)
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
por: Shen, Yiran, et al.
Publicado: (2025)
por: Shen, Yiran, et al.
Publicado: (2025)
Inconsistent dialogue responses and how to recover from them
por: Zhang, Mian, et al.
Publicado: (2024)
por: Zhang, Mian, et al.
Publicado: (2024)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
por: Wang, Peisong, et al.
Publicado: (2025)
por: Wang, Peisong, et al.
Publicado: (2025)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
por: Zhang, Xin, et al.
Publicado: (2026)
por: Zhang, Xin, et al.
Publicado: (2026)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
por: Li, Gaotang, et al.
Publicado: (2026)
por: Li, Gaotang, et al.
Publicado: (2026)
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
por: Chen, Xingyu, et al.
Publicado: (2024)
por: Chen, Xingyu, et al.
Publicado: (2024)
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification
por: Liu, Rui, et al.
Publicado: (2026)
por: Liu, Rui, et al.
Publicado: (2026)
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
por: Zhang, Yuheng, et al.
Publicado: (2025)
por: Zhang, Yuheng, et al.
Publicado: (2025)
DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
por: Liang, Tian, et al.
Publicado: (2025)
por: Liang, Tian, et al.
Publicado: (2025)
Blockwise Advantage Estimation for Multi-Objective RL with Verifiable Rewards
por: Pavlenko, Kirill, et al.
Publicado: (2026)
por: Pavlenko, Kirill, et al.
Publicado: (2026)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
por: Wang, Xiyao, et al.
Publicado: (2024)
por: Wang, Xiyao, et al.
Publicado: (2024)
On-Policy RL with Optimal Reward Baseline
por: Hao, Yaru, et al.
Publicado: (2025)
por: Hao, Yaru, et al.
Publicado: (2025)
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
por: Fan, Mingyuan, et al.
Publicado: (2026)
por: Fan, Mingyuan, et al.
Publicado: (2026)
A Knowledge Plug-and-Play Test Bed for Open-domain Dialogue Generation
por: Li, Xiangci, et al.
Publicado: (2024)
por: Li, Xiangci, et al.
Publicado: (2024)
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
por: Li, Ming, et al.
Publicado: (2025)
por: Li, Ming, et al.
Publicado: (2025)
MT-RewardTree: A Comprehensive Framework for Advancing LLM-Based Machine Translation via Reward Modeling
por: Feng, Zhaopeng, et al.
Publicado: (2025)
por: Feng, Zhaopeng, et al.
Publicado: (2025)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
por: Jiang, Yuxin, et al.
Publicado: (2026)
por: Jiang, Yuxin, et al.
Publicado: (2026)
Evolving Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
por: Zhou, Yujun, et al.
Publicado: (2025)
por: Zhou, Yujun, et al.
Publicado: (2025)
Sotopia-RL: Reward Design for Social Intelligence
por: Yu, Haofei, et al.
Publicado: (2025)
por: Yu, Haofei, et al.
Publicado: (2025)
Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
por: Jia, Ruipeng, et al.
Publicado: (2025)
por: Jia, Ruipeng, et al.
Publicado: (2025)
Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
por: Yu, Qinan, et al.
Publicado: (2026)
por: Yu, Qinan, et al.
Publicado: (2026)
Entropy Guided Extrapolative Decoding to Improve Factuality in Large Language Models
por: Das, Souvik, et al.
Publicado: (2024)
por: Das, Souvik, et al.
Publicado: (2024)
Collaborative decoding of critical tokens for boosting factuality of large language models
por: Jin, Lifeng, et al.
Publicado: (2024)
por: Jin, Lifeng, et al.
Publicado: (2024)
Save the Good Prefix: Precise Error Penalization via Process-Supervised RL to Enhance LLM Reasoning
por: Liu, Haolin, et al.
Publicado: (2026)
por: Liu, Haolin, et al.
Publicado: (2026)
Rethinking Sample Polarity in Reinforcement Learning with Verifiable Rewards
por: Tang, Xinyu, et al.
Publicado: (2025)
por: Tang, Xinyu, et al.
Publicado: (2025)
Ejemplares similares
-
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
por: Liu, Xiaoyuan, et al.
Publicado: (2025) -
Teaching LLMs to Refine with Tools
por: Yu, Dian, et al.
Publicado: (2024) -
Don't Get Lost in the Trees: Streamlining LLM Reasoning by Overcoming Tree Search Exploration Pitfalls
por: Wang, Ante, et al.
Publicado: (2025) -
SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning
por: Ding, Yuyang, et al.
Publicado: (2025) -
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
por: Yu, Dian, et al.
Publicado: (2024)