Co-rewarding: Stable Self-supervised RL for Eliciting Reasoning in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zizhuo, Zhu, Jianing, Ge, Xinmu, Zhao, Zihua, Zhou, Zhanke, Li, Xuan, Feng, Xiao, Yao, Jiangchao, Han, Bo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
by: Li, Xuan, et al.
Published: (2023)
by: Li, Xuan, et al.
Published: (2023)
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
Less is More: One-shot Subgraph Reasoning on Large-scale Knowledge Graphs
by: Zhou, Zhanke, et al.
Published: (2024)
by: Zhou, Zhanke, et al.
Published: (2024)
Towards Understanding Valuable Preference Data for Large Language Model Alignment
by: Zhang, Zizhuo, et al.
Published: (2025)
by: Zhang, Zizhuo, et al.
Published: (2025)
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
by: Wang, Yajing, et al.
Published: (2024)
by: Wang, Yajing, et al.
Published: (2024)
RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with Large Language Models
by: Feng, Xiao, et al.
Published: (2026)
by: Feng, Xiao, et al.
Published: (2026)
Self-Calibrated Tuning of Vision-Language Models for Out-of-Distribution Detection
by: Yu, Geng, et al.
Published: (2024)
by: Yu, Geng, et al.
Published: (2024)
Reference-guided Policy Optimization for Molecular Optimization via LLM Reasoning
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
Neural Atoms: Propagating Long-range Interaction in Molecular Graphs through Efficient Communication Channel
by: Li, Xuan, et al.
Published: (2023)
by: Li, Xuan, et al.
Published: (2023)
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales?
by: Zhou, Zhanke, et al.
Published: (2024)
by: Zhou, Zhanke, et al.
Published: (2024)
Model Inversion Attacks: A Survey of Approaches and Countermeasures
by: Zhou, Zhanke, et al.
Published: (2024)
by: Zhou, Zhanke, et al.
Published: (2024)
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory
by: Sun, Jingwei, et al.
Published: (2026)
by: Sun, Jingwei, et al.
Published: (2026)
Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards
by: Huang, Yu, et al.
Published: (2026)
by: Huang, Yu, et al.
Published: (2026)
Fast and Accurate Blind Flexible Docking
by: Zhang, Zizhuo, et al.
Published: (2025)
by: Zhang, Zizhuo, et al.
Published: (2025)
Physics Reasoner: Knowledge-Augmented Reasoning for Solving Physics Problems with Large Language Models
by: Pang, Xinyu, et al.
Published: (2024)
by: Pang, Xinyu, et al.
Published: (2024)
Decoupling the Class Label and the Target Concept in Machine Unlearning
by: Zhu, Jianing, et al.
Published: (2024)
by: Zhu, Jianing, et al.
Published: (2024)
Envisioning Outlier Exposure by Large Language Models for Out-of-Distribution Detection
by: Cao, Chentao, et al.
Published: (2024)
by: Cao, Chentao, et al.
Published: (2024)
Per-parameter Task Arithmetic for Unlearning in Large Language Models
by: Cai, Chengyi, et al.
Published: (2026)
by: Cai, Chengyi, et al.
Published: (2026)
Self-Training Elicits Concise Reasoning in Large Language Models
by: Munkhbat, Tergel, et al.
Published: (2025)
by: Munkhbat, Tergel, et al.
Published: (2025)
Self-Debias: Self-correcting for Debiasing Large Language Models
by: Feng, Xuan, et al.
Published: (2026)
by: Feng, Xuan, et al.
Published: (2026)
Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach
by: Li, Haolin, et al.
Published: (2026)
by: Li, Haolin, et al.
Published: (2026)
AlphaApollo: A System for Deep Agentic Reasoning
by: Zhou, Zhanke, et al.
Published: (2025)
by: Zhou, Zhanke, et al.
Published: (2025)
SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
by: Guo, Xiaojun, et al.
Published: (2025)
by: Guo, Xiaojun, et al.
Published: (2025)
Understanding Fairness Surrogate Functions in Algorithmic Fairness
by: Yao, Wei, et al.
Published: (2023)
by: Yao, Wei, et al.
Published: (2023)
Rethinking Reasoning Quality in Large Language Models through Enhanced Chain-of-Thought via RL
by: He, Haoyang, et al.
Published: (2025)
by: He, Haoyang, et al.
Published: (2025)
From Debate to Equilibrium: Belief-Driven Multi-Agent LLM Reasoning via Bayesian Nash Equilibrium
by: Yi, Xie, et al.
Published: (2025)
by: Yi, Xie, et al.
Published: (2025)
Mitigating Noisy Correspondence by Geometrical Structure Consistency Learning
by: Zhao, Zihua, et al.
Published: (2024)
by: Zhao, Zihua, et al.
Published: (2024)
Enhancing Sequential Recommendation with World Knowledge from Large Language Models
by: Dai, Tianjie, et al.
Published: (2025)
by: Dai, Tianjie, et al.
Published: (2025)
Eliciting In-context Retrieval and Reasoning for Long-context Large Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Dual-granularity Sinkhorn Distillation for Enhanced Learning from Long-tailed Noisy Data
by: Hong, Feng, et al.
Published: (2025)
by: Hong, Feng, et al.
Published: (2025)
Is There Only One Vacuum?
by: Zong, Xinmu, et al.
Published: (2025)
by: Zong, Xinmu, et al.
Published: (2025)
Self-supervised network distillation: an effective approach to exploration in sparse reward environments
by: Pecháč, Matej, et al.
Published: (2023)
by: Pecháč, Matej, et al.
Published: (2023)
Exploring Training on Heterogeneous Data with Mixture of Low-rank Adapters
by: Zhou, Yuhang, et al.
Published: (2024)
by: Zhou, Yuhang, et al.
Published: (2024)
Distributive fairness during the transition to adolescence: The role of peer comparison and social value orientation
by: Siqi Liu, et al.
Published: (2024)
by: Siqi Liu, et al.
Published: (2024)
Mind the Gap Between Prototypes and Images in Cross-domain Finetuning
by: Tian, Hongduan, et al.
Published: (2024)
by: Tian, Hongduan, et al.
Published: (2024)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
by: Liu, Shih-Yang, et al.
Published: (2026)
by: Liu, Shih-Yang, et al.
Published: (2026)
Self-supervised Hierarchical Visual Reasoning with World Model
by: Xu, Yuanfei, et al.
Published: (2026)
by: Xu, Yuanfei, et al.
Published: (2026)
Unraveling the Impact of Heterophilic Structures on Graph Positive-Unlabeled Learning
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Noisy Test-Time Adaptation in Vision-Language Models
by: Cao, Chentao, et al.
Published: (2025)
by: Cao, Chentao, et al.
Published: (2025)
Similar Items
-
DeepInception: Hypnotize Large Language Model to Be Jailbreaker
by: Li, Xuan, et al.
Published: (2023) -
From Passive to Active Reasoning: Can Large Language Models Ask the Right Questions under Incomplete Information?
by: Zhou, Zhanke, et al.
Published: (2025) -
Less is More: One-shot Subgraph Reasoning on Large-scale Knowledge Graphs
by: Zhou, Zhanke, et al.
Published: (2024) -
Towards Understanding Valuable Preference Data for Large Language Model Alignment
by: Zhang, Zizhuo, et al.
Published: (2025) -
Eliciting Causal Abilities in Large Language Models for Reasoning Tasks
by: Wang, Yajing, et al.
Published: (2024)