SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yifan, Li, Bolian, Cho, David, Zhang, Ruqi, Sui, Fanping, Grama, Ananth |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference-Time Code Selection via Symbolic Equivalence Partitioning
by: Cho, David, et al.
Published: (2026)
by: Cho, David, et al.
Published: (2026)
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024)
by: Li, Bolian, et al.
Published: (2024)
More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
No Free Lunch: Fundamental Limits of Learning Non-Hallucinating Generative Models
by: Wu, Changlong, et al.
Published: (2024)
by: Wu, Changlong, et al.
Published: (2024)
Addressing Performance Saturation for LLM RL via Precise Entropy Curve Control
by: Li, Bolian, et al.
Published: (2026)
by: Li, Bolian, et al.
Published: (2026)
Generalized Learning of Coefficients in Spectral Graph Convolutional Networks
by: Coşkun, Mustafa, et al.
Published: (2024)
by: Coşkun, Mustafa, et al.
Published: (2024)
SAL: Selective Adaptive Learning for Backpropagation-Free Training with Sparsification
by: Liu, Fanping, et al.
Published: (2026)
by: Liu, Fanping, et al.
Published: (2026)
LLMs Can Get "Brain Rot": A Pilot Study on Twitter/X
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
by: Zhao, Zijian, et al.
Published: (2026)
by: Zhao, Zijian, et al.
Published: (2026)
Structure-R1: Dynamically Leveraging Structural Knowledge in LLM Reasoning through Reinforcement Learning
by: Wu, Junlin, et al.
Published: (2025)
by: Wu, Junlin, et al.
Published: (2025)
Multi-Agent Collaborative Reward Design for Enhancing Reasoning in Reinforcement Learning
by: Yang, Pei, et al.
Published: (2025)
by: Yang, Pei, et al.
Published: (2025)
Think in Games: Learning to Reason in Games via Reinforcement Learning with Large Language Models
by: Liao, Yi, et al.
Published: (2025)
by: Liao, Yi, et al.
Published: (2025)
The Semantic Training Gap: Ontology-Grounded Tool Architectures for Industrial AI Agent Systems
by: Chethan, Grama
Published: (2026)
by: Chethan, Grama
Published: (2026)
Template-as-Ontology: Configurable Synthetic Data Infrastructure for Cross-Domain Manufacturing AI Validation
by: Chethan, Grama
Published: (2026)
by: Chethan, Grama
Published: (2026)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
by: Li, Mengqi, et al.
Published: (2025)
by: Li, Mengqi, et al.
Published: (2025)
Entropy-MCMC: Sampling from Flat Basins with Ease
by: Li, Bolian, et al.
Published: (2023)
by: Li, Bolian, et al.
Published: (2023)
Making Reliable and Flexible Decisions in Long-tailed Classification
by: Li, Bolian, et al.
Published: (2025)
by: Li, Bolian, et al.
Published: (2025)
Learning Reasoning Rewards from Expert Demonstrations with Inverse Reinforcement Learning
by: Fanconi, Claudio, et al.
Published: (2025)
by: Fanconi, Claudio, et al.
Published: (2025)
Learning Rewards, Not Labels: Adversarial Inverse Reinforcement Learning for Machinery Fault Detection
by: Neupane, Dhiraj, et al.
Published: (2026)
by: Neupane, Dhiraj, et al.
Published: (2026)
RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
by: Zhang, Zijing, et al.
Published: (2025)
by: Zhang, Zijing, et al.
Published: (2025)
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning
by: Zhang, Kongcheng, et al.
Published: (2025)
by: Zhang, Kongcheng, et al.
Published: (2025)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
by: Ye, Zhiling, et al.
Published: (2025)
by: Ye, Zhiling, et al.
Published: (2025)
Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
by: Wen, Xumeng, et al.
Published: (2025)
by: Wen, Xumeng, et al.
Published: (2025)
From Roots to Rewards: Dynamic Tree Reasoning with Reinforcement Learning
by: Bahloul, Ahmed, et al.
Published: (2025)
by: Bahloul, Ahmed, et al.
Published: (2025)
Reward-Shifted Speculative Sampling Is An Efficient Test-Time Weak-to-Strong Aligner
by: Li, Bolian, et al.
Published: (2025)
by: Li, Bolian, et al.
Published: (2025)
Scaling Medical Reasoning Verification via Tool-Integrated Reinforcement Learning
by: Zhang, Hang, et al.
Published: (2026)
by: Zhang, Hang, et al.
Published: (2026)
Label-Free Reinforcement Learning via Cross-Model Entropy
by: Gorbett, Matt, et al.
Published: (2026)
by: Gorbett, Matt, et al.
Published: (2026)
Attention-Guided Reward for Reinforcement Learning-based Jailbreak against Large Reasoning Models
by: Lin, Zheng, et al.
Published: (2026)
by: Lin, Zheng, et al.
Published: (2026)
Artificial Intelligence for Optimal Learning: A Comparative Approach towards AI-Enhanced Learning Environments
by: Hariharan, Ananth
Published: (2025)
by: Hariharan, Ananth
Published: (2025)
Training-Free Time Series Classification via In-Context Reasoning with LLM Agents
by: Sui, Songyuan, et al.
Published: (2025)
by: Sui, Songyuan, et al.
Published: (2025)
RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents
by: Wang, Peisong, et al.
Published: (2025)
by: Wang, Peisong, et al.
Published: (2025)
Reward-Decomposed Reinforcement Learning for Immersive Video Role-Playing
by: Wang, Miao, et al.
Published: (2026)
by: Wang, Miao, et al.
Published: (2026)
Efficient Reasoning via Reward Model
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
Reward Training Wheels: Adaptive Auxiliary Rewards for Robotics Reinforcement Learning
by: Wang, Linji, et al.
Published: (2025)
by: Wang, Linji, et al.
Published: (2025)
Reward Hacking in Rubric-Based Reinforcement Learning
by: Mahmoud, Anas, et al.
Published: (2026)
by: Mahmoud, Anas, et al.
Published: (2026)
RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
by: Feng, Sicheng, et al.
Published: (2025)
by: Feng, Sicheng, et al.
Published: (2025)
Single Agent Robust Deep Reinforcement Learning for Bus Fleet Control
by: Zhang, Yifan
Published: (2025)
by: Zhang, Yifan
Published: (2025)
Learning in Context, Guided by Choice: A Reward-Free Paradigm for Reinforcement Learning with Transformers
by: Dong, Juncheng, et al.
Published: (2026)
by: Dong, Juncheng, et al.
Published: (2026)
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
by: Gan, Siyuan, et al.
Published: (2026)
by: Gan, Siyuan, et al.
Published: (2026)
Similar Items
-
Inference-Time Code Selection via Symbolic Equivalence Partitioning
by: Cho, David, et al.
Published: (2026) -
DRIFT: Learning from Abundant User Dissatisfaction in Real-World Preference Learning
by: Wang, Yifan, et al.
Published: (2025) -
Cascade Reward Sampling for Efficient Decoding-Time Alignment
by: Li, Bolian, et al.
Published: (2024) -
More is Less: The Pitfalls of Multi-Model Synthetic Preference Data in DPO Safety Alignment
by: Wang, Yifan, et al.
Published: (2025) -
No Free Lunch: Fundamental Limits of Learning Non-Hallucinating Generative Models
by: Wu, Changlong, et al.
Published: (2024)