LongR: Unleashing Long-Context Reasoning via Reinforcement Learning with Dense Utility Rewards
Fuente:
arXiv
Saved in:
| Main Authors: | Ping, Bowen, Chen, Zijun, Yu, Yiyao, Hui, Tingfeng, Yan, Junchi, Chang, Baobao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
by: Ping, Bowen, et al.
Published: (2026)
by: Ping, Bowen, et al.
Published: (2026)
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
by: Chen, Guanzheng, et al.
Published: (2026)
by: Chen, Guanzheng, et al.
Published: (2026)
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
by: Ling, Zhan, et al.
Published: (2025)
by: Ling, Zhan, et al.
Published: (2025)
Reducing Distraction in Long-Context Language Models by Focused Learning
by: Wu, Zijun, et al.
Published: (2024)
by: Wu, Zijun, et al.
Published: (2024)
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
by: Wan, Fanqi, et al.
Published: (2025)
by: Wan, Fanqi, et al.
Published: (2025)
MemBuilder: Reinforcing LLMs for Long-Term Memory Construction via Attributed Dense Rewards
by: Shen, Zhiyu, et al.
Published: (2026)
by: Shen, Zhiyu, et al.
Published: (2026)
TAMTRL: Teacher-Aligned Reward Reshaping for Multi-Turn Reinforcement Learning in Long-Context Compression
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning
by: Hu, Ziyou, et al.
Published: (2025)
by: Hu, Ziyou, et al.
Published: (2025)
LoongRL: Reinforcement Learning for Advanced Reasoning over Long Contexts
by: Wang, Siyuan, et al.
Published: (2025)
by: Wang, Siyuan, et al.
Published: (2025)
GATEAU: Selecting Influential Samples for Long Context Alignment
by: Si, Shuzheng, et al.
Published: (2024)
by: Si, Shuzheng, et al.
Published: (2024)
Dynamic Long Context Reasoning over Compressed Memory via End-to-End Reinforcement Learning
by: Chen, Zhuoen, et al.
Published: (2026)
by: Chen, Zhuoen, et al.
Published: (2026)
LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
by: Bai, Yushi, et al.
Published: (2024)
by: Bai, Yushi, et al.
Published: (2024)
LongAttnComp: Cross-Family Context Compression for Long-Context Reasoning
by: Ji, Mengmeng, et al.
Published: (2026)
by: Ji, Mengmeng, et al.
Published: (2026)
Hybrid Reinforcement: When Reward Is Sparse, It's Better to Be Dense
by: Tao, Leitian, et al.
Published: (2025)
by: Tao, Leitian, et al.
Published: (2025)
Long-form RewardBench: Evaluating Reward Models for Long-form Generation
by: Huang, Hui, et al.
Published: (2026)
by: Huang, Hui, et al.
Published: (2026)
GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignment
by: Lv, Minxuan, et al.
Published: (2026)
by: Lv, Minxuan, et al.
Published: (2026)
Facilitating Long Context Understanding via Supervised Chain-of-Thought Reasoning
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
by: Guan, Xin, et al.
Published: (2026)
by: Guan, Xin, et al.
Published: (2026)
LongFaith: Enhancing Long-Context Reasoning in LLMs with Faithful Synthetic Data
by: Yang, Cehao, et al.
Published: (2025)
by: Yang, Cehao, et al.
Published: (2025)
ProtoReasoning: Prototypes as the Foundation for Generalizable Reasoning in LLMs
by: He, Feng, et al.
Published: (2025)
by: He, Feng, et al.
Published: (2025)
SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models
by: Yang, Ziyi, et al.
Published: (2025)
by: Yang, Ziyi, et al.
Published: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
by: Hui, Tingfeng, et al.
Published: (2024)
by: Hui, Tingfeng, et al.
Published: (2024)
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution
by: Li, Jiahui, et al.
Published: (2024)
by: Li, Jiahui, et al.
Published: (2024)
R-Search: Empowering LLM Reasoning with Search via Multi-Reward Reinforcement Learning
by: Zhao, Qingfei, et al.
Published: (2025)
by: Zhao, Qingfei, et al.
Published: (2025)
Long Context Pre-Training with Lighthouse Attention
by: Peng, Bowen, et al.
Published: (2026)
by: Peng, Bowen, et al.
Published: (2026)
ACE-RL: Adaptive Constraint-Enhanced Reward for Long-form Generation Reinforcement Learning
by: Chen, Jianghao, et al.
Published: (2025)
by: Chen, Jianghao, et al.
Published: (2025)
Hit-RAG: Learning to Reason with Long Contexts via Preference Alignment
by: Liu, Junming, et al.
Published: (2026)
by: Liu, Junming, et al.
Published: (2026)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
Joint Enhancement of Relational Reasoning for Long-Context LLMs
by: Chen, Zhirui, et al.
Published: (2025)
by: Chen, Zhirui, et al.
Published: (2025)
RM-R1: Reward Modeling as Reasoning
by: Chen, Xiusi, et al.
Published: (2025)
by: Chen, Xiusi, et al.
Published: (2025)
PERK: Long-Context Reasoning as Parameter-Efficient Test-Time Learning
by: Chen, Zeming, et al.
Published: (2025)
by: Chen, Zeming, et al.
Published: (2025)
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
by: Liu, Zhenghao, et al.
Published: (2026)
by: Liu, Zhenghao, et al.
Published: (2026)
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
by: Mitra, Purbesh, et al.
Published: (2025)
by: Mitra, Purbesh, et al.
Published: (2025)
Tournament-GRPO: Group-Wise Tournament Rewards for Reinforcement Learning in Open-Ended Long-Form Generation
by: Yang, Zixuan, et al.
Published: (2026)
by: Yang, Zixuan, et al.
Published: (2026)
Incentivizing In-depth Reasoning over Long Contexts with Process Advantage Shaping
by: Peng, Miao, et al.
Published: (2026)
by: Peng, Miao, et al.
Published: (2026)
LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information
by: Ping, Bowen, et al.
Published: (2025)
by: Ping, Bowen, et al.
Published: (2025)
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
by: Qiao, Zile, et al.
Published: (2025)
by: Qiao, Zile, et al.
Published: (2025)
LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
by: Tang, Zecheng, et al.
Published: (2025)
by: Tang, Zecheng, et al.
Published: (2025)
Beyond Isolated Capabilities: Bridging Long CoT Reasoning and Long-Context Understanding
by: Wang, Yifei
Published: (2025)
by: Wang, Yifei
Published: (2025)
Similar Items
-
LongAct: Harnessing Intrinsic Activation Patterns for Long-Context Reinforcement Learning
by: Ping, Bowen, et al.
Published: (2026) -
LongRLVR: Long-Context Reinforcement Learning Requires Verifiable Context Rewards
by: Chen, Guanzheng, et al.
Published: (2026) -
LongReason: A Synthetic Long-Context Reasoning Benchmark via Context Expansion
by: Ling, Zhan, et al.
Published: (2025) -
Reducing Distraction in Long-Context Language Models by Focused Learning
by: Wu, Zijun, et al.
Published: (2024) -
QwenLong-L1: Towards Long-Context Large Reasoning Models with Reinforcement Learning
by: Wan, Fanqi, et al.
Published: (2025)