PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Piché, Alexandre, Kamalloo, Ehsan, Pardinas, Rafael, Chen, Xiaoyin, Bahdanau, Dzmitry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026)
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026)
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
von: Piché, Alexandre, et al.
Veröffentlicht: (2024)
Self-Evolving Curriculum for LLM Reasoning
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
von: Patel, Arkil, et al.
Veröffentlicht: (2026)
Bridging the Gap Between Target Networks and Functional Regularization
von: Piche, Alexandre, et al.
Veröffentlicht: (2022)
von: Piche, Alexandre, et al.
Veröffentlicht: (2022)
BRIDGE: Predicting Human Task Completion Time From Model Performance
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
von: Liu, Fengyuan, et al.
Veröffentlicht: (2026)
TapeAgents: a Holistic Framework for Agent Development and Optimization
von: Bahdanau, Dzmitry, et al.
Veröffentlicht: (2024)
von: Bahdanau, Dzmitry, et al.
Veröffentlicht: (2024)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
von: Liu, Shaoteng, et al.
Veröffentlicht: (2024)
von: Liu, Shaoteng, et al.
Veröffentlicht: (2024)
Improved Off-policy Reinforcement Learning in Biological Sequence Design
von: Kim, Hyeonah, et al.
Veröffentlicht: (2024)
von: Kim, Hyeonah, et al.
Veröffentlicht: (2024)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
von: Wang, Jichao, et al.
Veröffentlicht: (2026)
von: Wang, Jichao, et al.
Veröffentlicht: (2026)
Exploring validation metrics for offline model-based optimisation with diffusion models
von: Beckham, Christopher, et al.
Veröffentlicht: (2022)
von: Beckham, Christopher, et al.
Veröffentlicht: (2022)
SPEED-RL: Faster Training of Reasoning Models via Online Curriculum Learning
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
von: Dong, Yihong, et al.
Veröffentlicht: (2025)
PokeRL: Reinforcement Learning for Pokemon Red
von: Mudireddy, Dheeraj, et al.
Veröffentlicht: (2026)
von: Mudireddy, Dheeraj, et al.
Veröffentlicht: (2026)
How to Get Your LLM to Generate Challenging Problems for Evaluation
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
von: Patel, Arkil, et al.
Veröffentlicht: (2025)
FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies
von: Gao, Chenxiao, et al.
Veröffentlicht: (2026)
von: Gao, Chenxiao, et al.
Veröffentlicht: (2026)
Conditional Sequence Modeling for Safe Reinforcement Learning
von: Bai, Wensong, et al.
Veröffentlicht: (2026)
von: Bai, Wensong, et al.
Veröffentlicht: (2026)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
von: Bhatia, Abhinav, et al.
Veröffentlicht: (2023)
von: Bhatia, Abhinav, et al.
Veröffentlicht: (2023)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
von: Oertell, Owen, et al.
Veröffentlicht: (2024)
ObjectRL: An Object-Oriented Reinforcement Learning Codebase
von: Baykal, Gulcin, et al.
Veröffentlicht: (2025)
von: Baykal, Gulcin, et al.
Veröffentlicht: (2025)
Automation and Feature Selection Enhancement with Reinforcement Learning (RL)
von: Nagaraju, Sumana Sanyasipura
Veröffentlicht: (2025)
von: Nagaraju, Sumana Sanyasipura
Veröffentlicht: (2025)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
von: Li, Yuhang, et al.
Veröffentlicht: (2026)
Evaluating In-Context Learning of Libraries for Code Generation
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
von: Patel, Arkil, et al.
Veröffentlicht: (2023)
Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning
von: Huang, Shengyi, et al.
Veröffentlicht: (2024)
von: Huang, Shengyi, et al.
Veröffentlicht: (2024)
Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing
von: Khan, Azal Ahmad, et al.
Veröffentlicht: (2026)
von: Khan, Azal Ahmad, et al.
Veröffentlicht: (2026)
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
von: Xia, Peng, et al.
Veröffentlicht: (2026)
von: Xia, Peng, et al.
Veröffentlicht: (2026)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
von: Rafailov, Rafael, et al.
Veröffentlicht: (2024)
RL-ACRGNet: Reinforcement Learning-Based Chest Radiology Report Generation Network
von: Meena, Yogesh Kumar, et al.
Veröffentlicht: (2026)
von: Meena, Yogesh Kumar, et al.
Veröffentlicht: (2026)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
von: Zhang, Xin, et al.
Veröffentlicht: (2026)
RL's Razor: Why Online Reinforcement Learning Forgets Less
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
von: Shenfeld, Idan, et al.
Veröffentlicht: (2025)
RL as Regressor: A Reinforcement Learning Approach for Function Approximation
von: Huang, Yongchao
Veröffentlicht: (2025)
von: Huang, Yongchao
Veröffentlicht: (2025)
Survival Reinforcement Learning: Toward Scalable Self-Supervised RL
von: Nguimatsia-Tiofack, Franki, et al.
Veröffentlicht: (2026)
von: Nguimatsia-Tiofack, Franki, et al.
Veröffentlicht: (2026)
Probabilistic Learning and Generation in Deep Sequence Models
von: Chen, Wenlong
Veröffentlicht: (2026)
von: Chen, Wenlong
Veröffentlicht: (2026)
Reinformer: Max-Return Sequence Modeling for Offline RL
von: Zhuang, Zifeng, et al.
Veröffentlicht: (2024)
von: Zhuang, Zifeng, et al.
Veröffentlicht: (2024)
FairReweighing: Density Estimation-Based Reweighing Framework for Improving Separation in Fair Regression
von: Xi, Xiaoyin, et al.
Veröffentlicht: (2025)
von: Xi, Xiaoyin, et al.
Veröffentlicht: (2025)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
von: Luo, Cheng, et al.
Veröffentlicht: (2024)
von: Luo, Cheng, et al.
Veröffentlicht: (2024)
Towards Sample-Efficiency and Generalization of Transfer and Inverse Reinforcement Learning: A Comprehensive Literature Review
von: Hassani, Hossein, et al.
Veröffentlicht: (2024)
von: Hassani, Hossein, et al.
Veröffentlicht: (2024)
MAGNNET: Multi-Agent Graph Neural Network-based Efficient Task Allocation for Autonomous Vehicles with Deep Reinforcement Learning
von: Ratnabala, Lavanya, et al.
Veröffentlicht: (2025)
von: Ratnabala, Lavanya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
von: Pardinas, Rafael, et al.
Veröffentlicht: (2026) -
LLMs can learn self-restraint through iterative self-reflection
von: Piché, Alexandre, et al.
Veröffentlicht: (2024) -
Self-Evolving Curriculum for LLM Reasoning
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025) -
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
von: Ardestani, MohamamdJavad, et al.
Veröffentlicht: (2025) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
von: Patel, Arkil, et al.
Veröffentlicht: (2026)