PipelineRL: Faster On-policy Reinforcement Learning for Long Sequence Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Piché, Alexandre, Kamalloo, Ehsan, Pardinas, Rafael, Chen, Xiaoyin, Bahdanau, Dzmitry |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
par: Pardinas, Rafael, et autres
Publié: (2026)
par: Pardinas, Rafael, et autres
Publié: (2026)
LLMs can learn self-restraint through iterative self-reflection
par: Piché, Alexandre, et autres
Publié: (2024)
par: Piché, Alexandre, et autres
Publié: (2024)
Self-Evolving Curriculum for LLM Reasoning
par: Chen, Xiaoyin, et autres
Publié: (2025)
par: Chen, Xiaoyin, et autres
Publié: (2025)
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
par: Ardestani, MohamamdJavad, et autres
Publié: (2025)
par: Ardestani, MohamamdJavad, et autres
Publié: (2025)
Forecasting Downstream Performance of LLMs With Proxy Metrics
par: Patel, Arkil, et autres
Publié: (2026)
par: Patel, Arkil, et autres
Publié: (2026)
Bridging the Gap Between Target Networks and Functional Regularization
par: Piche, Alexandre, et autres
Publié: (2022)
par: Piche, Alexandre, et autres
Publié: (2022)
BRIDGE: Predicting Human Task Completion Time From Model Performance
par: Liu, Fengyuan, et autres
Publié: (2026)
par: Liu, Fengyuan, et autres
Publié: (2026)
TapeAgents: a Holistic Framework for Agent Development and Optimization
par: Bahdanau, Dzmitry, et autres
Publié: (2024)
par: Bahdanau, Dzmitry, et autres
Publié: (2024)
RL-GPT: Integrating Reinforcement Learning and Code-as-policy
par: Liu, Shaoteng, et autres
Publié: (2024)
par: Liu, Shaoteng, et autres
Publié: (2024)
Improved Off-policy Reinforcement Learning in Biological Sequence Design
par: Kim, Hyeonah, et autres
Publié: (2024)
par: Kim, Hyeonah, et autres
Publié: (2024)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
par: Hu, Zhengding, et autres
Publié: (2026)
par: Hu, Zhengding, et autres
Publié: (2026)
SOLAR-RL: Semi-Online Long-horizon Assignment Reinforcement Learning
par: Wang, Jichao, et autres
Publié: (2026)
par: Wang, Jichao, et autres
Publié: (2026)
Exploring validation metrics for offline model-based optimisation with diffusion models
par: Beckham, Christopher, et autres
Publié: (2022)
par: Beckham, Christopher, et autres
Publié: (2022)
SPEED-RL: Faster Training of Reasoning Models via Online Curriculum Learning
par: Zhang, Ruiqi, et autres
Publié: (2025)
par: Zhang, Ruiqi, et autres
Publié: (2025)
RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
par: Dong, Yihong, et autres
Publié: (2025)
par: Dong, Yihong, et autres
Publié: (2025)
PokeRL: Reinforcement Learning for Pokemon Red
par: Mudireddy, Dheeraj, et autres
Publié: (2026)
par: Mudireddy, Dheeraj, et autres
Publié: (2026)
How to Get Your LLM to Generate Challenging Problems for Evaluation
par: Patel, Arkil, et autres
Publié: (2025)
par: Patel, Arkil, et autres
Publié: (2025)
FlowRL: A Taxonomy and Modular Framework for Reinforcement Learning with Diffusion Policies
par: Gao, Chenxiao, et autres
Publié: (2026)
par: Gao, Chenxiao, et autres
Publié: (2026)
Conditional Sequence Modeling for Safe Reinforcement Learning
par: Bai, Wensong, et autres
Publié: (2026)
par: Bai, Wensong, et autres
Publié: (2026)
RL$^3$: Boosting Meta Reinforcement Learning via RL inside RL$^2$
par: Bhatia, Abhinav, et autres
Publié: (2023)
par: Bhatia, Abhinav, et autres
Publié: (2023)
RL for Consistency Models: Faster Reward Guided Text-to-Image Generation
par: Oertell, Owen, et autres
Publié: (2024)
par: Oertell, Owen, et autres
Publié: (2024)
ObjectRL: An Object-Oriented Reinforcement Learning Codebase
par: Baykal, Gulcin, et autres
Publié: (2025)
par: Baykal, Gulcin, et autres
Publié: (2025)
Automation and Feature Selection Enhancement with Reinforcement Learning (RL)
par: Nagaraju, Sumana Sanyasipura
Publié: (2025)
par: Nagaraju, Sumana Sanyasipura
Publié: (2025)
QuRL: Efficient Reinforcement Learning with Quantized Rollout
par: Li, Yuhang, et autres
Publié: (2026)
par: Li, Yuhang, et autres
Publié: (2026)
Evaluating In-Context Learning of Libraries for Code Generation
par: Patel, Arkil, et autres
Publié: (2023)
par: Patel, Arkil, et autres
Publié: (2023)
Open RL Benchmark: Comprehensive Tracked Experiments for Reinforcement Learning
par: Huang, Shengyi, et autres
Publié: (2024)
par: Huang, Shengyi, et autres
Publié: (2024)
Faster Synchronous On-Policy RL via Straggler-Aware Group Sizing
par: Khan, Azal Ahmad, et autres
Publié: (2026)
par: Khan, Azal Ahmad, et autres
Publié: (2026)
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
par: Xia, Peng, et autres
Publié: (2026)
par: Xia, Peng, et autres
Publié: (2026)
D5RL: Diverse Datasets for Data-Driven Deep Reinforcement Learning
par: Rafailov, Rafael, et autres
Publié: (2024)
par: Rafailov, Rafael, et autres
Publié: (2024)
RL-ACRGNet: Reinforcement Learning-Based Chest Radiology Report Generation Network
par: Meena, Yogesh Kumar, et autres
Publié: (2026)
par: Meena, Yogesh Kumar, et autres
Publié: (2026)
Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewards
par: Zhang, Xin, et autres
Publié: (2026)
par: Zhang, Xin, et autres
Publié: (2026)
RL's Razor: Why Online Reinforcement Learning Forgets Less
par: Shenfeld, Idan, et autres
Publié: (2025)
par: Shenfeld, Idan, et autres
Publié: (2025)
RL as Regressor: A Reinforcement Learning Approach for Function Approximation
par: Huang, Yongchao
Publié: (2025)
par: Huang, Yongchao
Publié: (2025)
Survival Reinforcement Learning: Toward Scalable Self-Supervised RL
par: Nguimatsia-Tiofack, Franki, et autres
Publié: (2026)
par: Nguimatsia-Tiofack, Franki, et autres
Publié: (2026)
Probabilistic Learning and Generation in Deep Sequence Models
par: Chen, Wenlong
Publié: (2026)
par: Chen, Wenlong
Publié: (2026)
Reinformer: Max-Return Sequence Modeling for Offline RL
par: Zhuang, Zifeng, et autres
Publié: (2024)
par: Zhuang, Zifeng, et autres
Publié: (2024)
FairReweighing: Density Estimation-Based Reweighing Framework for Improving Separation in Fair Regression
par: Xi, Xiaoyin, et autres
Publié: (2025)
par: Xi, Xiaoyin, et autres
Publié: (2025)
Mini-Sequence Transformer: Optimizing Intermediate Memory for Long Sequences Training
par: Luo, Cheng, et autres
Publié: (2024)
par: Luo, Cheng, et autres
Publié: (2024)
Towards Sample-Efficiency and Generalization of Transfer and Inverse Reinforcement Learning: A Comprehensive Literature Review
par: Hassani, Hossein, et autres
Publié: (2024)
par: Hassani, Hossein, et autres
Publié: (2024)
MAGNNET: Multi-Agent Graph Neural Network-based Efficient Task Allocation for Autonomous Vehicles with Deep Reinforcement Learning
par: Ratnabala, Lavanya, et autres
Publié: (2025)
par: Ratnabala, Lavanya, et autres
Publié: (2025)
Documents similaires
-
Apriel-1.5-OpenReasoner: RL Post-Training for General-Purpose and Efficient Reasoning
par: Pardinas, Rafael, et autres
Publié: (2026) -
LLMs can learn self-restraint through iterative self-reflection
par: Piché, Alexandre, et autres
Publié: (2024) -
Self-Evolving Curriculum for LLM Reasoning
par: Chen, Xiaoyin, et autres
Publié: (2025) -
LongRecall: A Structured Approach for Robust Recall Evaluation in Long-Form Text
par: Ardestani, MohamamdJavad, et autres
Publié: (2025) -
Forecasting Downstream Performance of LLMs With Proxy Metrics
par: Patel, Arkil, et autres
Publié: (2026)