JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Bingxiang, Qu, Zekai, Liu, Zeyuan, Chen, Yinghao, Zuo, Yuxin, Qian, Cheng, Zhang, Kaiyan, Chen, Weize, Xiao, Chaojun, Cui, Ganqu, Ding, Ning, Liu, Zhiyuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
von: Li, Yaxuan, et al.
Veröffentlicht: (2026)
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
von: Yuan, Lifan, et al.
Veröffentlicht: (2025)
von: Yuan, Lifan, et al.
Veröffentlicht: (2025)
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
von: Li, Haozhan, et al.
Veröffentlicht: (2025)
von: Li, Haozhan, et al.
Veröffentlicht: (2025)
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning
von: Li, Ran, et al.
Veröffentlicht: (2026)
von: Li, Ran, et al.
Veröffentlicht: (2026)
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
von: Zhu, Xuekai, et al.
Veröffentlicht: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2026)
How Far Can Unsupervised RLVR Scale LLM Training?
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
von: He, Bingxiang, et al.
Veröffentlicht: (2026)
Free Process Rewards without Process Labels
von: Yuan, Lifan, et al.
Veröffentlicht: (2024)
von: Yuan, Lifan, et al.
Veröffentlicht: (2024)
Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
von: Lv, Xingtai, et al.
Veröffentlicht: (2024)
Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoning
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
von: Cheng, Qianjia, et al.
Veröffentlicht: (2026)
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
ReviewRL: Towards Automated Scientific Review with RL
von: Zeng, Sihang, et al.
Veröffentlicht: (2025)
von: Zeng, Sihang, et al.
Veröffentlicht: (2025)
The Right Time Matters: Data Arrangement Affects Zero-Shot Generalization in Instruction Tuning
von: He, Bingxiang, et al.
Veröffentlicht: (2024)
von: He, Bingxiang, et al.
Veröffentlicht: (2024)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
von: Wang, Sudong, et al.
Veröffentlicht: (2026)
von: Wang, Sudong, et al.
Veröffentlicht: (2026)
INTERVENOR: Prompting the Coding Ability of Large Language Models with the Interactive Chain of Repair
von: Wang, Hanbin, et al.
Veröffentlicht: (2023)
von: Wang, Hanbin, et al.
Veröffentlicht: (2023)
STO-RL: Offline RL under Sparse Rewards via LLM-Guided Subgoal Temporal Order
von: Gu, Chengyang, et al.
Veröffentlicht: (2026)
von: Gu, Chengyang, et al.
Veröffentlicht: (2026)
SE-Bench: Benchmarking Self-Evolution with Knowledge Internalization
von: Yuan, Jiarui, et al.
Veröffentlicht: (2026)
von: Yuan, Jiarui, et al.
Veröffentlicht: (2026)
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
von: Yu, Tianyu, et al.
Veröffentlicht: (2025)
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
UltraFeedback: Boosting Language Models with Scaled AI Feedback
von: Cui, Ganqu, et al.
Veröffentlicht: (2023)
von: Cui, Ganqu, et al.
Veröffentlicht: (2023)
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
von: Cen, Zhepeng, et al.
Veröffentlicht: (2025)
von: Cen, Zhepeng, et al.
Veröffentlicht: (2025)
TTRL: Test-Time Reinforcement Learning
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
von: Zeng, Weihao, et al.
Veröffentlicht: (2025)
Accelerating PDE Surrogates via RL-Guided Mesh Optimization
von: Meng, Yang, et al.
Veröffentlicht: (2026)
von: Meng, Yang, et al.
Veröffentlicht: (2026)
AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset
von: He, Bingxiang, et al.
Veröffentlicht: (2025)
von: He, Bingxiang, et al.
Veröffentlicht: (2025)
DiffusionRL: Efficient Training of Diffusion Policies for Robotic Grasping Using RL-Adapted Large-Scale Datasets
von: Makarova, Maria, et al.
Veröffentlicht: (2025)
von: Makarova, Maria, et al.
Veröffentlicht: (2025)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Exposing Vulnerabilities in RL: A Novel Stealthy Backdoor Attack through Reward Poisoning
von: Zhang, Bokang, et al.
Veröffentlicht: (2025)
von: Zhang, Bokang, et al.
Veröffentlicht: (2025)
The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE Training
von: Chen, Weize, et al.
Veröffentlicht: (2025)
von: Chen, Weize, et al.
Veröffentlicht: (2025)
ToRL: Scaling Tool-Integrated RL
von: Li, Xuefeng, et al.
Veröffentlicht: (2025)
von: Li, Xuefeng, et al.
Veröffentlicht: (2025)
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
von: Feng, Xuelu, et al.
Veröffentlicht: (2025)
Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
von: Chen, Zihan, et al.
Veröffentlicht: (2025)
UltraMedical: Building Specialized Generalists in Biomedicine
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyan, et al.
Veröffentlicht: (2024)
Mastering Text, Code and Math Simultaneously via Fusing Highly Specialized Language Models
von: Ding, Ning, et al.
Veröffentlicht: (2024)
von: Ding, Ning, et al.
Veröffentlicht: (2024)
SeeUPO: Sequence-Level Agentic-RL with Convergence Guarantees
von: Hu, Tianyi, et al.
Veröffentlicht: (2026)
von: Hu, Tianyi, et al.
Veröffentlicht: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
von: Hu, Zhengding, et al.
Veröffentlicht: (2026)
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
von: Zhang, Qiang, et al.
Veröffentlicht: (2026)
von: Zhang, Qiang, et al.
Veröffentlicht: (2026)
Process Reinforcement through Implicit Rewards
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
von: Cui, Ganqu, et al.
Veröffentlicht: (2025)
GAP-RL: Grasps As Points for RL Towards Dynamic Object Grasping
von: Xie, Pengwei, et al.
Veröffentlicht: (2024)
von: Xie, Pengwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
von: Li, Yaxuan, et al.
Veröffentlicht: (2026) -
From $f(x)$ and $g(x)$ to $f(g(x))$: LLMs Learn New Skills in RL by Composing Old Ones
von: Yuan, Lifan, et al.
Veröffentlicht: (2025) -
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
von: Li, Haozhan, et al.
Veröffentlicht: (2025) -
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning
von: Li, Ran, et al.
Veröffentlicht: (2026) -
RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)