FP4 Explore, BF16 Train: Diffusion Reinforcement Learning via Efficient Rollout Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yitong, Chen, Junsong, Xue, Shuchen, Zeren, Pengcuo, Fu, Siyuan, Yang, Dinghao, Tang, Yangyang, Bai, Junjie, Luo, Ping, Han, Song, Xie, Enze |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
by: Fujii, Kazuki, et al.
Published: (2024)
by: Fujii, Kazuki, et al.
Published: (2024)
SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
by: Chen, Junsong, et al.
Published: (2025)
by: Chen, Junsong, et al.
Published: (2025)
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
by: Xi, Haocheng, et al.
Published: (2026)
by: Xi, Haocheng, et al.
Published: (2026)
Accelerating Diffusion Sampling with Optimized Time Steps
by: Xue, Shuchen, et al.
Published: (2024)
by: Xue, Shuchen, et al.
Published: (2024)
Fast-dLLM v2: Efficient Block-Diffusion LLM
by: Wu, Chengyue, et al.
Published: (2025)
by: Wu, Chengyue, et al.
Published: (2025)
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
PixArt-$α$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
by: Chen, Junsong, et al.
Published: (2023)
by: Chen, Junsong, et al.
Published: (2023)
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer
by: Zhu, Haoyi, et al.
Published: (2026)
by: Zhu, Haoyi, et al.
Published: (2026)
Defeating the Training-Inference Mismatch via FP16
by: Qi, Penghui, et al.
Published: (2025)
by: Qi, Penghui, et al.
Published: (2025)
Learning Single Index Models with Diffusion Priors
by: Tang, Anqi, et al.
Published: (2025)
by: Tang, Anqi, et al.
Published: (2025)
Data-regularized Reinforcement Learning for Diffusion Models at Scale
by: Ye, Haotian, et al.
Published: (2025)
by: Ye, Haotian, et al.
Published: (2025)
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
by: Xie, Enze, et al.
Published: (2025)
by: Xie, Enze, et al.
Published: (2025)
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
by: Liu, Akide, et al.
Published: (2025)
by: Liu, Akide, et al.
Published: (2025)
DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent Space
by: Chen, Junyu, et al.
Published: (2025)
by: Chen, Junyu, et al.
Published: (2025)
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
by: Chen, Junyu, et al.
Published: (2024)
by: Chen, Junyu, et al.
Published: (2024)
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
by: Qiu, Zhaopeng, et al.
Published: (2026)
by: Qiu, Zhaopeng, et al.
Published: (2026)
Noise Consistency Training: A Native Approach for One-Step Generator in Learning Additional Controls
by: Luo, Yihong, et al.
Published: (2025)
by: Luo, Yihong, et al.
Published: (2025)
Towards Fully FP8 GEMM LLM Training at Scale
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
by: Hernández-Cano, Alejandro, et al.
Published: (2025)
SGEMM-cube: Precision-Recovery FP32 GEMM Approximation on Ascend NPUs with FP16 Matrix Engines
by: Xue, Weicheng, et al.
Published: (2025)
by: Xue, Weicheng, et al.
Published: (2025)
FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs
by: Dege, Pengcuo, et al.
Published: (2025)
by: Dege, Pengcuo, et al.
Published: (2025)
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
by: Xu, Yixuan Even, et al.
Published: (2025)
by: Xu, Yixuan Even, et al.
Published: (2025)
Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
by: Xing, Shangyu, et al.
Published: (2025)
by: Xing, Shangyu, et al.
Published: (2025)
$μ$nit Scaling: Simple and Scalable FP8 LLM Training
by: Narayan, Saaketh, et al.
Published: (2025)
by: Narayan, Saaketh, et al.
Published: (2025)
Each Prompt Matters: Scaling Reinforcement Learning Without Wasting Rollouts on Hundred-Billion-Scale MoE
by: Zeng, Anxiang, et al.
Published: (2025)
by: Zeng, Anxiang, et al.
Published: (2025)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models
by: Chen, Junsong, et al.
Published: (2024)
by: Chen, Junsong, et al.
Published: (2024)
Leveraging Error Diversity in Group Rollouts for Reinforcement Learning
by: Liu, Wenpu, et al.
Published: (2026)
by: Liu, Wenpu, et al.
Published: (2026)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
by: Yoo, Seungwoo, et al.
Published: (2026)
by: Yoo, Seungwoo, et al.
Published: (2026)
DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space
by: He, Wenkun, et al.
Published: (2025)
by: He, Wenkun, et al.
Published: (2025)
On Rollouts in Model-Based Reinforcement Learning
by: Frauenknecht, Bernd, et al.
Published: (2025)
by: Frauenknecht, Bernd, et al.
Published: (2025)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
by: Xie, Enze, et al.
Published: (2024)
by: Xie, Enze, et al.
Published: (2024)
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
by: Zhao, Yuyang, et al.
Published: (2026)
by: Zhao, Yuyang, et al.
Published: (2026)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
by: Zhang, Yu, et al.
Published: (2025)
by: Zhang, Yu, et al.
Published: (2025)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
by: Bai, Yuyang, et al.
Published: (2026)
by: Bai, Yuyang, et al.
Published: (2026)
Rollout-Training Co-Design for Efficient LLM-Based Multi-Agent Reinforcement Learning
by: Jiang, Zhida, et al.
Published: (2026)
by: Jiang, Zhida, et al.
Published: (2026)
Tape: A Cellular Automata Benchmark for Evaluating Rule-Shift Generalization in Reinforcement Learning
by: Pan, Enze
Published: (2026)
by: Pan, Enze
Published: (2026)
DyDiff: Long-Horizon Rollout via Dynamics Diffusion for Offline Reinforcement Learning
by: Zhao, Hanye, et al.
Published: (2024)
by: Zhao, Hanye, et al.
Published: (2024)
Guided Cooperation in Hierarchical Reinforcement Learning via Model-based Rollout
by: Wang, Haoran, et al.
Published: (2023)
by: Wang, Haoran, et al.
Published: (2023)
Towards Efficient Pre-training: Exploring FP4 Precision in Large Language Models
by: Zhou, Jiecheng, et al.
Published: (2025)
by: Zhou, Jiecheng, et al.
Published: (2025)
Model-Driven Subspaces for Large-Scale Optimization with Local Approximation Strategy
by: He, Yitong, et al.
Published: (2025)
by: He, Yitong, et al.
Published: (2025)
Similar Items
-
Balancing Speed and Stability: The Trade-offs of FP8 vs. BF16 Training in LLMs
by: Fujii, Kazuki, et al.
Published: (2024) -
SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency Distillation
by: Chen, Junsong, et al.
Published: (2025) -
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
by: Wu, Chengyue, et al.
Published: (2025) -
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
by: Xi, Haocheng, et al.
Published: (2026) -
Accelerating Diffusion Sampling with Optimized Time Steps
by: Xue, Shuchen, et al.
Published: (2024)