Precise: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zou, Jade, Huang, Tao, Kong, Weijie, Li, Junzhe, Wu, Yue, Tian, Qi, Xiong, Jiangfeng, Zhang, Jianwei, Bo, Liefeng, Zhong, Zhao |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025)
by: Li, Junzhe, et al.
Published: (2025)
OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
by: Pan, Kaihang, et al.
Published: (2026)
by: Pan, Kaihang, et al.
Published: (2026)
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
by: Tu, Shuyuan, et al.
Published: (2026)
by: Tu, Shuyuan, et al.
Published: (2026)
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
by: Bartosh, Grigory, et al.
Published: (2025)
by: Bartosh, Grigory, et al.
Published: (2025)
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
by: Bergmeister, Andreas, et al.
Published: (2026)
by: Bergmeister, Andreas, et al.
Published: (2026)
Flow Matching Posterior Sampling: A Training-free Conditional Generation for Flow Matching
by: Song, Kaiyu, et al.
Published: (2024)
by: Song, Kaiyu, et al.
Published: (2024)
SuperFlow: Training Flow Matching Models with RL on the Fly
by: Chen, Kaijie, et al.
Published: (2025)
by: Chen, Kaijie, et al.
Published: (2025)
Role-Based Fault Tolerance System for LLM RL Post-Training
by: Chen, Zhenqian, et al.
Published: (2025)
by: Chen, Zhenqian, et al.
Published: (2025)
Flow-GRPO: Training Flow Matching Models via Online RL
by: Liu, Jie, et al.
Published: (2025)
by: Liu, Jie, et al.
Published: (2025)
Neural Stochastic Flows: Solver-Free Modelling and Inference for SDE Solutions
by: Kiyohara, Naoki, et al.
Published: (2025)
by: Kiyohara, Naoki, et al.
Published: (2025)
EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions
by: Tian, Linrui, et al.
Published: (2024)
by: Tian, Linrui, et al.
Published: (2024)
FlowErase-RL: Rethinking Concept Erasure as Reward Optimization in Flow Matching Models
by: Sun, Yi, et al.
Published: (2026)
by: Sun, Yi, et al.
Published: (2026)
Small Generalizable Prompt Predictive Models Can Steer Efficient RL Post-Training of Large Reasoning Models
by: Qu, Yun, et al.
Published: (2026)
by: Qu, Yun, et al.
Published: (2026)
Integrating Local Precision With Global Consistency for Unsupervised Magnetic Resonance Image Registration
by: He Deng, et al.
Published: (2026)
by: He Deng, et al.
Published: (2026)
JigsawRL: Assembling RL Pipelines for Efficient LLM Post-Training
by: Hu, Zhengding, et al.
Published: (2026)
by: Hu, Zhengding, et al.
Published: (2026)
SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training
by: Wang, Chen, et al.
Published: (2025)
by: Wang, Chen, et al.
Published: (2025)
Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
Smart-GRPO: Smartly Sampling Noise for Efficient RL of Flow-Matching Models
by: Yu, Benjamin, et al.
Published: (2025)
by: Yu, Benjamin, et al.
Published: (2025)
Lagrangian Flow Matching: A Least-Action Framework for Principled Path Design
by: Du, Shukai, et al.
Published: (2026)
by: Du, Shukai, et al.
Published: (2026)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
by: Tian, Linrui, et al.
Published: (2025)
by: Tian, Linrui, et al.
Published: (2025)
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
by: Han, Zhenyu, et al.
Published: (2025)
by: Han, Zhenyu, et al.
Published: (2025)
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
by: McAllister, David, et al.
Published: (2026)
by: McAllister, David, et al.
Published: (2026)
Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?
by: Chen, Wanyi, et al.
Published: (2026)
by: Chen, Wanyi, et al.
Published: (2026)
SDE-Attention: Latent Attention in SDE-RNNs for Irregularly Sampled Time Series with Missing Data
by: Fang, Yuting, et al.
Published: (2025)
by: Fang, Yuting, et al.
Published: (2025)
Post-Training Language Models for Crosslingual Consistency
by: Liu, Tianyu, et al.
Published: (2026)
by: Liu, Tianyu, et al.
Published: (2026)
Consistency Flow Matching: Defining Straight Flows with Velocity Consistency
by: Yang, Ling, et al.
Published: (2024)
by: Yang, Ling, et al.
Published: (2024)
GSB: Group Superposition Binarization for Vision Transformer with Limited Training Samples
by: Gao, Tian, et al.
Published: (2023)
by: Gao, Tian, et al.
Published: (2023)
Manifold-Aware Exploration for Reinforcement Learning in Video Generation
by: Zheng, Mingzhe, et al.
Published: (2026)
by: Zheng, Mingzhe, et al.
Published: (2026)
Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching
by: Shi, Zhekun, et al.
Published: (2025)
by: Shi, Zhekun, et al.
Published: (2025)
floq: Training Critics via Flow-Matching for Scaling Compute in Value-Based RL
by: Agrawalla, Bhavya, et al.
Published: (2025)
by: Agrawalla, Bhavya, et al.
Published: (2025)
FlowRL: Matching Reward Distributions for LLM Reasoning
by: Zhu, Xuekai, et al.
Published: (2025)
by: Zhu, Xuekai, et al.
Published: (2025)
Stochastic Suspended Sediment Dynamics in Semi-Bounded Open Channel Flows: A Reflected SDE Approach
by: Kumbhakar, Manotosh, et al.
Published: (2024)
by: Kumbhakar, Manotosh, et al.
Published: (2024)
Training-Free Refinement of Flow Matching with Divergence-based Sampling
by: Cha, Yeonwoo, et al.
Published: (2026)
by: Cha, Yeonwoo, et al.
Published: (2026)
Balancing Signal and Variance: Adaptive Offline RL Post-Training for VLA Flow Models
by: Zhang, Hongyin, et al.
Published: (2025)
by: Zhang, Hongyin, et al.
Published: (2025)
Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling
by: Ma, Zhiyuan, et al.
Published: (2025)
by: Ma, Zhiyuan, et al.
Published: (2025)
LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
by: Liang, Zhanhao, et al.
Published: (2026)
by: Liang, Zhanhao, et al.
Published: (2026)
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
First SFT, Second RL, Third UPT: Continual Improving Multi-Modal LLM Reasoning via Unsupervised Post-Training
by: Wei, Lai, et al.
Published: (2025)
by: Wei, Lai, et al.
Published: (2025)
RollPacker: Mitigating Long-Tail Rollouts for Fast, Synchronous RL Post-Training
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
Similar Items
-
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
by: Li, Junzhe, et al.
Published: (2025) -
OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning
by: Pan, Kaihang, et al.
Published: (2026) -
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
by: Tu, Shuyuan, et al.
Published: (2026) -
SDE Matching: Scalable and Simulation-Free Training of Latent Stochastic Differential Equations
by: Bartosh, Grigory, et al.
Published: (2025) -
Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
by: Bergmeister, Andreas, et al.
Published: (2026)