Truncated Rectified Flow Policy for Reinforcement Learning with One-Step Sampling
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Xubin, Yang, Yipeng, Li, Zhan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
por: Nguyen, Thanh, et al.
Publicado: (2025)
por: Nguyen, Thanh, et al.
Publicado: (2025)
One-Step Flow Policy Mirror Descent
por: Chen, Tianyi, et al.
Publicado: (2025)
por: Chen, Tianyi, et al.
Publicado: (2025)
Latent Policy Steering through One-Step Flow Policies
por: Im, Hokyun, et al.
Publicado: (2026)
por: Im, Hokyun, et al.
Publicado: (2026)
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
por: Li, Zeqiao, et al.
Publicado: (2026)
por: Li, Zeqiao, et al.
Publicado: (2026)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
por: Wang, Zeyuan, et al.
Publicado: (2025)
por: Wang, Zeyuan, et al.
Publicado: (2025)
Rectified Robust Policy Optimization for Model-Uncertain Constrained Reinforcement Learning without Strong Duality
por: Ma, Shaocong, et al.
Publicado: (2025)
por: Ma, Shaocong, et al.
Publicado: (2025)
Rectifying Regression in Reinforcement Learning
por: Ayoub, Alex, et al.
Publicado: (2025)
por: Ayoub, Alex, et al.
Publicado: (2025)
Elucidating Rectified Flow with Deterministic Sampler: Polynomial Discretization Complexity for Multi and One-step Models
por: Yang, Ruofeng, et al.
Publicado: (2025)
por: Yang, Ruofeng, et al.
Publicado: (2025)
PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement Learning
por: Yang, Shunpeng, et al.
Publicado: (2026)
por: Yang, Shunpeng, et al.
Publicado: (2026)
Delta Rectified Flow Sampling for Text-to-Image Editing
por: Beaudouin, Gaspard, et al.
Publicado: (2025)
por: Beaudouin, Gaspard, et al.
Publicado: (2025)
Order-Optimal Sample Complexity of Rectified Flows
por: Sahoo, Hari Krishna, et al.
Publicado: (2026)
por: Sahoo, Hari Krishna, et al.
Publicado: (2026)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
por: Zhang, Chubin, et al.
Publicado: (2025)
por: Zhang, Chubin, et al.
Publicado: (2025)
ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
por: Hou, Hongru, et al.
Publicado: (2026)
por: Hou, Hongru, et al.
Publicado: (2026)
SR$^2$-LoRA: Self-Rectifying Inter-layer Relations in Low-Rank Adaptation for Class-Incremental Learning
por: Wan, Fengqiang, et al.
Publicado: (2026)
por: Wan, Fengqiang, et al.
Publicado: (2026)
Appeal: Allow Mislabeled Samples the Chance to be Rectified in Partial Label Learning
por: Si, Chongjie, et al.
Publicado: (2023)
por: Si, Chongjie, et al.
Publicado: (2023)
Policy Bifurcation in Safe Reinforcement Learning
por: Zou, Wenjun, et al.
Publicado: (2024)
por: Zou, Wenjun, et al.
Publicado: (2024)
On-Policy Policy Gradient Reinforcement Learning Without On-Policy Sampling
por: Corrado, Nicholas E., et al.
Publicado: (2023)
por: Corrado, Nicholas E., et al.
Publicado: (2023)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
por: Wang, Zeyuan, et al.
Publicado: (2026)
por: Wang, Zeyuan, et al.
Publicado: (2026)
Reinforcement Learning for Flow-Matching Policies
por: Pfrommer, Samuel, et al.
Publicado: (2025)
por: Pfrommer, Samuel, et al.
Publicado: (2025)
Improving Rectified Flow with Boundary Conditions
por: Hu, Xixi, et al.
Publicado: (2025)
por: Hu, Xixi, et al.
Publicado: (2025)
One Step Learning, One Step Review
por: Huang, Xiaolong, et al.
Publicado: (2024)
por: Huang, Xiaolong, et al.
Publicado: (2024)
On the Convergence and Straightness of Rectified Flow
por: Bansal, Vansh, et al.
Publicado: (2024)
por: Bansal, Vansh, et al.
Publicado: (2024)
A Unified Framework for Data-Free One-Step Sampling via Wasserstein Gradient Flows
por: Wang, Chenguang, et al.
Publicado: (2026)
por: Wang, Chenguang, et al.
Publicado: (2026)
Enhancing Safety in Reinforcement Learning with Human Feedback via Rectified Policy Optimization
por: Peng, Xiyue, et al.
Publicado: (2024)
por: Peng, Xiyue, et al.
Publicado: (2024)
One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation
por: Wang, Zhendong, et al.
Publicado: (2024)
por: Wang, Zhendong, et al.
Publicado: (2024)
Policy Improvement Reinforcement Learning
por: Wang, Huaiyang, et al.
Publicado: (2026)
por: Wang, Huaiyang, et al.
Publicado: (2026)
Optimal Flow Matching: Learning Straight Trajectories in Just One Step
por: Kornilov, Nikita, et al.
Publicado: (2024)
por: Kornilov, Nikita, et al.
Publicado: (2024)
Let's Rectify Step by Step: Improving Aspect-based Sentiment Analysis with Diffusion Models
por: Liu, Shunyu, et al.
Publicado: (2024)
por: Liu, Shunyu, et al.
Publicado: (2024)
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
por: Koo, Juil, et al.
Publicado: (2026)
por: Koo, Juil, et al.
Publicado: (2026)
FlowTS: Time Series Generation via Rectified Flow
por: Hu, Yang, et al.
Publicado: (2024)
por: Hu, Yang, et al.
Publicado: (2024)
SAC Flow: Sample-Efficient Reinforcement Learning of Flow-Based Policies via Velocity-Reparameterized Sequential Modeling
por: Zhang, Yixian, et al.
Publicado: (2025)
por: Zhang, Yixian, et al.
Publicado: (2025)
Open the Black Box: Step-based Policy Updates for Temporally-Correlated Episodic Reinforcement Learning
por: Li, Ge, et al.
Publicado: (2024)
por: Li, Ge, et al.
Publicado: (2024)
Rectified Flows for Fast Multiscale Fluid Flow Modeling
por: Armegioiu, Victor, et al.
Publicado: (2025)
por: Armegioiu, Victor, et al.
Publicado: (2025)
Flow-Based Policy for Online Reinforcement Learning
por: Lv, Lei, et al.
Publicado: (2025)
por: Lv, Lei, et al.
Publicado: (2025)
Sample-Efficient Policy Constraint Offline Deep Reinforcement Learning based on Sample Filtering
por: Chen, Yuanhao, et al.
Publicado: (2025)
por: Chen, Yuanhao, et al.
Publicado: (2025)
Query-Policy Misalignment in Preference-Based Reinforcement Learning
por: Hu, Xiao, et al.
Publicado: (2023)
por: Hu, Xiao, et al.
Publicado: (2023)
Riemannian MeanFlow for One-Step Generation on Manifolds
por: Zhong, Zichen, et al.
Publicado: (2026)
por: Zhong, Zichen, et al.
Publicado: (2026)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
por: Hao, Ruijie, et al.
Publicado: (2026)
por: Hao, Ruijie, et al.
Publicado: (2026)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
por: Zhan, Guojian, et al.
Publicado: (2026)
por: Zhan, Guojian, et al.
Publicado: (2026)
Finite-Sample Analysis of Policy Evaluation for Robust Average Reward Reinforcement Learning
por: Xu, Yang, et al.
Publicado: (2025)
por: Xu, Yang, et al.
Publicado: (2025)
Ejemplares similares
-
One-Step Flow Q-Learning: Addressing the Diffusion Policy Bottleneck in Offline Reinforcement Learning
por: Nguyen, Thanh, et al.
Publicado: (2025) -
One-Step Flow Policy Mirror Descent
por: Chen, Tianyi, et al.
Publicado: (2025) -
Latent Policy Steering through One-Step Flow Policies
por: Im, Hokyun, et al.
Publicado: (2026) -
Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
por: Li, Zeqiao, et al.
Publicado: (2026) -
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
por: Wang, Zeyuan, et al.
Publicado: (2025)