Boosting Maximum Entropy Reinforcement Learning via One-Step Flow Matching
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zeqiao, Wang, Yijing, Wang, Haoyu, Li, Zheng, Zuo, Zhiqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
From Learning to Mastery: Achieving Safe and Efficient Real-World Autonomous Driving with Human-In-The-Loop Reinforcement Learning
von: Zeqiao, Li, et al.
Veröffentlicht: (2025)
von: Zeqiao, Li, et al.
Veröffentlicht: (2025)
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving
von: Zeqiao, Li, et al.
Veröffentlicht: (2025)
von: Zeqiao, Li, et al.
Veröffentlicht: (2025)
MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation
von: Fu, Yuxiang, et al.
Veröffentlicht: (2025)
von: Fu, Yuxiang, et al.
Veröffentlicht: (2025)
Reward-Punishment Reinforcement Learning with Maximum Entropy
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
von: Wang, Jiexin, et al.
Veröffentlicht: (2024)
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
von: Lv, Lei, et al.
Veröffentlicht: (2026)
von: Lv, Lei, et al.
Veröffentlicht: (2026)
Scalable, Explainable and Provably Robust Anomaly Detection with One-Step Flow Matching
von: Li, Zhong, et al.
Veröffentlicht: (2025)
von: Li, Zhong, et al.
Veröffentlicht: (2025)
Maximum Entropy Reinforcement Learning with Diffusion Policy
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Dong, Xiaoyi, et al.
Veröffentlicht: (2025)
Controllable Flow Matching for Online Reinforcement Learning
von: Wang, Bin, et al.
Veröffentlicht: (2025)
von: Wang, Bin, et al.
Veröffentlicht: (2025)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
von: Kong, Lingkai, et al.
Veröffentlicht: (2025)
von: Kong, Lingkai, et al.
Veröffentlicht: (2025)
Mind Your Entropy: From Maximum Entropy to Trajectory Entropy-Constrained RL
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
von: Zhan, Guojian, et al.
Veröffentlicht: (2025)
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
Entropy-Regularized Adjoint Matching for Offline Reinforcement Learning
von: Ghanem, Abdelghani, et al.
Veröffentlicht: (2026)
von: Ghanem, Abdelghani, et al.
Veröffentlicht: (2026)
Diffusion-Augmented Markov Decision Processes for Maximum Entropy Reinforcement Learning
von: Sanokowski, Sebastian, et al.
Veröffentlicht: (2025)
von: Sanokowski, Sebastian, et al.
Veröffentlicht: (2025)
Path-Guided Flow Matching for Dataset Distillation
von: Li, Xuhui, et al.
Veröffentlicht: (2026)
von: Li, Xuhui, et al.
Veröffentlicht: (2026)
ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy
von: Li, Hongming, et al.
Veröffentlicht: (2024)
von: Li, Hongming, et al.
Veröffentlicht: (2024)
Maximum Entropy Inverse Reinforcement Learning of Diffusion Models with Energy-Based Models
von: Yoon, Sangwoong, et al.
Veröffentlicht: (2024)
von: Yoon, Sangwoong, et al.
Veröffentlicht: (2024)
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
von: Hu, Jiajun, et al.
Veröffentlicht: (2026)
von: Hu, Jiajun, et al.
Veröffentlicht: (2026)
Discrete Flow Matching for Offline-to-Online Reinforcement Learning
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
Mastering Massive Multi-Task Reinforcement Learning via Mixture-of-Expert Decision Transformer
von: Kong, Yilun, et al.
Veröffentlicht: (2025)
von: Kong, Yilun, et al.
Veröffentlicht: (2025)
Maximum Entropy On-Policy Actor-Critic via Entropy Advantage Estimation
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024)
von: Choe, Jean Seong Bjorn, et al.
Veröffentlicht: (2024)
Discrete MeanFlow: One-Step Generation via Conditional Transition Kernels
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
One-Step Graph-Structured Neural Flows for Irregular Multivariate Time Series Classification
von: Gao, Mengzhou, et al.
Veröffentlicht: (2026)
von: Gao, Mengzhou, et al.
Veröffentlicht: (2026)
Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
von: Shin, Yongjae, et al.
Veröffentlicht: (2026)
von: Shin, Yongjae, et al.
Veröffentlicht: (2026)
E-GRPO: High Entropy Steps Drive Effective Reinforcement Learning for Flow Models
von: Zhang, Shengjun, et al.
Veröffentlicht: (2026)
von: Zhang, Shengjun, et al.
Veröffentlicht: (2026)
Learning to Embed Distributions via Maximum Kernel Entropy
von: Kachaiev, Oleksii, et al.
Veröffentlicht: (2024)
von: Kachaiev, Oleksii, et al.
Veröffentlicht: (2024)
Deriving the Scaled-Dot-Function via Maximum Likelihood Estimation and Maximum Entropy Approach
von: Ma, Jiyong
Veröffentlicht: (2025)
von: Ma, Jiyong
Veröffentlicht: (2025)
Rethinking Plasticity in Deep Reinforcement Learning
von: He, Zhiqiang
Veröffentlicht: (2026)
von: He, Zhiqiang
Veröffentlicht: (2026)
Escaping Optimization Stagnation: Taking Steps Beyond Task Arithmetic via Difference Vectors
von: Wang, Jinping, et al.
Veröffentlicht: (2025)
von: Wang, Jinping, et al.
Veröffentlicht: (2025)
Learning Unbiased Permutations via Flow Matching
von: Min, Yimeng, et al.
Veröffentlicht: (2026)
von: Min, Yimeng, et al.
Veröffentlicht: (2026)
Epigraph-Guided Flow Matching for Safe and Performant Offline Reinforcement Learning
von: Tayal, Manan, et al.
Veröffentlicht: (2026)
von: Tayal, Manan, et al.
Veröffentlicht: (2026)
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks
von: Wang, Tianlong, et al.
Veröffentlicht: (2024)
von: Wang, Tianlong, et al.
Veröffentlicht: (2024)
Iterated $Q$-Network: Beyond One-Step Bellman Updates in Deep Reinforcement Learning
von: Vincent, Théo, et al.
Veröffentlicht: (2024)
von: Vincent, Théo, et al.
Veröffentlicht: (2024)
Approximate Subgraph Matching with Neural Graph Representations and Reinforcement Learning
von: Li, Kaiyang, et al.
Veröffentlicht: (2026)
von: Li, Kaiyang, et al.
Veröffentlicht: (2026)
Diffusion World Model: Future Modeling Beyond Step-by-Step Rollout for Offline Reinforcement Learning
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
von: Ding, Zihan, et al.
Veröffentlicht: (2024)
Compose Yourself: Average-Velocity Flow Matching for One-Step Speech Enhancement
von: Yang, Gang, et al.
Veröffentlicht: (2025)
von: Yang, Gang, et al.
Veröffentlicht: (2025)
TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesis
von: Bai, Sikai, et al.
Veröffentlicht: (2026)
von: Bai, Sikai, et al.
Veröffentlicht: (2026)
Boosting Deductive Reasoning with Step Signals In RLHF
von: Li, Jialian, et al.
Veröffentlicht: (2024)
von: Li, Jialian, et al.
Veröffentlicht: (2024)
WFR-MFM: One-Step Inference for Dynamic Unbalanced Optimal Transport
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
From Learning to Mastery: Achieving Safe and Efficient Real-World Autonomous Driving with Human-In-The-Loop Reinforcement Learning
von: Zeqiao, Li, et al.
Veröffentlicht: (2025) -
Confidence-Guided Human-AI Collaboration: Reinforcement Learning with Distributional Proxy Value Propagation for Autonomous Driving
von: Zeqiao, Li, et al.
Veröffentlicht: (2025) -
MoFlow: One-Step Flow Matching for Human Trajectory Forecasting via Implicit Maximum Likelihood Estimation based Distillation
von: Fu, Yuxiang, et al.
Veröffentlicht: (2025) -
Reward-Punishment Reinforcement Learning with Maximum Entropy
von: Wang, Jiexin, et al.
Veröffentlicht: (2024) -
FLAC: Maximum Entropy RL via Kinetic Energy Regularized Bridge Matching
von: Lv, Lei, et al.
Veröffentlicht: (2026)