Q-Flow: Stable and Expressive Reinforcement Learning with Flow-Based Policy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Doo, JaeHyeok, Jeon, Byeongguk, Ye, Seonghyeon, Lee, Kimin, Seo, Minjoon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
EXPO: Stable Reinforcement Learning with Expressive Policies
von: Dong, Perry, et al.
Veröffentlicht: (2025)
von: Dong, Perry, et al.
Veröffentlicht: (2025)
FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning
von: Alles, Marvin, et al.
Veröffentlicht: (2025)
von: Alles, Marvin, et al.
Veröffentlicht: (2025)
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
von: Kang, Sinjae, et al.
Veröffentlicht: (2026)
von: Kang, Sinjae, et al.
Veröffentlicht: (2026)
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
von: Jeon, WooJae, et al.
Veröffentlicht: (2024)
von: Jeon, WooJae, et al.
Veröffentlicht: (2024)
Flow-Based Policy for Online Reinforcement Learning
von: Lv, Lei, et al.
Veröffentlicht: (2025)
von: Lv, Lei, et al.
Veröffentlicht: (2025)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
Flow Q-Learning
von: Park, Seohong, et al.
Veröffentlicht: (2025)
von: Park, Seohong, et al.
Veröffentlicht: (2025)
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
von: Li, Mingxuan, et al.
Veröffentlicht: (2026)
von: Li, Mingxuan, et al.
Veröffentlicht: (2026)
AdaFlow: Imitation Learning with Variance-Adaptive Flow-Based Policies
von: Hu, Xixi, et al.
Veröffentlicht: (2024)
von: Hu, Xixi, et al.
Veröffentlicht: (2024)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
von: Hao, Ruijie, et al.
Veröffentlicht: (2026)
Towards Interpretable Reinforcement Learning with Constrained Normalizing Flow Policies
von: Rietz, Finn, et al.
Veröffentlicht: (2024)
von: Rietz, Finn, et al.
Veröffentlicht: (2024)
Rethinking the Role of Proxy Rewards in Language Model Alignment
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
von: Kim, Sungdong, et al.
Veröffentlicht: (2024)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning with Penalized Action Noise Injection
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025)
von: Oh, JunHyeok, et al.
Veröffentlicht: (2025)
Adversarial Reinforcement Learning Framework for ESP Cheater Simulation
von: Park, Inkyu, et al.
Veröffentlicht: (2025)
von: Park, Inkyu, et al.
Veröffentlicht: (2025)
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
von: Tiofack, Franki Nguimatsia, et al.
Veröffentlicht: (2025)
von: Tiofack, Franki Nguimatsia, et al.
Veröffentlicht: (2025)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
Score-Based One-step MeanFlow Policy Optimization
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026)
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026)
Aligning Large Language Models by On-Policy Self-Judgment
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
von: Lee, Sangkyu, et al.
Veröffentlicht: (2024)
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
von: Zhong, Shan, et al.
Veröffentlicht: (2025)
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
von: Won, Yunjae, et al.
Veröffentlicht: (2025)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
von: Seo, Younggyo, et al.
Veröffentlicht: (2024)
von: Seo, Younggyo, et al.
Veröffentlicht: (2024)
Flow Actor-Critic for Offline Reinforcement Learning
von: Chae, Jongseong, et al.
Veröffentlicht: (2026)
von: Chae, Jongseong, et al.
Veröffentlicht: (2026)
Reinforcement Learning via Value Gradient Flow
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
von: Xu, Haoran, et al.
Veröffentlicht: (2026)
Controllable Flow Matching for Online Reinforcement Learning
von: Wang, Bin, et al.
Veröffentlicht: (2025)
von: Wang, Bin, et al.
Veröffentlicht: (2025)
FlowPG: Action-constrained Policy Gradient with Normalizing Flows
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2024)
von: Brahmanage, Janaka Chathuranga, et al.
Veröffentlicht: (2024)
Reparameterization Flow Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
Decision Flow Policy Optimization
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
Model-based Offline Reinforcement Learning with Lower Expectile Q-Learning
von: Park, Kwanyoung, et al.
Veröffentlicht: (2024)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2024)
Discrete Flow Matching for Offline-to-Online Reinforcement Learning
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
von: Khan, Fairoz Nower, et al.
Veröffentlicht: (2026)
Learning More Expressive General Policies for Classical Planning Domains
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2024)
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2024)
Q-Policy: Quantum-Enhanced Policy Evaluation for Scalable Reinforcement Learning
von: Cherukuri, Kalyan, et al.
Veröffentlicht: (2025)
von: Cherukuri, Kalyan, et al.
Veröffentlicht: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
von: Lee, Hayeong, et al.
Veröffentlicht: (2026)
von: Lee, Hayeong, et al.
Veröffentlicht: (2026)
Exclusively Penalized Q-learning for Offline Reinforcement Learning
von: Yeom, Junghyuk, et al.
Veröffentlicht: (2024)
von: Yeom, Junghyuk, et al.
Veröffentlicht: (2024)
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs
von: Zheng, Haizhong, et al.
Veröffentlicht: (2026)
von: Zheng, Haizhong, et al.
Veröffentlicht: (2026)
Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
von: Shin, Yongjae, et al.
Veröffentlicht: (2026)
von: Shin, Yongjae, et al.
Veröffentlicht: (2026)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
von: Kong, Lingkai, et al.
Veröffentlicht: (2025)
von: Kong, Lingkai, et al.
Veröffentlicht: (2025)
SPQR: Controlling Q-ensemble Independence with Spiked Random Model for Reinforcement Learning
von: Lee, Dohyeok, et al.
Veröffentlicht: (2024)
von: Lee, Dohyeok, et al.
Veröffentlicht: (2024)
GEPO: Group Expectation Policy Optimization for Stable Heterogeneous Reinforcement Learning
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026) -
EXPO: Stable Reinforcement Learning with Expressive Policies
von: Dong, Perry, et al.
Veröffentlicht: (2025) -
FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning
von: Alles, Marvin, et al.
Veröffentlicht: (2025) -
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
von: Kang, Sinjae, et al.
Veröffentlicht: (2026) -
PG-Rainbow: Using Distributional Reinforcement Learning in Policy Gradient Methods
von: Jeon, WooJae, et al.
Veröffentlicht: (2024)