FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wan, Zhenglin, Wu, Jingxuan, Yu, Xingrui, Zhang, Chubin, Lei, Mingcong, An, Bo, Tsang, Ivor |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adversarial Dual On-Policy Distillation from Expressive Teacher
by: Wan, Zhenglin, et al.
Published: (2026)
by: Wan, Zhenglin, et al.
Published: (2026)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025)
by: Zhang, Chubin, et al.
Published: (2025)
Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
by: Wu, Jingxuan, et al.
Published: (2025)
by: Wu, Jingxuan, et al.
Published: (2025)
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
by: Yan, Xin, et al.
Published: (2026)
by: Yan, Xin, et al.
Published: (2026)
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)
by: Yu, Xingrui, et al.
Published: (2024)
Diversifying Policy Behaviors with Extrinsic Behavioral Curiosity
by: Wan, Zhenglin, et al.
Published: (2024)
by: Wan, Zhenglin, et al.
Published: (2024)
Flow-Factory: A Unified Framework for Reinforcement Learning in Flow-Matching Models
by: Ping, Bowen, et al.
Published: (2026)
by: Ping, Bowen, et al.
Published: (2026)
Advancing Analytic Class-Incremental Learning through Vision-Language Calibration
by: Zhao, Binyu, et al.
Published: (2026)
by: Zhao, Binyu, et al.
Published: (2026)
SA-VLA: Spatially-Aware Flow-Matching for Vision-Language-Action Reinforcement Learning
by: Pan, Xu, et al.
Published: (2026)
by: Pan, Xu, et al.
Published: (2026)
Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Models
by: Wu, Jingxuan, et al.
Published: (2026)
by: Wu, Jingxuan, et al.
Published: (2026)
Beyond-Expert Performance with Limited Demonstrations: Efficient Imitation Learning with Double Exploration
by: Zhao, Heyang, et al.
Published: (2025)
by: Zhao, Heyang, et al.
Published: (2025)
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
by: Zhang, Tonghe, et al.
Published: (2025)
by: Zhang, Tonghe, et al.
Published: (2025)
Reinforcement Learning for Flow-Matching Policies
by: Pfrommer, Samuel, et al.
Published: (2025)
by: Pfrommer, Samuel, et al.
Published: (2025)
Mitigating Mismatch within Reference-based Preference Optimization
by: Yuan, Suqin, et al.
Published: (2026)
by: Yuan, Suqin, et al.
Published: (2026)
Scouting By Reward: VLM-TO-IRL-Driven Player Selection For Esports
by: Yan, Qing, et al.
Published: (2026)
by: Yan, Qing, et al.
Published: (2026)
Translate Policy to Language: Flow Matching Generated Rewards for LLM Explanations
by: Yang, Xinyi, et al.
Published: (2025)
by: Yang, Xinyi, et al.
Published: (2025)
CoMI-IRL: Contrastive Multi-Intention Inverse Reinforcement Learning
by: Mone, Antonio, et al.
Published: (2026)
by: Mone, Antonio, et al.
Published: (2026)
ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
by: Castanyer, Roger Creus, et al.
Published: (2025)
by: Castanyer, Roger Creus, et al.
Published: (2025)
Transductive Reward Inference on Graph
by: Qu, Bohao, et al.
Published: (2024)
by: Qu, Bohao, et al.
Published: (2024)
Flow-Direct: Feedback-Efficient and Reusable Guidance for Flow Models via Non-Parametric Guidance Field
by: Tan, Kim Yong, et al.
Published: (2026)
by: Tan, Kim Yong, et al.
Published: (2026)
Do You Trust the Process?: Modeling Institutional Trust for Community Adoption of Reinforcement Learning Policies
by: Balepur, Naina, et al.
Published: (2025)
by: Balepur, Naina, et al.
Published: (2025)
Zeroth-Order Policy Gradient for Reinforcement Learning from Human Feedback without Reward Inference
by: Zhang, Qining, et al.
Published: (2024)
by: Zhang, Qining, et al.
Published: (2024)
SurgIRL: Towards Life-Long Learning for Surgical Automation by Incremental Reinforcement Learning
by: Ho, Yun-Jie, et al.
Published: (2024)
by: Ho, Yun-Jie, et al.
Published: (2024)
OAT-FM: Optimal Acceleration Transport for Improved Flow Matching
by: Yue, Angxiao, et al.
Published: (2025)
by: Yue, Angxiao, et al.
Published: (2025)
Hierarchical Spatial-Temporal Graph-Enhanced Model for Map-Matching
by: Gao, Anjun, et al.
Published: (2026)
by: Gao, Anjun, et al.
Published: (2026)
Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR
by: Zhang, Yuyang, et al.
Published: (2025)
by: Zhang, Yuyang, et al.
Published: (2025)
Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning
by: Gao, Chen-Xiao, et al.
Published: (2025)
by: Gao, Chen-Xiao, et al.
Published: (2025)
TreeIRL: Safe Urban Driving with Tree Search and Inverse Reinforcement Learning
by: Tomov, Momchil S., et al.
Published: (2025)
by: Tomov, Momchil S., et al.
Published: (2025)
Reward Models in Deep Reinforcement Learning: A Survey
by: Yu, Rui, et al.
Published: (2025)
by: Yu, Rui, et al.
Published: (2025)
Flow-Based Policy for Online Reinforcement Learning
by: Lv, Lei, et al.
Published: (2025)
by: Lv, Lei, et al.
Published: (2025)
Energy-Weighted Flow Matching for Offline Reinforcement Learning
by: Zhang, Shiyuan, et al.
Published: (2025)
by: Zhang, Shiyuan, et al.
Published: (2025)
DeepStock: Reinforcement Learning with Policy Regularizations for Inventory Management
by: Xie, Yaqi, et al.
Published: (2026)
by: Xie, Yaqi, et al.
Published: (2026)
Learning ORDER-Aware Multimodal Representations for Composite Materials Design
by: Li, Xinyao, et al.
Published: (2026)
by: Li, Xinyao, et al.
Published: (2026)
Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization
by: Yu, Xin, et al.
Published: (2026)
by: Yu, Xin, et al.
Published: (2026)
Deep Reinforcement Learning with Hybrid Intrinsic Reward Model
by: Yuan, Mingqi, et al.
Published: (2025)
by: Yuan, Mingqi, et al.
Published: (2025)
Mastering Continual Reinforcement Learning through Fine-Grained Sparse Network Allocation and Dormant Neuron Exploration
by: Zheng, Chengqi, et al.
Published: (2025)
by: Zheng, Chengqi, et al.
Published: (2025)
The Propagation Field: A Geometric Substrate Theory of Deep Learning
by: Gu, Xingrui
Published: (2026)
by: Gu, Xingrui
Published: (2026)
Conceptual Belief-Informed Reinforcement Learning
by: Gu, Xingrui, et al.
Published: (2024)
by: Gu, Xingrui, et al.
Published: (2024)
Uncertainty-Penalized Reinforcement Learning from Human Feedback with Diverse Reward LoRA Ensembles
by: Zhai, Yuanzhao, et al.
Published: (2023)
by: Zhai, Yuanzhao, et al.
Published: (2023)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
by: Lyu, Mingyang, et al.
Published: (2025)
by: Lyu, Mingyang, et al.
Published: (2025)
Similar Items
-
Adversarial Dual On-Policy Distillation from Expressive Teacher
by: Wan, Zhenglin, et al.
Published: (2026) -
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
by: Zhang, Chubin, et al.
Published: (2025) -
Letting Trajectories Spread: Quality-Preserving Control for Diverse Flow Matching
by: Wu, Jingxuan, et al.
Published: (2025) -
HBVLA: Pushing 1-Bit Post-Training Quantization for Vision-Language-Action Models
by: Yan, Xin, et al.
Published: (2026) -
Imitation from Diverse Behaviors: Wasserstein Quality Diversity Imitation Learning with Single-Step Archive Exploration
by: Yu, Xingrui, et al.
Published: (2024)