Reinforcement Learning for Flow-Matching Policies
Fuente:
arXiv
Guardado en:
| Autores principales: | Pfrommer, Samuel, Huang, Yixiao, Sojoudi, Somayeh |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
por: Pfrommer, Samuel, et al.
Publicado: (2025)
por: Pfrommer, Samuel, et al.
Publicado: (2025)
Transport of Algebraic Structure to Latent Embeddings
por: Pfrommer, Samuel, et al.
Publicado: (2024)
por: Pfrommer, Samuel, et al.
Publicado: (2024)
Pausing Policy Learning in Non-stationary Reinforcement Learning
por: Lee, Hyunin, et al.
Publicado: (2024)
por: Lee, Hyunin, et al.
Publicado: (2024)
Reinforcement Learning via Value Gradient Flow
por: Xu, Haoran, et al.
Publicado: (2026)
por: Xu, Haoran, et al.
Publicado: (2026)
Approximately Gaussian Replicator Flows: Nonconvex Optimization as a Nash-Convergent Evolutionary Game
por: Anderson, Brendon G., et al.
Publicado: (2024)
por: Anderson, Brendon G., et al.
Publicado: (2024)
Transformers Provably Learn to Internalize Chain-of-Thought
por: Huang, Yixiao, et al.
Publicado: (2026)
por: Huang, Yixiao, et al.
Publicado: (2026)
OVERT: A Benchmark for Over-Refusal Evaluation on Text-to-Image Models
por: Cheng, Ziheng, et al.
Publicado: (2025)
por: Cheng, Ziheng, et al.
Publicado: (2025)
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
por: Cheng, Ziheng, et al.
Publicado: (2026)
por: Cheng, Ziheng, et al.
Publicado: (2026)
Infinite-Horizon Reach-Avoid Zero-Sum Games via Deep Reinforcement Learning
por: Li, Jingqi, et al.
Publicado: (2022)
por: Li, Jingqi, et al.
Publicado: (2022)
Evolving Diffusion and Flow Matching Policies for Online Reinforcement Learning
por: Zhang, Chubin, et al.
Publicado: (2025)
por: Zhang, Chubin, et al.
Publicado: (2025)
Ranking Manipulation for Conversational Search Engines
por: Pfrommer, Samuel, et al.
Publicado: (2024)
por: Pfrommer, Samuel, et al.
Publicado: (2024)
Do Sparse Autoencoders Identify Reasoning Features in Language Models?
por: Ma, George, et al.
Publicado: (2026)
por: Ma, George, et al.
Publicado: (2026)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
por: Huang, Yixiao, et al.
Publicado: (2025)
por: Huang, Yixiao, et al.
Publicado: (2025)
ReinFlow: Fine-tuning Flow Matching Policy with Online Reinforcement Learning
por: Zhang, Tonghe, et al.
Publicado: (2025)
por: Zhang, Tonghe, et al.
Publicado: (2025)
Mixing Classifiers to Alleviate the Accuracy-Robustness Trade-Off
por: Bai, Yatong, et al.
Publicado: (2023)
por: Bai, Yatong, et al.
Publicado: (2023)
Absence of spurious solutions far from ground truth: A low-rank analysis with high-order losses
por: Ma, Ziye, et al.
Publicado: (2024)
por: Ma, Ziye, et al.
Publicado: (2024)
FM-IRL: Flow-Matching for Reward Modeling and Policy Regularization in Reinforcement Learning
por: Wan, Zhenglin, et al.
Publicado: (2025)
por: Wan, Zhenglin, et al.
Publicado: (2025)
Towards Optimal Branching of Linear and Semidefinite Relaxations for Neural Network Robustness Certification
por: Anderson, Brendon G., et al.
Publicado: (2021)
por: Anderson, Brendon G., et al.
Publicado: (2021)
Flow Matching for Offline Reinforcement Learning with Discrete Actions
por: Khan, Fairoz Nower, et al.
Publicado: (2026)
por: Khan, Fairoz Nower, et al.
Publicado: (2026)
DRAGON: Distributional Rewards Optimize Diffusion Generative Models
por: Bai, Yatong, et al.
Publicado: (2025)
por: Bai, Yatong, et al.
Publicado: (2025)
Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies
por: Li, Zeyang, et al.
Publicado: (2026)
por: Li, Zeyang, et al.
Publicado: (2026)
Efficient Global Optimization of Two-Layer ReLU Networks: Quadratic-Time Algorithms and Adversarial Training
por: Bai, Yatong, et al.
Publicado: (2022)
por: Bai, Yatong, et al.
Publicado: (2022)
Energy-Weighted Flow Matching for Offline Reinforcement Learning
por: Zhang, Shiyuan, et al.
Publicado: (2025)
por: Zhang, Shiyuan, et al.
Publicado: (2025)
PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement Learning
por: Yang, Shunpeng, et al.
Publicado: (2026)
por: Yang, Shunpeng, et al.
Publicado: (2026)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
por: Bai, Yatong, et al.
Publicado: (2023)
por: Bai, Yatong, et al.
Publicado: (2023)
Riemannian Flow Matching Policy for Robot Motion Learning
por: Braun, Max, et al.
Publicado: (2024)
por: Braun, Max, et al.
Publicado: (2024)
Flow Matching Policy Gradients
por: McAllister, David, et al.
Publicado: (2025)
por: McAllister, David, et al.
Publicado: (2025)
Reinforcement Fine-Tuning of Flow-Matching Policies for Vision-Language-Action Models
por: Lyu, Mingyang, et al.
Publicado: (2025)
por: Lyu, Mingyang, et al.
Publicado: (2025)
Controllable Flow Matching for Online Reinforcement Learning
por: Wang, Bin, et al.
Publicado: (2025)
por: Wang, Bin, et al.
Publicado: (2025)
Improving the Accuracy-Robustness Trade-Off of Classifiers via Adaptive Smoothing
por: Bai, Yatong, et al.
Publicado: (2023)
por: Bai, Yatong, et al.
Publicado: (2023)
MixedNUTS: Training-Free Accuracy-Robustness Balance via Nonlinearly Mixed Classifiers
por: Bai, Yatong, et al.
Publicado: (2024)
por: Bai, Yatong, et al.
Publicado: (2024)
Quantile-Coupled Flow Matching for Distributional Reinforcement Learning
por: Groom, Michael, et al.
Publicado: (2026)
por: Groom, Michael, et al.
Publicado: (2026)
FlowCritic: Bridging Value Estimation with Flow Matching in Reinforcement Learning
por: Zhong, Shan, et al.
Publicado: (2025)
por: Zhong, Shan, et al.
Publicado: (2025)
Online Matching via Reinforcement Learning: An Expert Policy Orchestration Strategy
por: Mignacco, Chiara, et al.
Publicado: (2025)
por: Mignacco, Chiara, et al.
Publicado: (2025)
Discrete Flow Matching for Offline-to-Online Reinforcement Learning
por: Khan, Fairoz Nower, et al.
Publicado: (2026)
por: Khan, Fairoz Nower, et al.
Publicado: (2026)
Flow-Based Policy for Online Reinforcement Learning
por: Lv, Lei, et al.
Publicado: (2025)
por: Lv, Lei, et al.
Publicado: (2025)
Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
por: Kong, Lingkai, et al.
Publicado: (2026)
por: Kong, Lingkai, et al.
Publicado: (2026)
Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
por: Kong, Lingkai, et al.
Publicado: (2025)
por: Kong, Lingkai, et al.
Publicado: (2025)
Flow Matching with Injected Noise for Offline-to-Online Reinforcement Learning
por: Shin, Yongjae, et al.
Publicado: (2026)
por: Shin, Yongjae, et al.
Publicado: (2026)
Max-Entropy Reinforcement Learning with Flow Matching and A Case Study on LQR
por: Zhang, Yuyang, et al.
Publicado: (2025)
por: Zhang, Yuyang, et al.
Publicado: (2025)
Ejemplares similares
-
Spooky Action at a Distance: Normalization Layers Enable Side-Channel Spatial Communication
por: Pfrommer, Samuel, et al.
Publicado: (2025) -
Transport of Algebraic Structure to Latent Embeddings
por: Pfrommer, Samuel, et al.
Publicado: (2024) -
Pausing Policy Learning in Non-stationary Reinforcement Learning
por: Lee, Hyunin, et al.
Publicado: (2024) -
Reinforcement Learning via Value Gradient Flow
por: Xu, Haoran, et al.
Publicado: (2026) -
Approximately Gaussian Replicator Flows: Nonconvex Optimization as a Nash-Convergent Evolutionary Game
por: Anderson, Brendon G., et al.
Publicado: (2024)