OM2P: Offline Multi-Agent Mean-Flow Policy
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Zhuoran, Wang, Xun, Zhong, Hai, Xia, Qingxin, Zhang, Lihua, Huang, Longbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Reparameterization Flow Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
von: Zhong, Hai, et al.
Veröffentlicht: (2026)
Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies
von: Li, Zhuoran, et al.
Veröffentlicht: (2026)
von: Li, Zhuoran, et al.
Veröffentlicht: (2026)
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
von: Wang, Xun, et al.
Veröffentlicht: (2025)
von: Wang, Xun, et al.
Veröffentlicht: (2025)
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
von: Zhong, Hai, et al.
Veröffentlicht: (2025)
Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration
von: Zhong, Hai, et al.
Veröffentlicht: (2024)
von: Zhong, Hai, et al.
Veröffentlicht: (2024)
Beyond the Proxy: Trajectory-Distilled Guidance for Offline GFlowNet Training
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2025)
Offline Critic-Guided Diffusion Policy for Multi-User Delay-Constrained Scheduling
von: Li, Zhuoran, et al.
Veröffentlicht: (2025)
von: Li, Zhuoran, et al.
Veröffentlicht: (2025)
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
von: Hu, Rui, et al.
Veröffentlicht: (2024)
von: Hu, Rui, et al.
Veröffentlicht: (2024)
From Solo to Symphony: Orchestrating Multi-Agent Collaboration with Single-Agent Demos
von: Wang, Xun, et al.
Veröffentlicht: (2025)
von: Wang, Xun, et al.
Veröffentlicht: (2025)
PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching
von: Chen, Ruishuo, et al.
Veröffentlicht: (2026)
von: Chen, Ruishuo, et al.
Veröffentlicht: (2026)
Offline Multi-Agent Reinforcement Learning via In-Sample Sequential Policy Optimization
von: Liu, Zongkai, et al.
Veröffentlicht: (2024)
von: Liu, Zongkai, et al.
Veröffentlicht: (2024)
Pessimistic Value Iteration for Multi-Task Data Sharing in Offline Reinforcement Learning
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
von: Bai, Chenjia, et al.
Veröffentlicht: (2024)
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
von: Zhan, Guojian, et al.
Veröffentlicht: (2026)
Finite-time Convergence Analysis of Actor-Critic with Evolving Reward
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
Layer-Aware Influence for Online Data Valuation Estimation
von: Yang, Ziao, et al.
Veröffentlicht: (2025)
von: Yang, Ziao, et al.
Veröffentlicht: (2025)
FlowQ: Energy-Guided Flow Policies for Offline Reinforcement Learning
von: Alles, Marvin, et al.
Veröffentlicht: (2025)
von: Alles, Marvin, et al.
Veröffentlicht: (2025)
CROP: Conservative Reward for Model-based Offline Policy Optimization
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
von: Wu, Yecheng, et al.
Veröffentlicht: (2026)
von: Wu, Yecheng, et al.
Veröffentlicht: (2026)
Real-Time Parallel Counterfactual Regret Minimization
von: Li, Boning, et al.
Veröffentlicht: (2026)
von: Li, Boning, et al.
Veröffentlicht: (2026)
Policy-regularized Offline Multi-objective Reinforcement Learning
von: Lin, Qian, et al.
Veröffentlicht: (2024)
von: Lin, Qian, et al.
Veröffentlicht: (2024)
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
von: Tayal, Mumuksh, et al.
Veröffentlicht: (2026)
Beyond State-Wise Mirror Descent: Offline Policy Optimization with Parametric Policies
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
One-Step Generative Policies with Q-Learning: A Reformulation of MeanFlow
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2025)
Guided Flow Policy: Learning from High-Value Actions in Offline Reinforcement Learning
von: Tiofack, Franki Nguimatsia, et al.
Veröffentlicht: (2025)
von: Tiofack, Franki Nguimatsia, et al.
Veröffentlicht: (2025)
Network Topology Optimization via Deep Reinforcement Learning
von: Li, Zhuoran, et al.
Veröffentlicht: (2022)
von: Li, Zhuoran, et al.
Veröffentlicht: (2022)
Score-Based One-step MeanFlow Policy Optimization
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026)
von: Kim, Kyungyoon, et al.
Veröffentlicht: (2026)
Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
von: Wang, Zeyuan, et al.
Veröffentlicht: (2026)
Offline Reinforcement Learning with Generative Trajectory Policies
von: Feng, Xinsong, et al.
Veröffentlicht: (2025)
von: Feng, Xinsong, et al.
Veröffentlicht: (2025)
Analytic Energy-Guided Policy Optimization for Offline Reinforcement Learning
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
von: Hu, Jifeng, et al.
Veröffentlicht: (2025)
Preferred-Action-Optimized Diffusion Policies for Offline Reinforcement Learning
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
von: Zhang, Tianle, et al.
Veröffentlicht: (2024)
Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents
von: Li, Ziming, et al.
Veröffentlicht: (2026)
von: Li, Ziming, et al.
Veröffentlicht: (2026)
Belief-Based Offline Reinforcement Learning for Delay-Robust Policy Optimization
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2025)
von: Zhan, Simon Sinong, et al.
Veröffentlicht: (2025)
Policy-Based Trajectory Clustering in Offline Reinforcement Learning
von: Hu, Hao, et al.
Veröffentlicht: (2025)
von: Hu, Hao, et al.
Veröffentlicht: (2025)
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
von: Li, Mingxuan, et al.
Veröffentlicht: (2026)
von: Li, Mingxuan, et al.
Veröffentlicht: (2026)
TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2026)
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
von: Lin, Hongqiang, et al.
Veröffentlicht: (2026)
Policy Constraint by Only Support Constraint for Offline Reinforcement Learning
von: Gao, Yunkai, et al.
Veröffentlicht: (2025)
von: Gao, Yunkai, et al.
Veröffentlicht: (2025)
Robust Policy Expansion for Offline-to-Online RL under Diverse Data Corruption
von: He, Longxiang, et al.
Veröffentlicht: (2025)
von: He, Longxiang, et al.
Veröffentlicht: (2025)
LLM-Driven Policy Diffusion: Enhancing Generalization in Offline Reinforcement Learning
von: Zhang, Hanping, et al.
Veröffentlicht: (2025)
von: Zhang, Hanping, et al.
Veröffentlicht: (2025)
Offline Policy Evaluation for Reinforcement Learning with Adaptively Collected Data
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
von: Madhow, Sunil, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Reparameterization Flow Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2026) -
Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion Policies
von: Li, Zhuoran, et al.
Veröffentlicht: (2026) -
Beyond Shallow Behavior: Task-Efficient Value-Based Multi-Task Offline MARL via Skill Discovery
von: Wang, Xun, et al.
Veröffentlicht: (2025) -
Reparameterization Proximal Policy Optimization
von: Zhong, Hai, et al.
Veröffentlicht: (2025) -
Offline-to-Online Multi-Agent Reinforcement Learning with Offline Value Function Memory and Sequential Exploration
von: Zhong, Hai, et al.
Veröffentlicht: (2024)