Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhan, Wenhao, Fujimoto, Scott, Zhu, Zheqing, Lee, Jason D., Jiang, Daniel R., Efroni, Yonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Aligned Multi Objective Optimization
von: Efroni, Yonathan, et al.
Veröffentlicht: (2025)
von: Efroni, Yonathan, et al.
Veröffentlicht: (2025)
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
von: Wu, Runzhe, et al.
Veröffentlicht: (2025)
Simple Optimizers for Convex Aligned Multi-Objective Optimization
von: Kretzu, Ben, et al.
Veröffentlicht: (2025)
von: Kretzu, Ben, et al.
Veröffentlicht: (2025)
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024)
Gradient Free Deep Reinforcement Learning With TabPFN
von: Schiff, David, et al.
Veröffentlicht: (2025)
von: Schiff, David, et al.
Veröffentlicht: (2025)
Pearl: A Production-ready Reinforcement Learning Agent
von: Zhu, Zheqing, et al.
Veröffentlicht: (2023)
von: Zhu, Zheqing, et al.
Veröffentlicht: (2023)
Generalizing Multi-Step Inverse Models for Representation Learning to Finite-Memory POMDPs
von: Wu, Lili, et al.
Veröffentlicht: (2024)
von: Wu, Lili, et al.
Veröffentlicht: (2024)
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
von: Roth, Amit, et al.
Veröffentlicht: (2026)
von: Roth, Amit, et al.
Veröffentlicht: (2026)
Are Expressive Models Truly Necessary for Offline RL?
von: Wang, Guan, et al.
Veröffentlicht: (2024)
von: Wang, Guan, et al.
Veröffentlicht: (2024)
Uncertainty of Joint Neural Contextual Bandit
von: Guo, Hongbo, et al.
Veröffentlicht: (2024)
von: Guo, Hongbo, et al.
Veröffentlicht: (2024)
Accelerating RL for LLM Reasoning with Optimal Advantage Regression
von: Brantley, Kianté, et al.
Veröffentlicht: (2025)
von: Brantley, Kianté, et al.
Veröffentlicht: (2025)
Meta-DT: Offline Meta-RL as Conditional Sequence Modeling with World Model Disentanglement
von: Wang, Zhi, et al.
Veröffentlicht: (2024)
von: Wang, Zhi, et al.
Veröffentlicht: (2024)
Time After Time: Deep-Q Effect Estimation for Interventions on When and What to do
von: Wald, Yoav, et al.
Veröffentlicht: (2025)
von: Wald, Yoav, et al.
Veröffentlicht: (2025)
Optimal Multi-Distribution Learning
von: Zhang, Zihan, et al.
Veröffentlicht: (2023)
von: Zhang, Zihan, et al.
Veröffentlicht: (2023)
Multi-Agent Path Finding via Offline RL and LLM Collaboration
von: Atasever, Merve, et al.
Veröffentlicht: (2025)
von: Atasever, Merve, et al.
Veröffentlicht: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
von: Kim, Juno, et al.
Veröffentlicht: (2025)
von: Kim, Juno, et al.
Veröffentlicht: (2025)
Exploiting Low-Rank Structure in Max-K-Cut Problems
von: Stevens, Ria, et al.
Veröffentlicht: (2026)
von: Stevens, Ria, et al.
Veröffentlicht: (2026)
Sparse-Reg: Improving Sample Complexity in Offline Reinforcement Learning using Sparsity
von: Arnob, Samin Yeasar, et al.
Veröffentlicht: (2025)
von: Arnob, Samin Yeasar, et al.
Veröffentlicht: (2025)
Low Rank Learning for Offline Query Optimization
von: Yi, Zixuan, et al.
Veröffentlicht: (2025)
von: Yi, Zixuan, et al.
Veröffentlicht: (2025)
DEAS: DEtached value learning with Action Sequence for Scalable Offline RL
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
von: Kim, Changyeon, et al.
Veröffentlicht: (2025)
Polynomial Expansion Rank Adaptation: Enhancing Low-Rank Fine-Tuning with High-Order Interactions
von: Zhang, Wenhao, et al.
Veröffentlicht: (2026)
von: Zhang, Wenhao, et al.
Veröffentlicht: (2026)
Improving Generative Ad Text on Facebook using Reinforcement Learning
von: Jiang, Daniel R., et al.
Veröffentlicht: (2025)
von: Jiang, Daniel R., et al.
Veröffentlicht: (2025)
Mean-Field Diffuser: Scaling Offline MARL to Thousands of Agents
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
von: Li, Wenhao, et al.
Veröffentlicht: (2026)
Offline Multi-task Transfer RL with Representational Penalization
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
von: Bose, Avinandan, et al.
Veröffentlicht: (2024)
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Language-Conditioned Offline RL for Multi-Robot Navigation
von: Morad, Steven, et al.
Veröffentlicht: (2024)
von: Morad, Steven, et al.
Veröffentlicht: (2024)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
von: Zhu, Lingwei, et al.
Veröffentlicht: (2025)
The Bias of Harmful Label Associations in Vision-Language Models
von: Hazirbas, Caner, et al.
Veröffentlicht: (2024)
von: Hazirbas, Caner, et al.
Veröffentlicht: (2024)
AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear Mapping
von: Dong, Haonan, et al.
Veröffentlicht: (2025)
von: Dong, Haonan, et al.
Veröffentlicht: (2025)
Scalable Offline Model-Based RL with Action Chunks
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
Budgeting Counterfactual for Offline RL
von: Liu, Yao, et al.
Veröffentlicht: (2023)
von: Liu, Yao, et al.
Veröffentlicht: (2023)
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
von: Zhang, Qiang, et al.
Veröffentlicht: (2026)
von: Zhang, Qiang, et al.
Veröffentlicht: (2026)
Improving Offline RL by Blending Heuristics
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
von: Geng, Sinong, et al.
Veröffentlicht: (2023)
Provable Reward-Agnostic Preference-Based Reinforcement Learning
von: Zhan, Wenhao, et al.
Veröffentlicht: (2023)
von: Zhan, Wenhao, et al.
Veröffentlicht: (2023)
Streetwise Agents: Empowering Offline RL Policies to Outsmart Exogenous Stochastic Disturbances in RTC
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
von: Soni, Aditya, et al.
Veröffentlicht: (2024)
H2O+: An Improved Framework for Hybrid Offline-and-Online RL with Dynamics Gaps
von: Niu, Haoyi, et al.
Veröffentlicht: (2023)
von: Niu, Haoyi, et al.
Veröffentlicht: (2023)
Self-Improvement of Language Models by Post-Training on Multi-Agent Debate
von: Samanta, Ankur, et al.
Veröffentlicht: (2025)
von: Samanta, Ankur, et al.
Veröffentlicht: (2025)
Learning Emergence of Interaction Patterns across Independent RL Agents in Multi-Agent Environments
von: Baddam, Vasanth Reddy, et al.
Veröffentlicht: (2024)
von: Baddam, Vasanth Reddy, et al.
Veröffentlicht: (2024)
[Re] FairDICE: A Fair Tradeoff in Multi-objective Offline RL
von: Adema, Peter, et al.
Veröffentlicht: (2026)
von: Adema, Peter, et al.
Veröffentlicht: (2026)
CtRL-Sim: Reactive and Controllable Driving Agents with Offline Reinforcement Learning
von: Rowe, Luke, et al.
Veröffentlicht: (2024)
von: Rowe, Luke, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Aligned Multi Objective Optimization
von: Efroni, Yonathan, et al.
Veröffentlicht: (2025) -
Imbalanced Gradients in RL Post-Training of Multi-Task LLMs
von: Wu, Runzhe, et al.
Veröffentlicht: (2025) -
Simple Optimizers for Convex Aligned Multi-Objective Optimization
von: Kretzu, Ben, et al.
Veröffentlicht: (2025) -
RL in Latent MDPs is Tractable: Online Guarantees via Off-Policy Evaluation
von: Kwon, Jeongyeol, et al.
Veröffentlicht: (2024) -
Gradient Free Deep Reinforcement Learning With TabPFN
von: Schiff, David, et al.
Veröffentlicht: (2025)