Difference Rewards Policy Gradients
Fuente:
arXiv
Guardado en:
| Autores principales: | Castellini, Jacopo, Devlin, Sam, Oliehoek, Frans A., Savani, Rahul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2020
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
por: Castellini, Jacopo, et al.
Publicado: (2019)
por: Castellini, Jacopo, et al.
Publicado: (2019)
Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach
por: Peralez, Johan, et al.
Publicado: (2024)
por: Peralez, Johan, et al.
Publicado: (2024)
On Convex Optimal Value Functions For POSGs
por: Cunha, Rafael F., et al.
Publicado: (2023)
por: Cunha, Rafael F., et al.
Publicado: (2023)
Towards General Negotiation Strategies with End-to-End Reinforcement Learning
por: Renting, Bram M., et al.
Publicado: (2024)
por: Renting, Bram M., et al.
Publicado: (2024)
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
por: Nath, Saptarshi, et al.
Publicado: (2025)
por: Nath, Saptarshi, et al.
Publicado: (2025)
Individual Contributions as Intrinsic Exploration Scaffolds for Multi-agent Reinforcement Learning
por: Li, Xinran, et al.
Publicado: (2024)
por: Li, Xinran, et al.
Publicado: (2024)
Distributed Value Decomposition Networks with Networked Agents
por: Varela, Guilherme S., et al.
Publicado: (2025)
por: Varela, Guilherme S., et al.
Publicado: (2025)
Representational Collapse in Multi-Agent LLM Committees: Measurement and Diversity-Aware Consensus
por: Patel, Dipkumar
Publicado: (2026)
por: Patel, Dipkumar
Publicado: (2026)
A Principle of Targeted Intervention for Multi-Agent Reinforcement Learning
por: Liu, Anjie, et al.
Publicado: (2025)
por: Liu, Anjie, et al.
Publicado: (2025)
Networked Agents in the Dark: Team Value Learning under Partial Observability
por: Varela, Guilherme S., et al.
Publicado: (2025)
por: Varela, Guilherme S., et al.
Publicado: (2025)
Collaboration Promotes Group Resilience in Multi-Agent RL
por: Shraga, Ilai, et al.
Publicado: (2021)
por: Shraga, Ilai, et al.
Publicado: (2021)
Risk-Sensitive Multi-Agent Reinforcement Learning in Network Aggregative Markov Games
por: Ghaemi, Hafez, et al.
Publicado: (2024)
por: Ghaemi, Hafez, et al.
Publicado: (2024)
Dynamic Attentional Context Scoping: Agent-Triggered Focus Sessions for Isolated Per-Agent Steering in Multi-Agent LLM Orchestration
por: Patel, Nickson
Publicado: (2026)
por: Patel, Nickson
Publicado: (2026)
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
por: Maryanskyy, Artem
Publicado: (2026)
por: Maryanskyy, Artem
Publicado: (2026)
QTypeMix: Enhancing Multi-Agent Cooperative Strategies through Heterogeneous and Homogeneous Value Decomposition
por: Fu, Songchen, et al.
Publicado: (2024)
por: Fu, Songchen, et al.
Publicado: (2024)
Differentiable Model Predictive Safety for Heterogeneous Mobility at Urban Intersections
por: Song, Wenzhe, et al.
Publicado: (2026)
por: Song, Wenzhe, et al.
Publicado: (2026)
MindGames Arena Generalization Track: In2AI Solution with Delayed Per-Step Reward Attribution
por: Korshuk, Aliaksei, et al.
Publicado: (2026)
por: Korshuk, Aliaksei, et al.
Publicado: (2026)
Enhancing Heterogeneous Multi-Agent Cooperation in Decentralized MARL via GNN-driven Intrinsic Rewards
por: Monon, Jahir Sadik, et al.
Publicado: (2024)
por: Monon, Jahir Sadik, et al.
Publicado: (2024)
Characterizing MARL for Energy Control: A Multi-KPI Benchmark on the CityLearn Environment
por: Khouja, Aymen, et al.
Publicado: (2026)
por: Khouja, Aymen, et al.
Publicado: (2026)
Discovering Antagonists in Networks of Systems: Robot Deployment
por: Wenger, Ingeborg, et al.
Publicado: (2025)
por: Wenger, Ingeborg, et al.
Publicado: (2025)
Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents
por: Tang, Wenjie, et al.
Publicado: (2026)
por: Tang, Wenjie, et al.
Publicado: (2026)
Knowledge Equivalence in Digital Twins of Intelligent Systems
por: Zhang, Nan, et al.
Publicado: (2022)
por: Zhang, Nan, et al.
Publicado: (2022)
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
por: Schipper, Olivier, et al.
Publicado: (2025)
por: Schipper, Olivier, et al.
Publicado: (2025)
Centrally Coordinated Multi-Agent Reinforcement Learning for Power Grid Topology Control
por: de Mol, Barbera, et al.
Publicado: (2025)
por: de Mol, Barbera, et al.
Publicado: (2025)
AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergence
por: Yu, Geunbin
Publicado: (2026)
por: Yu, Geunbin
Publicado: (2026)
Decentralized Aerial Manipulation of a Cable-Suspended Load using Multi-Agent Reinforcement Learning
por: Zeng, Jack, et al.
Publicado: (2025)
por: Zeng, Jack, et al.
Publicado: (2025)
Bimanual Robot Manipulation via Multi-Agent In-Context Learning
por: Palma, Alessio, et al.
Publicado: (2026)
por: Palma, Alessio, et al.
Publicado: (2026)
Generating Causal Explanations of Vehicular Agent Behavioural Interactions with Learnt Reward Profiles
por: Howard, Rhys, et al.
Publicado: (2025)
por: Howard, Rhys, et al.
Publicado: (2025)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
por: Akella, Aditya
Publicado: (2025)
por: Akella, Aditya
Publicado: (2025)
Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
por: Anand, Emile, et al.
Publicado: (2026)
por: Anand, Emile, et al.
Publicado: (2026)
How to Correctly do Semantic Backpropagation on Language-based Agentic Systems
por: Wang, Wenyi, et al.
Publicado: (2024)
por: Wang, Wenyi, et al.
Publicado: (2024)
David vs. Goliath: Verifiable Agent-to-Agent Jailbreaking via Reinforcement Learning
por: Nellessen, Samuel, et al.
Publicado: (2026)
por: Nellessen, Samuel, et al.
Publicado: (2026)
A Systematic Study of Multi-Agent Deep Reinforcement Learning for Safe and Robust Autonomous Highway Ramp Entry
por: Schester, Larry, et al.
Publicado: (2024)
por: Schester, Larry, et al.
Publicado: (2024)
ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork
por: Wang, Caroline, et al.
Publicado: (2025)
por: Wang, Caroline, et al.
Publicado: (2025)
Learning to Communicate Across Modalities: Perceptual Heterogeneity in Multi-Agent Systems
por: Pitzer, Naomi, et al.
Publicado: (2026)
por: Pitzer, Naomi, et al.
Publicado: (2026)
Learning To Help: Training Models to Assist Legacy Devices
por: Wu, Yu, et al.
Publicado: (2024)
por: Wu, Yu, et al.
Publicado: (2024)
ME-IGM: Individual-Global-Max in Maximum Entropy Multi-Agent Reinforcement Learning
por: Chen, Wen-Tse, et al.
Publicado: (2024)
por: Chen, Wen-Tse, et al.
Publicado: (2024)
Emergent Coordination in Multi-Agent Systems via Pressure Fields and Temporal Decay
por: Rodriguez, Roland
Publicado: (2026)
por: Rodriguez, Roland
Publicado: (2026)
SPIRAL: Self-Play Incremental Racing Algorithm for Learning in Multi-Drone Competitions
por: Akgün, Onur
Publicado: (2025)
por: Akgün, Onur
Publicado: (2025)
Curriculum-Based Iterative Self-Play for Scalable Multi-Drone Racing
por: Akgün, Onur
Publicado: (2025)
por: Akgün, Onur
Publicado: (2025)
Ejemplares similares
-
Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
por: Castellini, Jacopo, et al.
Publicado: (2019) -
Optimally Solving Simultaneous-Move Dec-POMDPs: The Sequential Central Planning Approach
por: Peralez, Johan, et al.
Publicado: (2024) -
On Convex Optimal Value Functions For POSGs
por: Cunha, Rafael F., et al.
Publicado: (2023) -
Towards General Negotiation Strategies with End-to-End Reinforcement Learning
por: Renting, Bram M., et al.
Publicado: (2024) -
Policy Search, Retrieval, and Composition via Task Similarity in Collaborative Agentic Systems
por: Nath, Saptarshi, et al.
Publicado: (2025)