Enhancing Multi-Agent Collaboration with Attention-Based Actor-Critic Policies
Fuente:
arXiv
Saved in:
| Main Authors: | Belinchon, Hugo Garrido-Lestache, Kedziora, Jeremy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Avoiding Death through Fear Intrinsic Conditioning
by: Sanchez, Rodney, et al.
Published: (2025)
by: Sanchez, Rodney, et al.
Published: (2025)
The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning
by: Vahdati, Sahar, et al.
Published: (2026)
by: Vahdati, Sahar, et al.
Published: (2026)
EduQate: Generating Adaptive Curricula through RMABs in Education Settings
by: Tio, Sidney, et al.
Published: (2024)
by: Tio, Sidney, et al.
Published: (2024)
DM$^2$: Decentralized Multi-Agent Reinforcement Learning for Distribution Matching
by: Wang, Caroline, et al.
Published: (2022)
by: Wang, Caroline, et al.
Published: (2022)
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
by: Benfeghoul, Martin, et al.
Published: (2024)
by: Benfeghoul, Martin, et al.
Published: (2024)
Quantum Abduction: A New Paradigm for Reasoning under Uncertainty
by: Pareschi, Remo
Published: (2025)
by: Pareschi, Remo
Published: (2025)
A survey of air combat behavior modeling using machine learning
by: Gorton, Patrick Ribu, et al.
Published: (2024)
by: Gorton, Patrick Ribu, et al.
Published: (2024)
Pioneer Agent: Continual Improvement of Small Language Models in Production
by: Atreja, Dhruv, et al.
Published: (2026)
by: Atreja, Dhruv, et al.
Published: (2026)
Prediction Instability in Machine Learning Ensembles
by: Kedziora, Jeremy
Published: (2024)
by: Kedziora, Jeremy
Published: (2024)
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)
by: Della Libera, Luca
Published: (2024)
Strategy Masking: A Method for Guardrails in Value-based Reinforcement Learning Agents
by: Keane, Jonathan, et al.
Published: (2025)
by: Keane, Jonathan, et al.
Published: (2025)
AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems
by: Koenigstein, Nicole
Published: (2026)
by: Koenigstein, Nicole
Published: (2026)
Efficient Contextual Preferential Bayesian Optimization with Historical Examples
by: Khan, Farha A., et al.
Published: (2022)
by: Khan, Farha A., et al.
Published: (2022)
Language Models, Graph Searching, and Supervision Adulteration: When More Supervision is Less and How to Make More More
by: Frydenlund, Arvid
Published: (2025)
by: Frydenlund, Arvid
Published: (2025)
Regret-Aware Policy Optimization: Environment-Level Memory for Replay Suppression under Delayed Harm
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
Score-informed Neural Operator for Enhancing Ordering-based Causal Discovery
by: Kang, Jiyeon, et al.
Published: (2025)
by: Kang, Jiyeon, et al.
Published: (2025)
A Tale of Two Systems: Characterizing Architectural Complexity on Machine Learning-Enabled Systems
by: Ferreira, Renato Cordeiro
Published: (2025)
by: Ferreira, Renato Cordeiro
Published: (2025)
A Metrics-Oriented Architectural Model to Characterize Complexity on Machine Learning-Enabled Systems
by: Ferreira, Renato Cordeiro
Published: (2025)
by: Ferreira, Renato Cordeiro
Published: (2025)
How Metacognitive Architectures Remember Their Own Thoughts: A Systematic Review
by: Nolte, Robin, et al.
Published: (2025)
by: Nolte, Robin, et al.
Published: (2025)
Adaptable Hindsight Experience Replay for Search-Based Learning
by: Vazaios, Alexandros, et al.
Published: (2025)
by: Vazaios, Alexandros, et al.
Published: (2025)
Algorithm Selection for Optimal Multi-Agent Path Finding via Graph Embedding
by: Shabalin, Carmel, et al.
Published: (2024)
by: Shabalin, Carmel, et al.
Published: (2024)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
by: Silue, Bram, et al.
Published: (2025)
by: Silue, Bram, et al.
Published: (2025)
Perceptron Collaborative Filtering
by: Chakraborty, Arya
Published: (2024)
by: Chakraborty, Arya
Published: (2024)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)
by: Steele, Brady, et al.
Published: (2026)
Differentiable Symbolic Planning: A Neural Architecture for Constraint Reasoning with Learned Feasibility
by: Oruganti, Venkatakrishna Reddy
Published: (2026)
by: Oruganti, Venkatakrishna Reddy
Published: (2026)
Embedded Safety-Aligned Intelligence via Differentiable Internal Alignment Embeddings
by: Rathva, Harsh, et al.
Published: (2025)
by: Rathva, Harsh, et al.
Published: (2025)
Safe Reinforcement Learning with Preference-based Constraint Inference
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Working Paper: Active Causal Structure Learning with Latent Variables: Towards Learning to Detour in Autonomous Robots
by: Riscos, Pablo de los, et al.
Published: (2024)
by: Riscos, Pablo de los, et al.
Published: (2024)
Fast and Precise: Adjusting Planning Horizon with Adaptive Subgoal Search
by: Zawalski, Michał, et al.
Published: (2022)
by: Zawalski, Michał, et al.
Published: (2022)
Umbrella Reinforcement Learning -- computationally efficient tool for hard non-linear problems
by: Nuzhin, Egor E., et al.
Published: (2024)
by: Nuzhin, Egor E., et al.
Published: (2024)
Dreaming Smoothly and Sample Efficiently with Gradient Penalized Latent Dynamics
by: Sonigra, Romil V., et al.
Published: (2026)
by: Sonigra, Romil V., et al.
Published: (2026)
GIRL: Generative Imagination Reinforcement Learning via Information-Theoretic Hallucination Control
by: Hiremath, Prakul Sunil
Published: (2026)
by: Hiremath, Prakul Sunil
Published: (2026)
Incentives for Responsiveness, Instrumental Control and Impact
by: Carey, Ryan, et al.
Published: (2020)
by: Carey, Ryan, et al.
Published: (2020)
When Words Change the Model: Sensitivity of LLMs for Constraint Programming Modelling
by: Pellegrino, Alessio, et al.
Published: (2025)
by: Pellegrino, Alessio, et al.
Published: (2025)
On the Generalization Gap in LLM Planning: Tests and Verifier-Reward RL
by: Belcamino, Valerio, et al.
Published: (2026)
by: Belcamino, Valerio, et al.
Published: (2026)
CORE: Towards Scalable and Efficient Causal Discovery with Reinforcement Learning
by: Sauter, Andreas W. M., et al.
Published: (2024)
by: Sauter, Andreas W. M., et al.
Published: (2024)
Not All Transitions Matter: Evidence from PPO
by: Basnet, Ajhesh
Published: (2026)
by: Basnet, Ajhesh
Published: (2026)
AGWM: Affordance-Grounded World Models for Environments with Compositional Prerequisites
by: Zhang, Qinshi, et al.
Published: (2026)
by: Zhang, Qinshi, et al.
Published: (2026)
What Do World Models Learn in RL? Probing Latent Representations in Learned Environment Simulators
by: Zhang, Xinyu
Published: (2026)
by: Zhang, Xinyu
Published: (2026)
Similar Items
-
Avoiding Death through Fear Intrinsic Conditioning
by: Sanchez, Rodney, et al.
Published: (2025) -
The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning
by: Vahdati, Sahar, et al.
Published: (2026) -
EduQate: Generating Adaptive Curricula through RMABs in Education Settings
by: Tio, Sidney, et al.
Published: (2024) -
DM$^2$: Decentralized Multi-Agent Reinforcement Learning for Distribution Matching
by: Wang, Caroline, et al.
Published: (2022) -
When in Doubt, Think Slow: Iterative Reasoning with Latent Imagination
by: Benfeghoul, Martin, et al.
Published: (2024)