InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909957461377024 |
|---|---|
| author | Li, Bin Zhang, Ruichi Liang, Han Zhang, Jingyan Zhang, Juze Chen, Xin Xu, Lan Yu, Jingyi Wang, Jingya |
| author_facet | Li, Bin Zhang, Ruichi Liang, Han Zhang, Jingyan Zhang, Juze Chen, Xin Xu, Lan Yu, Jingyi Wang, Jingya |
| contents | Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAgent, the first end-to-end framework for text-driven physics-based multi-agent humanoid control. At its core, we introduce an autoregressive diffusion transformer equipped with multi-stream blocks, which decouples proprioception, exteroception, and action to mitigate cross-modal interference while enabling synergistic coordination. We further propose a novel interaction graph exteroception representation that explicitly captures fine-grained joint-to-joint spatial dependencies to facilitate network learning. Additionally, within it we devise a sparse edge-based attention mechanism that dynamically prunes redundant connections and emphasizes critical inter-agent spatial relations, thereby enhancing the robustness of interaction modeling. Extensive experiments demonstrate that InterAgent consistently outperforms multiple strong baselines, achieving state-of-the-art performance. It enables producing coherent, physically plausible, and semantically faithful multi-agent behaviors from only text prompts. Our code and data will be released to facilitate future research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_07410 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs Li, Bin Zhang, Ruichi Liang, Han Zhang, Jingyan Zhang, Juze Chen, Xin Xu, Lan Yu, Jingyi Wang, Jingya Computer Vision and Pattern Recognition Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAgent, the first end-to-end framework for text-driven physics-based multi-agent humanoid control. At its core, we introduce an autoregressive diffusion transformer equipped with multi-stream blocks, which decouples proprioception, exteroception, and action to mitigate cross-modal interference while enabling synergistic coordination. We further propose a novel interaction graph exteroception representation that explicitly captures fine-grained joint-to-joint spatial dependencies to facilitate network learning. Additionally, within it we devise a sparse edge-based attention mechanism that dynamically prunes redundant connections and emphasizes critical inter-agent spatial relations, thereby enhancing the robustness of interaction modeling. Extensive experiments demonstrate that InterAgent consistently outperforms multiple strong baselines, achieving state-of-the-art performance. It enables producing coherent, physically plausible, and semantically faithful multi-agent behaviors from only text prompts. Our code and data will be released to facilitate future research. |
| title | InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.07410 |