MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909860083269632 |
|---|---|
| author | Ashraf, Tajamul Nawaz, Umair Shaker, Abdelrahman M. Anwer, Rao Torr, Philip Khan, Fahad Shahbaz Khan, Salman |
| author_facet | Ashraf, Tajamul Nawaz, Umair Shaker, Abdelrahman M. Anwer, Rao Torr, Philip Khan, Fahad Shahbaz Khan, Salman |
| contents | Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with a vision-centric agent tuning framework that automatically synthesizes multimodal trajectories, generates step-wise preference pairs, and trains a VLM controller for robust tool-use reasoning. Our pipeline first constructs M-TRACE, a large-scale dataset of 28.5K multimodal tasks with 177K verified trajectories, enabling imitation-based trajectory tuning. Building on this, we develop MATRIX Agent, a controller finetuned on M-TRACE for step-wise tool reasoning. To achieve finer alignment, we further introduce Pref-X, a set of 11K automatically generated preference pairs, and optimize MATRIX on it via step-wise preference learning. Across three benchmarks, Agent-X, GTA, and GAIA, MATRIX consistently surpasses both open- and closed-source VLMs, demonstrating scalable and effective multimodal tool use. Our data and code is avaliable at https://github.com/mbzuai-oryx/MATRIX. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_08567 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning Ashraf, Tajamul Nawaz, Umair Shaker, Abdelrahman M. Anwer, Rao Torr, Philip Khan, Fahad Shahbaz Khan, Salman Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with a vision-centric agent tuning framework that automatically synthesizes multimodal trajectories, generates step-wise preference pairs, and trains a VLM controller for robust tool-use reasoning. Our pipeline first constructs M-TRACE, a large-scale dataset of 28.5K multimodal tasks with 177K verified trajectories, enabling imitation-based trajectory tuning. Building on this, we develop MATRIX Agent, a controller finetuned on M-TRACE for step-wise tool reasoning. To achieve finer alignment, we further introduce Pref-X, a set of 11K automatically generated preference pairs, and optimize MATRIX on it via step-wise preference learning. Across three benchmarks, Agent-X, GTA, and GAIA, MATRIX consistently surpasses both open- and closed-source VLMs, demonstrating scalable and effective multimodal tool use. Our data and code is avaliable at https://github.com/mbzuai-oryx/MATRIX. |
| title | MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Computation and Language |
| url | https://arxiv.org/abs/2510.08567 |