MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ashraf, Tajamul, Nawaz, Umair, Shaker, Abdelrahman M., Anwer, Rao, Torr, Philip, Khan, Fahad Shahbaz, Khan, Salman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909860083269632
author Ashraf, Tajamul
Nawaz, Umair
Shaker, Abdelrahman M.
Anwer, Rao
Torr, Philip
Khan, Fahad Shahbaz
Khan, Salman
author_facet Ashraf, Tajamul
Nawaz, Umair
Shaker, Abdelrahman M.
Anwer, Rao
Torr, Philip
Khan, Fahad Shahbaz
Khan, Salman
contents Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with a vision-centric agent tuning framework that automatically synthesizes multimodal trajectories, generates step-wise preference pairs, and trains a VLM controller for robust tool-use reasoning. Our pipeline first constructs M-TRACE, a large-scale dataset of 28.5K multimodal tasks with 177K verified trajectories, enabling imitation-based trajectory tuning. Building on this, we develop MATRIX Agent, a controller finetuned on M-TRACE for step-wise tool reasoning. To achieve finer alignment, we further introduce Pref-X, a set of 11K automatically generated preference pairs, and optimize MATRIX on it via step-wise preference learning. Across three benchmarks, Agent-X, GTA, and GAIA, MATRIX consistently surpasses both open- and closed-source VLMs, demonstrating scalable and effective multimodal tool use. Our data and code is avaliable at https://github.com/mbzuai-oryx/MATRIX.
format Preprint
id arxiv_https___arxiv_org_abs_2510_08567
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
Ashraf, Tajamul
Nawaz, Umair
Shaker, Abdelrahman M.
Anwer, Rao
Torr, Philip
Khan, Fahad Shahbaz
Khan, Salman
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Vision language models (VLMs) are increasingly deployed as controllers with access to external tools for complex reasoning and decision-making, yet their effectiveness remains limited by the scarcity of high-quality multimodal trajectories and the cost of manual annotation. We address this challenge with a vision-centric agent tuning framework that automatically synthesizes multimodal trajectories, generates step-wise preference pairs, and trains a VLM controller for robust tool-use reasoning. Our pipeline first constructs M-TRACE, a large-scale dataset of 28.5K multimodal tasks with 177K verified trajectories, enabling imitation-based trajectory tuning. Building on this, we develop MATRIX Agent, a controller finetuned on M-TRACE for step-wise tool reasoning. To achieve finer alignment, we further introduce Pref-X, a set of 11K automatically generated preference pairs, and optimize MATRIX on it via step-wise preference learning. Across three benchmarks, Agent-X, GTA, and GAIA, MATRIX consistently surpasses both open- and closed-source VLMs, demonstrating scalable and effective multimodal tool use. Our data and code is avaliable at https://github.com/mbzuai-oryx/MATRIX.
title MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2510.08567