TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hasegawa, Kimihiro, Imrattanatrai, Wiradee, Asada, Masaki, Fukuda, Ken, Mitamura, Teruko
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909817964068864
author Hasegawa, Kimihiro
Imrattanatrai, Wiradee
Asada, Masaki
Fukuda, Ken
Mitamura, Teruko
author_facet Hasegawa, Kimihiro
Imrattanatrai, Wiradee
Asada, Masaki
Fukuda, Ken
Mitamura, Teruko
contents Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite its potential use cases, the system development tailored for such an assistant is still underexplored. In this paper, we propose a novel framework, called TAMA, a Tool-Augmented Multimodal Agent, for procedural activity understanding. TAMA enables interleaved multimodal reasoning by making use of multimedia-returning tools in a training-free setting. Our experimental result on the multimodal procedural QA dataset, ProMQA-Assembly, shows that our approach can improve the performance of vision-language models, especially GPT-5 and MiMo-VL. Furthermore, our ablation studies provide empirical support for the effectiveness of two features that characterize our framework, multimedia-returning tools and agentic flexible tool selection. We believe our proposed framework and experimental results facilitate the thinking with images paradigm for video and multimodal tasks, let alone the development of procedural activity assistants.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00161
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
Hasegawa, Kimihiro
Imrattanatrai, Wiradee
Asada, Masaki
Fukuda, Ken
Mitamura, Teruko
Computation and Language
Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological experiments. Despite its potential use cases, the system development tailored for such an assistant is still underexplored. In this paper, we propose a novel framework, called TAMA, a Tool-Augmented Multimodal Agent, for procedural activity understanding. TAMA enables interleaved multimodal reasoning by making use of multimedia-returning tools in a training-free setting. Our experimental result on the multimodal procedural QA dataset, ProMQA-Assembly, shows that our approach can improve the performance of vision-language models, especially GPT-5 and MiMo-VL. Furthermore, our ablation studies provide empirical support for the effectiveness of two features that characterize our framework, multimedia-returning tools and agentic flexible tool selection. We believe our proposed framework and experimental results facilitate the thinking with images paradigm for video and multimodal tasks, let alone the development of procedural activity assistants.
title TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
topic Computation and Language
url https://arxiv.org/abs/2510.00161