MMSkills: Towards Multimodal Skills for General Visual Agents

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kangning, Shao, Shuai, Li, Qingyao, Lin, Jianghao, Fu, Lingyue, Wang, Shijian, Jiao, Wenxiang, Lu, Yuan, Liu, Weiwen, Zhang, Weinan, Yu, Yong
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913178720403456
author Zhang, Kangning
Shao, Shuai
Li, Qingyao
Lin, Jianghao
Fu, Lingyue
Wang, Shijian
Jiao, Wenxiang
Lu, Yuan
Liu, Weiwen
Zhang, Weinan
Yu, Yong
author_facet Zhang, Kangning
Shao, Shuai
Li, Qingyao
Lin, Jianghao
Fu, Lingyue
Wang, Shijian
Jiao, Wenxiang
Lu, Yuan
Liu, Weiwen
Zhang, Weinan
Yu, Yong
contents Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agents, however, procedural knowledge is inherently multimodal: reuse depends not only on what operation to perform, but also on recognizing the relevant state, interpreting visual evidence of progress or failure, and deciding what to do next. We formalize this requirement as multimodal procedural knowledge and address three practical challenges: (I) what a multimodal skill package should contain; (II) where such packages can be derived from public interaction experience; and (III) how agents can consult multimodal evidence at inference time without excessive image context or over-anchoring to reference screenshots. We introduce MMSkills, a framework for representing, generating, and using reusable multimodal procedures for runtime visual decision making. Each MMSkill is a compact, state-conditioned package that couples a textual procedure with runtime state cards and multi-view keyframes. To construct these packages, we develop an agentic trajectory-to-skill Generator that transforms public non-evaluation trajectories into reusable multimodal skills through workflow grouping, procedure induction, visual grounding, and meta-skill-guided auditing. To use them, we introduce a branch-loaded multimodal skill agent: selected state cards and keyframes are inspected in a temporary branch, aligned with the live environment, and distilled into structured guidance for the main agent. Experiments across GUI and game-based visual-agent benchmarks show that MMSkills consistently improve both frontier and smaller multimodal agents, suggesting that external multimodal procedural knowledge complements model-internal priors.
format Preprint
id arxiv_https___arxiv_org_abs_2605_13527
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMSkills: Towards Multimodal Skills for General Visual Agents
Zhang, Kangning
Shao, Shuai
Li, Qingyao
Lin, Jianghao
Fu, Lingyue
Wang, Shijian
Jiao, Wenxiang
Lu, Yuan
Liu, Weiwen
Zhang, Weinan
Yu, Yong
Artificial Intelligence
Reusable skills have become a core substrate for improving agent capabilities, yet most existing skill packages encode reusable behavior primarily as textual prompts, executable code, or learned routines. For visual agents, however, procedural knowledge is inherently multimodal: reuse depends not only on what operation to perform, but also on recognizing the relevant state, interpreting visual evidence of progress or failure, and deciding what to do next. We formalize this requirement as multimodal procedural knowledge and address three practical challenges: (I) what a multimodal skill package should contain; (II) where such packages can be derived from public interaction experience; and (III) how agents can consult multimodal evidence at inference time without excessive image context or over-anchoring to reference screenshots. We introduce MMSkills, a framework for representing, generating, and using reusable multimodal procedures for runtime visual decision making. Each MMSkill is a compact, state-conditioned package that couples a textual procedure with runtime state cards and multi-view keyframes. To construct these packages, we develop an agentic trajectory-to-skill Generator that transforms public non-evaluation trajectories into reusable multimodal skills through workflow grouping, procedure induction, visual grounding, and meta-skill-guided auditing. To use them, we introduce a branch-loaded multimodal skill agent: selected state cards and keyframes are inspected in a temporary branch, aligned with the live environment, and distilled into structured guidance for the main agent. Experiments across GUI and game-based visual-agent benchmarks show that MMSkills consistently improve both frontier and smaller multimodal agents, suggesting that external multimodal procedural knowledge complements model-internal priors.
title MMSkills: Towards Multimodal Skills for General Visual Agents
topic Artificial Intelligence
url https://arxiv.org/abs/2605.13527