MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Yifan, Yin, Fei, Bai, Qingyan, Lin, Zicheng, Yang, Yujiu
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866916006871433216
author Chen, Yifan
Yin, Fei
Bai, Qingyan
Lin, Zicheng
Yang, Yujiu
author_facet Chen, Yifan
Yin, Fei
Bai, Qingyan
Lin, Zicheng
Yang, Yujiu
contents We introduce MMCL-Bench, a benchmark for multimodal context learning: learning task-local rules, procedures, and empirical patterns from visual or mixed-modality teaching context and applying them to new visual instances. Unlike text-only context learning or standard multimodal question answering, this setting requires models to recover and localize relevant evidence from images, screenshots, manuals, videos, and frame sequences before they can reason over the learned context. MMCL-Bench contains 102 tasks spanning three categories: rule system application, procedural task execution, and empirical discovery and induction. We evaluate frontier multimodal models with strict rubric-based scoring and find that current systems remain far from robust multimodal context learning, with even the strongest model solving fewer than one-third of tasks under strict evaluation. Diagnostic ablations and error analysis show that failures arise throughout the context-to-answer pipeline, including context anchoring, visual evidence extraction, context reasoning, and response construction. MMCL-Bench thus highlights multimodal context learning as an important unsolved capability bottleneck for current multimodal models.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12703
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
Chen, Yifan
Yin, Fei
Bai, Qingyan
Lin, Zicheng
Yang, Yujiu
Computer Vision and Pattern Recognition
Artificial Intelligence
We introduce MMCL-Bench, a benchmark for multimodal context learning: learning task-local rules, procedures, and empirical patterns from visual or mixed-modality teaching context and applying them to new visual instances. Unlike text-only context learning or standard multimodal question answering, this setting requires models to recover and localize relevant evidence from images, screenshots, manuals, videos, and frame sequences before they can reason over the learned context. MMCL-Bench contains 102 tasks spanning three categories: rule system application, procedural task execution, and empirical discovery and induction. We evaluate frontier multimodal models with strict rubric-based scoring and find that current systems remain far from robust multimodal context learning, with even the strongest model solving fewer than one-third of tasks under strict evaluation. Diagnostic ablations and error analysis show that failures arise throughout the context-to-answer pipeline, including context anchoring, visual evidence extraction, context reasoning, and response construction. MMCL-Bench thus highlights multimodal context learning as an important unsolved capability bottleneck for current multimodal models.
title MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2605.12703