Saved in:
Bibliographic Details
Main Authors: Dong, Yuhao, Tian, Shulin, Liu, Shuai, Ding, Shuangrui, Zang, Yuhang, Dong, Xiaoyi, Cao, Yuhang, Wang, Jiaqi, Liu, Ziwei
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2602.08439
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908822481666048
author Dong, Yuhao
Tian, Shulin
Liu, Shuai
Ding, Shuangrui
Zang, Yuhang
Dong, Xiaoyi
Cao, Yuhang
Wang, Jiaqi
Liu, Ziwei
author_facet Dong, Yuhao
Tian, Shulin
Liu, Shuai
Ding, Shuangrui
Zang, Yuhang
Dong, Xiaoyi
Cao, Yuhang
Wang, Jiaqi
Liu, Ziwei
contents Despite the growing video understanding capabilities of recent Multimodal Large Language Models (MLLMs), existing video benchmarks primarily assess understanding based on models' static, internal knowledge, rather than their ability to learn and adapt from dynamic, novel contexts from few examples. To bridge this gap, we present Demo-driven Video In-Context Learning, a novel task focused on learning from in-context demonstrations to answer questions about the target videos. Alongside this, we propose Demo-ICL-Bench, a challenging benchmark designed to evaluate demo-driven video in-context learning capabilities. Demo-ICL-Bench is constructed from 1200 instructional YouTube videos with associated questions, from which two types of demonstrations are derived: (i) summarizing video subtitles for text demonstration; and (ii) corresponding instructional videos as video demonstrations. To effectively tackle this new challenge, we develop Demo-ICL, an MLLM with a two-stage training strategy: video-supervised fine-tuning and information-assisted direct preference optimization, jointly enhancing the model's ability to learn from in-context examples. Extensive experiments with state-of-the-art MLLMs confirm the difficulty of Demo-ICL-Bench, demonstrate the effectiveness of Demo-ICL, and thereby unveil future research directions.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08439
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
Dong, Yuhao
Tian, Shulin
Liu, Shuai
Ding, Shuangrui
Zang, Yuhang
Dong, Xiaoyi
Cao, Yuhang
Wang, Jiaqi
Liu, Ziwei
Computer Vision and Pattern Recognition
Despite the growing video understanding capabilities of recent Multimodal Large Language Models (MLLMs), existing video benchmarks primarily assess understanding based on models' static, internal knowledge, rather than their ability to learn and adapt from dynamic, novel contexts from few examples. To bridge this gap, we present Demo-driven Video In-Context Learning, a novel task focused on learning from in-context demonstrations to answer questions about the target videos. Alongside this, we propose Demo-ICL-Bench, a challenging benchmark designed to evaluate demo-driven video in-context learning capabilities. Demo-ICL-Bench is constructed from 1200 instructional YouTube videos with associated questions, from which two types of demonstrations are derived: (i) summarizing video subtitles for text demonstration; and (ii) corresponding instructional videos as video demonstrations. To effectively tackle this new challenge, we develop Demo-ICL, an MLLM with a two-stage training strategy: video-supervised fine-tuning and information-assisted direct preference optimization, jointly enhancing the model's ability to learn from in-context examples. Extensive experiments with state-of-the-art MLLMs confirm the difficulty of Demo-ICL-Bench, demonstrate the effectiveness of Demo-ICL, and thereby unveil future research directions.
title Demo-ICL: In-Context Learning for Procedural Video Knowledge Acquisition
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.08439