ViSpeak: Visual Instruction Feedback in Streaming Videos

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Fu, Shenghao, Yang, Qize, Li, Yuan-Ming, Peng, Yi-Xing, Lin, Kun-Yu, Wei, Xihan, Hu, Jian-Fang, Xie, Xiaohua, Zheng, Wei-Shi
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916653908885504
author Fu, Shenghao
Yang, Qize
Li, Yuan-Ming
Peng, Yi-Xing
Lin, Kun-Yu
Wei, Xihan
Hu, Jian-Fang
Xie, Xiaohua
Zheng, Wei-Shi
author_facet Fu, Shenghao
Yang, Qize
Li, Yuan-Ming
Peng, Yi-Xing
Lin, Kun-Yu
Wei, Xihan
Hu, Jian-Fang
Xie, Xiaohua
Zheng, Wei-Shi
contents Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming video understanding from a new perspective and propose a novel task named Visual Instruction Feedback in which models should be aware of visual contents and learn to extract instructions from them. For example, when users wave their hands to agents, agents should recognize the gesture and start conversations with welcome information. Thus, following instructions in visual modality greatly enhances user-agent interactions. To facilitate research, we define seven key subtasks highly relevant to visual modality and collect the ViSpeak-Instruct dataset for training and the ViSpeak-Bench for evaluation. Further, we propose the ViSpeak model, which is a SOTA streaming video understanding LMM with GPT-4o-level performance on various streaming video understanding benchmarks. After finetuning on our ViSpeak-Instruct dataset, ViSpeak is equipped with basic visual instruction feedback ability, serving as a solid baseline for future research.
format Preprint
id arxiv_https___arxiv_org_abs_2503_12769
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ViSpeak: Visual Instruction Feedback in Streaming Videos
Fu, Shenghao
Yang, Qize
Li, Yuan-Ming
Peng, Yi-Xing
Lin, Kun-Yu
Wei, Xihan
Hu, Jian-Fang
Xie, Xiaohua
Zheng, Wei-Shi
Computer Vision and Pattern Recognition
Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming video understanding from a new perspective and propose a novel task named Visual Instruction Feedback in which models should be aware of visual contents and learn to extract instructions from them. For example, when users wave their hands to agents, agents should recognize the gesture and start conversations with welcome information. Thus, following instructions in visual modality greatly enhances user-agent interactions. To facilitate research, we define seven key subtasks highly relevant to visual modality and collect the ViSpeak-Instruct dataset for training and the ViSpeak-Bench for evaluation. Further, we propose the ViSpeak model, which is a SOTA streaming video understanding LMM with GPT-4o-level performance on various streaming video understanding benchmarks. After finetuning on our ViSpeak-Instruct dataset, ViSpeak is equipped with basic visual instruction feedback ability, serving as a solid baseline for future research.
title ViSpeak: Visual Instruction Feedback in Streaming Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.12769