VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Vasu, Pavan Kumar Anasosalu, Koc, Cem, Faghri, Fartash, Li, Chun-Liang, Feng, Bo, Lai, Zhengfeng, Cao, Meng, Tuzel, Oncel, Pouransari, Hadi
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910192555261952
author Vasu, Pavan Kumar Anasosalu
Koc, Cem
Faghri, Fartash
Li, Chun-Liang
Feng, Bo
Lai, Zhengfeng
Cao, Meng
Tuzel, Oncel
Pouransari, Hadi
author_facet Vasu, Pavan Kumar Anasosalu
Koc, Cem
Faghri, Fartash
Li, Chun-Liang
Feng, Bo
Lai, Zhengfeng
Cao, Meng
Tuzel, Oncel
Pouransari, Hadi
contents Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Existing VLM frameworks predominantly assess models in offline settings. In contrast, the performance of a streaming VLM depends on additional metrics beyond pure video understanding, including proactiveness, which reflects the timeliness of the model's responses, and consistency, which captures the robustness of its responses over time. To address this limitation, we propose VSAS-Bench, a new framework and benchmark for Visual Streaming Assistants. In contrast to prior benchmarks that primarily employ single-turn question answering on video inputs, VSAS-Bench features temporally dense annotations with over 18,000 annotations across diverse input domains and task types. We introduce standardized synchronous and asynchronous evaluation protocols, along with metrics that isolate and measure distinct capabilities of streaming VLMs. Using this framework, we conduct large-scale evaluations of recent video and streaming VLMs, analyzing the accuracy-latency trade-off under key design factors such as memory buffer length, memory access policy, and input resolution, yielding several practical insights. Finally, we show empirically that conventional VLMs can be adapted to streaming settings without additional training, and demonstrate that these adapted models outperform recent streaming VLMs. For example, Qwen3-VL-4B surpasses Dispider, the best streaming VLM on our benchmark, by 3% under the asynchronous protocol. The benchmark and code will be available at https://github.com/apple/ml-vsas-bench.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07634
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
Vasu, Pavan Kumar Anasosalu
Koc, Cem
Faghri, Fartash
Li, Chun-Liang
Feng, Bo
Lai, Zhengfeng
Cao, Meng
Tuzel, Oncel
Pouransari, Hadi
Computer Vision and Pattern Recognition
Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Existing VLM frameworks predominantly assess models in offline settings. In contrast, the performance of a streaming VLM depends on additional metrics beyond pure video understanding, including proactiveness, which reflects the timeliness of the model's responses, and consistency, which captures the robustness of its responses over time. To address this limitation, we propose VSAS-Bench, a new framework and benchmark for Visual Streaming Assistants. In contrast to prior benchmarks that primarily employ single-turn question answering on video inputs, VSAS-Bench features temporally dense annotations with over 18,000 annotations across diverse input domains and task types. We introduce standardized synchronous and asynchronous evaluation protocols, along with metrics that isolate and measure distinct capabilities of streaming VLMs. Using this framework, we conduct large-scale evaluations of recent video and streaming VLMs, analyzing the accuracy-latency trade-off under key design factors such as memory buffer length, memory access policy, and input resolution, yielding several practical insights. Finally, we show empirically that conventional VLMs can be adapted to streaming settings without additional training, and demonstrate that these adapted models outperform recent streaming VLMs. For example, Qwen3-VL-4B surpasses Dispider, the best streaming VLM on our benchmark, by 3% under the asynchronous protocol. The benchmark and code will be available at https://github.com/apple/ml-vsas-bench.
title VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.07634