VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Shenghui, Li, Po-han, Chinchali, Sandeep, Topcu, Ufuk
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914051798335488
author Chen, Shenghui
Li, Po-han
Chinchali, Sandeep
Topcu, Ufuk
author_facet Chen, Shenghui
Li, Po-han
Chinchali, Sandeep
Topcu, Ufuk
contents Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries' utility in downstream tasks. We address these gaps with Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method that scores VLM outputs using two metrics: grounding (how well the summary aligns with visual content) and utility (how informative it is for the task). VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on LearningPaper24, SUTD-TrafficQA, and LongVideoBench show that summaries selected by VIBE consistently improve performance-boosting task accuracy by up to 61.23% and reducing response time by 75.77% compared to naive VLM summaries or raw video.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17423
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
Chen, Shenghui
Li, Po-han
Chinchali, Sandeep
Topcu, Ufuk
Computer Vision and Pattern Recognition
Human-Computer Interaction
Information Theory
Many decision-making tasks, where both accuracy and efficiency matter, still require human supervision. For example, tasks like traffic officers reviewing hour-long dashcam footage or researchers screening conference videos can benefit from concise summaries that reduce cognitive load and save time. Yet current vision-language models (VLMs) often produce verbose, redundant outputs that hinder task performance. Existing video caption evaluation depends on costly human annotations and overlooks the summaries' utility in downstream tasks. We address these gaps with Video-to-text Information Bottleneck Evaluation (VIBE), an annotation-free method that scores VLM outputs using two metrics: grounding (how well the summary aligns with visual content) and utility (how informative it is for the task). VIBE selects from randomly sampled VLM outputs by ranking them according to the two scores to support effective human decision-making. Human studies on LearningPaper24, SUTD-TrafficQA, and LongVideoBench show that summaries selected by VIBE consistently improve performance-boosting task accuracy by up to 61.23% and reducing response time by 75.77% compared to naive VLM summaries or raw video.
title VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
topic Computer Vision and Pattern Recognition
Human-Computer Interaction
Information Theory
url https://arxiv.org/abs/2505.17423