VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Shenghui, Li, Po-han, Chinchali, Sandeep, Topcu, Ufuk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
by: Li, Po-han, et al.
Published: (2026)
by: Li, Po-han, et al.
Published: (2026)
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
by: Chen, Shenghui, et al.
Published: (2024)
by: Chen, Shenghui, et al.
Published: (2024)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
by: Li, Po-han, et al.
Published: (2024)
by: Li, Po-han, et al.
Published: (2024)
Task-aware Distributed Source Coding under Dynamic Bandwidth
by: Li, Po-han, et al.
Published: (2023)
by: Li, Po-han, et al.
Published: (2023)
Spacewalker: Traversing Representation Spaces for Fast Interactive Exploration and Annotation of Unstructured Data
by: Heine, Lukas, et al.
Published: (2024)
by: Heine, Lukas, et al.
Published: (2024)
Human-Agent Cooperation in Games under Incomplete Information through Natural Language Communication
by: Chen, Shenghui, et al.
Published: (2024)
by: Chen, Shenghui, et al.
Published: (2024)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
by: Gutiérrez, Juan, et al.
Published: (2025)
by: Gutiérrez, Juan, et al.
Published: (2025)
Text-to-Image Representativity Fairness Evaluation Framework
by: Yamani, Asma, et al.
Published: (2024)
by: Yamani, Asma, et al.
Published: (2024)
Learning Annotation Consensus for Continuous Emotion Recognition
by: Shoer, Ibrahim, et al.
Published: (2025)
by: Shoer, Ibrahim, et al.
Published: (2025)
Confidence Contours: Uncertainty-Aware Annotation for Medical Semantic Segmentation
by: Ye, Andre, et al.
Published: (2023)
by: Ye, Andre, et al.
Published: (2023)
Weak-Annotation of HAR Datasets using Vision Foundation Models
by: Bock, Marius, et al.
Published: (2024)
by: Bock, Marius, et al.
Published: (2024)
HaDR: Applying Domain Randomization for Generating Synthetic Multimodal Dataset for Hand Instance Segmentation in Cluttered Industrial Environments
by: Grushko, Stefan, et al.
Published: (2023)
by: Grushko, Stefan, et al.
Published: (2023)
Positive-First Most Ambiguous: A Simple Active Learning Criterion for Interactive Retrieval of Rare Categories
by: Zaher, Kawtar, et al.
Published: (2026)
by: Zaher, Kawtar, et al.
Published: (2026)
Revisiting Human-in-the-Loop Object Retrieval with Pre-Trained Vision Transformers
by: Zaher, Kawtar, et al.
Published: (2026)
by: Zaher, Kawtar, et al.
Published: (2026)
ERPA: Efficient RPA Model Integrating OCR and LLMs for Intelligent Document Processing
by: Abdellaif, Osama, et al.
Published: (2024)
by: Abdellaif, Osama, et al.
Published: (2024)
CLAS: A Machine Learning Enhanced Framework for Exploring Large 3D Design Datasets
by: Zhang, XiuYu, et al.
Published: (2024)
by: Zhang, XiuYu, et al.
Published: (2024)
A Versatile Dataset of Mouse and Eye Movements on Search Engine Results Pages
by: Latifzadeh, Kayhan, et al.
Published: (2025)
by: Latifzadeh, Kayhan, et al.
Published: (2025)
Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations
by: Lachmann, Tim, et al.
Published: (2026)
by: Lachmann, Tim, et al.
Published: (2026)
Civiverse: A Dataset for Analyzing User Engagement with Open-Source Text-to-Image Models
by: Palmini, Maria-Teresa De Rosa, et al.
Published: (2024)
by: Palmini, Maria-Teresa De Rosa, et al.
Published: (2024)
VideoMix: Aggregating How-To Videos for Task-Oriented Learning
by: Yang, Saelyne, et al.
Published: (2025)
by: Yang, Saelyne, et al.
Published: (2025)
Is Medieval Distant Viewing Possible? : Extending and Enriching Annotation of Legacy Image Collections using Visual Analytics
by: Meinecke, Christofer, et al.
Published: (2022)
by: Meinecke, Christofer, et al.
Published: (2022)
Towards Geographic Inclusion in the Evaluation of Text-to-Image Models
by: Hall, Melissa, et al.
Published: (2024)
by: Hall, Melissa, et al.
Published: (2024)
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
Semantic Draw Engineering for Text-to-Image Creation
by: Li, Yang, et al.
Published: (2023)
by: Li, Yang, et al.
Published: (2023)
VisionCAD: An Integration-Free Radiology Copilot Framework
by: Li, Jiaming, et al.
Published: (2025)
by: Li, Jiaming, et al.
Published: (2025)
Bridging Text and Image for Artist Style Transfer via Contrastive Learning
by: Liu, Zhi-Song, et al.
Published: (2024)
by: Liu, Zhi-Song, et al.
Published: (2024)
GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
by: Ramachandran, Akhil, et al.
Published: (2026)
by: Ramachandran, Akhil, et al.
Published: (2026)
POET: Supporting Prompting Creativity and Personalization with Automated Expansion of Text-to-Image Generation
by: Han, Evans Xu, et al.
Published: (2025)
by: Han, Evans Xu, et al.
Published: (2025)
NarrativeBridge: Enhancing Video Captioning with Causal-Temporal Narrative
by: Nadeem, Asmar, et al.
Published: (2024)
by: Nadeem, Asmar, et al.
Published: (2024)
Reframe Anything: LLM Agent for Open World Video Reframing
by: Cao, Jiawang, et al.
Published: (2024)
by: Cao, Jiawang, et al.
Published: (2024)
Vid2Coach: Transforming How-To Videos into Task Assistants
by: Huh, Mina, et al.
Published: (2025)
by: Huh, Mina, et al.
Published: (2025)
Analyzing Swimming Performance Using Drone Captured Aerial Videos
by: Tran, Thu, et al.
Published: (2025)
by: Tran, Thu, et al.
Published: (2025)
Video Joint-Embedding Predictive Architectures for Facial Expression Recognition
by: Eing, Lennart, et al.
Published: (2026)
by: Eing, Lennart, et al.
Published: (2026)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
by: Mazumdar, Amrita, et al.
Published: (2026)
by: Mazumdar, Amrita, et al.
Published: (2026)
Alt4Blind: A User Interface to Simplify Charts Alt-Text Creation
by: Moured, Omar, et al.
Published: (2024)
by: Moured, Omar, et al.
Published: (2024)
Design First, Code Later: Aesthetically Pleasing Template-Free Slides Generation
by: Cui, Zhiyao, et al.
Published: (2026)
by: Cui, Zhiyao, et al.
Published: (2026)
ConceptFactory: Facilitate 3D Object Knowledge Annotation with Object Conceptualization
by: Sun, Jianhua, et al.
Published: (2024)
by: Sun, Jianhua, et al.
Published: (2024)
MVTN: A Multiscale Video Transformer Network for Hand Gesture Recognition
by: Garg, Mallika, et al.
Published: (2024)
by: Garg, Mallika, et al.
Published: (2024)
Designing Multi-Robot Ground Video Sensemaking with Public Safety Professionals
by: Zhou, Puqi, et al.
Published: (2026)
by: Zhou, Puqi, et al.
Published: (2026)
Similar Items
-
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
by: Li, Po-han, et al.
Published: (2026) -
CSA: Data-efficient Mapping of Unimodal Features to Multimodal Features
by: Li, Po-han, et al.
Published: (2024) -
Human-Agent Coordination in Games under Incomplete Information via Multi-Step Intent
by: Chen, Shenghui, et al.
Published: (2024) -
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
by: Li, Po-han, et al.
Published: (2024) -
Task-aware Distributed Source Coding under Dynamic Bandwidth
by: Li, Po-han, et al.
Published: (2023)