Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Shvetsova, Nina, Nagrani, Arsha, Schiele, Bernt, Kuehne, Hilde, Rupprecht, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023)
by: Kahatapitiya, Kumara, et al.
Published: (2023)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
by: Sheludzko, Siarhei, et al.
Published: (2026)
by: Sheludzko, Siarhei, et al.
Published: (2026)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
MINERVA-Cultural: A Benchmark for Cultural and Multilingual Long Video Reasoning
by: Singh, Darshan, et al.
Published: (2026)
by: Singh, Darshan, et al.
Published: (2026)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
CAViAR: Critic-Augmented Video Agentic Reasoning
by: Menon, Sachit, et al.
Published: (2025)
by: Menon, Sachit, et al.
Published: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
by: Min, Juhong, et al.
Published: (2024)
by: Min, Juhong, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
by: Khandelwal, Eshika, et al.
Published: (2025)
by: Khandelwal, Eshika, et al.
Published: (2025)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
AutoAD III: The Prequel -- Back to the Pixels
by: Han, Tengda, et al.
Published: (2024)
by: Han, Tengda, et al.
Published: (2024)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024)
by: Ahmed, Noor, et al.
Published: (2024)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
by: Dong, Jiangxin, et al.
Published: (2021)
by: Dong, Jiangxin, et al.
Published: (2021)
Scribbles for All: Benchmarking Scribble Supervised Segmentation Across Datasets
by: Boettcher, Wolfgang, et al.
Published: (2024)
by: Boettcher, Wolfgang, et al.
Published: (2024)
Probing into Camera Control of Video Models
by: Hou, Chen, et al.
Published: (2026)
by: Hou, Chen, et al.
Published: (2026)
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
by: Böhle, Moritz, et al.
Published: (2021)
by: Böhle, Moritz, et al.
Published: (2021)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
by: Alshami, Eyad, et al.
Published: (2025)
by: Alshami, Eyad, et al.
Published: (2025)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
by: Gairola, Siddhartha, et al.
Published: (2025)
by: Gairola, Siddhartha, et al.
Published: (2025)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
by: Nath, Mriganka, et al.
Published: (2026)
by: Nath, Mriganka, et al.
Published: (2026)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
by: Shi, Shaoshuai, et al.
Published: (2023)
by: Shi, Shaoshuai, et al.
Published: (2023)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
by: Das, Anurag, et al.
Published: (2024)
by: Das, Anurag, et al.
Published: (2024)
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding
by: Nagrani, Arsha, et al.
Published: (2026)
by: Nagrani, Arsha, et al.
Published: (2026)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
by: Hu, Xinting, et al.
Published: (2025)
by: Hu, Xinting, et al.
Published: (2025)
SimNP: Learning Self-Similarity Priors Between Neural Points
by: Wewer, Christopher, et al.
Published: (2023)
by: Wewer, Christopher, et al.
Published: (2023)
Sp2360: Sparse-view 360 Scene Reconstruction using Cascaded 2D Diffusion Priors
by: Paul, Soumava, et al.
Published: (2024)
by: Paul, Soumava, et al.
Published: (2024)
Convolutional Differentiable Logic Gate Networks
by: Petersen, Felix, et al.
Published: (2024)
by: Petersen, Felix, et al.
Published: (2024)
Dual Guidance Semi-Supervised Action Detection
by: Singh, Ankit, et al.
Published: (2025)
by: Singh, Ankit, et al.
Published: (2025)
Similar Items
-
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023) -
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025) -
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025) -
VicTR: Video-conditioned Text Representations for Activity Recognition
by: Kahatapitiya, Kumara, et al.
Published: (2023) -
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
by: Sheludzko, Siarhei, et al.
Published: (2026)