HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
Fuente:
arXiv
Saved in:
| Main Authors: | Shvetsova, Nina, Kukleva, Anna, Hong, Xudong, Rupprecht, Christian, Schiele, Bernt, Kuehne, Hilde |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
by: Sheludzko, Siarhei, et al.
Published: (2026)
by: Sheludzko, Siarhei, et al.
Published: (2026)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024)
by: Ahmed, Noor, et al.
Published: (2024)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
by: Das, Anurag, et al.
Published: (2026)
by: Das, Anurag, et al.
Published: (2026)
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
What, when, and where? -- Self-Supervised Spatio-Temporal Grounding in Untrimmed Multi-Action Videos from Narrated Instructions
by: Chen, Brian, et al.
Published: (2023)
by: Chen, Brian, et al.
Published: (2023)
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
by: Kukleva, Anna, et al.
Published: (2024)
by: Kukleva, Anna, et al.
Published: (2024)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
by: Gairola, Siddhartha, et al.
Published: (2025)
by: Gairola, Siddhartha, et al.
Published: (2025)
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
by: Dong, Jiangxin, et al.
Published: (2021)
by: Dong, Jiangxin, et al.
Published: (2021)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
Probing into Camera Control of Video Models
by: Hou, Chen, et al.
Published: (2026)
by: Hou, Chen, et al.
Published: (2026)
Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
by: Böhle, Moritz, et al.
Published: (2021)
by: Böhle, Moritz, et al.
Published: (2021)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
by: Arya, Shreyash, et al.
Published: (2024)
by: Arya, Shreyash, et al.
Published: (2024)
Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
by: Shi, Shaoshuai, et al.
Published: (2023)
by: Shi, Shaoshuai, et al.
Published: (2023)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
by: Nath, Mriganka, et al.
Published: (2026)
by: Nath, Mriganka, et al.
Published: (2026)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
by: Alshami, Eyad, et al.
Published: (2025)
by: Alshami, Eyad, et al.
Published: (2025)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
by: Das, Anurag, et al.
Published: (2024)
by: Das, Anurag, et al.
Published: (2024)
Studying How to Efficiently and Effectively Guide Models with Explanations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
by: Wang, Yunnan, et al.
Published: (2026)
by: Wang, Yunnan, et al.
Published: (2026)
Convolutional Differentiable Logic Gate Networks
by: Petersen, Felix, et al.
Published: (2024)
by: Petersen, Felix, et al.
Published: (2024)
SimNP: Learning Self-Similarity Priors Between Neural Points
by: Wewer, Christopher, et al.
Published: (2023)
by: Wewer, Christopher, et al.
Published: (2023)
Sp2360: Sparse-view 360 Scene Reconstruction using Cascaded 2D Diffusion Priors
by: Paul, Soumava, et al.
Published: (2024)
by: Paul, Soumava, et al.
Published: (2024)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
by: Hu, Xinting, et al.
Published: (2025)
by: Hu, Xinting, et al.
Published: (2025)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
by: Bousselham, Walid, et al.
Published: (2026)
by: Bousselham, Walid, et al.
Published: (2026)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Similar Items
-
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025) -
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025) -
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
by: Sheludzko, Siarhei, et al.
Published: (2026) -
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024) -
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025)