CoVR-2: Automatic Data Construction for Composed Video Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Ventura, Lucas, Yang, Antoine, Schmid, Cordelia, Varol, Gül |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
by: Ventura, Lucas, et al.
Published: (2025)
by: Ventura, Lucas, et al.
Published: (2025)
CoVR-R:Reason-Aware Composed Video Retrieval
by: Thawakar, Omkar, et al.
Published: (2026)
by: Thawakar, Omkar, et al.
Published: (2026)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Reason-Then-Retrieve for CoVR-R with Structured Edit Prompts and Dense-Sparse Fusion
by: Liu, DongQing, et al.
Published: (2026)
by: Liu, DongQing, et al.
Published: (2026)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
A Cross-Dataset Study for Text-based 3D Human Motion Retrieval
by: Bensabath, Léore, et al.
Published: (2024)
by: Bensabath, Léore, et al.
Published: (2024)
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Large-scale Pre-training for Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2025)
by: Kazakos, Evangelos, et al.
Published: (2025)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
Retrieval-Enhanced Contrastive Vision-Text Models
by: Iscen, Ahmet, et al.
Published: (2023)
by: Iscen, Ahmet, et al.
Published: (2023)
Automatic Synthesis of High-Quality Triplet Data for Composed Image Retrieval
by: Li, Haiwen, et al.
Published: (2025)
by: Li, Haiwen, et al.
Published: (2025)
InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
by: Zhang, Yangsong, et al.
Published: (2025)
by: Zhang, Yangsong, et al.
Published: (2025)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
by: Han, Gyuwon, et al.
Published: (2026)
by: Han, Gyuwon, et al.
Published: (2026)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
Composed Object Retrieval: Object-level Retrieval via Composed Expressions
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
by: Ghosh, Partha, et al.
Published: (2024)
by: Ghosh, Partha, et al.
Published: (2024)
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
by: Serussi, Gabriele, et al.
Published: (2026)
by: Serussi, Gabriele, et al.
Published: (2026)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
by: Pacaud, Paul, et al.
Published: (2025)
by: Pacaud, Paul, et al.
Published: (2025)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
SINC: Spatial Composition of 3D Human Motions for Simultaneous Action Generation
by: Athanasiou, Nikos, et al.
Published: (2023)
by: Athanasiou, Nikos, et al.
Published: (2023)
Dense Optical Tracking: Connecting the Dots
by: Moing, Guillaume Le, et al.
Published: (2023)
by: Moing, Guillaume Le, et al.
Published: (2023)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
Composed Video Retrieval via Enriched Context and Discriminative Embeddings
by: Thawakar, Omkar, et al.
Published: (2024)
by: Thawakar, Omkar, et al.
Published: (2024)
Beyond Simple Edits: Composed Video Retrieval with Dense Modifications
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
by: Gupta, Animesh, et al.
Published: (2025)
by: Gupta, Animesh, et al.
Published: (2025)
Lost in Translation, Found in Context: Sign Language Translation with Contextual Cues
by: Jang, Youngjoon, et al.
Published: (2025)
by: Jang, Youngjoon, et al.
Published: (2025)
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval
by: Hummel, Thomas, et al.
Published: (2024)
by: Hummel, Thomas, et al.
Published: (2024)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
by: Chen, Zining, et al.
Published: (2025)
by: Chen, Zining, et al.
Published: (2025)
Text-Driven 3D Hand Motion Generation from Sign Language Data
by: Bensabath, Léore, et al.
Published: (2025)
by: Bensabath, Léore, et al.
Published: (2025)
HORT: Monocular Hand-held Objects Reconstruction with Transformers
by: Chen, Zerui, et al.
Published: (2025)
by: Chen, Zerui, et al.
Published: (2025)
Learning Correlation Structures for Vision Transformers
by: Kim, Manjin, et al.
Published: (2024)
by: Kim, Manjin, et al.
Published: (2024)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
Similar Items
-
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
by: Ventura, Lucas, et al.
Published: (2025) -
CoVR-R:Reason-Aware Composed Video Retrieval
by: Thawakar, Omkar, et al.
Published: (2026) -
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024) -
Reason-Then-Retrieve for CoVR-R with Structured Edit Prompts and Dense-Sparse Fusion
by: Liu, DongQing, et al.
Published: (2026) -
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)