Large-scale Pre-training for Grounded Video Caption Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kazakos, Evangelos, Schmid, Cordelia, Sivic, Josef |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024)
by: Kazakos, Evangelos, et al.
Published: (2024)
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023)
by: Zhou, Xingyi, et al.
Published: (2023)
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024)
by: Chen, Shizhe, et al.
Published: (2024)
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024)
by: Zhou, Xingyi, et al.
Published: (2024)
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)
by: Fiastre, Gabriel, et al.
Published: (2025)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
by: Wysoczańska, Monika, et al.
Published: (2025)
by: Wysoczańska, Monika, et al.
Published: (2025)
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
by: Zheng, Minghang, et al.
Published: (2024)
by: Zheng, Minghang, et al.
Published: (2024)
VoCap: Video Object Captioning and Segmentation from Any Prompt
by: Uijlings, Jasper, et al.
Published: (2025)
by: Uijlings, Jasper, et al.
Published: (2025)
GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
by: Souček, Tomáš, et al.
Published: (2023)
by: Souček, Tomáš, et al.
Published: (2023)
BrickNet: Graph-Backed Generative Brick Assembly
by: Kulits, Peter, et al.
Published: (2026)
by: Kulits, Peter, et al.
Published: (2026)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
by: Khan, Zeeshan, et al.
Published: (2025)
by: Khan, Zeeshan, et al.
Published: (2025)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
by: Ventura, Lucas, et al.
Published: (2025)
by: Ventura, Lucas, et al.
Published: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
by: Ghosh, Partha, et al.
Published: (2024)
by: Ghosh, Partha, et al.
Published: (2024)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
by: Ventura, Lucas, et al.
Published: (2023)
by: Ventura, Lucas, et al.
Published: (2023)
What Are You Doing? A Closer Look at Controllable Human Video Generation
by: Bugliarello, Emanuele, et al.
Published: (2025)
by: Bugliarello, Emanuele, et al.
Published: (2025)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
by: Zhuang, Weijun, et al.
Published: (2025)
by: Zhuang, Weijun, et al.
Published: (2025)
EditDuet: A Multi-Agent System for Video Non-Linear Editing
by: Sandoval-Castaneda, Marcelo, et al.
Published: (2025)
by: Sandoval-Castaneda, Marcelo, et al.
Published: (2025)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
by: Caron, Mathilde, et al.
Published: (2024)
by: Caron, Mathilde, et al.
Published: (2024)
TIM: A Time Interval Machine for Audio-Visual Action Recognition
by: Chalk, Jacob, et al.
Published: (2024)
by: Chalk, Jacob, et al.
Published: (2024)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
by: Chen, Shizhe, et al.
Published: (2025)
by: Chen, Shizhe, et al.
Published: (2025)
Foundation Model for Endoscopy Video Analysis via Large-scale Self-supervised Pre-train
by: Wang, Zhao, et al.
Published: (2023)
by: Wang, Zhao, et al.
Published: (2023)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
DreamLIP: Language-Image Pre-training with Long Captions
by: Zheng, Kecheng, et al.
Published: (2024)
by: Zheng, Kecheng, et al.
Published: (2024)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
Canny2Palm: Realistic and Controllable Palmprint Generation for Large-scale Pre-training
by: Lan, Xingzeng, et al.
Published: (2025)
by: Lan, Xingzeng, et al.
Published: (2025)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
by: Kim, Jae Myung, et al.
Published: (2025)
by: Kim, Jae Myung, et al.
Published: (2025)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
by: Souček, Tomáš, et al.
Published: (2024)
by: Souček, Tomáš, et al.
Published: (2024)
Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning
by: Bardhan, Jai, et al.
Published: (2026)
by: Bardhan, Jai, et al.
Published: (2026)
Med-GLIP: Advancing Medical Language-Image Pre-training with Large-scale Grounded Dataset
by: Deng, Ziye, et al.
Published: (2025)
by: Deng, Ziye, et al.
Published: (2025)
Dense Optical Tracking: Connecting the Dots
by: Moing, Guillaume Le, et al.
Published: (2023)
by: Moing, Guillaume Le, et al.
Published: (2023)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
by: Chen, Shizhe, et al.
Published: (2026)
by: Chen, Shizhe, et al.
Published: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
by: Garcia, Ricardo, et al.
Published: (2024)
by: Garcia, Ricardo, et al.
Published: (2024)
mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus
by: Futeral, Matthieu, et al.
Published: (2024)
by: Futeral, Matthieu, et al.
Published: (2024)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
by: Zhang, Xinsong, et al.
Published: (2025)
by: Zhang, Xinsong, et al.
Published: (2025)
NimbleD: Enhancing Self-supervised Monocular Depth Estimation with Pseudo-labels and Large-scale Video Pre-training
by: Luginov, Albert, et al.
Published: (2024)
by: Luginov, Albert, et al.
Published: (2024)
6D Object Pose Tracking in Internet Videos for Robotic Manipulation
by: Ponimatkin, Georgy, et al.
Published: (2025)
by: Ponimatkin, Georgy, et al.
Published: (2025)
AutoTVG: A New Vision-language Pre-training Paradigm for Temporal Video Grounding
by: Zhang, Xing, et al.
Published: (2024)
by: Zhang, Xing, et al.
Published: (2024)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
by: Zhang, Shi-Xue, et al.
Published: (2025)
by: Zhang, Shi-Xue, et al.
Published: (2025)
Similar Items
-
Grounded Video Caption Generation
by: Kazakos, Evangelos, et al.
Published: (2024) -
Dense Video Object Captioning from Disjoint Supervision
by: Zhou, Xingyi, et al.
Published: (2023) -
SUGAR: Pre-training 3D Visual Representations for Robotics
by: Chen, Shizhe, et al.
Published: (2024) -
Streaming Dense Video Captioning
by: Zhou, Xingyi, et al.
Published: (2024) -
CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
by: Fiastre, Gabriel, et al.
Published: (2025)