VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Munasinghe, Shehan, Gani, Hanan, Zhu, Wenqi, Cao, Jiale, Xing, Eric, Khan, Fahad Shahbaz, Khan, Salman
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!