xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ryoo, Michael S., Zhou, Honglu, Kendre, Shrikant, Qin, Can, Xue, Le, Shu, Manli, Park, Jongwoo, Ranasinghe, Kanchana, Savarese, Silvio, Xu, Ran, Xiong, Caiming, Niebles, Juan Carlos
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!