How Can Objects Help Video-Language Understanding?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zitian, Wang, Shijie, Cho, Junho, Yoo, Jaewook, Sun, Chen
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909722585595904
author Tang, Zitian
Wang, Shijie
Cho, Junho
Yoo, Jaewook
Sun, Chen
author_facet Tang, Zitian
Wang, Shijie
Cho, Junho
Yoo, Jaewook
Sun, Chen
contents Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves provide strong empirical performances for understanding tasks, despite missing fine-grained spatiotemporal information. To answer this question, we introduce ObjectMLLM, a framework capable of leveraging arbitrary computer vision algorithm to extract and integrate structured visual representation. Through extensive evaluations on six video question answering benchmarks, we confirm that explicit integration of object-centric representation remains necessary. Surprisingly, we observe that the simple approach of quantizing the continuous, structured object information and representing them as plain text performs the best, offering a data-efficient approach to integrate other visual perception modules into MLLM design. Our code and models are released at https://github.com/brown-palm/ObjectMLLM.
format Preprint
id arxiv_https___arxiv_org_abs_2504_07454
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Can Objects Help Video-Language Understanding?
Tang, Zitian
Wang, Shijie
Cho, Junho
Yoo, Jaewook
Sun, Chen
Computer Vision and Pattern Recognition
Do we still need to represent objects explicitly in multimodal large language models (MLLMs)? To one extreme, pre-trained encoders convert images into visual tokens, with which objects and spatiotemporal relationships may be implicitly modeled. To the other extreme, image captions by themselves provide strong empirical performances for understanding tasks, despite missing fine-grained spatiotemporal information. To answer this question, we introduce ObjectMLLM, a framework capable of leveraging arbitrary computer vision algorithm to extract and integrate structured visual representation. Through extensive evaluations on six video question answering benchmarks, we confirm that explicit integration of object-centric representation remains necessary. Surprisingly, we observe that the simple approach of quantizing the continuous, structured object information and representing them as plain text performs the best, offering a data-efficient approach to integrate other visual perception modules into MLLM design. Our code and models are released at https://github.com/brown-palm/ObjectMLLM.
title How Can Objects Help Video-Language Understanding?
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.07454