Do Vision-Language Transformers Exhibit Visual Commonsense? An Empirical Study of VCR

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhenyang, Guo, Yangyang, Wang, Kejie, Chen, Xiaolin, Nie, Liqiang, Kankanhalli, Mohan
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!