The Limits of Learning from Pictures and Text: Vision-Language Models and Embodied Scene Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rosenberg, Gillian, Stadhard, Skylar, Hansen, Bruce C., Greene, Michelle R.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!