Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lee, Seokmin, Lee, Yunghee, Pak, Byeonghyun, Woo, Byeongju
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!

Documents similaires