TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Zijian, Zheng, Xuhui, Wu, Xuecheng, Peng, Chong, Cao, Xuezhi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!