DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Yuzhong, Liu, Feng, Liu, Yue, Liao, Mingxiang, Gong, Chen, Ye, Qixiang, Wan, Fang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916637205069824
author Zhao, Yuzhong
Liu, Feng
Liu, Yue
Liao, Mingxiang
Gong, Chen
Ye, Qixiang
Wan, Fang
author_facet Zhao, Yuzhong
Liu, Feng
Liu, Yue
Liao, Mingxiang
Gong, Chen
Ye, Qixiang
Wan, Fang
contents One fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which hinders them to find out precise language descriptions. In this study, we propose a DynRefer approach, to pursue high-accuracy region-level referring through mimicking the resolution adaptability of human visual cognition. During training, DynRefer stochastically aligns language descriptions of multimodal tasks with images of multiple resolutions, which are constructed by nesting a set of random views around the referred region. During inference, DynRefer performs selectively multimodal referring by sampling proper region representations for tasks from the nested views based on image and task priors. This allows the visual information for referring to better match human preferences, thereby improving the representational adaptability of region-level multimodal models. Experiments show that DynRefer brings mutual improvement upon broad tasks including region-level captioning, open-vocabulary region recognition and attribute detection. Furthermore, DynRefer achieves state-of-the-art results on multiple region-level multimodal tasks using a single model. Code is available at https://github.com/callsys/DynRefer.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16071
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
Zhao, Yuzhong
Liu, Feng
Liu, Yue
Liao, Mingxiang
Gong, Chen
Ye, Qixiang
Wan, Fang
Computer Vision and Pattern Recognition
One fundamental task of multimodal models is to translate referred image regions to human preferred language descriptions. Existing methods, however, ignore the resolution adaptability needs of different tasks, which hinders them to find out precise language descriptions. In this study, we propose a DynRefer approach, to pursue high-accuracy region-level referring through mimicking the resolution adaptability of human visual cognition. During training, DynRefer stochastically aligns language descriptions of multimodal tasks with images of multiple resolutions, which are constructed by nesting a set of random views around the referred region. During inference, DynRefer performs selectively multimodal referring by sampling proper region representations for tasks from the nested views based on image and task priors. This allows the visual information for referring to better match human preferences, thereby improving the representational adaptability of region-level multimodal models. Experiments show that DynRefer brings mutual improvement upon broad tasks including region-level captioning, open-vocabulary region recognition and attribute detection. Furthermore, DynRefer achieves state-of-the-art results on multiple region-level multimodal tasks using a single model. Code is available at https://github.com/callsys/DynRefer.
title DynRefer: Delving into Region-level Multimodal Tasks via Dynamic Resolution
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.16071