Region-based Cluster Discrimination for Visual Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xie, Yin, Yang, Kaicheng, An, Xiang, Wu, Kun, Zhao, Yongle, Deng, Weimo, Ran, Zimin, Wang, Yumeng, Feng, Ziyong, Miles, Roy, Elezi, Ismail, Deng, Jiankang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912504092819456
author Xie, Yin
Yang, Kaicheng
An, Xiang
Wu, Kun
Zhao, Yongle
Deng, Weimo
Ran, Zimin
Wang, Yumeng
Feng, Ziyong
Miles, Roy
Elezi, Ismail
Deng, Jiankang
author_facet Xie, Yin
Yang, Kaicheng
An, Xiang
Wu, Kun
Zhao, Yongle
Deng, Weimo
Ran, Zimin
Wang, Yumeng
Feng, Ziyong
Miles, Roy
Elezi, Ismail
Deng, Jiankang
contents Learning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations constrains their effectiveness for dense prediction tasks, such as grounding, OCR, and segmentation. To address this gap, we introduce Region-Aware Cluster Discrimination (RICE), a novel method that enhances region-level visual and OCR capabilities. We first construct a billion-scale candidate region dataset and propose a Region Transformer layer to extract rich regional semantics. We further design a unified region cluster discrimination loss that jointly supports object and OCR learning within a single classification framework, enabling efficient and scalable distributed training on large-scale data. Extensive experiments show that RICE consistently outperforms previous methods on tasks, including segmentation, dense detection, and visual perception for Multimodal Large Language Models (MLLMs). The pre-trained models have been released at https://github.com/deepglint/MVT.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20025
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Region-based Cluster Discrimination for Visual Representation Learning
Xie, Yin
Yang, Kaicheng
An, Xiang
Wu, Kun
Zhao, Yongle
Deng, Weimo
Ran, Zimin
Wang, Yumeng
Feng, Ziyong
Miles, Roy
Elezi, Ismail
Deng, Jiankang
Computer Vision and Pattern Recognition
Learning visual representations is foundational for a broad spectrum of downstream tasks. Although recent vision-language contrastive models, such as CLIP and SigLIP, have achieved impressive zero-shot performance via large-scale vision-language alignment, their reliance on global representations constrains their effectiveness for dense prediction tasks, such as grounding, OCR, and segmentation. To address this gap, we introduce Region-Aware Cluster Discrimination (RICE), a novel method that enhances region-level visual and OCR capabilities. We first construct a billion-scale candidate region dataset and propose a Region Transformer layer to extract rich regional semantics. We further design a unified region cluster discrimination loss that jointly supports object and OCR learning within a single classification framework, enabling efficient and scalable distributed training on large-scale data. Extensive experiments show that RICE consistently outperforms previous methods on tasks, including segmentation, dense detection, and visual perception for Multimodal Large Language Models (MLLMs). The pre-trained models have been released at https://github.com/deepglint/MVT.
title Region-based Cluster Discrimination for Visual Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2507.20025