Saved in:
Bibliographic Details
Main Authors: Yu, Qiyang, Fang, Yu, Li, Tianrui, Cao, Xuemei, Chen, Yan, Li, Jianghao, Min, Fan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.19021
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909920674185216
author Yu, Qiyang
Fang, Yu
Li, Tianrui
Cao, Xuemei
Chen, Yan
Li, Jianghao
Min, Fan
author_facet Yu, Qiyang
Fang, Yu
Li, Tianrui
Cao, Xuemei
Chen, Yan
Li, Jianghao
Min, Fan
contents Vision Transformers (ViTs) have demonstrated strong capabilities in capturing global dependencies but often struggle to efficiently represent fine-grained local details. Existing multi-scale approaches alleviate this issue by integrating hierarchical or hybrid features; however, they rely on fixed patch sizes and introduce redundant computation. To address these limitations, we propose Granularity-driven Vision Transformer (Grc-ViT), a dynamic coarse-to-fine framework that adaptively adjusts visual granularity based on image complexity. It comprises two key stages: (1) Coarse Granularity Evaluation module, which assesses visual complexity using edge density, entropy, and frequency-domain cues to estimate suitable patch and window sizes; (2) Fine-grained Refinement module, which refines attention computation according to the selected granularity, enabling efficient and precise feature learning. Two learnable parameters, α and \b{eta}, are optimized end-to-end to balance global reasoning and local perception. Comprehensive evaluations demonstrate that Grc-ViT enhances fine-grained discrimination while achieving a superior trade-off between accuracy and computational efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19021
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
Yu, Qiyang
Fang, Yu
Li, Tianrui
Cao, Xuemei
Chen, Yan
Li, Jianghao
Min, Fan
Computer Vision and Pattern Recognition
Vision Transformers (ViTs) have demonstrated strong capabilities in capturing global dependencies but often struggle to efficiently represent fine-grained local details. Existing multi-scale approaches alleviate this issue by integrating hierarchical or hybrid features; however, they rely on fixed patch sizes and introduce redundant computation. To address these limitations, we propose Granularity-driven Vision Transformer (Grc-ViT), a dynamic coarse-to-fine framework that adaptively adjusts visual granularity based on image complexity. It comprises two key stages: (1) Coarse Granularity Evaluation module, which assesses visual complexity using edge density, entropy, and frequency-domain cues to estimate suitable patch and window sizes; (2) Fine-grained Refinement module, which refines attention computation according to the selected granularity, enabling efficient and precise feature learning. Two learnable parameters, α and \b{eta}, are optimized end-to-end to balance global reasoning and local perception. Comprehensive evaluations demonstrate that Grc-ViT enhances fine-grained discrimination while achieving a superior trade-off between accuracy and computational efficiency.
title Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.19021