Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Qiyang, Fang, Yu, Li, Tianrui, Cao, Xuemei, Chen, Yan, Li, Jianghao, Min, Fan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
di: Yu, Qiyang, et al.
Pubblicazione: (2025)
di: Yu, Qiyang, et al.
Pubblicazione: (2025)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
Brain-Inspired Stepwise Patch Merging for Vision Transformers
di: Yu, Yonghao, et al.
Pubblicazione: (2024)
di: Yu, Yonghao, et al.
Pubblicazione: (2024)
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
di: Peng, Jiankun, et al.
Pubblicazione: (2026)
di: Peng, Jiankun, et al.
Pubblicazione: (2026)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
di: Li, Yifan, et al.
Pubblicazione: (2025)
di: Li, Yifan, et al.
Pubblicazione: (2025)
Dynamic Texture Transfer using PatchMatch and Transformers
di: Pu, Guo, et al.
Pubblicazione: (2024)
di: Pu, Guo, et al.
Pubblicazione: (2024)
MGRQ: Post-Training Quantization For Vision Transformer With Mixed Granularity Reconstruction
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
di: Yang, Lianwei, et al.
Pubblicazione: (2024)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
di: Cao, Jianjian, et al.
Pubblicazione: (2024)
di: Cao, Jianjian, et al.
Pubblicazione: (2024)
Semantic Graph Consistency: Going Beyond Patches for Regularizing Self-Supervised Vision Transformers
di: Devaguptapu, Chaitanya, et al.
Pubblicazione: (2024)
di: Devaguptapu, Chaitanya, et al.
Pubblicazione: (2024)
Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches
di: Yu, Qing, et al.
Pubblicazione: (2024)
di: Yu, Qing, et al.
Pubblicazione: (2024)
Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness
di: Alonso, Pedro, et al.
Pubblicazione: (2026)
di: Alonso, Pedro, et al.
Pubblicazione: (2026)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
di: Xiong, Yuwen, et al.
Pubblicazione: (2024)
di: Xiong, Yuwen, et al.
Pubblicazione: (2024)
All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning
di: Yang, Zheng, et al.
Pubblicazione: (2025)
di: Yang, Zheng, et al.
Pubblicazione: (2025)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
di: Lin, Ci-Siang, et al.
Pubblicazione: (2025)
di: Lin, Ci-Siang, et al.
Pubblicazione: (2025)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
di: Shu, Yuyang, et al.
Pubblicazione: (2024)
di: Shu, Yuyang, et al.
Pubblicazione: (2024)
Context-Aware Token Selection and Packing for Enhanced Vision Transformer
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
di: Zhang, Tianyi, et al.
Pubblicazione: (2024)
A Timely Survey on Vision Transformer for Deepfake Detection
di: Wang, Zhikan, et al.
Pubblicazione: (2024)
di: Wang, Zhikan, et al.
Pubblicazione: (2024)
Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer
di: Wu, Junyi, et al.
Pubblicazione: (2024)
di: Wu, Junyi, et al.
Pubblicazione: (2024)
Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
di: Murrugarra-LLerena, Jeffri, et al.
Pubblicazione: (2025)
di: Murrugarra-LLerena, Jeffri, et al.
Pubblicazione: (2025)
A New Perspective on Privacy Protection in Federated Learning with Granular-Ball Computing
di: Lai, Guannan, et al.
Pubblicazione: (2025)
di: Lai, Guannan, et al.
Pubblicazione: (2025)
Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go Beyond
di: Zhai, Huiyu, et al.
Pubblicazione: (2025)
di: Zhai, Huiyu, et al.
Pubblicazione: (2025)
Split Adaptation for Pre-trained Vision Transformers
di: Wang, Lixu, et al.
Pubblicazione: (2025)
di: Wang, Lixu, et al.
Pubblicazione: (2025)
Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
Entropy Guided Dynamic Patch Segmentation for Time Series Transformers
di: Abeywickrama, Sachith, et al.
Pubblicazione: (2025)
di: Abeywickrama, Sachith, et al.
Pubblicazione: (2025)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
di: Zhao, Min, et al.
Pubblicazione: (2025)
di: Zhao, Min, et al.
Pubblicazione: (2025)
Hi-ResNet: Edge Detail Enhancement for High-Resolution Remote Sensing Segmentation
di: Chen, Yuxia, et al.
Pubblicazione: (2023)
di: Chen, Yuxia, et al.
Pubblicazione: (2023)
Rethinking Patch Dependence for Masked Autoencoders
di: Fu, Letian, et al.
Pubblicazione: (2024)
di: Fu, Letian, et al.
Pubblicazione: (2024)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
di: Huang, Shaofei, et al.
Pubblicazione: (2025)
di: Huang, Shaofei, et al.
Pubblicazione: (2025)
HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers
di: Li, Andy, et al.
Pubblicazione: (2026)
di: Li, Andy, et al.
Pubblicazione: (2026)
Stratify or Die: Rethinking Data Splits in Image Segmentation
di: Jami, Naga Venkata Sai Jitin, et al.
Pubblicazione: (2025)
di: Jami, Naga Venkata Sai Jitin, et al.
Pubblicazione: (2025)
From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
di: Tian, Huiyuan, et al.
Pubblicazione: (2025)
di: Tian, Huiyuan, et al.
Pubblicazione: (2025)
Multi-Granularity Language-Guided Training for Multi-Object Tracking
di: Li, Yuhao, et al.
Pubblicazione: (2024)
di: Li, Yuhao, et al.
Pubblicazione: (2024)
Rethinking Vision Transformer Depth via Structural Reparameterization
di: Zhou, Chengwei, et al.
Pubblicazione: (2025)
di: Zhou, Chengwei, et al.
Pubblicazione: (2025)
Patch-based Selection and Refinement for Early Object Detection
di: Zhang, Tianyi, et al.
Pubblicazione: (2023)
di: Zhang, Tianyi, et al.
Pubblicazione: (2023)
Rethinking Query-based Transformer for Continual Image Segmentation
di: Zhu, Yuchen, et al.
Pubblicazione: (2025)
di: Zhu, Yuchen, et al.
Pubblicazione: (2025)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
di: Su, Yongyi, et al.
Pubblicazione: (2025)
di: Su, Yongyi, et al.
Pubblicazione: (2025)
Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor Distillation
di: Han, Yu, et al.
Pubblicazione: (2024)
di: Han, Yu, et al.
Pubblicazione: (2024)
Patch-Fool: Are Vision Transformers Always Robust Against Adversarial Perturbations?
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
di: Fu, Yonggan, et al.
Pubblicazione: (2022)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
di: Yin, Jianghao, et al.
Pubblicazione: (2026)
MSVM-UNet: Multi-Scale Vision Mamba UNet for Medical Image Segmentation
di: Chen, Chaowei, et al.
Pubblicazione: (2024)
di: Chen, Chaowei, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
di: Yu, Qiyang, et al.
Pubblicazione: (2025) -
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
di: Gee, Leonidas, et al.
Pubblicazione: (2024) -
Brain-Inspired Stepwise Patch Merging for Vision Transformers
di: Yu, Yonghao, et al.
Pubblicazione: (2024) -
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
di: Peng, Jiankun, et al.
Pubblicazione: (2026) -
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
di: Li, Yifan, et al.
Pubblicazione: (2025)