Saved in:
| Main Authors: | Yu, Qiyang, Fang, Yu, Li, Tianrui, Cao, Xuemei, Chen, Yan, Li, Jianghao, Min, Fan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.19021 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
by: Yu, Qiyang, et al.
Published: (2025)
by: Yu, Qiyang, et al.
Published: (2025)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Brain-Inspired Stepwise Patch Merging for Vision Transformers
by: Yu, Yonghao, et al.
Published: (2024)
by: Yu, Yonghao, et al.
Published: (2024)
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
by: Peng, Jiankun, et al.
Published: (2026)
by: Peng, Jiankun, et al.
Published: (2026)
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
by: Cao, Jianjian, et al.
Published: (2024)
by: Cao, Jianjian, et al.
Published: (2024)
Dynamic Texture Transfer using PatchMatch and Transformers
by: Pu, Guo, et al.
Published: (2024)
by: Pu, Guo, et al.
Published: (2024)
MGRQ: Post-Training Quantization For Vision Transformer With Mixed Granularity Reconstruction
by: Yang, Lianwei, et al.
Published: (2024)
by: Yang, Lianwei, et al.
Published: (2024)
Beyond Defenses: Manifold-Aligned Regularization for Intrinsic 3D Point Cloud Robustness
by: Alonso, Pedro, et al.
Published: (2026)
by: Alonso, Pedro, et al.
Published: (2026)
Exploring Vision Transformers for 3D Human Motion-Language Models with Motion Patches
by: Yu, Qing, et al.
Published: (2024)
by: Yu, Qing, et al.
Published: (2024)
Semantic Graph Consistency: Going Beyond Patches for Regularizing Self-Supervised Vision Transformers
by: Devaguptapu, Chaitanya, et al.
Published: (2024)
by: Devaguptapu, Chaitanya, et al.
Published: (2024)
Efficient Deformable ConvNets: Rethinking Dynamic and Sparse Operator for Vision Applications
by: Xiong, Yuwen, et al.
Published: (2024)
by: Xiong, Yuwen, et al.
Published: (2024)
All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning
by: Yang, Zheng, et al.
Published: (2025)
by: Yang, Zheng, et al.
Published: (2025)
A New Perspective on Privacy Protection in Federated Learning with Granular-Ball Computing
by: Lai, Guannan, et al.
Published: (2025)
by: Lai, Guannan, et al.
Published: (2025)
A Timely Survey on Vision Transformer for Deepfake Detection
by: Wang, Zhikan, et al.
Published: (2024)
by: Wang, Zhikan, et al.
Published: (2024)
Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
by: Lin, Ci-Siang, et al.
Published: (2025)
by: Lin, Ci-Siang, et al.
Published: (2025)
Context-Aware Token Selection and Packing for Enhanced Vision Transformer
by: Zhang, Tianyi, et al.
Published: (2024)
by: Zhang, Tianyi, et al.
Published: (2024)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
by: Shu, Yuyang, et al.
Published: (2024)
by: Shu, Yuyang, et al.
Published: (2024)
Hi-ResNet: Edge Detail Enhancement for High-Resolution Remote Sensing Segmentation
by: Chen, Yuxia, et al.
Published: (2023)
by: Chen, Yuxia, et al.
Published: (2023)
Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer
by: Wu, Junyi, et al.
Published: (2024)
by: Wu, Junyi, et al.
Published: (2024)
Entropy Guided Dynamic Patch Segmentation for Time Series Transformers
by: Abeywickrama, Sachith, et al.
Published: (2025)
by: Abeywickrama, Sachith, et al.
Published: (2025)
Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
by: Murrugarra-LLerena, Jeffri, et al.
Published: (2025)
by: Murrugarra-LLerena, Jeffri, et al.
Published: (2025)
Split Adaptation for Pre-trained Vision Transformers
by: Wang, Lixu, et al.
Published: (2025)
by: Wang, Lixu, et al.
Published: (2025)
Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
by: Zhu, Qinfeng, et al.
Published: (2024)
by: Zhu, Qinfeng, et al.
Published: (2024)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
by: Zhao, Min, et al.
Published: (2025)
by: Zhao, Min, et al.
Published: (2025)
Rethinking Occlusion in FER: A Semantic-Aware Perspective and Go Beyond
by: Zhai, Huiyu, et al.
Published: (2025)
by: Zhai, Huiyu, et al.
Published: (2025)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
by: Huang, Shaofei, et al.
Published: (2025)
by: Huang, Shaofei, et al.
Published: (2025)
Rethinking Patch Dependence for Masked Autoencoders
by: Fu, Letian, et al.
Published: (2024)
by: Fu, Letian, et al.
Published: (2024)
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
by: Yin, Jianghao, et al.
Published: (2026)
by: Yin, Jianghao, et al.
Published: (2026)
HiAP: A Multi-Granular Stochastic Auto-Pruning Framework for Vision Transformers
by: Li, Andy, et al.
Published: (2026)
by: Li, Andy, et al.
Published: (2026)
Multi-Granularity Language-Guided Training for Multi-Object Tracking
by: Li, Yuhao, et al.
Published: (2024)
by: Li, Yuhao, et al.
Published: (2024)
Patch-based Selection and Refinement for Early Object Detection
by: Zhang, Tianyi, et al.
Published: (2023)
by: Zhang, Tianyi, et al.
Published: (2023)
Generalized Correspondence Matching via Flexible Hierarchical Refinement and Patch Descriptor Distillation
by: Han, Yu, et al.
Published: (2024)
by: Han, Yu, et al.
Published: (2024)
Stratify or Die: Rethinking Data Splits in Image Segmentation
by: Jami, Naga Venkata Sai Jitin, et al.
Published: (2025)
by: Jami, Naga Venkata Sai Jitin, et al.
Published: (2025)
Unseen No More: Unlocking the Potential of CLIP for Generative Zero-shot HOI Detection
by: Guo, Yixin, et al.
Published: (2024)
by: Guo, Yixin, et al.
Published: (2024)
HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via Transformer
by: Zhang, Zhang, et al.
Published: (2025)
by: Zhang, Zhang, et al.
Published: (2025)
Rethinking Query-based Transformer for Continual Image Segmentation
by: Zhu, Yuchen, et al.
Published: (2025)
by: Zhu, Yuchen, et al.
Published: (2025)
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
by: Su, Yongyi, et al.
Published: (2025)
by: Su, Yongyi, et al.
Published: (2025)
From Per-Image Low-Rank to Encoding Mismatch: Rethinking Feature Distillation in Vision Transformers
by: Tian, Huiyuan, et al.
Published: (2025)
by: Tian, Huiyuan, et al.
Published: (2025)
On the Evaluation and Refinement of Vision-Language Instruction Tuning Datasets
by: Liao, Ning, et al.
Published: (2023)
by: Liao, Ning, et al.
Published: (2023)
Similar Items
-
Granular Computing-driven SAM: From Coarse-to-Fine Guidance for Prompt-Free Segmentation
by: Yu, Qiyang, et al.
Published: (2025) -
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024) -
Brain-Inspired Stepwise Patch Merging for Vision Transformers
by: Yu, Yonghao, et al.
Published: (2024) -
Dynamic Topology Awareness: Breaking the Granularity Rigidity in Vision-Language Navigation
by: Peng, Jiankun, et al.
Published: (2026) -
ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
by: Li, Yifan, et al.
Published: (2025)