Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Qihang, Huang, Huaibo, Chen, Mingrui, He, Ran |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
RMT: Retentive Networks Meet Vision Transformers
by: Fan, Qihang, et al.
Published: (2023)
by: Fan, Qihang, et al.
Published: (2023)
Advancing Vision Transformer with Enhanced Spatial Priors
by: Fan, Qihang, et al.
Published: (2026)
by: Fan, Qihang, et al.
Published: (2026)
Lightweight Vision Transformer with Bidirectional Interaction
by: Fan, Qihang, et al.
Published: (2023)
by: Fan, Qihang, et al.
Published: (2023)
Breaking the Low-Rank Dilemma of Linear Attention
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
Vision Transformer with Sparse Scan Prior
by: Zhang, Yuguang, et al.
Published: (2024)
by: Zhang, Yuguang, et al.
Published: (2024)
Rectifying Magnitude Neglect in Linear Attention
by: Fan, Qihang, et al.
Published: (2025)
by: Fan, Qihang, et al.
Published: (2025)
Vision Transformer with Super Token Sampling
by: Huang, Huaibo, et al.
Published: (2022)
by: Huang, Huaibo, et al.
Published: (2022)
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
by: Ge, Shiran, et al.
Published: (2025)
by: Ge, Shiran, et al.
Published: (2025)
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention
by: Ai, Yuang, et al.
Published: (2025)
by: Ai, Yuang, et al.
Published: (2025)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
by: Chen, Mingrui, et al.
Published: (2026)
by: Chen, Mingrui, et al.
Published: (2026)
ViTAR: Vision Transformer with Any Resolution
by: Fan, Qihang, et al.
Published: (2024)
by: Fan, Qihang, et al.
Published: (2024)
DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
by: Ai, Yuang, et al.
Published: (2025)
by: Ai, Yuang, et al.
Published: (2025)
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
by: Chen, Mingrui, et al.
Published: (2025)
by: Chen, Mingrui, et al.
Published: (2025)
LoRA-IR: Taming Low-Rank Experts for Efficient All-in-One Image Restoration
by: Ai, Yuang, et al.
Published: (2024)
by: Ai, Yuang, et al.
Published: (2024)
ZePo: Zero-Shot Portrait Stylization with Faster Sampling
by: Liu, Jin, et al.
Published: (2024)
by: Liu, Jin, et al.
Published: (2024)
NOFT: Test-Time Noise Finetune via Information Bottleneck for Highly Correlated Asset Creation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Simple Yet Effective Selective Imputation for Incomplete Multi-view Clustering
by: Xu, Cai, et al.
Published: (2025)
by: Xu, Cai, et al.
Published: (2025)
DeVAn: Dense Video Annotation for Video-Language Models
by: Liu, Tingkai, et al.
Published: (2023)
by: Liu, Tingkai, et al.
Published: (2023)
Agglomerative Token Clustering
by: Haurum, Joakim Bruslund, et al.
Published: (2024)
by: Haurum, Joakim Bruslund, et al.
Published: (2024)
Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
by: Ai, Yuang, et al.
Published: (2023)
by: Ai, Yuang, et al.
Published: (2023)
Marmot: Object-Level Self-Correction via Multi-Agent Reasoning
by: Sun, Jiayang, et al.
Published: (2025)
by: Sun, Jiayang, et al.
Published: (2025)
InfoBFR: Real-World Blind Face Restoration via Information Bottleneck
by: Gao, Nan, et al.
Published: (2025)
by: Gao, Nan, et al.
Published: (2025)
Parallel Augmentation and Dual Enhancement for Occluded Person Re-identification
by: Wang, Zi, et al.
Published: (2022)
by: Wang, Zi, et al.
Published: (2022)
Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
by: Ai, Yuang, et al.
Published: (2023)
by: Ai, Yuang, et al.
Published: (2023)
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models
by: Li, Chenyang, et al.
Published: (2026)
by: Li, Chenyang, et al.
Published: (2026)
MacNet: An End-to-End Manifold-Constrained Adaptive Clustering Network for Interpretable Whole Slide Image Classification
by: Ma, Mingrui, et al.
Published: (2026)
by: Ma, Mingrui, et al.
Published: (2026)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
by: Bai, Purui, et al.
Published: (2026)
by: Bai, Purui, et al.
Published: (2026)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
by: He, Ju, et al.
Published: (2025)
by: He, Ju, et al.
Published: (2025)
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
by: Zhou, Qihang, et al.
Published: (2025)
by: Zhou, Qihang, et al.
Published: (2025)
InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
by: Han, Xiaotian, et al.
Published: (2024)
by: Han, Xiaotian, et al.
Published: (2024)
TCFormer: Visual Recognition via Token Clustering Transformer
by: Zeng, Wang, et al.
Published: (2024)
by: Zeng, Wang, et al.
Published: (2024)
Contrastive Prompt Clustering for Weakly Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2025)
by: Wu, Wangyu, et al.
Published: (2025)
ClusterMark: Towards Robust Watermarking for Autoregressive Image Generators with Visual Token Clustering
by: Lukovnikov, Denis, et al.
Published: (2025)
by: Lukovnikov, Denis, et al.
Published: (2025)
A Lightweight Clustering Framework for Unsupervised Semantic Segmentation
by: Cheung, Yau Shing Jonathan, et al.
Published: (2023)
by: Cheung, Yau Shing Jonathan, et al.
Published: (2023)
LVFace: Progressive Cluster Optimization for Large Vision Models in Face Recognition
by: You, Jinghan, et al.
Published: (2025)
by: You, Jinghan, et al.
Published: (2025)
Prompt Categories Cluster for Weakly Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2024)
by: Wu, Wangyu, et al.
Published: (2024)
Towards Efficient and Effective Deep Clustering with Dynamic Grouping and Prototype Aggregation
by: Zhang, Haixin, et al.
Published: (2024)
by: Zhang, Haixin, et al.
Published: (2024)
FTCFormer: Fuzzy Token Clustering Transformer for Image Classification
by: Bao, Muyi, et al.
Published: (2025)
by: Bao, Muyi, et al.
Published: (2025)
Similar Items
-
Random Wins All: Rethinking Grouping Strategies for Vision Tokens
by: Fan, Qihang, et al.
Published: (2026) -
RMT: Retentive Networks Meet Vision Transformers
by: Fan, Qihang, et al.
Published: (2023) -
Advancing Vision Transformer with Enhanced Spatial Priors
by: Fan, Qihang, et al.
Published: (2026) -
Lightweight Vision Transformer with Bidirectional Interaction
by: Fan, Qihang, et al.
Published: (2023) -
Breaking the Low-Rank Dilemma of Linear Attention
by: Fan, Qihang, et al.
Published: (2024)