Random Wins All: Rethinking Grouping Strategies for Vision Tokens
Fuente:
arXiv
Salvato in:
| Autori principali: | Fan, Qihang, Ai, Yuang, Huang, Huaibo, He, Ran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rectifying Magnitude Neglect in Linear Attention
di: Fan, Qihang, et al.
Pubblicazione: (2025)
di: Fan, Qihang, et al.
Pubblicazione: (2025)
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
LoRA-IR: Taming Low-Rank Experts for Efficient All-in-One Image Restoration
di: Ai, Yuang, et al.
Pubblicazione: (2024)
di: Ai, Yuang, et al.
Pubblicazione: (2024)
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention
di: Ai, Yuang, et al.
Pubblicazione: (2025)
di: Ai, Yuang, et al.
Pubblicazione: (2025)
DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
di: Ai, Yuang, et al.
Pubblicazione: (2025)
di: Ai, Yuang, et al.
Pubblicazione: (2025)
Lightweight Vision Transformer with Bidirectional Interaction
di: Fan, Qihang, et al.
Pubblicazione: (2023)
di: Fan, Qihang, et al.
Pubblicazione: (2023)
Expand and Prune: Maximizing Trajectory Diversity for Effective GRPO in Generative Models
di: Ge, Shiran, et al.
Pubblicazione: (2025)
di: Ge, Shiran, et al.
Pubblicazione: (2025)
Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
di: Ai, Yuang, et al.
Pubblicazione: (2023)
di: Ai, Yuang, et al.
Pubblicazione: (2023)
Breaking the Low-Rank Dilemma of Linear Attention
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
Advancing Vision Transformer with Enhanced Spatial Priors
di: Fan, Qihang, et al.
Pubblicazione: (2026)
di: Fan, Qihang, et al.
Pubblicazione: (2026)
RMT: Retentive Networks Meet Vision Transformers
di: Fan, Qihang, et al.
Pubblicazione: (2023)
di: Fan, Qihang, et al.
Pubblicazione: (2023)
Vision Transformer with Sparse Scan Prior
di: Zhang, Yuguang, et al.
Pubblicazione: (2024)
di: Zhang, Yuguang, et al.
Pubblicazione: (2024)
Uncertainty-Aware Source-Free Adaptive Image Super-Resolution with Wavelet Augmentation Transformer
di: Ai, Yuang, et al.
Pubblicazione: (2023)
di: Ai, Yuang, et al.
Pubblicazione: (2023)
Vision Transformer with Super Token Sampling
di: Huang, Huaibo, et al.
Pubblicazione: (2022)
di: Huang, Huaibo, et al.
Pubblicazione: (2022)
ViTAR: Vision Transformer with Any Resolution
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
di: Han, Xiaotian, et al.
Pubblicazione: (2024)
di: Han, Xiaotian, et al.
Pubblicazione: (2024)
DreamClear: High-Capacity Real-World Image Restoration with Privacy-Safe Dataset Curation
di: Ai, Yuang, et al.
Pubblicazione: (2024)
di: Ai, Yuang, et al.
Pubblicazione: (2024)
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
di: Guo, Yiwei, et al.
Pubblicazione: (2026)
ZePo: Zero-Shot Portrait Stylization with Faster Sampling
di: Liu, Jin, et al.
Pubblicazione: (2024)
di: Liu, Jin, et al.
Pubblicazione: (2024)
NOFT: Test-Time Noise Finetune via Information Bottleneck for Highly Correlated Asset Creation
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
DeVAn: Dense Video Annotation for Video-Language Models
di: Liu, Tingkai, et al.
Pubblicazione: (2023)
di: Liu, Tingkai, et al.
Pubblicazione: (2023)
BitDance: Scaling Autoregressive Generative Models with Binary Tokens
di: Ai, Yuang, et al.
Pubblicazione: (2026)
di: Ai, Yuang, et al.
Pubblicazione: (2026)
Think 360°: Evaluating the Width-centric Reasoning Capability of MLLMs Beyond Depth
di: Chen, Mingrui, et al.
Pubblicazione: (2026)
di: Chen, Mingrui, et al.
Pubblicazione: (2026)
Marmot: Object-Level Self-Correction via Multi-Agent Reasoning
di: Sun, Jiayang, et al.
Pubblicazione: (2025)
di: Sun, Jiayang, et al.
Pubblicazione: (2025)
InfoBFR: Real-World Blind Face Restoration via Information Bottleneck
di: Gao, Nan, et al.
Pubblicazione: (2025)
di: Gao, Nan, et al.
Pubblicazione: (2025)
Parallel Augmentation and Dual Enhancement for Occluded Person Re-identification
di: Wang, Zi, et al.
Pubblicazione: (2022)
di: Wang, Zi, et al.
Pubblicazione: (2022)
MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding
di: Bai, Purui, et al.
Pubblicazione: (2026)
di: Bai, Purui, et al.
Pubblicazione: (2026)
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
di: Chen, Mingrui, et al.
Pubblicazione: (2025)
di: Chen, Mingrui, et al.
Pubblicazione: (2025)
TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection
di: Zhou, Qihang, et al.
Pubblicazione: (2025)
di: Zhou, Qihang, et al.
Pubblicazione: (2025)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
di: He, Ju, et al.
Pubblicazione: (2025)
di: He, Ju, et al.
Pubblicazione: (2025)
Rethinking Token Reduction for Large Vision-Language Models
di: Wang, Yi, et al.
Pubblicazione: (2026)
di: Wang, Yi, et al.
Pubblicazione: (2026)
Win-Win: Training High-Resolution Vision Transformers from Two Windows
di: Leroy, Vincent, et al.
Pubblicazione: (2023)
di: Leroy, Vincent, et al.
Pubblicazione: (2023)
DiffMAC: Diffusion Manifold Hallucination Correction for High Generalization Blind Face Restoration
di: Gao, Nan, et al.
Pubblicazione: (2024)
di: Gao, Nan, et al.
Pubblicazione: (2024)
Visual Anchors Are Strong Information Aggregators For Multimodal Large Language Model
di: Liu, Haogeng, et al.
Pubblicazione: (2024)
di: Liu, Haogeng, et al.
Pubblicazione: (2024)
GenVideoLens: Where LVLMs Fall Short in AI-Generated Video Detection?
di: Zou, Yueying, et al.
Pubblicazione: (2026)
di: Zou, Yueying, et al.
Pubblicazione: (2026)
Rethinking Scanning Strategies with Vision Mamba in Semantic Segmentation of Remote Sensing Imagery: An Experimental Study
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
di: Takezoe, Rinyoichi, et al.
Pubblicazione: (2026)
di: Takezoe, Rinyoichi, et al.
Pubblicazione: (2026)
UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model
di: Zhuang, Shaobin, et al.
Pubblicazione: (2026)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2026)
Randomized Autoregressive Visual Generation
di: Yu, Qihang, et al.
Pubblicazione: (2024)
di: Yu, Qihang, et al.
Pubblicazione: (2024)
Survey on AI-Generated Media Detection: From Non-MLLM to MLLM
di: Zou, Yueying, et al.
Pubblicazione: (2025)
di: Zou, Yueying, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Rectifying Magnitude Neglect in Linear Attention
di: Fan, Qihang, et al.
Pubblicazione: (2025) -
Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens
di: Fan, Qihang, et al.
Pubblicazione: (2024) -
LoRA-IR: Taming Low-Rank Experts for Efficient All-in-One Image Restoration
di: Ai, Yuang, et al.
Pubblicazione: (2024) -
Breaking Complexity Barriers: High-Resolution Image Restoration with Rank Enhanced Linear Attention
di: Ai, Yuang, et al.
Pubblicazione: (2025) -
DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling
di: Ai, Yuang, et al.
Pubblicazione: (2025)