Efficient Token Compression for Vision Transformer with Spatial Information Preserved
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Junzhu, Shen, Yang, Guo, Jinyang, Yao, Yazhou, Hua, Xiansheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
by: Yin, Jianjian, et al.
Published: (2025)
by: Yin, Jianjian, et al.
Published: (2025)
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
by: Zhang, Wang, et al.
Published: (2024)
by: Zhang, Wang, et al.
Published: (2024)
Context Guided Transformer Entropy Modeling for Video Compression
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
Serial Low-rank Adaptation of Vision Transformer
by: Zhong, Houqiang, et al.
Published: (2025)
by: Zhong, Houqiang, et al.
Published: (2025)
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
by: Wang, Yingquan, et al.
Published: (2024)
by: Wang, Yingquan, et al.
Published: (2024)
A Preprocessing Framework for Video Machine Vision under Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
GeoLink: A 3D-Aware Framework Towards Better Generalization in Cross-View Geo-Localization
by: Zhang, Hongyang, et al.
Published: (2026)
by: Zhang, Hongyang, et al.
Published: (2026)
Learning Compact Vision Tokens for Efficient Large Multimodal Models
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
by: Yao, Ting, et al.
Published: (2024)
by: Yao, Ting, et al.
Published: (2024)
Enhancing 3D Gaussian Splatting Compression via Spatial Condition-based Prediction
by: Ma, Jingui, et al.
Published: (2025)
by: Ma, Jingui, et al.
Published: (2025)
Diversity-Guided MLP Reduction for Efficient Large Vision Transformers
by: Shen, Chengchao, et al.
Published: (2025)
by: Shen, Chengchao, et al.
Published: (2025)
Relating CNN-Transformer Fusion Network for Change Detection
by: Gao, Yuhao, et al.
Published: (2024)
by: Gao, Yuhao, et al.
Published: (2024)
Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
by: Zhan, Wengyi, et al.
Published: (2025)
by: Zhan, Wengyi, et al.
Published: (2025)
Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023)
by: Baraldi, Lorenzo, et al.
Published: (2023)
ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
by: Zhu, Xilei, et al.
Published: (2024)
by: Zhu, Xilei, et al.
Published: (2024)
SPC-NeRF: Spatial Predictive Compression for Voxel Based Radiance Field
by: Song, Zetian, et al.
Published: (2024)
by: Song, Zetian, et al.
Published: (2024)
Rendering-Oriented 3D Point Cloud Attribute Compression using Sparse Tensor-based Transformer
by: Huo, Xiao, et al.
Published: (2024)
by: Huo, Xiao, et al.
Published: (2024)
AToken: A Unified Tokenizer for Vision
by: Lu, Jiasen, et al.
Published: (2025)
by: Lu, Jiasen, et al.
Published: (2025)
ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision
by: Liang, Xie, et al.
Published: (2025)
by: Liang, Xie, et al.
Published: (2025)
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
by: Zeng, YangChen
Published: (2025)
by: Zeng, YangChen
Published: (2025)
A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
by: Huang, He, et al.
Published: (2024)
by: Huang, He, et al.
Published: (2024)
Efficient and Generic Point Model for Lossless Point Cloud Attribute Compression
by: You, Kang, et al.
Published: (2024)
by: You, Kang, et al.
Published: (2024)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
by: Zheng, Guangting, et al.
Published: (2025)
by: Zheng, Guangting, et al.
Published: (2025)
GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
by: Wang, Longan, et al.
Published: (2025)
by: Wang, Longan, et al.
Published: (2025)
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
by: Li, Bingzhou, et al.
Published: (2026)
by: Li, Bingzhou, et al.
Published: (2026)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
by: Wang, Feng, et al.
Published: (2024)
by: Wang, Feng, et al.
Published: (2024)
SmartSplat: Feature-Smart Gaussians for Scalable Compression of Ultra-High-Resolution Images
by: Li, Linfei, et al.
Published: (2025)
by: Li, Linfei, et al.
Published: (2025)
SMC++: Masked Learning of Unsupervised Video Semantic Compression
by: Tian, Yuan, et al.
Published: (2024)
by: Tian, Yuan, et al.
Published: (2024)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
by: Zhang, Xuesong, et al.
Published: (2024)
by: Zhang, Xuesong, et al.
Published: (2024)
A Multimodal Transformer for Live Streaming Highlight Prediction
by: Deng, Jiaxin, et al.
Published: (2024)
by: Deng, Jiaxin, et al.
Published: (2024)
A Tri-Dynamic Preprocessing Framework for UGC Video Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
Rate-aware Compression for NeRF-based Volumetric Video
by: Zhang, Zhiyu, et al.
Published: (2024)
by: Zhang, Zhiyu, et al.
Published: (2024)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
by: Chen, Yangneng, et al.
Published: (2026)
by: Chen, Yangneng, et al.
Published: (2026)
Hybrid Local-Global Context Learning for Neural Video Compression
by: Zhai, Yongqi, et al.
Published: (2024)
by: Zhai, Yongqi, et al.
Published: (2024)
Multimodal Learning on Low-Quality Data with Conformal Predictive Self-Calibration
by: Jiang, Xun, et al.
Published: (2026)
by: Jiang, Xun, et al.
Published: (2026)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
by: Qian, Shun, et al.
Published: (2024)
by: Qian, Shun, et al.
Published: (2024)
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
by: Cai, Qi, et al.
Published: (2025)
by: Cai, Qi, et al.
Published: (2025)
BiTDiff: Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion
by: Jia, Tianzhi, et al.
Published: (2026)
by: Jia, Tianzhi, et al.
Published: (2026)
Reversing the Damage: A QP-Aware Transformer-Diffusion Approach for 8K Video Restoration under Codec Compression
by: Dehaghi, Ali Mollaahmadi, et al.
Published: (2024)
by: Dehaghi, Ali Mollaahmadi, et al.
Published: (2024)
Similar Items
-
Semi-supervised Semantic Segmentation with Multi-Constraint Consistency Learning
by: Yin, Jianjian, et al.
Published: (2025) -
A Light-weight Transformer-based Self-supervised Matching Network for Heterogeneous Images
by: Zhang, Wang, et al.
Published: (2024) -
Context Guided Transformer Entropy Modeling for Video Compression
by: Tong, Junlong, et al.
Published: (2025) -
Serial Low-rank Adaptation of Vision Transformer
by: Zhong, Houqiang, et al.
Published: (2025) -
Other Tokens Matter: Exploring Global and Local Features of Vision Transformers for Object Re-Identification
by: Wang, Yingquan, et al.
Published: (2024)