Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Phat, Cheung, Ngai-Man |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging
by: Nguyen, Phat, et al.
Published: (2026)
by: Nguyen, Phat, et al.
Published: (2026)
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026)
by: Liu, Chao, et al.
Published: (2026)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
Text to Image Generation and Editing: A Survey
by: Yang, Pengfei, et al.
Published: (2025)
by: Yang, Pengfei, et al.
Published: (2025)
Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
by: Zeng, Fanhu, et al.
Published: (2025)
by: Zeng, Fanhu, et al.
Published: (2025)
Frequency Masking for Universal Deepfake Detection
by: Doloriel, Chandler Timm, et al.
Published: (2024)
by: Doloriel, Chandler Timm, et al.
Published: (2024)
Vision Transformers on the Edge: A Comprehensive Survey of Model Compression and Acceleration Strategies
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
MRT: Learning Compact Representations with Mixed RWKV-Transformer for Extreme Image Compression
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
AIR: Zero-shot Generative Model Adaptation with Iterative Refinement
by: Liu, Guimeng, et al.
Published: (2025)
by: Liu, Guimeng, et al.
Published: (2025)
Extreme Model Compression for Edge Vision-Language Models: Sparse Temporal Token Fusion and Adaptive Neural Compression
by: Tanvir, Md Tasnin, et al.
Published: (2025)
by: Tanvir, Md Tasnin, et al.
Published: (2025)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
by: Mao, Junzhu, et al.
Published: (2025)
by: Mao, Junzhu, et al.
Published: (2025)
Urban Air Temperature Prediction using Conditional Diffusion Models
by: Dai, Siyang, et al.
Published: (2024)
by: Dai, Siyang, et al.
Published: (2024)
A Survey on Generative Modeling with Limited Data, Few Shots, and Zero Shot
by: Abdollahzadeh, Milad, et al.
Published: (2023)
by: Abdollahzadeh, Milad, et al.
Published: (2023)
Comprehensive Survey of Model Compression and Speed up for Vision Transformers
by: Chen, Feiyang, et al.
Published: (2024)
by: Chen, Feiyang, et al.
Published: (2024)
SegMaFormer: A Hybrid State-Space and Transformer Model for Efficient Segmentation
by: Nguyen, Duy D., et al.
Published: (2026)
by: Nguyen, Duy D., et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression
by: Su, Zunhai, et al.
Published: (2026)
by: Su, Zunhai, et al.
Published: (2026)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
by: Wang, Ziyao, et al.
Published: (2026)
by: Wang, Ziyao, et al.
Published: (2026)
Analyzing Transformer Models and Knowledge Distillation Approaches for Image Captioning on Edge AI
by: Kwok, Wing Man Casca, et al.
Published: (2025)
by: Kwok, Wing Man Casca, et al.
Published: (2025)
SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
by: Li, Zhaoxu, et al.
Published: (2025)
by: Li, Zhaoxu, et al.
Published: (2025)
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
by: Qiao, Bingtian, et al.
Published: (2026)
by: Qiao, Bingtian, et al.
Published: (2026)
Exploring Self-Supervised Vision Transformers for Deepfake Detection: A Comparative Analysis
by: Nguyen, Huy H., et al.
Published: (2024)
by: Nguyen, Huy H., et al.
Published: (2024)
ButterflyViT: 354$\times$ Expert Compression for Edge Vision Transformers
by: Karmore, Aryan
Published: (2026)
by: Karmore, Aryan
Published: (2026)
A Survey of Token Compression for Efficient Multimodal Large Language Models
by: Shao, Kele, et al.
Published: (2025)
by: Shao, Kele, et al.
Published: (2025)
Vision Transformer with Super Token Sampling
by: Huang, Huaibo, et al.
Published: (2022)
by: Huang, Huaibo, et al.
Published: (2022)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
by: Yu, Hanxun, et al.
Published: (2026)
by: Yu, Hanxun, et al.
Published: (2026)
RMT: Retentive Networks Meet Vision Transformers
by: Fan, Qihang, et al.
Published: (2023)
by: Fan, Qihang, et al.
Published: (2023)
On the Vulnerability of Skip Connections to Model Inversion Attacks
by: Koh, Jun Hao, et al.
Published: (2024)
by: Koh, Jun Hao, et al.
Published: (2024)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Compress image to patches for Vision Transformer
by: Zhao, Xinfeng, et al.
Published: (2025)
by: Zhao, Xinfeng, et al.
Published: (2025)
SUCCESS-GS: Survey of Compactness and Compression for Efficient Static and Dynamic Gaussian Splatting
by: Youn, Seokhyun, et al.
Published: (2025)
by: Youn, Seokhyun, et al.
Published: (2025)
Vision-centric Token Compression in Large Language Model
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
Wavelet-Based Image Tokenizer for Vision Transformers
by: Zhu, Zhenhai, et al.
Published: (2024)
by: Zhu, Zhenhai, et al.
Published: (2024)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
by: Shams, Montasir, et al.
Published: (2025)
by: Shams, Montasir, et al.
Published: (2025)
Towards Sustainable Universal Deepfake Detection with Frequency-Domain Masking
by: Doloriel, Chandler Timm C., et al.
Published: (2025)
by: Doloriel, Chandler Timm C., et al.
Published: (2025)
One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
by: Miwa, Keita, et al.
Published: (2025)
by: Miwa, Keita, et al.
Published: (2025)
CoCAViT: Compact Vision Transformer with Robust Global Coordination
by: Wang, Xuyang, et al.
Published: (2025)
by: Wang, Xuyang, et al.
Published: (2025)
Similar Items
-
Joint Architecture-Token-Bitwidth Multi-Axis Optimization of Vision Transformers for Semiconductor IC Packaging
by: Nguyen, Phat, et al.
Published: (2026) -
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026) -
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024) -
Text to Image Generation and Editing: A Survey
by: Yang, Pengfei, et al.
Published: (2025) -
Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
by: Zeng, Fanhu, et al.
Published: (2025)