Frequency-Aware Token Reduction for Efficient Vision Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Dong-Jae, Hur, Jiwan, Choi, Jaehyun, Yu, Jaemyung, Kim, Junmo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
Self-supervised Transformation Learning for Equivariant Representations
by: Yu, Jaemyung, et al.
Published: (2025)
by: Yu, Jaemyung, et al.
Published: (2025)
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
by: Choi, Jaehyun, et al.
Published: (2024)
by: Choi, Jaehyun, et al.
Published: (2024)
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
by: Lee, Dong-Jae, et al.
Published: (2026)
by: Lee, Dong-Jae, et al.
Published: (2026)
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)
by: Lee, Dongyeun, et al.
Published: (2025)
Learning Neural Deformation Representation for 4D Dynamic Shape Generation
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
DAM: Domain-Aware Module for Multi-Domain Dataset Condensation
by: Choi, Jaehyun, et al.
Published: (2025)
by: Choi, Jaehyun, et al.
Published: (2025)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
by: Hur, Jiwan, et al.
Published: (2024)
by: Hur, Jiwan, et al.
Published: (2024)
Inlier-Centric Post-Training Quantization for Object Detection Models
by: Kim, Minsu, et al.
Published: (2026)
by: Kim, Minsu, et al.
Published: (2026)
IMSE: Intrinsic Mixture of Spectral Experts Fine-tuning for Test-Time Adaptation
by: Baek, Sunghyun, et al.
Published: (2026)
by: Baek, Sunghyun, et al.
Published: (2026)
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
by: Shin, Youngwoo, et al.
Published: (2026)
by: Shin, Youngwoo, et al.
Published: (2026)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
by: Lee, Dong Hoon, et al.
Published: (2024)
by: Lee, Dong Hoon, et al.
Published: (2024)
Can visual language models resolve textual ambiguity with visual cues? Let visual puns tell you!
by: Chung, Jiwan, et al.
Published: (2024)
by: Chung, Jiwan, et al.
Published: (2024)
Global Geometry Is Not Enough for Vision Representations
by: Chung, Jiwan, et al.
Published: (2026)
by: Chung, Jiwan, et al.
Published: (2026)
Inspecting Explainability of Transformer Models with Additional Statistical Information
by: Nguyen, Hoang C., et al.
Published: (2023)
by: Nguyen, Hoang C., et al.
Published: (2023)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
Beta Sampling is All You Need: Efficient Image Generation Strategy for Diffusion Models using Stepwise Spectral Analysis
by: Lee, Haeil, et al.
Published: (2024)
by: Lee, Haeil, et al.
Published: (2024)
Modeling Stereo-Confidence Out of the End-to-End Stereo-Matching Network via Disparity Plane Sweep
by: Lee, Jae Young, et al.
Published: (2024)
by: Lee, Jae Young, et al.
Published: (2024)
Stereo-Matching Knowledge Distilled Monocular Depth Estimation Filtered by Multiple Disparity Consistency
by: Ka, Woonghyun, et al.
Published: (2024)
by: Ka, Woonghyun, et al.
Published: (2024)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
by: Lee, Gayoung, et al.
Published: (2025)
by: Lee, Gayoung, et al.
Published: (2025)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
The Effects of Mixed Sample Data Augmentation are Class Dependent
by: Lee, Haeil, et al.
Published: (2023)
by: Lee, Haeil, et al.
Published: (2023)
Faster Parameter-Efficient Tuning with Token Redundancy Reduction
by: Kim, Kwonyoung, et al.
Published: (2025)
by: Kim, Kwonyoung, et al.
Published: (2025)
Spectral Vision Transformer for Efficient Tokenization with Limited Data
by: Roberts, Alexandra G., et al.
Published: (2026)
by: Roberts, Alexandra G., et al.
Published: (2026)
LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
by: Shang, Yuzhang, et al.
Published: (2024)
by: Shang, Yuzhang, et al.
Published: (2024)
AT-SNN: Adaptive Tokens for Vision Transformer on Spiking Neural Network
by: Kang, Donghwa, et al.
Published: (2024)
by: Kang, Donghwa, et al.
Published: (2024)
Token Pruning using a Lightweight Background Aware Vision Transformer
by: Sah, Sudhakar, et al.
Published: (2024)
by: Sah, Sudhakar, et al.
Published: (2024)
MATE: Meet At The Embedding -- Connecting Images with Long Texts
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
PAR: Prompt-Aware Token Reduction Method for Efficient Large Multimodal Models
by: Liu, Yingen, et al.
Published: (2024)
by: Liu, Yingen, et al.
Published: (2024)
Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
by: Han, Gyojin, et al.
Published: (2026)
by: Han, Gyojin, et al.
Published: (2026)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
by: Chen, Siyi, et al.
Published: (2026)
by: Chen, Siyi, et al.
Published: (2026)
Test-Time Mixup Augmentation for Data and Class-Specific Uncertainty Estimation in Deep Learning Image Classification
by: Lee, Hansang, et al.
Published: (2022)
by: Lee, Hansang, et al.
Published: (2022)
Beyond Semantics: Disentangling Information Scope in Sparse Autoencoders for CLIP
by: Ro, Yusung, et al.
Published: (2026)
by: Ro, Yusung, et al.
Published: (2026)
See and Fix the Flaws: Enabling VLMs and Diffusion Models to Comprehend Visual Artifacts via Agentic Data Synthesis
by: Park, Jaehyun, et al.
Published: (2026)
by: Park, Jaehyun, et al.
Published: (2026)
Rethinking Token Reduction for Large Vision-Language Models
by: Wang, Yi, et al.
Published: (2026)
by: Wang, Yi, et al.
Published: (2026)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
by: Xu, Xuwei, et al.
Published: (2023)
by: Xu, Xuwei, et al.
Published: (2023)
Patch Rebirth: Toward Fast and Transferable Model Inversion of Vision Transformers
by: Heo, Seongsoo, et al.
Published: (2025)
by: Heo, Seongsoo, et al.
Published: (2025)
CLIP Tricks You: Training-free Token Pruning for Efficient Pixel Grounding in Large VIsion-Language Models
by: Lee, Sangin, et al.
Published: (2026)
by: Lee, Sangin, et al.
Published: (2026)
Dynamic Token Reduction during Generation for Vision Language Models
by: Liang, Xiaoyu, et al.
Published: (2025)
by: Liang, Xiaoyu, et al.
Published: (2025)
Towards Visual Text Design Transfer Across Languages
by: Choi, Yejin, et al.
Published: (2024)
by: Choi, Yejin, et al.
Published: (2024)
Similar Items
-
PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion
by: Choi, Jaehyun, et al.
Published: (2025) -
Self-supervised Transformation Learning for Equivariant Representations
by: Yu, Jaemyung, et al.
Published: (2025) -
AH-OCDA: Amplitude-based Curriculum Learning and Hopfield Segmentation Model for Open Compound Domain Adaptation
by: Choi, Jaehyun, et al.
Published: (2024) -
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
by: Lee, Dong-Jae, et al.
Published: (2026) -
DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization
by: Lee, Dongyeun, et al.
Published: (2025)