Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Sanghyeok, Choi, Joonmyung, Kim, Hyunwoo J. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
di: Lee, Sanghyeok, et al.
Pubblicazione: (2024)
di: Lee, Sanghyeok, et al.
Pubblicazione: (2024)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
di: Choi, Joonmyung, et al.
Pubblicazione: (2024)
di: Choi, Joonmyung, et al.
Pubblicazione: (2024)
Representation Shift: Unifying Token Compression with FlashAttention
di: Choi, Joonmyung, et al.
Pubblicazione: (2025)
di: Choi, Joonmyung, et al.
Pubblicazione: (2025)
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
di: Choi, Joonmyung, et al.
Pubblicazione: (2026)
di: Choi, Joonmyung, et al.
Pubblicazione: (2026)
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
di: Kim, Jongha, et al.
Pubblicazione: (2025)
di: Kim, Jongha, et al.
Pubblicazione: (2025)
Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
di: Park, Jihwan, et al.
Pubblicazione: (2025)
di: Park, Jihwan, et al.
Pubblicazione: (2025)
MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language Models
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
di: Ko, Dohwan, et al.
Pubblicazione: (2026)
ToSA: Token Selective Attention for Efficient Vision Transformers
di: Singh, Manish Kumar, et al.
Pubblicazione: (2024)
di: Singh, Manish Kumar, et al.
Pubblicazione: (2024)
Frequency-Aware Token Reduction for Efficient Vision Transformer
di: Lee, Dong-Jae, et al.
Pubblicazione: (2025)
di: Lee, Dong-Jae, et al.
Pubblicazione: (2025)
Robust Multimodal 3D Object Detection via Modality-Agnostic Decoding and Proximity-based Modality Ensemble
di: Cha, Juhan, et al.
Pubblicazione: (2024)
di: Cha, Juhan, et al.
Pubblicazione: (2024)
Efficient multi-view training for 3D Gaussian Splatting
di: Choi, Minhyuk, et al.
Pubblicazione: (2025)
di: Choi, Minhyuk, et al.
Pubblicazione: (2025)
Lossless Token Merging Even Without Fine-Tuning in Vision Transformers
di: Lee, Jaeyeon, et al.
Pubblicazione: (2025)
di: Lee, Jaeyeon, et al.
Pubblicazione: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
di: Lew, Jaihyun, et al.
Pubblicazione: (2024)
di: Lew, Jaihyun, et al.
Pubblicazione: (2024)
Sprint: Sparse-Dense Residual Fusion for Efficient Diffusion Transformers
di: Park, Dogyun, et al.
Pubblicazione: (2025)
di: Park, Dogyun, et al.
Pubblicazione: (2025)
Probabilistic Vision-Language Representation for Weakly Supervised Temporal Action Localization
di: Lim, Geuntaek, et al.
Pubblicazione: (2024)
di: Lim, Geuntaek, et al.
Pubblicazione: (2024)
MVFormer: Diversifying Feature Normalization and Token Mixing for Efficient Vision Transformers
di: Bae, Jongseong, et al.
Pubblicazione: (2024)
di: Bae, Jongseong, et al.
Pubblicazione: (2024)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
di: Uddin, Mohammad Helal, et al.
Pubblicazione: (2025)
di: Uddin, Mohammad Helal, et al.
Pubblicazione: (2025)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
di: Xia, Zunhui, et al.
Pubblicazione: (2025)
di: Xia, Zunhui, et al.
Pubblicazione: (2025)
Retinal Layer Segmentation in OCT Images With 2.5D Cross-slice Feature Fusion Module for Glaucoma Assessment
di: Kim, Hyunwoo, et al.
Pubblicazione: (2026)
di: Kim, Hyunwoo, et al.
Pubblicazione: (2026)
ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
di: Chytas, Sotirios Panagiotis, et al.
Pubblicazione: (2025)
di: Chytas, Sotirios Panagiotis, et al.
Pubblicazione: (2025)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
di: Xiao, Chaodong, et al.
Pubblicazione: (2026)
di: Xiao, Chaodong, et al.
Pubblicazione: (2026)
Speed-up of Vision Transformer Models by Attention-aware Token Filtering
di: Naruko, Takahiro, et al.
Pubblicazione: (2025)
di: Naruko, Takahiro, et al.
Pubblicazione: (2025)
PPT: Token Pruning and Pooling for Efficient Vision Transformers
di: Wu, Xinjian, et al.
Pubblicazione: (2023)
di: Wu, Xinjian, et al.
Pubblicazione: (2023)
Multi-manifold Attention for Vision Transformers
di: Konstantinidis, Dimitrios, et al.
Pubblicazione: (2022)
di: Konstantinidis, Dimitrios, et al.
Pubblicazione: (2022)
Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
di: Ko, Dohwan, et al.
Pubblicazione: (2025)
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation
di: Cho, Yubin, et al.
Pubblicazione: (2024)
di: Cho, Yubin, et al.
Pubblicazione: (2024)
Self-Supervised Multi-Scale Transformer with Attention-Guided Fusion for Efficient Crack Detection
di: Kyem, Blessing Agyei, et al.
Pubblicazione: (2025)
di: Kyem, Blessing Agyei, et al.
Pubblicazione: (2025)
GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer
di: Jia, Ding, et al.
Pubblicazione: (2024)
di: Jia, Ding, et al.
Pubblicazione: (2024)
Groupwise Query Specialization and Quality-Aware Multi-Assignment for Transformer-based Visual Relationship Detection
di: Kim, Jongha, et al.
Pubblicazione: (2024)
di: Kim, Jongha, et al.
Pubblicazione: (2024)
Learning to Merge Tokens via Decoupled Embedding for Efficient Vision Transformers
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
di: Lee, Dong Hoon, et al.
Pubblicazione: (2024)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
di: Galagain, Calvin, et al.
Pubblicazione: (2026)
di: Galagain, Calvin, et al.
Pubblicazione: (2026)
MSVIT: Improving Spiking Vision Transformer Using Multi-scale Attention Fusion
di: Hua, Wei, et al.
Pubblicazione: (2025)
di: Hua, Wei, et al.
Pubblicazione: (2025)
Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality Generation
di: Park, Dogyun, et al.
Pubblicazione: (2025)
di: Park, Dogyun, et al.
Pubblicazione: (2025)
HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models
di: Liu, Jizhihui, et al.
Pubblicazione: (2025)
di: Liu, Jizhihui, et al.
Pubblicazione: (2025)
Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention
di: Qiao, Bingtian, et al.
Pubblicazione: (2026)
di: Qiao, Bingtian, et al.
Pubblicazione: (2026)
UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
di: Yeom, Seul-Ki, et al.
Pubblicazione: (2024)
di: Yeom, Seul-Ki, et al.
Pubblicazione: (2024)
Spectral Vision Transformer for Efficient Tokenization with Limited Data
di: Roberts, Alexandra G., et al.
Pubblicazione: (2026)
di: Roberts, Alexandra G., et al.
Pubblicazione: (2026)
Vote&Mix: Plug-and-Play Token Reduction for Efficient Vision Transformer
di: Peng, Shuai, et al.
Pubblicazione: (2024)
di: Peng, Shuai, et al.
Pubblicazione: (2024)
Block-based Symmetric Pruning and Fusion for Efficient Vision Transformers
di: Hsieh, Yi-Kuan, et al.
Pubblicazione: (2025)
di: Hsieh, Yi-Kuan, et al.
Pubblicazione: (2025)
ExFusion: Efficient Transformer Training via Multi-Experts Fusion
di: Ruan, Jiacheng, et al.
Pubblicazione: (2026)
di: Ruan, Jiacheng, et al.
Pubblicazione: (2026)
Documenti analoghi
-
EfficientViM: Efficient Vision Mamba with Hidden State Mixer based State Space Duality
di: Lee, Sanghyeok, et al.
Pubblicazione: (2024) -
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
di: Choi, Joonmyung, et al.
Pubblicazione: (2024) -
Representation Shift: Unifying Token Compression with FlashAttention
di: Choi, Joonmyung, et al.
Pubblicazione: (2025) -
DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning
di: Choi, Joonmyung, et al.
Pubblicazione: (2026) -
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
di: Kim, Jongha, et al.
Pubblicazione: (2025)