EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Han, Li, Junyan, Hu, Muyan, Gan, Chuang, Han, Song |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2022
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss
di: Zhang, Zhuoyang, et al.
Pubblicazione: (2024)
di: Zhang, Zhuoyang, et al.
Pubblicazione: (2024)
FlexAttention for Efficient High-Resolution Vision-Language Models
di: Li, Junyan, et al.
Pubblicazione: (2024)
di: Li, Junyan, et al.
Pubblicazione: (2024)
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
di: Zou, Dongyun, et al.
Pubblicazione: (2026)
di: Zou, Dongyun, et al.
Pubblicazione: (2026)
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
di: Xie, Enze, et al.
Pubblicazione: (2024)
di: Xie, Enze, et al.
Pubblicazione: (2024)
MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
di: Lin, Ji, et al.
Pubblicazione: (2021)
di: Lin, Ji, et al.
Pubblicazione: (2021)
Attention to Detail: Global-Local Attention for High-Resolution AI-Generated Image Detection
di: Han, Lawrence
Pubblicazione: (2026)
di: Han, Lawrence
Pubblicazione: (2026)
HIRI-ViT: Scaling Vision Transformer with High Resolution Inputs
di: Yao, Ting, et al.
Pubblicazione: (2024)
di: Yao, Ting, et al.
Pubblicazione: (2024)
Efficient Image Super-Resolution with Multi-Scale Spatial Adaptive Attention Networks
di: Rao, Sushi, et al.
Pubblicazione: (2026)
di: Rao, Sushi, et al.
Pubblicazione: (2026)
Efficient Domain-Adaptive Multi-Task Dense Prediction with Vision Foundation Models
di: Kang, Beomseok, et al.
Pubblicazione: (2025)
di: Kang, Beomseok, et al.
Pubblicazione: (2025)
Agent Attention: On the Integration of Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2023)
di: Han, Dongchen, et al.
Pubblicazione: (2023)
LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
di: Li, Xiaohui, et al.
Pubblicazione: (2025)
di: Li, Xiaohui, et al.
Pubblicazione: (2025)
Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
di: Chen, Junyu, et al.
Pubblicazione: (2024)
di: Chen, Junyu, et al.
Pubblicazione: (2024)
ViT-AdaLA: Adapting Vision Transformers with Linear Attention
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
di: Ding, Xinpeng, et al.
Pubblicazione: (2023)
di: Ding, Xinpeng, et al.
Pubblicazione: (2023)
ViTAR: Vision Transformer with Any Resolution
di: Fan, Qihang, et al.
Pubblicazione: (2024)
di: Fan, Qihang, et al.
Pubblicazione: (2024)
DistriFusion: Distributed Parallel Inference for High-Resolution Diffusion Models
di: Li, Muyang, et al.
Pubblicazione: (2024)
di: Li, Muyang, et al.
Pubblicazione: (2024)
ViT-CoMer: Vision Transformer with Convolutional Multi-scale Feature Interaction for Dense Predictions
di: Xia, Chunlong, et al.
Pubblicazione: (2024)
di: Xia, Chunlong, et al.
Pubblicazione: (2024)
AllTracker: Efficient Dense Point Tracking at High Resolution
di: Harley, Adam W., et al.
Pubblicazione: (2025)
di: Harley, Adam W., et al.
Pubblicazione: (2025)
Scaling Vision Pre-Training to 4K Resolution
di: Shi, Baifeng, et al.
Pubblicazione: (2025)
di: Shi, Baifeng, et al.
Pubblicazione: (2025)
Multi-Scale Representations by Varying Window Attention for Semantic Segmentation
di: Yan, Haotian, et al.
Pubblicazione: (2024)
di: Yan, Haotian, et al.
Pubblicazione: (2024)
Bridging the Divide: Reconsidering Softmax and Linear Attention
di: Han, Dongchen, et al.
Pubblicazione: (2024)
di: Han, Dongchen, et al.
Pubblicazione: (2024)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
di: Chen, Weiming, et al.
Pubblicazione: (2026)
di: Chen, Weiming, et al.
Pubblicazione: (2026)
UniViTAR: Unified Vision Transformer with Native Resolution
di: Qiao, Limeng, et al.
Pubblicazione: (2025)
di: Qiao, Limeng, et al.
Pubblicazione: (2025)
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
di: Chen, Zhan, et al.
Pubblicazione: (2025)
di: Chen, Zhan, et al.
Pubblicazione: (2025)
MTLSI-Net: A Linear Semantic Interaction Network for Parameter-Efficient Multi-Task Dense Prediction
di: Liu, Chen, et al.
Pubblicazione: (2026)
di: Liu, Chen, et al.
Pubblicazione: (2026)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
di: Fang, Tongcheng, et al.
Pubblicazione: (2026)
di: Fang, Tongcheng, et al.
Pubblicazione: (2026)
Demystify Mamba in Vision: A Linear Attention Perspective
di: Han, Dongchen, et al.
Pubblicazione: (2024)
di: Han, Dongchen, et al.
Pubblicazione: (2024)
Sparse Refinement for Efficient High-Resolution Semantic Segmentation
di: Liu, Zhijian, et al.
Pubblicazione: (2024)
di: Liu, Zhijian, et al.
Pubblicazione: (2024)
HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
di: Han, Chengxi, et al.
Pubblicazione: (2024)
di: Han, Chengxi, et al.
Pubblicazione: (2024)
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
di: Xie, Enze, et al.
Pubblicazione: (2025)
di: Xie, Enze, et al.
Pubblicazione: (2025)
DELTA: Dense Efficient Long-range 3D Tracking for any video
di: Ngo, Tuan Duc, et al.
Pubblicazione: (2024)
di: Ngo, Tuan Duc, et al.
Pubblicazione: (2024)
Alias-Free ViT: Fractional Shift Invariance via Linear Attention
di: Michaeli, Hagay, et al.
Pubblicazione: (2025)
di: Michaeli, Hagay, et al.
Pubblicazione: (2025)
MAT: Multi-Range Attention Transformer for Efficient Image Super-Resolution
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
di: Xie, Chengxing, et al.
Pubblicazione: (2024)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
di: Liu, Haixu, et al.
Pubblicazione: (2025)
di: Liu, Haixu, et al.
Pubblicazione: (2025)
Cross-Resolution Attention Network for High-Resolution PM2.5 Prediction
di: Kheder, Ammar, et al.
Pubblicazione: (2026)
di: Kheder, Ammar, et al.
Pubblicazione: (2026)
EdgeCrafter: Compact ViTs for Edge Dense Prediction via Task-Specialized Distillation
di: Liu, Longfei, et al.
Pubblicazione: (2026)
di: Liu, Longfei, et al.
Pubblicazione: (2026)
LetheViT: Selective Machine Unlearning for Vision Transformers via Attention-Guided Contrastive Learning
di: Tong, Yujia, et al.
Pubblicazione: (2025)
di: Tong, Yujia, et al.
Pubblicazione: (2025)
Multi-Layer Dense Attention Decoder for Polyp Segmentation
di: Patel, Krushi, et al.
Pubblicazione: (2024)
di: Patel, Krushi, et al.
Pubblicazione: (2024)
MuViT: Multi-Resolution Vision Transformers for Learning Across Scales in Microscopy
di: Mantes, Albert Dominguez, et al.
Pubblicazione: (2026)
di: Mantes, Albert Dominguez, et al.
Pubblicazione: (2026)
MagNet: Multi-Level Attention Graph Network for Predicting High-Resolution Spatial Transcriptomics
di: Zhu, Junchao, et al.
Pubblicazione: (2025)
di: Zhu, Junchao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss
di: Zhang, Zhuoyang, et al.
Pubblicazione: (2024) -
FlexAttention for Efficient High-Resolution Vision-Language Models
di: Li, Junyan, et al.
Pubblicazione: (2024) -
JetViT: Efficient High-Resolution Vision Transformer with Post-Training Attention Search
di: Zou, Dongyun, et al.
Pubblicazione: (2026) -
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers
di: Xie, Enze, et al.
Pubblicazione: (2024) -
MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning
di: Lin, Ji, et al.
Pubblicazione: (2021)