UniForm: A Reuse Attention Mechanism Optimized for Efficient Vision Transformers on Edge Devices
Fuente:
arXiv
Saved in:
| Main Authors: | Yeom, Seul-Ki, Kim, Tae-Ho |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
by: Simon, Marcel, et al.
Published: (2025)
by: Simon, Marcel, et al.
Published: (2025)
VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization
by: Yeom, Seul-Ki, et al.
Published: (2026)
by: Yeom, Seul-Ki, et al.
Published: (2026)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
by: Setyawan, Novendra, et al.
Published: (2025)
by: Setyawan, Novendra, et al.
Published: (2025)
Factorized Multi-Resolution HashGrid for Efficient Neural Radiance Fields: Execution on Edge-Devices
by: Jun-Seong, Kim, et al.
Published: (2026)
by: Jun-Seong, Kim, et al.
Published: (2026)
Cooperative Inference for Real-Time 3D Human Pose Estimation in Multi-Device Edge Networks
by: Choi, Hyun-Ho, et al.
Published: (2025)
by: Choi, Hyun-Ho, et al.
Published: (2025)
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
by: Liu, Xiang, et al.
Published: (2024)
by: Liu, Xiang, et al.
Published: (2024)
Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers
by: Lee, Sanghyeok, et al.
Published: (2024)
by: Lee, Sanghyeok, et al.
Published: (2024)
ToFe: Lagged Token Freezing and Reusing for Efficient Vision Transformer Inference
by: Zhang, Haoyue, et al.
Published: (2025)
by: Zhang, Haoyue, et al.
Published: (2025)
Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
EfficientQuant: An Efficient Post-Training Quantization for CNN-Transformer Hybrid Models on Edge Devices
by: Saha, Shaibal, et al.
Published: (2025)
by: Saha, Shaibal, et al.
Published: (2025)
EdgeFusion: On-Device Text-to-Image Generation
by: Castells, Thibault, et al.
Published: (2024)
by: Castells, Thibault, et al.
Published: (2024)
UniViTAR: Unified Vision Transformer with Native Resolution
by: Qiao, Limeng, et al.
Published: (2025)
by: Qiao, Limeng, et al.
Published: (2025)
ToSA: Token Selective Attention for Efficient Vision Transformers
by: Singh, Manish Kumar, et al.
Published: (2024)
by: Singh, Manish Kumar, et al.
Published: (2024)
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
by: Lu, Andrew, et al.
Published: (2025)
by: Lu, Andrew, et al.
Published: (2025)
S2AFormer: Strip Self-Attention for Efficient Vision Transformer
by: Xu, Guoan, et al.
Published: (2025)
by: Xu, Guoan, et al.
Published: (2025)
PADRe: A Unifying Polynomial Attention Drop-in Replacement for Efficient Vision Transformer
by: Letourneau, Pierre-David, et al.
Published: (2024)
by: Letourneau, Pierre-David, et al.
Published: (2024)
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
by: Lee, Jewon, et al.
Published: (2025)
by: Lee, Jewon, et al.
Published: (2025)
SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
by: Hu, Dongting, et al.
Published: (2026)
by: Hu, Dongting, et al.
Published: (2026)
DuoFormer: Leveraging Hierarchical Representations by Local and Global Attention Vision Transformer
by: Tang, Xiaoya, et al.
Published: (2025)
by: Tang, Xiaoya, et al.
Published: (2025)
Representative Attention For Vision Transformers
by: Li, Yuntong, et al.
Published: (2026)
by: Li, Yuntong, et al.
Published: (2026)
Vision Transformers with Hierarchical Attention
by: Liu, Yun, et al.
Published: (2021)
by: Liu, Yun, et al.
Published: (2021)
NuWa: Deriving Lightweight Task-Specific Vision Transformers for Edge Devices
by: Wei, Ziteng, et al.
Published: (2025)
by: Wei, Ziteng, et al.
Published: (2025)
Reusing Attention for One-stage Lane Topology Understanding
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
xEdgeFace: Efficient Cross-Spectral Face Recognition for Edge Devices
by: George, Anjith, et al.
Published: (2025)
by: George, Anjith, et al.
Published: (2025)
Comprehensive Information Bottleneck for Unveiling Universal Attribution to Interpret Vision Transformers
by: Hong, Jung-Ho, et al.
Published: (2025)
by: Hong, Jung-Ho, et al.
Published: (2025)
ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification
by: Kim, Ga-Eun, et al.
Published: (2023)
by: Kim, Ga-Eun, et al.
Published: (2023)
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
by: Mao, Weijia, et al.
Published: (2025)
by: Mao, Weijia, et al.
Published: (2025)
EdgeFM: Efficient Edge Inference for Vision-Language Models
by: Deng, Mengling, et al.
Published: (2026)
by: Deng, Mengling, et al.
Published: (2026)
Structured Initialization for Attention in Vision Transformers
by: Zheng, Jianqiao, et al.
Published: (2024)
by: Zheng, Jianqiao, et al.
Published: (2024)
Vision Transformers are Circulant Attention Learners
by: Han, Dongchen, et al.
Published: (2025)
by: Han, Dongchen, et al.
Published: (2025)
Multi-manifold Attention for Vision Transformers
by: Konstantinidis, Dimitrios, et al.
Published: (2022)
by: Konstantinidis, Dimitrios, et al.
Published: (2022)
HAViT: Historical Attention Vision Transformer
by: Banik, Swarnendu, et al.
Published: (2026)
by: Banik, Swarnendu, et al.
Published: (2026)
EViT-Unet: U-Net Like Efficient Vision Transformer for Medical Image Segmentation on Mobile and Edge Devices
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
EcoLens: Leveraging Multi-Objective Bayesian Optimization for Energy-Efficient Video Processing on Edge Devices
by: Civjan, Benjamin, et al.
Published: (2025)
by: Civjan, Benjamin, et al.
Published: (2025)
EdgeDiT: Hardware-Aware Diffusion Transformers for Efficient On-Device Image Generation
by: Kodavanti, Sravanth, et al.
Published: (2026)
by: Kodavanti, Sravanth, et al.
Published: (2026)
Reusing Computation in Text-to-Image Diffusion for Efficient Generation of Image Sets
by: Decatur, Dale, et al.
Published: (2025)
by: Decatur, Dale, et al.
Published: (2025)
HEART-VIT: Hessian-Guided Efficient Dynamic Attention and Token Pruning in Vision Transformer
by: Uddin, Mohammad Helal, et al.
Published: (2025)
by: Uddin, Mohammad Helal, et al.
Published: (2025)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
Similar Items
-
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
by: Simon, Marcel, et al.
Published: (2025) -
VINO: Video-driven Invariance for Non-contextual Objects via Structural Prior Guided De-contextualization
by: Yeom, Seul-Ki, et al.
Published: (2026) -
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022) -
UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation
by: Zhao, Lei, et al.
Published: (2025) -
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
by: Setyawan, Novendra, et al.
Published: (2025)