Low-latency vision transformers via large-scale multi-head attention
Fuente:
arXiv
Saved in:
| Main Authors: | Gross, Ronit D., Halevi, Tal, Koresh, Ella, Tzach, Yarden, Kanter, Ido |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
by: Koresh, Ella, et al.
Published: (2025)
by: Koresh, Ella, et al.
Published: (2025)
Advanced deep architecture pruning using single filter performance
by: Tzach, Yarden, et al.
Published: (2025)
by: Tzach, Yarden, et al.
Published: (2025)
Tiny language models
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
Self-attention vector output similarities reveal how machines pay attention
by: Halevi, Tal, et al.
Published: (2025)
by: Halevi, Tal, et al.
Published: (2025)
Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
by: Tzach, Yarden, et al.
Published: (2025)
by: Tzach, Yarden, et al.
Published: (2025)
Towards a universal mechanism for successful deep learning
by: Meir, Yuval, et al.
Published: (2023)
by: Meir, Yuval, et al.
Published: (2023)
Single-Nodal Spontaneous Symmetry Breaking in NLP Models
by: Rosner, Shalom, et al.
Published: (2026)
by: Rosner, Shalom, et al.
Published: (2026)
Joint multi-dimensional dynamic attention and transformer for general image restoration
by: Zhang, Huan, et al.
Published: (2024)
by: Zhang, Huan, et al.
Published: (2024)
Cross-Layer Attentive Feature Upsampling for Low-latency Semantic Segmentation
by: Cheng, Tianheng, et al.
Published: (2026)
by: Cheng, Tianheng, et al.
Published: (2026)
LiteTracker: Leveraging Temporal Causality for Accurate Low-latency Tissue Tracking
by: Karaoglu, Mert Asim, et al.
Published: (2025)
by: Karaoglu, Mert Asim, et al.
Published: (2025)
Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention
by: Hao, Haiqing, et al.
Published: (2026)
by: Hao, Haiqing, et al.
Published: (2026)
Hierarchical Event Memory for Accurate and Low-latency Online Video Temporal Grounding
by: Zheng, Minghang, et al.
Published: (2025)
by: Zheng, Minghang, et al.
Published: (2025)
A recurrent vision transformer shows signatures of primate visual attention
by: Morgan, Jonathan, et al.
Published: (2025)
by: Morgan, Jonathan, et al.
Published: (2025)
A multi-scale vision transformer-based multimodal GeoAI model for mapping Arctic permafrost thaw
by: Li, Wenwen, et al.
Published: (2025)
by: Li, Wenwen, et al.
Published: (2025)
Metonymy in vision models undermines attention-based interpretability
by: Aniraj, Ananthu, et al.
Published: (2026)
by: Aniraj, Ananthu, et al.
Published: (2026)
HyperSLICE: HyperBand optimized Spiral for Low-latency Interactive Cardiac Examination
by: Jaubert, Olivier, et al.
Published: (2023)
by: Jaubert, Olivier, et al.
Published: (2023)
Translation Entropy: A Statistical Framework for Evaluating Translation Systems
by: Gross, Ronit D., et al.
Published: (2025)
by: Gross, Ronit D., et al.
Published: (2025)
Memory-efficient Low-latency Remote Photoplethysmography through Temporal-Spatial State Space Duality
by: Wang, Kegang, et al.
Published: (2025)
by: Wang, Kegang, et al.
Published: (2025)
Improving underwater semantic segmentation with underwater image quality attention and muti-scale aggregation attention
by: Zuo, Xin, et al.
Published: (2025)
by: Zuo, Xin, et al.
Published: (2025)
Multi-modal and Multi-view Fundus Image Fusion for Retinopathy Diagnosis via Multi-scale Cross-attention and Shifted Window Self-attention
by: Huang, Yonghao, et al.
Published: (2025)
by: Huang, Yonghao, et al.
Published: (2025)
Decentralized LoRA augmented transformer with multi-scale feature learning for secured eye diagnosis
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
by: Borno, Md. Naimur Asif, et al.
Published: (2025)
GFFE: G-buffer Free Frame Extrapolation for Low-latency Real-time Rendering
by: Wu, Songyin, et al.
Published: (2024)
by: Wu, Songyin, et al.
Published: (2024)
Enhancing compact convolutional transformers with super attention
by: Leandre, Simpenzwe Honore, et al.
Published: (2025)
by: Leandre, Simpenzwe Honore, et al.
Published: (2025)
Towards Low-latency Event-based Visual Recognition with Hybrid Step-wise Distillation Spiking Neural Networks
by: Zhong, Xian, et al.
Published: (2024)
by: Zhong, Xian, et al.
Published: (2024)
Interpreting vision transformers via residual replacement model
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
For a semiotic AI: Bridging computer vision and visual semiotics for computational observation of large scale facial image archives
by: Morra, Lia, et al.
Published: (2024)
by: Morra, Lia, et al.
Published: (2024)
Cross multiscale vision transformer for deep fake detection
by: P, Akhshan, et al.
Published: (2025)
by: P, Akhshan, et al.
Published: (2025)
Adapting MIMO video restoration networks to low latency constraints
by: Dewil, Valéry, et al.
Published: (2024)
by: Dewil, Valéry, et al.
Published: (2024)
LILAC: Long-sequence Incremental Low-latency Arbitrary Motion Stylization via Streaming VAE-Diffusion with Causal Decoding
by: Ren, Peng, et al.
Published: (2025)
by: Ren, Peng, et al.
Published: (2025)
Zero-to-Hero: Enhancing Zero-Shot Novel View Synthesis via Attention Map Filtering
by: Sobol, Ido, et al.
Published: (2024)
by: Sobol, Ido, et al.
Published: (2024)
Implicit Style-Content Separation using B-LoRA
by: Frenkel, Yarden, et al.
Published: (2024)
by: Frenkel, Yarden, et al.
Published: (2024)
Learning rigid-body simulators over implicit shapes for large-scale scenes and vision
by: Rubanova, Yulia, et al.
Published: (2024)
by: Rubanova, Yulia, et al.
Published: (2024)
an interpretable vision transformer framework for automated brain tumor classification
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
Beyond the final layer: Attentive multilayer fusion for vision transformers
by: Ciernik, Laure, et al.
Published: (2026)
by: Ciernik, Laure, et al.
Published: (2026)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction
by: Smock, Brandon, et al.
Published: (2025)
by: Smock, Brandon, et al.
Published: (2025)
A multi-temporal multi-spectral attention-augmented deep convolution neural network with contrastive learning for crop yield prediction
by: Dangi, Shalini, et al.
Published: (2025)
by: Dangi, Shalini, et al.
Published: (2025)
A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
by: Papa, Lorenzo, et al.
Published: (2023)
by: Papa, Lorenzo, et al.
Published: (2023)
METER: a mobile vision transformer architecture for monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
Similar Items
-
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
by: Koresh, Ella, et al.
Published: (2025) -
Advanced deep architecture pruning using single filter performance
by: Tzach, Yarden, et al.
Published: (2025) -
Tiny language models
by: Gross, Ronit D., et al.
Published: (2025) -
Self-attention vector output similarities reveal how machines pay attention
by: Halevi, Tal, et al.
Published: (2025) -
Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning
by: Tzach, Yarden, et al.
Published: (2025)