A survey on efficient vision transformers: algorithms, techniques, and performance benchmarking
Fuente:
arXiv
Saved in:
| Main Authors: | Papa, Lorenzo, Russo, Paolo, Amerini, Irene, Zhou, Luping |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
METER: a mobile vision transformer architecture for monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
Shedding Light on Depth: Explainability Assessment in Monocular Depth Estimation
by: Cirillo, Lorenzo, et al.
Published: (2025)
by: Cirillo, Lorenzo, et al.
Published: (2025)
D4D: An RGBD diffusion model to boost monocular depth estimation
by: Papa, L., et al.
Published: (2024)
by: Papa, L., et al.
Published: (2024)
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
by: Zanchetta, Luca, et al.
Published: (2025)
by: Zanchetta, Luca, et al.
Published: (2025)
DepthFake: a depth-based strategy for detecting Deepfake videos
by: Maiano, Luca, et al.
Published: (2022)
by: Maiano, Luca, et al.
Published: (2022)
On the impact of key design aspects in simulated Hybrid Quantum Neural Networks for Earth Observation
by: Papa, Lorenzo, et al.
Published: (2024)
by: Papa, Lorenzo, et al.
Published: (2024)
Diffusion Models for Earth Observation Use-cases: from cloud removal to urban change detection
by: Sanguigni, Fulvio, et al.
Published: (2023)
by: Sanguigni, Fulvio, et al.
Published: (2023)
Continuous fake media detection: adapting deepfake detectors to new generative techniques
by: Tassone, Francesco, et al.
Published: (2024)
by: Tassone, Francesco, et al.
Published: (2024)
Enhancing Ground-to-Aerial Image Matching for Visual Misinformation Detection Using Semantic Segmentation
by: Mule, Emanuele, et al.
Published: (2025)
by: Mule, Emanuele, et al.
Published: (2025)
Z-SASLM: Zero-Shot Style-Aligned SLI Blending Latent Manipulation
by: Borgi, Alessio, et al.
Published: (2025)
by: Borgi, Alessio, et al.
Published: (2025)
Enhancing Abnormality Identification: Robust Out-of-Distribution Strategies for Deepfake Detection
by: Maiano, Luca, et al.
Published: (2025)
by: Maiano, Luca, et al.
Published: (2025)
LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation
by: Cardullo, Daniele, et al.
Published: (2025)
by: Cardullo, Daniele, et al.
Published: (2025)
R3ST: A Synthetic 3D Dataset With Realistic Trajectories
by: Teglia, Simone, et al.
Published: (2025)
by: Teglia, Simone, et al.
Published: (2025)
STLight: a Fully Convolutional Approach for Efficient Predictive Learning by Spatio-Temporal joint Processing
by: Alfarano, Andrea, et al.
Published: (2024)
by: Alfarano, Andrea, et al.
Published: (2024)
A Semantic Segmentation-guided Approach for Ground-to-Aerial Image Matching
by: Pro, Francesco, et al.
Published: (2024)
by: Pro, Francesco, et al.
Published: (2024)
Learning from Unlabelled Data with Transformers: Domain Adaptation for Semantic Segmentation of High Resolution Aerial Images
by: Dionelis, Nikolaos, et al.
Published: (2024)
by: Dionelis, Nikolaos, et al.
Published: (2024)
Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video Compression
by: Chen, Zhenghao, et al.
Published: (2024)
by: Chen, Zhenghao, et al.
Published: (2024)
Comics Datasets Framework: Mix of Comics datasets for detection benchmarking
by: Vivoli, Emanuele, et al.
Published: (2024)
by: Vivoli, Emanuele, et al.
Published: (2024)
Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark
by: Santos-Villafranca, Maria, et al.
Published: (2026)
by: Santos-Villafranca, Maria, et al.
Published: (2026)
UniBrain: A Unified Model for Cross-Subject Brain Decoding
by: Wang, Zicheng, et al.
Published: (2024)
by: Wang, Zicheng, et al.
Published: (2024)
A survey of datasets for computer vision in agriculture
by: Heider, Nico, et al.
Published: (2025)
by: Heider, Nico, et al.
Published: (2025)
Spintronics for image recognition: performance benchmarking via ultrafast data-driven simulations
by: Moureaux, Anatole, et al.
Published: (2023)
by: Moureaux, Anatole, et al.
Published: (2023)
A survey of facial recognition techniques
by: Bahjat, Aya Kaysan
Published: (2026)
by: Bahjat, Aya Kaysan
Published: (2026)
Robust CLIP-Based Detector for Exposing Diffusion Model-Generated Images
by: Santosh, et al.
Published: (2024)
by: Santosh, et al.
Published: (2024)
Cross multiscale vision transformer for deep fake detection
by: P, Akhshan, et al.
Published: (2025)
by: P, Akhshan, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Improving Weakly Supervised Temporal Action Localization by Exploiting Multi-resolution Information in Temporal Domain
by: Su, Rui, et al.
Published: (2025)
by: Su, Rui, et al.
Published: (2025)
TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
by: Xu, Wenting, et al.
Published: (2024)
by: Xu, Wenting, et al.
Published: (2024)
Progressive Cross-Stream Cooperation in Spatial and Temporal Domain for Action Localization
by: Su, Rui, et al.
Published: (2019)
by: Su, Rui, et al.
Published: (2019)
A comprehensive survey of oracle character recognition: challenges, benchmarks, and beyond
by: Li, Jing, et al.
Published: (2024)
by: Li, Jing, et al.
Published: (2024)
Event-based vision on FPGAs -- a survey
by: Kryjak, Tomasz
Published: (2024)
by: Kryjak, Tomasz
Published: (2024)
Multispectral airborne laser scanning for tree species classification: a benchmark of machine learning and deep learning algorithms
by: Taher, Josef, et al.
Published: (2025)
by: Taher, Josef, et al.
Published: (2025)
an interpretable vision transformer framework for automated brain tumor classification
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
by: Mbonu, Chinedu Emmanuel, et al.
Published: (2026)
Beyond the final layer: Attentive multilayer fusion for vision transformers
by: Ciernik, Laure, et al.
Published: (2026)
by: Ciernik, Laure, et al.
Published: (2026)
MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation
by: Yang, Ling, et al.
Published: (2023)
by: Yang, Ling, et al.
Published: (2023)
Hyperspectral data augmentation with transformer-based diffusion models
by: Ferrari, Mattia, et al.
Published: (2025)
by: Ferrari, Mattia, et al.
Published: (2025)
A Review of Longitudinal Radiology Report Generation: Dataset Composition, Methods, and Performance Evaluation
by: Zhou, Shaoyang, et al.
Published: (2025)
by: Zhou, Shaoyang, et al.
Published: (2025)
A comprehensive review of datasets and deep learning techniques for vision in Unmanned Surface Vehicles
by: Trinh, Linh, et al.
Published: (2024)
by: Trinh, Linh, et al.
Published: (2024)
More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era
by: Li, Yingtai, et al.
Published: (2025)
by: Li, Yingtai, et al.
Published: (2025)
Self-supervised pretraining for an iterative image size agnostic vision transformer
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
by: Prisadnikov, Nedyalko, et al.
Published: (2026)
Similar Items
-
METER: a mobile vision transformer architecture for monocular depth estimation
by: Papa, L., et al.
Published: (2024) -
Shedding Light on Depth: Explainability Assessment in Monocular Depth Estimation
by: Cirillo, Lorenzo, et al.
Published: (2025) -
D4D: An RGBD diffusion model to boost monocular depth estimation
by: Papa, L., et al.
Published: (2024) -
VidCLearn: A Continual Learning Approach for Text-to-Video Generation
by: Zanchetta, Luca, et al.
Published: (2025) -
DepthFake: a depth-based strategy for detecting Deepfake videos
by: Maiano, Luca, et al.
Published: (2022)