Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
Fuente:
arXiv
Guardado en:
| Autores principales: | Nottebaum, Moritz, Dunnhofer, Matteo, Micheloni, Christian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CPUBone: Efficient Vision Backbone Design for Devices with Low Parallelization Capabilities
por: Nottebaum, Moritz, et al.
Publicado: (2026)
por: Nottebaum, Moritz, et al.
Publicado: (2026)
LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
por: Nottebaum, Moritz, et al.
Publicado: (2024)
por: Nottebaum, Moritz, et al.
Publicado: (2024)
Tracking Skiers from the Top to the Bottom
por: Dunnhofer, Matteo, et al.
Publicado: (2023)
por: Dunnhofer, Matteo, et al.
Publicado: (2023)
Is Tracking really more challenging in First Person Egocentric Vision?
por: Dunnhofer, Matteo, et al.
Publicado: (2025)
por: Dunnhofer, Matteo, et al.
Publicado: (2025)
Better, But Not Sufficient: Testing Video ANNs Against Macaque IT Dynamics
por: Dunnhofer, Matteo, et al.
Publicado: (2026)
por: Dunnhofer, Matteo, et al.
Publicado: (2026)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
por: Manigrasso, Zaira, et al.
Publicado: (2024)
por: Manigrasso, Zaira, et al.
Publicado: (2024)
SkelMamba: A State Space Model for Efficient Skeleton Action Recognition of Neurological Disorders
por: Martinel, Niki, et al.
Publicado: (2024)
por: Martinel, Niki, et al.
Publicado: (2024)
Revisiting the Integration of Convolution and Attention for Vision Backbone
por: Zhu, Lei, et al.
Publicado: (2024)
por: Zhu, Lei, et al.
Publicado: (2024)
ViR: Towards Efficient Vision Retention Backbones
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
ReMAR-DS: Recalibrated Feature Learning for Metal Artifact Reduction and CT Domain Transformation
por: Rehman, Mubashara, et al.
Publicado: (2025)
por: Rehman, Mubashara, et al.
Publicado: (2025)
Vision Backbone Efficient Selection for Image Classification in Low-Data Regimes
por: Guerin, Joris, et al.
Publicado: (2024)
por: Guerin, Joris, et al.
Publicado: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
por: Alkin, Benedikt, et al.
Publicado: (2024)
por: Alkin, Benedikt, et al.
Publicado: (2024)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
por: Liu, Yunze, et al.
Publicado: (2024)
por: Liu, Yunze, et al.
Publicado: (2024)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
por: Safdar, Aon, et al.
Publicado: (2025)
por: Safdar, Aon, et al.
Publicado: (2025)
Quasar-ViT: Hardware-Oriented Quantization-Aware Architecture Search for Vision Transformers
por: Li, Zhengang, et al.
Publicado: (2024)
por: Li, Zhengang, et al.
Publicado: (2024)
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
por: Apedo, Yvon, et al.
Publicado: (2026)
por: Apedo, Yvon, et al.
Publicado: (2026)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
por: Böhle, Moritz, et al.
Publicado: (2025)
por: Böhle, Moritz, et al.
Publicado: (2025)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
por: Horawalavithana, Sameera, et al.
Publicado: (2026)
por: Horawalavithana, Sameera, et al.
Publicado: (2026)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
por: Barsellotti, Luca, et al.
Publicado: (2024)
por: Barsellotti, Luca, et al.
Publicado: (2024)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
por: Wang, Yulin, et al.
Publicado: (2024)
por: Wang, Yulin, et al.
Publicado: (2024)
CHOSEN: Compilation to Hardware Optimization Stack for Efficient Vision Transformer Inference
por: Sadeghi, Mohammad Erfan, et al.
Publicado: (2024)
por: Sadeghi, Mohammad Erfan, et al.
Publicado: (2024)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
por: Khazem, Salim
Publicado: (2026)
por: Khazem, Salim
Publicado: (2026)
Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising
por: Saied, Youssef, et al.
Publicado: (2026)
por: Saied, Youssef, et al.
Publicado: (2026)
A Survey on Backbones for Deep Video Action Recognition
por: Tang, Zixuan, et al.
Publicado: (2024)
por: Tang, Zixuan, et al.
Publicado: (2024)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
por: Liu, Haixu, et al.
Publicado: (2025)
por: Liu, Haixu, et al.
Publicado: (2025)
ForestProtector: An IoT Architecture Integrating Machine Vision and Deep Reinforcement Learning for Efficient Wildfire Monitoring
por: Bonilla-Ormachea, Kenneth, et al.
Publicado: (2025)
por: Bonilla-Ormachea, Kenneth, et al.
Publicado: (2025)
Self-Supervised Backbone Framework for Diverse Agricultural Vision Tasks
por: Sornapudi, Sudhir, et al.
Publicado: (2024)
por: Sornapudi, Sudhir, et al.
Publicado: (2024)
RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone
por: Munir, Mustafa, et al.
Publicado: (2024)
por: Munir, Mustafa, et al.
Publicado: (2024)
A Generic Shared Attention Mechanism for Various Backbone Neural Networks
por: Huang, Zhongzhan, et al.
Publicado: (2022)
por: Huang, Zhongzhan, et al.
Publicado: (2022)
A Survey on Mamba Architecture for Vision Applications
por: Ibrahim, Fady, et al.
Publicado: (2025)
por: Ibrahim, Fady, et al.
Publicado: (2025)
EdgeDiT: Hardware-Aware Diffusion Transformers for Efficient On-Device Image Generation
por: Kodavanti, Sravanth, et al.
Publicado: (2026)
por: Kodavanti, Sravanth, et al.
Publicado: (2026)
Generalized Large-Scale Data Condensation via Various Backbone and Statistical Matching
por: Shao, Shitong, et al.
Publicado: (2023)
por: Shao, Shitong, et al.
Publicado: (2023)
Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design
por: Duvvuri, Raghuvir, et al.
Publicado: (2025)
por: Duvvuri, Raghuvir, et al.
Publicado: (2025)
Beyond ZOH: Advanced Discretization Strategies for Vision Mamba
por: Ibrahim, Fady, et al.
Publicado: (2026)
por: Ibrahim, Fady, et al.
Publicado: (2026)
Pretraining Objective Matters in Extreme Low-Data FGVC: A Backbone-Controlled Study
por: Hackett, Alexander, et al.
Publicado: (2026)
por: Hackett, Alexander, et al.
Publicado: (2026)
X-Distill: Cross-Architecture Vision Distillation for Visuomotor Learning
por: Shao, Maanping, et al.
Publicado: (2026)
por: Shao, Maanping, et al.
Publicado: (2026)
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
por: Li, Wenbo, et al.
Publicado: (2024)
por: Li, Wenbo, et al.
Publicado: (2024)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
por: Li, Siyuan, et al.
Publicado: (2023)
por: Li, Siyuan, et al.
Publicado: (2023)
Beyond the Visible: Multispectral Vision-Language Learning for Earth Observation
por: Marimo, Clive Tinashe, et al.
Publicado: (2025)
por: Marimo, Clive Tinashe, et al.
Publicado: (2025)
Data Metabolism: An Efficient Data Design Schema For Vision Language Model
por: Zhang, Jingyuan, et al.
Publicado: (2025)
por: Zhang, Jingyuan, et al.
Publicado: (2025)
Ejemplares similares
-
CPUBone: Efficient Vision Backbone Design for Devices with Low Parallelization Capabilities
por: Nottebaum, Moritz, et al.
Publicado: (2026) -
LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
por: Nottebaum, Moritz, et al.
Publicado: (2024) -
Tracking Skiers from the Top to the Bottom
por: Dunnhofer, Matteo, et al.
Publicado: (2023) -
Is Tracking really more challenging in First Person Egocentric Vision?
por: Dunnhofer, Matteo, et al.
Publicado: (2025) -
Better, But Not Sufficient: Testing Video ANNs Against Macaque IT Dynamics
por: Dunnhofer, Matteo, et al.
Publicado: (2026)