Guardado en:
| Autores principales: | Nottebaum, Moritz, Dunnhofer, Matteo, Micheloni, Christian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.26425 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
por: Nottebaum, Moritz, et al.
Publicado: (2026)
por: Nottebaum, Moritz, et al.
Publicado: (2026)
LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
por: Nottebaum, Moritz, et al.
Publicado: (2024)
por: Nottebaum, Moritz, et al.
Publicado: (2024)
Tracking Skiers from the Top to the Bottom
por: Dunnhofer, Matteo, et al.
Publicado: (2023)
por: Dunnhofer, Matteo, et al.
Publicado: (2023)
Is Tracking really more challenging in First Person Egocentric Vision?
por: Dunnhofer, Matteo, et al.
Publicado: (2025)
por: Dunnhofer, Matteo, et al.
Publicado: (2025)
Better, But Not Sufficient: Testing Video ANNs Against Macaque IT Dynamics
por: Dunnhofer, Matteo, et al.
Publicado: (2026)
por: Dunnhofer, Matteo, et al.
Publicado: (2026)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
por: Manigrasso, Zaira, et al.
Publicado: (2024)
por: Manigrasso, Zaira, et al.
Publicado: (2024)
SkelMamba: A State Space Model for Efficient Skeleton Action Recognition of Neurological Disorders
por: Martinel, Niki, et al.
Publicado: (2024)
por: Martinel, Niki, et al.
Publicado: (2024)
Vision Backbone Efficient Selection for Image Classification in Low-Data Regimes
por: Guerin, Joris, et al.
Publicado: (2024)
por: Guerin, Joris, et al.
Publicado: (2024)
ReMAR-DS: Recalibrated Feature Learning for Metal Artifact Reduction and CT Domain Transformation
por: Rehman, Mubashara, et al.
Publicado: (2025)
por: Rehman, Mubashara, et al.
Publicado: (2025)
ViR: Towards Efficient Vision Retention Backbones
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
por: Hatamizadeh, Ali, et al.
Publicado: (2023)
Revisiting the Integration of Convolution and Attention for Vision Backbone
por: Zhu, Lei, et al.
Publicado: (2024)
por: Zhu, Lei, et al.
Publicado: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
por: Alkin, Benedikt, et al.
Publicado: (2024)
por: Alkin, Benedikt, et al.
Publicado: (2024)
CoMViT: An Efficient Vision Backbone for Supervised Classification in Medical Imaging
por: Safdar, Aon, et al.
Publicado: (2025)
por: Safdar, Aon, et al.
Publicado: (2025)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
por: Liu, Yunze, et al.
Publicado: (2024)
por: Liu, Yunze, et al.
Publicado: (2024)
Efficient Partitioning Vision Transformer on Edge Devices for Distributed Inference
por: Liu, Xiang, et al.
Publicado: (2024)
por: Liu, Xiang, et al.
Publicado: (2024)
Pretraining Objective Matters in Extreme Low-Data FGVC: A Backbone-Controlled Study
por: Hackett, Alexander, et al.
Publicado: (2026)
por: Hackett, Alexander, et al.
Publicado: (2026)
Back to the Barn with LLAMAs: Evolving Pretrained LLM Backbones in Finetuning Vision Language Models
por: Horawalavithana, Sameera, et al.
Publicado: (2026)
por: Horawalavithana, Sameera, et al.
Publicado: (2026)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
por: Barsellotti, Luca, et al.
Publicado: (2024)
por: Barsellotti, Luca, et al.
Publicado: (2024)
EfficientTrain++: Generalized Curriculum Learning for Efficient Visual Backbone Training
por: Wang, Yulin, et al.
Publicado: (2024)
por: Wang, Yulin, et al.
Publicado: (2024)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
por: Khazem, Salim
Publicado: (2026)
por: Khazem, Salim
Publicado: (2026)
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
por: Böhle, Moritz, et al.
Publicado: (2025)
por: Böhle, Moritz, et al.
Publicado: (2025)
Self-Supervised Backbone Framework for Diverse Agricultural Vision Tasks
por: Sornapudi, Sudhir, et al.
Publicado: (2024)
por: Sornapudi, Sudhir, et al.
Publicado: (2024)
Normalization Equivariance for Arbitrary Backbones, with Application to Image Denoising
por: Saied, Youssef, et al.
Publicado: (2026)
por: Saied, Youssef, et al.
Publicado: (2026)
A Survey on Backbones for Deep Video Action Recognition
por: Tang, Zixuan, et al.
Publicado: (2024)
por: Tang, Zixuan, et al.
Publicado: (2024)
Tighnari: Multi-modal Plant Species Prediction Based on Hierarchical Cross-Attention Using Graph-Based and Vision Backbone-Extracted Features
por: Liu, Haixu, et al.
Publicado: (2025)
por: Liu, Haixu, et al.
Publicado: (2025)
Evaluation of Safety Cognition Capability in Vision-Language Models for Autonomous Driving
por: Zhang, Enming, et al.
Publicado: (2025)
por: Zhang, Enming, et al.
Publicado: (2025)
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
por: Pan, Chenbin, et al.
Publicado: (2025)
por: Pan, Chenbin, et al.
Publicado: (2025)
An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
por: Jiang, Siyang, et al.
Publicado: (2025)
por: Jiang, Siyang, et al.
Publicado: (2025)
RapidNet: Multi-Level Dilated Convolution Based Mobile Backbone
por: Munir, Mustafa, et al.
Publicado: (2024)
por: Munir, Mustafa, et al.
Publicado: (2024)
A Generic Shared Attention Mechanism for Various Backbone Neural Networks
por: Huang, Zhongzhan, et al.
Publicado: (2022)
por: Huang, Zhongzhan, et al.
Publicado: (2022)
Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
por: Li, Junjie, et al.
Publicado: (2025)
por: Li, Junjie, et al.
Publicado: (2025)
Generalized Large-Scale Data Condensation via Various Backbone and Statistical Matching
por: Shao, Shitong, et al.
Publicado: (2023)
por: Shao, Shitong, et al.
Publicado: (2023)
Vision-DeepResearch: Incentivizing DeepResearch Capability in Multimodal Large Language Models
por: Huang, Wenxuan, et al.
Publicado: (2026)
por: Huang, Wenxuan, et al.
Publicado: (2026)
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
por: Zhang, Zhuoyang, et al.
Publicado: (2025)
por: Zhang, Zhuoyang, et al.
Publicado: (2025)
Latent-Info and Low-Dimensional Learning for Human Mesh Recovery and Parallel Optimization
por: Zhang, Xiang, et al.
Publicado: (2025)
por: Zhang, Xiang, et al.
Publicado: (2025)
Data Metabolism: An Efficient Data Design Schema For Vision Language Model
por: Zhang, Jingyuan, et al.
Publicado: (2025)
por: Zhang, Jingyuan, et al.
Publicado: (2025)
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
por: Li, Wenbo, et al.
Publicado: (2024)
por: Li, Wenbo, et al.
Publicado: (2024)
NuWa: Deriving Lightweight Task-Specific Vision Transformers for Edge Devices
por: Wei, Ziteng, et al.
Publicado: (2025)
por: Wei, Ziteng, et al.
Publicado: (2025)
A Review of Intelligent Device Fault Diagnosis Technologies Based on Machine Vision
por: Liu, Guiran, et al.
Publicado: (2024)
por: Liu, Guiran, et al.
Publicado: (2024)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
por: Kim, Eunki, et al.
Publicado: (2025)
por: Kim, Eunki, et al.
Publicado: (2025)
Ejemplares similares
-
Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
por: Nottebaum, Moritz, et al.
Publicado: (2026) -
LowFormer: Hardware Efficient Design for Convolutional Transformer Backbones
por: Nottebaum, Moritz, et al.
Publicado: (2024) -
Tracking Skiers from the Top to the Bottom
por: Dunnhofer, Matteo, et al.
Publicado: (2023) -
Is Tracking really more challenging in First Person Egocentric Vision?
por: Dunnhofer, Matteo, et al.
Publicado: (2025) -
Better, But Not Sufficient: Testing Video ANNs Against Macaque IT Dynamics
por: Dunnhofer, Matteo, et al.
Publicado: (2026)