Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jeffrey, Gregory, Jonathan, Chrysos, Grigorios G. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MetaFormer Baselines for Vision
by: Yu, Weihao, et al.
Published: (2022)
by: Yu, Weihao, et al.
Published: (2022)
High-throughput digital twin framework for predicting neurite deterioration using MetaFormer attention
by: Qian, Kuanren, et al.
Published: (2024)
by: Qian, Kuanren, et al.
Published: (2024)
Multilinear Operator Networks
by: Cheng, Yixin, et al.
Published: (2024)
by: Cheng, Yixin, et al.
Published: (2024)
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
by: Triaridis, Kostas, et al.
Published: (2025)
by: Triaridis, Kostas, et al.
Published: (2025)
Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
by: Keuth, Ron, et al.
Published: (2025)
by: Keuth, Ron, et al.
Published: (2025)
The Last Mile to Supervised Performance: Semi-Supervised Domain Adaptation for Semantic Segmentation
by: Morales-Brotons, Daniel, et al.
Published: (2024)
by: Morales-Brotons, Daniel, et al.
Published: (2024)
MetaFormer-driven Encoding Network for Robust Medical Semantic Segmentation
by: Tran, Le-Anh, et al.
Published: (2026)
by: Tran, Le-Anh, et al.
Published: (2026)
mRadNet: A Compact Radar Object Detector with MetaFormer
by: Chen, Huaiyu, et al.
Published: (2025)
by: Chen, Huaiyu, et al.
Published: (2025)
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
by: Oldfield, James, et al.
Published: (2024)
by: Oldfield, James, et al.
Published: (2024)
MetaSeg: MetaFormer-based Global Contexts-aware Network for Efficient Semantic Segmentation
by: Kang, Beoungwoo, et al.
Published: (2024)
by: Kang, Beoungwoo, et al.
Published: (2024)
VisionTS++: Cross-Modal Time Series Foundation Model with Continual Pre-trained Vision Backbones
by: Shen, Lefei, et al.
Published: (2025)
by: Shen, Lefei, et al.
Published: (2025)
Rethinking Attention: Polynomial Alternatives to Softmax in Transformers
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Vision Backbone Efficient Selection for Image Classification in Low-Data Regimes
by: Guerin, Joris, et al.
Published: (2024)
by: Guerin, Joris, et al.
Published: (2024)
PNeRV: A Polynomial Neural Representation for Videos
by: Gupta, Sonam, et al.
Published: (2024)
by: Gupta, Sonam, et al.
Published: (2024)
Vision-LSTM: xLSTM as Generic Vision Backbone
by: Alkin, Benedikt, et al.
Published: (2024)
by: Alkin, Benedikt, et al.
Published: (2024)
MI-NeRF: Learning a Single Face NeRF from Multiple Identities
by: Chatziagapi, Aggelina, et al.
Published: (2024)
by: Chatziagapi, Aggelina, et al.
Published: (2024)
MIGS: Multi-Identity Gaussian Splatting via Tensor Decomposition
by: Chatziagapi, Aggelina, et al.
Published: (2024)
by: Chatziagapi, Aggelina, et al.
Published: (2024)
AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer
by: Shan, Jiquan, et al.
Published: (2025)
by: Shan, Jiquan, et al.
Published: (2025)
Explaining the Impact of Training on Vision Models via Activation Clustering
by: Boubekki, Ahcène, et al.
Published: (2024)
by: Boubekki, Ahcène, et al.
Published: (2024)
ViR: Towards Efficient Vision Retention Backbones
by: Hatamizadeh, Ali, et al.
Published: (2023)
by: Hatamizadeh, Ali, et al.
Published: (2023)
Pre-Trained Vision Models as Perception Backbones for Safety Filters in Autonomous Driving
by: Yang, Yuxuan, et al.
Published: (2024)
by: Yang, Yuxuan, et al.
Published: (2024)
Exploiting the Exact Denoising Posterior Score in Training-Free Guidance of Diffusion Models
by: Bellchambers, Gregory
Published: (2025)
by: Bellchambers, Gregory
Published: (2025)
Semi-Supervised Fine-Tuning of Vision Foundation Models with Content-Style Decomposition
by: Drozdova, Mariia, et al.
Published: (2024)
by: Drozdova, Mariia, et al.
Published: (2024)
When Does Sparse MoE Help in Vision? The Role of Backbone Compute Leverage in Sparse Routing
by: Sun, Libo, et al.
Published: (2026)
by: Sun, Libo, et al.
Published: (2026)
Unlocking Pre-trained Image Backbones for Semantic Image Synthesis
by: Berrada, Tariq, et al.
Published: (2023)
by: Berrada, Tariq, et al.
Published: (2023)
Polyp Segmentation Generalisability of Pretrained Backbones
by: Sanderson, Edward, et al.
Published: (2024)
by: Sanderson, Edward, et al.
Published: (2024)
PerFormer: A Permutation Based Vision Transformer for Remaining Useful Life Prediction
by: Fan, Zhengyang, et al.
Published: (2025)
by: Fan, Zhengyang, et al.
Published: (2025)
SpecFormer: Guarding Vision Transformer Robustness via Maximum Singular Value Penalization
by: Hu, Xixu, et al.
Published: (2024)
by: Hu, Xixu, et al.
Published: (2024)
GeoFormer: A Vision and Sequence Transformer-based Approach for Greenhouse Gas Monitoring
by: Khirwar, Madhav, et al.
Published: (2024)
by: Khirwar, Madhav, et al.
Published: (2024)
Enhancing Zero-Shot Image Recognition in Vision-Language Models through Human-like Concept Guidance
by: Liu, Hui, et al.
Published: (2025)
by: Liu, Hui, et al.
Published: (2025)
SAFE-KD: Risk-Controlled Early-Exit Distillation for Vision Backbones
by: Khazem, Salim
Published: (2026)
by: Khazem, Salim
Published: (2026)
Graph Neural Networks in Vision-Language Image Understanding: A Survey
by: Senior, Henry, et al.
Published: (2023)
by: Senior, Henry, et al.
Published: (2023)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
by: Tschannen, Michael, et al.
Published: (2024)
by: Tschannen, Michael, et al.
Published: (2024)
ParFormer: A Vision Transformer with Parallel Mixer and Sparse Channel Attention Patch Embedding
by: Setyawan, Novendra, et al.
Published: (2024)
by: Setyawan, Novendra, et al.
Published: (2024)
MetaMix: Meta-state Precision Searcher for Mixed-precision Activation Quantization
by: Kim, Han-Byul, et al.
Published: (2023)
by: Kim, Han-Byul, et al.
Published: (2023)
Enhancing Fingerprint Image Synthesis with GANs, Diffusion Models, and Style Transfer Techniques
by: Tang, W., et al.
Published: (2024)
by: Tang, W., et al.
Published: (2024)
Balanced Image Stylization with Style Matching Score
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
Text-Guided Alternative Image Clustering
by: Stephan, Andreas, et al.
Published: (2024)
by: Stephan, Andreas, et al.
Published: (2024)
Using Backbone Foundation Model for Evaluating Fairness in Chest Radiography Without Demographic Data
by: Queiroz, Dilermando, et al.
Published: (2024)
by: Queiroz, Dilermando, et al.
Published: (2024)
Flatness Improves Backbone Generalisation in Few-shot Classification
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Similar Items
-
MetaFormer Baselines for Vision
by: Yu, Weihao, et al.
Published: (2022) -
High-throughput digital twin framework for predicting neurite deterioration using MetaFormer attention
by: Qian, Kuanren, et al.
Published: (2024) -
Multilinear Operator Networks
by: Cheng, Yixin, et al.
Published: (2024) -
Mitigating Diffusion Model Hallucinations with Dynamic Guidance
by: Triaridis, Kostas, et al.
Published: (2025) -
Shaken or Stirred? An Analysis of MetaFormer's Token Mixing for Medical Imaging
by: Keuth, Ron, et al.
Published: (2025)