Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jeffrey, Gregory, Jonathan, Chrysos, Grigorios G.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910239879593984
author Wang, Jeffrey
Gregory, Jonathan
Chrysos, Grigorios G.
author_facet Wang, Jeffrey
Gregory, Jonathan
Chrysos, Grigorios G.
contents Modern vision backbones treat pointwise activations (e.g., ReLU, GELU) and exponential softmax as essential sources of nonlinearity, but we demonstrate they are not required within MetaFormer-style vision backbones. We design activation-free polynomial alternatives for three core primitives (MLPs, convolutions, and attention), where Hadamard products replace standard nonlinearities to yield polynomial functions of the input. These modules integrate seamlessly into existing architectures: instantiated within MetaFormer, a modular framework for vision backbones, our PolyNeXt models match or exceed activation-based counterparts across model scales on ImageNet classification, ADE20K semantic segmentation, and out-of-distribution robustness. We also substantially outperform prior polynomial networks at reduced computational cost, showing that polynomial variants of standard modules beat complex custom architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2605_20839
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models
Wang, Jeffrey
Gregory, Jonathan
Chrysos, Grigorios G.
Computer Vision and Pattern Recognition
Machine Learning
Modern vision backbones treat pointwise activations (e.g., ReLU, GELU) and exponential softmax as essential sources of nonlinearity, but we demonstrate they are not required within MetaFormer-style vision backbones. We design activation-free polynomial alternatives for three core primitives (MLPs, convolutions, and attention), where Hadamard products replace standard nonlinearities to yield polynomial functions of the input. These modules integrate seamlessly into existing architectures: instantiated within MetaFormer, a modular framework for vision backbones, our PolyNeXt models match or exceed activation-based counterparts across model scales on ImageNet classification, ADE20K semantic segmentation, and out-of-distribution robustness. We also substantially outperform prior polynomial networks at reduced computational cost, showing that polynomial variants of standard modules beat complex custom architectures.
title Activation-Free Backbones for Image Recognition: Polynomial Alternatives within MetaFormer-Style Vision Models
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.20839