ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Burgert, Tom, Stoll, Oliver, Rota, Paolo, Demir, Begüm
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917192205860864
author Burgert, Tom
Stoll, Oliver
Rota, Paolo
Demir, Begüm
author_facet Burgert, Tom
Stoll, Oliver
Rota, Paolo
Demir, Begüm
contents The hypothesis that Convolutional Neural Networks (CNNs) are inherently texture-biased has shaped much of the discourse on feature use in deep learning. We revisit this hypothesis by examining limitations in the cue-conflict experiment by Geirhos et al. To address these limitations, we propose a domain-agnostic framework that quantifies feature reliance through systematic suppression of shape, texture, and color cues, avoiding the confounds of forced-choice conflicts. By evaluating humans and neural networks under controlled suppression conditions, we find that CNNs are not inherently texture-biased but predominantly rely on local shape features. Nonetheless, this reliance can be substantially mitigated through modern training strategies or architectures (ConvNeXt, ViTs). We further extend the analysis across computer vision, medical imaging, and remote sensing, revealing that reliance patterns differ systematically: computer vision models prioritize shape, medical imaging models emphasize color, and remote sensing models exhibit a stronger reliance on texture. Code is available at https://github.com/tomburgert/feature-reliance.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20234
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
Burgert, Tom
Stoll, Oliver
Rota, Paolo
Demir, Begüm
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
The hypothesis that Convolutional Neural Networks (CNNs) are inherently texture-biased has shaped much of the discourse on feature use in deep learning. We revisit this hypothesis by examining limitations in the cue-conflict experiment by Geirhos et al. To address these limitations, we propose a domain-agnostic framework that quantifies feature reliance through systematic suppression of shape, texture, and color cues, avoiding the confounds of forced-choice conflicts. By evaluating humans and neural networks under controlled suppression conditions, we find that CNNs are not inherently texture-biased but predominantly rely on local shape features. Nonetheless, this reliance can be substantially mitigated through modern training strategies or architectures (ConvNeXt, ViTs). We further extend the analysis across computer vision, medical imaging, and remote sensing, revealing that reliance patterns differ systematically: computer vision models prioritize shape, medical imaging models emphasize color, and remote sensing models exhibit a stronger reliance on texture. Code is available at https://github.com/tomburgert/feature-reliance.
title ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.20234