On the Influence of Shape, Texture and Color for Learning Semantic Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mütze, Annika, Grabowsky, Natalie, Heinert, Edgar, Rottmann, Matthias, Gottschalk, Hanno
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909866362142720
author Mütze, Annika
Grabowsky, Natalie
Heinert, Edgar
Rottmann, Matthias
Gottschalk, Hanno
author_facet Mütze, Annika
Grabowsky, Natalie
Heinert, Edgar
Rottmann, Matthias
Gottschalk, Hanno
contents Recent research has investigated the shape and texture biases of pre-trained deep neural networks (DNNs) in image classification. Those works test how much a trained DNN relies on specific image cues like texture. The present study shifts the focus to understanding the cue influence during training, analyzing what DNNs can learn from shape, texture, and color cues in absence of the others; investigating their individual and combined influence on the learning success. We analyze these cue influences at multiple levels by decomposing datasets into cue-specific versions. Addressing semantic segmentation, we learn the given task from these reduced cue datasets, creating cue experts. Early fusion of cues is performed by constructing appropriate datasets. This is complemented by a late fusion of experts which allows us to study cue influence location-dependent on pixel level. Experiments on Cityscapes, PASCAL Context, and a synthetic CARLA dataset show that while no single cue dominates, the shape + color expert predominantly improves the prediction of small objects and border pixels. The cue performance order is consistent for the tested convolutional and transformer architecture, indicating similar cue extraction capabilities, although pre-trained transformers are said to be more biased towards shape than convolutional neural networks.
format Preprint
id arxiv_https___arxiv_org_abs_2410_14878
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle On the Influence of Shape, Texture and Color for Learning Semantic Segmentation
Mütze, Annika
Grabowsky, Natalie
Heinert, Edgar
Rottmann, Matthias
Gottschalk, Hanno
Computer Vision and Pattern Recognition
Recent research has investigated the shape and texture biases of pre-trained deep neural networks (DNNs) in image classification. Those works test how much a trained DNN relies on specific image cues like texture. The present study shifts the focus to understanding the cue influence during training, analyzing what DNNs can learn from shape, texture, and color cues in absence of the others; investigating their individual and combined influence on the learning success. We analyze these cue influences at multiple levels by decomposing datasets into cue-specific versions. Addressing semantic segmentation, we learn the given task from these reduced cue datasets, creating cue experts. Early fusion of cues is performed by constructing appropriate datasets. This is complemented by a late fusion of experts which allows us to study cue influence location-dependent on pixel level. Experiments on Cityscapes, PASCAL Context, and a synthetic CARLA dataset show that while no single cue dominates, the shape + color expert predominantly improves the prediction of small objects and border pixels. The cue performance order is consistent for the tested convolutional and transformer architecture, indicating similar cue extraction capabilities, although pre-trained transformers are said to be more biased towards shape than convolutional neural networks.
title On the Influence of Shape, Texture and Color for Learning Semantic Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.14878