Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hümmer, Christoph, Schwonberg, Manuel, Zhou, Liangwei, Cao, Hu, Knoll, Alois, Gottschalk, Hanno
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915057402642432
author Hümmer, Christoph
Schwonberg, Manuel
Zhou, Liangwei
Cao, Hu
Knoll, Alois
Gottschalk, Hanno
author_facet Hümmer, Christoph
Schwonberg, Manuel
Zhou, Liangwei
Cao, Hu
Knoll, Alois
Gottschalk, Hanno
contents Domain generalization (DG) remains a significant challenge for perception based on deep neural networks (DNNs), where domain shifts occur due to synthetic data, lighting, weather, or location changes. Vision-language models (VLMs) marked a large step for the generalization capabilities and have been already applied to various tasks. Very recently, first approaches utilized VLMs for domain generalized segmentation and object detection and obtained strong generalization. However, all these approaches rely on complex modules, feature augmentation frameworks or additional models. Surprisingly and in contrast to that, we found that simple fine-tuning of vision-language pre-trained models yields competitive or even stronger generalization results while being extremely simple to apply. Moreover, we found that vision-language pre-training consistently provides better generalization than the previous standard of vision-only pre-training. This challenges the standard of using ImageNet-based transfer learning for domain generalization. Fully fine-tuning a vision-language pre-trained model is capable of reaching the domain generalization SOTA when training on the synthetic GTA5 dataset. Moreover, we confirm this observation for object detection on a novel synthetic-to-real benchmark. We further obtain superior generalization capabilities by reaching 77.9% mIoU on the popular Cityscapes-to-ACDC benchmark. We also found improved in-domain generalization, leading to an improved SOTA of 86.4% mIoU on the Cityscapes test set marking the first place on the leaderboard.
format Preprint
id arxiv_https___arxiv_org_abs_2312_02021
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning
Hümmer, Christoph
Schwonberg, Manuel
Zhou, Liangwei
Cao, Hu
Knoll, Alois
Gottschalk, Hanno
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Domain generalization (DG) remains a significant challenge for perception based on deep neural networks (DNNs), where domain shifts occur due to synthetic data, lighting, weather, or location changes. Vision-language models (VLMs) marked a large step for the generalization capabilities and have been already applied to various tasks. Very recently, first approaches utilized VLMs for domain generalized segmentation and object detection and obtained strong generalization. However, all these approaches rely on complex modules, feature augmentation frameworks or additional models. Surprisingly and in contrast to that, we found that simple fine-tuning of vision-language pre-trained models yields competitive or even stronger generalization results while being extremely simple to apply. Moreover, we found that vision-language pre-training consistently provides better generalization than the previous standard of vision-only pre-training. This challenges the standard of using ImageNet-based transfer learning for domain generalization. Fully fine-tuning a vision-language pre-trained model is capable of reaching the domain generalization SOTA when training on the synthetic GTA5 dataset. Moreover, we confirm this observation for object detection on a novel synthetic-to-real benchmark. We further obtain superior generalization capabilities by reaching 77.9% mIoU on the popular Cityscapes-to-ACDC benchmark. We also found improved in-domain generalization, leading to an improved SOTA of 86.4% mIoU on the Cityscapes test set marking the first place on the leaderboard.
title Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.02021