IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yaming, Gao, Chenqiang, Liu, Fangcen, Guo, Junjie, Wang, Lan, Peng, Xinggan, Meng, Deyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918356905361408
author Zhang, Yaming
Gao, Chenqiang
Liu, Fangcen
Guo, Junjie
Wang, Lan
Peng, Xinggan
Meng, Deyu
author_facet Zhang, Yaming
Gao, Chenqiang
Liu, Fangcen
Guo, Junjie
Wang, Lan
Peng, Xinggan
Meng, Deyu
contents Existing infrared and visible (IR-VIS) methods inherit the general representations of Pre-trained Visual Models (PVMs) to facilitate complementary learning. However, our analysis indicates that under the full fine-tuning paradigm, the feature space becomes highly constrained and low-ranked, which has been proven to seriously impair generalization. One remedy is to freeze the parameters, which preserves pretrained knowledge and helps maintain feature diversity. To this end, we propose IV-tuning, to parameter-efficiently harness PVMs for various IR-VIS downstream tasks, including salient object detection, semantic segmentation, and object detection. Extensive experiments across various settings demonstrate that IV-tuning outperforms previous state-of-the-art methods, and exhibits superior generalization and scalability. Remarkably, with only a single backbone, IV-tuning effectively facilitates the complementary learning of infrared and visible modalities with merely 3% trainable backbone parameters, and achieves superior computational efficiency compared to conventional IR-VIS paradigms.
format Preprint
id arxiv_https___arxiv_org_abs_2412_16654
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks
Zhang, Yaming
Gao, Chenqiang
Liu, Fangcen
Guo, Junjie
Wang, Lan
Peng, Xinggan
Meng, Deyu
Computer Vision and Pattern Recognition
Existing infrared and visible (IR-VIS) methods inherit the general representations of Pre-trained Visual Models (PVMs) to facilitate complementary learning. However, our analysis indicates that under the full fine-tuning paradigm, the feature space becomes highly constrained and low-ranked, which has been proven to seriously impair generalization. One remedy is to freeze the parameters, which preserves pretrained knowledge and helps maintain feature diversity. To this end, we propose IV-tuning, to parameter-efficiently harness PVMs for various IR-VIS downstream tasks, including salient object detection, semantic segmentation, and object detection. Extensive experiments across various settings demonstrate that IV-tuning outperforms previous state-of-the-art methods, and exhibits superior generalization and scalability. Remarkably, with only a single backbone, IV-tuning effectively facilitates the complementary learning of infrared and visible modalities with merely 3% trainable backbone parameters, and achieves superior computational efficiency compared to conventional IR-VIS paradigms.
title IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.16654