IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918356905361408 |
|---|---|
| author | Zhang, Yaming Gao, Chenqiang Liu, Fangcen Guo, Junjie Wang, Lan Peng, Xinggan Meng, Deyu |
| author_facet | Zhang, Yaming Gao, Chenqiang Liu, Fangcen Guo, Junjie Wang, Lan Peng, Xinggan Meng, Deyu |
| contents | Existing infrared and visible (IR-VIS) methods inherit the general representations of Pre-trained Visual Models (PVMs) to facilitate complementary learning. However, our analysis indicates that under the full fine-tuning paradigm, the feature space becomes highly constrained and low-ranked, which has been proven to seriously impair generalization. One remedy is to freeze the parameters, which preserves pretrained knowledge and helps maintain feature diversity. To this end, we propose IV-tuning, to parameter-efficiently harness PVMs for various IR-VIS downstream tasks, including salient object detection, semantic segmentation, and object detection. Extensive experiments across various settings demonstrate that IV-tuning outperforms previous state-of-the-art methods, and exhibits superior generalization and scalability. Remarkably, with only a single backbone, IV-tuning effectively facilitates the complementary learning of infrared and visible modalities with merely 3% trainable backbone parameters, and achieves superior computational efficiency compared to conventional IR-VIS paradigms. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_16654 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks Zhang, Yaming Gao, Chenqiang Liu, Fangcen Guo, Junjie Wang, Lan Peng, Xinggan Meng, Deyu Computer Vision and Pattern Recognition Existing infrared and visible (IR-VIS) methods inherit the general representations of Pre-trained Visual Models (PVMs) to facilitate complementary learning. However, our analysis indicates that under the full fine-tuning paradigm, the feature space becomes highly constrained and low-ranked, which has been proven to seriously impair generalization. One remedy is to freeze the parameters, which preserves pretrained knowledge and helps maintain feature diversity. To this end, we propose IV-tuning, to parameter-efficiently harness PVMs for various IR-VIS downstream tasks, including salient object detection, semantic segmentation, and object detection. Extensive experiments across various settings demonstrate that IV-tuning outperforms previous state-of-the-art methods, and exhibits superior generalization and scalability. Remarkably, with only a single backbone, IV-tuning effectively facilitates the complementary learning of infrared and visible modalities with merely 3% trainable backbone parameters, and achieves superior computational efficiency compared to conventional IR-VIS paradigms. |
| title | IV-tuning: Parameter-Efficient Transfer Learning for Infrared-Visible Tasks |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.16654 |