Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Zhong, Zihan, Tang, Zhiqiang, He, Tong, Fang, Haoyang, Yuan, Chun
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866917579599118336
author Zhong, Zihan
Tang, Zhiqiang
He, Tong
Fang, Haoyang
Yuan, Chun
author_facet Zhong, Zihan
Tang, Zhiqiang
He, Tong
Fang, Haoyang
Yuan, Chun
contents The Segment Anything Model (SAM) stands as a foundational framework for image segmentation. While it exhibits remarkable zero-shot generalization in typical scenarios, its advantage diminishes when applied to specialized domains like medical imagery and remote sensing. To address this limitation, this paper introduces Conv-LoRA, a simple yet effective parameter-efficient fine-tuning approach. By integrating ultra-lightweight convolutional parameters into Low-Rank Adaptation (LoRA), Conv-LoRA can inject image-related inductive biases into the plain ViT encoder, further reinforcing SAM's local prior assumption. Notably, Conv-LoRA not only preserves SAM's extensive segmentation knowledge but also revives its capacity of learning high-level image semantics, which is constrained by SAM's foreground-background segmentation pretraining. Comprehensive experimentation across diverse benchmarks spanning multiple domains underscores Conv-LoRA's superiority in adapting SAM to real-world semantic segmentation tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2401_17868
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model
Zhong, Zihan
Tang, Zhiqiang
He, Tong
Fang, Haoyang
Yuan, Chun
Computer Vision and Pattern Recognition
Machine Learning
The Segment Anything Model (SAM) stands as a foundational framework for image segmentation. While it exhibits remarkable zero-shot generalization in typical scenarios, its advantage diminishes when applied to specialized domains like medical imagery and remote sensing. To address this limitation, this paper introduces Conv-LoRA, a simple yet effective parameter-efficient fine-tuning approach. By integrating ultra-lightweight convolutional parameters into Low-Rank Adaptation (LoRA), Conv-LoRA can inject image-related inductive biases into the plain ViT encoder, further reinforcing SAM's local prior assumption. Notably, Conv-LoRA not only preserves SAM's extensive segmentation knowledge but also revives its capacity of learning high-level image semantics, which is constrained by SAM's foreground-background segmentation pretraining. Comprehensive experimentation across diverse benchmarks spanning multiple domains underscores Conv-LoRA's superiority in adapting SAM to real-world semantic segmentation tasks.
title Convolution Meets LoRA: Parameter Efficient Finetuning for Segment Anything Model
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2401.17868