DiffCLIP: Differential Attention Meets CLIP

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hammoud, Hasan Abed Al Kader, Ghanem, Bernard
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929750750003200
author Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
author_facet Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
contents We propose DiffCLIP, a novel vision-language model that extends the differential attention mechanism to CLIP architectures. Differential attention was originally developed for large language models to amplify relevant context while canceling out noisy information. In this work, we integrate this mechanism into CLIP's dual encoder (image and text) framework. With minimal additional parameters, DiffCLIP achieves superior performance on image-text understanding tasks. Across zero-shot classification, retrieval, and robustness benchmarks, DiffCLIP consistently outperforms baseline CLIP models. Notably, these gains come with negligible computational overhead, demonstrating that differential attention can significantly enhance multi-modal representations without sacrificing efficiency. Code can be found at https://github.com/hammoudhasan/DiffCLIP.
format Preprint
id arxiv_https___arxiv_org_abs_2503_06626
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DiffCLIP: Differential Attention Meets CLIP
Hammoud, Hasan Abed Al Kader
Ghanem, Bernard
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
We propose DiffCLIP, a novel vision-language model that extends the differential attention mechanism to CLIP architectures. Differential attention was originally developed for large language models to amplify relevant context while canceling out noisy information. In this work, we integrate this mechanism into CLIP's dual encoder (image and text) framework. With minimal additional parameters, DiffCLIP achieves superior performance on image-text understanding tasks. Across zero-shot classification, retrieval, and robustness benchmarks, DiffCLIP consistently outperforms baseline CLIP models. Notably, these gains come with negligible computational overhead, demonstrating that differential attention can significantly enhance multi-modal representations without sacrificing efficiency. Code can be found at https://github.com/hammoudhasan/DiffCLIP.
title DiffCLIP: Differential Attention Meets CLIP
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2503.06626