CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913582254391296 |
|---|---|
| author | Imam, Raza Alam, Mohammed Talha Rahman, Umaima Guizani, Mohsen Karray, Fakhri |
| author_facet | Imam, Raza Alam, Mohammed Talha Rahman, Umaima Guizani, Mohsen Karray, Fakhri |
| contents | Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushing unrelated pairs apart. However, astronomical image-label datasets are significantly smaller compared to general image and label datasets available from the internet. We introduce CosmoCLIP, an astronomical image-text contrastive learning framework precisely fine-tuned on the pre-trained CLIP model using SpaceNet and BLIP-based captions. SpaceNet, attained via FLARE, constitutes ~13k optimally distributed images, while BLIP acts as a rich knowledge extractor. The rich semantics derived from this SpaceNet and BLIP descriptions, when learned contrastively, enable CosmoCLIP to achieve superior generalization across various in-domain and out-of-domain tasks. Our results demonstrate that CosmoCLIP is a straightforward yet powerful framework, significantly outperforming CLIP in zero-shot classification and image-text retrieval tasks. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2407_07315 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging Imam, Raza Alam, Mohammed Talha Rahman, Umaima Guizani, Mohsen Karray, Fakhri Computer Vision and Pattern Recognition Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushing unrelated pairs apart. However, astronomical image-label datasets are significantly smaller compared to general image and label datasets available from the internet. We introduce CosmoCLIP, an astronomical image-text contrastive learning framework precisely fine-tuned on the pre-trained CLIP model using SpaceNet and BLIP-based captions. SpaceNet, attained via FLARE, constitutes ~13k optimally distributed images, while BLIP acts as a rich knowledge extractor. The rich semantics derived from this SpaceNet and BLIP descriptions, when learned contrastively, enable CosmoCLIP to achieve superior generalization across various in-domain and out-of-domain tasks. Our results demonstrate that CosmoCLIP is a straightforward yet powerful framework, significantly outperforming CLIP in zero-shot classification and image-text retrieval tasks. |
| title | CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2407.07315 |