CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Imam, Raza, Alam, Mohammed Talha, Rahman, Umaima, Guizani, Mohsen, Karray, Fakhri
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913582254391296
author Imam, Raza
Alam, Mohammed Talha
Rahman, Umaima
Guizani, Mohsen
Karray, Fakhri
author_facet Imam, Raza
Alam, Mohammed Talha
Rahman, Umaima
Guizani, Mohsen
Karray, Fakhri
contents Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushing unrelated pairs apart. However, astronomical image-label datasets are significantly smaller compared to general image and label datasets available from the internet. We introduce CosmoCLIP, an astronomical image-text contrastive learning framework precisely fine-tuned on the pre-trained CLIP model using SpaceNet and BLIP-based captions. SpaceNet, attained via FLARE, constitutes ~13k optimally distributed images, while BLIP acts as a rich knowledge extractor. The rich semantics derived from this SpaceNet and BLIP descriptions, when learned contrastively, enable CosmoCLIP to achieve superior generalization across various in-domain and out-of-domain tasks. Our results demonstrate that CosmoCLIP is a straightforward yet powerful framework, significantly outperforming CLIP in zero-shot classification and image-text retrieval tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2407_07315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
Imam, Raza
Alam, Mohammed Talha
Rahman, Umaima
Guizani, Mohsen
Karray, Fakhri
Computer Vision and Pattern Recognition
Existing vision-text contrastive learning models enhance representation transferability and support zero-shot prediction by matching paired image and caption embeddings while pushing unrelated pairs apart. However, astronomical image-label datasets are significantly smaller compared to general image and label datasets available from the internet. We introduce CosmoCLIP, an astronomical image-text contrastive learning framework precisely fine-tuned on the pre-trained CLIP model using SpaceNet and BLIP-based captions. SpaceNet, attained via FLARE, constitutes ~13k optimally distributed images, while BLIP acts as a rich knowledge extractor. The rich semantics derived from this SpaceNet and BLIP descriptions, when learned contrastively, enable CosmoCLIP to achieve superior generalization across various in-domain and out-of-domain tasks. Our results demonstrate that CosmoCLIP is a straightforward yet powerful framework, significantly outperforming CLIP in zero-shot classification and image-text retrieval tasks.
title CosmoCLIP: Generalizing Large Vision-Language Models for Astronomical Imaging
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.07315