PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mishra-Sharma, Siddharth, Song, Yiding, Thaler, Jesse
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916159512641536
author Mishra-Sharma, Siddharth
Song, Yiding
Thaler, Jesse
author_facet Mishra-Sharma, Siddharth
Song, Yiding
Thaler, Jesse
contents We present PAPERCLIP (Proposal Abstracts Provide an Effective Representation for Contrastive Language-Image Pre-training), a method which associates astronomical observations imaged by telescopes with natural language using a neural network model. The model is fine-tuned from a pre-trained Contrastive Language-Image Pre-training (CLIP) model using successful observing proposal abstracts and corresponding downstream observations, with the abstracts optionally summarized via guided generation using large language models (LLMs). Using observations from the Hubble Space Telescope (HST) as an example, we show that the fine-tuned model embodies a meaningful joint representation between observations and natural language through tests targeting image retrieval (i.e., finding the most relevant observations using natural language queries) and description retrieval (i.e., querying for astrophysical object classes and use cases most relevant to a given observation). Our study demonstrates the potential for using generalist foundation models rather than task-specific models for interacting with astronomical data by leveraging text as an interface.
format Preprint
id arxiv_https___arxiv_org_abs_2403_08851
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models
Mishra-Sharma, Siddharth
Song, Yiding
Thaler, Jesse
Instrumentation and Methods for Astrophysics
Computation and Language
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
We present PAPERCLIP (Proposal Abstracts Provide an Effective Representation for Contrastive Language-Image Pre-training), a method which associates astronomical observations imaged by telescopes with natural language using a neural network model. The model is fine-tuned from a pre-trained Contrastive Language-Image Pre-training (CLIP) model using successful observing proposal abstracts and corresponding downstream observations, with the abstracts optionally summarized via guided generation using large language models (LLMs). Using observations from the Hubble Space Telescope (HST) as an example, we show that the fine-tuned model embodies a meaningful joint representation between observations and natural language through tests targeting image retrieval (i.e., finding the most relevant observations using natural language queries) and description retrieval (i.e., querying for astrophysical object classes and use cases most relevant to a given observation). Our study demonstrates the potential for using generalist foundation models rather than task-specific models for interacting with astronomical data by leveraging text as an interface.
title PAPERCLIP: Associating Astronomical Observations and Natural Language with Multi-Modal Models
topic Instrumentation and Methods for Astrophysics
Computation and Language
Computer Vision and Pattern Recognition
Information Retrieval
Machine Learning
url https://arxiv.org/abs/2403.08851