radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Riggi, S., Cecconello, T., Pilzer, A., Palazzo, S., Gupta, N., Hopkins, A. M., Trigilio, C., Umana, G.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918108987392000
author Riggi, S.
Cecconello, T.
Pilzer, A.
Palazzo, S.
Gupta, N.
Hopkins, A. M.
Trigilio, C.
Umana, G.
author_facet Riggi, S.
Cecconello, T.
Pilzer, A.
Palazzo, S.
Gupta, N.
Hopkins, A. M.
Trigilio, C.
Umana, G.
contents The advent of next-generation radio telescopes is set to transform radio astronomy by producing massive data volumes that challenge traditional processing methods. Deep learning techniques have shown strong potential in automating radio analysis tasks, yet are often constrained by the limited availability of large annotated datasets. Recent progress in self-supervised learning has led to foundational radio vision models, but adapting them for new tasks typically requires coding expertise, limiting their accessibility to a broader astronomical community. Text-based AI interfaces offer a promising alternative by enabling task-specific queries and example-driven learning. In this context, Large Language Models (LLMs), with their remarkable zero-shot capabilities, are increasingly used in scientific domains. However, deploying large-scale models remains resource-intensive, and there is a growing demand for AI systems that can reason over both visual and textual data in astronomical analysis. This study explores small-scale Vision-Language Models (VLMs) as AI assistants for radio astronomy, combining LLM capabilities with vision transformers. We fine-tuned the LLaVA VLM on a dataset of 59k radio images from multiple surveys, enriched with 38k image-caption pairs from the literature. The fine-tuned models show clear improvements over base models in radio-specific tasks, achieving ~30% F1-score gains in extended source detection, but they underperform vision-only classifiers and exhibit ~20% drop on general multimodal tasks. Inclusion of caption data and LoRA fine-tuning enhances instruction-following and helps recover ~10% accuracy on multimodal benchmarks. This work lays the foundation for future advancements in radio VLMs, highlighting their potential and limitations, such as the need for better multimodal alignment, higher-quality datasets, and mitigation of catastrophic forgetting.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis
Riggi, S.
Cecconello, T.
Pilzer, A.
Palazzo, S.
Gupta, N.
Hopkins, A. M.
Trigilio, C.
Umana, G.
Instrumentation and Methods for Astrophysics
The advent of next-generation radio telescopes is set to transform radio astronomy by producing massive data volumes that challenge traditional processing methods. Deep learning techniques have shown strong potential in automating radio analysis tasks, yet are often constrained by the limited availability of large annotated datasets. Recent progress in self-supervised learning has led to foundational radio vision models, but adapting them for new tasks typically requires coding expertise, limiting their accessibility to a broader astronomical community. Text-based AI interfaces offer a promising alternative by enabling task-specific queries and example-driven learning. In this context, Large Language Models (LLMs), with their remarkable zero-shot capabilities, are increasingly used in scientific domains. However, deploying large-scale models remains resource-intensive, and there is a growing demand for AI systems that can reason over both visual and textual data in astronomical analysis. This study explores small-scale Vision-Language Models (VLMs) as AI assistants for radio astronomy, combining LLM capabilities with vision transformers. We fine-tuned the LLaVA VLM on a dataset of 59k radio images from multiple surveys, enriched with 38k image-caption pairs from the literature. The fine-tuned models show clear improvements over base models in radio-specific tasks, achieving ~30% F1-score gains in extended source detection, but they underperform vision-only classifiers and exhibit ~20% drop on general multimodal tasks. Inclusion of caption data and LoRA fine-tuning enhances instruction-following and helps recover ~10% accuracy on multimodal benchmarks. This work lays the foundation for future advancements in radio VLMs, highlighting their potential and limitations, such as the need for better multimodal alignment, higher-quality datasets, and mitigation of catastrophic forgetting.
title radio-llava: Advancing Vision-Language Models for Radio Astronomical Source Analysis
topic Instrumentation and Methods for Astrophysics
url https://arxiv.org/abs/2503.23859