SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Qiaolin, Jiang, Xilin, He, Linyang, Wu, Junkai, Mesgarani, Nima
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866909797571362816
author Wang, Qiaolin
Jiang, Xilin
He, Linyang
Wu, Junkai
Mesgarani, Nima
author_facet Wang, Qiaolin
Jiang, Xilin
He, Linyang
Wu, Junkai
Mesgarani, Nima
contents While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio data to teach LALM stepwise reasoning. To circumvent this data and modality gap, we present SightSound-R1, a cross-modal distillation framework that transfers advanced reasoning from a stronger LVLM teacher to a weaker LALM student on the same audio-visual question answering (AVQA) dataset. SightSound-R1 consists of three core steps: (i) test-time scaling to generate audio-focused chains of thought (CoT) from an LVLM teacher, (ii) audio-grounded validation to filter hallucinations, and (iii) a distillation pipeline with supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) for the LALM student. Results show that SightSound-R1 improves LALM reasoning performance both in the in-domain AVQA test set as well as in unseen auditory scenes and questions, outperforming both pretrained and label-only distilled baselines. Thus, we conclude that vision reasoning can be effectively transferred to audio models and scaled with abundant audio-visual data.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
Wang, Qiaolin
Jiang, Xilin
He, Linyang
Wu, Junkai
Mesgarani, Nima
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
While large audio-language models (LALMs) have demonstrated state-of-the-art audio understanding, their reasoning capability in complex soundscapes still falls behind large vision-language models (LVLMs). Compared to the visual domain, one bottleneck is the lack of large-scale chain-of-thought audio data to teach LALM stepwise reasoning. To circumvent this data and modality gap, we present SightSound-R1, a cross-modal distillation framework that transfers advanced reasoning from a stronger LVLM teacher to a weaker LALM student on the same audio-visual question answering (AVQA) dataset. SightSound-R1 consists of three core steps: (i) test-time scaling to generate audio-focused chains of thought (CoT) from an LVLM teacher, (ii) audio-grounded validation to filter hallucinations, and (iii) a distillation pipeline with supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) for the LALM student. Results show that SightSound-R1 improves LALM reasoning performance both in the in-domain AVQA test set as well as in unseen auditory scenes and questions, outperforming both pretrained and label-only distilled baselines. Thus, we conclude that vision reasoning can be effectively transferred to audio models and scaled with abundant audio-visual data.
title SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2509.15661