Advancing Speech Understanding in Speech-Aware Language Models with GRPO
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866911166330044416 |
|---|---|
| author | Elmakies, Avishai Aronowitz, Hagai Shabtay, Nimrod Schwartz, Eli Hoory, Ron Dekel, Avihu |
| author_facet | Elmakies, Avishai Aronowitz, Hagai Shabtay, Nimrod Schwartz, Eli Hoory, Ron Dekel, Avihu |
| contents | In this paper, we introduce a Group Relative Policy Optimization (GRPO)-based method for training Speech-Aware Large Language Models (SALLMs) on open-format speech understanding tasks, such as Spoken Question Answering and Automatic Speech Translation. SALLMs have proven highly effective for speech understanding tasks. GRPO has recently gained traction for its efficiency in training LLMs, and prior work has explored its application to SALLMs, primarily in multiple-choice tasks. Building on this, we focus on open-format tasks that better reflect the generative abilities of the models. Our approach leverages GRPO with BLEU as the reward signal to optimize SALLMs, and we demonstrate empirically that it surpasses standard SFT across several key metrics. Finally, we explore the potential of incorporating off-policy samples within GRPO for these tasks, highlighting avenues for further improvement and further research. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_16990 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Advancing Speech Understanding in Speech-Aware Language Models with GRPO Elmakies, Avishai Aronowitz, Hagai Shabtay, Nimrod Schwartz, Eli Hoory, Ron Dekel, Avihu Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing In this paper, we introduce a Group Relative Policy Optimization (GRPO)-based method for training Speech-Aware Large Language Models (SALLMs) on open-format speech understanding tasks, such as Spoken Question Answering and Automatic Speech Translation. SALLMs have proven highly effective for speech understanding tasks. GRPO has recently gained traction for its efficiency in training LLMs, and prior work has explored its application to SALLMs, primarily in multiple-choice tasks. Building on this, we focus on open-format tasks that better reflect the generative abilities of the models. Our approach leverages GRPO with BLEU as the reward signal to optimize SALLMs, and we demonstrate empirically that it surpasses standard SFT across several key metrics. Finally, we explore the potential of incorporating off-policy samples within GRPO for these tasks, highlighting avenues for further improvement and further research. |
| title | Advancing Speech Understanding in Speech-Aware Language Models with GRPO |
| topic | Computation and Language Artificial Intelligence Machine Learning Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.16990 |