LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Yirong, Geng, Yizhong, Wei, Peidong, Chen, Yanjun, Yang, Jinghan, Chen, Rongfei, Zhang, Wei, Shen, Xiaoyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909747011125248
author Sun, Yirong
Geng, Yizhong
Wei, Peidong
Chen, Yanjun
Yang, Jinghan
Chen, Rongfei
Zhang, Wei
Shen, Xiaoyu
author_facet Sun, Yirong
Geng, Yizhong
Wei, Peidong
Chen, Yanjun
Yang, Jinghan
Chen, Rongfei
Zhang, Wei
Shen, Xiaoyu
contents The development of Large Speech-Language Models (LSLMs) has been slowed by fragmented architectures and a lack of transparency, hindering the systematic comparison and reproducibility of research. Unlike in the vision-language domain, the LSLM field suffers from the common practice of releasing model weights without their corresponding training data and configurations. To address these critical gaps, we introduce LLaSO, the first fully open, end-to-end framework for large-scale speech-language modeling. LLaSO provides the community with three essential resources: (1) LLaSO-Align, a 12M-instance speech-text alignment corpus; (2) LLaSO-Instruct, a 13.5M-instance multi-task instruction-tuning dataset; and (3) LLaSO-Eval, a reproducible benchmark for standardized evaluation. To validate our framework, we build and release LLaSO-Base, a 3.8B-parameter reference model trained exclusively on our public data. It achieves a normalized score of 0.72, establishing a strong, reproducible baseline that surpasses comparable models. Our analysis reveals that while broader training coverage enhances performance, significant generalization gaps persist on unseen tasks, particularly in pure audio scenarios. By releasing the complete stack of data, benchmarks, and models, LLaSO establishes a foundational open standard to unify research efforts and accelerate community-driven progress in LSLMs. We release the code, dataset, pretrained models, and results in https://github.com/EIT-NLP/LLaSO.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15418
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
Sun, Yirong
Geng, Yizhong
Wei, Peidong
Chen, Yanjun
Yang, Jinghan
Chen, Rongfei
Zhang, Wei
Shen, Xiaoyu
Computation and Language
Artificial Intelligence
Machine Learning
Multimedia
Sound
The development of Large Speech-Language Models (LSLMs) has been slowed by fragmented architectures and a lack of transparency, hindering the systematic comparison and reproducibility of research. Unlike in the vision-language domain, the LSLM field suffers from the common practice of releasing model weights without their corresponding training data and configurations. To address these critical gaps, we introduce LLaSO, the first fully open, end-to-end framework for large-scale speech-language modeling. LLaSO provides the community with three essential resources: (1) LLaSO-Align, a 12M-instance speech-text alignment corpus; (2) LLaSO-Instruct, a 13.5M-instance multi-task instruction-tuning dataset; and (3) LLaSO-Eval, a reproducible benchmark for standardized evaluation. To validate our framework, we build and release LLaSO-Base, a 3.8B-parameter reference model trained exclusively on our public data. It achieves a normalized score of 0.72, establishing a strong, reproducible baseline that surpasses comparable models. Our analysis reveals that while broader training coverage enhances performance, significant generalization gaps persist on unseen tasks, particularly in pure audio scenarios. By releasing the complete stack of data, benchmarks, and models, LLaSO establishes a foundational open standard to unify research efforts and accelerate community-driven progress in LSLMs. We release the code, dataset, pretrained models, and results in https://github.com/EIT-NLP/LLaSO.
title LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
topic Computation and Language
Artificial Intelligence
Machine Learning
Multimedia
Sound
url https://arxiv.org/abs/2508.15418