SSR: Alignment-Aware Modality Connector for Speech Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910949001134080 |
|---|---|
| author | Tan, Weiting Inaguma, Hirofumi Dong, Ning Tomasello, Paden Ma, Xutai |
| author_facet | Tan, Weiting Inaguma, Hirofumi Dong, Ning Tomasello, Paden Ma, Xutai |
| contents | Fusing speech into pre-trained language model (SpeechLM) usually suffers from inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality. We propose SSR-Connector (Segmented Speech Representation Connector) for better modality fusion. Leveraging speech-text alignments, our approach segments and compresses speech features to match the granularity of text embeddings. Additionally, we introduce a two-stage training pipeline that includes the distillation and fine-tuning phases to mitigate catastrophic forgetting. SSR-Connector outperforms existing mechanism for speech-text modality fusion, consistently achieving better speech understanding (e.g., +10 accuracy on StoryCloze and +20 on Speech-MMLU) while preserving pre-trained text ability. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_00168 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | SSR: Alignment-Aware Modality Connector for Speech Language Models Tan, Weiting Inaguma, Hirofumi Dong, Ning Tomasello, Paden Ma, Xutai Computation and Language Sound Audio and Speech Processing Fusing speech into pre-trained language model (SpeechLM) usually suffers from inefficient encoding of long-form speech and catastrophic forgetting of pre-trained text modality. We propose SSR-Connector (Segmented Speech Representation Connector) for better modality fusion. Leveraging speech-text alignments, our approach segments and compresses speech features to match the granularity of text embeddings. Additionally, we introduce a two-stage training pipeline that includes the distillation and fine-tuning phases to mitigate catastrophic forgetting. SSR-Connector outperforms existing mechanism for speech-text modality fusion, consistently achieving better speech understanding (e.g., +10 accuracy on StoryCloze and +20 on Speech-MMLU) while preserving pre-trained text ability. |
| title | SSR: Alignment-Aware Modality Connector for Speech Language Models |
| topic | Computation and Language Sound Audio and Speech Processing |
| url | https://arxiv.org/abs/2410.00168 |