Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Jianfeng, Salazar, Julian, Yao, Kaisheng, Li, Haoqi, Cai, Jinglun
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909091050291200
author He, Jianfeng
Salazar, Julian
Yao, Kaisheng
Li, Haoqi
Cai, Jinglun
author_facet He, Jianfeng
Salazar, Julian
Yao, Kaisheng
Li, Haoqi
Cai, Jinglun
contents End-to-end (E2E) spoken language understanding (SLU) is constrained by the cost of collecting speech-semantics pairs, especially when label domains change. Hence, we explore \textit{zero-shot} E2E SLU, which learns E2E SLU without speech-semantics pairs, instead using only speech-text and text-semantics pairs. Previous work achieved zero-shot by pseudolabeling all speech-text transcripts with a natural language understanding (NLU) model learned on text-semantics corpora. However, this method requires the domains of speech-text and text-semantics to match, which often mismatch due to separate collections. Furthermore, using the entire collected speech-text corpus from any domains leads to \textit{imbalance} and \textit{noise} issues. To address these, we propose \textit{cross-modal selective self-training} (CMSST). CMSST tackles imbalance by clustering in a joint space of the three modalities (speech, text, and semantics) and handles label noise with a selection network. We also introduce two benchmarks for zero-shot E2E SLU, covering matched and found speech (mismatched) settings. Experiments show that CMSST improves performance in both two settings, with significantly reduced sample sizes and training time. Our code and data are released in https://github.com/amazon-science/zero-shot-E2E-slu.
format Preprint
id arxiv_https___arxiv_org_abs_2305_12793
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
He, Jianfeng
Salazar, Julian
Yao, Kaisheng
Li, Haoqi
Cai, Jinglun
Audio and Speech Processing
Computation and Language
Multimedia
Sound
End-to-end (E2E) spoken language understanding (SLU) is constrained by the cost of collecting speech-semantics pairs, especially when label domains change. Hence, we explore \textit{zero-shot} E2E SLU, which learns E2E SLU without speech-semantics pairs, instead using only speech-text and text-semantics pairs. Previous work achieved zero-shot by pseudolabeling all speech-text transcripts with a natural language understanding (NLU) model learned on text-semantics corpora. However, this method requires the domains of speech-text and text-semantics to match, which often mismatch due to separate collections. Furthermore, using the entire collected speech-text corpus from any domains leads to \textit{imbalance} and \textit{noise} issues. To address these, we propose \textit{cross-modal selective self-training} (CMSST). CMSST tackles imbalance by clustering in a joint space of the three modalities (speech, text, and semantics) and handles label noise with a selection network. We also introduce two benchmarks for zero-shot E2E SLU, covering matched and found speech (mismatched) settings. Experiments show that CMSST improves performance in both two settings, with significantly reduced sample sizes and training time. Our code and data are released in https://github.com/amazon-science/zero-shot-E2E-slu.
title Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
topic Audio and Speech Processing
Computation and Language
Multimedia
Sound
url https://arxiv.org/abs/2305.12793