RCT: Random Consistency Training for Semi-supervised Sound Event Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2021
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913180817555456 |
|---|---|
| author | Shao, Nian Loweimi, Erfan Li, Xiaofei |
| author_facet | Shao, Nian Loweimi, Erfan Li, Xiaofei |
| contents | Sound event detection (SED), as a core module of acoustic environmental analysis, suffers from the problem of data deficiency. The integration of semi-supervised learning (SSL) largely mitigates such problem while bringing no extra annotation budget. This paper researches on several core modules of SSL, and introduces a random consistency training (RCT) strategy. First, a self-consistency loss is proposed to fuse with the teacher-student model to stabilize the training. Second, a hard mixup data augmentation is proposed to account for the additive property of sounds. Third, a random augmentation scheme is applied to flexibly combine different types of data augmentations. Experiments show that the proposed strategy outperform other widely-used strategies. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2110_11144 |
| institution | arXiv |
| publishDate | 2021 |
| record_format | arxiv |
| spellingShingle | RCT: Random Consistency Training for Semi-supervised Sound Event Detection Shao, Nian Loweimi, Erfan Li, Xiaofei Audio and Speech Processing Machine Learning Sound Sound event detection (SED), as a core module of acoustic environmental analysis, suffers from the problem of data deficiency. The integration of semi-supervised learning (SSL) largely mitigates such problem while bringing no extra annotation budget. This paper researches on several core modules of SSL, and introduces a random consistency training (RCT) strategy. First, a self-consistency loss is proposed to fuse with the teacher-student model to stabilize the training. Second, a hard mixup data augmentation is proposed to account for the additive property of sounds. Third, a random augmentation scheme is applied to flexibly combine different types of data augmentations. Experiments show that the proposed strategy outperform other widely-used strategies. |
| title | RCT: Random Consistency Training for Semi-supervised Sound Event Detection |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2110.11144 |