B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917252560846848 |
|---|---|
| author | Gao, Yingying Zhang, Shilei Yang, Runyan Cui, Zihao Feng, Junlan |
| author_facet | Gao, Yingying Zhang, Shilei Yang, Runyan Cui, Zihao Feng, Junlan |
| contents | Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising method which enhances the performance through rule-based or model-based verification functions rather than human annotations. We treat the sample selection during the learning process as a long-term procedure and whether to select a sample as the action to make policy, thus achieving the application of RL to measure sample quality in SER. We propose a modified Group Relative Policy Optimization (GRPO) to adapt it to classification problems, which takes the samples in a batch as a group and uses the average reward of these samples as the baseline to calculate the advantage. And rather than using a verifiable reward function as in GRPO, we put forward self-reward functions and teacher-reward functions to encourage the model to produce high-confidence outputs. Experiments indicate that the proposed method improves the performance of baseline without RL by 19.8%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_06290 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization Gao, Yingying Zhang, Shilei Yang, Runyan Cui, Zihao Feng, Junlan Audio and Speech Processing Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising method which enhances the performance through rule-based or model-based verification functions rather than human annotations. We treat the sample selection during the learning process as a long-term procedure and whether to select a sample as the action to make policy, thus achieving the application of RL to measure sample quality in SER. We propose a modified Group Relative Policy Optimization (GRPO) to adapt it to classification problems, which takes the samples in a batch as a group and uses the average reward of these samples as the baseline to calculate the advantage. And rather than using a verifiable reward function as in GRPO, we put forward self-reward functions and teacher-reward functions to encourage the model to produce high-confidence outputs. Experiments indicate that the proposed method improves the performance of baseline without RL by 19.8%. |
| title | B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2602.06290 |