B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gao, Yingying, Zhang, Shilei, Yang, Runyan, Cui, Zihao, Feng, Junlan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917252560846848
author Gao, Yingying
Zhang, Shilei
Yang, Runyan
Cui, Zihao
Feng, Junlan
author_facet Gao, Yingying
Zhang, Shilei
Yang, Runyan
Cui, Zihao
Feng, Junlan
contents Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising method which enhances the performance through rule-based or model-based verification functions rather than human annotations. We treat the sample selection during the learning process as a long-term procedure and whether to select a sample as the action to make policy, thus achieving the application of RL to measure sample quality in SER. We propose a modified Group Relative Policy Optimization (GRPO) to adapt it to classification problems, which takes the samples in a batch as a group and uses the average reward of these samples as the baseline to calculate the advantage. And rather than using a verifiable reward function as in GRPO, we put forward self-reward functions and teacher-reward functions to encourage the model to produce high-confidence outputs. Experiments indicate that the proposed method improves the performance of baseline without RL by 19.8%.
format Preprint
id arxiv_https___arxiv_org_abs_2602_06290
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
Gao, Yingying
Zhang, Shilei
Yang, Runyan
Cui, Zihao
Feng, Junlan
Audio and Speech Processing
Unsupervised speech emotion recognition (SER) focuses on addressing the problem of data sparsity and annotation bias of emotional speech. Reinforcement learning (RL) is a promising method which enhances the performance through rule-based or model-based verification functions rather than human annotations. We treat the sample selection during the learning process as a long-term procedure and whether to select a sample as the action to make policy, thus achieving the application of RL to measure sample quality in SER. We propose a modified Group Relative Policy Optimization (GRPO) to adapt it to classification problems, which takes the samples in a batch as a group and uses the average reward of these samples as the baseline to calculate the advantage. And rather than using a verifiable reward function as in GRPO, we put forward self-reward functions and teacher-reward functions to encourage the model to produce high-confidence outputs. Experiments indicate that the proposed method improves the performance of baseline without RL by 19.8%.
title B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
topic Audio and Speech Processing
url https://arxiv.org/abs/2602.06290