ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Lou, Haowei, Paik, Hye-young, Hu, Wen, Yao, Lina
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911383769055232
author Lou, Haowei
Paik, Hye-young
Hu, Wen
Yao, Lina
author_facet Lou, Haowei
Paik, Hye-young
Hu, Wen
Yao, Lina
contents Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each type of style. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and language classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking styles while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12289
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
Lou, Haowei
Paik, Hye-young
Hu, Wen
Yao, Lina
Sound
Machine Learning
Audio and Speech Processing
Learning representative embeddings for different types of speaking styles, such as emotion, age, and gender, is critical for both recognition tasks (e.g., cognitive computing and human-computer interaction) and generative tasks (e.g., style-controllable speech generation). In this work, we introduce ParaMETA, a unified and flexible framework for learning and controlling speaking styles directly from speech. Unlike existing methods that rely on single-task models or cross-modal alignment, ParaMETA learns disentangled, task-specific embeddings by projecting speech into dedicated subspaces for each type of style. This design reduces inter-task interference, mitigates negative transfer, and allows a single model to handle multiple paralinguistic tasks such as emotion, gender, age, and language classification. Beyond recognition, ParaMETA enables fine-grained style control in Text-To-Speech (TTS) generative models. It supports both speech- and text-based prompting and allows users to modify one speaking styles while preserving others. Extensive experiments demonstrate that ParaMETA outperforms strong baselines in classification accuracy and generates more natural and expressive speech, while maintaining a lightweight and efficient model suitable for real-world applications.
title ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2601.12289