Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Rui, Ma, Zening |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
von: Li, Hengyu, et al.
Veröffentlicht: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
von: Chung, Raymond
Veröffentlicht: (2026)
von: Chung, Raymond
Veröffentlicht: (2026)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
A Large-Scale Probing Analysis of Speaker-Specific Attributes in Self-Supervised Speech Representations
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
von: Chiu, Aemon Yat Fei, et al.
Veröffentlicht: (2025)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2025)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
Facial Expression-Enhanced TTS: Combining Face Representation and Emotion Intensity for Adaptive Speech
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
von: Chu, Yunji, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
Analysis of Self-Supervised Speech Models on Children's Speech and Infant Vocalizations
von: Li, Jialu, et al.
Veröffentlicht: (2024)
von: Li, Jialu, et al.
Veröffentlicht: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
von: Wilkins, Julia, et al.
Veröffentlicht: (2025)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Progressive Residual Extraction based Pre-training for Speech Representation Learning
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
von: Wang, Tianrui, et al.
Veröffentlicht: (2024)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
von: Ueda, Lucas H., et al.
Veröffentlicht: (2026)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
von: Chen, Yafeng, et al.
Veröffentlicht: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
von: Chen, Yafeng, et al.
Veröffentlicht: (2023)
von: Chen, Yafeng, et al.
Veröffentlicht: (2023)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
Audio-Visual Representation Learning via Knowledge Distillation from Speech Foundation Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Rate-Aware Learned Speech Compression
von: Xu, Jun, et al.
Veröffentlicht: (2025)
von: Xu, Jun, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Models for Automatic Whispered Speech Recognition
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2024)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
GigaAM: Efficient Self-Supervised Learner for Speech Recognition
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Kutsakov, Aleksandr, et al.
Veröffentlicht: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
von: Brima, Yusuf, et al.
Veröffentlicht: (2023)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
von: Hyeon, Jonghwan, et al.
Veröffentlicht: (2024)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
von: Wang, Shih-heng, et al.
Veröffentlicht: (2024)
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
von: Poncelet, Jakob, et al.
Veröffentlicht: (2021)
von: Poncelet, Jakob, et al.
Veröffentlicht: (2021)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
Noise-Aware Speech Separation with Contrastive Learning
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
von: Zhang, Zizheng, et al.
Veröffentlicht: (2023)
AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text
von: Wu, Jingyao, et al.
Veröffentlicht: (2026)
von: Wu, Jingyao, et al.
Veröffentlicht: (2026)
Self-Supervised Speech Quality Assessment (S3QA): Leveraging Speech Foundation Models for a Scalable Speech Quality Metric
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
von: Ogg, Mattson, et al.
Veröffentlicht: (2025)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Refining Self-Supervised Learnt Speech Representation using Brain Activations
von: Li, Hengyu, et al.
Veröffentlicht: (2024) -
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
von: Yamashita, Natsuo, et al.
Veröffentlicht: (2024) -
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024) -
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
von: Chung, Raymond
Veröffentlicht: (2026) -
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)