Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Wallbridge, Sarenne, Minixhofer, Christoph, Lai, Catherine, Bell, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
by: Minixhofer, Christoph, et al.
Published: (2024)
by: Minixhofer, Christoph, et al.
Published: (2024)
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
by: Minixhofer, Christoph, et al.
Published: (2025)
by: Minixhofer, Christoph, et al.
Published: (2025)
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025)
by: Storey, Edward, et al.
Published: (2025)
TTSDS -- Text-to-Speech Distribution Score
by: Minixhofer, Christoph, et al.
Published: (2024)
by: Minixhofer, Christoph, et al.
Published: (2024)
Linear-Complexity Self-Supervised Learning for Speech Processing
by: Zhang, Shucong, et al.
Published: (2024)
by: Zhang, Shucong, et al.
Published: (2024)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
NEST-RQ: Next Token Prediction for Speech Self-Supervised Pre-Training
by: Han, Minglun, et al.
Published: (2024)
by: Han, Minglun, et al.
Published: (2024)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
by: Li, Yuxin, et al.
Published: (2025)
by: Li, Yuxin, et al.
Published: (2025)
ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
by: Wang, Junyu, et al.
Published: (2025)
by: Wang, Junyu, et al.
Published: (2025)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
by: Rao, Rajath, et al.
Published: (2025)
by: Rao, Rajath, et al.
Published: (2025)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
by: Khaertdinov, Bulat, et al.
Published: (2024)
by: Khaertdinov, Bulat, et al.
Published: (2024)
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
by: Li, Gang, et al.
Published: (2025)
by: Li, Gang, et al.
Published: (2025)
Layer-Wise Analysis of Self-Supervised Acoustic Word Embeddings: A Study on Speech Emotion Recognition
by: Saliba, Alexandra, et al.
Published: (2024)
by: Saliba, Alexandra, et al.
Published: (2024)
Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts
by: Jin, Hojun, et al.
Published: (2025)
by: Jin, Hojun, et al.
Published: (2025)
How Should We Extract Discrete Audio Tokens from Self-Supervised Models?
by: Mousavi, Pooneh, et al.
Published: (2024)
by: Mousavi, Pooneh, et al.
Published: (2024)
AG-LSEC: Audio Grounded Lexical Speaker Error Correction
by: Paturi, Rohit, et al.
Published: (2024)
by: Paturi, Rohit, et al.
Published: (2024)
Explaining Spectrograms in Machine Learning: A Study on Neural Networks for Speech Classification
by: James, Jesin, et al.
Published: (2024)
by: James, Jesin, et al.
Published: (2024)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
by: Siriwardhana, Shamane, et al.
Published: (2020)
by: Siriwardhana, Shamane, et al.
Published: (2020)
AV-data2vec: Self-supervised Learning of Audio-Visual Speech Representations with Contextualized Target Representations
by: Lian, Jiachen, et al.
Published: (2023)
by: Lian, Jiachen, et al.
Published: (2023)
Crossmodal ASR Error Correction with Discrete Speech Units
by: Li, Yuanchao, et al.
Published: (2024)
by: Li, Yuanchao, et al.
Published: (2024)
TIPAA-SSL: Text Independent Phone-to-Audio Alignment based on Self-Supervised Learning and Knowledge Transfer
by: Tits, Noé, et al.
Published: (2024)
by: Tits, Noé, et al.
Published: (2024)
PSST! Prosodic Speech Segmentation with Transformers
by: Roll, Nathan, et al.
Published: (2023)
by: Roll, Nathan, et al.
Published: (2023)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
by: Huzaifah, Muhammad, et al.
Published: (2024)
by: Huzaifah, Muhammad, et al.
Published: (2024)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
Sounding Like a Winner? Prosodic Differences in Post-Match Interviews
by: Kakouros, Sofoklis, et al.
Published: (2025)
by: Kakouros, Sofoklis, et al.
Published: (2025)
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
by: Choi, Kwanghee, et al.
Published: (2025)
by: Choi, Kwanghee, et al.
Published: (2025)
Which Prosodic Features Matter Most for Pragmatics?
by: Ward, Nigel G., et al.
Published: (2024)
by: Ward, Nigel G., et al.
Published: (2024)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
by: Sun, Yujia, et al.
Published: (2024)
by: Sun, Yujia, et al.
Published: (2024)
An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems
by: Tulsiani, Hitesh, et al.
Published: (2024)
by: Tulsiani, Hitesh, et al.
Published: (2024)
Conversational Speech Reveals Structural Robustness Failures in SpeechLLM Backbones
by: Teleki, Maria, et al.
Published: (2025)
by: Teleki, Maria, et al.
Published: (2025)
Retrieval-Augmented Self-Taught Reasoning Model with Adaptive Chain-of-Thought for ASR Named Entity Correction
by: An, Junjie, et al.
Published: (2026)
by: An, Junjie, et al.
Published: (2026)
ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning
by: Lu, Yichen, et al.
Published: (2025)
by: Lu, Yichen, et al.
Published: (2025)
Phonology-Guided Speech-to-Speech Translation for African Languages
by: Ochieng, Peter, et al.
Published: (2024)
by: Ochieng, Peter, et al.
Published: (2024)
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
by: Zhu, Haina, et al.
Published: (2025)
by: Zhu, Haina, et al.
Published: (2025)
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study
by: Fan, Xiaoran, et al.
Published: (2025)
by: Fan, Xiaoran, et al.
Published: (2025)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
by: Sanders, Nicholas, et al.
Published: (2025)
by: Sanders, Nicholas, et al.
Published: (2025)
Controlling Surprisal in Music Generation via Information Content Curve Matching
by: Bjare, Mathias Rose, et al.
Published: (2024)
by: Bjare, Mathias Rose, et al.
Published: (2024)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
by: Ahasan, Md Mubtasim, et al.
Published: (2025)
by: Ahasan, Md Mubtasim, et al.
Published: (2025)
Similar Items
-
Beyond Oversmoothing: Evaluating DDPM and MSE for Scalable Speech Synthesis in ASR
by: Minixhofer, Christoph, et al.
Published: (2024) -
TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems
by: Minixhofer, Christoph, et al.
Published: (2025) -
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
by: Li, Yuanchao, et al.
Published: (2024) -
Language Bias in Self-Supervised Learning For Automatic Speech Recognition
by: Storey, Edward, et al.
Published: (2025) -
TTSDS -- Text-to-Speech Distribution Score
by: Minixhofer, Christoph, et al.
Published: (2024)