Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Yujia, Zhao, Zeyu, Richmond, Korin, Li, Yuanchao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
von: Han, Zhichen, et al.
Veröffentlicht: (2024)
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
Mega-ASR: Towards In-the-wild^2 Speech Recognition via Scaling up Real-world Acoustic Simulation
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
von: Ji, Shulei, et al.
Veröffentlicht: (2023)
von: Ji, Shulei, et al.
Veröffentlicht: (2023)
DeepResonance: Enhancing Multimodal Music Understanding via Music-centric Multi-way Instruction Tuning
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition
von: Pan, Yu, et al.
Veröffentlicht: (2023)
von: Pan, Yu, et al.
Veröffentlicht: (2023)
GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
von: Zuo, Heda, et al.
Veröffentlicht: (2025)
It's Never Too Late: Fusing Acoustic Information into Large Language Models for Automatic Speech Recognition
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
von: Ma, Ziyang, et al.
Veröffentlicht: (2024)
Cross-Modal Learning for Music-to-Music-Video Description Generation
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Mao, Zhuoyuan, et al.
Veröffentlicht: (2025)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
von: Qian, Zhiwen, et al.
Veröffentlicht: (2025)
MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
von: Zhang, Yixiao, et al.
Veröffentlicht: (2024)
Sync-TVA: A Graph-Attention Framework for Multimodal Emotion Recognition with Cross-Modal Fusion
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
von: Deng, Zeyu, et al.
Veröffentlicht: (2025)
OpenMU: Your Swiss Army Knife for Music Understanding
von: Zhao, Mengjie, et al.
Veröffentlicht: (2024)
von: Zhao, Mengjie, et al.
Veröffentlicht: (2024)
Exploring Classical Piano Performance Generation with Expressive Music Variational AutoEncoder
von: Luo, Jing, et al.
Veröffentlicht: (2025)
von: Luo, Jing, et al.
Veröffentlicht: (2025)
MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
von: Deng, Zihao, et al.
Veröffentlicht: (2023)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
von: Min, Anna, et al.
Veröffentlicht: (2025)
von: Min, Anna, et al.
Veröffentlicht: (2025)
Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
von: Wu, Junxian, et al.
Veröffentlicht: (2025)
von: Wu, Junxian, et al.
Veröffentlicht: (2025)
Frechet Music Distance: A Metric For Generative Symbolic Music Evaluation
von: Retkowski, Jan, et al.
Veröffentlicht: (2024)
von: Retkowski, Jan, et al.
Veröffentlicht: (2024)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
Zero-Effort Image-to-Music Generation: An Interpretable RAG-based VLM Approach
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
von: Choi, Suhwan, et al.
Veröffentlicht: (2025)
A Survey of Foundation Models for Music Understanding
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
von: Li, Wenjun, et al.
Veröffentlicht: (2024)
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
von: Shahin, Mostafa, et al.
Veröffentlicht: (2025)
von: Shahin, Mostafa, et al.
Veröffentlicht: (2025)
Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
von: Ye, Zhen, et al.
Veröffentlicht: (2025)
Whispering LLaMA: A Cross-Modal Generative Error Correction Framework for Speech Recognition
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
von: Radhakrishnan, Srijith, et al.
Veröffentlicht: (2023)
CoComposer: LLM Multi-agent Collaborative Music Composition
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
From Sound to Sight: Towards AI-authored Music Videos
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
von: Vitasovic, Leo, et al.
Veröffentlicht: (2025)
Deciphering GunType Hierarchy through Acoustic Analysis of Gunshot Recordings
von: Shah, Ankit, et al.
Veröffentlicht: (2025)
von: Shah, Ankit, et al.
Veröffentlicht: (2025)
Exploring Adapter Design Tradeoffs for Low Resource Music Generation
von: Mehta, Atharva, et al.
Veröffentlicht: (2025)
von: Mehta, Atharva, et al.
Veröffentlicht: (2025)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Semi-Supervised Cognitive State Classification from Speech with Multi-View Pseudo-Labeling
von: Li, Yuanchao, et al.
Veröffentlicht: (2024) -
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
von: Han, Zhichen, et al.
Veröffentlicht: (2024) -
Revise, Reason, and Recognize: LLM-Based Emotion Recognition via Emotion-Specific Prompts and ASR Error Correction
von: Li, Yuanchao, et al.
Veröffentlicht: (2024) -
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
von: Li, Yuanchao, et al.
Veröffentlicht: (2024) -
Speech Emotion Recognition with ASR Transcripts: A Comprehensive Study on Word Error Rate and Fusion Techniques
von: Li, Yuanchao, et al.
Veröffentlicht: (2024)