Dual Information Speech Language Models for Emotional Conversations
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Chun, Liu, Chenyang, Xu, Wenze, Deng, Weihong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
di: Ognjen, et al.
Pubblicazione: (2024)
di: Ognjen, et al.
Pubblicazione: (2024)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
di: Diwan, Anuj, et al.
Pubblicazione: (2026)
di: Diwan, Anuj, et al.
Pubblicazione: (2026)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
di: Le, Chenyang, et al.
Pubblicazione: (2024)
di: Le, Chenyang, et al.
Pubblicazione: (2024)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
di: Wang, Qichao, et al.
Pubblicazione: (2025)
di: Wang, Qichao, et al.
Pubblicazione: (2025)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
di: Yao, Wenhan, et al.
Pubblicazione: (2025)
di: Yao, Wenhan, et al.
Pubblicazione: (2025)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
di: Zhou, Siyi, et al.
Pubblicazione: (2025)
Roadmap towards Superhuman Speech Understanding using Large Language Models
di: Bu, Fan, et al.
Pubblicazione: (2024)
di: Bu, Fan, et al.
Pubblicazione: (2024)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
di: Vu, Tai
Pubblicazione: (2025)
di: Vu, Tai
Pubblicazione: (2025)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
di: Zhou, Li, et al.
Pubblicazione: (2026)
di: Zhou, Li, et al.
Pubblicazione: (2026)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
di: Gong, Hongyu, et al.
Pubblicazione: (2024)
di: Gong, Hongyu, et al.
Pubblicazione: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models
di: Li, Bo, et al.
Pubblicazione: (2025)
di: Li, Bo, et al.
Pubblicazione: (2025)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2020)
di: Siriwardhana, Shamane, et al.
Pubblicazione: (2020)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
di: Lin, Yueqian, et al.
Pubblicazione: (2024)
di: Lin, Yueqian, et al.
Pubblicazione: (2024)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
di: Zhang, Qinglin, et al.
Pubblicazione: (2024)
di: Zhang, Qinglin, et al.
Pubblicazione: (2024)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
di: Yoon, Jinsung, et al.
Pubblicazione: (2025)
di: Yoon, Jinsung, et al.
Pubblicazione: (2025)
Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
di: Feng, Bo-Han, et al.
Pubblicazione: (2025)
di: Feng, Bo-Han, et al.
Pubblicazione: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
di: Pan, Yu, et al.
Pubblicazione: (2025)
di: Pan, Yu, et al.
Pubblicazione: (2025)
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
di: Wang, Siyin, et al.
Pubblicazione: (2025)
di: Wang, Siyin, et al.
Pubblicazione: (2025)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
di: Luo, Jiachen, et al.
Pubblicazione: (2025)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
di: Gaido, Marco, et al.
Pubblicazione: (2024)
di: Gaido, Marco, et al.
Pubblicazione: (2024)
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
di: Gedeon, Máté, et al.
Pubblicazione: (2025)
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
di: Wang, Chen, et al.
Pubblicazione: (2025)
di: Wang, Chen, et al.
Pubblicazione: (2025)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
di: Khaertdinov, Bulat, et al.
Pubblicazione: (2024)
di: Khaertdinov, Bulat, et al.
Pubblicazione: (2024)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
di: Patel, Fagun, et al.
Pubblicazione: (2025)
di: Patel, Fagun, et al.
Pubblicazione: (2025)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
di: Yang, Xiaoran, et al.
Pubblicazione: (2024)
di: Yang, Xiaoran, et al.
Pubblicazione: (2024)
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
di: Liu, Tianchi, et al.
Pubblicazione: (2024)
di: Liu, Tianchi, et al.
Pubblicazione: (2024)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
di: Ghosh, Sreyan, et al.
Pubblicazione: (2026)
di: Ghosh, Sreyan, et al.
Pubblicazione: (2026)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2024)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
di: Shen, Maohao, et al.
Pubblicazione: (2024)
di: Shen, Maohao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023) -
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
di: Ognjen, et al.
Pubblicazione: (2024) -
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
di: Diwan, Anuj, et al.
Pubblicazione: (2026) -
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025) -
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
di: Le, Chenyang, et al.
Pubblicazione: (2024)