Dual Information Speech Language Models for Emotional Conversations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Chun, Liu, Chenyang, Xu, Wenze, Deng, Weihong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
von: Ognjen, et al.
Veröffentlicht: (2024)
von: Ognjen, et al.
Veröffentlicht: (2024)
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
von: Diwan, Anuj, et al.
Veröffentlicht: (2026)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech
von: Zhou, Siyi, et al.
Veröffentlicht: (2025)
von: Zhou, Siyi, et al.
Veröffentlicht: (2025)
Roadmap towards Superhuman Speech Understanding using Large Language Models
von: Bu, Fan, et al.
Veröffentlicht: (2024)
von: Bu, Fan, et al.
Veröffentlicht: (2024)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
von: Saengthong, Phurich, et al.
Veröffentlicht: (2025)
Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
von: Vu, Tai
Veröffentlicht: (2025)
von: Vu, Tai
Veröffentlicht: (2025)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
von: Zhou, Li, et al.
Veröffentlicht: (2026)
von: Zhou, Li, et al.
Veröffentlicht: (2026)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
von: Gong, Hongyu, et al.
Veröffentlicht: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models
von: Li, Bo, et al.
Veröffentlicht: (2025)
von: Li, Bo, et al.
Veröffentlicht: (2025)
Jointly Fine-Tuning "BERT-like" Self Supervised Models to Improve Multimodal Speech Emotion Recognition
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2020)
von: Siriwardhana, Shamane, et al.
Veröffentlicht: (2020)
SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
von: Lin, Yueqian, et al.
Veröffentlicht: (2024)
Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024)
von: Zhang, Qinglin, et al.
Veröffentlicht: (2024)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
von: Liu, Rui, et al.
Veröffentlicht: (2025)
von: Liu, Rui, et al.
Veröffentlicht: (2025)
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
von: Yoon, Jinsung, et al.
Veröffentlicht: (2025)
von: Yoon, Jinsung, et al.
Veröffentlicht: (2025)
Investigating Safety Vulnerabilities of Large Audio-Language Models Under Speaker Emotional Variations
von: Feng, Bo-Han, et al.
Veröffentlicht: (2025)
von: Feng, Bo-Han, et al.
Veröffentlicht: (2025)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
Bimodal Connection Attention Fusion for Speech Emotion Recognition
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
von: Luo, Jiachen, et al.
Veröffentlicht: (2025)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
Toward Conversational Hungarian Speech Recognition: Introducing the BEA-Large and BEA-Dialogue Datasets
von: Gedeon, Máté, et al.
Veröffentlicht: (2025)
von: Gedeon, Máté, et al.
Veröffentlicht: (2025)
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
Exploring Self-Supervised Multi-view Contrastive Learning for Speech Emotion Recognition with Limited Annotations
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
von: Khaertdinov, Bulat, et al.
Veröffentlicht: (2024)
The Sound of Syntax: Finetuning and Comprehensive Evaluation of Language Models for Speech Pathology
von: Patel, Fagun, et al.
Veröffentlicht: (2025)
von: Patel, Fagun, et al.
Veröffentlicht: (2025)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
Towards Quantifying and Reducing Language Mismatch Effects in Cross-Lingual Speech Anti-Spoofing
von: Liu, Tianchi, et al.
Veröffentlicht: (2024)
von: Liu, Tianchi, et al.
Veröffentlicht: (2024)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2024)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
Get Large Language Models Ready to Speak: A Late-fusion Approach for Speech Generation
von: Shen, Maohao, et al.
Veröffentlicht: (2024)
von: Shen, Maohao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023) -
Device-Directed Speech Detection for Follow-up Conversations Using Large Language Models
von: Ognjen, et al.
Veröffentlicht: (2024) -
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining
von: Diwan, Anuj, et al.
Veröffentlicht: (2026) -
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025) -
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
von: Le, Chenyang, et al.
Veröffentlicht: (2024)