XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zuo, Tianlun, Hu, Jingbin, Li, Yuke, Zhu, Xinfa, Li, Hai, Yan, Ying, Liu, Junhui, Xie, Danming, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
von: Guo, Dake, et al.
Veröffentlicht: (2024)
von: Guo, Dake, et al.
Veröffentlicht: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
Towards Attribution of Generators and Emotional Manipulation in Cross-Lingual Synthetic Speech using Geometric Learning
von: Girish, et al.
Veröffentlicht: (2025)
von: Girish, et al.
Veröffentlicht: (2025)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
von: Li, Hanzhao, et al.
Veröffentlicht: (2024)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
UrduSpeech: A 156-Hour Urdu Speech Corpus with 12-Dimension Paralinguistic Annotations
von: Haq, Attia Nafees ul, et al.
Veröffentlicht: (2026)
von: Haq, Attia Nafees ul, et al.
Veröffentlicht: (2026)
Enhancing Non-Core Language Instruction-Following in Speech LLMs via Semi-Implicit Cross-Lingual CoT Reasoning
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
von: Xue, Hongfei, et al.
Veröffentlicht: (2025)
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
von: Hu, Jingbin, et al.
Veröffentlicht: (2026)
von: Hu, Jingbin, et al.
Veröffentlicht: (2026)
Quantifying Cross-Lingual Transfer in Paralinguistic Speech Tasks
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
von: Buitrago, Pol, et al.
Veröffentlicht: (2026)
Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
von: Li, Guojian, et al.
Veröffentlicht: (2026)
von: Li, Guojian, et al.
Veröffentlicht: (2026)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
KALL-E:Autoregressive Speech Synthesis with Next-Distribution Prediction
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2024)
RefXVC: Cross-Lingual Voice Conversion with Enhanced Reference Leveraging
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingyang, et al.
Veröffentlicht: (2024)
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
von: Wang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Wang, Pengcheng, et al.
Veröffentlicht: (2025)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
von: Yang, Haoyuan, et al.
Veröffentlicht: (2026)
von: Yang, Haoyuan, et al.
Veröffentlicht: (2026)
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
Coarse-to-fine Alignment Makes Better Speech-image Retrieval
von: Zhou, Lifeng, et al.
Veröffentlicht: (2024)
von: Zhou, Lifeng, et al.
Veröffentlicht: (2024)
SpeechT-RAG: Reliable Depression Detection in LLMs with Retrieval-Augmented Generation Using Speech Timing Information
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2025)
LA-RAG:Enhancing LLM-based ASR Accuracy with Retrieval-Augmented Generation
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval
von: Zhou, Lifeng, et al.
Veröffentlicht: (2024)
von: Zhou, Lifeng, et al.
Veröffentlicht: (2024)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Weakly Supervised Data Refinement and Flexible Sequence Compression for Efficient Thai LLM-based ASR
von: Shao, Mingchen, et al.
Veröffentlicht: (2025) -
FreeV: Free Lunch For Vocoders Through Pseudo Inversed Mel Filter
von: Lv, Yuanjun, et al.
Veröffentlicht: (2024) -
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025) -
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023) -
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)