ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Aoduo, Lv, Haoran, Xu, Hongjian, Li, Shengmin, Qin, Sihao, Li, Zimeng, Pun, Chi Man, Chen, Xuhang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VEDAL: Variational Error-Driven Asynchronous Learning for 3D Gaussian Splatting Pruning
di: Li, Aoduo, et al.
Pubblicazione: (2026)
di: Li, Aoduo, et al.
Pubblicazione: (2026)
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
di: Huang, Lian, et al.
Pubblicazione: (2024)
di: Huang, Lian, et al.
Pubblicazione: (2024)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
di: Zhang, Shaozuo, et al.
Pubblicazione: (2025)
di: Zhang, Shaozuo, et al.
Pubblicazione: (2025)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Multi-Channel Speech Enhancement for Cocktail Party Speech Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2026)
di: Chen, Youjun, et al.
Pubblicazione: (2026)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
di: Huang, Dawei, et al.
Pubblicazione: (2026)
di: Huang, Dawei, et al.
Pubblicazione: (2026)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
di: Li, Bohan, et al.
Pubblicazione: (2025)
di: Li, Bohan, et al.
Pubblicazione: (2025)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
di: Wang, Dingdong, et al.
Pubblicazione: (2026)
di: Wang, Dingdong, et al.
Pubblicazione: (2026)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
di: Feng, Pengchao, et al.
Pubblicazione: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
di: Lv, Sihan, et al.
Pubblicazione: (2026)
di: Lv, Sihan, et al.
Pubblicazione: (2026)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
BridgeCode: A Dual Speech Representation Paradigm for Autoregressive Zero-Shot Text-to-Speech Synthesis
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2026)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2026)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
di: Zhou, Li, et al.
Pubblicazione: (2026)
di: Zhou, Li, et al.
Pubblicazione: (2026)
FS-RWKV: Leveraging Frequency Spatial-Aware RWKV for 3T-to-7T MRI Translation
di: Lei, Yingtie, et al.
Pubblicazione: (2025)
di: Lei, Yingtie, et al.
Pubblicazione: (2025)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
di: Wang, Cong, et al.
Pubblicazione: (2025)
di: Wang, Cong, et al.
Pubblicazione: (2025)
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
di: Meng, Zhaoyang, et al.
Pubblicazione: (2026)
di: Meng, Zhaoyang, et al.
Pubblicazione: (2026)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
di: Shi, Xiaohan, et al.
Pubblicazione: (2023)
di: Shi, Xiaohan, et al.
Pubblicazione: (2023)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
di: Liang, Qifan, et al.
Pubblicazione: (2026)
di: Liang, Qifan, et al.
Pubblicazione: (2026)
Persian Speech Emotion Recognition by Fine-Tuning Transformers
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
di: Shayaninasab, Minoo, et al.
Pubblicazione: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
Explainable Transformer-CNN Fusion for Noise-Robust Speech Emotion Recognition
di: Chakrabarty, Sudip, et al.
Pubblicazione: (2025)
di: Chakrabarty, Sudip, et al.
Pubblicazione: (2025)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
di: Gao, Yingxue, et al.
Pubblicazione: (2024)
di: Gao, Yingxue, et al.
Pubblicazione: (2024)
Emotion-Aware Quantization for Discrete Speech Representations: An Analysis of Emotion Preservation
di: Zhou, Haoguang, et al.
Pubblicazione: (2026)
di: Zhou, Haoguang, et al.
Pubblicazione: (2026)
SOLIDO: A Robust Watermarking Method for Speech Synthesis via Low-Rank Adaptation
di: Li, Yue, et al.
Pubblicazione: (2025)
di: Li, Yue, et al.
Pubblicazione: (2025)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
di: Derington, Anna, et al.
Pubblicazione: (2023)
di: Derington, Anna, et al.
Pubblicazione: (2023)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
di: Tang, Haobin, et al.
Pubblicazione: (2024)
di: Tang, Haobin, et al.
Pubblicazione: (2024)
MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
di: Sailor, Hardik B., et al.
Pubblicazione: (2025)
di: Sailor, Hardik B., et al.
Pubblicazione: (2025)
UniVocal: Unified Speech-Singing Code-Switching Synthesis
di: Shi, Yufei, et al.
Pubblicazione: (2026)
di: Shi, Yufei, et al.
Pubblicazione: (2026)
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
di: Shi, Jiacheng, et al.
Pubblicazione: (2026)
di: Shi, Jiacheng, et al.
Pubblicazione: (2026)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
di: Shi, Haoxiang, et al.
Pubblicazione: (2024)
di: Shi, Haoxiang, et al.
Pubblicazione: (2024)
Rare Word Recognition and Translation Without Fine-Tuning via Task Vector in Speech Models
di: Jing, Ruihao, et al.
Pubblicazione: (2025)
di: Jing, Ruihao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VEDAL: Variational Error-Driven Asynchronous Learning for 3D Gaussian Splatting Pruning
di: Li, Aoduo, et al.
Pubblicazione: (2026) -
Self-Attention and Hybrid Features for Replay and Deep-Fake Audio Detection
di: Huang, Lian, et al.
Pubblicazione: (2024) -
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
di: Li, Haoxun, et al.
Pubblicazione: (2025) -
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024) -
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
di: Zhang, Shaozuo, et al.
Pubblicazione: (2025)