Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Rui, Gao, Pu, Xi, Jiatian, Sisman, Berrak, Busso, Carlos, Li, Haizhou |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
por: Ulgen, Ismail Rasim, et al.
Publicado: (2024)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2024)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
por: Rajapakshe, Thejan, et al.
Publicado: (2024)
por: Rajapakshe, Thejan, et al.
Publicado: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
por: Lam, Perry, et al.
Publicado: (2022)
por: Lam, Perry, et al.
Publicado: (2022)
Can Emotion Fool Anti-spoofing?
por: Mahapatra, Aurosweta, et al.
Publicado: (2025)
por: Mahapatra, Aurosweta, et al.
Publicado: (2025)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
por: Mahapatra, Aurosweta, et al.
Publicado: (2025)
por: Mahapatra, Aurosweta, et al.
Publicado: (2025)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
por: Melechovsky, Jan, et al.
Publicado: (2022)
por: Melechovsky, Jan, et al.
Publicado: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
por: Rajapakshe, Thejan, et al.
Publicado: (2023)
por: Rajapakshe, Thejan, et al.
Publicado: (2023)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
por: Du, Zongyang, et al.
Publicado: (2025)
por: Du, Zongyang, et al.
Publicado: (2025)
Fine-Grained Quantitative Emotion Editing for Speech Generation
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
por: Jawaid, Ahad, et al.
Publicado: (2024)
por: Jawaid, Ahad, et al.
Publicado: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2025)
por: Inoue, Sho, et al.
Publicado: (2025)
EmoFake: An Initial Dataset for Emotion Fake Audio Detection
por: Zhao, Yan, et al.
Publicado: (2022)
por: Zhao, Yan, et al.
Publicado: (2022)
Versatile audio-visual learning for emotion recognition
por: Goncalves, Lucas, et al.
Publicado: (2023)
por: Goncalves, Lucas, et al.
Publicado: (2023)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
por: Hasan, Rashedul, et al.
Publicado: (2025)
por: Hasan, Rashedul, et al.
Publicado: (2025)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
por: Du, Zongyang, et al.
Publicado: (2024)
por: Du, Zongyang, et al.
Publicado: (2024)
Universal Speech Content Factorization
por: Xinyuan, Henry Li, et al.
Publicado: (2026)
por: Xinyuan, Henry Li, et al.
Publicado: (2026)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
por: Ulgen, Ismail Rasim, et al.
Publicado: (2025)
por: Ulgen, Ismail Rasim, et al.
Publicado: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
por: Zhou, Li, et al.
Publicado: (2026)
por: Zhou, Li, et al.
Publicado: (2026)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
por: Yang, Yiqing, et al.
Publicado: (2025)
por: Yang, Yiqing, et al.
Publicado: (2025)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
por: Lee, Philip H., et al.
Publicado: (2024)
por: Lee, Philip H., et al.
Publicado: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
por: Cho, Deok-Hyeon, et al.
Publicado: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
por: Qi, Tianhua, et al.
Publicado: (2026)
por: Qi, Tianhua, et al.
Publicado: (2026)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
por: Bian, Weizhen, et al.
Publicado: (2024)
por: Bian, Weizhen, et al.
Publicado: (2024)
DiTSE: High-Fidelity Generative Speech Enhancement via Latent Diffusion Transformers
por: Guimarães, Heitor R., et al.
Publicado: (2025)
por: Guimarães, Heitor R., et al.
Publicado: (2025)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
por: Ren, Yong, et al.
Publicado: (2026)
por: Ren, Yong, et al.
Publicado: (2026)
EmoTale: An Enacted Speech-emotion Dataset in Danish
por: Hjuler, Maja J., et al.
Publicado: (2025)
por: Hjuler, Maja J., et al.
Publicado: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
por: Cong, Gaoxiang, et al.
Publicado: (2024)
por: Cong, Gaoxiang, et al.
Publicado: (2024)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
por: Xie, Tianxin, et al.
Publicado: (2025)
por: Xie, Tianxin, et al.
Publicado: (2025)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
por: Tian, Wenjie, et al.
Publicado: (2026)
por: Tian, Wenjie, et al.
Publicado: (2026)
Bridging Speech Emotion Recognition and Personality: Dataset and Temporal Interaction Condition Network
por: Gao, Yuan, et al.
Publicado: (2025)
por: Gao, Yuan, et al.
Publicado: (2025)
EmoTech: A Multi-modal Speech Emotion Recognition Using Multi-source Low-level Information with Hybrid Recurrent Network
por: Avro, Shamin Bin Habib, et al.
Publicado: (2025)
por: Avro, Shamin Bin Habib, et al.
Publicado: (2025)
Ejemplares similares
-
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
por: Liu, Rui, et al.
Publicado: (2024) -
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
por: Ulgen, Ismail Rasim, et al.
Publicado: (2024) -
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
por: Rajapakshe, Thejan, et al.
Publicado: (2024) -
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
por: Lam, Perry, et al.
Publicado: (2022) -
Can Emotion Fool Anti-spoofing?
por: Mahapatra, Aurosweta, et al.
Publicado: (2025)