FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Rui, Xi, Jiatian, Jiang, Ziyue, Li, Haizhou |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
por: Liu, Rui, et al.
Publicado: (2025)
por: Liu, Rui, et al.
Publicado: (2025)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
por: Oh, Hyung-Seok, et al.
Publicado: (2023)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
por: Eren, Eray, et al.
Publicado: (2025)
por: Eren, Eray, et al.
Publicado: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
por: Zhou, Xuehao, et al.
Publicado: (2024)
por: Zhou, Xuehao, et al.
Publicado: (2024)
Benchmarking Prosody Encoding in Discrete Speech Tokens
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
por: Inoue, Sho, et al.
Publicado: (2025)
por: Inoue, Sho, et al.
Publicado: (2025)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
por: Ren, Yong, et al.
Publicado: (2026)
por: Ren, Yong, et al.
Publicado: (2026)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
por: Liu, Jiaxuan, et al.
Publicado: (2024)
por: Liu, Jiaxuan, et al.
Publicado: (2024)
Generative Expressive Conversational Speech Synthesis
por: Liu, Rui, et al.
Publicado: (2024)
por: Liu, Rui, et al.
Publicado: (2024)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
por: Lei, Shun, et al.
Publicado: (2023)
por: Lei, Shun, et al.
Publicado: (2023)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
por: Cui, Wenqian, et al.
Publicado: (2026)
por: Cui, Wenqian, et al.
Publicado: (2026)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
por: Borodin, Kirill, et al.
Publicado: (2025)
por: Borodin, Kirill, et al.
Publicado: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
por: He, Xiangheng, et al.
Publicado: (2024)
por: He, Xiangheng, et al.
Publicado: (2024)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
por: Ao, Junyi, et al.
Publicado: (2025)
por: Ao, Junyi, et al.
Publicado: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
por: Jiang, Feng, et al.
Publicado: (2025)
por: Jiang, Feng, et al.
Publicado: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
por: Jiang, Yuepeng, et al.
Publicado: (2024)
por: Jiang, Yuepeng, et al.
Publicado: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
por: Zeng, Aohan, et al.
Publicado: (2024)
por: Zeng, Aohan, et al.
Publicado: (2024)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
por: Koriyama, Tomoki
Publicado: (2025)
por: Koriyama, Tomoki
Publicado: (2025)
Speech Editing -- a Summary
por: Kässmann, Tobias, et al.
Publicado: (2024)
por: Kässmann, Tobias, et al.
Publicado: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2025)
por: Inoue, Sho, et al.
Publicado: (2025)
Scaling Analysis of Interleaved Speech-Text Language Models
por: Maimon, Gallil, et al.
Publicado: (2025)
por: Maimon, Gallil, et al.
Publicado: (2025)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
por: Wang, Chunhui, et al.
Publicado: (2024)
por: Wang, Chunhui, et al.
Publicado: (2024)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
por: Manakul, Potsawee, et al.
Publicado: (2026)
por: Manakul, Potsawee, et al.
Publicado: (2026)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
por: Liu, Henglyu, et al.
Publicado: (2025)
por: Liu, Henglyu, et al.
Publicado: (2025)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
Sequential Editing for Lifelong Training of Speech Recognition Models
por: Kulshreshtha, Devang, et al.
Publicado: (2024)
por: Kulshreshtha, Devang, et al.
Publicado: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
por: Karapiperis, Sotirios, et al.
Publicado: (2024)
por: Karapiperis, Sotirios, et al.
Publicado: (2024)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
por: Li, Xiang, et al.
Publicado: (2024)
por: Li, Xiang, et al.
Publicado: (2024)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
por: Chen, Yushen, et al.
Publicado: (2024)
por: Chen, Yushen, et al.
Publicado: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
por: Hwang, Min-Jae, et al.
Publicado: (2024)
por: Hwang, Min-Jae, et al.
Publicado: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
por: Xie, Jingran, et al.
Publicado: (2025)
por: Xie, Jingran, et al.
Publicado: (2025)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
por: Yang, Dongchao, et al.
Publicado: (2024)
por: Yang, Dongchao, et al.
Publicado: (2024)
Multilingual Prosody Transfer: Comparing Supervised & Transfer Learning
por: Goel, Arnav, et al.
Publicado: (2024)
por: Goel, Arnav, et al.
Publicado: (2024)
Phone-Level Prosody Modelling with GMM-Based MDN for Diverse and Controllable Speech Synthesis
por: Du, Chenpeng, et al.
Publicado: (2021)
por: Du, Chenpeng, et al.
Publicado: (2021)
Ejemplares similares
-
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
por: Liu, Rui, et al.
Publicado: (2025) -
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
por: Liu, Rui, et al.
Publicado: (2024) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
por: Oh, Hyung-Seok, et al.
Publicado: (2023) -
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
por: Chen, Yang, et al.
Publicado: (2024) -
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)