Computational Narrative Understanding for Expressive Text-to-Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Michel, Gaspard, Epure, Elena V., Cerisara, Christophe |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speech Language Models for Under-Represented Languages: Insights from Wolof
di: Sy, Yaya, et al.
Pubblicazione: (2025)
di: Sy, Yaya, et al.
Pubblicazione: (2025)
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
di: Liu, Sen, et al.
Pubblicazione: (2024)
di: Liu, Sen, et al.
Pubblicazione: (2024)
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
di: Sy, Yaya, et al.
Pubblicazione: (2025)
di: Sy, Yaya, et al.
Pubblicazione: (2025)
Generative Expressive Conversational Speech Synthesis
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
di: Hwang, Min-Jae, et al.
Pubblicazione: (2024)
di: Hwang, Min-Jae, et al.
Pubblicazione: (2024)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
di: Kang, Wonjune, et al.
Pubblicazione: (2025)
di: Kang, Wonjune, et al.
Pubblicazione: (2025)
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
Continuous Speech Tokenizer in Text To Speech
di: Li, Yixing, et al.
Pubblicazione: (2024)
di: Li, Yixing, et al.
Pubblicazione: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
di: Mousavi, Pooneh, et al.
Pubblicazione: (2025)
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM
di: Song, Yaodong, et al.
Pubblicazione: (2025)
di: Song, Yaodong, et al.
Pubblicazione: (2025)
Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion
di: Frohmann, Markus, et al.
Pubblicazione: (2025)
di: Frohmann, Markus, et al.
Pubblicazione: (2025)
SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
di: Gong, Hongyu, et al.
Pubblicazione: (2024)
di: Gong, Hongyu, et al.
Pubblicazione: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
di: Hu, Jiliang, et al.
Pubblicazione: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
Beyond Musical Descriptors: Extracting Preference-Bearing Intent in Music Queries
di: Baranes, Marion, et al.
Pubblicazione: (2026)
di: Baranes, Marion, et al.
Pubblicazione: (2026)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
di: Tsiamas, Ioannis, et al.
Pubblicazione: (2024)
di: Tsiamas, Ioannis, et al.
Pubblicazione: (2024)
Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison
di: Lam, Tsz Kin, et al.
Pubblicazione: (2025)
di: Lam, Tsz Kin, et al.
Pubblicazione: (2025)
Scaling Analysis of Interleaved Speech-Text Language Models
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
End-to-End Speech-to-Text Translation: A Survey
di: Sethiya, Nivedita, et al.
Pubblicazione: (2023)
di: Sethiya, Nivedita, et al.
Pubblicazione: (2023)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
MunTTS: A Text-to-Speech System for Mundari
di: Gumma, Varun, et al.
Pubblicazione: (2024)
di: Gumma, Varun, et al.
Pubblicazione: (2024)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
di: Fucci, Dennis, et al.
Pubblicazione: (2024)
di: Fucci, Dennis, et al.
Pubblicazione: (2024)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
di: Bott, Thomas, et al.
Pubblicazione: (2024)
di: Bott, Thomas, et al.
Pubblicazione: (2024)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
di: Arora, Siddhant, et al.
Pubblicazione: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
di: Xie, Jingran, et al.
Pubblicazione: (2025)
di: Xie, Jingran, et al.
Pubblicazione: (2025)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
di: Tian, Jinchuan, et al.
Pubblicazione: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
di: Zeng, Aohan, et al.
Pubblicazione: (2024)
Communication-Efficient Personalized Federated Learning for Speech-to-Text Tasks
di: Du, Yichao, et al.
Pubblicazione: (2024)
di: Du, Yichao, et al.
Pubblicazione: (2024)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
di: Lou, Haowei, et al.
Pubblicazione: (2025)
di: Lou, Haowei, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Speech Language Models for Under-Represented Languages: Insights from Wolof
di: Sy, Yaya, et al.
Pubblicazione: (2025) -
StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations
di: Liu, Sen, et al.
Pubblicazione: (2024) -
BaldWhisper: Faster Whisper with Head Shearing and Layer Merging
di: Sy, Yaya, et al.
Pubblicazione: (2025) -
Generative Expressive Conversational Speech Synthesis
di: Liu, Rui, et al.
Pubblicazione: (2024) -
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)