Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
Fuente:
arXiv
Salvato in:
| Autori principali: | Tsiamas, Ioannis, Sperber, Matthias, Finch, Andrew, Garg, Sarthak |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Benchmarking Prosody Encoding in Discrete Speech Tokens
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
di: He, Xiangheng, et al.
Pubblicazione: (2024)
di: He, Xiangheng, et al.
Pubblicazione: (2024)
Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech
di: Borodin, Kirill, et al.
Pubblicazione: (2025)
di: Borodin, Kirill, et al.
Pubblicazione: (2025)
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
di: Cui, Wenqian, et al.
Pubblicazione: (2026)
End-to-End Speech-to-Text Translation: A Survey
di: Sethiya, Nivedita, et al.
Pubblicazione: (2023)
di: Sethiya, Nivedita, et al.
Pubblicazione: (2023)
High-Fidelity Simultaneous Speech-To-Speech Translation
di: Labiausse, Tom, et al.
Pubblicazione: (2025)
di: Labiausse, Tom, et al.
Pubblicazione: (2025)
Direct Speech to Speech Translation: A Review
di: Sarim, Mohammad, et al.
Pubblicazione: (2025)
di: Sarim, Mohammad, et al.
Pubblicazione: (2025)
Continuous Speech Tokenizer in Text To Speech
di: Li, Yixing, et al.
Pubblicazione: (2024)
di: Li, Yixing, et al.
Pubblicazione: (2024)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
di: Moslem, Yasmin
Pubblicazione: (2024)
di: Moslem, Yasmin
Pubblicazione: (2024)
Simultaneous Speech-to-Speech Translation Without Aligned Data
di: Labiausse, Tom, et al.
Pubblicazione: (2026)
di: Labiausse, Tom, et al.
Pubblicazione: (2026)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
di: Park, Chanho, et al.
Pubblicazione: (2023)
di: Park, Chanho, et al.
Pubblicazione: (2023)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
di: Deng, Keqi, et al.
Pubblicazione: (2025)
di: Deng, Keqi, et al.
Pubblicazione: (2025)
NAIST Simultaneous Speech Translation System for IWSLT 2024
di: Ko, Yuka, et al.
Pubblicazione: (2024)
di: Ko, Yuka, et al.
Pubblicazione: (2024)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
di: Gupta, Mahendra, et al.
Pubblicazione: (2024)
di: Gupta, Mahendra, et al.
Pubblicazione: (2024)
Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
di: Akarsh, Sai, et al.
Pubblicazione: (2024)
di: Akarsh, Sai, et al.
Pubblicazione: (2024)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
di: Zhang, Yuhao, et al.
Pubblicazione: (2025)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
di: Li, Yingting, et al.
Pubblicazione: (2024)
di: Li, Yingting, et al.
Pubblicazione: (2024)
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
Compact Speech Translation Models via Discrete Speech Units Pretraining
di: Lam, Tsz Kin, et al.
Pubblicazione: (2024)
di: Lam, Tsz Kin, et al.
Pubblicazione: (2024)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
di: Vesterbacka, Leonora, et al.
Pubblicazione: (2025)
di: Vesterbacka, Leonora, et al.
Pubblicazione: (2025)
MunTTS: A Text-to-Speech System for Mundari
di: Gumma, Varun, et al.
Pubblicazione: (2024)
di: Gumma, Varun, et al.
Pubblicazione: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
di: Wotherspoon, Shannon, et al.
Pubblicazione: (2024)
di: Wotherspoon, Shannon, et al.
Pubblicazione: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
Spatial Speech Translation: Translating Across Space With Binaural Hearables
di: Chen, Tuochao, et al.
Pubblicazione: (2025)
di: Chen, Tuochao, et al.
Pubblicazione: (2025)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
di: Wang, Jianjin, et al.
Pubblicazione: (2025)
di: Wang, Jianjin, et al.
Pubblicazione: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
di: Huang, Wuwei, et al.
Pubblicazione: (2025)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2026)
di: Satish, Shree Harsha Bokkahalli, et al.
Pubblicazione: (2026)
Representation Purification for End-to-End Speech Translation
di: Zhang, Chengwei, et al.
Pubblicazione: (2024)
di: Zhang, Chengwei, et al.
Pubblicazione: (2024)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Benchmarking Prosody Encoding in Discrete Speech Tokens
di: Onda, Kentaro, et al.
Pubblicazione: (2025) -
DiffProsody: Diffusion-based Latent Prosody Generation for Expressive Speech Synthesis with Prosody Conditional Adversarial Training
di: Oh, Hyung-Seok, et al.
Pubblicazione: (2023) -
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025) -
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
di: Liu, Rui, et al.
Pubblicazione: (2024) -
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)