Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Akarsh, Sai, Raghusimha, Vamshi, Mondal, Anindita, Vuppala, Anil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
von: Mondal, Anindita, et al.
Veröffentlicht: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Direct Speech to Speech Translation: A Review
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Simultaneous Speech-to-Speech Translation Without Aligned Data
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
von: Deng, Keqi, et al.
Veröffentlicht: (2025)
Compact Speech Translation Models via Discrete Speech Units Pretraining
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2024)
von: Lam, Tsz Kin, et al.
Veröffentlicht: (2024)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
von: Liu, Henglyu, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
Advancing Speech Translation: A Corpus of Mandarin-English Conversational Telephone Speech
von: Wotherspoon, Shannon, et al.
Veröffentlicht: (2024)
von: Wotherspoon, Shannon, et al.
Veröffentlicht: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
Spatial Speech Translation: Translating Across Space With Binaural Hearables
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
von: Wang, Jianjin, et al.
Veröffentlicht: (2025)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
Speech-to-Text Translation with Phoneme-Augmented CoT: Enhancing Cross-Lingual Transfer in Low-Resource Scenarios
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
von: Gállego, Gerard I., et al.
Veröffentlicht: (2025)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
NAIST Simultaneous Speech Translation System for IWSLT 2024
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
von: Ko, Yuka, et al.
Veröffentlicht: (2024)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
DiariST: Streaming Speech Translation with Speaker Diarization
von: Yang, Mu, et al.
Veröffentlicht: (2023)
von: Yang, Mu, et al.
Veröffentlicht: (2023)
Efficient Speech Translation through Model Compression and Knowledge Distillation
von: Moslem, Yasmin
Veröffentlicht: (2025)
von: Moslem, Yasmin
Veröffentlicht: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
A Few-Shot Approach to Dysarthric Speech Intelligibility Level Classification Using Transformers
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
von: Duret, Jarod, et al.
Veröffentlicht: (2024)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
von: Kim, Seung-Bin, et al.
Veröffentlicht: (2024)
MunTTS: A Text-to-Speech System for Mundari
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
von: Gumma, Varun, et al.
Veröffentlicht: (2024)
Investigating Decoder-only Large Language Models for Speech-to-text Translation
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
Bemba Speech Translation: Exploring a Low-Resource African Language
von: Farouq, Muhammad Hazim Al, et al.
Veröffentlicht: (2025)
von: Farouq, Muhammad Hazim Al, et al.
Veröffentlicht: (2025)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings
von: Mondal, Anindita, et al.
Veröffentlicht: (2024) -
Direct Speech-to-Speech Neural Machine Translation: A Survey
von: Gupta, Mahendra, et al.
Veröffentlicht: (2024) -
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025) -
Direct Speech to Speech Translation: A Review
von: Sarim, Mohammad, et al.
Veröffentlicht: (2025) -
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)