Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
Fuente:
arXiv
Guardado en:
| Autores principales: | Duret, Jarod, Estève, Yannick, Parcollet, Titouan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Early Prediction of Self-Supervised Speech Model Performance
por: Whetten, Ryan, et al.
Publicado: (2025)
por: Whetten, Ryan, et al.
Publicado: (2025)
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
por: Whetten, Ryan, et al.
Publicado: (2026)
por: Whetten, Ryan, et al.
Publicado: (2026)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
por: Whetten, Ryan, et al.
Publicado: (2024)
por: Whetten, Ryan, et al.
Publicado: (2024)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
por: Zhang, Yuhao, et al.
Publicado: (2025)
por: Zhang, Yuhao, et al.
Publicado: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
por: Choi, Jeongsoo, et al.
Publicado: (2025)
por: Choi, Jeongsoo, et al.
Publicado: (2025)
Simultaneous Speech-to-Speech Translation Without Aligned Data
por: Labiausse, Tom, et al.
Publicado: (2026)
por: Labiausse, Tom, et al.
Publicado: (2026)
MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
por: Duret, Jarod, et al.
Publicado: (2024)
por: Duret, Jarod, et al.
Publicado: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
por: Hwang, Min-Jae, et al.
Publicado: (2024)
por: Hwang, Min-Jae, et al.
Publicado: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
SummaryMixing: A Linear-Complexity Alternative to Self-Attention for Speech Recognition and Understanding
por: Parcollet, Titouan, et al.
Publicado: (2023)
por: Parcollet, Titouan, et al.
Publicado: (2023)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
por: Parcollet, Titouan, et al.
Publicado: (2025)
por: Parcollet, Titouan, et al.
Publicado: (2025)
SimulTron: On-Device Simultaneous Speech to Speech Translation
por: Agranovich, Alex, et al.
Publicado: (2024)
por: Agranovich, Alex, et al.
Publicado: (2024)
Translatotron 3: Speech to Speech Translation with Monolingual Data
por: Nachmani, Eliya, et al.
Publicado: (2023)
por: Nachmani, Eliya, et al.
Publicado: (2023)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
por: Zaiem, Salah, et al.
Publicado: (2023)
por: Zaiem, Salah, et al.
Publicado: (2023)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
por: Shankar, Bhavani, et al.
Publicado: (2024)
por: Shankar, Bhavani, et al.
Publicado: (2024)
Compact Speech Translation Models via Discrete Speech Units Pretraining
por: Lam, Tsz Kin, et al.
Publicado: (2024)
por: Lam, Tsz Kin, et al.
Publicado: (2024)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
por: Fang, Qingkai, et al.
Publicado: (2024)
por: Fang, Qingkai, et al.
Publicado: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
por: Rossenbach, Nick, et al.
Publicado: (2024)
por: Rossenbach, Nick, et al.
Publicado: (2024)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
por: Mdhaffar, Salima, et al.
Publicado: (2024)
por: Mdhaffar, Salima, et al.
Publicado: (2024)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
por: Zaiem, Salah, et al.
Publicado: (2024)
por: Zaiem, Salah, et al.
Publicado: (2024)
Textless Unit-to-Unit training for Many-to-Many Multilingual Speech-to-Speech Translation
por: Kim, Minsu, et al.
Publicado: (2023)
por: Kim, Minsu, et al.
Publicado: (2023)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
por: Lin, Hsi-Che, et al.
Publicado: (2024)
por: Lin, Hsi-Che, et al.
Publicado: (2024)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
por: Dutta, Soumya, et al.
Publicado: (2025)
por: Dutta, Soumya, et al.
Publicado: (2025)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
por: Nagpal, Chirag, et al.
Publicado: (2024)
por: Nagpal, Chirag, et al.
Publicado: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
por: Wang, Chien-Chun, et al.
Publicado: (2026)
por: Wang, Chien-Chun, et al.
Publicado: (2026)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
por: Futami, Hayato, et al.
Publicado: (2025)
por: Futami, Hayato, et al.
Publicado: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
por: Labiausse, Tom, et al.
Publicado: (2025)
por: Labiausse, Tom, et al.
Publicado: (2025)
Direct Speech to Speech Translation: A Review
por: Sarim, Mohammad, et al.
Publicado: (2025)
por: Sarim, Mohammad, et al.
Publicado: (2025)
Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems
por: Platnick, Daniel, et al.
Publicado: (2024)
por: Platnick, Daniel, et al.
Publicado: (2024)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
por: Wang, Xiaofei, et al.
Publicado: (2023)
por: Wang, Xiaofei, et al.
Publicado: (2023)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
por: Puvvada, Krishna C., et al.
Publicado: (2024)
por: Puvvada, Krishna C., et al.
Publicado: (2024)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
por: van Dalen, Rogier C., et al.
Publicado: (2025)
por: van Dalen, Rogier C., et al.
Publicado: (2025)
AlignAtt: Using Attention-based Audio-Translation Alignments as a Guide for Simultaneous Speech Translation
por: Papi, Sara, et al.
Publicado: (2023)
por: Papi, Sara, et al.
Publicado: (2023)
Textless Speech-to-Speech Translation With Limited Parallel Data
por: Diwan, Anuj, et al.
Publicado: (2023)
por: Diwan, Anuj, et al.
Publicado: (2023)
The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
por: Ljubešić, Nikola, et al.
Publicado: (2024)
por: Ljubešić, Nikola, et al.
Publicado: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
por: Deng, Keqi, et al.
Publicado: (2025)
por: Deng, Keqi, et al.
Publicado: (2025)
Modeling Overlapped Speech with Shuffles
por: Wiesner, Matthew, et al.
Publicado: (2026)
por: Wiesner, Matthew, et al.
Publicado: (2026)
Streaming Speech-to-Text Translation with a SpeechLLM
por: Parcollet, Titouan, et al.
Publicado: (2026)
por: Parcollet, Titouan, et al.
Publicado: (2026)
Ejemplares similares
-
Towards Early Prediction of Self-Supervised Speech Model Performance
por: Whetten, Ryan, et al.
Publicado: (2025) -
A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models
por: Whetten, Ryan, et al.
Publicado: (2026) -
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
por: Whetten, Ryan, et al.
Publicado: (2024) -
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
por: Zhang, Yuhao, et al.
Publicado: (2025) -
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
por: Choi, Jeongsoo, et al.
Publicado: (2025)