Guardado en:
| Autores principales: | Deng, Keqi, Woodland, Philip C. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2406.04541 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2023)
por: Deng, Keqi, et al.
Publicado: (2023)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
por: Deng, Keqi, et al.
Publicado: (2025)
por: Deng, Keqi, et al.
Publicado: (2025)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2024)
por: Deng, Keqi, et al.
Publicado: (2024)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
por: Deng, Keqi, et al.
Publicado: (2024)
por: Deng, Keqi, et al.
Publicado: (2024)
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
por: Hussein, Amir, et al.
Publicado: (2025)
por: Hussein, Amir, et al.
Publicado: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
por: Labiausse, Tom, et al.
Publicado: (2025)
por: Labiausse, Tom, et al.
Publicado: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
por: Le, Chenyang, et al.
Publicado: (2025)
por: Le, Chenyang, et al.
Publicado: (2025)
Transducer Consistency Regularization for Speech to Text Applications
por: Tseng, Cindy, et al.
Publicado: (2024)
por: Tseng, Cindy, et al.
Publicado: (2024)
Simultaneous Speech-to-Speech Translation Without Aligned Data
por: Labiausse, Tom, et al.
Publicado: (2026)
por: Labiausse, Tom, et al.
Publicado: (2026)
Minimising Biasing Word Errors for Contextual ASR with the Tree-Constrained Pointer Generator
por: Sun, Guangzhi, et al.
Publicado: (2022)
por: Sun, Guangzhi, et al.
Publicado: (2022)
Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
por: Wu, Wen, et al.
Publicado: (2023)
por: Wu, Wen, et al.
Publicado: (2023)
Distribution-based Emotion Recognition in Conversation
por: Wu, Wen, et al.
Publicado: (2022)
por: Wu, Wen, et al.
Publicado: (2022)
NAIST Simultaneous Speech Translation System for IWSLT 2024
por: Ko, Yuka, et al.
Publicado: (2024)
por: Ko, Yuka, et al.
Publicado: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
por: Huang, Wuwei, et al.
Publicado: (2025)
por: Huang, Wuwei, et al.
Publicado: (2025)
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
por: Pothula, Aishwarya, et al.
Publicado: (2025)
por: Pothula, Aishwarya, et al.
Publicado: (2025)
SimulTron: On-Device Simultaneous Speech to Speech Translation
por: Agranovich, Alex, et al.
Publicado: (2024)
por: Agranovich, Alex, et al.
Publicado: (2024)
A Modular-based Strategy for Mitigating Gradient Conflicts in Simultaneous Speech Translation
por: Liu, Xiaoqian, et al.
Publicado: (2024)
por: Liu, Xiaoqian, et al.
Publicado: (2024)
Direct Speech-to-Speech Neural Machine Translation: A Survey
por: Gupta, Mahendra, et al.
Publicado: (2024)
por: Gupta, Mahendra, et al.
Publicado: (2024)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
por: Zhang, Shaolei, et al.
Publicado: (2024)
por: Zhang, Shaolei, et al.
Publicado: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
por: Cheng, Shanbo, et al.
Publicado: (2025)
por: Cheng, Shanbo, et al.
Publicado: (2025)
Self-Supervised Learning for Multi-Channel Neural Transducer
por: Kojima, Atsushi
Publicado: (2024)
por: Kojima, Atsushi
Publicado: (2024)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
por: Wang, Peidong, et al.
Publicado: (2025)
por: Wang, Peidong, et al.
Publicado: (2025)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
por: Djanibekov, Amirbek, et al.
Publicado: (2026)
por: Djanibekov, Amirbek, et al.
Publicado: (2026)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
por: Zhang, Tian-Hao, et al.
Publicado: (2023)
por: Zhang, Tian-Hao, et al.
Publicado: (2023)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
por: Cheng, Shanbo, et al.
Publicado: (2024)
por: Cheng, Shanbo, et al.
Publicado: (2024)
Textless Speech-to-Speech Translation With Limited Parallel Data
por: Diwan, Anuj, et al.
Publicado: (2023)
por: Diwan, Anuj, et al.
Publicado: (2023)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
por: Liu, Xiaoqian, et al.
Publicado: (2024)
por: Liu, Xiaoqian, et al.
Publicado: (2024)
Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition
por: Moritz, Niko, et al.
Publicado: (2024)
por: Moritz, Niko, et al.
Publicado: (2024)
Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction
por: Kim, Minchan, et al.
Publicado: (2024)
por: Kim, Minchan, et al.
Publicado: (2024)
REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation
por: Hirschkind, Nameer, et al.
Publicado: (2025)
por: Hirschkind, Nameer, et al.
Publicado: (2025)
SimulSeamless: FBK at IWSLT 2024 Simultaneous Speech Translation
por: Papi, Sara, et al.
Publicado: (2024)
por: Papi, Sara, et al.
Publicado: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
por: Ma, Zhengrui, et al.
Publicado: (2024)
por: Ma, Zhengrui, et al.
Publicado: (2024)
Zero-resource Speech Translation and Recognition with LLMs
por: Mundnich, Karel, et al.
Publicado: (2024)
por: Mundnich, Karel, et al.
Publicado: (2024)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
por: Choi, Jeongsoo, et al.
Publicado: (2025)
por: Choi, Jeongsoo, et al.
Publicado: (2025)
Direct Speech to Speech Translation: A Review
por: Sarim, Mohammad, et al.
Publicado: (2025)
por: Sarim, Mohammad, et al.
Publicado: (2025)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
por: Wang, Peng, et al.
Publicado: (2023)
por: Wang, Peng, et al.
Publicado: (2023)
Promptformer: Prompted Conformer Transducer for ASR
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
por: Duarte-Torres, Sergio, et al.
Publicado: (2024)
Ejemplares similares
-
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2023) -
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
por: Deng, Keqi, et al.
Publicado: (2025) -
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
por: Deng, Keqi, et al.
Publicado: (2024) -
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
por: Deng, Keqi, et al.
Publicado: (2024) -
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
por: Hussein, Amir, et al.
Publicado: (2025)