What does it take to get state of the art in simultaneous speech-to-speech translation?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wilmet, Vincent, Du, Johnson |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language translation, and change of accent for speech-to-speech task using diffusion model
von: Mishra, Abhishek, et al.
Veröffentlicht: (2025)
von: Mishra, Abhishek, et al.
Veröffentlicht: (2025)
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
Addressing speaker gender bias in large scale speech translation systems
von: Bansal, Shubham, et al.
Veröffentlicht: (2025)
von: Bansal, Shubham, et al.
Veröffentlicht: (2025)
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024)
What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
von: Verhoef, Tessa, et al.
Veröffentlicht: (2024)
von: Verhoef, Tessa, et al.
Veröffentlicht: (2024)
Prominence-aware automatic speech recognition for conversational speech
von: Linke, Julian, et al.
Veröffentlicht: (2025)
von: Linke, Julian, et al.
Veröffentlicht: (2025)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
von: Akkiraju, Bhavana, et al.
Veröffentlicht: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
What is the social benefit of hate speech detection research? A Systematic Review
von: Wong, Sidney Gig-Jan
Veröffentlicht: (2024)
von: Wong, Sidney Gig-Jan
Veröffentlicht: (2024)
Limit cycles for speech
von: Gafos, Adamantios I., et al.
Veröffentlicht: (2025)
von: Gafos, Adamantios I., et al.
Veröffentlicht: (2025)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
von: Kesiraju, Santosh, et al.
Veröffentlicht: (2023)
Code-mixed Sentiment and Hate-speech Prediction
von: Yadav, Anjali, et al.
Veröffentlicht: (2024)
von: Yadav, Anjali, et al.
Veröffentlicht: (2024)
Sustainable self-supervised learning for speech representations
von: Lugo, Luis, et al.
Veröffentlicht: (2024)
von: Lugo, Luis, et al.
Veröffentlicht: (2024)
Domain adapted machine translation: What does catastrophic forgetting forget and why?
von: Saunders, Danielle, et al.
Veröffentlicht: (2024)
von: Saunders, Danielle, et al.
Veröffentlicht: (2024)
SpeakGer: A meta-data enriched speech corpus of German state and federal parliaments
von: Lange, Kai-Robin, et al.
Veröffentlicht: (2024)
von: Lange, Kai-Robin, et al.
Veröffentlicht: (2024)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
Translating speech with just images
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
von: Oneata, Dan, et al.
Veröffentlicht: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
von: Zheng, Naijun, et al.
Veröffentlicht: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
von: Okocha, Chibuzor, et al.
Veröffentlicht: (2025)
SCALAR: A Part-of-speech Tagger for Identifiers
von: Newman, Christian D., et al.
Veröffentlicht: (2025)
von: Newman, Christian D., et al.
Veröffentlicht: (2025)
Discovering dynamical laws for speech gestures
von: Kirkham, Sam
Veröffentlicht: (2025)
von: Kirkham, Sam
Veröffentlicht: (2025)
Exploring the topics, sentiments and hate speech in the Spanish information environment
von: LOPEZ, ALEJANDRO BUITRAGO, et al.
Veröffentlicht: (2024)
von: LOPEZ, ALEJANDRO BUITRAGO, et al.
Veröffentlicht: (2024)
Child-directed speech facilitates production, not comprehension, in BabyLMs
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2026)
von: Bunzeck, Bastian, et al.
Veröffentlicht: (2026)
Emergent morpho-phonological representations in self-supervised speech models
von: Gauthier, Jon, et al.
Veröffentlicht: (2025)
von: Gauthier, Jon, et al.
Veröffentlicht: (2025)
Neural inhibition during speech planning contributes to contrastive hyperarticulation
von: Stern, Michael C., et al.
Veröffentlicht: (2022)
von: Stern, Michael C., et al.
Veröffentlicht: (2022)
A stylometric analysis of speaker attribution from speech transcripts
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025)
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2025)
End-to-end Speech Recognition with similar length speech and text
von: Fan, Peng, et al.
Veröffentlicht: (2025)
von: Fan, Peng, et al.
Veröffentlicht: (2025)
Hate speech detection in algerian dialect using deep learning
von: Lanasri, Dihia, et al.
Veröffentlicht: (2023)
von: Lanasri, Dihia, et al.
Veröffentlicht: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
On the reliability of feature attribution methods for speech classification
von: Shen, Gaofei, et al.
Veröffentlicht: (2025)
von: Shen, Gaofei, et al.
Veröffentlicht: (2025)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
von: Araiza-Illan, Gloria, et al.
Veröffentlicht: (2023)
Part-of-speech tagging for Nagamese Language using CRF
von: Shohe, Alovi N, et al.
Veröffentlicht: (2025)
von: Shohe, Alovi N, et al.
Veröffentlicht: (2025)
Tagarela - A Portuguese speech dataset from podcasts
von: de Oliveira, Frederico Santos, et al.
Veröffentlicht: (2026)
von: de Oliveira, Frederico Santos, et al.
Veröffentlicht: (2026)
What does it take to certify a conversion checker?
von: Lennon-Bertrand, Meven
Veröffentlicht: (2025)
von: Lennon-Bertrand, Meven
Veröffentlicht: (2025)
Covertly improving intelligibility with data-driven adaptations of speech timing
von: Tuttösí, Paige, et al.
Veröffentlicht: (2026)
von: Tuttösí, Paige, et al.
Veröffentlicht: (2026)
Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
von: Mohamed, Mukhtar, et al.
Veröffentlicht: (2024)
von: Mohamed, Mukhtar, et al.
Veröffentlicht: (2024)
Code-switching in text and speech challenges information-theoretic speaker design
von: Bhattacharya, Debasmita, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Debasmita, et al.
Veröffentlicht: (2024)
negativas: a prototype for searching and classifying sentential negation in speech data
von: de Gois, Túlio Sousa, et al.
Veröffentlicht: (2025)
von: de Gois, Túlio Sousa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Language translation, and change of accent for speech-to-speech task using diffusion model
von: Mishra, Abhishek, et al.
Veröffentlicht: (2025) -
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
von: Moell, Birger, et al.
Veröffentlicht: (2025) -
Addressing speaker gender bias in large scale speech translation systems
von: Bansal, Shubham, et al.
Veröffentlicht: (2025) -
Direct Punjabi to English speech translation using discrete units
von: Kaur, Prabhjot, et al.
Veröffentlicht: (2024) -
What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
von: Verhoef, Tessa, et al.
Veröffentlicht: (2024)