Translating speech with just images
Fuente:
arXiv
Salvato in:
| Autori principali: | Oneata, Dan, Kamper, Herman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024)
di: Kamper, Herman, et al.
Pubblicazione: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
The mutual exclusivity bias of bilingual visually grounded speech models
di: Oneata, Dan, et al.
Pubblicazione: (2025)
di: Oneata, Dan, et al.
Pubblicazione: (2025)
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
Disentanglement in a GAN for Unconditional Speech Synthesis
di: Baas, Matthew, et al.
Pubblicazione: (2023)
di: Baas, Matthew, et al.
Pubblicazione: (2023)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
di: Malan, Simon, et al.
Pubblicazione: (2024)
di: Malan, Simon, et al.
Pubblicazione: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
di: Malan, Simon, et al.
Pubblicazione: (2025)
di: Malan, Simon, et al.
Pubblicazione: (2025)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
di: Lyth, Dan, et al.
Pubblicazione: (2024)
di: Lyth, Dan, et al.
Pubblicazione: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
TranSentence: Speech-to-speech Translation via Language-agnostic Sentence-level Speech Encoding without Language-parallel Data
di: Kim, Seung-Bin, et al.
Pubblicazione: (2024)
di: Kim, Seung-Bin, et al.
Pubblicazione: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023)
di: Araiza-Illan, Gloria, et al.
Pubblicazione: (2023)
Semantic enrichment towards efficient speech representations
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
Can Whisper perform speech-based in-context learning?
di: Wang, Siyin, et al.
Pubblicazione: (2023)
di: Wang, Siyin, et al.
Pubblicazione: (2023)
Self-consistent context aware conformer transducer for speech recognition
di: Kolokolov, Konstantin, et al.
Pubblicazione: (2024)
di: Kolokolov, Konstantin, et al.
Pubblicazione: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
di: Zheng, Naijun, et al.
Pubblicazione: (2024)
Direct Punjabi to English speech translation using discrete units
di: Kaur, Prabhjot, et al.
Pubblicazione: (2024)
di: Kaur, Prabhjot, et al.
Pubblicazione: (2024)
Transferable speech-to-text large language model alignment module
di: Wu, Boyong, et al.
Pubblicazione: (2024)
di: Wu, Boyong, et al.
Pubblicazione: (2024)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
di: Dong, Lukuang, et al.
Pubblicazione: (2026)
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
di: Ma, Te, et al.
Pubblicazione: (2025)
di: Ma, Te, et al.
Pubblicazione: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
di: Carbonneau, Marc-André, et al.
Pubblicazione: (2025)
di: Carbonneau, Marc-André, et al.
Pubblicazione: (2025)
Beyond the binary: Limitations and possibilities of gender-related speech technology research
di: Sanchez, Ariadna, et al.
Pubblicazione: (2024)
di: Sanchez, Ariadna, et al.
Pubblicazione: (2024)
Transferring speech-generic and depression-specific knowledge for Alzheimer's disease detection
di: Cui, Ziyun, et al.
Pubblicazione: (2023)
di: Cui, Ziyun, et al.
Pubblicazione: (2023)
An experiment on an automated literature survey of data-driven speech enhancement methods
di: Santos, Arthur dos, et al.
Pubblicazione: (2023)
di: Santos, Arthur dos, et al.
Pubblicazione: (2023)
Convoifilter: A case study of doing cocktail party speech recognition
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2023)
di: Nguyen, Thai-Binh, et al.
Pubblicazione: (2023)
Exploring the limits of decoder-only models trained on public speech recognition corpora
di: Gupta, Ankit, et al.
Pubblicazione: (2024)
di: Gupta, Ankit, et al.
Pubblicazione: (2024)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation
di: Sedukhin, Oleg, et al.
Pubblicazione: (2026)
di: Sedukhin, Oleg, et al.
Pubblicazione: (2026)
Automatic speech recognition for the Nepali language using CNN, bidirectional LSTM and ResNet
di: Dhakal, Manish, et al.
Pubblicazione: (2024)
di: Dhakal, Manish, et al.
Pubblicazione: (2024)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
Korean aegyo speech shows systematic F1 increase to signal childlike qualities
di: Kim, Ji-eun, et al.
Pubblicazione: (2026)
di: Kim, Ji-eun, et al.
Pubblicazione: (2026)
Neighbors and relatives: How do speech embeddings reflect linguistic connections across the world?
di: Törö, Tuukka, et al.
Pubblicazione: (2025)
di: Törö, Tuukka, et al.
Pubblicazione: (2025)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
di: Sanchez, Ariadna, et al.
Pubblicazione: (2025)
di: Sanchez, Ariadna, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024) -
Unsupervised lexicon learning from speech is limited by representations rather than clustering
di: Slabbert, Danel, et al.
Pubblicazione: (2025) -
The mutual exclusivity bias of bilingual visually grounded speech models
di: Oneata, Dan, et al.
Pubblicazione: (2025) -
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023) -
Disentanglement in a GAN for Unconditional Speech Synthesis
di: Baas, Matthew, et al.
Pubblicazione: (2023)