Self-Train Before You Transcribe
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Flynn, Robert, Ragni, Anton |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
von: Flynn, Robert, et al.
Veröffentlicht: (2023)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
von: Flynn, Robert, et al.
Veröffentlicht: (2026)
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
von: Sun, Wanli, et al.
Veröffentlicht: (2025)
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024)
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024)
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)
TOGGL: Transcribing Overlapping Speech with Staggered Labeling
von: Li, Chak-Fai, et al.
Veröffentlicht: (2024)
von: Li, Chak-Fai, et al.
Veröffentlicht: (2024)
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
Transcribing Rhythmic Patterns of the Guitar Track in Polyphonic Music
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
von: Lukoianov, Aleksandr, et al.
Veröffentlicht: (2025)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
von: Caubrière, Antoine, et al.
Veröffentlicht: (2024)
Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
von: Kanda, Naoyuki, et al.
Veröffentlicht: (2024)
von: Kanda, Naoyuki, et al.
Veröffentlicht: (2024)
Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
von: Leem, Seong-Gyun, et al.
Veröffentlicht: (2024)
von: Leem, Seong-Gyun, et al.
Veröffentlicht: (2024)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
von: Chen, Guoguo, et al.
Veröffentlicht: (2021)
von: Chen, Guoguo, et al.
Veröffentlicht: (2021)
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
von: Li, Song, et al.
Veröffentlicht: (2024)
von: Li, Song, et al.
Veröffentlicht: (2024)
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2025)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Self-supervised learning of speech representations with Dutch archival data
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
von: Vaessen, Nik, et al.
Veröffentlicht: (2025)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
von: Kim, Eungbeom, et al.
Veröffentlicht: (2024)
von: Kim, Eungbeom, et al.
Veröffentlicht: (2024)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Lyrics Transcription for Humans: A Readability-Aware Benchmark
von: Cífka, Ondřej, et al.
Veröffentlicht: (2024)
von: Cífka, Ondřej, et al.
Veröffentlicht: (2024)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
What You Read Isn't What You Hear: Linguistic Sensitivity in Deepfake Speech Detection
von: Nguyen, Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Binh, et al.
Veröffentlicht: (2025)
Anatomy of Industrial Scale Multilingual ASR
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
von: Ramirez, Francis McCann, et al.
Veröffentlicht: (2024)
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
von: Park, Taejin, et al.
Veröffentlicht: (2024)
von: Park, Taejin, et al.
Veröffentlicht: (2024)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
von: Kim, Jongsuk, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How Much Context Does My Attention-Based ASR System Need?
von: Flynn, Robert, et al.
Veröffentlicht: (2023) -
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
von: Flynn, Robert, et al.
Veröffentlicht: (2026) -
Score-Based Training for Energy-Based TTS Models
von: Sun, Wanli, et al.
Veröffentlicht: (2025) -
Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
von: Galimzianov, Dmitrii, et al.
Veröffentlicht: (2024) -
ISPA: Inter-Species Phonetic Alphabet for Transcribing Animal Sounds
von: Hagiwara, Masato, et al.
Veröffentlicht: (2024)