A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History Archives
Fuente:
arXiv
Salvato in:
| Autori principali: | Lehečka, Jan, Psutka, Josef V., Šmídl, Luboš, Ircing, Pavel, Psutka, Josef |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
di: Lehečka, Jan, et al.
Pubblicazione: (2023)
di: Lehečka, Jan, et al.
Pubblicazione: (2023)
Speech Technology Services for Oral History Research
di: Draxler, Christoph, et al.
Pubblicazione: (2024)
di: Draxler, Christoph, et al.
Pubblicazione: (2024)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
di: Pražák, Aleš, et al.
Pubblicazione: (2025)
di: Pražák, Aleš, et al.
Pubblicazione: (2025)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2025)
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2025)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
Augmenting Automatic Speech Recognition Models with Disfluency Detection
di: Amann, Robin, et al.
Pubblicazione: (2024)
di: Amann, Robin, et al.
Pubblicazione: (2024)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
di: Della Libera, Luca, et al.
Pubblicazione: (2026)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
AutoPsyC: Automatic Recognition of Psychodynamic Conflicts from Semi-structured Interviews with Large Language Models
di: Hossain, Sayed Muddashir, et al.
Pubblicazione: (2025)
di: Hossain, Sayed Muddashir, et al.
Pubblicazione: (2025)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
di: Yang, Yuchen, et al.
Pubblicazione: (2025)
di: Yang, Yuchen, et al.
Pubblicazione: (2025)
Error-preserving Automatic Speech Recognition of Young English Learners' Language
di: Michot, Janick, et al.
Pubblicazione: (2024)
di: Michot, Janick, et al.
Pubblicazione: (2024)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
di: Li, Yuqi, et al.
Pubblicazione: (2025)
di: Li, Yuqi, et al.
Pubblicazione: (2025)
Towards Explainability in Legal Outcome Prediction Models
di: Valvoda, Josef, et al.
Pubblicazione: (2024)
di: Valvoda, Josef, et al.
Pubblicazione: (2024)
Automatic Speech Recognition for Sanskrit with Transfer Learning
di: Sadhukhan, Bidit, et al.
Pubblicazione: (2025)
di: Sadhukhan, Bidit, et al.
Pubblicazione: (2025)
Automatic Speech Recognition for Greek Medical Dictation
di: Georgilas, Vardis, et al.
Pubblicazione: (2025)
di: Georgilas, Vardis, et al.
Pubblicazione: (2025)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
di: Hoffmann, Michael, et al.
Pubblicazione: (2025)
di: Hoffmann, Michael, et al.
Pubblicazione: (2025)
ViSpeR: Multilingual Audio-Visual Speech Recognition
di: Narayan, Sanath, et al.
Pubblicazione: (2024)
di: Narayan, Sanath, et al.
Pubblicazione: (2024)
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
di: Sayeed, Mohammad Amaan, et al.
Pubblicazione: (2023)
di: Sayeed, Mohammad Amaan, et al.
Pubblicazione: (2023)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)
di: Adnan, Renas, et al.
Pubblicazione: (2025)
di: Adnan, Renas, et al.
Pubblicazione: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
di: Lehečka, Jan, et al.
Pubblicazione: (2024)
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
di: Wunderle, Julia, et al.
Pubblicazione: (2025)
di: Wunderle, Julia, et al.
Pubblicazione: (2025)
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
di: Taguchi, Chihiro, et al.
Pubblicazione: (2024)
di: Taguchi, Chihiro, et al.
Pubblicazione: (2024)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
di: Tathe, Aniket, et al.
Pubblicazione: (2024)
di: Tathe, Aniket, et al.
Pubblicazione: (2024)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
di: Min, Do June, et al.
Pubblicazione: (2024)
di: Min, Do June, et al.
Pubblicazione: (2024)
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
di: Yang, Guanrou, et al.
Pubblicazione: (2026)
di: Yang, Guanrou, et al.
Pubblicazione: (2026)
WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
di: Zhu, Yi, et al.
Pubblicazione: (2024)
di: Zhu, Yi, et al.
Pubblicazione: (2024)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
di: Chen, William, et al.
Pubblicazione: (2025)
di: Chen, William, et al.
Pubblicazione: (2025)
Handling Numeric Expressions in Automatic Speech Recognition
di: Huber, Christian, et al.
Pubblicazione: (2024)
di: Huber, Christian, et al.
Pubblicazione: (2024)
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
di: Asano, Shunta, et al.
Pubblicazione: (2026)
di: Asano, Shunta, et al.
Pubblicazione: (2026)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
di: Saengthong, Phurich, et al.
Pubblicazione: (2025)
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
di: Jain, Yash, et al.
Pubblicazione: (2024)
di: Jain, Yash, et al.
Pubblicazione: (2024)
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025)
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025)
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
di: De Cristofaro, Domenico, et al.
Pubblicazione: (2025)
di: De Cristofaro, Domenico, et al.
Pubblicazione: (2025)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families
di: Chen, Jianan, et al.
Pubblicazione: (2026)
di: Chen, Jianan, et al.
Pubblicazione: (2026)
Large Language Models for Oral History Understanding with Text Classification and Sentiment Analysis
di: Cherukuri, Komala Subramanyam, et al.
Pubblicazione: (2025)
di: Cherukuri, Komala Subramanyam, et al.
Pubblicazione: (2025)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
di: Zhang, Shucong, et al.
Pubblicazione: (2025)
di: Zhang, Shucong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
di: Lehečka, Jan, et al.
Pubblicazione: (2023) -
Speech Technology Services for Oral History Research
di: Draxler, Christoph, et al.
Pubblicazione: (2024) -
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
di: Pražák, Aleš, et al.
Pubblicazione: (2025) -
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
di: Nguyen, Tuan, et al.
Pubblicazione: (2024) -
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2025)