Saved in:
| Main Authors: | Lehečka, Jan, Psutka, Josef V., Šmídl, Luboš, Ircing, Pavel, Psutka, Josef |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.17160 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
by: Lehečka, Jan, et al.
Published: (2023)
by: Lehečka, Jan, et al.
Published: (2023)
Speech Technology Services for Oral History Research
by: Draxler, Christoph, et al.
Published: (2024)
by: Draxler, Christoph, et al.
Published: (2024)
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
by: Pražák, Aleš, et al.
Published: (2025)
by: Pražák, Aleš, et al.
Published: (2025)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
by: Nguyen, Tuan, et al.
Published: (2024)
by: Nguyen, Tuan, et al.
Published: (2024)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
by: Kloots, Marianne de Heer, et al.
Published: (2024)
by: Kloots, Marianne de Heer, et al.
Published: (2024)
Augmenting Automatic Speech Recognition Models with Disfluency Detection
by: Amann, Robin, et al.
Published: (2024)
by: Amann, Robin, et al.
Published: (2024)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
by: Della Libera, Luca, et al.
Published: (2026)
by: Della Libera, Luca, et al.
Published: (2026)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
by: Attia, Ahmed Adel, et al.
Published: (2024)
by: Attia, Ahmed Adel, et al.
Published: (2024)
AutoPsyC: Automatic Recognition of Psychodynamic Conflicts from Semi-structured Interviews with Large Language Models
by: Hossain, Sayed Muddashir, et al.
Published: (2025)
by: Hossain, Sayed Muddashir, et al.
Published: (2025)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
by: Lehečka, Jan, et al.
Published: (2024)
by: Lehečka, Jan, et al.
Published: (2024)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
by: Yang, Yuchen, et al.
Published: (2025)
by: Yang, Yuchen, et al.
Published: (2025)
Towards Explainability in Legal Outcome Prediction Models
by: Valvoda, Josef, et al.
Published: (2024)
by: Valvoda, Josef, et al.
Published: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
by: Hu, Shujie, et al.
Published: (2024)
by: Hu, Shujie, et al.
Published: (2024)
Error-preserving Automatic Speech Recognition of Young English Learners' Language
by: Michot, Janick, et al.
Published: (2024)
by: Michot, Janick, et al.
Published: (2024)
Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian
by: Hoffmann, Michael, et al.
Published: (2025)
by: Hoffmann, Michael, et al.
Published: (2025)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
by: Stuhlmann, Linus, et al.
Published: (2025)
by: Stuhlmann, Linus, et al.
Published: (2025)
Automatic Speech Recognition for Sanskrit with Transfer Learning
by: Sadhukhan, Bidit, et al.
Published: (2025)
by: Sadhukhan, Bidit, et al.
Published: (2025)
Automatic Speech Recognition for Greek Medical Dictation
by: Georgilas, Vardis, et al.
Published: (2025)
by: Georgilas, Vardis, et al.
Published: (2025)
Spoken Word2Vec: Learning Skipgram Embeddings from Speech
by: Sayeed, Mohammad Amaan, et al.
Published: (2023)
by: Sayeed, Mohammad Amaan, et al.
Published: (2023)
ViSpeR: Multilingual Audio-Visual Speech Recognition
by: Narayan, Sanath, et al.
Published: (2024)
by: Narayan, Sanath, et al.
Published: (2024)
Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)
by: Adnan, Renas, et al.
Published: (2025)
by: Adnan, Renas, et al.
Published: (2025)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
by: Tathe, Aniket, et al.
Published: (2024)
by: Tathe, Aniket, et al.
Published: (2024)
New Encoders for German Trained from Scratch: Comparing ModernGBERT with Converted LLM2Vec Models
by: Wunderle, Julia, et al.
Published: (2025)
by: Wunderle, Julia, et al.
Published: (2025)
Beyond Bilingual Transfer: Multilingual Code-Switching in Instruction Tuning
by: Asano, Shunta, et al.
Published: (2026)
by: Asano, Shunta, et al.
Published: (2026)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
by: Min, Do June, et al.
Published: (2024)
by: Min, Do June, et al.
Published: (2024)
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
by: Yang, Guanrou, et al.
Published: (2026)
by: Yang, Guanrou, et al.
Published: (2026)
WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
by: Chen, William, et al.
Published: (2025)
by: Chen, William, et al.
Published: (2025)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
by: Taguchi, Chihiro, et al.
Published: (2024)
by: Taguchi, Chihiro, et al.
Published: (2024)
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
by: Saengthong, Phurich, et al.
Published: (2025)
by: Saengthong, Phurich, et al.
Published: (2025)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
by: Nguyen, Tuan, et al.
Published: (2024)
by: Nguyen, Tuan, et al.
Published: (2024)
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
by: De Cristofaro, Domenico, et al.
Published: (2025)
by: De Cristofaro, Domenico, et al.
Published: (2025)
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
by: Jain, Yash, et al.
Published: (2024)
by: Jain, Yash, et al.
Published: (2024)
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
by: Choi, Anna Seo Gyeong, et al.
Published: (2025)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
by: Nfissi, Alaa, et al.
Published: (2025)
by: Nfissi, Alaa, et al.
Published: (2025)
Towards Empowering Consumers through Sentence-level Readability Scoring in German ESG Reports
by: Schüßler, Benjamin Josef, et al.
Published: (2026)
by: Schüßler, Benjamin Josef, et al.
Published: (2026)
Epicure: Navigating the Emergent Geometry of Food Ingredient Embeddings
by: Radzikowski, Jakub, et al.
Published: (2026)
by: Radzikowski, Jakub, et al.
Published: (2026)
Similar Items
-
Transfer Learning of Transformer-based Speech Recognition Models from Czech to Slovak
by: Lehečka, Jan, et al.
Published: (2023) -
Speech Technology Services for Oral History Research
by: Draxler, Christoph, et al.
Published: (2024) -
Lightweight Target-Speaker-Based Overlap Transcription for Practical Streaming ASR
by: Pražák, Aleš, et al.
Published: (2025) -
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
by: Nguyen, Tuan, et al.
Published: (2024) -
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)