Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kounadis-Bastian, Dionyssos, Schrüfer, Oliver, Derington, Anna, Wierstorf, Hagen, Eyben, Florian, Burkhardt, Felix, Schuller, Björn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
von: Derington, Anna, et al.
Veröffentlicht: (2023)
von: Derington, Anna, et al.
Veröffentlicht: (2023)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024)
Teaching Wav2Vec2 the Language of the Brain
von: Fiedler, Tobias, et al.
Veröffentlicht: (2025)
von: Fiedler, Tobias, et al.
Veröffentlicht: (2025)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)
A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement
von: Shankar, Ravi, et al.
Veröffentlicht: (2024)
von: Shankar, Ravi, et al.
Veröffentlicht: (2024)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
von: Alexey, Protopopov
Veröffentlicht: (2026)
von: Alexey, Protopopov
Veröffentlicht: (2026)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
von: Fan, Wei, et al.
Veröffentlicht: (2025)
von: Fan, Wei, et al.
Veröffentlicht: (2025)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
Which one Performs Better? Wav2Vec or Whisper? Applying both in Badini Kurdish Speech to Text (BKSTT)
von: Adnan, Renas, et al.
Veröffentlicht: (2025)
von: Adnan, Renas, et al.
Veröffentlicht: (2025)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
A Comparative Analysis of Bilingual and Trilingual Wav2Vec Models for Automatic Speech Recognition in Multilingual Oral History Archives
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
von: Lehečka, Jan, et al.
Veröffentlicht: (2024)
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2024)
von: Singh, Vishwanath Pratap, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
Transcription and translation of videos using fine-tuned XLSR Wav2Vec2 on custom dataset and mBART
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
Using voice analysis as an early indicator of risk for depression in young adults
von: Scherer, Klaus R., et al.
Veröffentlicht: (2024)
von: Scherer, Klaus R., et al.
Veröffentlicht: (2024)
Wav2Gloss: Generating Interlinear Glossed Text from Speech
von: He, Taiqi, et al.
Veröffentlicht: (2024)
von: He, Taiqi, et al.
Veröffentlicht: (2024)
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
von: De Cristofaro, Domenico, et al.
Veröffentlicht: (2025)
von: De Cristofaro, Domenico, et al.
Veröffentlicht: (2025)
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
von: Tathe, Aniket, et al.
Veröffentlicht: (2024)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
Enrolment-based personalisation for improving individual-level fairness in speech emotion recognition
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
von: Triantafyllopoulos, Andreas, et al.
Veröffentlicht: (2024)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
von: Yang, Yuchen, et al.
Veröffentlicht: (2025)
von: Yang, Yuchen, et al.
Veröffentlicht: (2025)
WavMark: Watermarking for Audio Generation
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
BrainWavLM: Fine-tuning Speech Representations with Brain Responses to Language
von: Vattikonda, Nishitha, et al.
Veröffentlicht: (2025)
von: Vattikonda, Nishitha, et al.
Veröffentlicht: (2025)
Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
ManWav: The First Manchu ASR Model
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
von: Derington, Anna, et al.
Veröffentlicht: (2023) -
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
von: Schrüfer, Oliver, et al.
Veröffentlicht: (2024) -
Teaching Wav2Vec2 the Language of the Brain
von: Fiedler, Tobias, et al.
Veröffentlicht: (2025) -
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025) -
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
von: Della Libera, Luca, et al.
Veröffentlicht: (2026)