A Closer Look at Wav2Vec2 Embeddings for On-Device Single-Channel Speech Enhancement
Fuente:
arXiv
Guardado en:
| Autores principales: | Shankar, Ravi, Tan, Ke, Xu, Buye, Kumar, Anurag |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
por: Li, Yuqi, et al.
Publicado: (2025)
por: Li, Yuqi, et al.
Publicado: (2025)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
por: Jafarzadeh, Pourya, et al.
Publicado: (2024)
por: Jafarzadeh, Pourya, et al.
Publicado: (2024)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
por: Nguyen, Tuan, et al.
Publicado: (2024)
por: Nguyen, Tuan, et al.
Publicado: (2024)
Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
por: Shankar, Ravi, et al.
Publicado: (2024)
por: Shankar, Ravi, et al.
Publicado: (2024)
AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
por: Kalkhorani, Vahid Ahmadi, et al.
Publicado: (2024)
por: Kalkhorani, Vahid Ahmadi, et al.
Publicado: (2024)
xLSTM-SENet: xLSTM for Single-Channel Speech Enhancement
por: Kühne, Nikolai Lund, et al.
Publicado: (2025)
por: Kühne, Nikolai Lund, et al.
Publicado: (2025)
End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2
por: Tathe, Aniket, et al.
Publicado: (2024)
por: Tathe, Aniket, et al.
Publicado: (2024)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
por: Kloots, Marianne de Heer, et al.
Publicado: (2024)
por: Kloots, Marianne de Heer, et al.
Publicado: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
por: Stuhlmann, Linus, et al.
Publicado: (2025)
por: Stuhlmann, Linus, et al.
Publicado: (2025)
MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement
por: Kühne, Nikolai Lund, et al.
Publicado: (2025)
por: Kühne, Nikolai Lund, et al.
Publicado: (2025)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
por: Alexey, Protopopov
Publicado: (2026)
por: Alexey, Protopopov
Publicado: (2026)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
por: Kounadis-Bastian, Dionyssos, et al.
Publicado: (2024)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
por: Fan, Wei, et al.
Publicado: (2025)
por: Fan, Wei, et al.
Publicado: (2025)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
por: Pandey, Ashutosh, et al.
Publicado: (2024)
por: Pandey, Ashutosh, et al.
Publicado: (2024)
ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement
por: Xu, Zhongweiyang, et al.
Publicado: (2026)
por: Xu, Zhongweiyang, et al.
Publicado: (2026)
Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation
por: Xu, Zhongweiyang, et al.
Publicado: (2026)
por: Xu, Zhongweiyang, et al.
Publicado: (2026)
Speech Enhancement Using Continuous Embeddings of Neural Audio Codec
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Single and Few-step Diffusion for Generative Speech Enhancement
por: Lay, Bunlong, et al.
Publicado: (2023)
por: Lay, Bunlong, et al.
Publicado: (2023)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
por: Ridoy, Md Sazzadul Islam, et al.
Publicado: (2025)
por: Ridoy, Md Sazzadul Islam, et al.
Publicado: (2025)
Personalized Speech Enhancement Without a Separate Speaker Embedding Model
por: Pärnamaa, Tanel, et al.
Publicado: (2024)
por: Pärnamaa, Tanel, et al.
Publicado: (2024)
Aligning Generative Speech Enhancement with Perceptual Feedback
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
por: Nguyen, Tuan, et al.
Publicado: (2024)
por: Nguyen, Tuan, et al.
Publicado: (2024)
Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
por: Cheng, Longbiao, et al.
Publicado: (2024)
por: Cheng, Longbiao, et al.
Publicado: (2024)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
por: Baser, Oguzhan, et al.
Publicado: (2025)
por: Baser, Oguzhan, et al.
Publicado: (2025)
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
por: Yang, Guanrou, et al.
Publicado: (2026)
por: Yang, Guanrou, et al.
Publicado: (2026)
Improving Resource-Efficient Speech Enhancement via Neural Differentiable DSP Vocoder Refinement
por: Guimarães, Heitor R., et al.
Publicado: (2025)
por: Guimarães, Heitor R., et al.
Publicado: (2025)
Continued Pretraining for Domain Adaptation of Wav2vec2.0 in Automatic Speech Recognition for Elementary Math Classroom Settings
por: Attia, Ahmed Adel, et al.
Publicado: (2024)
por: Attia, Ahmed Adel, et al.
Publicado: (2024)
Active Speech Enhancement: Active Speech Denoising Decliping and Deveraberation
por: Yaish, Ofir, et al.
Publicado: (2025)
por: Yaish, Ofir, et al.
Publicado: (2025)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
por: Li, Haoyang, et al.
Publicado: (2024)
por: Li, Haoyang, et al.
Publicado: (2024)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
por: Pietroń, Marcin, et al.
Publicado: (2026)
por: Pietroń, Marcin, et al.
Publicado: (2026)
BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
por: Mattursun, Alimjan, et al.
Publicado: (2024)
por: Mattursun, Alimjan, et al.
Publicado: (2024)
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
por: Ji, Shengpeng, et al.
Publicado: (2025)
por: Ji, Shengpeng, et al.
Publicado: (2025)
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
por: Singh, Vishwanath Pratap, et al.
Publicado: (2024)
por: Singh, Vishwanath Pratap, et al.
Publicado: (2024)
Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech
por: Fu, Szu-Wei, et al.
Publicado: (2024)
por: Fu, Szu-Wei, et al.
Publicado: (2024)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
por: Hu, Guoqiang, et al.
Publicado: (2024)
por: Hu, Guoqiang, et al.
Publicado: (2024)
EffiFusion-GAN: Efficient Fusion Generative Adversarial Network for Speech Enhancement
por: Wen, Bin, et al.
Publicado: (2025)
por: Wen, Bin, et al.
Publicado: (2025)
SEED: Speaker Embedding Enhancement Diffusion Model
por: Nam, KiHyun, et al.
Publicado: (2025)
por: Nam, KiHyun, et al.
Publicado: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
por: Ruggiero, Giuseppe, et al.
Publicado: (2025)
por: Ruggiero, Giuseppe, et al.
Publicado: (2025)
CMGAN: Conformer-based Metric GAN for Speech Enhancement
por: Cao, Ruizhe, et al.
Publicado: (2022)
por: Cao, Ruizhe, et al.
Publicado: (2022)
Schrödinger Bridge Mamba for One-Step Speech Enhancement
por: Yang, Jing, et al.
Publicado: (2025)
por: Yang, Jing, et al.
Publicado: (2025)
Ejemplares similares
-
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
por: Li, Yuqi, et al.
Publicado: (2025) -
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
por: Jafarzadeh, Pourya, et al.
Publicado: (2024) -
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
por: Nguyen, Tuan, et al.
Publicado: (2024) -
Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
por: Shankar, Ravi, et al.
Publicado: (2024) -
AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling
por: Kalkhorani, Vahid Ahmadi, et al.
Publicado: (2024)