Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Alexey, Protopopov |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
von: Diatlova, Daria, et al.
Veröffentlicht: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
von: Stourbe, Theophile, et al.
Veröffentlicht: (2024)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
von: Fan, Wei, et al.
Veröffentlicht: (2025)
von: Fan, Wei, et al.
Veröffentlicht: (2025)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
von: Hu, Yuchen, et al.
Veröffentlicht: (2023)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
von: Sun, Zhaokai, et al.
Veröffentlicht: (2025)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2024)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2024)
TRNet: Two-level Refinement Network leveraging Speech Enhancement for Noise Robust Speech Emotion Recognition
von: Chen, Chengxin, et al.
Veröffentlicht: (2024)
von: Chen, Chengxin, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
Quantifying Quanvolutional Neural Networks Robustness for Speech in Healthcare Applications
von: Tran, Ha, et al.
Veröffentlicht: (2026)
von: Tran, Ha, et al.
Veröffentlicht: (2026)
Multi-blank Transducers for Speech Recognition
von: Xu, Hainan, et al.
Veröffentlicht: (2022)
von: Xu, Hainan, et al.
Veröffentlicht: (2022)
Dynamic Gated Recurrent Neural Network for Compute-efficient Speech Enhancement
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
von: Cheng, Longbiao, et al.
Veröffentlicht: (2024)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
Test-Time Adaptation for Speech Emotion Recognition
von: Dong, Jiaheng, et al.
Veröffentlicht: (2026)
von: Dong, Jiaheng, et al.
Veröffentlicht: (2026)
Drax: Speech Recognition with Discrete Flow Matching
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
von: Navon, Aviv, et al.
Veröffentlicht: (2025)
Keyword-Guided Adaptation of Automatic Speech Recognition
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024)
von: Shamsian, Aviv, et al.
Veröffentlicht: (2024)
Are Modern Speech Enhancement Systems Vulnerable to Adversarial Attacks?
von: Makarov, Rostislav, et al.
Veröffentlicht: (2025)
von: Makarov, Rostislav, et al.
Veröffentlicht: (2025)
Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance
von: Li, Victor OK, et al.
Veröffentlicht: (2025)
von: Li, Victor OK, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2023)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
von: Lashkarashvili, Nineli, et al.
Veröffentlicht: (2024)
von: Lashkarashvili, Nineli, et al.
Veröffentlicht: (2024)
Neural Speech Extraction with Human Feedback
von: Itani, Malek, et al.
Veröffentlicht: (2025)
von: Itani, Malek, et al.
Veröffentlicht: (2025)
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
von: Ravenscroft, William, et al.
Veröffentlicht: (2024)
von: Ravenscroft, William, et al.
Veröffentlicht: (2024)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
von: Tripathi, Suraj, et al.
Veröffentlicht: (2019)
von: Tripathi, Suraj, et al.
Veröffentlicht: (2019)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
von: Yang, Zijian, et al.
Veröffentlicht: (2026)
von: Yang, Zijian, et al.
Veröffentlicht: (2026)
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
von: Sundar, Anirudh S., et al.
Veröffentlicht: (2023)
Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2025)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2025)
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications
von: Wee, Marcus Yu Zhe, et al.
Veröffentlicht: (2025)
von: Wee, Marcus Yu Zhe, et al.
Veröffentlicht: (2025)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
Latent-Domain Predictive Neural Speech Coding
von: Jiang, Xue, et al.
Veröffentlicht: (2022)
von: Jiang, Xue, et al.
Veröffentlicht: (2022)
Enhancing Speech Emotion Recognition using Dynamic Spectral Features and Kalman Smoothing
von: Hizabri, Marouane El, et al.
Veröffentlicht: (2026)
von: Hizabri, Marouane El, et al.
Veröffentlicht: (2026)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
von: Feng, Chen, et al.
Veröffentlicht: (2025)
von: Feng, Chen, et al.
Veröffentlicht: (2025)
Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2022)
von: Rajapakshe, Thejan, et al.
Veröffentlicht: (2022)
A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
von: Zhou, Nanjun, et al.
Veröffentlicht: (2025)
von: Zhou, Nanjun, et al.
Veröffentlicht: (2025)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
von: Upadhyay, Shreya G., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adapting WavLM for Speech Emotion Recognition
von: Diatlova, Daria, et al.
Veröffentlicht: (2024) -
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024) -
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024) -
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026) -
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)