SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yuqi, Zheng, Yuanzhong, Guo, Zhongtian, Wang, Yaoxuan, Yin, Jianjun, Fei, Haojun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
di: Kounadis-Bastian, Dionyssos, et al.
Pubblicazione: (2024)
di: Kounadis-Bastian, Dionyssos, et al.
Pubblicazione: (2024)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
di: Alexey, Protopopov
Pubblicazione: (2026)
di: Alexey, Protopopov
Pubblicazione: (2026)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024)
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
di: Pietroń, Marcin, et al.
Pubblicazione: (2026)
di: Pietroń, Marcin, et al.
Pubblicazione: (2026)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
di: Chen, Jinming, et al.
Pubblicazione: (2024)
di: Chen, Jinming, et al.
Pubblicazione: (2024)
Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2024)
QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
di: Wu, Zhiyu, et al.
Pubblicazione: (2025)
di: Wu, Zhiyu, et al.
Pubblicazione: (2025)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
Adapting WavLM for Speech Emotion Recognition
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
di: Diatlova, Daria, et al.
Pubblicazione: (2024)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
di: Jafarzadeh, Pourya, et al.
Pubblicazione: (2024)
di: Jafarzadeh, Pourya, et al.
Pubblicazione: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
WavInWav: Time-domain Speech Hiding via Invertible Neural Network
di: Fan, Wei, et al.
Pubblicazione: (2025)
di: Fan, Wei, et al.
Pubblicazione: (2025)
Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
di: Stourbe, Theophile, et al.
Pubblicazione: (2024)
di: Stourbe, Theophile, et al.
Pubblicazione: (2024)
WavMark: Watermarking for Audio Generation
di: Chen, Guangyu, et al.
Pubblicazione: (2023)
di: Chen, Guangyu, et al.
Pubblicazione: (2023)
Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech
di: Nakazawa, Kazushi
Pubblicazione: (2026)
di: Nakazawa, Kazushi
Pubblicazione: (2026)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
di: Li, Feng, et al.
Pubblicazione: (2024)
di: Li, Feng, et al.
Pubblicazione: (2024)
ManWav: The First Manchu ASR Model
di: Seo, Jean, et al.
Pubblicazione: (2024)
di: Seo, Jean, et al.
Pubblicazione: (2024)
Wav2code: Restore Clean Speech Representations via Codebook Lookup for Noise-Robust ASR
di: Hu, Yuchen, et al.
Pubblicazione: (2023)
di: Hu, Yuchen, et al.
Pubblicazione: (2023)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2025)
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2025)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
di: Nfissi, Alaa, et al.
Pubblicazione: (2025)
di: Nfissi, Alaa, et al.
Pubblicazione: (2025)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
di: Baser, Oguzhan, et al.
Pubblicazione: (2025)
di: Baser, Oguzhan, et al.
Pubblicazione: (2025)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
di: Hu, Shujie, et al.
Pubblicazione: (2024)
di: Hu, Shujie, et al.
Pubblicazione: (2024)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2025)
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2025)
Towards Robust Overlapping Speech Detection: A Speaker-Aware Progressive Approach Using WavLM
di: Sun, Zhaokai, et al.
Pubblicazione: (2025)
di: Sun, Zhaokai, et al.
Pubblicazione: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
ROAR: Reinforcing Original to Augmented Data Ratio Dynamics for Wav2Vec2.0 Based ASR
di: Singh, Vishwanath Pratap, et al.
Pubblicazione: (2024)
di: Singh, Vishwanath Pratap, et al.
Pubblicazione: (2024)
WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models
di: Chen, Yifu, et al.
Pubblicazione: (2025)
di: Chen, Yifu, et al.
Pubblicazione: (2025)
XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
di: Zhang, Qishan, et al.
Pubblicazione: (2024)
di: Zhang, Qishan, et al.
Pubblicazione: (2024)
Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial Animation
di: Li, Hao, et al.
Pubblicazione: (2025)
di: Li, Hao, et al.
Pubblicazione: (2025)
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2025)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
di: Comunità, Marco, et al.
Pubblicazione: (2024)
di: Comunità, Marco, et al.
Pubblicazione: (2024)
Transcription and translation of videos using fine-tuned XLSR Wav2Vec2 on custom dataset and mBART
di: Tathe, Aniket, et al.
Pubblicazione: (2024)
di: Tathe, Aniket, et al.
Pubblicazione: (2024)
Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
di: Zhang, Yanzhe, et al.
Pubblicazione: (2024)
di: Zhang, Yanzhe, et al.
Pubblicazione: (2024)
BAST: Binaural Audio Spectrogram Transformer for Binaural Sound Localization
di: Kuang, Sheng, et al.
Pubblicazione: (2022)
di: Kuang, Sheng, et al.
Pubblicazione: (2022)
WavChat: A Survey of Spoken Dialogue Models
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
di: Kounadis-Bastian, Dionyssos, et al.
Pubblicazione: (2024) -
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
di: Alexey, Protopopov
Pubblicazione: (2026) -
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024) -
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
di: Pietroń, Marcin, et al.
Pubblicazione: (2026) -
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
di: Chen, Jinming, et al.
Pubblicazione: (2024)