Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Raina, Vyas, Gales, Mark |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
von: Raina, Vyas, et al.
Veröffentlicht: (2024)
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Spontaneous Speech-Based Suicide Risk Detection Using Whisper and Large Language Models
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
von: Cui, Ziyun, et al.
Veröffentlicht: (2024)
Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Multilingual DistilWhisper: Efficient Distillation of Multi-task Speech Models via Language-Specific Experts
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
von: Ferraz, Thomas Palmeira, et al.
Veröffentlicht: (2023)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
von: Peng, Yifan, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
von: Meng, Lingwei, et al.
Veröffentlicht: (2024)
PI-Whisper: Designing an Adaptive and Incremental Automatic Speech Recognition System for Edge Devices
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
von: Nassereldine, Amir, et al.
Veröffentlicht: (2024)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
Simul-Whisper: Attention-Guided Streaming Whisper with Truncation Detection
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
von: Wang, Haoyu, et al.
Veröffentlicht: (2024)
Hallucination Benchmark for Speech Foundation Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
von: Li, Jinpeng, et al.
Veröffentlicht: (2024)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
Kid-Whisper: Towards Bridging the Performance Gap in Automatic Speech Recognition for Children VS. Adults
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2023)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2024)
SpeechGuard: Exploring the Adversarial Robustness of Multimodal Large Language Models
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
von: Peri, Raghuveer, et al.
Veröffentlicht: (2024)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Speech Foundation Models and Crowdsourcing for Efficient, High-Quality Data Collection
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
von: Lee, Beomseok, et al.
Veröffentlicht: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
von: Ao, Junyi, et al.
Veröffentlicht: (2025)
von: Ao, Junyi, et al.
Veröffentlicht: (2025)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Transfer Learning from Whisper for Microscopic Intelligibility Prediction
von: Best, Paul, et al.
Veröffentlicht: (2024)
von: Best, Paul, et al.
Veröffentlicht: (2024)
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
Extending Whisper with prompt tuning to target-speaker ASR
von: Ma, Hao, et al.
Veröffentlicht: (2023)
von: Ma, Hao, et al.
Veröffentlicht: (2023)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
von: Raina, Vyas, et al.
Veröffentlicht: (2024) -
Universal Acoustic Adversarial Attacks for Flexible Control of Speech-LLMs
von: Ma, Rao, et al.
Veröffentlicht: (2025) -
Assessment of L2 Oral Proficiency using Speech Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2025) -
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024) -
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)