MORE: Multi-Objective Adversarial Attacks on Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Xiaoxue, Li, Zexin, Chen, Yiming, Chen, Nancy F. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Transferable Adversarial Attacks against ASR
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
di: Cohen, Ohad, et al.
Pubblicazione: (2024)
di: Cohen, Ohad, et al.
Pubblicazione: (2024)
Collaborative Watermarking for Adversarial Speech Synthesis
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
di: Juvela, Lauri, et al.
Pubblicazione: (2023)
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
di: Nagase, Ryotaro, et al.
Pubblicazione: (2024)
di: Nagase, Ryotaro, et al.
Pubblicazione: (2024)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
di: Li, Haoyang, et al.
Pubblicazione: (2024)
di: Li, Haoyang, et al.
Pubblicazione: (2024)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
di: Vu, Tai
Pubblicazione: (2025)
di: Vu, Tai
Pubblicazione: (2025)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
di: Muaz, Muhammad, et al.
Pubblicazione: (2024)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
di: Singh, Satwinder, et al.
Pubblicazione: (2025)
di: Singh, Satwinder, et al.
Pubblicazione: (2025)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
di: Chen, William, et al.
Pubblicazione: (2025)
di: Chen, William, et al.
Pubblicazione: (2025)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
di: Jin, Zengrui, et al.
Pubblicazione: (2022)
di: Jin, Zengrui, et al.
Pubblicazione: (2022)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
di: Singh, Karamvir
Pubblicazione: (2025)
di: Singh, Karamvir
Pubblicazione: (2025)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
di: Kamahori, Keisuke, et al.
Pubblicazione: (2025)
di: Kamahori, Keisuke, et al.
Pubblicazione: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
di: Akinrintoyo, Emmanuel, et al.
Pubblicazione: (2025)
di: Akinrintoyo, Emmanuel, et al.
Pubblicazione: (2025)
VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification
di: Wang, Pengyu, et al.
Pubblicazione: (2025)
di: Wang, Pengyu, et al.
Pubblicazione: (2025)
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
di: Bao, Chen, et al.
Pubblicazione: (2025)
di: Bao, Chen, et al.
Pubblicazione: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
di: Izotov, Yuriy, et al.
Pubblicazione: (2025)
di: Izotov, Yuriy, et al.
Pubblicazione: (2025)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
di: Nigar, Nishargo
Pubblicazione: (2024)
di: Nigar, Nishargo
Pubblicazione: (2024)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
di: Tripathi, Suraj, et al.
Pubblicazione: (2019)
di: Tripathi, Suraj, et al.
Pubblicazione: (2019)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
di: Zhang, Wenda, et al.
Pubblicazione: (2026)
Aligning Generative Speech Enhancement with Perceptual Feedback
di: Li, Haoyang, et al.
Pubblicazione: (2025)
di: Li, Haoyang, et al.
Pubblicazione: (2025)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
di: Hono, Yukiya, et al.
Pubblicazione: (2023)
di: Hono, Yukiya, et al.
Pubblicazione: (2023)
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
di: Zheng, Haolong, et al.
Pubblicazione: (2025)
di: Zheng, Haolong, et al.
Pubblicazione: (2025)
Semantically Corrected Amharic Automatic Speech Recognition
di: Adnew, Samuael, et al.
Pubblicazione: (2024)
di: Adnew, Samuael, et al.
Pubblicazione: (2024)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
di: Gao, Lingyun, et al.
Pubblicazione: (2024)
di: Gao, Lingyun, et al.
Pubblicazione: (2024)
Multi-Metric Preference Alignment for Generative Speech Restoration
di: Zhang, Junan, et al.
Pubblicazione: (2025)
di: Zhang, Junan, et al.
Pubblicazione: (2025)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
di: Kundu, Niloy Kumar, et al.
Pubblicazione: (2024)
di: Kundu, Niloy Kumar, et al.
Pubblicazione: (2024)
Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
di: Kang, Zuheng, et al.
Pubblicazione: (2024)
di: Kang, Zuheng, et al.
Pubblicazione: (2024)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
di: Jeong, Iksoon, et al.
Pubblicazione: (2025)
di: Jeong, Iksoon, et al.
Pubblicazione: (2025)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
di: Wu, Linzhi, et al.
Pubblicazione: (2026)
di: Wu, Linzhi, et al.
Pubblicazione: (2026)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
di: Saif, A F M, et al.
Pubblicazione: (2025)
di: Saif, A F M, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Transferable Adversarial Attacks against ASR
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024) -
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025) -
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
di: Gao, Xiaoxue, et al.
Pubblicazione: (2025) -
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024) -
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
di: Cohen, Ohad, et al.
Pubblicazione: (2024)