MORE: Multi-Objective Adversarial Attacks on Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Gao, Xiaoxue, Li, Zexin, Chen, Yiming, Chen, Nancy F. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Transferable Adversarial Attacks against ASR
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
por: Gao, Xiaoxue, et al.
Publicado: (2025)
por: Gao, Xiaoxue, et al.
Publicado: (2025)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
por: Gao, Xiaoxue, et al.
Publicado: (2025)
por: Gao, Xiaoxue, et al.
Publicado: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
por: Cohen, Ohad, et al.
Publicado: (2024)
por: Cohen, Ohad, et al.
Publicado: (2024)
Collaborative Watermarking for Adversarial Speech Synthesis
por: Juvela, Lauri, et al.
Publicado: (2023)
por: Juvela, Lauri, et al.
Publicado: (2023)
Can We Estimate Purchase Intention Based on Zero-shot Speech Emotion Recognition?
por: Nagase, Ryotaro, et al.
Publicado: (2024)
por: Nagase, Ryotaro, et al.
Publicado: (2024)
From KAN to GR-KAN: Advancing Speech Enhancement with KAN-Based Methodology
por: Li, Haoyang, et al.
Publicado: (2024)
por: Li, Haoyang, et al.
Publicado: (2024)
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
por: Vu, Tai
Publicado: (2025)
por: Vu, Tai
Publicado: (2025)
Bridging Modalities: Knowledge Distillation and Masked Training for Translating Multi-Modal Emotion Recognition to Uni-Modal, Speech-Only Emotion Recognition
por: Muaz, Muhammad, et al.
Publicado: (2024)
por: Muaz, Muhammad, et al.
Publicado: (2024)
Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition
por: Singh, Satwinder, et al.
Publicado: (2025)
por: Singh, Satwinder, et al.
Publicado: (2025)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
por: Chen, William, et al.
Publicado: (2025)
por: Chen, William, et al.
Publicado: (2025)
Personalized Adversarial Data Augmentation for Dysarthric and Elderly Speech Recognition
por: Jin, Zengrui, et al.
Publicado: (2022)
por: Jin, Zengrui, et al.
Publicado: (2022)
Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
por: Singh, Karamvir
Publicado: (2025)
por: Singh, Karamvir
Publicado: (2025)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
por: Kamahori, Keisuke, et al.
Publicado: (2025)
por: Kamahori, Keisuke, et al.
Publicado: (2025)
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
por: Akinrintoyo, Emmanuel, et al.
Publicado: (2025)
VINP: Variational Bayesian Inference with Neural Speech Prior for Joint ASR-Effective Speech Dereverberation and Blind RIR Identification
por: Wang, Pengyu, et al.
Publicado: (2025)
por: Wang, Pengyu, et al.
Publicado: (2025)
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
por: Bao, Chen, et al.
Publicado: (2025)
por: Bao, Chen, et al.
Publicado: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
Speech Command Recognition Using LogNNet Reservoir Computing for Embedded Systems
por: Izotov, Yuriy, et al.
Publicado: (2025)
por: Izotov, Yuriy, et al.
Publicado: (2025)
Speech Emotion Recognition Using CNN and Its Use Case in Digital Healthcare
por: Nigar, Nishargo
Publicado: (2024)
por: Nigar, Nishargo
Publicado: (2024)
Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
por: Tripathi, Suraj, et al.
Publicado: (2019)
por: Tripathi, Suraj, et al.
Publicado: (2019)
Scaling Ambiguity: Augmenting Human Annotation in Speech Emotion Recognition with Audio-Language Models
por: Zhang, Wenda, et al.
Publicado: (2026)
por: Zhang, Wenda, et al.
Publicado: (2026)
Aligning Generative Speech Enhancement with Perceptual Feedback
por: Li, Haoyang, et al.
Publicado: (2025)
por: Li, Haoyang, et al.
Publicado: (2025)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
por: Hono, Yukiya, et al.
Publicado: (2023)
por: Hono, Yukiya, et al.
Publicado: (2023)
TICL+: A Case Study On Speech In-Context Learning for Children's Speech Recognition
por: Zheng, Haolong, et al.
Publicado: (2025)
por: Zheng, Haolong, et al.
Publicado: (2025)
Semantically Corrected Amharic Automatic Speech Recognition
por: Adnew, Samuael, et al.
Publicado: (2024)
por: Adnew, Samuael, et al.
Publicado: (2024)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
por: Chi, Hyung Gun, et al.
Publicado: (2025)
por: Chi, Hyung Gun, et al.
Publicado: (2025)
Reading Miscue Detection in Primary School through Automatic Speech Recognition
por: Gao, Lingyun, et al.
Publicado: (2024)
por: Gao, Lingyun, et al.
Publicado: (2024)
Multi-Metric Preference Alignment for Generative Speech Restoration
por: Zhang, Junan, et al.
Publicado: (2025)
por: Zhang, Junan, et al.
Publicado: (2025)
Enhanced Speech Emotion Recognition with Efficient Channel Attention Guided Deep CNN-BiLSTM Framework
por: Kundu, Niloy Kumar, et al.
Publicado: (2024)
por: Kundu, Niloy Kumar, et al.
Publicado: (2024)
Efficient Multi-Model Fusion with Adversarial Complementary Representation Learning
por: Kang, Zuheng, et al.
Publicado: (2024)
por: Kang, Zuheng, et al.
Publicado: (2024)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Alternating Approach-Putt Models for Multi-Stage Speech Enhancement
por: Jeong, Iksoon, et al.
Publicado: (2025)
por: Jeong, Iksoon, et al.
Publicado: (2025)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
por: Hu, Yuchen, et al.
Publicado: (2024)
por: Hu, Yuchen, et al.
Publicado: (2024)
TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition
por: Wu, Linzhi, et al.
Publicado: (2026)
por: Wu, Linzhi, et al.
Publicado: (2026)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
por: Yang, Chao-Han Huck, et al.
Publicado: (2024)
por: Yang, Chao-Han Huck, et al.
Publicado: (2024)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
por: Du, Zhihao, et al.
Publicado: (2024)
por: Du, Zhihao, et al.
Publicado: (2024)
Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
por: Saif, A F M, et al.
Publicado: (2025)
por: Saif, A F M, et al.
Publicado: (2025)
Ejemplares similares
-
Transferable Adversarial Attacks against ASR
por: Gao, Xiaoxue, et al.
Publicado: (2024) -
MultiGen: Child-Friendly Multilingual Speech Generator with LLMs
por: Gao, Xiaoxue, et al.
Publicado: (2025) -
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
por: Gao, Xiaoxue, et al.
Publicado: (2025) -
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
por: Gao, Xiaoxue, et al.
Publicado: (2024) -
Multi-Microphone Speech Emotion Recognition using the Hierarchical Token-semantic Audio Transformer Architecture
por: Cohen, Ohad, et al.
Publicado: (2024)