EventTrojan: Manipulating Non-Intrusive Speech Quality Assessment via Imperceptible Events
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ren, Ying, Shen, Kailai, Ye, Zhe, Yan, Diqun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Breaking Speaker Recognition with PaddingBack
von: Ye, Zhe, et al.
Veröffentlicht: (2023)
von: Ye, Zhe, et al.
Veröffentlicht: (2023)
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
von: Cao, Boxuan, et al.
Veröffentlicht: (2025)
von: Cao, Boxuan, et al.
Veröffentlicht: (2025)
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024)
Imperceptible Rhythm Backdoor Attacks: Exploring Rhythm Transformation for Embedding Undetectable Vulnerabilities on Speech Recognition
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
EvMic: Event-based Non-contact sound recovery from effective spatial-temporal modeling
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
CleanMel: Mel-Spectrogram Enhancement for Improving Both Speech Quality and ASR
von: Shao, Nian, et al.
Veröffentlicht: (2025)
von: Shao, Nian, et al.
Veröffentlicht: (2025)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
Leveraging Language Model Capabilities for Sound Event Detection
von: Wang, Hualei, et al.
Veröffentlicht: (2023)
von: Wang, Hualei, et al.
Veröffentlicht: (2023)
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
von: Hu, Xintong, et al.
Veröffentlicht: (2025)
von: Hu, Xintong, et al.
Veröffentlicht: (2025)
Interpreting Pretrained Speech Models for Automatic Speech Assessment of Voice Disorders
von: Lau, Hok-Shing, et al.
Veröffentlicht: (2024)
von: Lau, Hok-Shing, et al.
Veröffentlicht: (2024)
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
von: Chen, Sijing, et al.
Veröffentlicht: (2024)
von: Chen, Sijing, et al.
Veröffentlicht: (2024)
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
FlexSED: Towards Open-Vocabulary Sound Event Detection
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
Temporal Information Reconstruction and Non-Aligned Residual in Spiking Neural Networks for Speech Classification
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
AudioScene: Integrating Object-Event Audio into 3D Scenes
von: Yuan, Shuaihang, et al.
Veröffentlicht: (2025)
von: Yuan, Shuaihang, et al.
Veröffentlicht: (2025)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
Speech-Forensics: Towards Comprehensive Synthetic Speech Dataset Establishment and Analysis
von: Ji, Zhoulin, et al.
Veröffentlicht: (2024)
von: Ji, Zhoulin, et al.
Veröffentlicht: (2024)
Enhancing Speech Quality through the Integration of BGRU and Transformer Architectures
von: Alghnam, Souliman, et al.
Veröffentlicht: (2025)
von: Alghnam, Souliman, et al.
Veröffentlicht: (2025)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Assessment of Personality Dimensions Across Situations Using Conversational Speech
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
von: Zhang, Alice, et al.
Veröffentlicht: (2025)
Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
von: Yamamoto, Katsuhiko, et al.
Veröffentlicht: (2025)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
Incremental FastPitch: Chunk-based High Quality Text to Speech
von: Du, Muyang, et al.
Veröffentlicht: (2024)
von: Du, Muyang, et al.
Veröffentlicht: (2024)
Improving Voice Quality in Speech Anonymization With Just Perception-Informed Losses
von: Ghosh, Suhita, et al.
Veröffentlicht: (2024)
von: Ghosh, Suhita, et al.
Veröffentlicht: (2024)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
von: Xiong, Chenxu, et al.
Veröffentlicht: (2024)
Investigation of Whisper ASR Hallucinations Induced by Non-Speech Audio
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
von: Barański, Mateusz, et al.
Veröffentlicht: (2025)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2025)
SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation
von: Mu, Da, et al.
Veröffentlicht: (2024)
von: Mu, Da, et al.
Veröffentlicht: (2024)
SingMOS-Pro: An Comprehensive Benchmark for Singing Quality Assessment
von: Tang, Yuxun, et al.
Veröffentlicht: (2025)
von: Tang, Yuxun, et al.
Veröffentlicht: (2025)
QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025)
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
von: Mai, Jialong, et al.
Veröffentlicht: (2025)
von: Mai, Jialong, et al.
Veröffentlicht: (2025)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Breaking Speaker Recognition with PaddingBack
von: Ye, Zhe, et al.
Veröffentlicht: (2023) -
Leveraging Multiple Speech Enhancers for Non-Intrusive Intelligibility Prediction for Hearing-Impaired Listeners
von: Cao, Boxuan, et al.
Veröffentlicht: (2025) -
Non-Intrusive Speech Intelligibility Prediction for Hearing-Impaired Users using Intermediate ASR Features and Human Memory Models
von: Mogridge, Rhiannon, et al.
Veröffentlicht: (2024) -
Imperceptible Rhythm Backdoor Attacks: Exploring Rhythm Transformation for Embedding Undetectable Vulnerabilities on Speech Recognition
von: Yao, Wenhan, et al.
Veröffentlicht: (2024) -
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
von: Zezario, Ryandhimas E., et al.
Veröffentlicht: (2023)