Saved in:
| Main Authors: | Xue, Xiangyuan, Wang, Yuyu, Yao, Ruijie, Ni, Xiaoyue, Jiang, Xiaofan, Nie, Jingping |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.27508 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Breathing and Semantic Pause Detection and Exertion-Level Classification in Post-Exercise Speech
by: Wang, Yuyu, et al.
Published: (2025)
by: Wang, Yuyu, et al.
Published: (2025)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
by: Huang, Dawei, et al.
Published: (2026)
by: Huang, Dawei, et al.
Published: (2026)
SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
by: Yao, Lingfeng, et al.
Published: (2025)
by: Yao, Lingfeng, et al.
Published: (2025)
Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
by: Xue, Xiangyuan, et al.
Published: (2026)
by: Xue, Xiangyuan, et al.
Published: (2026)
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
by: Wang, Chengyou, et al.
Published: (2026)
by: Wang, Chengyou, et al.
Published: (2026)
Foundation Model Hidden Representations for Heart Rate Estimation from Auscultation
by: Nie, Jingping, et al.
Published: (2025)
by: Nie, Jingping, et al.
Published: (2025)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
by: Huang, Mengcheng, et al.
Published: (2026)
by: Huang, Mengcheng, et al.
Published: (2026)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
by: Xie, Xurong, et al.
Published: (2022)
by: Xie, Xurong, et al.
Published: (2022)
Muting Whisper: A Universal Acoustic Adversarial Attack on Speech Foundation Models
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
by: Raina, Vyas, et al.
Published: (2024)
by: Raina, Vyas, et al.
Published: (2024)
OSUM-Pangu: An Open-Source Multidimension Speech Understanding Foundation Model Built upon OpenPangu on Ascend NPUs
by: Liao, Yujie, et al.
Published: (2026)
by: Liao, Yujie, et al.
Published: (2026)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
by: Hwang, Min-Jae, et al.
Published: (2024)
by: Hwang, Min-Jae, et al.
Published: (2024)
Optimizing Speech Language Models for Acoustic Consistency
by: Rohanian, Morteza, et al.
Published: (2025)
by: Rohanian, Morteza, et al.
Published: (2025)
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
by: Phukan, Orchid Chetia, et al.
Published: (2025)
by: Phukan, Orchid Chetia, et al.
Published: (2025)
Can We Hear from Events? Generating Speech from Event Camera
by: Fang, Jingping, et al.
Published: (2026)
by: Fang, Jingping, et al.
Published: (2026)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
by: Peng, Yuezhang, et al.
Published: (2025)
by: Peng, Yuezhang, et al.
Published: (2025)
Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
by: Nie, Jingping, et al.
Published: (2024)
by: Nie, Jingping, et al.
Published: (2024)
Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
by: Sun, Haoqin, et al.
Published: (2026)
by: Sun, Haoqin, et al.
Published: (2026)
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
by: Truong, Duc-Tuan, et al.
Published: (2025)
by: Truong, Duc-Tuan, et al.
Published: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
by: Li, Jiaqi, et al.
Published: (2024)
by: Li, Jiaqi, et al.
Published: (2024)
TADA: A Generative Framework for Speech Modeling via Text-Acoustic Dual Alignment
by: Dang, Trung, et al.
Published: (2026)
by: Dang, Trung, et al.
Published: (2026)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
by: Li, Haoxun, et al.
Published: (2025)
by: Li, Haoxun, et al.
Published: (2025)
A Neural Speech Codec for Noise Robust Speech Coding
by: Huang, Jiayi, et al.
Published: (2023)
by: Huang, Jiayi, et al.
Published: (2023)
HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot Text-to-Speech with Model and Data Scaling
by: Wang, Chunhui, et al.
Published: (2024)
by: Wang, Chunhui, et al.
Published: (2024)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
by: Wagner, Dominik, et al.
Published: (2024)
by: Wagner, Dominik, et al.
Published: (2024)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
by: Cao, Yubing, et al.
Published: (2025)
by: Cao, Yubing, et al.
Published: (2025)
UrgentMOS: Unified Multi-Metric and Preference Learning for Robust Speech Quality Assessment
by: Wang, Wei, et al.
Published: (2026)
by: Wang, Wei, et al.
Published: (2026)
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification
by: Zhao, Yiyang, et al.
Published: (2025)
by: Zhao, Yiyang, et al.
Published: (2025)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
by: Ren, Wenze, et al.
Published: (2024)
by: Ren, Wenze, et al.
Published: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
by: Shen, Feiyu, et al.
Published: (2023)
by: Shen, Feiyu, et al.
Published: (2023)
MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
by: Jin, Yutong, et al.
Published: (2026)
by: Jin, Yutong, et al.
Published: (2026)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
by: Lei, Shun, et al.
Published: (2023)
by: Lei, Shun, et al.
Published: (2023)
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
by: Zhang, Zixing, et al.
Published: (2024)
by: Zhang, Zixing, et al.
Published: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
by: Sutherland, Robert, et al.
Published: (2024)
by: Sutherland, Robert, et al.
Published: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
by: Ku, Pin-Jui, et al.
Published: (2024)
by: Ku, Pin-Jui, et al.
Published: (2024)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Hallucination Benchmark for Speech Foundation Models
by: Koudounas, Alkis, et al.
Published: (2025)
by: Koudounas, Alkis, et al.
Published: (2025)
Global Rotation Equivariant Phase Modeling for Speech Enhancement with Deep Magnitude-Phase Interaction
by: Wang, Chengzhong, et al.
Published: (2026)
by: Wang, Chengzhong, et al.
Published: (2026)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
by: Koriyama, Tomoki
Published: (2025)
by: Koriyama, Tomoki
Published: (2025)
Similar Items
-
Breathing and Semantic Pause Detection and Exertion-Level Classification in Post-Exercise Speech
by: Wang, Yuyu, et al.
Published: (2025) -
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
by: Huang, Dawei, et al.
Published: (2026) -
SpeechVerifier: Robust Acoustic Fingerprint against Tampering Attacks via Watermarking
by: Yao, Lingfeng, et al.
Published: (2025) -
Edge-Cloud Collaborative Speech Emotion Captioning via Token-Level Speculative Decoding in Audio-Language Models
by: Xue, Xiangyuan, et al.
Published: (2026) -
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
by: Wang, Chengyou, et al.
Published: (2026)