Saved in:
| Main Authors: | Yang, Xingjian, Yang, Yudong, Guo, Zhixing, Zhou, Yongjie, Yan, Nan, Wang, Lan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.10161 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
by: Yang, Yudong, et al.
Published: (2024)
by: Yang, Yudong, et al.
Published: (2024)
Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection
by: Yu, Hangbin, et al.
Published: (2026)
by: Yu, Hangbin, et al.
Published: (2026)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
by: Wang, Hui, et al.
Published: (2025)
by: Wang, Hui, et al.
Published: (2025)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
by: Xue, Jun, et al.
Published: (2026)
by: Xue, Jun, et al.
Published: (2026)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
by: Feng, Tiantian, et al.
Published: (2025)
by: Feng, Tiantian, et al.
Published: (2025)
Utilizing Speaker Profiles for Impersonation Audio Detection
by: Gu, Hao, et al.
Published: (2024)
by: Gu, Hao, et al.
Published: (2024)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
by: Jin, Jiawei, et al.
Published: (2025)
by: Jin, Jiawei, et al.
Published: (2025)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
by: Fang, Qingkai, et al.
Published: (2025)
by: Fang, Qingkai, et al.
Published: (2025)
VividVoice: A Unified Framework for Scene-Aware Visually-Driven Speech Synthesis
by: Ma, Chengyuan, et al.
Published: (2026)
by: Ma, Chengyuan, et al.
Published: (2026)
Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
by: Yan, Canxiang, et al.
Published: (2025)
by: Yan, Canxiang, et al.
Published: (2025)
X-Talk: On the Underestimated Potential of Modular Speech-to-Speech Dialogue System
by: Liu, Zhanxun, et al.
Published: (2025)
by: Liu, Zhanxun, et al.
Published: (2025)
The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents
by: Wang, Lixu, et al.
Published: (2025)
by: Wang, Lixu, et al.
Published: (2025)
When Tone and Words Disagree: Towards Robust Speech Emotion Recognition under Acoustic-Semantic Conflict
by: Huang, Dawei, et al.
Published: (2026)
by: Huang, Dawei, et al.
Published: (2026)
Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
ParaGSE: Parallel Generative Speech Enhancement with Group-Vector-Quantization-based Neural Speech Codec
by: Liu, Fei, et al.
Published: (2026)
by: Liu, Fei, et al.
Published: (2026)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
by: Luong, Hieu-Thi, et al.
Published: (2024)
by: Luong, Hieu-Thi, et al.
Published: (2024)
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
by: Xie, Hanke, et al.
Published: (2025)
by: Xie, Hanke, et al.
Published: (2025)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
by: Zhou, Yang-Hao, et al.
Published: (2026)
by: Zhou, Yang-Hao, et al.
Published: (2026)
Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
by: Yang, Yudong, et al.
Published: (2024)
by: Yang, Yudong, et al.
Published: (2024)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
by: Yang, Fei, et al.
Published: (2026)
by: Yang, Fei, et al.
Published: (2026)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
by: Bai, Ye, et al.
Published: (2024)
by: Bai, Ye, et al.
Published: (2024)
Efficient Training for Cross-lingual Speech Language Models
by: Zhou, Yan, et al.
Published: (2026)
by: Zhou, Yan, et al.
Published: (2026)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
by: Pan, Yu, et al.
Published: (2025)
by: Pan, Yu, et al.
Published: (2025)
An Efficient Transfer Learning Method Based on Adapter with Local Attributes for Speech Emotion Recognition
by: Song, Haoyu, et al.
Published: (2025)
by: Song, Haoyu, et al.
Published: (2025)
Zero-Day Audio DeepFake Detection via Retrieval Augmentation and Profile Matching
by: Liu, Xuechen, et al.
Published: (2025)
by: Liu, Xuechen, et al.
Published: (2025)
Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
by: Liu, Yanyan, et al.
Published: (2025)
by: Liu, Yanyan, et al.
Published: (2025)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
by: Li, Hanzhao, et al.
Published: (2025)
by: Li, Hanzhao, et al.
Published: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
by: Li, Tao, et al.
Published: (2025)
by: Li, Tao, et al.
Published: (2025)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
SLM-SS: Speech Language Model for Generative Speech Separation
by: Li, Tianhua, et al.
Published: (2026)
by: Li, Tianhua, et al.
Published: (2026)
Few-Shot Keyword Spotting from Mixed Speech
by: Yuan, Junming, et al.
Published: (2024)
by: Yuan, Junming, et al.
Published: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
by: Lu, Wenhuan, et al.
Published: (2025)
by: Lu, Wenhuan, et al.
Published: (2025)
MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
by: Qiu, Ke, et al.
Published: (2026)
by: Qiu, Ke, et al.
Published: (2026)
From Sharpness to Better Generalization for Speech Deepfake Detection
by: Huang, Wen, et al.
Published: (2025)
by: Huang, Wen, et al.
Published: (2025)
IntMeanFlow: Few-step Speech Generation with Integral Velocity Distillation
by: Wang, Wei, et al.
Published: (2025)
by: Wang, Wei, et al.
Published: (2025)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
by: Jiang, Yicong, et al.
Published: (2024)
by: Jiang, Yicong, et al.
Published: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
by: Du, Chenpeng, et al.
Published: (2024)
by: Du, Chenpeng, et al.
Published: (2024)
SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
by: Monjur, Mahathir, et al.
Published: (2025)
by: Monjur, Mahathir, et al.
Published: (2025)
VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables
by: He, Lixing, et al.
Published: (2025)
by: He, Lixing, et al.
Published: (2025)
Similar Items
-
UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
by: Yang, Yudong, et al.
Published: (2025) -
An Audio-textual Diffusion Model For Converting Speech Signals Into Ultrasound Tongue Imaging Data
by: Yang, Yudong, et al.
Published: (2024) -
Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection
by: Yu, Hangbin, et al.
Published: (2026) -
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
by: Wang, Hui, et al.
Published: (2025) -
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
by: Xue, Jun, et al.
Published: (2026)