Unlocking Speech Instruction Data Potential with Query Rewriting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hei, Yonghua, Yan, Yibo, Liu, Shuliang, Zhou, Huiyu, Zhang, Linfeng, Hu, Xuming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
von: Lin, Zijian, et al.
Veröffentlicht: (2025)
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
von: Zhou, Xuanru, et al.
Veröffentlicht: (2026)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2026)
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)
Analysis and Evaluation of Synthetic Data Generation in Speech Dysfluency Detection
von: Zhang, Jinming, et al.
Veröffentlicht: (2025)
von: Zhang, Jinming, et al.
Veröffentlicht: (2025)
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
von: Qin, Ruiyang, et al.
Veröffentlicht: (2024)
Voice Cloning for Dysarthric Speech Synthesis: Addressing Data Scarcity in Speech-Language Pathology
von: Moell, Birger, et al.
Veröffentlicht: (2025)
von: Moell, Birger, et al.
Veröffentlicht: (2025)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Unsupervised Speech Enhancement using Data-defined Priors
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
von: Klement, Dominik, et al.
Veröffentlicht: (2025)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
CaSNet: Compress-and-Send Network Based Multi-Device Speech Enhancement Model for Distributed Microphone Arrays
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
von: Jiang, Chengqian, et al.
Veröffentlicht: (2026)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyoung, et al.
Veröffentlicht: (2024)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
A Mel Spectrogram Enhancement Paradigm Based on CWT in Speech Synthesis
von: Hu, Guoqiang, et al.
Veröffentlicht: (2024)
von: Hu, Guoqiang, et al.
Veröffentlicht: (2024)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
von: Chang, Yi, et al.
Veröffentlicht: (2024)
von: Chang, Yi, et al.
Veröffentlicht: (2024)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
von: Chen, Sijing, et al.
Veröffentlicht: (2024)
von: Chen, Sijing, et al.
Veröffentlicht: (2024)
MEBM-Speech: Multi-scale Enhanced BrainMagic for Robust MEG Speech Detection
von: Songyi, Li, et al.
Veröffentlicht: (2026)
von: Songyi, Li, et al.
Veröffentlicht: (2026)
TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
DeepASMR: LLM-Based Zero-Shot ASMR Speech Generation for Anyone of Any Voice
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
von: Zhang, Leying, et al.
Veröffentlicht: (2026)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
A Comparison of Speech Data Augmentation Methods Using S3PRL Toolkit
von: Huh, Mina, et al.
Veröffentlicht: (2023)
von: Huh, Mina, et al.
Veröffentlicht: (2023)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
von: Xu, Mohan, et al.
Veröffentlicht: (2024)
von: Xu, Mohan, et al.
Veröffentlicht: (2024)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
von: Wang, Yuxiang, et al.
Veröffentlicht: (2026)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
When LLMs Meets Acoustic Landmarks: An Efficient Approach to Integrate Speech into Large Language Models for Depression Detection
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhang, Xiangyu, et al.
Veröffentlicht: (2024)
DSFlow: Dual Supervision and Step-Aware Architecture for One-Step Flow Matching Speech Synthesis
von: Lin, Bin, et al.
Veröffentlicht: (2026)
von: Lin, Bin, et al.
Veröffentlicht: (2026)
A Tutorial on Clinical Speech AI Development: From Data Collection to Model Validation
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2024)
von: Ng, Si-Ioi, et al.
Veröffentlicht: (2024)
Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech Separation
von: Chen, Guo, et al.
Veröffentlicht: (2025)
von: Chen, Guo, et al.
Veröffentlicht: (2025)
EventTrojan: Manipulating Non-Intrusive Speech Quality Assessment via Imperceptible Events
von: Ren, Ying, et al.
Veröffentlicht: (2023)
von: Ren, Ying, et al.
Veröffentlicht: (2023)
No Audiogram: Leveraging Existing Scores for Personalized Speech Intelligibility Prediction
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025) -
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding
von: Lin, Zijian, et al.
Veröffentlicht: (2025) -
JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
von: Zhang, Leying, et al.
Veröffentlicht: (2026) -
Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
von: Zhou, Xuanru, et al.
Veröffentlicht: (2026) -
EmoSpeech: A Corpus of Emotionally Rich and Contextually Detailed Speech Annotations
von: Bian, Weizhen, et al.
Veröffentlicht: (2024)