PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Runyan, Yang, Huibao, Zhang, Xiqing, Ye, Tiantian, Liu, Ying, Gao, Yingying, Zhang, Shilei, Deng, Chao, Feng, Junlan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
von: Si, Yuke, et al.
Veröffentlicht: (2025)
von: Si, Yuke, et al.
Veröffentlicht: (2025)
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
von: Yang, Runyan, et al.
Veröffentlicht: (2025)
von: Yang, Runyan, et al.
Veröffentlicht: (2025)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
von: Chen, Yanan, et al.
Veröffentlicht: (2024)
On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
von: Hao, Yaqian, et al.
Veröffentlicht: (2024)
von: Hao, Yaqian, et al.
Veröffentlicht: (2024)
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
von: Gao, Yingying, et al.
Veröffentlicht: (2024)
von: Gao, Yingying, et al.
Veröffentlicht: (2024)
B-GRPO: Unsupervised Speech Emotion Recognition based on Batched-Group Relative Policy Optimization
von: Gao, Yingying, et al.
Veröffentlicht: (2026)
von: Gao, Yingying, et al.
Veröffentlicht: (2026)
MFSN: Multi-perspective Fusion Search Network For Pre-training Knowledge in Speech Emotion Recognition
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
von: Sun, Haiyang, et al.
Veröffentlicht: (2023)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
CEC: A Noisy Label Detection Method for Speaker Recognition
von: Shen, Yao, et al.
Veröffentlicht: (2024)
von: Shen, Yao, et al.
Veröffentlicht: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
UniArray: Unified Spectral-Spatial Modeling for Array-Geometry-Agnostic Speech Separation
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
von: Chen, Weiguang, et al.
Veröffentlicht: (2025)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
Speech Denoising with Auditory Models
von: Saddler, Mark R., et al.
Veröffentlicht: (2020)
von: Saddler, Mark R., et al.
Veröffentlicht: (2020)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
URGENT-PK: Perceptually-Aligned Ranking Model Designed for Speech Enhancement Competition
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
von: Wang, Jiahe, et al.
Veröffentlicht: (2025)
Losses Can Be Blessings: Routing Self-Supervised Speech Representations Towards Efficient Multilingual and Multitask Speech Processing
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
von: Fu, Yonggan, et al.
Veröffentlicht: (2022)
StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model
von: Xiang, Yang, et al.
Veröffentlicht: (2025)
von: Xiang, Yang, et al.
Veröffentlicht: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
von: Yang, Yiqing, et al.
Veröffentlicht: (2025)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxuan, et al.
Veröffentlicht: (2024)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
Efficient Compression of Multitask Multilingual Speech Models
von: Ferraz, Thomas Palmeira
Veröffentlicht: (2024)
von: Ferraz, Thomas Palmeira
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Adaptive Central Frequencies Locally Competitive Algorithm for Speech
von: Bahadi, Soufiyan, et al.
Veröffentlicht: (2025)
von: Bahadi, Soufiyan, et al.
Veröffentlicht: (2025)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
von: Li, Jizhen, et al.
Veröffentlicht: (2025)
von: Li, Jizhen, et al.
Veröffentlicht: (2025)
Complex-Cycle-Consistent Diffusion Model for Monaural Speech Enhancement
von: Li, Yi, et al.
Veröffentlicht: (2024)
von: Li, Yi, et al.
Veröffentlicht: (2024)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HarmoniFuse: A Component-Selective and Prompt-Adaptive Framework for Multi-Task Speech Language Modeling
von: Si, Yuke, et al.
Veröffentlicht: (2025) -
Teaching Audio Models to Reason: A Unified Framework for Source- and Layer-wise Distillation
von: Yang, Runyan, et al.
Veröffentlicht: (2025) -
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
von: Chen, Yanan, et al.
Veröffentlicht: (2024) -
On Calibration of Speech Classification Models: Insights from Energy-Based Model Investigations
von: Hao, Yaqian, et al.
Veröffentlicht: (2024) -
GenDistiller: Distilling Pre-trained Language Models based on an Autoregressive Generative Model
von: Gao, Yingying, et al.
Veröffentlicht: (2024)