Gespeichert in:
| Hauptverfasser: | Li, Hanzhao, Li, Yuke, Wang, Xinsheng, Hu, Jingbin, Xie, Qicong, Yang, Shan, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2501.04644 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
von: Xie, Hanke, et al.
Veröffentlicht: (2025)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
von: Chen, Huakang, et al.
Veröffentlicht: (2026)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
von: Chen, Peikun, et al.
Veröffentlicht: (2024)
SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
von: Li, Yue, et al.
Veröffentlicht: (2024)
von: Li, Yue, et al.
Veröffentlicht: (2024)
WenetSpeech-Wu: Datasets, Benchmarks, and Models for a Unified Chinese Wu Dialect Speech Processing Ecosystem
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
von: Wang, Chengyou, et al.
Veröffentlicht: (2026)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
EDSep: An Effective Diffusion-Based Method for Speech Source Separation
von: Dong, Jinwei, et al.
Veröffentlicht: (2025)
von: Dong, Jinwei, et al.
Veröffentlicht: (2025)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
von: Li, Yangze, et al.
Veröffentlicht: (2024)
von: Li, Yangze, et al.
Veröffentlicht: (2024)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
von: Xie, Kun, et al.
Veröffentlicht: (2025)
von: Xie, Kun, et al.
Veröffentlicht: (2025)
SponTTS: modeling and transferring spontaneous style for TTS
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
von: Li, Hanzhao, et al.
Veröffentlicht: (2023)
Adaptive Data Augmentation with NaturalSpeech3 for Far-field Speaker Verification
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
von: Lin, Yuke, et al.
Veröffentlicht: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
Steering Language Model to Stable Speech Emotion Recognition via Contextual Perception and Chain of Thought
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
von: Guo, Hao-Han, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
von: Xie, Tianxin, et al.
Veröffentlicht: (2025)
Coarse-to-fine Alignment Makes Better Speech-image Retrieval
von: Zhou, Lifeng, et al.
Veröffentlicht: (2024)
von: Zhou, Lifeng, et al.
Veröffentlicht: (2024)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
von: Yao, Jixun, et al.
Veröffentlicht: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
von: Tian, Wenjie, et al.
Veröffentlicht: (2025)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
Acoustic BPE for Speech Generation with Discrete Tokens
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
von: Shen, Feiyu, et al.
Veröffentlicht: (2023)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
von: Shan, Weiqiao, et al.
Veröffentlicht: (2025)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
DialoSpeech: Dual-Speaker Dialogue Generation with LLM and Flow Matching
von: Xie, Hanke, et al.
Veröffentlicht: (2025) -
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023) -
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
von: Chen, Huakang, et al.
Veröffentlicht: (2026) -
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
von: Chen, Peikun, et al.
Veröffentlicht: (2024) -
SCDNet: Self-supervised Learning Feature-based Speaker Change Detection
von: Li, Yue, et al.
Veröffentlicht: (2024)