Gespeichert in:
| Hauptverfasser: | Li, Zhipeng, Xing, Xiaofen, Wang, Jun, Chen, Shuaiqi, Yu, Guoqiao, Wan, Guanglu, Xu, Xiangmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2409.05730 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Long-Context Speech Synthesis with Context-Aware Memory
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
von: Li, Zhipeng, et al.
Veröffentlicht: (2025)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
von: Li, Zhipeng, et al.
Veröffentlicht: (2024)
Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
von: Chen, Weidong, et al.
Veröffentlicht: (2023)
von: Chen, Weidong, et al.
Veröffentlicht: (2023)
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)
von: Xin, Detai, et al.
Veröffentlicht: (2026)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
von: Mai, Jialong, et al.
Veröffentlicht: (2025)
von: Mai, Jialong, et al.
Veröffentlicht: (2025)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
von: Nie, Sihang, et al.
Veröffentlicht: (2025)
von: Nie, Sihang, et al.
Veröffentlicht: (2025)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2025)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
ISSE: An Instruction-Guided Speech Style Editing Dataset And Benchmark
von: Chen, Yun, et al.
Veröffentlicht: (2025)
von: Chen, Yun, et al.
Veröffentlicht: (2025)
MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research
von: Li, Song, et al.
Veröffentlicht: (2024)
von: Li, Song, et al.
Veröffentlicht: (2024)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025)
von: Zhang, Tian-Hao, et al.
Veröffentlicht: (2025)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
von: Deng, Yimin, et al.
Veröffentlicht: (2024)
SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis
von: Wang, Huimeng, et al.
Veröffentlicht: (2026)
von: Wang, Huimeng, et al.
Veröffentlicht: (2026)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
von: Chung, Raymond
Veröffentlicht: (2026)
von: Chung, Raymond
Veröffentlicht: (2026)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
von: Kuhlmann, Michael, et al.
Veröffentlicht: (2026)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
von: Tao, Dehua, et al.
Veröffentlicht: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
von: Chen, Weidong, et al.
Veröffentlicht: (2025)
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
von: Tao, Dehua, et al.
Veröffentlicht: (2026)
von: Tao, Dehua, et al.
Veröffentlicht: (2026)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
Adapting Speech Foundation Models for Unified Multimodal Speech Recognition with Large Language Models
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
von: Zhang, Jing-Xuan, et al.
Veröffentlicht: (2025)
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
von: Huang, Yiqiao, et al.
Veröffentlicht: (2024)
von: Huang, Yiqiao, et al.
Veröffentlicht: (2024)
UniTalker: Conversational Speech-Visual Synthesis
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
von: Ma, Linhan, et al.
Veröffentlicht: (2025)
von: Ma, Linhan, et al.
Veröffentlicht: (2025)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
von: Liu, Wenrui, et al.
Veröffentlicht: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
A Survey on Speech Large Language Models for Understanding
von: Peng, Jing, et al.
Veröffentlicht: (2024)
von: Peng, Jing, et al.
Veröffentlicht: (2024)
Adaptive Speaker Embedding Self-Augmentation for Personal Voice Activity Detection with Short Enrollment Speech
von: Feng, Fuyuan, et al.
Veröffentlicht: (2026)
von: Feng, Fuyuan, et al.
Veröffentlicht: (2026)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2023)
EmoSSLSphere: Multilingual Emotional Speech Synthesis with Spherical Vectors and Discrete Speech Tokens
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Long-Context Speech Synthesis with Context-Aware Memory
von: Li, Zhipeng, et al.
Veröffentlicht: (2025) -
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
von: Xing, Jingyuan, et al.
Veröffentlicht: (2025) -
Multi-Scale Temporal Transformer For Speech Emotion Recognition
von: Li, Zhipeng, et al.
Veröffentlicht: (2024) -
Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
von: Chen, Weidong, et al.
Veröffentlicht: (2023) -
LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space
von: Xin, Detai, et al.
Veröffentlicht: (2026)