Long-Context Speech Synthesis with Context-Aware Memory
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Zhipeng, Xing, Xiaofen, Xing, Jingyuan, Hu, Hangrui, Lu, Heng, Xu, Xiangmin |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
par: Xing, Jingyuan, et autres
Publié: (2025)
par: Xing, Jingyuan, et autres
Publié: (2025)
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
par: Nie, Sihang, et autres
Publié: (2025)
par: Nie, Sihang, et autres
Publié: (2025)
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
par: Mai, Jialong, et autres
Publié: (2025)
par: Mai, Jialong, et autres
Publié: (2025)
Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
par: Chen, Weidong, et autres
Publié: (2023)
par: Chen, Weidong, et autres
Publié: (2023)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
par: Fang, Yuanbo, et autres
Publié: (2025)
par: Fang, Yuanbo, et autres
Publié: (2025)
AS-Speech: Adaptive Style For Speech Synthesis
par: Li, Zhipeng, et autres
Publié: (2024)
par: Li, Zhipeng, et autres
Publié: (2024)
Seeing the Context: Rich Visual Context-Aware Speech Recognition via Multimodal Reasoning
par: Tian, Wenjie, et autres
Publié: (2026)
par: Tian, Wenjie, et autres
Publié: (2026)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
par: Yen, Hao, et autres
Publié: (2024)
par: Yen, Hao, et autres
Publié: (2024)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
par: Xue, Jinlong, et autres
Publié: (2024)
par: Xue, Jinlong, et autres
Publié: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
par: Gao, Xiaoxue, et autres
Publié: (2024)
par: Gao, Xiaoxue, et autres
Publié: (2024)
Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
par: Bai, Ye, et autres
Publié: (2024)
par: Bai, Ye, et autres
Publié: (2024)
Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation
par: Wang, Wupeng, et autres
Publié: (2025)
par: Wang, Wupeng, et autres
Publié: (2025)
Beyond the Utterance: An Empirical Study of Very Long Context Speech Recognition
par: Flynn, Robert, et autres
Publié: (2026)
par: Flynn, Robert, et autres
Publié: (2026)
Multi-Scale Temporal Transformer For Speech Emotion Recognition
par: Li, Zhipeng, et autres
Publié: (2024)
par: Li, Zhipeng, et autres
Publié: (2024)
SoCov: Semi-Orthogonal Parametric Pooling of Covariance Matrix for Speaker Recognition
par: Li, Rongjin, et autres
Publié: (2025)
par: Li, Rongjin, et autres
Publié: (2025)
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
par: Yu, Fan, et autres
Publié: (2025)
par: Yu, Fan, et autres
Publié: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
par: Liao, Shijia, et autres
Publié: (2024)
par: Liao, Shijia, et autres
Publié: (2024)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
par: Yu, Fan, et autres
Publié: (2024)
par: Yu, Fan, et autres
Publié: (2024)
Rate-Aware Learned Speech Compression
par: Xu, Jun, et autres
Publié: (2025)
par: Xu, Jun, et autres
Publié: (2025)
STSR: High-Fidelity Speech Super-Resolution via Spectral-Transient Context Modeling
par: Yuan, Jiajun, et autres
Publié: (2025)
par: Yuan, Jiajun, et autres
Publié: (2025)
End-to-End Target Speaker Speech Recognition Using Context-Aware Attention Mechanisms for Challenging Enrollment Scenario
par: Ghane, Mohsen, et autres
Publié: (2025)
par: Ghane, Mohsen, et autres
Publié: (2025)
Exploring In-Context Learning Capabilities of ChatGPT for Pathological Speech Detection
par: Amiri, Mahdi, et autres
Publié: (2025)
par: Amiri, Mahdi, et autres
Publié: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
par: Lu, Ye-Xin, et autres
Publié: (2025)
par: Lu, Ye-Xin, et autres
Publié: (2025)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
par: Zhao, Shengkui, et autres
Publié: (2025)
par: Zhao, Shengkui, et autres
Publié: (2025)
Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
par: Chung, Soo-Whan, et autres
Publié: (2025)
par: Chung, Soo-Whan, et autres
Publié: (2025)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
par: Li, Haoxun, et autres
Publié: (2025)
par: Li, Haoxun, et autres
Publié: (2025)
FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
par: Xie, Kun, et autres
Publié: (2025)
par: Xie, Kun, et autres
Publié: (2025)
Intra- and Inter-modal Context Interaction Modeling for Conversational Speech Synthesis
par: Jia, Zhenqi, et autres
Publié: (2024)
par: Jia, Zhenqi, et autres
Publié: (2024)
Improving Streaming Speech Recognition With Time-Shifted Contextual Attention And Dynamic Right Context Masking
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation
par: Huang, Ziling, et autres
Publié: (2025)
par: Huang, Ziling, et autres
Publié: (2025)
Spectral Masking with Explicit Time-Context Windowing for Neural Network-Based Monaural Speech Enhancement
par: Fiorio, Luan Vinícius, et autres
Publié: (2024)
par: Fiorio, Luan Vinícius, et autres
Publié: (2024)
Towards a Generalizable Speech Marker for Parkinson's Disease Diagnosis
par: Siniukov, Maksim, et autres
Publié: (2025)
par: Siniukov, Maksim, et autres
Publié: (2025)
Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
par: Yu, Chin-Yun, et autres
Publié: (2024)
par: Yu, Chin-Yun, et autres
Publié: (2024)
Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
par: Ao, Junyi, et autres
Publié: (2025)
par: Ao, Junyi, et autres
Publié: (2025)
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
par: Jung, Jaemin, et autres
Publié: (2024)
par: Jung, Jaemin, et autres
Publié: (2024)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
par: Hu, Yifan, et autres
Publié: (2025)
par: Hu, Yifan, et autres
Publié: (2025)
Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
par: Sato, Ryo, et autres
Publié: (2025)
par: Sato, Ryo, et autres
Publié: (2025)
Noise-Aware Speech Separation with Contrastive Learning
par: Zhang, Zizheng, et autres
Publié: (2023)
par: Zhang, Zizheng, et autres
Publié: (2023)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
par: Cha, Jun-Hyeok, et autres
Publié: (2025)
par: Cha, Jun-Hyeok, et autres
Publié: (2025)
Documents similaires
-
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
par: Xing, Jingyuan, et autres
Publié: (2025) -
HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS
par: Nie, Sihang, et autres
Publié: (2025) -
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
par: Mai, Jialong, et autres
Publié: (2025) -
Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition
par: Chen, Weidong, et autres
Publié: (2023) -
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
par: Fang, Yuanbo, et autres
Publié: (2025)