FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Tian-Hao, Zhang, Jiawei, Wang, Jun, Qian, Xinyuan, Yin, Xu-Cheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
di: Zhang, Jiawei, et al.
Pubblicazione: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
di: Deng, Yimin, et al.
Pubblicazione: (2024)
di: Deng, Yimin, et al.
Pubblicazione: (2024)
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
di: Kang, Wonjune, et al.
Pubblicazione: (2025)
di: Kang, Wonjune, et al.
Pubblicazione: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
di: Guan, Wenhao, et al.
Pubblicazione: (2023)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024)
di: Yang, Qian, et al.
Pubblicazione: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
di: Li, Yinghao Aaron, et al.
Pubblicazione: (2024)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
di: Lou, Haowei, et al.
Pubblicazione: (2025)
di: Lou, Haowei, et al.
Pubblicazione: (2025)
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
di: Kondo, Yuto, et al.
Pubblicazione: (2025)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
di: Narain, Jaya, et al.
Pubblicazione: (2025)
di: Narain, Jaya, et al.
Pubblicazione: (2025)
Generative Expressive Conversational Speech Synthesis
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
di: Chen, Weidong, et al.
Pubblicazione: (2025)
di: Chen, Weidong, et al.
Pubblicazione: (2025)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
di: Jawaid, Ahad, et al.
Pubblicazione: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
di: Guo, Yinlin, et al.
Pubblicazione: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
di: Zhu, Xinfa, et al.
Pubblicazione: (2023)
di: Zhu, Xinfa, et al.
Pubblicazione: (2023)
Joint Multi-scale Cross-lingual Speaking Style Transfer with Bidirectional Attention Mechanism for Automatic Dubbing
di: Li, Jingbei, et al.
Pubblicazione: (2023)
di: Li, Jingbei, et al.
Pubblicazione: (2023)
Breaking Through the Spike: Spike Window Decoding for Accelerated and Precise Automatic Speech Recognition
di: Zhang, Wei, et al.
Pubblicazione: (2025)
di: Zhang, Wei, et al.
Pubblicazione: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
di: Kim, Nam-Gyu, et al.
Pubblicazione: (2025)
di: Kim, Nam-Gyu, et al.
Pubblicazione: (2025)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
di: Shi, Haoxiang, et al.
Pubblicazione: (2024)
di: Shi, Haoxiang, et al.
Pubblicazione: (2024)
VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
di: Zhou, Yixuan, et al.
Pubblicazione: (2024)
Mamba in Speech: Towards an Alternative to Self-Attention
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
di: Zhang, Xiangyu, et al.
Pubblicazione: (2024)
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
di: Anastassiou, Philip, et al.
Pubblicazione: (2024)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
di: Chung, Raymond
Pubblicazione: (2026)
di: Chung, Raymond
Pubblicazione: (2026)
Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
di: Zhou, Xukun, et al.
Pubblicazione: (2024)
di: Zhou, Xukun, et al.
Pubblicazione: (2024)
Speech Synthesis along Perceptual Voice Quality Dimensions
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
di: Lou, Haowei, et al.
Pubblicazione: (2026)
di: Lou, Haowei, et al.
Pubblicazione: (2026)
MambaRate: Speech Quality Assessment Across Different Sampling Rates
di: Kakoulidis, Panos, et al.
Pubblicazione: (2025)
di: Kakoulidis, Panos, et al.
Pubblicazione: (2025)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
di: Dutta, Soumya, et al.
Pubblicazione: (2025)
di: Dutta, Soumya, et al.
Pubblicazione: (2025)
Improving Speech Enhancement with Multi-Metric Supervision from Learned Quality Assessment
di: Wang, Wei, et al.
Pubblicazione: (2025)
di: Wang, Wei, et al.
Pubblicazione: (2025)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
di: Liao, Shijia, et al.
Pubblicazione: (2024)
di: Liao, Shijia, et al.
Pubblicazione: (2024)
CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness Condition
di: Wang, Jianzong, et al.
Pubblicazione: (2024)
di: Wang, Jianzong, et al.
Pubblicazione: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
Enhancing Expressiveness in Dance Generation via Integrating Frequency and Music Style Information
di: Huang, Qiaochu, et al.
Pubblicazione: (2024)
di: Huang, Qiaochu, et al.
Pubblicazione: (2024)
Generative Speech Foundation Model Pretraining for High-Quality Speech Extraction and Restoration
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
di: Ku, Pin-Jui, et al.
Pubblicazione: (2024)
Factor-Conditioned Speaking-Style Captioning
di: Ando, Atsushi, et al.
Pubblicazione: (2024)
di: Ando, Atsushi, et al.
Pubblicazione: (2024)
Vision-Integrated High-Quality Neural Speech Coding
di: Guo, Yao, et al.
Pubblicazione: (2025)
di: Guo, Yao, et al.
Pubblicazione: (2025)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
di: Jiang, Ziyue, et al.
Pubblicazione: (2023)
DurIAN-E 2: Duration Informed Attention Network with Adaptive Variational Autoencoder and Adversarial Learning for Expressive Text-to-Speech Synthesis
di: Gu, Yu, et al.
Pubblicazione: (2024)
di: Gu, Yu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception
di: Zhang, Jiawei, et al.
Pubblicazione: (2024) -
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
di: Deng, Yimin, et al.
Pubblicazione: (2024) -
Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style
di: Kang, Wonjune, et al.
Pubblicazione: (2025) -
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
di: Guan, Wenhao, et al.
Pubblicazione: (2023) -
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
di: Yang, Qian, et al.
Pubblicazione: (2024)