HiStyle: Hierarchical Style Embedding Predictor for Text-Prompt-Guided Controllable Speech Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Ziyu, Li, Hanzhao, Hu, Jingbin, Li, Wenhao, Xie, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
by: Li, Haitao, et al.
Published: (2026)
by: Li, Haitao, et al.
Published: (2026)
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
by: Liu, Jiaxuan, et al.
Published: (2024)
by: Liu, Jiaxuan, et al.
Published: (2024)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
by: Luo, Dan, et al.
Published: (2025)
by: Luo, Dan, et al.
Published: (2025)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
by: Li, Hanzhao, et al.
Published: (2025)
by: Li, Hanzhao, et al.
Published: (2025)
Scaling Rich Style-Prompted Text-to-Speech Datasets
by: Diwan, Anuj, et al.
Published: (2025)
by: Diwan, Anuj, et al.
Published: (2025)
Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
by: Kim, Nam-Gyu
Published: (2025)
by: Kim, Nam-Gyu
Published: (2025)
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
by: Kang, Jaehoon, et al.
Published: (2026)
by: Kang, Jaehoon, et al.
Published: (2026)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
by: Lou, Haowei, et al.
Published: (2024)
by: Lou, Haowei, et al.
Published: (2024)
StyleGuard: Preventing Text-to-Image-Model-based Style Mimicry Attacks by Style Perturbations
by: Li, Yanjie, et al.
Published: (2025)
by: Li, Yanjie, et al.
Published: (2025)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
by: Xiong, Chenxu, et al.
Published: (2024)
by: Xiong, Chenxu, et al.
Published: (2024)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
by: Kim, Nam-Gyu, et al.
Published: (2025)
by: Kim, Nam-Gyu, et al.
Published: (2025)
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
by: Li, Yinghao Aaron, et al.
Published: (2024)
by: Li, Yinghao Aaron, et al.
Published: (2024)
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation
by: Wu, Yi, et al.
Published: (2025)
by: Wu, Yi, et al.
Published: (2025)
StyleForge: Enhancing Text-to-Image Synthesis for Any Artistic Styles with Dual Binding
by: Park, Junseo, et al.
Published: (2024)
by: Park, Junseo, et al.
Published: (2024)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
by: Wang, Helin, et al.
Published: (2025)
by: Wang, Helin, et al.
Published: (2025)
HiCAST: Highly Customized Arbitrary Style Transfer with Adapter Enhanced Diffusion Models
by: Wang, Hanzhang, et al.
Published: (2024)
by: Wang, Hanzhang, et al.
Published: (2024)
Route-Induced Density and Stability (RIDE): Controlled Intervention and Mechanism Analysis of Routing-Style Meta Prompts on LLM Internal States
by: Zhang, Dianxing, et al.
Published: (2026)
by: Zhang, Dianxing, et al.
Published: (2026)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
by: Qi, Xin, et al.
Published: (2024)
by: Qi, Xin, et al.
Published: (2024)
ArtCrafter: Text-Image Aligning Style Transfer via Embedding Reframing
by: Huang, Nisha, et al.
Published: (2025)
by: Huang, Nisha, et al.
Published: (2025)
StyleVAR: Controllable Image Style Transfer via Visual Autoregressive Modeling
by: Jing, Liqi, et al.
Published: (2026)
by: Jing, Liqi, et al.
Published: (2026)
StyleCrafter: Enhancing Stylized Text-to-Video Generation with Style Adapter
by: Liu, Gongye, et al.
Published: (2023)
by: Liu, Gongye, et al.
Published: (2023)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
by: Wang, Yongqi, et al.
Published: (2023)
by: Wang, Yongqi, et al.
Published: (2023)
Unsupervised Text Style Transfer via LLMs and Attention Masking with Multi-way Interactions
by: Pan, Lei, et al.
Published: (2024)
by: Pan, Lei, et al.
Published: (2024)
StyleLoco: Generative Adversarial Distillation for Natural Humanoid Robot Locomotion
by: Ma, Le, et al.
Published: (2025)
by: Ma, Le, et al.
Published: (2025)
PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
by: Han, Tianshun, et al.
Published: (2025)
by: Han, Tianshun, et al.
Published: (2025)
FonTS: Text Rendering with Typography and Style Controls
by: Shi, Wenda, et al.
Published: (2024)
by: Shi, Wenda, et al.
Published: (2024)
Causality Guided Representation Learning for Cross-Style Hate Speech Detection
by: Zhao, Chengshuai, et al.
Published: (2025)
by: Zhao, Chengshuai, et al.
Published: (2025)
Distilling Text Style Transfer With Self-Explanation From LLMs
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
StyleRec: A Benchmark Dataset for Prompt Recovery in Writing Style Transformation
by: Liu, Shenyang, et al.
Published: (2025)
by: Liu, Shenyang, et al.
Published: (2025)
Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
by: Wu, Zhe, et al.
Published: (2025)
by: Wu, Zhe, et al.
Published: (2025)
Mitigating Intra-Speaker Variability in Diarization with Style-Controllable Speech Augmentation
by: Kim, Miseul, et al.
Published: (2025)
by: Kim, Miseul, et al.
Published: (2025)
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
by: Park, Hyun Joon, et al.
Published: (2024)
by: Park, Hyun Joon, et al.
Published: (2024)
Unsupervised Disentanglement of Content and Style via Variance-Invariance Constraints
by: Wu, Yuxuan, et al.
Published: (2024)
by: Wu, Yuxuan, et al.
Published: (2024)
StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles
by: Zhang, Tian-Hao, et al.
Published: (2025)
by: Zhang, Tian-Hao, et al.
Published: (2025)
Measuring Embedding Sensitivity to Authorial Style in French: Comparing Literary Texts with Language Model Rewritings
by: Icard, Benjamin, et al.
Published: (2026)
by: Icard, Benjamin, et al.
Published: (2026)
HiNet: Novel Multi-Scenario & Multi-Task Learning with Hierarchical Information Extraction
by: Zhou, Jie, et al.
Published: (2023)
by: Zhou, Jie, et al.
Published: (2023)
Style-Preserving Policy Optimization for Game Agents
by: Li, Lingfeng, et al.
Published: (2025)
by: Li, Lingfeng, et al.
Published: (2025)
StyleMamba : State Space Model for Efficient Text-driven Image Style Transfer
by: Wang, Zijia, et al.
Published: (2024)
by: Wang, Zijia, et al.
Published: (2024)
StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer
by: Zheng, Guantian
Published: (2026)
by: Zheng, Guantian
Published: (2026)
Similar Items
-
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
by: Li, Haitao, et al.
Published: (2026) -
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
by: Liu, Jiaxuan, et al.
Published: (2024) -
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
by: Luo, Dan, et al.
Published: (2025) -
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
by: Li, Hanzhao, et al.
Published: (2025) -
Scaling Rich Style-Prompted Text-to-Speech Datasets
by: Diwan, Anuj, et al.
Published: (2025)