Lina-Speech: Gated Linear Attention and Initial-State Tuning for Multi-Sample Prompting Text-To-Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lemerle, Théodor, Guichoux, Téo, Roebel, Axel, Obin, Nicolas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024)
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
von: Guichoux, Téo, et al.
Veröffentlicht: (2025)
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025)
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025)
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
von: Guichoux, Téo, et al.
Veröffentlicht: (2024)
von: Guichoux, Téo, et al.
Veröffentlicht: (2024)
FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
von: Seki, Kentaro, et al.
Veröffentlicht: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
MusicGen-Stem: Multi-stem music generation and edition through autoregressive modeling
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
PitchFlower: A flow-based neural audio codec with pitch controllability
von: Torres, Diego, et al.
Veröffentlicht: (2025)
von: Torres, Diego, et al.
Veröffentlicht: (2025)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
von: Leung, Wing-Zin, et al.
Veröffentlicht: (2024)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
von: Chen, Zhengyang, et al.
Veröffentlicht: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
Continuous Audio Language Models
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
von: Rouard, Simon, et al.
Veröffentlicht: (2025)
Retrieval Augmented Generation in Prompt-based Text-to-Speech Synthesis with Context-Aware Contrastive Language-Audio Pretraining
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
von: Xue, Jinlong, et al.
Veröffentlicht: (2024)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Evaluating Text-to-Speech Synthesis from a Large Discrete Token-based Speech Language Model
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
von: Wang, Siyang, et al.
Veröffentlicht: (2024)
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
von: Li, Hanzhao, et al.
Veröffentlicht: (2025)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
von: Jiang, Ziyue, et al.
Veröffentlicht: (2023)
Comprehend and Talk: Text to Speech Synthesis via Dual Language Modeling
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
von: Cao, Junjie, et al.
Veröffentlicht: (2025)
Enabling Beam Search for Language Model-Based Text-to-Speech Synthesis
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
von: Tu, Zehai, et al.
Veröffentlicht: (2024)
CLEP-DG: Contrastive Learning for Speech Emotion Domain Generalization via Soft Prompt Tuning
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2025)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
von: Ma, Ding, et al.
Veröffentlicht: (2026)
von: Ma, Ding, et al.
Veröffentlicht: (2026)
Outlier Reduction with Gated Attention for Improved Post-training Quantization in Large Sequence-to-sequence Speech Foundation Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Unified Pathological Speech Analysis with Prompt Tuning
von: Yang, Fei, et al.
Veröffentlicht: (2024)
von: Yang, Fei, et al.
Veröffentlicht: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
von: Ahmad, Hawraz A., et al.
Veröffentlicht: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
von: Lemerle, Théodor, et al.
Veröffentlicht: (2024) -
Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
von: Guichoux, Téo, et al.
Veröffentlicht: (2025) -
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025) -
2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
von: Guichoux, Téo, et al.
Veröffentlicht: (2024) -
FLASepformer: Efficient Speech Separation with Gated Focused Linear Attention Transformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2025)