Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Ye, Zhen, Zhu, Xinfa, Chan, Chi-Min, Wang, Xinsheng, Tan, Xu, Lei, Jiahe, Peng, Yi, Liu, Haohe, Jin, Yizhu, Dai, Zheqi, Lin, Hongzhan, Chen, Jianyi, Du, Xingjian, Xue, Liumeng, Chen, Yunlin, Li, Zhifei, Xie, Lei, Kong, Qiuqiang, Guo, Yike, Xue, Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
MusicScore: A Dataset for Music Score Modeling and Generation
di: Lin, Yuheng, et al.
Pubblicazione: (2024)
di: Lin, Yuheng, et al.
Pubblicazione: (2024)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
di: Tian, Wenjie, et al.
Pubblicazione: (2025)
AudioX: A Unified Framework for Anything-to-Audio Generation
di: Tian, Zeyue, et al.
Pubblicazione: (2025)
di: Tian, Zeyue, et al.
Pubblicazione: (2025)
Learning Temporal Resolution in Spectrogram for Audio Classification
di: Liu, Haohe, et al.
Pubblicazione: (2022)
di: Liu, Haohe, et al.
Pubblicazione: (2022)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
di: Li, Hanzhao, et al.
Pubblicazione: (2024)
di: Li, Hanzhao, et al.
Pubblicazione: (2024)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
di: Chen, Jianyi, et al.
Pubblicazione: (2024)
di: Chen, Jianyi, et al.
Pubblicazione: (2024)
pTSE-T: Presentation Target Speaker Extraction using Unaligned Text Cues
di: Jiang, Ziyang, et al.
Pubblicazione: (2024)
di: Jiang, Ziyang, et al.
Pubblicazione: (2024)
SponTTS: modeling and transferring spontaneous style for TTS
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
di: Li, Hanzhao, et al.
Pubblicazione: (2023)
Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
di: Ye, Zhen, et al.
Pubblicazione: (2024)
di: Ye, Zhen, et al.
Pubblicazione: (2024)
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
di: Mei, Xinhao, et al.
Pubblicazione: (2023)
di: Mei, Xinhao, et al.
Pubblicazione: (2023)
Audio-FLAN: A Preliminary Release
di: Xue, Liumeng, et al.
Pubblicazione: (2025)
di: Xue, Liumeng, et al.
Pubblicazione: (2025)
Inference-time Scaling for Diffusion-based Audio Super-resolution
di: Jin, Yizhu, et al.
Pubblicazione: (2025)
di: Jin, Yizhu, et al.
Pubblicazione: (2025)
Smart Fitting Room: A One-stop Framework for Matching-aware Virtual Try-on
di: Yu, Mingzhe, et al.
Pubblicazione: (2024)
di: Yu, Mingzhe, et al.
Pubblicazione: (2024)
Separate Anything You Describe
di: Liu, Xubo, et al.
Pubblicazione: (2023)
di: Liu, Xubo, et al.
Pubblicazione: (2023)
FashionDPO:Fine-tune Fashion Outfit Generation Model using Direct Preference Optimization
di: Yu, Mingzhe, et al.
Pubblicazione: (2025)
di: Yu, Mingzhe, et al.
Pubblicazione: (2025)
EVA: An Embodied World Model for Future Video Anticipation
di: Chi, Xiaowei, et al.
Pubblicazione: (2024)
di: Chi, Xiaowei, et al.
Pubblicazione: (2024)
Low-latency Speech Enhancement via Speech Token Generation
di: Xue, Huaying, et al.
Pubblicazione: (2023)
di: Xue, Huaying, et al.
Pubblicazione: (2023)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
di: Liu, Haohe, et al.
Pubblicazione: (2023)
di: Liu, Haohe, et al.
Pubblicazione: (2023)
SLAM-LLM: A Modular, Open-Source Multimodal Large Language Model Framework and Best Practice for Speech, Language, Audio and Music Processing
di: Ma, Ziyang, et al.
Pubblicazione: (2026)
di: Ma, Ziyang, et al.
Pubblicazione: (2026)
Transforming Video Subjective Testing with Training, Engagement, and Real-Time Feedback
di: Rahul, Kumar, et al.
Pubblicazione: (2026)
di: Rahul, Kumar, et al.
Pubblicazione: (2026)
Volume Tracking Based Reference Mesh Extraction for Time-Varying Mesh Compression
di: Chen, Guodong, et al.
Pubblicazione: (2024)
di: Chen, Guodong, et al.
Pubblicazione: (2024)
Efficient Speech Watermarking for Speech Synthesis via Progressive Knowledge Distillation
di: Cui, Yang, et al.
Pubblicazione: (2025)
di: Cui, Yang, et al.
Pubblicazione: (2025)
Real-Time Interactive Hybrid Ocean: Spectrum-Consistent Wave Particle-FFT Coupling
di: Xue, Shengze, et al.
Pubblicazione: (2025)
di: Xue, Shengze, et al.
Pubblicazione: (2025)
Identity-Driven Multimedia Forgery Detection via Reference Assistance
di: Xu, Junhao, et al.
Pubblicazione: (2024)
di: Xu, Junhao, et al.
Pubblicazione: (2024)
DeepStream: Prototyping Deep Joint Source-Channel Coding for Real-Time Multimedia Transmissions
di: Chi, Kaiyi, et al.
Pubblicazione: (2025)
di: Chi, Kaiyi, et al.
Pubblicazione: (2025)
SpaceMeta: Global-Scale Massive Multi-User Virtual Interaction over LEO Satellite Constellations
di: Huang, Jiahe, et al.
Pubblicazione: (2024)
di: Huang, Jiahe, et al.
Pubblicazione: (2024)
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation
di: Wang, Yongqi, et al.
Pubblicazione: (2025)
di: Wang, Yongqi, et al.
Pubblicazione: (2025)
Time-RA: Towards Time Series Reasoning for Anomaly Diagnosis with LLM Feedback
di: Yang, Yiyuan, et al.
Pubblicazione: (2025)
di: Yang, Yiyuan, et al.
Pubblicazione: (2025)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
di: Chen, Siran, et al.
Pubblicazione: (2025)
di: Chen, Siran, et al.
Pubblicazione: (2025)
Can We Hear from Events? Generating Speech from Event Camera
di: Fang, Jingping, et al.
Pubblicazione: (2026)
di: Fang, Jingping, et al.
Pubblicazione: (2026)
Emotion-Driven Personalized Recommendation for AI-Generated Content Using Multi-Modal Sentiment and Intent Analysis
di: Hu, Zheqi, et al.
Pubblicazione: (2025)
di: Hu, Zheqi, et al.
Pubblicazione: (2025)
Encoding Time and Energy Model for SVT-AV1 based on Video Complexity
di: Eichermüller, Lena, et al.
Pubblicazione: (2024)
di: Eichermüller, Lena, et al.
Pubblicazione: (2024)
HeadsetOff: Enabling Photorealistic Video Conferencing on Economical VR Headsets
di: Jin, Yili, et al.
Pubblicazione: (2024)
di: Jin, Yili, et al.
Pubblicazione: (2024)
Calibration & Reconstruction: Deep Integrated Language for Referring Image Segmentation
di: Yan, Yichen, et al.
Pubblicazione: (2024)
di: Yan, Yichen, et al.
Pubblicazione: (2024)
SpeechEE: A Novel Benchmark for Speech Event Extraction
di: Wang, Bin, et al.
Pubblicazione: (2024)
di: Wang, Bin, et al.
Pubblicazione: (2024)
Text-aware and Context-aware Expressive Audiobook Speech Synthesis
di: Guo, Dake, et al.
Pubblicazione: (2024)
di: Guo, Dake, et al.
Pubblicazione: (2024)
Trusted Fake Audio Detection Based on Dirichlet Distribution
di: Ding, Chi, et al.
Pubblicazione: (2025)
di: Ding, Chi, et al.
Pubblicazione: (2025)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
di: Chen, Siran, et al.
Pubblicazione: (2025)
di: Chen, Siran, et al.
Pubblicazione: (2025)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
di: Cui, Meng, et al.
Pubblicazione: (2023)
di: Cui, Meng, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025) -
MusicScore: A Dataset for Music Score Modeling and Generation
di: Lin, Yuheng, et al.
Pubblicazione: (2024) -
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
di: Tian, Wenjie, et al.
Pubblicazione: (2025) -
AudioX: A Unified Framework for Anything-to-Audio Generation
di: Tian, Zeyue, et al.
Pubblicazione: (2025) -
Learning Temporal Resolution in Spectrogram for Audio Classification
di: Liu, Haohe, et al.
Pubblicazione: (2022)