Enregistré dans:
Détails bibliographiques
Auteurs principaux: Chen, Jiatao, Tang, Xing, Duan, Xiaoyue, Feng, Yutang, Zhang, Jinchao, Zhou, Jie
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:https://arxiv.org/abs/2602.08233
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914315430264832
author Chen, Jiatao
Tang, Xing
Duan, Xiaoyue
Feng, Yutang
Zhang, Jinchao
Zhou, Jie
author_facet Chen, Jiatao
Tang, Xing
Duan, Xiaoyue
Feng, Yutang
Zhang, Jinchao
Zhou, Jie
contents While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address this, we propose Tutti, a unified framework designed for structured multi-singer generation. Specifically, we introduce a Structure-Aware Singer Prompt to enable flexible singer scheduling evolving with musical structure, and propose Complementary Texture Learning via Condition-Guided VAE to capture implicit acoustic textures (e.g., spatial reverberation and spectral fusion) that are complementary to explicit controls. Experiments demonstrate that Tutti excels in precise multi-singer scheduling and significantly enhances the acoustic realism of choral generation, offering a novel paradigm for complex multi-singer arrangement. Audio samples are available at https://annoauth123-ctrl.github.io/Tutii_Demo/.
format Preprint
id arxiv_https___arxiv_org_abs_2602_08233
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Tutti: Expressive Multi-Singer Synthesis via Structure-Level Timbre Control and Vocal Texture Modeling
Chen, Jiatao
Tang, Xing
Duan, Xiaoyue
Feng, Yutang
Zhang, Jinchao
Zhou, Jie
Sound
Artificial Intelligence
While existing Singing Voice Synthesis systems achieve high-fidelity solo performances, they are constrained by global timbre control, failing to address dynamic multi-singer arrangement and vocal texture within a single song. To address this, we propose Tutti, a unified framework designed for structured multi-singer generation. Specifically, we introduce a Structure-Aware Singer Prompt to enable flexible singer scheduling evolving with musical structure, and propose Complementary Texture Learning via Condition-Guided VAE to capture implicit acoustic textures (e.g., spatial reverberation and spectral fusion) that are complementary to explicit controls. Experiments demonstrate that Tutti excels in precise multi-singer scheduling and significantly enhances the acoustic realism of choral generation, offering a novel paradigm for complex multi-singer arrangement. Audio samples are available at https://annoauth123-ctrl.github.io/Tutii_Demo/.
title Tutti: Expressive Multi-Singer Synthesis via Structure-Level Timbre Control and Vocal Texture Modeling
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2602.08233