SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Jianyi, Zhong, Rongxiu, Zhang, Shilei, Qian, Kun, Liu, Jinglei, Guo, Yike, Xue, Wei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915880868249600
author Chen, Jianyi
Zhong, Rongxiu
Zhang, Shilei
Qian, Kun
Liu, Jinglei
Guo, Yike
Xue, Wei
author_facet Chen, Jianyi
Zhong, Rongxiu
Zhang, Shilei
Qian, Kun
Liu, Jinglei
Guo, Yike
Xue, Wei
contents Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this work, we propose a simple yet powerful trick: we assume that AI models can understand and generate time-accelerated (speeded-up) audio at rates such as 2x, 4x, or even 8x. By first generating a high-speed version of the music, we greatly reduce the temporal length and resource requirements, making it feasible to handle long-form music that would otherwise exceed memory or computational limits. The generated audio is then restored to its original speed, recovering the full temporal structure. This temporal speed-up and slow-down strategy naturally follows the principle of hierarchical generation from abstract to detailed content, and can be conveniently applied to existing music generation models to enable long-form music generation. We instantiate this idea in SqueezeComposer, a framework that employs diffusion models for generation in the accelerated domain and refinement in the restored domain. We validate the effectiveness of this approach on two tasks: long-form music generation, which evaluates temporal-wise control (including continuation, completion, and generation from scratch), and whole-song singing accompaniment generation, which evaluates track-wise control. Experimental results demonstrate that our simple temporal speed-up trick enables efficient, scalable, and high-quality long-form music generation. Audio samples are available at https://SqueezeComposer.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21073
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing
Chen, Jianyi
Zhong, Rongxiu
Zhang, Shilei
Qian, Kun
Liu, Jinglei
Guo, Yike
Xue, Wei
Audio and Speech Processing
Computation and Language
Sound
Composing coherent long-form music remains a significant challenge due to the complexity of modeling long-range dependencies and the prohibitive memory and computational requirements associated with lengthy audio representations. In this work, we propose a simple yet powerful trick: we assume that AI models can understand and generate time-accelerated (speeded-up) audio at rates such as 2x, 4x, or even 8x. By first generating a high-speed version of the music, we greatly reduce the temporal length and resource requirements, making it feasible to handle long-form music that would otherwise exceed memory or computational limits. The generated audio is then restored to its original speed, recovering the full temporal structure. This temporal speed-up and slow-down strategy naturally follows the principle of hierarchical generation from abstract to detailed content, and can be conveniently applied to existing music generation models to enable long-form music generation. We instantiate this idea in SqueezeComposer, a framework that employs diffusion models for generation in the accelerated domain and refinement in the restored domain. We validate the effectiveness of this approach on two tasks: long-form music generation, which evaluates temporal-wise control (including continuation, completion, and generation from scratch), and whole-song singing accompaniment generation, which evaluates track-wise control. Experimental results demonstrate that our simple temporal speed-up trick enables efficient, scalable, and high-quality long-form music generation. Audio samples are available at https://SqueezeComposer.github.io/.
title SqueezeComposer: Temporal Speed-up is A Simple Trick for Long-form Music Composing
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2603.21073