MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ristori, Eleonora, Bindini, Luca, Frasconi, Paolo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918450208702464
author Ristori, Eleonora
Bindini, Luca
Frasconi, Paolo
author_facet Ristori, Eleonora
Bindini, Luca
Frasconi, Paolo
contents Research on audio generation has progressively developed along both waveform-based and spectrogram-based directions, giving rise to diverse strategies for representing and generating audio. At the same time, advances in image synthesis have shown that autoregression across scales, rather than tokens, improves coherence and detail. Building on these ideas, we introduce MARS (Multi-channel AutoRegression on Spectrograms), which, to the best of our knowledge, is the first adaptation of next-scale autoregressive modeling to the spectrogram domain. MARS treats spectrograms as multi-channel images and employs channel multiplexing (CMX), a reshaping strategy that reduces spatial resolution without information loss. A shared tokenizer provides consistent discrete representations across scales, enabling a transformer-based autoregressor to refine spectrograms from coarse to fine resolutions efficiently. Experiments on a large-scale dataset demonstrate that MARS performs comparably or better than state-of-the-art baselines across multiple evaluation metrics, establishing an efficient and scalable paradigm for high-fidelity sound generation.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26007
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
Ristori, Eleonora
Bindini, Luca
Frasconi, Paolo
Sound
Artificial Intelligence
Machine Learning
Research on audio generation has progressively developed along both waveform-based and spectrogram-based directions, giving rise to diverse strategies for representing and generating audio. At the same time, advances in image synthesis have shown that autoregression across scales, rather than tokens, improves coherence and detail. Building on these ideas, we introduce MARS (Multi-channel AutoRegression on Spectrograms), which, to the best of our knowledge, is the first adaptation of next-scale autoregressive modeling to the spectrogram domain. MARS treats spectrograms as multi-channel images and employs channel multiplexing (CMX), a reshaping strategy that reduces spatial resolution without information loss. A shared tokenizer provides consistent discrete representations across scales, enabling a transformer-based autoregressor to refine spectrograms from coarse to fine resolutions efficiently. Experiments on a large-scale dataset demonstrate that MARS performs comparably or better than state-of-the-art baselines across multiple evaluation metrics, establishing an efficient and scalable paradigm for high-fidelity sound generation.
title MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
topic Sound
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2509.26007