MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
Fuente:
arXiv
Saved in:
| Main Authors: | Ristori, Eleonora, Bindini, Luca, Frasconi, Paolo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dealing with Uncertainty in Contextual Anomaly Detection
by: Bindini, Luca, et al.
Published: (2025)
by: Bindini, Luca, et al.
Published: (2025)
Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
by: Işık, Atakan, et al.
Published: (2025)
by: Işık, Atakan, et al.
Published: (2025)
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
by: Zhang, Zihan, et al.
Published: (2025)
by: Zhang, Zihan, et al.
Published: (2025)
Woosh: A Sound Effects Foundation Model
by: Hadjeres, Gaëtan, et al.
Published: (2026)
by: Hadjeres, Gaëtan, et al.
Published: (2026)
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)
by: Mancini, Eleonora, et al.
Published: (2025)
Steering Autoregressive Music Generation with Recursive Feature Machines
by: Zhao, Daniel, et al.
Published: (2025)
by: Zhao, Daniel, et al.
Published: (2025)
Masked Audio Generation using a Single Non-Autoregressive Transformer
by: Ziv, Alon, et al.
Published: (2024)
by: Ziv, Alon, et al.
Published: (2024)
Investigating the Effectiveness of Explainability Methods in Parkinson's Detection from Speech
by: Mancini, Eleonora, et al.
Published: (2024)
by: Mancini, Eleonora, et al.
Published: (2024)
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
by: Lu, Yuxin, et al.
Published: (2026)
by: Lu, Yuxin, et al.
Published: (2026)
Beyond Fixed Frames: Dynamic Character-Aligned Speech Tokenization
by: Della Libera, Luca, et al.
Published: (2026)
by: Della Libera, Luca, et al.
Published: (2026)
Exploring Token-Space Manipulation in Latent Audio Tokenizers
by: Paissan, Francesco, et al.
Published: (2026)
by: Paissan, Francesco, et al.
Published: (2026)
Cross-Attention with Confidence Weighting for Multi-Channel Audio Alignment
by: Nihal, Ragib Amin, et al.
Published: (2025)
by: Nihal, Ragib Amin, et al.
Published: (2025)
FastAST: Accelerating Audio Spectrogram Transformer via Token Merging and Cross-Model Knowledge Distillation
by: Behera, Swarup Ranjan, et al.
Published: (2024)
by: Behera, Swarup Ranjan, et al.
Published: (2024)
Epic-Sounds: A Large-scale Dataset of Actions That Sound
by: Huh, Jaesung, et al.
Published: (2023)
by: Huh, Jaesung, et al.
Published: (2023)
Multi-Task Learning for Lung sound & Lung disease classification
by: K V, Suma, et al.
Published: (2024)
by: K V, Suma, et al.
Published: (2024)
Survey on the Evaluation of Generative Models in Music
by: Lerch, Alexander, et al.
Published: (2025)
by: Lerch, Alexander, et al.
Published: (2025)
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
by: Della Libera, Luca, et al.
Published: (2026)
by: Della Libera, Luca, et al.
Published: (2026)
Continuous Autoregressive Models with Noise Augmentation Avoid Error Accumulation
by: Pasini, Marco, et al.
Published: (2024)
by: Pasini, Marco, et al.
Published: (2024)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
by: Jia, Dongya, et al.
Published: (2025)
by: Jia, Dongya, et al.
Published: (2025)
Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
by: Marincione, Davide, et al.
Published: (2025)
by: Marincione, Davide, et al.
Published: (2025)
Explainable Multi-Modal Deep Learning for Automatic Detection of Lung Diseases from Respiratory Audio Signals
by: Saky, S M Asiful Islam, et al.
Published: (2025)
by: Saky, S M Asiful Islam, et al.
Published: (2025)
DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
by: Novack, Zachary, et al.
Published: (2024)
by: Novack, Zachary, et al.
Published: (2024)
Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation
by: Si, Chen, et al.
Published: (2025)
by: Si, Chen, et al.
Published: (2025)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
by: Lee, Junwon, et al.
Published: (2024)
by: Lee, Junwon, et al.
Published: (2024)
Of All StrIPEs: Investigating Structure-informed Positional Encoding for Efficient Music Generation
by: Agarwal, Manvi, et al.
Published: (2025)
by: Agarwal, Manvi, et al.
Published: (2025)
Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
by: Pasini, Marco, et al.
Published: (2025)
by: Pasini, Marco, et al.
Published: (2025)
Fusing Memory and Attention: A study on LSTM, Transformer and Hybrid Architectures for Symbolic Music Generation
by: Ghoshal, Soudeep, et al.
Published: (2026)
by: Ghoshal, Soudeep, et al.
Published: (2026)
QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
by: Wang, Chien-Chun, et al.
Published: (2025)
by: Wang, Chien-Chun, et al.
Published: (2025)
Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic Rewards
by: Fang, Linghan, et al.
Published: (2026)
by: Fang, Linghan, et al.
Published: (2026)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
by: Bruford, Fred, et al.
Published: (2024)
by: Bruford, Fred, et al.
Published: (2024)
H-Infinity Filter Enhanced CNN-LSTM for Arrhythmia Detection from Heart Sound Recordings
by: Kumar, Rohith Shinoj, et al.
Published: (2025)
by: Kumar, Rohith Shinoj, et al.
Published: (2025)
Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling
by: Choudhary, Yash, et al.
Published: (2025)
by: Choudhary, Yash, et al.
Published: (2025)
Data-Driven Analysis of Text-Conditioned AI-Generated Music: A Case Study with Suno and Udio
by: Casini, Luca, et al.
Published: (2025)
by: Casini, Luca, et al.
Published: (2025)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
by: Della Libera, Luca, et al.
Published: (2025)
by: Della Libera, Luca, et al.
Published: (2025)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
by: Della Libera, Luca, et al.
Published: (2025)
by: Della Libera, Luca, et al.
Published: (2025)
Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
by: Yang, Yiyuan, et al.
Published: (2025)
by: Yang, Yiyuan, et al.
Published: (2025)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
by: Zhang, Shiqi, et al.
Published: (2025)
by: Zhang, Shiqi, et al.
Published: (2025)
Let There Be Sound: Reconstructing High Quality Speech from Silent Videos
by: Kim, Ji-Hoon, et al.
Published: (2023)
by: Kim, Ji-Hoon, et al.
Published: (2023)
FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
by: Fan, Pingyi, et al.
Published: (2025)
by: Fan, Pingyi, et al.
Published: (2025)
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
by: Passoni, Riccardo, et al.
Published: (2025)
by: Passoni, Riccardo, et al.
Published: (2025)
Similar Items
-
Dealing with Uncertainty in Contextual Anomaly Detection
by: Bindini, Luca, et al.
Published: (2025) -
Geometry-Aware Optimization for Respiratory Sound Classification: Enhancing Sensitivity with SAM-Optimized Audio Spectrogram Transformers
by: Işık, Atakan, et al.
Published: (2025) -
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
by: Zhang, Zihan, et al.
Published: (2025) -
Woosh: A Sound Effects Foundation Model
by: Hadjeres, Gaëtan, et al.
Published: (2026) -
Leveraging Whisper Embeddings for Audio-based Lyrics Matching
by: Mancini, Eleonora, et al.
Published: (2025)