MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zihan, Cheng, Xize, Jiang, Zhennan, Fu, Dongjie, Chen, Jingyuan, Zhao, Zhou, Jin, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
by: Ristori, Eleonora, et al.
Published: (2025)
by: Ristori, Eleonora, et al.
Published: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024)
by: Cheng, Xize, et al.
Published: (2024)
CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
by: Han, Runduo, et al.
Published: (2025)
by: Han, Runduo, et al.
Published: (2025)
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
by: Chen, Jun, et al.
Published: (2025)
by: Chen, Jun, et al.
Published: (2025)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
by: Li, Yuqi, et al.
Published: (2025)
by: Li, Yuqi, et al.
Published: (2025)
Interpretable All-Type Audio Deepfake Detection with Audio LLMs via Frequency-Time Reinforcement Learning
by: Xie, Yuankun, et al.
Published: (2026)
by: Xie, Yuankun, et al.
Published: (2026)
ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
by: Zhao, Junqi, et al.
Published: (2024)
by: Zhao, Junqi, et al.
Published: (2024)
Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
by: Cheng, Xize, et al.
Published: (2025)
by: Cheng, Xize, et al.
Published: (2025)
OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
Analytic Incremental Learning For Sound Source Localization With Imbalance Rectification
by: Fan, Zexia, et al.
Published: (2026)
by: Fan, Zexia, et al.
Published: (2026)
SLM-SS: Speech Language Model for Generative Speech Separation
by: Li, Tianhua, et al.
Published: (2026)
by: Li, Tianhua, et al.
Published: (2026)
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
by: Yang, Dongchao, et al.
Published: (2025)
by: Yang, Dongchao, et al.
Published: (2025)
AudioCapBench: Quick Evaluation on Audio Captioning across Sound, Music, and Speech
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought
by: Li, Xuanchen, et al.
Published: (2026)
by: Li, Xuanchen, et al.
Published: (2026)
Speech-Audio Compositional Attacks on Multimodal LLMs and Their Mitigation with SALMONN-Guard
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
by: Qiang, Chunyu, et al.
Published: (2025)
by: Qiang, Chunyu, et al.
Published: (2025)
SAME: A Semantically-Aligned Music Autoencoder
by: Parker, Julian D., et al.
Published: (2026)
by: Parker, Julian D., et al.
Published: (2026)
SpeakerLM: End-to-End Versatile Speaker Diarization and Recognition with Multimodal Large Language Models
by: Yin, Han, et al.
Published: (2025)
by: Yin, Han, et al.
Published: (2025)
Towards Open World Sound Event Detection
by: Hai, P. H., et al.
Published: (2026)
by: Hai, P. H., et al.
Published: (2026)
Speech Separation for Hearing-Impaired Children in the Classroom
by: Olalere, Feyisayo, et al.
Published: (2025)
by: Olalere, Feyisayo, et al.
Published: (2025)
'Studies for': A Human-AI Co-Creative Sound Artwork Using a Real-time Multi-channel Sound Generation Model
by: Nagashima, Chihiro, et al.
Published: (2025)
by: Nagashima, Chihiro, et al.
Published: (2025)
Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
by: Bjare, Mathias Rose, et al.
Published: (2025)
by: Bjare, Mathias Rose, et al.
Published: (2025)
Back to Ear: Perceptually Driven High Fidelity Music Reconstruction
by: Wang, Kangdi, et al.
Published: (2025)
by: Wang, Kangdi, et al.
Published: (2025)
The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
by: Zhang, Ruixing, et al.
Published: (2026)
by: Zhang, Ruixing, et al.
Published: (2026)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
by: Wang, Jingyuan, et al.
Published: (2024)
by: Wang, Jingyuan, et al.
Published: (2024)
TopSeg: A Multi-Scale Topological Framework for Data-Efficient Heart Sound Segmentation
by: Zhang, Peihong, et al.
Published: (2025)
by: Zhang, Peihong, et al.
Published: (2025)
Environmental Sound Deepfake Detection Using Deep-Learning Framework
by: Pham, Lam, et al.
Published: (2026)
by: Pham, Lam, et al.
Published: (2026)
Dynamic Fusion Multimodal Network for SpeechWellness Detection
by: Sun, Wenqiang, et al.
Published: (2025)
by: Sun, Wenqiang, et al.
Published: (2025)
Structure-Aware Piano Accompaniment via Style Planning and Dataset-Aligned Pattern Retrieval
by: Zang, Wanyu, et al.
Published: (2026)
by: Zang, Wanyu, et al.
Published: (2026)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
by: Xiong, Chenxu, et al.
Published: (2024)
by: Xiong, Chenxu, et al.
Published: (2024)
AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan
by: Xie, Yuankun, et al.
Published: (2026)
by: Xie, Yuankun, et al.
Published: (2026)
Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
by: Shibata, Yuto, et al.
Published: (2025)
by: Shibata, Yuto, et al.
Published: (2025)
One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
by: Stylianou, Ioannis, et al.
Published: (2026)
by: Stylianou, Ioannis, et al.
Published: (2026)
Robust Target Speaker Diarization and Separation via Augmented Speaker Embedding Sampling
by: Jalal, Md Asif, et al.
Published: (2025)
by: Jalal, Md Asif, et al.
Published: (2025)
Sub-Band Spectral Matching with Localized Score Aggregation for Robust Anomalous Sound Detection
by: Saengthong, Phurich, et al.
Published: (2026)
by: Saengthong, Phurich, et al.
Published: (2026)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
by: Cai, Pengfei, et al.
Published: (2025)
by: Cai, Pengfei, et al.
Published: (2025)
TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
by: Yang, Xiaoda, et al.
Published: (2026)
by: Yang, Xiaoda, et al.
Published: (2026)
DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic Gating Fusion
by: Yu, Yinfeng, et al.
Published: (2025)
by: Yu, Yinfeng, et al.
Published: (2025)
Generating Separated Singing Vocals Using a Diffusion Model Conditioned on Music Mixtures
by: Plaja-Roglans, Genís, et al.
Published: (2025)
by: Plaja-Roglans, Genís, et al.
Published: (2025)
Similar Items
-
MARS: Sound Generation via Multi-Channel Autoregression on Spectrograms
by: Ristori, Eleonora, et al.
Published: (2025) -
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
by: Cheng, Xize, et al.
Published: (2024) -
CabinSep: IR-Augmented Mask-Based MVDR for Real-Time In-Car Speech Separation with Distributed Heterogeneous Arrays
by: Han, Runduo, et al.
Published: (2025) -
LSZone: A Lightweight Spatial Information Modeling Architecture for Real-time In-car Multi-zone Speech Separation
by: Chen, Jun, et al.
Published: (2025) -
SepPrune: Structured Pruning for Efficient Deep Speech Separation
by: Li, Yuqi, et al.
Published: (2025)