SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model
Fuente:
arXiv
Saved in:
| Main Authors: | Shams, Siavash, Dindar, Sukru Samet, Jiang, Xilin, Mesgarani, Nima |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
by: Jiang, Xilin, et al.
Published: (2023)
by: Jiang, Xilin, et al.
Published: (2023)
Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
by: Jiang, Xilin, et al.
Published: (2024)
by: Jiang, Xilin, et al.
Published: (2024)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
by: Jiang, Xilin, et al.
Published: (2024)
by: Jiang, Xilin, et al.
Published: (2024)
MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
by: Shimizu, Riki, et al.
Published: (2025)
by: Shimizu, Riki, et al.
Published: (2025)
Exploring Finetuned Audio-LLM on Heart Murmur Features
by: Florea, Adrian, et al.
Published: (2025)
by: Florea, Adrian, et al.
Published: (2025)
Layer-wise Minimal Pair Probing Reveals Contextual Grammatical-Conceptual Hierarchy in Speech Representations
by: He, Linyang, et al.
Published: (2025)
by: He, Linyang, et al.
Published: (2025)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
by: Yadav, Sarthak, et al.
Published: (2024)
by: Yadav, Sarthak, et al.
Published: (2024)
Neuro2Semantic: A Transfer Learning Framework for Semantic Reconstruction of Continuous Language from Human Intracranial EEG
by: Shams, Siavash, et al.
Published: (2025)
by: Shams, Siavash, et al.
Published: (2025)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
by: Li, Yinghao Aaron, et al.
Published: (2024)
by: Li, Yinghao Aaron, et al.
Published: (2024)
Style-Talker: Finetuning Audio Language Model and Style-Based Text-to-Speech Model for Fast Spoken Dialogue Generation
by: Li, Yinghao Aaron, et al.
Published: (2024)
by: Li, Yinghao Aaron, et al.
Published: (2024)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
by: Wang, Qiaolin, et al.
Published: (2025)
by: Wang, Qiaolin, et al.
Published: (2025)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
by: Jiang, Xilin, et al.
Published: (2025)
by: Jiang, Xilin, et al.
Published: (2025)
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis
by: Li, Yinghao Aaron, et al.
Published: (2025)
by: Li, Yinghao Aaron, et al.
Published: (2025)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
by: Sinha, Anshuman, et al.
Published: (2024)
by: Sinha, Anshuman, et al.
Published: (2024)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
by: Erol, Mehmet Hamza, et al.
Published: (2024)
by: Erol, Mehmet Hamza, et al.
Published: (2024)
Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience
by: Jiang, Xilin, et al.
Published: (2024)
by: Jiang, Xilin, et al.
Published: (2024)
Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
by: Jiang, Xilin, et al.
Published: (2025)
by: Jiang, Xilin, et al.
Published: (2025)
Singer Identity Representation Learning using Self-Supervised Techniques
by: Torres, Bernardo, et al.
Published: (2024)
by: Torres, Bernardo, et al.
Published: (2024)
Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection
by: Gulzar, Kashaf, et al.
Published: (2026)
by: Gulzar, Kashaf, et al.
Published: (2026)
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
by: Vaessen, Nik, et al.
Published: (2024)
by: Vaessen, Nik, et al.
Published: (2024)
Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection
by: Sudan, Jaskirat, et al.
Published: (2026)
by: Sudan, Jaskirat, et al.
Published: (2026)
Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
by: Khorrami, Khazar, et al.
Published: (2023)
by: Khorrami, Khazar, et al.
Published: (2023)
On the Transferability of Large-Scale Self-Supervision to Few-Shot Audio Classification
by: Heggan, Calum, et al.
Published: (2024)
by: Heggan, Calum, et al.
Published: (2024)
Additive Margin in Contrastive Self-Supervised Frameworks to Learn Discriminative Speaker Representations
by: Lepage, Theo, et al.
Published: (2024)
by: Lepage, Theo, et al.
Published: (2024)
MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations
by: Heggan, Calum, et al.
Published: (2023)
by: Heggan, Calum, et al.
Published: (2023)
Understanding Self-Supervised Learning of Speech Representation via Invariance and Redundancy Reduction
by: Brima, Yusuf, et al.
Published: (2023)
by: Brima, Yusuf, et al.
Published: (2023)
Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
by: Arcos-Holzinger, Sandra, et al.
Published: (2026)
by: Arcos-Holzinger, Sandra, et al.
Published: (2026)
Learning Music Audio Representations With Limited Data
by: Plachouras, Christos, et al.
Published: (2025)
by: Plachouras, Christos, et al.
Published: (2025)
Improving Perceptual Audio Aesthetic Assessment via Triplet Loss and Self-Supervised Embeddings
by: Wisnu, Dyah A. M. G., et al.
Published: (2025)
by: Wisnu, Dyah A. M. G., et al.
Published: (2025)
Learning Disentangled Audio Representations through Controlled Synthesis
by: Brima, Yusuf, et al.
Published: (2024)
by: Brima, Yusuf, et al.
Published: (2024)
Comparing Self-Supervised Learning Models Pre-Trained on Human Speech and Animal Vocalizations for Bioacoustics Processing
by: Sarkar, Eklavya, et al.
Published: (2025)
by: Sarkar, Eklavya, et al.
Published: (2025)
SepMamba: State-space models for speaker separation using Mamba
by: Avenstrup, Thor Højhus, et al.
Published: (2024)
by: Avenstrup, Thor Højhus, et al.
Published: (2024)
COCOLA: Coherence-Oriented Contrastive Learning of Musical Audio Representations
by: Ciranni, Ruben, et al.
Published: (2024)
by: Ciranni, Ruben, et al.
Published: (2024)
Interpretable Embeddings of Speech Enhance and Explain Brain Encoding Performance of Audio Models
by: Shimizu, Riki, et al.
Published: (2025)
by: Shimizu, Riki, et al.
Published: (2025)
Unsupervised Composable Representations for Audio
by: Bindi, Giovanni, et al.
Published: (2024)
by: Bindi, Giovanni, et al.
Published: (2024)
Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
by: Miara, Victor, et al.
Published: (2024)
by: Miara, Victor, et al.
Published: (2024)
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
by: Alex, Tony, et al.
Published: (2025)
by: Alex, Tony, et al.
Published: (2025)
Singing Voice Conversion with Accompaniment Using Self-Supervised Representation-Based Melody Features
by: Chen, Wei, et al.
Published: (2025)
by: Chen, Wei, et al.
Published: (2025)
The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation
by: Collins, Nick
Published: (2024)
by: Collins, Nick
Published: (2024)
SAM: A Mamba-2 State-Space Audio-Language Model
by: Lee, Taehan, et al.
Published: (2025)
by: Lee, Taehan, et al.
Published: (2025)
Similar Items
-
DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
by: Jiang, Xilin, et al.
Published: (2023) -
Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
by: Jiang, Xilin, et al.
Published: (2024) -
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
by: Jiang, Xilin, et al.
Published: (2024) -
MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
by: Shimizu, Riki, et al.
Published: (2025) -
Exploring Finetuned Audio-LLM on Heart Murmur Features
by: Florea, Adrian, et al.
Published: (2025)