SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes
Fuente:
arXiv
Salvato in:
| Autori principali: | Alex, Tony, Ahmed, Sara, Mustafa, Armin, Awais, Muhammad, Jackson, Philip JB |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
di: Alex, Tony, et al.
Pubblicazione: (2025)
di: Alex, Tony, et al.
Pubblicazione: (2025)
End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding
di: Zeng, Wei, et al.
Pubblicazione: (2024)
di: Zeng, Wei, et al.
Pubblicazione: (2024)
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
di: Chen, Yuanjian, et al.
Pubblicazione: (2026)
di: Chen, Yuanjian, et al.
Pubblicazione: (2026)
EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
di: Guo, Xiaoxuan, et al.
Pubblicazione: (2025)
di: Guo, Xiaoxuan, et al.
Pubblicazione: (2025)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
di: Lee, Dongheon, et al.
Pubblicazione: (2024)
di: Lee, Dongheon, et al.
Pubblicazione: (2024)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
di: Yadav, Sarthak, et al.
Pubblicazione: (2024)
Generating Moving 3D Soundscapes with Latent Diffusion Models
di: Templin, Christian, et al.
Pubblicazione: (2025)
di: Templin, Christian, et al.
Pubblicazione: (2025)
Transferable Adversarial Attacks on Audio Deepfake Detection
di: Farooq, Muhammad Umar, et al.
Pubblicazione: (2025)
di: Farooq, Muhammad Umar, et al.
Pubblicazione: (2025)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
di: Cai, Pengfei, et al.
Pubblicazione: (2024)
di: Cai, Pengfei, et al.
Pubblicazione: (2024)
Improving Speech Inversion Through Self-Supervised Embeddings and Enhanced Tract Variables
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2023)
RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
di: Wei, Haojie, et al.
Pubblicazione: (2023)
di: Wei, Haojie, et al.
Pubblicazione: (2023)
Soundscape Captioning using Sound Affective Quality Network and Large Language Model
di: Hou, Yuanbo, et al.
Pubblicazione: (2024)
di: Hou, Yuanbo, et al.
Pubblicazione: (2024)
Self-Supervised Audio-Visual Soundscape Stylization
di: Li, Tingle, et al.
Pubblicazione: (2024)
di: Li, Tingle, et al.
Pubblicazione: (2024)
Disambiguation of Chinese Polyphones in an End-to-End Framework with Semantic Features Extracted by Pre-trained BERT
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
di: Dai, Dongyang, et al.
Pubblicazione: (2025)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
di: Han, Bing, et al.
Pubblicazione: (2025)
di: Han, Bing, et al.
Pubblicazione: (2025)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
di: Shan, Weiqiao, et al.
Pubblicazione: (2025)
di: Shan, Weiqiao, et al.
Pubblicazione: (2025)
CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
di: Luong, Justin, et al.
Pubblicazione: (2025)
di: Luong, Justin, et al.
Pubblicazione: (2025)
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
di: Kang, Minjae, et al.
Pubblicazione: (2025)
di: Kang, Minjae, et al.
Pubblicazione: (2025)
Bridging ASR and LLMs for Dysarthric Speech Recognition: Benchmarking Self-Supervised and Generative Approaches
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
di: Aboeitta, Ahmed, et al.
Pubblicazione: (2025)
A Neural Score Follower for Computer Accompaniment of Polyphonic Musical Instruments
di: Pillay, Ashwin
Pubblicazione: (2025)
di: Pillay, Ashwin
Pubblicazione: (2025)
Enhancing Crowdsourced Audio for Text-to-Speech Models
di: Giraldo, José, et al.
Pubblicazione: (2024)
di: Giraldo, José, et al.
Pubblicazione: (2024)
Semi-Supervised Self-Learning Enhanced Music Emotion Recognition
di: Sun, Yifu, et al.
Pubblicazione: (2024)
di: Sun, Yifu, et al.
Pubblicazione: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
di: Wilkins, Julia, et al.
Pubblicazione: (2024)
di: Wilkins, Julia, et al.
Pubblicazione: (2024)
Contrastive Loss Based Frame-wise Feature disentanglement for Polyphonic Sound Event Detection
di: Guan, Yadong, et al.
Pubblicazione: (2024)
di: Guan, Yadong, et al.
Pubblicazione: (2024)
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
di: Li, Haowen, et al.
Pubblicazione: (2026)
di: Li, Haowen, et al.
Pubblicazione: (2026)
Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
di: Sinha, Anshuman, et al.
Pubblicazione: (2024)
di: Sinha, Anshuman, et al.
Pubblicazione: (2024)
ARAUS: A Large-Scale Dataset and Baseline Models of Affective Responses to Augmented Urban Soundscapes
di: Ooi, Kenneth, et al.
Pubblicazione: (2022)
di: Ooi, Kenneth, et al.
Pubblicazione: (2022)
Feature Selection via Graph Topology Inference for Soundscape Emotion Recognition
di: Rey, Samuel, et al.
Pubblicazione: (2025)
di: Rey, Samuel, et al.
Pubblicazione: (2025)
Autonomous Soundscape Augmentation with Multimodal Fusion of Visual and Participant-linked Inputs
di: Ooi, Kenneth, et al.
Pubblicazione: (2023)
di: Ooi, Kenneth, et al.
Pubblicazione: (2023)
Robust Bioacoustic Detection via Richly Labelled Synthetic Soundscape Augmentation
di: Soltero, Kaspar, et al.
Pubblicazione: (2025)
di: Soltero, Kaspar, et al.
Pubblicazione: (2025)
Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models
di: Kabir, Muhammad Ashad, et al.
Pubblicazione: (2026)
di: Kabir, Muhammad Ashad, et al.
Pubblicazione: (2026)
AudioScene: Integrating Object-Event Audio into 3D Scenes
di: Yuan, Shuaihang, et al.
Pubblicazione: (2025)
di: Yuan, Shuaihang, et al.
Pubblicazione: (2025)
Room Impulse Response Synthesis via Differentiable Feedback Delay Networks for Efficient Spatial Audio Rendering
di: Gerami, Armin, et al.
Pubblicazione: (2025)
di: Gerami, Armin, et al.
Pubblicazione: (2025)
MATS: An Audio Language Model under Text-only Supervision
di: Wang, Wen, et al.
Pubblicazione: (2025)
di: Wang, Wen, et al.
Pubblicazione: (2025)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
di: Lin, Jiaju, et al.
Pubblicazione: (2024)
di: Lin, Jiaju, et al.
Pubblicazione: (2024)
Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
di: Zhang, Ellie L., et al.
Pubblicazione: (2025)
di: Zhang, Ellie L., et al.
Pubblicazione: (2025)
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
Enhancing Speech Emotion Recognition through Segmental Average Pooling of Self-Supervised Learning Features
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
di: Hyeon, Jonghwan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
di: Alex, Tony, et al.
Pubblicazione: (2025) -
End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding
di: Zeng, Wei, et al.
Pubblicazione: (2024) -
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
di: Chen, Yuanjian, et al.
Pubblicazione: (2026) -
EnvSSLAM-FFN: Lightweight Layer-Fused System for ESDD 2026 Challenge
di: Guo, Xiaoxuan, et al.
Pubblicazione: (2025) -
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
di: Lee, Dongheon, et al.
Pubblicazione: (2024)