Exploring Federated Self-Supervised Learning for General Purpose Audio Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rehman, Yasar Abbas Ur, Lau, Kin Wai, Xie, Yuyang, Ma, Lan, Shen, Jiajun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
von: Rehman, Yasar Abbas Ur, et al.
Veröffentlicht: (2025)
von: Rehman, Yasar Abbas Ur, et al.
Veröffentlicht: (2025)
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
von: Lau, Kin Wai, et al.
Veröffentlicht: (2024)
von: Lau, Kin Wai, et al.
Veröffentlicht: (2024)
Natural Language Supervision for General-Purpose Audio Representations
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
von: Lau, Kin Wai, et al.
Veröffentlicht: (2026)
von: Lau, Kin Wai, et al.
Veröffentlicht: (2026)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
MT2KD: Towards A General-Purpose Encoder for Speech, Speaker, and Audio Events
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
Towards Weakly Supervised Text-to-Audio Grounding
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
von: Wilkins, Julia, et al.
Veröffentlicht: (2024)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
von: Zhou, Xuanru, et al.
Veröffentlicht: (2026)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2026)
DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
von: Lu, Ke-Han, et al.
Veröffentlicht: (2025)
von: Lu, Ke-Han, et al.
Veröffentlicht: (2025)
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2026)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
Pitch Estimation With Mean Averaging Smoothed Product Spectrum And Musical Consonance Evaluation Using MASP
von: Baskin, Murat Yasar
Veröffentlicht: (2025)
von: Baskin, Murat Yasar
Veröffentlicht: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Can Large Language Models Understand Spatial Audio?
von: Tang, Changli, et al.
Veröffentlicht: (2024)
von: Tang, Changli, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
CardioLive: Empowering Video Streaming with Online Cardiac Monitoring
von: Lyu, Sheng, et al.
Veröffentlicht: (2025)
von: Lyu, Sheng, et al.
Veröffentlicht: (2025)
Robust Audio Tagging under Class-wise Supervision Unreliability
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
MATS: An Audio Language Model under Text-only Supervision
von: Wang, Wen, et al.
Veröffentlicht: (2025)
von: Wang, Wen, et al.
Veröffentlicht: (2025)
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
von: Du, Chenpeng, et al.
Veröffentlicht: (2022)
Towards Spatial Audio Understanding via Question Answering
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
Exploring Perceptual Audio Quality Measurement on Stereo Processing Using the Open Dataset of Audio Quality
von: Delgado, Pablo M., et al.
Veröffentlicht: (2025)
von: Delgado, Pablo M., et al.
Veröffentlicht: (2025)
Improving Audio Captioning Models with Fine-grained Audio Features, Text Embedding Supervision, and LLM Mix-up Augmentation
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
von: Wu, Shih-Lun, et al.
Veröffentlicht: (2023)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Xi+: Uncertainty Supervision for Robust Speaker Embedding
von: Li, Junjie, et al.
Veröffentlicht: (2025)
von: Li, Junjie, et al.
Veröffentlicht: (2025)
Automatic Music Mixing using a Generative Model of Effect Embeddings
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
von: Moliner, Eloi, et al.
Veröffentlicht: (2025)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
von: Jia, Yuhang, et al.
Veröffentlicht: (2025)
High-Fidelity Generative Audio Compression at 0.275kbps
von: Ma, Hao, et al.
Veröffentlicht: (2026)
von: Ma, Hao, et al.
Veröffentlicht: (2026)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FSSUAVL: A Discriminative Framework using Vision Models for Federated Self-Supervised Audio and Image Understanding
von: Rehman, Yasar Abbas Ur, et al.
Veröffentlicht: (2025) -
AudioRepInceptionNeXt: A lightweight single-stream architecture for efficient audio recognition
von: Lau, Kin Wai, et al.
Veröffentlicht: (2024) -
Natural Language Supervision for General-Purpose Audio Representations
von: Elizalde, Benjamin, et al.
Veröffentlicht: (2023) -
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
von: Han, Bing, et al.
Veröffentlicht: (2025) -
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
von: Lau, Kin Wai, et al.
Veröffentlicht: (2026)