Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mahmud, Tanvir, Amizadeh, Saeed, Koishida, Kazuhito, Marculescu, Diana |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
von: Hu, Jing, et al.
Veröffentlicht: (2026)
von: Hu, Jing, et al.
Veröffentlicht: (2026)
Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
von: Jang, Jaehyuk, et al.
Veröffentlicht: (2026)
von: Jang, Jaehyuk, et al.
Veröffentlicht: (2026)
Zero-Shot Text-to-Speech from Continuous Text Streams
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
Single-channel speech enhancement using learnable loss mixup
von: Chang, Oscar, et al.
Veröffentlicht: (2023)
von: Chang, Oscar, et al.
Veröffentlicht: (2023)
Generalizable Audio Spoofing Detection using Non-Semantic Representations
von: Das, Arnab, et al.
Veröffentlicht: (2025)
von: Das, Arnab, et al.
Veröffentlicht: (2025)
DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
von: Lee, Geonyoung, et al.
Veröffentlicht: (2025)
von: Lee, Geonyoung, et al.
Veröffentlicht: (2025)
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
von: Deng, Qixin, et al.
Veröffentlicht: (2025)
von: Deng, Qixin, et al.
Veröffentlicht: (2025)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
AudioRouter: Data Efficient Audio Understanding via RL based Dual Reasoning
von: Chen, Liyang, et al.
Veröffentlicht: (2026)
von: Chen, Liyang, et al.
Veröffentlicht: (2026)
Learning Semantic Information from Raw Audio Signal Using Both Contextual and Phonetic Representations
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
Text-Queried Audio Source Separation via Hierarchical Modeling
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
von: Yin, Xinlei, et al.
Veröffentlicht: (2025)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
von: Batlle-Roca, Roser, et al.
Veröffentlicht: (2024)
Audio Atlas: Visualizing and Exploring Audio Datasets
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2024)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2024)
Efficient Autoregressive Audio Modeling via Next-Scale Prediction
von: Qiu, Kai, et al.
Veröffentlicht: (2024)
von: Qiu, Kai, et al.
Veröffentlicht: (2024)
Performance Improvement of Language-Queried Audio Source Separation Based on Caption Augmentation From Large Language Models for DCASE Challenge 2024 Task 9
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024)
von: Lee, Do Hyun, et al.
Veröffentlicht: (2024)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
von: Han, Bing, et al.
Veröffentlicht: (2026)
von: Han, Bing, et al.
Veröffentlicht: (2026)
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
MoWE-Audio: Multitask AudioLLMs with Mixture of Weak Encoders
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
von: Zhang, Wenyu, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
Audio Mamba: Bidirectional State Space Model for Audio Representation Learning
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
von: Erol, Mehmet Hamza, et al.
Veröffentlicht: (2024)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2024)
AudioScene: Integrating Object-Event Audio into 3D Scenes
von: Yuan, Shuaihang, et al.
Veröffentlicht: (2025)
von: Yuan, Shuaihang, et al.
Veröffentlicht: (2025)
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
von: Dang, Ting, et al.
Veröffentlicht: (2025)
von: Dang, Ting, et al.
Veröffentlicht: (2025)
Stable Audio Open
von: Evans, Zach, et al.
Veröffentlicht: (2024)
von: Evans, Zach, et al.
Veröffentlicht: (2024)
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
von: Watcharasupat, Karn N., et al.
Veröffentlicht: (2024)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024) -
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024) -
T-VSL: Text-Guided Visual Sound Source Localization in Mixtures
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024) -
MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024) -
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
von: Tabassum, Afrina, et al.
Veröffentlicht: (2024)