BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Xuenan, Zhang, Zhiling, Zhou, Zelin, Zhang, Pingyue, Xie, Zeyu, Wu, Mengyue, Zhu, Kenny Q. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)
PicoAudio: Enabling Precise Timestamp and Frequency Controllability of Audio Events in Text-to-audio Generation
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Enhancing Audio Generation Diversity with Visual Information
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Towards Weakly Supervised Text-to-Audio Grounding
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
von: Li, Baihan, et al.
Veröffentlicht: (2024)
von: Li, Baihan, et al.
Veröffentlicht: (2024)
PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
von: Zheng, Zihao, et al.
Veröffentlicht: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Yaoyun, et al.
Veröffentlicht: (2024)
STAR: Speech-to-Audio Generation via Representation Learning
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
von: Xie, Zeyu, et al.
Veröffentlicht: (2025)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
FakeSound: Deepfake General Audio Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
von: Xie, Zeyu, et al.
Veröffentlicht: (2024)
Streaming Audio Transformers for Online Audio Tagging
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
von: Sun, Luoyi, et al.
Veröffentlicht: (2023)
von: Sun, Luoyi, et al.
Veröffentlicht: (2023)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Can Audio Large Language Models Verify Speaker Identity?
von: Ren, Yiming, et al.
Veröffentlicht: (2025)
von: Ren, Yiming, et al.
Veröffentlicht: (2025)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Robust Audio Tagging under Class-wise Supervision Unreliability
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
von: Chen, Wenxi, et al.
Veröffentlicht: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
Audio Avatar Fingerprinting: An Approach for Authorized Use of Voice Cloning in the Era of Synthetic Audio
von: Gerstner, Candice R.
Veröffentlicht: (2026)
von: Gerstner, Candice R.
Veröffentlicht: (2026)
Synthetic Audio Forensics Evaluation (SAFE) Challenge
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)
von: Trapeznikov, Kirill, et al.
Veröffentlicht: (2025)
ALDAS: Audio-Linguistic Data Augmentation for Spoofed Audio Detection
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
von: Khanjani, Zahra, et al.
Veröffentlicht: (2024)
Perceptual Musical Features for Interpretable Audio Tagging
von: Lyberatos, Vassilis, et al.
Veröffentlicht: (2023)
von: Lyberatos, Vassilis, et al.
Veröffentlicht: (2023)
ProLAP: Probabilistic Language-Audio Pre-Training
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2025)
von: Manabe, Toranosuke, et al.
Veröffentlicht: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
Open-Set Source Tracing of Audio Deepfake Systems
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
von: Klein, Nicholas, et al.
Veröffentlicht: (2025)
SynHate: Detecting Hate Speech in Synthetic Deepfake Audio
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
von: Ranjan, Rishabh, et al.
Veröffentlicht: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Weighted-Sampling Audio Adversarial Example Attack
von: Liu, Xiaolei, et al.
Veröffentlicht: (2019)
von: Liu, Xiaolei, et al.
Veröffentlicht: (2019)
Code Drift: Towards Idempotent Neural Audio Codecs
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
Semantic Proximity Alignment: Towards Human Perception-consistent Audio Tagging by Aligning with Label Text Description
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
von: Liu, Wuyang, et al.
Veröffentlicht: (2023)
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
von: Xu, Xuenan, et al.
Veröffentlicht: (2024) -
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
von: Xu, Xuenan, et al.
Veröffentlicht: (2024) -
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023) -
AudioTime: A Temporally-aligned Audio-text Benchmark Dataset
von: Xie, Zeyu, et al.
Veröffentlicht: (2024) -
Hierarchical Label Propagation: A Model-Size-Dependent Performance Booster for AudioSet Tagging
von: Tuncay, Ludovic, et al.
Veröffentlicht: (2025)