Audiocards: Structured Metadata Improves Audio Language Models For Sound Design
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sridhar, Sripathi, Seetharaman, Prem, Nieto, Oriol, Cartwright, Mark, Salamon, Justin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generative Audio Extension and Morphing
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
Compositional Audio Representation Learning
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024)
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024)
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
von: García, Hugo Flores, et al.
Veröffentlicht: (2024)
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
von: Chen, William, et al.
Veröffentlicht: (2026)
von: Chen, William, et al.
Veröffentlicht: (2026)
Mix2Morph: Learning Sound Morphing from Noisy Mixes
von: Chu, Annie, et al.
Veröffentlicht: (2026)
von: Chu, Annie, et al.
Veröffentlicht: (2026)
FLAM: Frame-Wise Language-Audio Modeling
von: Wu, Yusong, et al.
Veröffentlicht: (2025)
von: Wu, Yusong, et al.
Veröffentlicht: (2025)
Video-Guided Foley Sound Generation with Multimodal Controls
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
von: Chen, Ziyang, et al.
Veröffentlicht: (2024)
TAC: Timestamped Audio Captioning
von: Kumar, Sonal, et al.
Veröffentlicht: (2026)
von: Kumar, Sonal, et al.
Veröffentlicht: (2026)
PromptSep: Generative Audio Separation via Multimodal Prompting
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
von: Manco, Ilaria, et al.
Veröffentlicht: (2024)
von: Manco, Ilaria, et al.
Veröffentlicht: (2024)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
Taming Audio VAEs via Target-KL Regularization
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Code Drift: Towards Idempotent Neural Audio Codecs
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2024)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2024)
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2024)
Expressive Range Characterization of Open Text-to-Audio Models
von: Morse, Jonathan, et al.
Veröffentlicht: (2025)
von: Morse, Jonathan, et al.
Veröffentlicht: (2025)
First-Shot Unsupervised Anomalous Sound Detection With Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
von: Manivannan, Mithun, et al.
Veröffentlicht: (2024)
von: Manivannan, Mithun, et al.
Veröffentlicht: (2024)
ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling
von: Xie, Hao-Hui, et al.
Veröffentlicht: (2026)
von: Xie, Hao-Hui, et al.
Veröffentlicht: (2026)
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning in Large Audio-Language Models
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
von: Xiong, Zhen, et al.
Veröffentlicht: (2025)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
SpotSound: Enhancing Large Audio-Language Models with Fine-Grained Temporal Grounding
von: Sun, Luoyi, et al.
Veröffentlicht: (2026)
von: Sun, Luoyi, et al.
Veröffentlicht: (2026)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
AudioGS: Spectrogram-Based Audio Gaussian Splatting for Sound Field Reconstruction
von: Bi, Chunhao, et al.
Veröffentlicht: (2026)
von: Bi, Chunhao, et al.
Veröffentlicht: (2026)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
MUKA: Multi Kernel Audio Adaptation Of Audio-Language Models
von: Bensaid, Reda, et al.
Veröffentlicht: (2026)
von: Bensaid, Reda, et al.
Veröffentlicht: (2026)
AudioToolAgent: An Agentic Framework for Audio-Language Models
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2025)
von: Wijngaard, Gijs, et al.
Veröffentlicht: (2025)
Spatial Audio Question Answering and Reasoning on Dynamic Source Movements
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2026)
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2026)
BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification
von: Kim, June-Woo, et al.
Veröffentlicht: (2024)
von: Kim, June-Woo, et al.
Veröffentlicht: (2024)
ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
von: Luo, Kaiwen, et al.
Veröffentlicht: (2026)
von: Luo, Kaiwen, et al.
Veröffentlicht: (2026)
Waveform-Logmel Audio Neural Networks for Respiratory Sound Classification
von: Xie, Jiadong, et al.
Veröffentlicht: (2025)
von: Xie, Jiadong, et al.
Veröffentlicht: (2025)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
von: Kim, Inho, et al.
Veröffentlicht: (2025)
von: Kim, Inho, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generative Audio Extension and Morphing
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026) -
Compositional Audio Representation Learning
von: Sridhar, Sripathi, et al.
Veröffentlicht: (2024) -
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024) -
Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic Imitations
von: García, Hugo Flores, et al.
Veröffentlicht: (2024) -
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
von: Chen, William, et al.
Veröffentlicht: (2026)