Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Jinhua, Liu, Xubo, Wang, Wenwu, Plumbley, Mark D., Phan, Huy, Benetos, Emmanouil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WavCraft: Audio Editing and Generation with Large Language Models
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
von: Zhang, Huan, et al.
Veröffentlicht: (2025)
von: Zhang, Huan, et al.
Veröffentlicht: (2025)
Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
von: Xu, Yanze, et al.
Veröffentlicht: (2026)
von: Xu, Yanze, et al.
Veröffentlicht: (2026)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Acoustic identification of individual animals with hierarchical contrastive learning
von: Nolasco, Ines, et al.
Veröffentlicht: (2024)
von: Nolasco, Ines, et al.
Veröffentlicht: (2024)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
von: Liu, Haohe, et al.
Veröffentlicht: (2025)
von: Liu, Haohe, et al.
Veröffentlicht: (2025)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
Domain-Invariant Representation Learning of Bird Sounds
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
von: Moummad, Ilyass, et al.
Veröffentlicht: (2024)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
Towards Reliable Objective Evaluation Metrics for Generative Singing Voice Separation Models
von: Bereuter, Paul A., et al.
Veröffentlicht: (2025)
von: Bereuter, Paul A., et al.
Veröffentlicht: (2025)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
von: Huang, Jiawen, et al.
Veröffentlicht: (2025)
von: Huang, Jiawen, et al.
Veröffentlicht: (2025)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Separate Anything You Describe
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
Universal Music Representations? Evaluating Foundation Models on World Music Corpora
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2025)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2025)
BioDCASE 2026 Challenge Baseline for Cross-Domain Mosquito Species Classification
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2024)
SCRAPL: Scattering Transform with Random Paths for Machine Learning
von: Mitcheltree, Christopher, et al.
Veröffentlicht: (2026)
von: Mitcheltree, Christopher, et al.
Veröffentlicht: (2026)
Generalized Multi-Source Inference for Text Conditioned Music Diffusion Models
von: Postolache, Emilian, et al.
Veröffentlicht: (2024)
von: Postolache, Emilian, et al.
Veröffentlicht: (2024)
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
von: He, Haolin, et al.
Veröffentlicht: (2025)
von: He, Haolin, et al.
Veröffentlicht: (2025)
A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
von: Xiao, Feiyang, et al.
Veröffentlicht: (2024)
von: Xiao, Feiyang, et al.
Veröffentlicht: (2024)
Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task
von: Phan, Dang Thoai
Veröffentlicht: (2024)
von: Phan, Dang Thoai
Veröffentlicht: (2024)
RUMAA: Repeat-Aware Unified Music Audio Analysis for Score-Performance Alignment, Transcription, and Mistake Detection
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
von: Chang, Sungkyun, et al.
Veröffentlicht: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WavCraft: Audio Editing and Generation with Large Language Models
von: Liang, Jinhua, et al.
Veröffentlicht: (2024) -
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
von: Zhang, Huan, et al.
Veröffentlicht: (2025) -
Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection
von: Liang, Jinhua, et al.
Veröffentlicht: (2024) -
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
von: Yuan, Yi, et al.
Veröffentlicht: (2024) -
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)