WavCraft: Audio Editing and Generation with Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Jinhua, Zhang, Huan, Liu, Haohe, Cao, Yin, Kong, Qiuqiang, Liu, Xubo, Wang, Wenwu, Plumbley, Mark D., Phan, Huy, Benetos, Emmanouil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities
von: Liang, Jinhua, et al.
Veröffentlicht: (2023)
von: Liang, Jinhua, et al.
Veröffentlicht: (2023)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
von: Zhang, Huan, et al.
Veröffentlicht: (2025)
von: Zhang, Huan, et al.
Veröffentlicht: (2025)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
von: Mei, Xinhao, et al.
Veröffentlicht: (2023)
von: Mei, Xinhao, et al.
Veröffentlicht: (2023)
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
Separate Anything You Describe
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
von: Singh, Shubhr, et al.
Veröffentlicht: (2025)
Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
von: Liu, Haohe, et al.
Veröffentlicht: (2025)
von: Liu, Haohe, et al.
Veröffentlicht: (2025)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
First-Shot Unsupervised Anomalous Sound Detection With Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
Selective-Memory Meta-Learning with Environment Representations for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
Inference-time Scaling for Diffusion-based Audio Super-resolution
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
von: He, Haolin, et al.
Veröffentlicht: (2025)
von: He, Haolin, et al.
Veröffentlicht: (2025)
EnvSDD: Benchmarking Environmental Sound Deepfake Detection
von: Yin, Han, et al.
Veröffentlicht: (2025)
von: Yin, Han, et al.
Veröffentlicht: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Learning Music Audio Representations With Limited Data
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
von: Plachouras, Christos, et al.
Veröffentlicht: (2025)
Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Survey
von: Cui, Meng, et al.
Veröffentlicht: (2024)
von: Cui, Meng, et al.
Veröffentlicht: (2024)
Music Source Restoration
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
LC-Protonets: Multi-Label Few-Shot Learning for World Music Audio Tagging
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
von: Papaioannou, Charilaos, et al.
Veröffentlicht: (2024)
Towards Out-of-Distribution Detection in Vocoder Recognition via Latent Feature Reconstruction
von: Du, Renmingyue, et al.
Veröffentlicht: (2024)
von: Du, Renmingyue, et al.
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Text-Queried Target Sound Event Localization
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
From Audio Encoders to Piano Judges: Benchmarking Performance Understanding for Solo Piano
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
von: Zhang, Huan, et al.
Veröffentlicht: (2024)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
von: Cui, Meng, et al.
Veröffentlicht: (2023)
von: Cui, Meng, et al.
Veröffentlicht: (2023)
PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
WavMark: Watermarking for Audio Generation
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Acoustic Prompt Tuning: Empowering Large Language Models with Audition Capabilities
von: Liang, Jinhua, et al.
Veröffentlicht: (2023) -
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022) -
From Aesthetics to Human Preferences: Comparative Perspectives of Evaluating Text-to-Music Systems
von: Zhang, Huan, et al.
Veröffentlicht: (2025) -
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023) -
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
von: Mei, Xinhao, et al.
Veröffentlicht: (2023)