Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Haohe, Deacon, Thomas, Wang, Wenwu, Paradis, Matt, Plumbley, Mark D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
EnvSDD: Benchmarking Environmental Sound Deepfake Detection
von: Yin, Han, et al.
Veröffentlicht: (2025)
von: Yin, Han, et al.
Veröffentlicht: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Learning Temporal Resolution in Spectrogram for Audio Classification
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
von: Liu, Haohe, et al.
Veröffentlicht: (2022)
Retrieval-Augmented Text-to-Audio Generation
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
von: Yuan, Yi, et al.
Veröffentlicht: (2023)
First-Shot Unsupervised Anomalous Sound Detection With Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
von: Zhang, Hejing, et al.
Veröffentlicht: (2023)
DreamAudio: Customized Text-to-Audio Generation with Diffusion Models
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
von: Yuan, Yi, et al.
Veröffentlicht: (2025)
PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
AudioTurbo: Fast Text-to-Audio Generation with Rectified Diffusion
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
von: Zhao, Junqi, et al.
Veröffentlicht: (2025)
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
von: Mei, Xinhao, et al.
Veröffentlicht: (2023)
von: Mei, Xinhao, et al.
Veröffentlicht: (2023)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
BioDCASE 2026 Challenge Baseline for Cross-Domain Mosquito Species Classification
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2026)
Selective-Memory Meta-Learning with Environment Representations for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
Separate Anything You Describe
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
Towards Generating Diverse Audio Captions via Adversarial Training
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
von: Mei, Xinhao, et al.
Veröffentlicht: (2022)
T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
von: Liu, Haohe, et al.
Veröffentlicht: (2023)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
von: Burchett-Vass, Rhys, et al.
Veröffentlicht: (2024)
von: Burchett-Vass, Rhys, et al.
Veröffentlicht: (2024)
Exploring Differences between Human Perception and Model Inference in Audio Event Recognition
von: Tan, Yizhou, et al.
Veröffentlicht: (2024)
von: Tan, Yizhou, et al.
Veröffentlicht: (2024)
Environmental Sound Classification on An Embedded Hardware Platform
von: Bibbo, Gabriel, et al.
Veröffentlicht: (2023)
von: Bibbo, Gabriel, et al.
Veröffentlicht: (2023)
Fish Tracking, Counting, and Behaviour Analysis in Digital Aquaculture: A Comprehensive Survey
von: Cui, Meng, et al.
Veröffentlicht: (2024)
von: Cui, Meng, et al.
Veröffentlicht: (2024)
A decade of DCASE: Achievements, practices, evaluations and future challenges
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024)
von: Mesaros, Annamaria, et al.
Veröffentlicht: (2024)
AI in Music and Sound: Pedagogical Reflections, Post-Structuralist Approaches and Creative Outcomes in Seminar Practice
von: Coelho, Guilherme
Veröffentlicht: (2025)
von: Coelho, Guilherme
Veröffentlicht: (2025)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
von: Cui, Meng, et al.
Veröffentlicht: (2023)
von: Cui, Meng, et al.
Veröffentlicht: (2023)
ASiT: Local-Global Audio Spectrogram vIsion Transformer for Event Classification
von: Atito, Sara, et al.
Veröffentlicht: (2022)
von: Atito, Sara, et al.
Veröffentlicht: (2022)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
von: Singh, Arshdeep, et al.
Veröffentlicht: (2025)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
von: Han, Bing, et al.
Veröffentlicht: (2025)
von: Han, Bing, et al.
Veröffentlicht: (2025)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
von: Tian, Wenjie, et al.
Veröffentlicht: (2026)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
von: Li, Baihan, et al.
Veröffentlicht: (2024)
von: Li, Baihan, et al.
Veröffentlicht: (2024)
Soundscape Captioning using Sound Affective Quality Network and Large Language Model
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
Description on IEEE ICME 2024 Grand Challenge: Semi-supervised Acoustic Scene Classification under Domain Shift
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
von: Bai, Jisheng, et al.
Veröffentlicht: (2024)
MixAssist: An Audio-Language Dataset for Co-Creative AI Assistance in Music Mixing
von: Clemens, Michael, et al.
Veröffentlicht: (2025)
von: Clemens, Michael, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
von: Yuan, Yi, et al.
Veröffentlicht: (2024) -
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
von: Xu, Xuenan, et al.
Veröffentlicht: (2024) -
Region-Specific Audio Tagging for Spatial Sound
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2025) -
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024) -
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
von: Yuan, Yi, et al.
Veröffentlicht: (2023)