Unleashing the Power of Natural Audio Featuring Multiple Sound Sources
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheng, Xize, Wang, Slytherin, Wang, Zehan, Huang, Rongjie, Jin, Tao, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
Audio Geolocation: A Natural Sounds Benchmark
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
Audio Simulation for Sound Source Localization in Virtual Evironment
von: Di Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Di Yuan, Yi, et al.
Veröffentlicht: (2024)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
PACE: Pretrained Audio Continual Learning
von: Li, Chang, et al.
Veröffentlicht: (2026)
von: Li, Chang, et al.
Veröffentlicht: (2026)
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2025)
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2025)
Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search
von: Yu, Tao, et al.
Veröffentlicht: (2026)
von: Yu, Tao, et al.
Veröffentlicht: (2026)
An Experimental Study on Joint Modeling for Sound Event Localization and Detection with Source Distance Estimation
von: Dong, Yuxuan, et al.
Veröffentlicht: (2025)
von: Dong, Yuxuan, et al.
Veröffentlicht: (2025)
Echo: Towards Advanced Audio Comprehension via Audio-Interleaved Reasoning
von: Wu, Daiqing, et al.
Veröffentlicht: (2026)
von: Wu, Daiqing, et al.
Veröffentlicht: (2026)
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
von: Wang, Yaoting, et al.
Veröffentlicht: (2023)
von: Wang, Yaoting, et al.
Veröffentlicht: (2023)
Learning Interpretable Features in Audio Latent Spaces via Sparse Autoencoders
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
von: Paek, Nathan, et al.
Veröffentlicht: (2025)
Integrating Audio, Visual, and Semantic Information for Enhanced Multimodal Speaker Diarization
von: Cheng, Luyao, et al.
Veröffentlicht: (2024)
von: Cheng, Luyao, et al.
Veröffentlicht: (2024)
SAO-Instruct: Free-form Audio Editing using Natural Language Instructions
von: Ungersböck, Michael, et al.
Veröffentlicht: (2025)
von: Ungersböck, Michael, et al.
Veröffentlicht: (2025)
AudioMosaic: Contrastive Masked Audio Representation Learning
von: Huang, Hanxun, et al.
Veröffentlicht: (2026)
von: Huang, Hanxun, et al.
Veröffentlicht: (2026)
AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation
von: Wang, Lu, et al.
Veröffentlicht: (2025)
von: Wang, Lu, et al.
Veröffentlicht: (2025)
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
von: Zhang, Zihan, et al.
Veröffentlicht: (2025)
von: Zhang, Zihan, et al.
Veröffentlicht: (2025)
Audio Super-Resolution with Latent Bridge Models
von: Li, Chang, et al.
Veröffentlicht: (2025)
von: Li, Chang, et al.
Veröffentlicht: (2025)
Transformer Redesign for Late Fusion of Audio-Text Features on Ultra-Low-Power Edge Hardware
von: Mitsis, Stavros, et al.
Veröffentlicht: (2025)
von: Mitsis, Stavros, et al.
Veröffentlicht: (2025)
Synthesizer Sound Matching Using Audio Spectrogram Transformers
von: Bruford, Fred, et al.
Veröffentlicht: (2024)
von: Bruford, Fred, et al.
Veröffentlicht: (2024)
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects?
von: Li, Jia, et al.
Veröffentlicht: (2025)
von: Li, Jia, et al.
Veröffentlicht: (2025)
wav2pos: Sound Source Localization using Masked Autoencoders
von: Berg, Axel, et al.
Veröffentlicht: (2024)
von: Berg, Axel, et al.
Veröffentlicht: (2024)
How to Label Resynthesized Audio: The Dual Role of Neural Audio Codecs in Audio Deepfake Detection
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
von: Xiao, Yixuan, et al.
Veröffentlicht: (2026)
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
SoundReactor: Frame-level Online Video-to-Audio Generation
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
von: Saito, Koichi, et al.
Veröffentlicht: (2025)
ADNAC: Audio Denoiser using Neural Audio Codec
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
von: Jimon, Daniel, et al.
Veröffentlicht: (2025)
Descriptor-Injected Cross-Modal Learning: A Systematic Exploration of Audio-MIDI Alignment via Spectral and Melodic Features
von: Méndez, Mariano Fernández
Veröffentlicht: (2026)
von: Méndez, Mariano Fernández
Veröffentlicht: (2026)
QAMRO: Quality-aware Adaptive Margin Ranking Optimization for Human-aligned Assessment of Audio Generation Systems
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
von: Wang, Chien-Chun, et al.
Veröffentlicht: (2025)
Heterogeneity-Aware Dataset Scheduling for Efficient Audio Large Language Model Training
von: Wu, Yanru, et al.
Veröffentlicht: (2026)
von: Wu, Yanru, et al.
Veröffentlicht: (2026)
Abnormal Respiratory Sound Identification Using Audio-Spectrogram Vision Transformer
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2024)
von: Ariyanti, Whenty, et al.
Veröffentlicht: (2024)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
ALMGuard: Safety Shortcuts and Where to Find Them as Guardrails for Audio-Language Models
von: Jin, Weifei, et al.
Veröffentlicht: (2025)
von: Jin, Weifei, et al.
Veröffentlicht: (2025)
Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
von: Bae, Sangmin, et al.
Veröffentlicht: (2023)
Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
von: Park, Sangwook, et al.
Veröffentlicht: (2021)
von: Park, Sangwook, et al.
Veröffentlicht: (2021)
Virtual Consistency for Audio Editing
von: Cervera, Matthieu, et al.
Veröffentlicht: (2025)
von: Cervera, Matthieu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024) -
ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024) -
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024) -
Prompt-Singer: Controllable Singing-Voice-Synthesis with Natural Language Prompt
von: Wang, Yongqi, et al.
Veröffentlicht: (2024) -
Audio Geolocation: A Natural Sounds Benchmark
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)