ZeroSep: Separate Anything in Audio with Zero Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huang, Chao, Ma, Yuesheng, Huang, Junxuan, Liang, Susan, Tang, Yunlong, Bi, Jing, Liu, Wenqiang, Mesgarani, Nima, Xu, Chenliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024)
DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
von: Jiang, Xilin, et al.
Veröffentlicht: (2023)
von: Jiang, Xilin, et al.
Veröffentlicht: (2023)
Exploring Finetuned Audio-LLM on Heart Murmur Features
von: Florea, Adrian, et al.
Veröffentlicht: (2025)
von: Florea, Adrian, et al.
Veröffentlicht: (2025)
PromptSep: Generative Audio Separation via Multimodal Prompting
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
von: Wen, Yutong, et al.
Veröffentlicht: (2025)
Interpretable Embeddings of Speech Enhance and Explain Brain Encoding Performance of Audio Models
von: Shimizu, Riki, et al.
Veröffentlicht: (2025)
von: Shimizu, Riki, et al.
Veröffentlicht: (2025)
UniSep: Universal Target Audio Separation with Language Models at Scale
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
High-Quality Visually-Guided Sound Separation from Diverse Categories
von: Huang, Chao, et al.
Veröffentlicht: (2023)
von: Huang, Chao, et al.
Veröffentlicht: (2023)
Can Sound Replace Vision in LLaVA With Token Substitution?
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
von: Vosoughi, Ali, et al.
Veröffentlicht: (2025)
Zero-Shot Audio Captioning Using Soft and Hard Prompts
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
von: Zhang, Yiming, et al.
Veröffentlicht: (2024)
Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
von: Tavares, Tiago, et al.
Veröffentlicht: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Zero-Shot Duet Singing Voices Separation with Diffusion Models
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2023)
Semantic visually-guided acoustic highlighting with large vision-language models
von: Huang, Junhua, et al.
Veröffentlicht: (2026)
von: Huang, Junhua, et al.
Veröffentlicht: (2026)
Multi-label Zero-Shot Audio Classification with Temporal Attention
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
von: Dogan, Duygu, et al.
Veröffentlicht: (2024)
SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
von: Wang, Qiaolin, et al.
Veröffentlicht: (2025)
A Fast and Lightweight Model for Causal Audio-Visual Speech Separation
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
von: Sang, Wendi, et al.
Veröffentlicht: (2025)
Can Quantized Audio Language Models Perform Zero-Shot Spoofing Detection?
von: Dutta, Bikash, et al.
Veröffentlicht: (2025)
von: Dutta, Bikash, et al.
Veröffentlicht: (2025)
SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2025)
von: Mu, Zhaoxi, et al.
Veröffentlicht: (2025)
Listen, Chat, and Remix: Text-Guided Soundscape Remixing for Enhanced Auditory Experience
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
von: Jiang, Xilin, et al.
Veröffentlicht: (2024)
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
von: Dutta, Soumya, et al.
Veröffentlicht: (2024)
DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
von: Li, Xiquan, et al.
Veröffentlicht: (2024)
OpenSep: Leveraging Large Language Models with Textual Inversion for Open World Audio Separation
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
von: Mahmud, Tanvir, et al.
Veröffentlicht: (2024)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
von: Yang, Mu, et al.
Veröffentlicht: (2024)
von: Yang, Mu, et al.
Veröffentlicht: (2024)
ConSep: a Noise- and Reverberation-Robust Speech Separation Framework by Magnitude Conditioning
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
von: Ho, Kuan-Hsun, et al.
Veröffentlicht: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
von: Dang, Trung, et al.
Veröffentlicht: (2024)
von: Dang, Trung, et al.
Veröffentlicht: (2024)
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
von: Steinhardt, Cynthia R., et al.
Veröffentlicht: (2024)
von: Steinhardt, Cynthia R., et al.
Veröffentlicht: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
von: Janiczek, John, et al.
Veröffentlicht: (2024)
von: Janiczek, John, et al.
Veröffentlicht: (2024)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2025)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
Erasing Your Voice Before It's Heard: Training-free Speaker Unlearning for Zero-shot Text-to-Speech
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
von: Lee, Myungjin, et al.
Veröffentlicht: (2026)
Unsupervised Single-Channel Audio Separation with Diffusion Source Priors
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
von: Chao, Rong, et al.
Veröffentlicht: (2025)
von: Chao, Rong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dual-path Mamba: Short and Long-term Bidirectional Selective Structured State Space Models for Speech Separation
von: Jiang, Xilin, et al.
Veröffentlicht: (2024) -
StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
von: Li, Yinghao Aaron, et al.
Veröffentlicht: (2024) -
DeCoR: Defy Knowledge Forgetting by Predicting Earlier Audio Codes
von: Jiang, Xilin, et al.
Veröffentlicht: (2023) -
Exploring Finetuned Audio-LLM on Heart Murmur Features
von: Florea, Adrian, et al.
Veröffentlicht: (2025) -
PromptSep: Generative Audio Separation via Multimodal Prompting
von: Wen, Yutong, et al.
Veröffentlicht: (2025)