Cross-attention Inspired Selective State Space Models for Target Sound Extraction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Donghang, Wang, Yiwen, Wu, Xihong, Qu, Tianshu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leveraging Sound Source Trajectories for Universal Sound Separation
von: Wu, Donghang, et al.
Veröffentlicht: (2024)
von: Wu, Donghang, et al.
Veröffentlicht: (2024)
TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
DENSE: Dynamic Embedding Causal Target Speech Extraction
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
von: Wang, Yiwen, et al.
Veröffentlicht: (2024)
MambaFoley: Foley Sound Generation using Selective State-Space Models
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Enhancing Target Speaker Extraction with Explicit Speaker Consistency Modeling
von: Wu, Shu, et al.
Veröffentlicht: (2025)
von: Wu, Shu, et al.
Veröffentlicht: (2025)
Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
von: Saijo, Kohei, et al.
Veröffentlicht: (2024)
Language-Queried Target Sound Extraction Without Parallel Training Data
von: Ma, Hao, et al.
Veröffentlicht: (2024)
von: Ma, Hao, et al.
Veröffentlicht: (2024)
Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
von: Tao, Ruijie, et al.
Veröffentlicht: (2024)
Context-Aware Query Refinement for Target Sound Extraction: Handling Partially Matched Queries
von: Sato, Ryo, et al.
Veröffentlicht: (2025)
von: Sato, Ryo, et al.
Veröffentlicht: (2025)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
Binaural Selective Attention Model for Target Speaker Extraction
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
von: Meng, Hanyu, et al.
Veröffentlicht: (2024)
Learning Control of Neural Sound Effects Synthesis from Physically Inspired Models
von: Zong, Yisu, et al.
Veröffentlicht: (2025)
von: Zong, Yisu, et al.
Veröffentlicht: (2025)
Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models
von: Zeng, Bang, et al.
Veröffentlicht: (2026)
von: Zeng, Bang, et al.
Veröffentlicht: (2026)
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
von: Zhao, He, et al.
Veröffentlicht: (2024)
von: Zhao, He, et al.
Veröffentlicht: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
von: Sato, Hiroshi, et al.
Veröffentlicht: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Online Similarity-and-Independence-Aware Beamformer for Low-latency Target Sound Extraction
von: Hiroe, Atsuo
Veröffentlicht: (2023)
von: Hiroe, Atsuo
Veröffentlicht: (2023)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Leveraging Language Information for Target Language Extraction
von: Yıldırım, Mehmet Sinan, et al.
Veröffentlicht: (2025)
von: Yıldırım, Mehmet Sinan, et al.
Veröffentlicht: (2025)
Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
von: Li, Jizhen, et al.
Veröffentlicht: (2025)
von: Li, Jizhen, et al.
Veröffentlicht: (2025)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Selective-Memory Meta-Learning with Environment Representations for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
von: Hu, Jinbo, et al.
Veröffentlicht: (2023)
Target Speaker Extraction with Curriculum Learning
von: Liu, Yun, et al.
Veröffentlicht: (2024)
von: Liu, Yun, et al.
Veröffentlicht: (2024)
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation
von: Mu, Da, et al.
Veröffentlicht: (2024)
von: Mu, Da, et al.
Veröffentlicht: (2024)
On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
UniSRM: A Unified Speech Reward Model for Reasoning-Based Fine-grained Assessment
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2026)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
von: Peng, Junyi, et al.
Veröffentlicht: (2024)
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Listen to Extract: Onset-Prompted Target Speaker Extraction
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
von: Shen, Pengjie, et al.
Veröffentlicht: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
von: Jin, Xutong, et al.
Veröffentlicht: (2024)
von: Jin, Xutong, et al.
Veröffentlicht: (2024)
Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
von: Dai, Wang, et al.
Veröffentlicht: (2025)
von: Dai, Wang, et al.
Veröffentlicht: (2025)
Beyond Speaker Identity: Text Guided Target Speech Extraction
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
von: Huo, Mingyue, et al.
Veröffentlicht: (2025)
Single-Channel Target Speech Extraction Utilizing Distance and Room Clues
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
von: Shi, Runwu, et al.
Veröffentlicht: (2025)
Distance Based Single-Channel Target Speech Extraction
von: Shi, Runwu, et al.
Veröffentlicht: (2024)
von: Shi, Runwu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Leveraging Sound Source Trajectories for Universal Sound Separation
von: Wu, Donghang, et al.
Veröffentlicht: (2024) -
TSE-PI: Target Sound Extraction under Reverberant Environments with Pitch Information
von: Wang, Yiwen, et al.
Veröffentlicht: (2024) -
DENSE: Dynamic Embedding Causal Target Speech Extraction
von: Wang, Yiwen, et al.
Veröffentlicht: (2024) -
MambaFoley: Foley Sound Generation using Selective State-Space Models
von: Colombo, Marco Furio, et al.
Veröffentlicht: (2024) -
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
von: Hernandez-Olivan, Carlos, et al.
Veröffentlicht: (2024)