MC-SEMamba: A Simple Multi-channel Extension of SEMamba
Fuente:
arXiv
Salvato in:
| Autori principali: | Ting, Wen-Yuan, Ren, Wenze, Chao, Rong, Lin, Hsin-Yi, Tsao, Yu, Zeng, Fan-Gang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
di: Lee, Yongjoon, et al.
Pubblicazione: (2026)
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
di: Chen, Xuanjun, et al.
Pubblicazione: (2025)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
di: Chen, Chih-Ning, et al.
Pubblicazione: (2026)
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
di: Du, Jiawei, et al.
Pubblicazione: (2024)
di: Du, Jiawei, et al.
Pubblicazione: (2024)
A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2024)
A Study on Incorporating Whisper for Robust Speech Assessment
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
Universal Speech Enhancement with Regression and Generative Mamba
di: Chao, Rong, et al.
Pubblicazione: (2025)
di: Chao, Rong, et al.
Pubblicazione: (2025)
Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Speech Intelligibility Assessment with Uncertainty-Aware Whisper Embeddings and sLSTM
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
A Study on Zero-Shot Non-Intrusive Speech Intelligibility for Hearing Aids Using Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2025)
Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2023)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
ConPCO: Preserving Phoneme Characteristics for Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2024)
di: Yan, Bi-Cheng, et al.
Pubblicazione: (2024)
Unsupervised Multi-channel Separation and Adaptation
di: Han, Cong, et al.
Pubblicazione: (2023)
di: Han, Cong, et al.
Pubblicazione: (2023)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
di: Hussain, Tassadaq, et al.
Pubblicazione: (2024)
Few-Shot and Pseudo-Label Guided Speech Quality Evaluation with Large Language Models
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2026)
SVSNet+: Enhancing Speaker Voice Similarity Assessment Models with Representations from Speech Foundation Models
di: Yin, Chun, et al.
Pubblicazione: (2024)
di: Yin, Chun, et al.
Pubblicazione: (2024)
A Study on Speech Assessment with Visual Cues
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
di: Ahmed, Shafique, et al.
Pubblicazione: (2025)
PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
di: Tan, Chih-Pin, et al.
Pubblicazione: (2024)
di: Tan, Chih-Pin, et al.
Pubblicazione: (2024)
STSM-FiLM: A FiLM-Conditioned Neural Architecture for Time-Scale Modification of Speech
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
di: Wisnu, Dyah A. M. G., et al.
Pubblicazione: (2025)
Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
di: Ng, Dianwen, et al.
Pubblicazione: (2025)
di: Ng, Dianwen, et al.
Pubblicazione: (2025)
Unsupervised Multi-channel Speech Dereverberation via Diffusion
di: Wu, Yulun, et al.
Pubblicazione: (2025)
di: Wu, Yulun, et al.
Pubblicazione: (2025)
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
di: Fang, Yuan, et al.
Pubblicazione: (2024)
di: Fang, Yuan, et al.
Pubblicazione: (2024)
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
di: Wang, Ruoyu, et al.
Pubblicazione: (2024)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
di: Wang, Kuan-Chen, et al.
Pubblicazione: (2024)
SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
di: Guo, Yu-Ren, et al.
Pubblicazione: (2025)
di: Guo, Yu-Ren, et al.
Pubblicazione: (2025)
Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
An Investigation of Incorporating Mamba for Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2024)
di: Chao, Rong, et al.
Pubblicazione: (2024)
Deep Learning-based Non-Intrusive Multi-Objective Speech Assessment Model with Cross-Domain Features
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
di: Zezario, Ryandhimas E., et al.
Pubblicazione: (2021)
Towards Environmental Preference Based Speech Enhancement For Individualised Multi-Modal Hearing Aids
di: Kirton-Wingate, Jasper, et al.
Pubblicazione: (2024)
di: Kirton-Wingate, Jasper, et al.
Pubblicazione: (2024)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
di: Lin, Meng-Ping, et al.
Pubblicazione: (2025)
Towards Robust Assessment of Pathological Voices via Combined Low-Level Descriptors and Foundation Model Representations
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
di: Ariyanti, Whenty, et al.
Pubblicazione: (2025)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2023)
di: Wang, Zhichao, et al.
Pubblicazione: (2023)
Neuro-MSBG: An End-to-End Neural Model for Hearing Loss Simulation
di: Yuan, Hui-Guan, et al.
Pubblicazione: (2025)
di: Yuan, Hui-Guan, et al.
Pubblicazione: (2025)
Controlling the Parameterized Multi-channel Wiener Filter using a tiny neural network
di: Grinstein, Eric, et al.
Pubblicazione: (2025)
di: Grinstein, Eric, et al.
Pubblicazione: (2025)
AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
di: Rong, Yan, et al.
Pubblicazione: (2025)
di: Rong, Yan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
di: Lee, Yongjoon, et al.
Pubblicazione: (2026) -
Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
di: Ren, Wenze, et al.
Pubblicazione: (2024) -
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
di: Ren, Wenze, et al.
Pubblicazione: (2024) -
Leveraging Mamba with Full-Face Vision for Audio-Visual Speech Enhancement
di: Chao, Rong, et al.
Pubblicazione: (2025) -
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
di: Ren, Wenze, et al.
Pubblicazione: (2024)