BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Zihan, Sun, Jiayao, Xia, Xianjun, Huang, Chuanzeng, Xiao, Yijian, Xie, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
The IEEE-IS2 2024 Music Packet Loss Concealment Challenge
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2024)
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2024)
The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)
DualSep: A Light-weight dual-encoder convolutional recurrent network for real-time in-car speech separation
von: Wang, Ziqian, et al.
Veröffentlicht: (2024)
von: Wang, Ziqian, et al.
Veröffentlicht: (2024)
An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
von: Xu, Linping, et al.
Veröffentlicht: (2024)
von: Xu, Linping, et al.
Veröffentlicht: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
von: Wang, Ziqian, et al.
Veröffentlicht: (2026)
von: Wang, Ziqian, et al.
Veröffentlicht: (2026)
MSU-Bench: Towards Understanding the Conversational Multi-talker Scenarios
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
von: Mu, Bingshen, et al.
Veröffentlicht: (2025)
CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation
von: Kalkhorani, Vahid Ahmadi, et al.
Veröffentlicht: (2024)
von: Kalkhorani, Vahid Ahmadi, et al.
Veröffentlicht: (2024)
Moises-Light: Resource-efficient Band-split U-Net For Music Source Separation
von: Yun-Ning, et al.
Veröffentlicht: (2025)
von: Yun-Ning, et al.
Veröffentlicht: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
von: Zhu, Xinfa, et al.
Veröffentlicht: (2023)
Réduire le bruit grâce à la réalité augmentée sonore -- Auditory Concealer
von: Boukhemia, Clara
Veröffentlicht: (2025)
von: Boukhemia, Clara
Veröffentlicht: (2025)
Unseen but not Unknown: Using Dataset Concealment to Robustly Evaluate Speech Quality Estimation Models
von: Pieper, Jaden, et al.
Veröffentlicht: (2026)
von: Pieper, Jaden, et al.
Veröffentlicht: (2026)
Experimental Results of Underwater Sound Speed Profile Inversion by Few-shot Multi-task Learning
von: Huang, Wei, et al.
Veröffentlicht: (2023)
von: Huang, Wei, et al.
Veröffentlicht: (2023)
Improving Speaker Representations Using Contrastive Losses on Multi-scale Features
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
von: Dixit, Satvik, et al.
Veröffentlicht: (2024)
Advancing Robust Underwater Acoustic Target Recognition through Multi-task Learning and Multi-Gate Mixture-of-Experts
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
von: Xie, Yuan, et al.
Veröffentlicht: (2024)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
Distil-DCCRN: A Small-footprint DCCRN Leveraging Feature-based Knowledge Distillation in Speech Enhancement
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
A Lightweight Slot-Attention Framework for Multi-Instrument Multi-Pitch Estimation
von: Taenzer, Michael
Veröffentlicht: (2026)
von: Taenzer, Michael
Veröffentlicht: (2026)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
End-to-End Multi-Task Learning for Adjustable Joint Noise Reduction and Hearing Loss Compensation
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2026)
von: Gonzalez, Philippe, et al.
Veröffentlicht: (2026)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
AV-SSAN: Audio-Visual Selective DoA Estimation through Explicit Multi-Band Semantic-Spatial Alignment
von: Chen, Yu, et al.
Veröffentlicht: (2025)
von: Chen, Yu, et al.
Veröffentlicht: (2025)
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
von: Xia, Kangxiang, et al.
Veröffentlicht: (2025)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
A Multi-task Learning Balanced Attention Convolutional Neural Network Model for Few-shot Underwater Acoustic Target Recognition
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
MeanFlowSE: One-Step Generative Speech Enhancement via MeanFlow
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
von: Zhu, Yike, et al.
Veröffentlicht: (2025)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
The THU-HCSI Multi-Speaker Multi-Lingual Few-Shot Voice Cloning System for LIMMITS'24 Challenge
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2024)
ICMC-ASR: The ICASSP 2024 In-Car Multi-Channel Automatic Speech Recognition Challenge
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
A Multi-Stage Framework for Multimodal Controllable Speech Synthesis
von: Niu, Rui, et al.
Veröffentlicht: (2025)
von: Niu, Rui, et al.
Veröffentlicht: (2025)
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
von: Wang, Shuiyuan, et al.
Veröffentlicht: (2026)
M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
von: Fan, Cunhang, et al.
Veröffentlicht: (2025)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
A Two-Stage Band-Split Mamba-2 Network For Music Separation
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
von: Bai, Jinglin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
von: Zhang, Zihan, et al.
Veröffentlicht: (2024) -
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024) -
RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024) -
The IEEE-IS2 2024 Music Packet Loss Concealment Challenge
von: Mezza, Alessandro Ilic, et al.
Veröffentlicht: (2024) -
The ICASSP 2024 Audio Deep Packet Loss Concealment Challenge
von: Diener, Lorenz, et al.
Veröffentlicht: (2024)