LDCodec: A high quality neural audio codec with low-complexity decoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Jiawei, Xu, Linping, Zhang, Dejun, Huang, Qingbo, Xia, Xianjun, Xiao, Yijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
von: Xu, Linping, et al.
Veröffentlicht: (2024)
von: Xu, Linping, et al.
Veröffentlicht: (2024)
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025)
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)
Real-time speech enhancement in noise for throat microphone using neural audio codec as foundation model
von: Hauret, Julien, et al.
Veröffentlicht: (2025)
von: Hauret, Julien, et al.
Veröffentlicht: (2025)
BS-PLCNet 2: Two-stage Band-split Packet Loss Concealment Network with Intra-model Knowledge Distillation
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
PitchFlower: A flow-based neural audio codec with pitch controllability
von: Torres, Diego, et al.
Veröffentlicht: (2025)
von: Torres, Diego, et al.
Veröffentlicht: (2025)
EnCodecMAE: Leveraging neural codecs for universal audio representation learning
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
von: Pepino, Leonardo, et al.
Veröffentlicht: (2023)
FlowDec: A flow-based full-band general audio codec with high perceptual quality
von: Welker, Simon, et al.
Veröffentlicht: (2025)
von: Welker, Simon, et al.
Veröffentlicht: (2025)
FreeCodec: A disentangled neural speech codec with fewer tokens
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
BS-PLCNet: Band-split Packet Loss Concealment Network with Multi-task Learning Framework and Multi-discriminators
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
von: Zhang, Zihan, et al.
Veröffentlicht: (2024)
RaD-Net 2: A causal two-stage repairing and denoising speech enhancement network with knowledge distillation and complex axial self-attention
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
U-SAM: An audio language Model for Unified Speech, Audio, and Music Understanding
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
RaD-Net: A Repairing and Denoising Network for Speech Signal Improvement
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
von: Liu, Mingshuai, et al.
Veröffentlicht: (2024)
MBCodec:Thorough disentangle for high-fidelity audio compression
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
von: Wang, Ziqian, et al.
Veröffentlicht: (2026)
von: Wang, Ziqian, et al.
Veröffentlicht: (2026)
Towards predicting binaural audio quality in listeners with normal and impaired hearing
von: Biberger, Thomas, et al.
Veröffentlicht: (2025)
von: Biberger, Thomas, et al.
Veröffentlicht: (2025)
Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
von: Siuzdak, Hubert
Veröffentlicht: (2023)
von: Siuzdak, Hubert
Veröffentlicht: (2023)
Online incremental learning for audio classification using a pretrained audio model
von: Mulimani, Manjunath, et al.
Veröffentlicht: (2025)
von: Mulimani, Manjunath, et al.
Veröffentlicht: (2025)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
ICGAN: An implicit conditioning method for interpretable feature control of neural audio synthesis
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
von: Liu, Yunyi, et al.
Veröffentlicht: (2024)
Meta-learning-based percussion transcription and $t\bar{a}la$ identification from low-resource audio
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2025)
von: Kodag, Rahul Bapusaheb, et al.
Veröffentlicht: (2025)
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
A Universal Harmonic Discriminator for High-quality GAN-based Vocoder
von: Xu, Nan, et al.
Veröffentlicht: (2025)
von: Xu, Nan, et al.
Veröffentlicht: (2025)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
von: Fang, Yangui, et al.
Veröffentlicht: (2025)
WavLM model ensemble for audio deepfake detection
von: Combei, David, et al.
Veröffentlicht: (2024)
von: Combei, David, et al.
Veröffentlicht: (2024)
Cryfish: On deep audio analysis with Large Language Models
von: Mitrofanov, Anton, et al.
Veröffentlicht: (2025)
von: Mitrofanov, Anton, et al.
Veröffentlicht: (2025)
Multiple Hankel matrix rank minimization for audio inpainting
von: Záviška, Pavel, et al.
Veröffentlicht: (2023)
von: Záviška, Pavel, et al.
Veröffentlicht: (2023)
Unmasking real-world audio deepfakes: A data-centric approach
von: Combei, David, et al.
Veröffentlicht: (2025)
von: Combei, David, et al.
Veröffentlicht: (2025)
SCORE: Scaling audio generation using Standardized COmposite REwards
von: Jung, Jaemin, et al.
Veröffentlicht: (2025)
von: Jung, Jaemin, et al.
Veröffentlicht: (2025)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
von: Ma, Wenbo, et al.
Veröffentlicht: (2024)
UniFlow: Unifying Speech Front-End Tasks via Continuous Generative Modeling
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
von: Wang, Ziqian, et al.
Veröffentlicht: (2025)
STASE: A spatialized text-to-audio synthesis engine for music generation
von: Chi, Tutti, et al.
Veröffentlicht: (2025)
von: Chi, Tutti, et al.
Veröffentlicht: (2025)
Sound event detection with audio-text models and heterogeneous temporal annotations
von: Harju, Manu, et al.
Veröffentlicht: (2025)
von: Harju, Manu, et al.
Veröffentlicht: (2025)
Multi-label audio classification with a noisy zero-shot teacher
von: Braun, Sebastian, et al.
Veröffentlicht: (2024)
von: Braun, Sebastian, et al.
Veröffentlicht: (2024)
Understanding the strengths and weaknesses of SSL models for audio deepfake model attribution
von: Pîrlogeanu, Gabriel, et al.
Veröffentlicht: (2026)
von: Pîrlogeanu, Gabriel, et al.
Veröffentlicht: (2026)
AudioMorphix: Training-free audio editing with diffusion probabilistic models
von: Liang, Jinhua, et al.
Veröffentlicht: (2025)
von: Liang, Jinhua, et al.
Veröffentlicht: (2025)
Towards audio language modeling -- an overview
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Positive and negative sampling strategies for self-supervised learning on audio-video data
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
von: Wang, Shanshan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An Intra-BRNN and GB-RVQ Based END-TO-END Neural Audio Codec
von: Xu, Linping, et al.
Veröffentlicht: (2024) -
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
von: Mack, Wolfgang, et al.
Veröffentlicht: (2025) -
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023) -
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024) -
Modeling strategies for speech enhancement in the latent space of a neural audio codec
von: Kammoun, Sofiene, et al.
Veröffentlicht: (2025)