Speaking Clearly: A Simplified Whisper-Based Codec for Low-Bitrate Speech Coding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xin, Li, Lin, Lu, Xiangni, Liu, Jianquan, Lee, Kong Aik |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
par: Xin, Detai, et autres
Publié: (2024)
par: Xin, Detai, et autres
Publié: (2024)
MuCodec: Ultra Low-Bitrate Music Codec
par: Xu, Yaoxun, et autres
Publié: (2024)
par: Xu, Yaoxun, et autres
Publié: (2024)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
par: Gong, Yitian, et autres
Publié: (2025)
par: Gong, Yitian, et autres
Publié: (2025)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
par: Della Libera, Luca, et autres
Publié: (2025)
par: Della Libera, Luca, et autres
Publié: (2025)
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
par: Dong, Zhongren, et autres
Publié: (2025)
par: Dong, Zhongren, et autres
Publié: (2025)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
par: Jiang, Xue, et autres
Publié: (2025)
par: Jiang, Xue, et autres
Publié: (2025)
FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
par: Della Libera, Luca, et autres
Publié: (2025)
par: Della Libera, Luca, et autres
Publié: (2025)
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
par: Yang, Peiji, et autres
Publié: (2024)
par: Yang, Peiji, et autres
Publié: (2024)
Towards Bitrate-Efficient and Noise-Robust Speech Coding with Variable Bitrate RVQ
par: Chae, Yunkee, et autres
Publié: (2025)
par: Chae, Yunkee, et autres
Publié: (2025)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
par: Liu, Haohe, et autres
Publié: (2024)
par: Liu, Haohe, et autres
Publié: (2024)
UniSRCodec: Unified and Low-Bitrate Single Codebook Codec with Sub-Band Reconstruction
par: Zhang, Zhisheng, et autres
Publié: (2026)
par: Zhang, Zhisheng, et autres
Publié: (2026)
U3-xi: Pushing the Boundaries of Speaker Recognition by Incorporating Uncertainty
par: Li, Junjie, et autres
Publié: (2026)
par: Li, Junjie, et autres
Publié: (2026)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
par: Jiang, Xiao-Hang, et autres
Publié: (2024)
par: Jiang, Xiao-Hang, et autres
Publié: (2024)
Scaling Transformers for Low-Bitrate High-Quality Speech Coding
par: Parker, Julian D, et autres
Publié: (2024)
par: Parker, Julian D, et autres
Publié: (2024)
FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
par: Li, Jiaqi, et autres
Publié: (2025)
par: Li, Jiaqi, et autres
Publié: (2025)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
par: Chen, Xuanjun, et autres
Publié: (2025)
par: Chen, Xuanjun, et autres
Publié: (2025)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
par: Li, Jiaqi, et autres
Publié: (2025)
par: Li, Jiaqi, et autres
Publié: (2025)
PhoenixCodec: Taming Neural Speech Coding for Extreme Low-Resource Scenarios
par: Wan, Zixiang, et autres
Publié: (2025)
par: Wan, Zixiang, et autres
Publié: (2025)
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
par: Shi, Jiacheng, et autres
Publié: (2026)
par: Shi, Jiacheng, et autres
Publié: (2026)
LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
par: Luong, Hieu-Thi, et autres
Publié: (2024)
par: Luong, Hieu-Thi, et autres
Publié: (2024)
A Neural Speech Codec for Noise Robust Speech Coding
par: Huang, Jiayi, et autres
Publié: (2023)
par: Huang, Jiayi, et autres
Publié: (2023)
Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
par: Guo, Haohan, et autres
Publié: (2024)
par: Guo, Haohan, et autres
Publié: (2024)
Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
par: Zhang, Li, et autres
Publié: (2024)
par: Zhang, Li, et autres
Publié: (2024)
SpatialCodec: Neural Spatial Speech Coding
par: Xu, Zhongweiyang, et autres
Publié: (2023)
par: Xu, Zhongweiyang, et autres
Publié: (2023)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
par: Casanova, Edresson, et autres
Publié: (2024)
par: Casanova, Edresson, et autres
Publié: (2024)
Fewer-token Neural Speech Codec with Time-invariant Codes
par: Ren, Yong, et autres
Publié: (2023)
par: Ren, Yong, et autres
Publié: (2023)
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
par: Meng, Zhaoyang, et autres
Publié: (2026)
par: Meng, Zhaoyang, et autres
Publié: (2026)
STCTS: Generative Semantic Compression for Ultra-Low Bitrate Speech via Explicit Text-Prosody-Timbre Decomposition
par: Wang, Siyu, et autres
Publié: (2025)
par: Wang, Siyu, et autres
Publié: (2025)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
par: Li, Tao, et autres
Publié: (2025)
par: Li, Tao, et autres
Publié: (2025)
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
par: Yang, Xusheng, et autres
Publié: (2025)
par: Yang, Xusheng, et autres
Publié: (2025)
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
par: Truong, Duc-Tuan, et autres
Publié: (2025)
par: Truong, Duc-Tuan, et autres
Publié: (2025)
QAMO: Quality-aware Multi-centroid One-class Learning For Speech Deepfake Detection
par: Truong, Duc-Tuan, et autres
Publié: (2025)
par: Truong, Duc-Tuan, et autres
Publié: (2025)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
par: Li, Jiaqi, et autres
Publié: (2024)
par: Li, Jiaqi, et autres
Publié: (2024)
Neural Speech Coding for Real-time Communications using Constant Bitrate Scalar Quantization
par: Brendel, Andreas, et autres
Publié: (2024)
par: Brendel, Andreas, et autres
Publié: (2024)
VoxGenesis: Unsupervised Discovery of Latent Speaker Manifold for Speech Synthesis
par: Lin, Weiwei, et autres
Publié: (2024)
par: Lin, Weiwei, et autres
Publié: (2024)
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
par: Chen, Xuanjun, et autres
Publié: (2025)
par: Chen, Xuanjun, et autres
Publié: (2025)
Adversarial speech for voice privacy protection from Personalized Speech generation
par: Chen, Shihao, et autres
Publié: (2024)
par: Chen, Shihao, et autres
Publié: (2024)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
par: Shi, Jiatong, et autres
Publié: (2025)
par: Shi, Jiatong, et autres
Publié: (2025)
Cosine Scoring with Uncertainty for Neural Speaker Embedding
par: Wang, Qiongqiong, et autres
Publié: (2024)
par: Wang, Qiongqiong, et autres
Publié: (2024)
Documents similaires
-
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
par: Xin, Detai, et autres
Publié: (2024) -
MuCodec: Ultra Low-Bitrate Music Codec
par: Xu, Yaoxun, et autres
Publié: (2024) -
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
par: Gong, Yitian, et autres
Publié: (2025) -
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024) -
FocalCodec: Low-Bitrate Speech Coding via Focal Modulation Networks
par: Della Libera, Luca, et autres
Publié: (2025)