DM-Codec: Distilling Multimodal Representations for Speech Tokenization
Fuente:
arXiv
Saved in:
| Main Authors: | Ahasan, Md Mubtasim, Fahim, Md, Mohiuddin, Tasnim, Rahman, A K M Mahbubur, Chadha, Aman, Iqbal, Tariq, Amin, M Ashraful, Islam, Md Mofijul, Ali, Amin Ahsan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
by: Ahasan, Md Mubtasim, et al.
Published: (2025)
by: Ahasan, Md Mubtasim, et al.
Published: (2025)
BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
by: Ahmed, Fahim, et al.
Published: (2025)
by: Ahmed, Fahim, et al.
Published: (2025)
SloMo-Fast: Slow-Momentum and Fast-Adaptive Teachers for Source-Free Continual Test-Time Adaptation
by: Iftee, Md Akil Raihan, et al.
Published: (2025)
by: Iftee, Md Akil Raihan, et al.
Published: (2025)
Embodied Referring Expression Comprehension in Human-Robot Interaction
by: Islam, Md Mofijul, et al.
Published: (2025)
by: Islam, Md Mofijul, et al.
Published: (2025)
RepCodec: A Speech Representation Codec for Speech Tokenization
by: Huang, Zhichao, et al.
Published: (2023)
by: Huang, Zhichao, et al.
Published: (2023)
Cognitively Inspired Energy-Based World Models
by: Gladstone, Alexi, et al.
Published: (2024)
by: Gladstone, Alexi, et al.
Published: (2024)
Cyclostationarity Analysis as a Complement to Self-Supervised Representations for Speech Deepfake Detection
by: Hanilçi, Cemal, et al.
Published: (2026)
by: Hanilçi, Cemal, et al.
Published: (2026)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2025)
BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech
by: Mohammad, Mir Sayeed, et al.
Published: (2024)
by: Mohammad, Mir Sayeed, et al.
Published: (2024)
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
by: Ridoy, Md Sazzadul Islam, et al.
Published: (2026)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
by: Ji, Shengpeng, et al.
Published: (2024)
by: Ji, Shengpeng, et al.
Published: (2024)
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
by: Ioannides, Georgios, et al.
Published: (2025)
by: Ioannides, Georgios, et al.
Published: (2025)
Multi-Level Embedding Conformer Framework for Bengali Automatic Speech Recognition
by: Sakib, Md. Nazmus, et al.
Published: (2025)
by: Sakib, Md. Nazmus, et al.
Published: (2025)
MathMist: A Parallel Multilingual Benchmark Dataset for Mathematical Problem Solving and Reasoning
by: Sobhani, Mahbub E, et al.
Published: (2025)
by: Sobhani, Mahbub E, et al.
Published: (2025)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
by: Jiang, Xue, et al.
Published: (2025)
by: Jiang, Xue, et al.
Published: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
by: Tseng, Wei-Cheng, et al.
Published: (2025)
by: Tseng, Wei-Cheng, et al.
Published: (2025)
Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
by: Nayeem, Md., et al.
Published: (2025)
by: Nayeem, Md., et al.
Published: (2025)
BanglaDialecto: An End-to-End AI-Powered Regional Speech Standardization
by: Samin, Md. Nazmus Sadat, et al.
Published: (2024)
by: Samin, Md. Nazmus Sadat, et al.
Published: (2024)
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains
by: Sayeedi, Md. Faiyaz Abdullah, et al.
Published: (2025)
by: Sayeedi, Md. Faiyaz Abdullah, et al.
Published: (2025)
On the Relation Between Speech Quality and Quantized Latent Representations of Neural Codecs
by: Halimeh, Mhd Modar, et al.
Published: (2025)
by: Halimeh, Mhd Modar, et al.
Published: (2025)
Distinctive Feature Codec: An Adaptive Efficient Speech Representation for Depression Detection
by: Zhang, Xiangyu, et al.
Published: (2025)
by: Zhang, Xiangyu, et al.
Published: (2025)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
by: Zhang, En-Wei, et al.
Published: (2025)
by: Zhang, En-Wei, et al.
Published: (2025)
Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
by: Li, Hanzhao, et al.
Published: (2024)
by: Li, Hanzhao, et al.
Published: (2024)
VoCodec: An Efficient Lightweight Low-Bitrate Speech Codec
by: Yang, Leyan, et al.
Published: (2026)
by: Yang, Leyan, et al.
Published: (2026)
CodeSep: Low-Bitrate Codec-Driven Speech Separation with Base-Token Disentanglement and Auxiliary-Token Serial Prediction
by: Du, Hui-Peng, et al.
Published: (2026)
by: Du, Hui-Peng, et al.
Published: (2026)
SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
by: Wang, Linqin, et al.
Published: (2024)
by: Wang, Linqin, et al.
Published: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
by: Wang, Yuancheng, et al.
Published: (2025)
by: Wang, Yuancheng, et al.
Published: (2025)
Frequency-Specific Neural Response and Cross-Correlation Analysis of Envelope Following Responses to Native Speech and Music Using Multichannel EEG Signals: A Case Study
by: Hasan, Md. Mahbub, et al.
Published: (2025)
by: Hasan, Md. Mahbub, et al.
Published: (2025)
Cochleagram-based Noise Adapted Speaker Identification System for Distorted Speech
by: Ahmed, Sabbir, et al.
Published: (2025)
by: Ahmed, Sabbir, et al.
Published: (2025)
WhisQ: Cross-Modal Representation Learning for Text-to-Music MOS Prediction
by: Emon, Jakaria Islam, et al.
Published: (2025)
by: Emon, Jakaria Islam, et al.
Published: (2025)
A Scoping Review of Deep Learning for Urban Visual Pollution and Proposal of a Real-Time Monitoring Framework with a Visual Pollution Index
by: Rahman, Mohammad Masudur, et al.
Published: (2026)
by: Rahman, Mohammad Masudur, et al.
Published: (2026)
Overview of Automatic Speech Analysis and Technologies for Neurodegenerative Disorders: Diagnosis and Assistive Applications
by: Sheikh, Shakeel A., et al.
Published: (2025)
by: Sheikh, Shakeel A., et al.
Published: (2025)
Density Adaptive Attention-based Speech Network: Enhancing Feature Understanding for Mental Health Disorders
by: Ioannides, Georgios, et al.
Published: (2024)
by: Ioannides, Georgios, et al.
Published: (2024)
Personalized Neural Speech Codec
by: Jang, Inseon, et al.
Published: (2024)
by: Jang, Inseon, et al.
Published: (2024)
A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
by: Islam, Rafiul, et al.
Published: (2026)
by: Islam, Rafiul, et al.
Published: (2026)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
by: Nguyen, Tan Dat, et al.
Published: (2024)
by: Nguyen, Tan Dat, et al.
Published: (2024)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
by: Gong, Yitian, et al.
Published: (2025)
by: Gong, Yitian, et al.
Published: (2025)
Exploring Disentangled Neural Speech Codecs from Self-Supervised Representations
by: Aihara, Ryo, et al.
Published: (2025)
by: Aihara, Ryo, et al.
Published: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
by: Xin, Detai, et al.
Published: (2024)
by: Xin, Detai, et al.
Published: (2024)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
by: Shi, Jiatong, et al.
Published: (2025)
by: Shi, Jiatong, et al.
Published: (2025)
Similar Items
-
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
by: Ahasan, Md Mubtasim, et al.
Published: (2025) -
BAPPA: Benchmarking Agents, Plans, and Pipelines for Automated Text-to-SQL Generation
by: Ahmed, Fahim, et al.
Published: (2025) -
SloMo-Fast: Slow-Momentum and Fast-Adaptive Teachers for Source-Free Continual Test-Time Adaptation
by: Iftee, Md Akil Raihan, et al.
Published: (2025) -
Embodied Referring Expression Comprehension in Human-Robot Interaction
by: Islam, Md Mofijul, et al.
Published: (2025) -
RepCodec: A Speech Representation Codec for Speech Tokenization
by: Huang, Zhichao, et al.
Published: (2023)