VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yuhao, Cheng, Ziyang, Liu, Heyang, Wu, Ronghua, Gu, Qunshan, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
von: Wang, Yuhao, et al.
Veröffentlicht: (2025)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
von: Hou, Yixuan, et al.
Veröffentlicht: (2025)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
von: Liu, Heyang, et al.
Veröffentlicht: (2025)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
Kanade: A Simple Disentangled Tokenizer for Spoken Language Modeling
von: Huang, Zhijie, et al.
Veröffentlicht: (2026)
von: Huang, Zhijie, et al.
Veröffentlicht: (2026)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means
von: Onda, Kentaro, et al.
Veröffentlicht: (2026)
von: Onda, Kentaro, et al.
Veröffentlicht: (2026)
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
von: Feng, Sheng, et al.
Veröffentlicht: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
LaSR: Context-Aware Speech Recognition via Latent Reasoning
von: Liu, Heyang, et al.
Veröffentlicht: (2026)
von: Liu, Heyang, et al.
Veröffentlicht: (2026)
MelTok: 2D Tokenization for Single-Codebook Audio Compression
von: Li, Jingyi, et al.
Veröffentlicht: (2025)
von: Li, Jingyi, et al.
Veröffentlicht: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
DuoTok: Source-Aware Dual-Track Tokenization for Multi-Track Music Language Modeling
von: Lin, Rui, et al.
Veröffentlicht: (2025)
von: Lin, Rui, et al.
Veröffentlicht: (2025)
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
von: Fang, Zheng, et al.
Veröffentlicht: (2026)
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
von: Yang, Shu-wen, et al.
Veröffentlicht: (2025)
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models
von: Lin, Yu-Xiang, et al.
Veröffentlicht: (2025)
von: Lin, Yu-Xiang, et al.
Veröffentlicht: (2025)
Multi-Class-Token Transformer for Multitask Self-supervised Music Information Retrieval
von: Kong, Yuexuan, et al.
Veröffentlicht: (2025)
von: Kong, Yuexuan, et al.
Veröffentlicht: (2025)
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
von: Xue, Liumeng, et al.
Veröffentlicht: (2026)
von: Xue, Liumeng, et al.
Veröffentlicht: (2026)
Decoding Linguistic Representations of Human Brain
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
Addressing Index Collapse of Large-Codebook Speech Tokenizer with Dual-Decoding Product-Quantized Variational Auto-Encoder
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
von: Song, Yuhan, et al.
Veröffentlicht: (2025)
von: Song, Yuhan, et al.
Veröffentlicht: (2025)
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
von: Geng, Xuelong, et al.
Veröffentlicht: (2025)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
Geolocation-Aware Robust Spoken Language Identification
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
von: Zhou, Yixuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yixuan, et al.
Veröffentlicht: (2025)
TokenChain: A Discrete Speech Chain via Semantic Token Modeling
von: Wang, Mingxuan, et al.
Veröffentlicht: (2025)
von: Wang, Mingxuan, et al.
Veröffentlicht: (2025)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Unmute the Patch Tokens: Rethinking Probing in Multi-Label Audio Classification
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
von: Rauch, Lukas, et al.
Veröffentlicht: (2025)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
von: Yang, Dongchao, et al.
Veröffentlicht: (2026)
von: Yang, Dongchao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
von: Wang, Yuhao, et al.
Veröffentlicht: (2025) -
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
von: Cheng, Ziyang, et al.
Veröffentlicht: (2026) -
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
von: Hou, Yixuan, et al.
Veröffentlicht: (2025) -
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
von: Liu, Heyang, et al.
Veröffentlicht: (2025) -
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
von: Liu, Heyang, et al.
Veröffentlicht: (2025)