Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Dingdong, Li, Junan, Cui, Mingyu, Yang, Dongchao, Chen, Xueyuan, Meng, Helen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
von: Wang, Dingdong, et al.
Veröffentlicht: (2024)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
von: Song, Yuhan, et al.
Veröffentlicht: (2025)
von: Song, Yuhan, et al.
Veröffentlicht: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Speech-Worthy Alignment for Japanese SpeechLLMs via Direct Preference Optimization
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)
von: Zhao, Mengjie, et al.
Veröffentlicht: (2026)
Exploring SSL Discrete Tokens for Multilingual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2026)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
von: Arora, Siddhant, et al.
Veröffentlicht: (2024)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2025)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
von: Mousavi, Pooneh, et al.
Veröffentlicht: (2025)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
von: Tseng, Liang-Hsuan, et al.
Veröffentlicht: (2026)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
DOA: Training-Free Decoder-Only Attention Policy for Long-Form Simultaneous Translation with SpeechLLMs
von: Papi, Sara, et al.
Veröffentlicht: (2026)
von: Papi, Sara, et al.
Veröffentlicht: (2026)
Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2025)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
von: Wang, Qichao, et al.
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
von: Yang, Dongchao, et al.
Veröffentlicht: (2024)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
von: Jung, Yeonjoon, et al.
Veröffentlicht: (2024)
Rubric-Guided Fine-tuning of SpeechLLMs for Multi-Aspect, Multi-Rater L2 Reading-Speech Assessment
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2026)
Continual Speech Learning with Fused Speech Features
von: Wang, Guitao, et al.
Veröffentlicht: (2025)
von: Wang, Guitao, et al.
Veröffentlicht: (2025)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
von: Chang, Heng-Jui, et al.
Veröffentlicht: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2025)
Benchmarking Prosody Encoding in Discrete Speech Tokens
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
Do Bias Benchmarks Generalise? Evidence from Voice-based Evaluation of Gender Bias in SpeechLLMs
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
von: Satish, Shree Harsha Bokkahalli, et al.
Veröffentlicht: (2025)
Long-Form Speech Generation with Spoken Language Models
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
EmotionThinker: Prosody-Aware Reinforcement Learning for Explainable Speech Emotion Reasoning
von: Wang, Dingdong, et al.
Veröffentlicht: (2026)
von: Wang, Dingdong, et al.
Veröffentlicht: (2026)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Rethinking Discrete Speech Representation Tokens for Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
Children's Speech Recognition through Discrete Token Enhancement
von: Sukhadia, Vrunda N., et al.
Veröffentlicht: (2024)
von: Sukhadia, Vrunda N., et al.
Veröffentlicht: (2024)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
S2ST-Omni: Hierarchical Language-Aware SpeechLLM Adaptation for Multilingual Speech-to-Speech Translation
von: Pan, Yu, et al.
Veröffentlicht: (2025)
von: Pan, Yu, et al.
Veröffentlicht: (2025)
Evaluating and Improving Continual Learning in Spoken Language Understanding
von: Yang, Muqiao, et al.
Veröffentlicht: (2024)
von: Yang, Muqiao, et al.
Veröffentlicht: (2024)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
von: Wang, Dingdong, et al.
Veröffentlicht: (2024) -
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
von: Wang, Dingdong, et al.
Veröffentlicht: (2025) -
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs
von: Song, Yuhan, et al.
Veröffentlicht: (2025) -
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
von: Yang, Dongchao, et al.
Veröffentlicht: (2024) -
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)