HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Chen, Yang, Peiji, Zhong, Yicheng, Yu, Jianxing, Wang, Zhisheng, Gou, Zihao, Chen, Wenqing, Yin, Jian |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
par: Zhong, Yicheng, et autres
Publié: (2025)
par: Zhong, Yicheng, et autres
Publié: (2025)
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
par: Yang, Peiji, et autres
Publié: (2024)
par: Yang, Peiji, et autres
Publié: (2024)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
par: Li, Weiqin, et autres
Publié: (2024)
par: Li, Weiqin, et autres
Publié: (2024)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
par: Wang, Dingdong, et autres
Publié: (2025)
par: Wang, Dingdong, et autres
Publié: (2025)
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
par: Arora, Siddhant, et autres
Publié: (2024)
par: Arora, Siddhant, et autres
Publié: (2024)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
par: Wang, Dingdong, et autres
Publié: (2025)
par: Wang, Dingdong, et autres
Publié: (2025)
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
par: Yang, Fei, et autres
Publié: (2026)
par: Yang, Fei, et autres
Publié: (2026)
Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
par: Yin, Han, et autres
Publié: (2025)
par: Yin, Han, et autres
Publié: (2025)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
par: Huang, Kuan-Po, et autres
Publié: (2023)
par: Huang, Kuan-Po, et autres
Publié: (2023)
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
par: Li, Mohan, et autres
Publié: (2024)
par: Li, Mohan, et autres
Publié: (2024)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
par: Zhang, Xueyao, et autres
Publié: (2025)
par: Zhang, Xueyao, et autres
Publié: (2025)
OSUM-EChat: Enhancing End-to-End Empathetic Spoken Chatbot via Understanding-Driven Spoken Dialogue
par: Geng, Xuelong, et autres
Publié: (2025)
par: Geng, Xuelong, et autres
Publié: (2025)
SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words
par: Ao, Junyi, et autres
Publié: (2024)
par: Ao, Junyi, et autres
Publié: (2024)
A Unified Spoken Language Model with Injected Emotional-Attribution Thinking for Human-like Interaction
par: Wang, Qing, et autres
Publié: (2026)
par: Wang, Qing, et autres
Publié: (2026)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
par: Jung, Yeonjoon, et autres
Publié: (2024)
par: Jung, Yeonjoon, et autres
Publié: (2024)
RSA-Bench: Benchmarking Audio Large Models in Real-World Acoustic Scenarios
par: Zhang, Yibo, et autres
Publié: (2026)
par: Zhang, Yibo, et autres
Publié: (2026)
On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation
par: Cheng, Changhao, et autres
Publié: (2026)
par: Cheng, Changhao, et autres
Publié: (2026)
VoxMind: An End-to-End Agentic Spoken Dialogue System
par: Liang, Tianle, et autres
Publié: (2026)
par: Liang, Tianle, et autres
Publié: (2026)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
par: Mousavi, Pooneh, et autres
Publié: (2025)
par: Mousavi, Pooneh, et autres
Publié: (2025)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
par: Ma, Ziyang, et autres
Publié: (2024)
par: Ma, Ziyang, et autres
Publié: (2024)
How Well Do Current Speech Deepfake Detection Methods Generalize to the Real World?
par: Li, Daixian, et autres
Publié: (2026)
par: Li, Daixian, et autres
Publié: (2026)
R2-SVC: Towards Real-World Robust and Expressive Zero-shot Singing Voice Conversion
par: Zheng, Junjie, et autres
Publié: (2025)
par: Zheng, Junjie, et autres
Publié: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
par: Tseng, Liang-Hsuan, et autres
Publié: (2025)
par: Tseng, Liang-Hsuan, et autres
Publié: (2025)
"Alexa, can you forget me?" Machine Unlearning Benchmark in Spoken Language Understanding
par: Koudounas, Alkis, et autres
Publié: (2025)
par: Koudounas, Alkis, et autres
Publié: (2025)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
par: Cui, Wenqian, et autres
Publié: (2025)
par: Cui, Wenqian, et autres
Publié: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
par: Hu, Jiliang, et autres
Publié: (2025)
par: Hu, Jiliang, et autres
Publié: (2025)
TurnGuide: Enhancing Meaningful Full Duplex Spoken Interactions via Dynamic Turn-Level Text-Speech Interleaving
par: Cui, Wenqian, et autres
Publié: (2025)
par: Cui, Wenqian, et autres
Publié: (2025)
Evaluating and Improving Continual Learning in Spoken Language Understanding
par: Yang, Muqiao, et autres
Publié: (2024)
par: Yang, Muqiao, et autres
Publié: (2024)
Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India
par: Bhogale, Kaushal, et autres
Publié: (2026)
par: Bhogale, Kaushal, et autres
Publié: (2026)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
par: Wang, Yujin, et autres
Publié: (2022)
par: Wang, Yujin, et autres
Publié: (2022)
MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix
par: Ma, Ziyang, et autres
Publié: (2025)
par: Ma, Ziyang, et autres
Publié: (2025)
TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
par: Li, Zehan, et autres
Publié: (2025)
par: Li, Zehan, et autres
Publié: (2025)
Long-Form Speech Generation with Spoken Language Models
par: Park, Se Jin, et autres
Publié: (2024)
par: Park, Se Jin, et autres
Publié: (2024)
RTCFake: Speech Deepfake Detection in Real-Time Communication
par: Xue, Jun, et autres
Publié: (2026)
par: Xue, Jun, et autres
Publié: (2026)
LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
par: Fang, Qingkai, et autres
Publié: (2025)
par: Fang, Qingkai, et autres
Publié: (2025)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
par: Lin, Chyi-Jiunn, et autres
Publié: (2024)
par: Lin, Chyi-Jiunn, et autres
Publié: (2024)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
par: Shon, Suwon, et autres
Publié: (2024)
par: Shon, Suwon, et autres
Publié: (2024)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
par: Li, Tao, et autres
Publié: (2025)
par: Li, Tao, et autres
Publié: (2025)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
par: Lou, Yuxuan, et autres
Publié: (2026)
par: Lou, Yuxuan, et autres
Publié: (2026)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
par: Ma, Ziyang, et autres
Publié: (2023)
par: Ma, Ziyang, et autres
Publié: (2023)
Documents similaires
-
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
par: Zhong, Yicheng, et autres
Publié: (2025) -
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
par: Yang, Peiji, et autres
Publié: (2024) -
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
par: Li, Weiqin, et autres
Publié: (2024) -
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
par: Wang, Dingdong, et autres
Publié: (2025) -
On the Evaluation of Speech Foundation Models for Spoken Language Understanding
par: Arora, Siddhant, et autres
Publié: (2024)