VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Heyang, Wang, Yuhao, Cheng, Ziyang, Liu, Hongcheng, Li, Yiqi, Hou, Yixuan, Wu, Ronghua, Gu, Qunshan, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
by: Cheng, Ziyang, et al.
Published: (2026)
by: Cheng, Ziyang, et al.
Published: (2026)
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
by: Hou, Yixuan, et al.
Published: (2025)
by: Hou, Yixuan, et al.
Published: (2025)
VocalNet-M2: Advancing Low-Latency Spoken Language Modeling via Integrated Multi-Codebook Tokenization and Multi-Token Prediction
by: Wang, Yuhao, et al.
Published: (2025)
by: Wang, Yuhao, et al.
Published: (2025)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
LaSR: Context-Aware Speech Recognition via Latent Reasoning
by: Liu, Heyang, et al.
Published: (2026)
by: Liu, Heyang, et al.
Published: (2026)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Decoding Linguistic Representations of Human Brain
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
by: Liu, Heyang, et al.
Published: (2024)
by: Liu, Heyang, et al.
Published: (2024)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
by: Liu, Hongcheng, et al.
Published: (2026)
by: Liu, Hongcheng, et al.
Published: (2026)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
Pet-Bench: Benchmarking the Abilities of Large Language Models as E-Pets in Social Network Services
by: Guo, Hongcheng, et al.
Published: (2025)
by: Guo, Hongcheng, et al.
Published: (2025)
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
by: Ni, Qinke, et al.
Published: (2026)
by: Ni, Qinke, et al.
Published: (2026)
Coding Speech through Vocal Tract Kinematics
by: Cho, Cheol Jun, et al.
Published: (2024)
by: Cho, Cheol Jun, et al.
Published: (2024)
Toward a Consensus Description of Vocal Effort, Vocal Load, Vocal Loading, and Vocal Fatigue
by: Hunter, Eric J., et al.
Published: (2020)
by: Hunter, Eric J., et al.
Published: (2020)
Towards an End-to-End Framework for Invasive Brain Signal Decoding with Large Language Models
by: Feng, Sheng, et al.
Published: (2024)
by: Feng, Sheng, et al.
Published: (2024)
LibriSQA: A Novel Dataset and Framework for Spoken Question Answering with Large Language Models
by: Zhao, Zihan, et al.
Published: (2023)
by: Zhao, Zihan, et al.
Published: (2023)
NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
by: Xue, Liumeng, et al.
Published: (2026)
by: Xue, Liumeng, et al.
Published: (2026)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
Infrequent Child-Directed Speech Is Bursty and May Draw Infant Vocalizations
by: Cychosz, Margaret, et al.
Published: (2026)
by: Cychosz, Margaret, et al.
Published: (2026)
MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation
by: Chen, Szu-Chi, et al.
Published: (2026)
by: Chen, Szu-Chi, et al.
Published: (2026)
Multimodal Segmentation for Vocal Tract Modeling
by: Jain, Rishi, et al.
Published: (2024)
by: Jain, Rishi, et al.
Published: (2024)
WebExpert: domain-aware web agents with critic-guided expert experience for high-precision search
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
by: Zhang, Linhao, et al.
Published: (2025)
by: Zhang, Linhao, et al.
Published: (2025)
SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
by: Guo, Hongcheng, et al.
Published: (2025)
by: Guo, Hongcheng, et al.
Published: (2025)
TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games
by: Mishra, Prakamya, et al.
Published: (2025)
by: Mishra, Prakamya, et al.
Published: (2025)
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
by: Chen, Zhe, et al.
Published: (2025)
by: Chen, Zhe, et al.
Published: (2025)
Vocalization in Group Writing
by: Kilani, Marwan
Published: (2020)
by: Kilani, Marwan
Published: (2020)
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
by: Yang, Siwei, et al.
Published: (2024)
by: Yang, Siwei, et al.
Published: (2024)
Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances
by: Wu, Zehui, et al.
Published: (2024)
by: Wu, Zehui, et al.
Published: (2024)
Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
by: Wang, Pingjie, et al.
Published: (2025)
by: Wang, Pingjie, et al.
Published: (2025)
LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
by: Gui, Jiayi, et al.
Published: (2024)
by: Gui, Jiayi, et al.
Published: (2024)
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
by: Qiu, Pengcheng, et al.
Published: (2025)
by: Qiu, Pengcheng, et al.
Published: (2025)
Similar Items
-
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025) -
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025) -
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
by: Cheng, Ziyang, et al.
Published: (2026) -
VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation
by: Wang, Yuhao, et al.
Published: (2025) -
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
by: Hou, Yixuan, et al.
Published: (2025)