MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Selvakumar, Ramaneswaran, Seth, Ashish, Anand, Nishit, Tyagi, Utkarsh, Kumar, Sonal, Ghosh, Sreyan, Manocha, Dinesh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912605755408384
author Selvakumar, Ramaneswaran
Seth, Ashish
Anand, Nishit
Tyagi, Utkarsh
Kumar, Sonal
Ghosh, Sreyan
Manocha, Dinesh
author_facet Selvakumar, Ramaneswaran
Seth, Ashish
Anand, Nishit
Tyagi, Utkarsh
Kumar, Sonal
Ghosh, Sreyan
Manocha, Dinesh
contents The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchmarks fall short in comprehensively evaluating how well these models generate context-aware responses, particularly when it comes to implicitly understanding fine-grained speech characteristics, such as pitch, emotion, timbre, and volume or the environmental acoustic context such as background sounds. Additionally, they inadequately assess the ability of models to align paralinguistic cues with complementary visual signals to inform their responses. To address these gaps, we introduce MultiVox, the first omni voice assistant benchmark designed to evaluate the ability of voice assistants to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. Specifically, MultiVox includes 1000 human-annotated and recorded speech dialogues that encompass diverse paralinguistic features and a range of visual cues such as images and videos. Our evaluation on 10 state-of-the-art models reveals that, although humans excel at these tasks, current models consistently struggle to produce contextually grounded responses.
format Preprint
id arxiv_https___arxiv_org_abs_2507_10859
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
Selvakumar, Ramaneswaran
Seth, Ashish
Anand, Nishit
Tyagi, Utkarsh
Kumar, Sonal
Ghosh, Sreyan
Manocha, Dinesh
Multimedia
Computation and Language
Human-Computer Interaction
The rapid progress of Large Language Models (LLMs) has empowered omni models to act as voice assistants capable of understanding spoken dialogues. These models can process multimodal inputs beyond text, such as speech and visual data, enabling more context-aware interactions. However, current benchmarks fall short in comprehensively evaluating how well these models generate context-aware responses, particularly when it comes to implicitly understanding fine-grained speech characteristics, such as pitch, emotion, timbre, and volume or the environmental acoustic context such as background sounds. Additionally, they inadequately assess the ability of models to align paralinguistic cues with complementary visual signals to inform their responses. To address these gaps, we introduce MultiVox, the first omni voice assistant benchmark designed to evaluate the ability of voice assistants to integrate spoken and visual cues including paralinguistic speech features for truly multimodal understanding. Specifically, MultiVox includes 1000 human-annotated and recorded speech dialogues that encompass diverse paralinguistic features and a range of visual cues such as images and videos. Our evaluation on 10 state-of-the-art models reveals that, although humans excel at these tasks, current models consistently struggle to produce contextually grounded responses.
title MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
topic Multimedia
Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2507.10859