MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
Fuente:
arXiv
Saved in:
| Main Authors: | Selvakumar, Ramaneswaran, Seth, Ashish, Anand, Nishit, Tyagi, Utkarsh, Kumar, Sonal, Ghosh, Sreyan, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025)
by: Seth, Ashish, et al.
Published: (2025)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
by: Sakshi, S, et al.
Published: (2024)
by: Sakshi, S, et al.
Published: (2024)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
by: Seth, Ashish, et al.
Published: (2024)
by: Seth, Ashish, et al.
Published: (2024)
Do Audio-Language Models Understand Linguistic Variations?
by: Selvakumar, Ramaneswaran, et al.
Published: (2024)
by: Selvakumar, Ramaneswaran, et al.
Published: (2024)
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP
by: Evuru, Chandra Kiran Reddy, et al.
Published: (2024)
by: Evuru, Chandra Kiran Reddy, et al.
Published: (2024)
Do Vision-Language Models Understand Compound Nouns?
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
by: Ghosh, Sreyan, et al.
Published: (2024)
by: Ghosh, Sreyan, et al.
Published: (2024)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
by: Ghosh, Sreyan, et al.
Published: (2023)
by: Ghosh, Sreyan, et al.
Published: (2023)
Memento: Augmenting Personalized Memory via Practical Multimodal Wearable Sensing in Visual Search and Wayfinding Navigation
by: Ghosh, Indrajeet, et al.
Published: (2025)
by: Ghosh, Indrajeet, et al.
Published: (2025)
Towards Aligning Multimodal LLMs with Human Experts: A Focus on Parent-Child Interaction
by: Shi, Weiyan, et al.
Published: (2025)
by: Shi, Weiyan, et al.
Published: (2025)
Vocalize: Lead Acquisition and User Engagement through Gamified Voice Competitions
by: Teskeredzic, Edvin, et al.
Published: (2025)
by: Teskeredzic, Edvin, et al.
Published: (2025)
INDCOR white paper 4: Evaluation of Interactive Narrative Design For Complexity Representations
by: Roth, Christian, et al.
Published: (2023)
by: Roth, Christian, et al.
Published: (2023)
Towards Interactive Multimodal Representation of ML Functions for Human Understanding of ML
by: Wang, Bokang, et al.
Published: (2026)
by: Wang, Bokang, et al.
Published: (2026)
Private Chat in a Public Space of Metaverse Systems
by: Chen, Jiarui, et al.
Published: (2025)
by: Chen, Jiarui, et al.
Published: (2025)
Coordinated 2D-3D Visualization of Volumetric Medical Data in XR with Multimodal Interactions
by: Liu, Qixuan, et al.
Published: (2025)
by: Liu, Qixuan, et al.
Published: (2025)
Explainable Multimodal Emotion Recognition
by: Lian, Zheng, et al.
Published: (2023)
by: Lian, Zheng, et al.
Published: (2023)
Multimodal Infusion Tuning for Large Models
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
by: Chen, Maximillian, et al.
Published: (2026)
by: Chen, Maximillian, et al.
Published: (2026)
Ichiyo: Fragile and Transient Interaction in Neighborhood
by: Shibata, Hirofumi, et al.
Published: (2025)
by: Shibata, Hirofumi, et al.
Published: (2025)
Language-Guided Multimodal Texture Authoring via Generative Models
by: Qian, Wanli, et al.
Published: (2026)
by: Qian, Wanli, et al.
Published: (2026)
From Multimodal Signals to Adaptive XR Experiences for De-escalation Training
by: Nierula, Birgit, et al.
Published: (2026)
by: Nierula, Birgit, et al.
Published: (2026)
THE WASTIVE: An Interactive Ebb and Flow of Digital Fabrication Waste
by: Shan, Yifan, et al.
Published: (2025)
by: Shan, Yifan, et al.
Published: (2025)
Whispering Water: Materializing Human-AI Dialogue as Interactive Ripples
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
MuDoC: An Interactive Multimodal Document-grounded Conversational AI System
by: Taneja, Karan, et al.
Published: (2025)
by: Taneja, Karan, et al.
Published: (2025)
When Drawing Is Not Enough: Exploring Spontaneous Speech with Sketch for Intent Alignment in Multimodal LLMs
by: Shi, Weiyan, et al.
Published: (2026)
by: Shi, Weiyan, et al.
Published: (2026)
MULTI-CASE: A Transformer-based Ethics-aware Multimodal Investigative Intelligence Framework
by: Fischer, Maximilian T., et al.
Published: (2024)
by: Fischer, Maximilian T., et al.
Published: (2024)
Winds Through Time: Interactive Data Visualization and Physicalization for Paleoclimate Communication
by: Hunter, David, et al.
Published: (2025)
by: Hunter, David, et al.
Published: (2025)
ICE: Interactive 3D Game Character Editing via Dialogue
by: Wu, Haoqian, et al.
Published: (2024)
by: Wu, Haoqian, et al.
Published: (2024)
Dynamic Multimodal Expression Generation for LLM-Driven Pedagogical Agents: From User Experience Perspective
by: Wan, Ninghao, et al.
Published: (2026)
by: Wan, Ninghao, et al.
Published: (2026)
The Rhythm of Tai Chi: Revitalizing Cultural Heritage in Virtual Reality through Interactive Visuals
by: Wang, Xianghan
Published: (2025)
by: Wang, Xianghan
Published: (2025)
BATON: A Multimodal Benchmark for Bidirectional Automation Transition Observation in Naturalistic Driving
by: Wang, Yuhang, et al.
Published: (2026)
by: Wang, Yuhang, et al.
Published: (2026)
Fostering Emotional Perspective-Taking: An Exploration of Affective Face-Tracking Interactions in the VR Narrative Rekindle
by: Fan, Hector, et al.
Published: (2026)
by: Fan, Hector, et al.
Published: (2026)
From Perception to Cognition: How Latency Affects Interaction Fluency and Social Presence in VR Conferencing
by: Song, Jiarun, et al.
Published: (2026)
by: Song, Jiarun, et al.
Published: (2026)
SPICA: Interactive Video Content Exploration through Augmented Audio Descriptions for Blind or Low-Vision Viewers
by: Ning, Zheng, et al.
Published: (2024)
by: Ning, Zheng, et al.
Published: (2024)
Multimodal Digital Sensing of Early-Life Laying Hens: A Pilot Study Integrating Thermal, Acoustic, Optical-Flow and Environmental Data
by: Dhaliwal, Yashan, et al.
Published: (2026)
by: Dhaliwal, Yashan, et al.
Published: (2026)
Evaluating the Usability of Microgestures for Text Editing Tasks in Virtual Reality
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Multimodal Dialog Systems with Dual Knowledge-enhanced Generative Pretrained Language Model
by: Chen, Xiaolin, et al.
Published: (2022)
by: Chen, Xiaolin, et al.
Published: (2022)
Similar Items
-
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2025) -
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
by: Sakshi, S, et al.
Published: (2024) -
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
by: Seth, Ashish, et al.
Published: (2024) -
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
by: Seth, Ashish, et al.
Published: (2024) -
Do Audio-Language Models Understand Linguistic Variations?
by: Selvakumar, Ramaneswaran, et al.
Published: (2024)