OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kim, Seunghee, Bang, Ingyu, Jang, Seokgyu, Kim, Changhyeon, Bae, Sanghwan, Choi, Jihun, Xuan, Richeng, Kim, Taeuk
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908996262166528
author Kim, Seunghee
Bang, Ingyu
Jang, Seokgyu
Kim, Changhyeon
Bae, Sanghwan
Choi, Jihun
Xuan, Richeng
Kim, Taeuk
author_facet Kim, Seunghee
Bang, Ingyu
Jang, Seokgyu
Kim, Changhyeon
Bae, Sanghwan
Choi, Jihun
Xuan, Richeng
Kim, Taeuk
contents Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models suffer from critical limitations, including modality shortcuts and biased reasoning paths. To address these challenges, we propose OMHBench, a novel benchmark designed to rigorously evaluate omni-modal multi-hop reasoning. It consists of 6,144 questions with balanced reasoning paths that are jointly grounded across all three modalities. Extensive evaluation of 13 state-of-the-art models reveals that (1) a large performance gap exists between proprietary and open-source MLLMs and (2) even proprietary models exhibit high sensitivity to reasoning path variations, resulting in asymmetric omni-modal grounding. Notably, models struggle when processing the speech modality, underscoring the need for balanced, multi-hop evaluation of omni-modal intelligence.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16198
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning
Kim, Seunghee
Bang, Ingyu
Jang, Seokgyu
Kim, Changhyeon
Bae, Sanghwan
Choi, Jihun
Xuan, Richeng
Kim, Taeuk
Computation and Language
Multimodal Large Language Models (MLLMs) have increasingly supported omni-modal processing across text, vision, and speech. However, existing evaluation frameworks for such models suffer from critical limitations, including modality shortcuts and biased reasoning paths. To address these challenges, we propose OMHBench, a novel benchmark designed to rigorously evaluate omni-modal multi-hop reasoning. It consists of 6,144 questions with balanced reasoning paths that are jointly grounded across all three modalities. Extensive evaluation of 13 state-of-the-art models reveals that (1) a large performance gap exists between proprietary and open-source MLLMs and (2) even proprietary models exhibit high sensitivity to reasoning path variations, resulting in asymmetric omni-modal grounding. Notably, models struggle when processing the speech modality, underscoring the need for balanced, multi-hop evaluation of omni-modal intelligence.
title OMHBench: Benchmarking Balanced and Grounded Omni-Modal Multi-Hop Reasoning
topic Computation and Language
url https://arxiv.org/abs/2508.16198