Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gu, Xiaojie, Tong, Sherry T., Feng, Aosong, Han, Sophia Simeng, Lu, Jinghui, Chen, Yingjian, Iwasawa, Yusuke, Matsuo, Yutaka, Park, Chanjun, Ying, Rex, Li, Irene
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913163649220608
author Gu, Xiaojie
Tong, Sherry T.
Feng, Aosong
Han, Sophia Simeng
Lu, Jinghui
Chen, Yingjian
Iwasawa, Yusuke
Matsuo, Yutaka
Park, Chanjun
Ying, Rex
Li, Irene
author_facet Gu, Xiaojie
Tong, Sherry T.
Feng, Aosong
Han, Sophia Simeng
Lu, Jinghui
Chen, Yingjian
Iwasawa, Yusuke
Matsuo, Yutaka
Park, Chanjun
Ying, Rex
Li, Irene
contents Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where reasoning breaks down. Omanic contains 10,296 machine-generated training examples (OmanicSynth) and 967 expert-reviewed human-annotated evaluation examples (OmanicBench), with each evaluation question decomposed into single-hop sub-questions, intermediate answers, and structured graph topologies. Experiments with proprietary and open-source LLMs show that Omanic is challenging, while step-wise analysis reveals a later-hop bottleneck, factual knowledge floor, and error propagation along reasoning chains. Fine-tuning on OmanicSynth transfers to six reasoning and mathematics benchmarks, yielding a 7.41-point average gain and validating its effectiveness as supervision for reasoning-capability transfer. We release the data at https://huggingface.co/datasets/li-lab/Omanic and the code at https://github.com/XiaojieGu/Omanic.
format Preprint
id arxiv_https___arxiv_org_abs_2603_16654
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
Gu, Xiaojie
Tong, Sherry T.
Feng, Aosong
Han, Sophia Simeng
Lu, Jinghui
Chen, Yingjian
Iwasawa, Yusuke
Matsuo, Yutaka
Park, Chanjun
Ying, Rex
Li, Irene
Computation and Language
Artificial Intelligence
Machine Learning
Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where reasoning breaks down. Omanic contains 10,296 machine-generated training examples (OmanicSynth) and 967 expert-reviewed human-annotated evaluation examples (OmanicBench), with each evaluation question decomposed into single-hop sub-questions, intermediate answers, and structured graph topologies. Experiments with proprietary and open-source LLMs show that Omanic is challenging, while step-wise analysis reveals a later-hop bottleneck, factual knowledge floor, and error propagation along reasoning chains. Fine-tuning on OmanicSynth transfers to six reasoning and mathematics benchmarks, yielding a 7.41-point average gain and validating its effectiveness as supervision for reasoning-capability transfer. We release the data at https://huggingface.co/datasets/li-lab/Omanic and the code at https://github.com/XiaojieGu/Omanic.
title Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2603.16654