Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913163649220608 |
|---|---|
| author | Gu, Xiaojie Tong, Sherry T. Feng, Aosong Han, Sophia Simeng Lu, Jinghui Chen, Yingjian Iwasawa, Yusuke Matsuo, Yutaka Park, Chanjun Ying, Rex Li, Irene |
| author_facet | Gu, Xiaojie Tong, Sherry T. Feng, Aosong Han, Sophia Simeng Lu, Jinghui Chen, Yingjian Iwasawa, Yusuke Matsuo, Yutaka Park, Chanjun Ying, Rex Li, Irene |
| contents | Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where reasoning breaks down. Omanic contains 10,296 machine-generated training examples (OmanicSynth) and 967 expert-reviewed human-annotated evaluation examples (OmanicBench), with each evaluation question decomposed into single-hop sub-questions, intermediate answers, and structured graph topologies. Experiments with proprietary and open-source LLMs show that Omanic is challenging, while step-wise analysis reveals a later-hop bottleneck, factual knowledge floor, and error propagation along reasoning chains. Fine-tuning on OmanicSynth transfers to six reasoning and mathematics benchmarks, yielding a 7.41-point average gain and validating its effectiveness as supervision for reasoning-capability transfer. We release the data at https://huggingface.co/datasets/li-lab/Omanic and the code at https://github.com/XiaojieGu/Omanic. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_16654 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models Gu, Xiaojie Tong, Sherry T. Feng, Aosong Han, Sophia Simeng Lu, Jinghui Chen, Yingjian Iwasawa, Yusuke Matsuo, Yutaka Park, Chanjun Ying, Rex Li, Irene Computation and Language Artificial Intelligence Machine Learning Evaluating the reasoning abilities of large language models (LLMs) solely from final answers can obscure failures in intermediate steps, especially in multi-hop QA benchmarks without step-level annotations. To address this gap, we introduce Omanic, an open-domain 4-hop QA benchmark designed not only to measure final-answer accuracy but also to diagnose where reasoning breaks down. Omanic contains 10,296 machine-generated training examples (OmanicSynth) and 967 expert-reviewed human-annotated evaluation examples (OmanicBench), with each evaluation question decomposed into single-hop sub-questions, intermediate answers, and structured graph topologies. Experiments with proprietary and open-source LLMs show that Omanic is challenging, while step-wise analysis reveals a later-hop bottleneck, factual knowledge floor, and error propagation along reasoning chains. Fine-tuning on OmanicSynth transfers to six reasoning and mathematics benchmarks, yielding a 7.41-point average gain and validating its effectiveness as supervision for reasoning-capability transfer. We release the data at https://huggingface.co/datasets/li-lab/Omanic and the code at https://github.com/XiaojieGu/Omanic. |
| title | Omanic: Towards Step-wise Evaluation of Multi-hop Reasoning in Large Language Models |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2603.16654 |