Saved in:
Bibliographic Details
Main Authors: Zhiren, Gong, Wu, Tiantong, Zhang, Jiaming, Zhang, Fuyao, Wang, Che, Hao, Yurong, Hou, Yikun, Ping, Foo, Zhao, Yilei, Huang, Fei, Yuen, Chau, Lim, Wei Yang Bryan
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2605.14754
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909043533021184
author Zhiren, Gong
Wu, Tiantong
Zhang, Jiaming
Zhang, Fuyao
Wang, Che
Hao, Yurong
Hou, Yikun
Ping, Foo
Zhao, Yilei
Huang, Fei
Yuen, Chau
Lim, Wei Yang Bryan
author_facet Zhiren, Gong
Wu, Tiantong
Zhang, Jiaming
Zhang, Fuyao
Wang, Che
Hao, Yurong
Hou, Yikun
Ping, Foo
Zhao, Yilei
Huang, Fei
Yuen, Chau
Lim, Wei Yang Bryan
contents Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14754
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
Zhiren, Gong
Wu, Tiantong
Zhang, Jiaming
Zhang, Fuyao
Wang, Che
Hao, Yurong
Hou, Yikun
Ping, Foo
Zhao, Yilei
Huang, Fei
Yuen, Chau
Lim, Wei Yang Bryan
Artificial Intelligence
Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse.
title XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition
topic Artificial Intelligence
url https://arxiv.org/abs/2605.14754