Saved in:
| Main Authors: | , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.14754 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909043533021184 |
|---|---|
| author | Zhiren, Gong Wu, Tiantong Zhang, Jiaming Zhang, Fuyao Wang, Che Hao, Yurong Hou, Yikun Ping, Foo Zhao, Yilei Huang, Fei Yuen, Chau Lim, Wei Yang Bryan |
| author_facet | Zhiren, Gong Wu, Tiantong Zhang, Jiaming Zhang, Fuyao Wang, Che Hao, Yurong Hou, Yikun Ping, Foo Zhao, Yilei Huang, Fei Yuen, Chau Lim, Wei Yang Bryan |
| contents | Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2605_14754 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition Zhiren, Gong Wu, Tiantong Zhang, Jiaming Zhang, Fuyao Wang, Che Hao, Yurong Hou, Yikun Ping, Foo Zhao, Yilei Huang, Fei Yuen, Chau Lim, Wei Yang Bryan Artificial Intelligence Large Language Models (LLMs) are increasingly deployed for knowledge synthesis, yet their capacity for compositional generalization in scientific knowledge remains under-characterized. Existing benchmarks primarily focus on single-turn restricted scenarios, failing to capture the capability boundaries exposed by real-world interactive scientific workflows. To address this, we introduce XDomainBench, a diagnostic benchmark for interactive interdisciplinary scientific reasoning. We formalize the composition order and mixture structure to enable systematic stress-testing from single-discipline to inter-disciplinary, comprising 8,598 interactive sessions across 20 domains and 4 task categories, with 8 realistic trajectory patterns covering difficulty and domain-mixture dynamics, simulating real AI4S scenarios. Large-scale evaluation of LLMs reveals a systematic reasoning collapse as composition order increases, stemming from two root causes: (i) direct difficulty increases induced by domain composition, and (ii) indirect interaction-amplified failures where trajectory patterns trigger error accumulation, reasoning breaks, and domain confusion, ultimately leading to session collapse. |
| title | XDomainBench: Diagnosing Reasoning Collapse in High-Dimensional Scientific Knowledge Composition |
| topic | Artificial Intelligence |
| url | https://arxiv.org/abs/2605.14754 |