How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866912598101852160 |
|---|---|
| author | Yang, Minglai Huang, Ethan Zhang, Liang Surdeanu, Mihai Wang, William Pan, Liangming |
| author_facet | Yang, Minglai Huang, Ethan Zhang, Liang Surdeanu, Mihai Wang, William Pan, Liangming |
| contents | We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC). GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation. Our experiments demonstrate that LLMs are significantly sensitive to IC, affecting both reasoning path selection and arithmetic accuracy. Additionally, training models with strong distractors improves performance in both in-distribution and out-of-distribution scenarios. We further propose a stepwise tree search guided by a process reward model, which notably enhances robustness in out-of-distribution conditions. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_18761 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark Yang, Minglai Huang, Ethan Zhang, Liang Surdeanu, Mihai Wang, William Pan, Liangming Computation and Language Artificial Intelligence Machine Learning We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC). GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation. Our experiments demonstrate that LLMs are significantly sensitive to IC, affecting both reasoning path selection and arithmetic accuracy. Additionally, training models with strong distractors improves performance in both in-distribution and out-of-distribution scenarios. We further propose a stepwise tree search guided by a process reward model, which notably enhances robustness in out-of-distribution conditions. |
| title | How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark |
| topic | Computation and Language Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2505.18761 |