How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yang, Minglai, Huang, Ethan, Zhang, Liang, Surdeanu, Mihai, Wang, William, Pan, Liangming
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912598101852160
author Yang, Minglai
Huang, Ethan
Zhang, Liang
Surdeanu, Mihai
Wang, William
Pan, Liangming
author_facet Yang, Minglai
Huang, Ethan
Zhang, Liang
Surdeanu, Mihai
Wang, William
Pan, Liangming
contents We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC). GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation. Our experiments demonstrate that LLMs are significantly sensitive to IC, affecting both reasoning path selection and arithmetic accuracy. Additionally, training models with strong distractors improves performance in both in-distribution and out-of-distribution scenarios. We further propose a stepwise tree search guided by a process reward model, which notably enhances robustness in out-of-distribution conditions.
format Preprint
id arxiv_https___arxiv_org_abs_2505_18761
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
Yang, Minglai
Huang, Ethan
Zhang, Liang
Surdeanu, Mihai
Wang, William
Pan, Liangming
Computation and Language
Artificial Intelligence
Machine Learning
We introduce Grade School Math with Distracting Context (GSM-DC), a synthetic benchmark to evaluate Large Language Models' (LLMs) reasoning robustness against systematically controlled irrelevant context (IC). GSM-DC constructs symbolic reasoning graphs with precise distractor injections, enabling rigorous, reproducible evaluation. Our experiments demonstrate that LLMs are significantly sensitive to IC, affecting both reasoning path selection and arithmetic accuracy. Additionally, training models with strong distractors improves performance in both in-distribution and out-of-distribution scenarios. We further propose a stepwise tree search guided by a process reward model, which notably enhances robustness in out-of-distribution conditions.
title How Is LLM Reasoning Distracted by Irrelevant Context? An Analysis Using a Controlled Benchmark
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.18761