The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Su, Zeli, Xu, Zhankai, Chen, Tianlei, Zheng, Longfei, Zhang, Xiaolu, Zhou, Jun, Zhang, Wentao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911726781333504
author Su, Zeli
Xu, Zhankai
Chen, Tianlei
Zheng, Longfei
Zhang, Xiaolu
Zhou, Jun
Zhang, Wentao
author_facet Su, Zeli
Xu, Zhankai
Chen, Tianlei
Zheng, Longfei
Zhang, Xiaolu
Zhou, Jun
Zhang, Wentao
contents Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specified tasks over externally provided reference text. In practice, such context is often unstructured and contaminated with benign but instruction-like semantic noise, such as editorial comments and system traces, which should be treated strictly as data. We introduce DistractionIF, a benchmark designed to evaluate robustness against such distractor instructions in reference text. Across a broad range of models, we observe a consistent inverse scaling phenomenon: larger models are often less robust, with performance dropping by up to 30 points as scale increases. Mechanistically, our perplexity analysis reveals that scaling erodes the probabilistic boundary between robust and distracted behaviors, making models increasingly prone to over-interpreting noise as instructions. To address this, we demonstrate that reinforcement learning, specifically Group Relative Policy Optimization (GRPO), can restore this boundary, improving robustness by up to 15.5% without compromising general instruction-following capability. Our findings highlight a critical instruction-following robustness gap in reference-grounded tasks and establish reinforcement learning as a promising path for enforcing strict data-instruction separation at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2605_29491
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF
Su, Zeli
Xu, Zhankai
Chen, Tianlei
Zheng, Longfei
Zhang, Xiaolu
Zhou, Jun
Zhang, Wentao
Artificial Intelligence
Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specified tasks over externally provided reference text. In practice, such context is often unstructured and contaminated with benign but instruction-like semantic noise, such as editorial comments and system traces, which should be treated strictly as data. We introduce DistractionIF, a benchmark designed to evaluate robustness against such distractor instructions in reference text. Across a broad range of models, we observe a consistent inverse scaling phenomenon: larger models are often less robust, with performance dropping by up to 30 points as scale increases. Mechanistically, our perplexity analysis reveals that scaling erodes the probabilistic boundary between robust and distracted behaviors, making models increasingly prone to over-interpreting noise as instructions. To address this, we demonstrate that reinforcement learning, specifically Group Relative Policy Optimization (GRPO), can restore this boundary, improving robustness by up to 15.5% without compromising general instruction-following capability. Our findings highlight a critical instruction-following robustness gap in reference-grounded tasks and establish reinforcement learning as a promising path for enforcing strict data-instruction separation at scale.
title The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF
topic Artificial Intelligence
url https://arxiv.org/abs/2605.29491