The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guan, Bryan, Roosta, Tanya, Passban, Peyman, Rezagholizadeh, Mehdi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913827843473408
author Guan, Bryan
Roosta, Tanya
Passban, Peyman
Rezagholizadeh, Mehdi
author_facet Guan, Bryan
Roosta, Tanya
Passban, Peyman
Rezagholizadeh, Mehdi
contents As large language models (LLMs) become integral to diverse applications, ensuring their reliability under varying input conditions is crucial. One key issue affecting this reliability is order sensitivity, wherein slight variations in the input arrangement can lead to inconsistent or biased outputs. Although recent advances have reduced this sensitivity, the problem remains unresolved. This paper investigates the extent of order sensitivity in LLMs whose internal components are hidden from users (such as closed-source models or those accessed via API calls). We conduct experiments across multiple tasks, including paraphrasing, relevance judgment, and multiple-choice questions. Our results show that input order significantly affects performance across tasks, with shuffled inputs leading to measurable declines in output accuracy. Few-shot prompting demonstrates mixed effectiveness and offers partial mitigation; however, fails to fully resolve the problem. These findings highlight persistent risks, particularly in high-stakes applications, and point to the need for more robust LLMs or improved input-handling techniques in future development.
format Preprint
id arxiv_https___arxiv_org_abs_2502_04134
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
Guan, Bryan
Roosta, Tanya
Passban, Peyman
Rezagholizadeh, Mehdi
Computation and Language
As large language models (LLMs) become integral to diverse applications, ensuring their reliability under varying input conditions is crucial. One key issue affecting this reliability is order sensitivity, wherein slight variations in the input arrangement can lead to inconsistent or biased outputs. Although recent advances have reduced this sensitivity, the problem remains unresolved. This paper investigates the extent of order sensitivity in LLMs whose internal components are hidden from users (such as closed-source models or those accessed via API calls). We conduct experiments across multiple tasks, including paraphrasing, relevance judgment, and multiple-choice questions. Our results show that input order significantly affects performance across tasks, with shuffled inputs leading to measurable declines in output accuracy. Few-shot prompting demonstrates mixed effectiveness and offers partial mitigation; however, fails to fully resolve the problem. These findings highlight persistent risks, particularly in high-stakes applications, and point to the need for more robust LLMs or improved input-handling techniques in future development.
title The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
topic Computation and Language
url https://arxiv.org/abs/2502.04134