When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Seleznyov, Mikhail, Chaichuk, Mikhail, Ershov, Gleb, Panchenko, Alexander, Tutubalina, Elena, Somov, Oleg
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915447569383424
author Seleznyov, Mikhail
Chaichuk, Mikhail
Ershov, Gleb
Panchenko, Alexander
Tutubalina, Elena
Somov, Oleg
author_facet Seleznyov, Mikhail
Chaichuk, Mikhail
Ershov, Gleb
Panchenko, Alexander
Tutubalina, Elena
Somov, Oleg
contents Large Language Models (LLMs) are highly sensitive to subtle, non-semantic variations in prompt phrasing and formatting. In this work, we present the first systematic evaluation of 5 methods for improving prompt robustness within a unified experimental framework. We benchmark these techniques on 8 models from Llama, Qwen and Gemma families across 52 tasks from Natural Instructions dataset. Our evaluation covers robustness methods from both fine-tuned and in-context learning paradigms, and tests their generalization against multiple types of distribution shifts. Finally, we extend our analysis to GPT-4.1 and DeepSeek V3 to assess frontier models' current robustness to format perturbations. Our findings offer actionable insights into the relative effectiveness of these robustness methods, enabling practitioners to make informed decisions when aiming for stable and reliable LLM performance in real-world applications. Code: https://github.com/AIRI-Institute/when-punctuation-matters.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11383
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
Seleznyov, Mikhail
Chaichuk, Mikhail
Ershov, Gleb
Panchenko, Alexander
Tutubalina, Elena
Somov, Oleg
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) are highly sensitive to subtle, non-semantic variations in prompt phrasing and formatting. In this work, we present the first systematic evaluation of 5 methods for improving prompt robustness within a unified experimental framework. We benchmark these techniques on 8 models from Llama, Qwen and Gemma families across 52 tasks from Natural Instructions dataset. Our evaluation covers robustness methods from both fine-tuned and in-context learning paradigms, and tests their generalization against multiple types of distribution shifts. Finally, we extend our analysis to GPT-4.1 and DeepSeek V3 to assess frontier models' current robustness to format perturbations. Our findings offer actionable insights into the relative effectiveness of these robustness methods, enabling practitioners to make informed decisions when aiming for stable and reliable LLM performance in real-world applications. Code: https://github.com/AIRI-Institute/when-punctuation-matters.
title When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2508.11383