Silent Tokens, Loud Effects: Padding in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Himelstein, Rom, LeVi, Amit, Belinkov, Yonatan, Mendelson, Avi
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916991244173312
author Himelstein, Rom
LeVi, Amit
Belinkov, Yonatan
Mendelson, Avi
author_facet Himelstein, Rom
LeVi, Amit
Belinkov, Yonatan
Mendelson, Avi
contents Padding tokens are widely used in large language models (LLMs) to equalize sequence lengths during batched inference. While they should be fully masked, implementation errors can cause them to influence computation, and the extent of this influence is not well understood. We systematically study this effect across three open-source model families (Llama, Gemma, Qwen), inserting controlled amounts of padding and evaluating outcomes along four axes: activations, generation quality, bias, and safety. Even small amounts of padding shift hidden representations, degrade quality in smaller models, alter bias in unpredictable ways, and weaken safety guardrails. These findings demonstrate that padding is not a harmless detail but a robustness risk that must be carefully handled in deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2510_01238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Silent Tokens, Loud Effects: Padding in LLMs
Himelstein, Rom
LeVi, Amit
Belinkov, Yonatan
Mendelson, Avi
Computation and Language
Machine Learning
Padding tokens are widely used in large language models (LLMs) to equalize sequence lengths during batched inference. While they should be fully masked, implementation errors can cause them to influence computation, and the extent of this influence is not well understood. We systematically study this effect across three open-source model families (Llama, Gemma, Qwen), inserting controlled amounts of padding and evaluating outcomes along four axes: activations, generation quality, bias, and safety. Even small amounts of padding shift hidden representations, degrade quality in smaller models, alter bias in unpredictable ways, and weaken safety guardrails. These findings demonstrate that padding is not a harmless detail but a robustness risk that must be carefully handled in deployment.
title Silent Tokens, Loud Effects: Padding in LLMs
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2510.01238