Silent Tokens, Loud Effects: Padding in LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866916991244173312 |
|---|---|
| author | Himelstein, Rom LeVi, Amit Belinkov, Yonatan Mendelson, Avi |
| author_facet | Himelstein, Rom LeVi, Amit Belinkov, Yonatan Mendelson, Avi |
| contents | Padding tokens are widely used in large language models (LLMs) to equalize sequence lengths during batched inference. While they should be fully masked, implementation errors can cause them to influence computation, and the extent of this influence is not well understood. We systematically study this effect across three open-source model families (Llama, Gemma, Qwen), inserting controlled amounts of padding and evaluating outcomes along four axes: activations, generation quality, bias, and safety. Even small amounts of padding shift hidden representations, degrade quality in smaller models, alter bias in unpredictable ways, and weaken safety guardrails. These findings demonstrate that padding is not a harmless detail but a robustness risk that must be carefully handled in deployment. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_01238 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Silent Tokens, Loud Effects: Padding in LLMs Himelstein, Rom LeVi, Amit Belinkov, Yonatan Mendelson, Avi Computation and Language Machine Learning Padding tokens are widely used in large language models (LLMs) to equalize sequence lengths during batched inference. While they should be fully masked, implementation errors can cause them to influence computation, and the extent of this influence is not well understood. We systematically study this effect across three open-source model families (Llama, Gemma, Qwen), inserting controlled amounts of padding and evaluating outcomes along four axes: activations, generation quality, bias, and safety. Even small amounts of padding shift hidden representations, degrade quality in smaller models, alter bias in unpredictable ways, and weaken safety guardrails. These findings demonstrate that padding is not a harmless detail but a robustness risk that must be carefully handled in deployment. |
| title | Silent Tokens, Loud Effects: Padding in LLMs |
| topic | Computation and Language Machine Learning |
| url | https://arxiv.org/abs/2510.01238 |