From independent patches to coordinated attention: Controlling information flow in vision transformers

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Murphy, Kieran A.
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866912877650116608
author Murphy, Kieran A.
author_facet Murphy, Kieran A.
contents We make the information transmitted by attention an explicit, measurable quantity in vision transformers. By inserting variational information bottlenecks on all attention-mediated writes to the residual stream -- without other architectural changes -- we train models with an explicit information cost and obtain a controllable spectrum from independent patch processing to fully expressive global attention. On ImageNet-100, we characterize how classification behavior and information routing evolve across this spectrum, and provide initial insights into how global visual representations emerge from local patch processing by analyzing the first attention heads that transmit information. By biasing learning toward solutions with constrained internal communication, our approach yields models that are more tractable for mechanistic analysis and more amenable to control.
format Preprint
id arxiv_https___arxiv_org_abs_2602_04784
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle From independent patches to coordinated attention: Controlling information flow in vision transformers
Murphy, Kieran A.
Machine Learning
We make the information transmitted by attention an explicit, measurable quantity in vision transformers. By inserting variational information bottlenecks on all attention-mediated writes to the residual stream -- without other architectural changes -- we train models with an explicit information cost and obtain a controllable spectrum from independent patch processing to fully expressive global attention. On ImageNet-100, we characterize how classification behavior and information routing evolve across this spectrum, and provide initial insights into how global visual representations emerge from local patch processing by analyzing the first attention heads that transmit information. By biasing learning toward solutions with constrained internal communication, our approach yields models that are more tractable for mechanistic analysis and more amenable to control.
title From independent patches to coordinated attention: Controlling information flow in vision transformers
topic Machine Learning
url https://arxiv.org/abs/2602.04784