From independent patches to coordinated attention: Controlling information flow in vision transformers
Fuente:
arXiv
Saved in:
| Main Author: | Murphy, Kieran A. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Latent attention on masked patches for flow reconstruction
by: Eze, Ben, et al.
Published: (2026)
by: Eze, Ben, et al.
Published: (2026)
Comparing the information content of probabilistic representation spaces
by: Murphy, Kieran A., et al.
Published: (2024)
by: Murphy, Kieran A., et al.
Published: (2024)
Integrating attention into explanation frameworks for language and vision transformers
by: Eggen, Marte, et al.
Published: (2025)
by: Eggen, Marte, et al.
Published: (2025)
InfoChess: A Game of Adversarial Inference and a Laboratory for Quantifiable Information Control
by: Murphy, Kieran A.
Published: (2026)
by: Murphy, Kieran A.
Published: (2026)
Machine-learning optimized measurements of chaotic dynamical systems via the information bottleneck
by: Murphy, Kieran A., et al.
Published: (2023)
by: Murphy, Kieran A., et al.
Published: (2023)
Which bits went where? Past and future transfer entropy decomposition with the information bottleneck
by: Murphy, Kieran A., et al.
Published: (2024)
by: Murphy, Kieran A., et al.
Published: (2024)
Easy attention: A simple attention mechanism for temporal predictions with transformers
by: Sanchis-Agudo, Marcial, et al.
Published: (2023)
by: Sanchis-Agudo, Marcial, et al.
Published: (2023)
Pi-transformer: A prior-informed dual-attention model for multivariate time-series anomaly detection
by: Maleki, Sepehr, et al.
Published: (2025)
by: Maleki, Sepehr, et al.
Published: (2025)
Spatially-informed transformers: Injecting geostatistical covariance biases into self-attention for spatio-temporal forecasting
by: Calleo, Yuri
Published: (2025)
by: Calleo, Yuri
Published: (2025)
Integrated electro-optic attention nonlinearities for transformers
by: Mickeler, Luis, et al.
Published: (2026)
by: Mickeler, Luis, et al.
Published: (2026)
Sensorformer: Cross-patch attention with global-patch compression is effective for high-dimensional multivariate time series forecasting
by: Qin, Liyang, et al.
Published: (2025)
by: Qin, Liyang, et al.
Published: (2025)
Information decomposition in complex systems via machine learning
by: Murphy, Kieran A., et al.
Published: (2023)
by: Murphy, Kieran A., et al.
Published: (2023)
Controlling changes to attention logits
by: Anson, Ben, et al.
Published: (2025)
by: Anson, Ben, et al.
Published: (2025)
Enhancing compact convolutional transformers with super attention
by: Leandre, Simpenzwe Honore, et al.
Published: (2025)
by: Leandre, Simpenzwe Honore, et al.
Published: (2025)
Surveying the space of descriptions of a composite system with machine learning
by: Murphy, Kieran A., et al.
Published: (2024)
by: Murphy, Kieran A., et al.
Published: (2024)
Static and multivariate-temporal attentive fusion transformer for readmission risk prediction
by: Sun, Zhe, et al.
Published: (2024)
by: Sun, Zhe, et al.
Published: (2024)
Physics-informed GNN for medium-high voltage AC power flow with edge-aware attention and line search correction operator
by: Kim, Changhun, et al.
Published: (2025)
by: Kim, Changhun, et al.
Published: (2025)
Visualizing the loss landscapes of physics-informed neural networks
by: Rowan, Conor, et al.
Published: (2026)
by: Rowan, Conor, et al.
Published: (2026)
A standard transformer and attention with linear biases for molecular conformer generation
by: Gurev, Viatcheslav, et al.
Published: (2025)
by: Gurev, Viatcheslav, et al.
Published: (2025)
Sparse patches adversarial attacks via extrapolating point-wise information
by: Nemcovsky, Yaniv, et al.
Published: (2024)
by: Nemcovsky, Yaniv, et al.
Published: (2024)
Critical attention scaling in long-context transformers
by: Chen, Shi, et al.
Published: (2025)
by: Chen, Shi, et al.
Published: (2025)
Linear attention is (maybe) all you need (to understand transformer optimization)
by: Ahn, Kwangjun, et al.
Published: (2023)
by: Ahn, Kwangjun, et al.
Published: (2023)
Learning spatially structured open quantum dynamics with regional-attention transformers
by: Du, Dounan, et al.
Published: (2025)
by: Du, Dounan, et al.
Published: (2025)
How to use and interpret activation patching
by: Heimersheim, Stefan, et al.
Published: (2024)
by: Heimersheim, Stefan, et al.
Published: (2024)
Linear cost mutual information estimation and independence test of similar performance as HSIC
by: Duda, Jarek, et al.
Published: (2025)
by: Duda, Jarek, et al.
Published: (2025)
Physics-informed Discretization-independent Deep Compositional Operator Network
by: Zhong, Weiheng, et al.
Published: (2024)
by: Zhong, Weiheng, et al.
Published: (2024)
Scalable iterative pruning of large language and vision models using block coordinate descent
by: Rosenberg, Gili, et al.
Published: (2024)
by: Rosenberg, Gili, et al.
Published: (2024)
Self-attention-based non-linear basis transformations for compact latent space modelling of dynamic optical fibre transmission matrices
by: Zheng, Yijie, et al.
Published: (2024)
by: Zheng, Yijie, et al.
Published: (2024)
An Adaptive Volatility-based Learning Rate Scheduler
by: Ren, Kieran Chai Kai
Published: (2025)
by: Ren, Kieran Chai Kai
Published: (2025)
A hybrid transformer and attention based recurrent neural network for robust and interpretable sentiment analysis of tweets
by: Jahin, Md Abrar, et al.
Published: (2024)
by: Jahin, Md Abrar, et al.
Published: (2024)
A physics-informed and attention-based graph learning approach for regional electric vehicle charging demand prediction
by: Qu, Haohao, et al.
Published: (2023)
by: Qu, Haohao, et al.
Published: (2023)
Convergence theory for Hermite approximations under adaptive coordinate transformations
by: Saleh, Yahya
Published: (2026)
by: Saleh, Yahya
Published: (2026)
DEFT: Efficient Fine-Tuning of Diffusion Models by Learning the Generalised $h$-transform
by: Denker, Alexander, et al.
Published: (2024)
by: Denker, Alexander, et al.
Published: (2024)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
by: Smart, Matthew, et al.
Published: (2025)
by: Smart, Matthew, et al.
Published: (2025)
Reorganizing attention-space geometry with expressive attention
by: Gros, Claudius
Published: (2024)
by: Gros, Claudius
Published: (2024)
Physics-integrated generative modeling using attentive planar normalizing flow based variational autoencoder
by: Akhtar, Sheikh Waqas
Published: (2024)
by: Akhtar, Sheikh Waqas
Published: (2024)
Unified CNNs and transformers underlying learning mechanism reveals multi-head attention modus vivendi
by: Koresh, Ella, et al.
Published: (2025)
by: Koresh, Ella, et al.
Published: (2025)
Physics-informed neural particle flow for the Bayesian update step
by: Csuzdi, Domonkos, et al.
Published: (2026)
by: Csuzdi, Domonkos, et al.
Published: (2026)
Speech transformer models for extracting information from baby cries
by: Bonafos, Guillem, et al.
Published: (2025)
by: Bonafos, Guillem, et al.
Published: (2025)
Clustering in pure-attention hardmax transformers and its role in sentiment analysis
by: Alcalde, Albert, et al.
Published: (2024)
by: Alcalde, Albert, et al.
Published: (2024)
Similar Items
-
Latent attention on masked patches for flow reconstruction
by: Eze, Ben, et al.
Published: (2026) -
Comparing the information content of probabilistic representation spaces
by: Murphy, Kieran A., et al.
Published: (2024) -
Integrating attention into explanation frameworks for language and vision transformers
by: Eggen, Marte, et al.
Published: (2025) -
InfoChess: A Game of Adversarial Inference and a Laboratory for Quantifiable Information Control
by: Murphy, Kieran A.
Published: (2026) -
Machine-learning optimized measurements of chaotic dynamical systems via the information bottleneck
by: Murphy, Kieran A., et al.
Published: (2023)