An extension of linear self-attention for in-context learning
Fuente:
arXiv
Guardado en:
| Autor principal: | Hagiwara, Katsuyuki |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A semi-supervised learning using over-parameterized regression
por: Hagiwara, Katsuyuki
Publicado: (2024)
por: Hagiwara, Katsuyuki
Publicado: (2024)
Asymptotic theory of in-context learning by linear attention
por: Lu, Yue M., et al.
Publicado: (2024)
por: Lu, Yue M., et al.
Publicado: (2024)
Re-examining learning linear functions in context
por: Naim, Omar, et al.
Publicado: (2024)
por: Naim, Omar, et al.
Publicado: (2024)
Sparks of cognitive flexibility: self-guided context inference for flexible stimulus-response mapping by attentional routing
por: Sommers, Rowan P., et al.
Publicado: (2025)
por: Sommers, Rowan P., et al.
Publicado: (2025)
Poly-attention: a general scheme for higher-order self-attention
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
por: Chakrabarti, Sayak, et al.
Publicado: (2026)
The emergence of clusters in self-attention dynamics
por: Geshkovski, Borjan, et al.
Publicado: (2023)
por: Geshkovski, Borjan, et al.
Publicado: (2023)
Inpainting physics: self-supervised learning for context-driven fluid simulation
por: Weidner, Jonas, et al.
Publicado: (2026)
por: Weidner, Jonas, et al.
Publicado: (2026)
Supervised learning pays attention
por: Craig, Erin, et al.
Publicado: (2025)
por: Craig, Erin, et al.
Publicado: (2025)
Dynamic metastability in the self-attention model
por: Geshkovski, Borjan, et al.
Publicado: (2024)
por: Geshkovski, Borjan, et al.
Publicado: (2024)
Meta-reinforcement learning with minimum attention
por: Gupta, Shashank, et al.
Publicado: (2025)
por: Gupta, Shashank, et al.
Publicado: (2025)
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training
por: Saponati, Matteo, et al.
Publicado: (2025)
por: Saponati, Matteo, et al.
Publicado: (2025)
A low latency attention module for streaming self-supervised speech representation learning
por: Ma, Jianbo, et al.
Publicado: (2023)
por: Ma, Jianbo, et al.
Publicado: (2023)
Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
por: Vialard, François-Xavier, et al.
Publicado: (2025)
por: Vialard, François-Xavier, et al.
Publicado: (2025)
Critical attention scaling in long-context transformers
por: Chen, Shi, et al.
Publicado: (2025)
por: Chen, Shi, et al.
Publicado: (2025)
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
por: Barnfield, Nicholas, et al.
Publicado: (2026)
por: Barnfield, Nicholas, et al.
Publicado: (2026)
Probing self-attention in self-supervised speech models for cross-linguistic differences
por: Gopinath, Sai, et al.
Publicado: (2024)
por: Gopinath, Sai, et al.
Publicado: (2024)
Online learning in bandits with predicted context
por: Guo, Yongyi, et al.
Publicado: (2023)
por: Guo, Yongyi, et al.
Publicado: (2023)
Transformer learns the cross-task prior and regularization for in-context learning
por: Lu, Fei, et al.
Publicado: (2025)
por: Lu, Fei, et al.
Publicado: (2025)
Simple linear attention language models balance the recall-throughput tradeoff
por: Arora, Simran, et al.
Publicado: (2024)
por: Arora, Simran, et al.
Publicado: (2024)
An end-to-end attention-based approach for learning on graphs
por: Buterez, David, et al.
Publicado: (2024)
por: Buterez, David, et al.
Publicado: (2024)
In-context learning agents are asymmetric belief updaters
por: Schubert, Johannes A., et al.
Publicado: (2024)
por: Schubert, Johannes A., et al.
Publicado: (2024)
Analyzing limits for in-context learning
por: Naim, Omar, et al.
Publicado: (2025)
por: Naim, Omar, et al.
Publicado: (2025)
The broader spectrum of in-context learning
por: Lampinen, Andrew Kyle, et al.
Publicado: (2024)
por: Lampinen, Andrew Kyle, et al.
Publicado: (2024)
In-context learning and Occam's razor
por: Elmoznino, Eric, et al.
Publicado: (2024)
por: Elmoznino, Eric, et al.
Publicado: (2024)
Safe reinforcement learning in uncertain contexts
por: Baumann, Dominik, et al.
Publicado: (2024)
por: Baumann, Dominik, et al.
Publicado: (2024)
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
por: Graef, Nils, et al.
Publicado: (2025)
por: Graef, Nils, et al.
Publicado: (2025)
A standard transformer and attention with linear biases for molecular conformer generation
por: Gurev, Viatcheslav, et al.
Publicado: (2025)
por: Gurev, Viatcheslav, et al.
Publicado: (2025)
Spatially-informed transformers: Injecting geostatistical covariance biases into self-attention for spatio-temporal forecasting
por: Calleo, Yuri
Publicado: (2025)
por: Calleo, Yuri
Publicado: (2025)
In-context learning to predict critical transitions in dynamical systems
por: Sevinchan, Yunus, et al.
Publicado: (2026)
por: Sevinchan, Yunus, et al.
Publicado: (2026)
Provably learning a multi-head attention layer
por: Chen, Sitan, et al.
Publicado: (2024)
por: Chen, Sitan, et al.
Publicado: (2024)
Efficient and accurate steering of Large Language Models through attention-guided feature learning
por: Davarmanesh, Parmida, et al.
Publicado: (2026)
por: Davarmanesh, Parmida, et al.
Publicado: (2026)
Universal and efficient graph neural networks with dynamic attention for machine learning interatomic potentials
por: Bi, Shuyu, et al.
Publicado: (2026)
por: Bi, Shuyu, et al.
Publicado: (2026)
Approximate learning of parsimonious Bayesian context trees
por: Ghani, Daniyar, et al.
Publicado: (2024)
por: Ghani, Daniyar, et al.
Publicado: (2024)
CausalLM is not optimal for in-context learning
por: Ding, Nan, et al.
Publicado: (2023)
por: Ding, Nan, et al.
Publicado: (2023)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
por: Smart, Matthew, et al.
Publicado: (2025)
por: Smart, Matthew, et al.
Publicado: (2025)
Efficient generative adversarial networks using linear additive-attention Transformers
por: Morales-Juarez, Emilio, et al.
Publicado: (2024)
por: Morales-Juarez, Emilio, et al.
Publicado: (2024)
Test time training enhances in-context learning of nonlinear functions
por: Kuwataka, Kento, et al.
Publicado: (2025)
por: Kuwataka, Kento, et al.
Publicado: (2025)
Regime-aware financial volatility forecasting via in-context learning
por: Asaad, Saba, et al.
Publicado: (2026)
por: Asaad, Saba, et al.
Publicado: (2026)
Optimal cross-learning for contextual bandits with unknown context distributions
por: Schneider, Jon, et al.
Publicado: (2024)
por: Schneider, Jon, et al.
Publicado: (2024)
Easy attention: A simple attention mechanism for temporal predictions with transformers
por: Sanchis-Agudo, Marcial, et al.
Publicado: (2023)
por: Sanchis-Agudo, Marcial, et al.
Publicado: (2023)
Ejemplares similares
-
A semi-supervised learning using over-parameterized regression
por: Hagiwara, Katsuyuki
Publicado: (2024) -
Asymptotic theory of in-context learning by linear attention
por: Lu, Yue M., et al.
Publicado: (2024) -
Re-examining learning linear functions in context
por: Naim, Omar, et al.
Publicado: (2024) -
Sparks of cognitive flexibility: self-guided context inference for flexible stimulus-response mapping by attentional routing
por: Sommers, Rowan P., et al.
Publicado: (2025) -
Poly-attention: a general scheme for higher-order self-attention
por: Chakrabarti, Sayak, et al.
Publicado: (2026)