An extension of linear self-attention for in-context learning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Hagiwara, Katsuyuki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A semi-supervised learning using over-parameterized regression
von: Hagiwara, Katsuyuki
Veröffentlicht: (2024)
von: Hagiwara, Katsuyuki
Veröffentlicht: (2024)
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
von: Lu, Yue M., et al.
Veröffentlicht: (2024)
Re-examining learning linear functions in context
von: Naim, Omar, et al.
Veröffentlicht: (2024)
von: Naim, Omar, et al.
Veröffentlicht: (2024)
Sparks of cognitive flexibility: self-guided context inference for flexible stimulus-response mapping by attentional routing
von: Sommers, Rowan P., et al.
Veröffentlicht: (2025)
von: Sommers, Rowan P., et al.
Veröffentlicht: (2025)
Poly-attention: a general scheme for higher-order self-attention
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)
The emergence of clusters in self-attention dynamics
von: Geshkovski, Borjan, et al.
Veröffentlicht: (2023)
von: Geshkovski, Borjan, et al.
Veröffentlicht: (2023)
Inpainting physics: self-supervised learning for context-driven fluid simulation
von: Weidner, Jonas, et al.
Veröffentlicht: (2026)
von: Weidner, Jonas, et al.
Veröffentlicht: (2026)
Supervised learning pays attention
von: Craig, Erin, et al.
Veröffentlicht: (2025)
von: Craig, Erin, et al.
Veröffentlicht: (2025)
Dynamic metastability in the self-attention model
von: Geshkovski, Borjan, et al.
Veröffentlicht: (2024)
von: Geshkovski, Borjan, et al.
Veröffentlicht: (2024)
Meta-reinforcement learning with minimum attention
von: Gupta, Shashank, et al.
Veröffentlicht: (2025)
von: Gupta, Shashank, et al.
Veröffentlicht: (2025)
The underlying structures of self-attention: symmetry, directionality, and emergent dynamics in Transformer training
von: Saponati, Matteo, et al.
Veröffentlicht: (2025)
von: Saponati, Matteo, et al.
Veröffentlicht: (2025)
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
Sliced ReLU attention: Quasi-linear contextual expressivity via sorting
von: Vialard, François-Xavier, et al.
Veröffentlicht: (2025)
von: Vialard, François-Xavier, et al.
Veröffentlicht: (2025)
Critical attention scaling in long-context transformers
von: Chen, Shi, et al.
Veröffentlicht: (2025)
von: Chen, Shi, et al.
Veröffentlicht: (2025)
Multi-layer Cross-attention is Provably Optimal for Multi-modal In-context Learning
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026)
von: Barnfield, Nicholas, et al.
Veröffentlicht: (2026)
Probing self-attention in self-supervised speech models for cross-linguistic differences
von: Gopinath, Sai, et al.
Veröffentlicht: (2024)
von: Gopinath, Sai, et al.
Veröffentlicht: (2024)
Online learning in bandits with predicted context
von: Guo, Yongyi, et al.
Veröffentlicht: (2023)
von: Guo, Yongyi, et al.
Veröffentlicht: (2023)
Transformer learns the cross-task prior and regularization for in-context learning
von: Lu, Fei, et al.
Veröffentlicht: (2025)
von: Lu, Fei, et al.
Veröffentlicht: (2025)
Simple linear attention language models balance the recall-throughput tradeoff
von: Arora, Simran, et al.
Veröffentlicht: (2024)
von: Arora, Simran, et al.
Veröffentlicht: (2024)
An end-to-end attention-based approach for learning on graphs
von: Buterez, David, et al.
Veröffentlicht: (2024)
von: Buterez, David, et al.
Veröffentlicht: (2024)
In-context learning agents are asymmetric belief updaters
von: Schubert, Johannes A., et al.
Veröffentlicht: (2024)
von: Schubert, Johannes A., et al.
Veröffentlicht: (2024)
Analyzing limits for in-context learning
von: Naim, Omar, et al.
Veröffentlicht: (2025)
von: Naim, Omar, et al.
Veröffentlicht: (2025)
The broader spectrum of in-context learning
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
In-context learning and Occam's razor
von: Elmoznino, Eric, et al.
Veröffentlicht: (2024)
von: Elmoznino, Eric, et al.
Veröffentlicht: (2024)
Safe reinforcement learning in uncertain contexts
von: Baumann, Dominik, et al.
Veröffentlicht: (2024)
von: Baumann, Dominik, et al.
Veröffentlicht: (2024)
Slim attention: cut your context memory in half without loss -- K-cache is all you need for MHA
von: Graef, Nils, et al.
Veröffentlicht: (2025)
von: Graef, Nils, et al.
Veröffentlicht: (2025)
A standard transformer and attention with linear biases for molecular conformer generation
von: Gurev, Viatcheslav, et al.
Veröffentlicht: (2025)
von: Gurev, Viatcheslav, et al.
Veröffentlicht: (2025)
Spatially-informed transformers: Injecting geostatistical covariance biases into self-attention for spatio-temporal forecasting
von: Calleo, Yuri
Veröffentlicht: (2025)
von: Calleo, Yuri
Veröffentlicht: (2025)
In-context learning to predict critical transitions in dynamical systems
von: Sevinchan, Yunus, et al.
Veröffentlicht: (2026)
von: Sevinchan, Yunus, et al.
Veröffentlicht: (2026)
Provably learning a multi-head attention layer
von: Chen, Sitan, et al.
Veröffentlicht: (2024)
von: Chen, Sitan, et al.
Veröffentlicht: (2024)
Efficient and accurate steering of Large Language Models through attention-guided feature learning
von: Davarmanesh, Parmida, et al.
Veröffentlicht: (2026)
von: Davarmanesh, Parmida, et al.
Veröffentlicht: (2026)
Universal and efficient graph neural networks with dynamic attention for machine learning interatomic potentials
von: Bi, Shuyu, et al.
Veröffentlicht: (2026)
von: Bi, Shuyu, et al.
Veröffentlicht: (2026)
Approximate learning of parsimonious Bayesian context trees
von: Ghani, Daniyar, et al.
Veröffentlicht: (2024)
von: Ghani, Daniyar, et al.
Veröffentlicht: (2024)
CausalLM is not optimal for in-context learning
von: Ding, Nan, et al.
Veröffentlicht: (2023)
von: Ding, Nan, et al.
Veröffentlicht: (2023)
In-context denoising with one-layer transformers: connections between attention and associative memory retrieval
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
von: Smart, Matthew, et al.
Veröffentlicht: (2025)
Efficient generative adversarial networks using linear additive-attention Transformers
von: Morales-Juarez, Emilio, et al.
Veröffentlicht: (2024)
von: Morales-Juarez, Emilio, et al.
Veröffentlicht: (2024)
Test time training enhances in-context learning of nonlinear functions
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
von: Kuwataka, Kento, et al.
Veröffentlicht: (2025)
Regime-aware financial volatility forecasting via in-context learning
von: Asaad, Saba, et al.
Veröffentlicht: (2026)
von: Asaad, Saba, et al.
Veröffentlicht: (2026)
Optimal cross-learning for contextual bandits with unknown context distributions
von: Schneider, Jon, et al.
Veröffentlicht: (2024)
von: Schneider, Jon, et al.
Veröffentlicht: (2024)
Easy attention: A simple attention mechanism for temporal predictions with transformers
von: Sanchis-Agudo, Marcial, et al.
Veröffentlicht: (2023)
von: Sanchis-Agudo, Marcial, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
A semi-supervised learning using over-parameterized regression
von: Hagiwara, Katsuyuki
Veröffentlicht: (2024) -
Asymptotic theory of in-context learning by linear attention
von: Lu, Yue M., et al.
Veröffentlicht: (2024) -
Re-examining learning linear functions in context
von: Naim, Omar, et al.
Veröffentlicht: (2024) -
Sparks of cognitive flexibility: self-guided context inference for flexible stimulus-response mapping by attentional routing
von: Sommers, Rowan P., et al.
Veröffentlicht: (2025) -
Poly-attention: a general scheme for higher-order self-attention
von: Chakrabarti, Sayak, et al.
Veröffentlicht: (2026)