In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Goel, Ayush, Kohli, Arjun, Somvanshi, Sarvagya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Predicting Post Virality with Temporal Cross-Attention over Trend Signals
by: Somvanshi, Sarvagya, et al.
Published: (2026)
by: Somvanshi, Sarvagya, et al.
Published: (2026)
Causal Pre-training Under the Fairness Lens: An Empirical Study of TabPFN
by: Liu, Qinyi, et al.
Published: (2026)
by: Liu, Qinyi, et al.
Published: (2026)
Noise Immunity in In-Context Tabular Learning: An Empirical Robustness Analysis of TabPFN's Attention Mechanisms
by: Hu, James, et al.
Published: (2026)
by: Hu, James, et al.
Published: (2026)
Tabular Data with Class Imbalance: Predicting Electric Vehicle Crash Severity with Pretrained Transformers (TabPFN) and Mamba-Based Models
by: Somvanshi, Shriyank, et al.
Published: (2025)
by: Somvanshi, Shriyank, et al.
Published: (2025)
Grokking in Linear Models for Logistic Regression
by: Das, Nataraj, et al.
Published: (2026)
by: Das, Nataraj, et al.
Published: (2026)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
by: Joshi, Sahil, et al.
Published: (2025)
by: Joshi, Sahil, et al.
Published: (2025)
A Survey on Deep Tabular Learning
by: Somvanshi, Shriyank, et al.
Published: (2024)
by: Somvanshi, Shriyank, et al.
Published: (2024)
Can Custom Models Learn In-Context? An Exploration of Hybrid Architecture Performance on In-Context Learning Tasks
by: Campbell, Ryan, et al.
Published: (2024)
by: Campbell, Ryan, et al.
Published: (2024)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
by: Chen, Lei, et al.
Published: (2024)
by: Chen, Lei, et al.
Published: (2024)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024)
by: Liang, Haodong, et al.
Published: (2024)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
by: Kim, Sunghwan, et al.
Published: (2026)
by: Kim, Sunghwan, et al.
Published: (2026)
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
Mimic In-Context Learning for Multimodal Tasks
by: Jiang, Yuchu, et al.
Published: (2025)
by: Jiang, Yuchu, et al.
Published: (2025)
DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks
by: Jiang, Qile, et al.
Published: (2025)
by: Jiang, Qile, et al.
Published: (2025)
Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
by: Wu, Gang, et al.
Published: (2025)
by: Wu, Gang, et al.
Published: (2025)
Evaluating Computational Accuracy of Large Language Models in Numerical Reasoning Tasks for Healthcare Applications
by: Malghan, Arjun R.
Published: (2025)
by: Malghan, Arjun R.
Published: (2025)
Automatic Piecewise Linear Regression for Predicting Student Learning Satisfaction
by: Choi, Haemin, et al.
Published: (2025)
by: Choi, Haemin, et al.
Published: (2025)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
by: Cao, Hoang T. H., et al.
Published: (2026)
by: Cao, Hoang T. H., et al.
Published: (2026)
Linear Attention for Efficient Bidirectional Sequence Modeling
by: Afzal, Arshia, et al.
Published: (2025)
by: Afzal, Arshia, et al.
Published: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Exact Linear Attention
by: Ou, Weinuo
Published: (2026)
by: Ou, Weinuo
Published: (2026)
Kaczmarz Linear Attention
by: Zou, Jiaxuan, et al.
Published: (2026)
by: Zou, Jiaxuan, et al.
Published: (2026)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
Does Faithfulness Conflict with Plausibility? An Empirical Study in Explainable AI across NLP Tasks
by: Lu, Xiaolei, et al.
Published: (2024)
by: Lu, Xiaolei, et al.
Published: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
by: Zhou, Yufa, et al.
Published: (2025)
by: Zhou, Yufa, et al.
Published: (2025)
A Multi-Task Learning Approach to Linear Multivariate Forecasting
by: Nochumsohn, Liran, et al.
Published: (2025)
by: Nochumsohn, Liran, et al.
Published: (2025)
Unsupervised Learning for Quadratic Assignment
by: Min, Yimeng, et al.
Published: (2025)
by: Min, Yimeng, et al.
Published: (2025)
Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
by: Danassis, Panayiotis, et al.
Published: (2025)
by: Danassis, Panayiotis, et al.
Published: (2025)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
by: Wi, Hyowon, et al.
Published: (2025)
by: Wi, Hyowon, et al.
Published: (2025)
LNUCB-TA: Linear-nonlinear Hybrid Bandit Learning with Temporal Attention
by: Khosravi, Hamed, et al.
Published: (2025)
by: Khosravi, Hamed, et al.
Published: (2025)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
by: Bouadi, Mohamed, et al.
Published: (2025)
by: Bouadi, Mohamed, et al.
Published: (2025)
An Empirical Study on Context Length for Open-Domain Dialog Generation
by: Shen, Xinyi, et al.
Published: (2024)
by: Shen, Xinyi, et al.
Published: (2024)
Similar Items
-
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024) -
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025) -
Predicting Post Virality with Temporal Cross-Attention over Trend Signals
by: Somvanshi, Sarvagya, et al.
Published: (2026) -
Causal Pre-training Under the Fairness Lens: An Empirical Study of TabPFN
by: Liu, Qinyi, et al.
Published: (2026) -
Noise Immunity in In-Context Tabular Learning: An Empirical Robustness Analysis of TabPFN's Attention Mechanisms
by: Hu, James, et al.
Published: (2026)