In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goel, Ayush, Kohli, Arjun, Somvanshi, Sarvagya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
Predicting Post Virality with Temporal Cross-Attention over Trend Signals
von: Somvanshi, Sarvagya, et al.
Veröffentlicht: (2026)
von: Somvanshi, Sarvagya, et al.
Veröffentlicht: (2026)
Causal Pre-training Under the Fairness Lens: An Empirical Study of TabPFN
von: Liu, Qinyi, et al.
Veröffentlicht: (2026)
von: Liu, Qinyi, et al.
Veröffentlicht: (2026)
Noise Immunity in In-Context Tabular Learning: An Empirical Robustness Analysis of TabPFN's Attention Mechanisms
von: Hu, James, et al.
Veröffentlicht: (2026)
von: Hu, James, et al.
Veröffentlicht: (2026)
Tabular Data with Class Imbalance: Predicting Electric Vehicle Crash Severity with Pretrained Transformers (TabPFN) and Mamba-Based Models
von: Somvanshi, Shriyank, et al.
Veröffentlicht: (2025)
von: Somvanshi, Shriyank, et al.
Veröffentlicht: (2025)
Grokking in Linear Models for Logistic Regression
von: Das, Nataraj, et al.
Veröffentlicht: (2026)
von: Das, Nataraj, et al.
Veröffentlicht: (2026)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
RACE Attention: A Strictly Linear-Time Attention Layer for Training on Outrageously Large Contexts
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
von: Joshi, Sahil, et al.
Veröffentlicht: (2025)
A Survey on Deep Tabular Learning
von: Somvanshi, Shriyank, et al.
Veröffentlicht: (2024)
von: Somvanshi, Shriyank, et al.
Veröffentlicht: (2024)
Can Custom Models Learn In-Context? An Exploration of Hybrid Architecture Performance on In-Context Learning Tasks
von: Campbell, Ryan, et al.
Veröffentlicht: (2024)
von: Campbell, Ryan, et al.
Veröffentlicht: (2024)
Distributional Associations vs In-Context Reasoning: A Study of Feed-forward and Attention Layers
von: Chen, Lei, et al.
Veröffentlicht: (2024)
von: Chen, Lei, et al.
Veröffentlicht: (2024)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
Transformers Handle Endogeneity in In-Context Linear Regression
von: Liang, Haodong, et al.
Veröffentlicht: (2024)
von: Liang, Haodong, et al.
Veröffentlicht: (2024)
Enhancing Linear Attention with Residual Learning
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
von: Lai, Xunhao, et al.
Veröffentlicht: (2025)
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
von: Kim, Sunghwan, et al.
Veröffentlicht: (2026)
Learning under Quantization for High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
Transformers Don't In-Context Learn Least Squares Regression
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
Mimic In-Context Learning for Multimodal Tasks
von: Jiang, Yuchu, et al.
Veröffentlicht: (2025)
von: Jiang, Yuchu, et al.
Veröffentlicht: (2025)
DeepSeek vs. ChatGPT vs. Claude: A Comparative Study for Scientific Computing and Scientific Machine Learning Tasks
von: Jiang, Qile, et al.
Veröffentlicht: (2025)
von: Jiang, Qile, et al.
Veröffentlicht: (2025)
Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
von: Wu, Gang, et al.
Veröffentlicht: (2025)
von: Wu, Gang, et al.
Veröffentlicht: (2025)
Evaluating Computational Accuracy of Large Language Models in Numerical Reasoning Tasks for Healthcare Applications
von: Malghan, Arjun R.
Veröffentlicht: (2025)
von: Malghan, Arjun R.
Veröffentlicht: (2025)
Automatic Piecewise Linear Regression for Predicting Student Learning Satisfaction
von: Choi, Haemin, et al.
Veröffentlicht: (2025)
von: Choi, Haemin, et al.
Veröffentlicht: (2025)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
Linear Attention for Efficient Bidirectional Sequence Modeling
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
von: Afzal, Arshia, et al.
Veröffentlicht: (2025)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
von: MiniCPM Team, et al.
Veröffentlicht: (2026)
Exact Linear Attention
von: Ou, Weinuo
Veröffentlicht: (2026)
von: Ou, Weinuo
Veröffentlicht: (2026)
Kaczmarz Linear Attention
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Zou, Jiaxuan, et al.
Veröffentlicht: (2026)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
von: Li, Yanjie, et al.
Veröffentlicht: (2024)
von: Li, Yanjie, et al.
Veröffentlicht: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
Orion-Bix: Bi-Axial Attention for Tabular In-Context Learning
von: Bouadi, Mohamed, et al.
Veröffentlicht: (2025)
von: Bouadi, Mohamed, et al.
Veröffentlicht: (2025)
Does Faithfulness Conflict with Plausibility? An Empirical Study in Explainable AI across NLP Tasks
von: Lu, Xiaolei, et al.
Veröffentlicht: (2024)
von: Lu, Xiaolei, et al.
Veröffentlicht: (2024)
Why Do Transformers Fail to Forecast Time Series In-Context?
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
von: Zhou, Yufa, et al.
Veröffentlicht: (2025)
A Multi-Task Learning Approach to Linear Multivariate Forecasting
von: Nochumsohn, Liran, et al.
Veröffentlicht: (2025)
von: Nochumsohn, Liran, et al.
Veröffentlicht: (2025)
Unsupervised Learning for Quadratic Assignment
von: Min, Yimeng, et al.
Veröffentlicht: (2025)
von: Min, Yimeng, et al.
Veröffentlicht: (2025)
Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
von: Danassis, Panayiotis, et al.
Veröffentlicht: (2025)
von: Danassis, Panayiotis, et al.
Veröffentlicht: (2025)
Learning Advanced Self-Attention for Linear Transformers in the Singular Value Domain
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
von: Wi, Hyowon, et al.
Veröffentlicht: (2025)
LNUCB-TA: Linear-nonlinear Hybrid Bandit Learning with Temporal Attention
von: Khosravi, Hamed, et al.
Veröffentlicht: (2025)
von: Khosravi, Hamed, et al.
Veröffentlicht: (2025)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
von: Bouadi, Mohamed, et al.
Veröffentlicht: (2025)
von: Bouadi, Mohamed, et al.
Veröffentlicht: (2025)
An Empirical Study on Context Length for Open-Domain Dialog Generation
von: Shen, Xinyi, et al.
Veröffentlicht: (2024)
von: Shen, Xinyi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024) -
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
von: Zuo, Yifei, et al.
Veröffentlicht: (2025) -
Predicting Post Virality with Temporal Cross-Attention over Trend Signals
von: Somvanshi, Sarvagya, et al.
Veröffentlicht: (2026) -
Causal Pre-training Under the Fairness Lens: An Empirical Study of TabPFN
von: Liu, Qinyi, et al.
Veröffentlicht: (2026) -
Noise Immunity in In-Context Tabular Learning: An Empirical Robustness Analysis of TabPFN's Attention Mechanisms
von: Hu, James, et al.
Veröffentlicht: (2026)