Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Deqing, Chen, Tian-Qi, Jia, Robin, Sharan, Vatsal |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026)
by: Fu, Deqing, et al.
Published: (2026)
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024)
by: Vasudeva, Bhavya, et al.
Published: (2024)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
by: Zhou, Tianyi, et al.
Published: (2024)
by: Zhou, Tianyi, et al.
Published: (2024)
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
by: Ye, Qilin, et al.
Published: (2025)
by: Ye, Qilin, et al.
Published: (2025)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
Latent Concept Disentanglement in Transformer-based Language Models
by: Hong, Guan Zhe, et al.
Published: (2025)
by: Hong, Guan Zhe, et al.
Published: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
Limitations on Accurate, Trusted, Human-level Reasoning
by: Panigrahy, Rina, et al.
Published: (2025)
by: Panigrahy, Rina, et al.
Published: (2025)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
by: Ramjee, Sharan
Published: (2026)
by: Ramjee, Sharan
Published: (2026)
Can GPT Redefine Medical Understanding? Evaluating GPT on Biomedical Machine Reading Comprehension
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Transformers Can Achieve Length Generalization But Not Robustly
by: Zhou, Yongchao, et al.
Published: (2024)
by: Zhou, Yongchao, et al.
Published: (2024)
DeLLMa: Decision Making Under Uncertainty with Large Language Models
by: Liu, Ollie, et al.
Published: (2024)
by: Liu, Ollie, et al.
Published: (2024)
Understanding Emergent In-Context Learning from a Kernel Regression Perspective
by: Han, Chi, et al.
Published: (2023)
by: Han, Chi, et al.
Published: (2023)
Supernova: Achieving More with Less in Transformer Architectures
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
by: Tanase, Andrei-Valentin, et al.
Published: (2025)
Can GPT Improve the State of Prior Authorization via Guideline Based Automated Question Answering?
by: Vatsal, Shubham, et al.
Published: (2024)
by: Vatsal, Shubham, et al.
Published: (2024)
Emotion Classification in Low and Moderate Resource Languages
by: Tafreshi, Shabnam, et al.
Published: (2024)
by: Tafreshi, Shabnam, et al.
Published: (2024)
Multilingual Prompt Engineering in Large Language Models: A Survey Across NLP Tasks
by: Vatsal, Shubham, et al.
Published: (2025)
by: Vatsal, Shubham, et al.
Published: (2025)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
Do pretrained Transformers Learn In-Context by Gradient Descent?
by: Shen, Lingfeng, et al.
Published: (2023)
by: Shen, Lingfeng, et al.
Published: (2023)
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
by: Zhang, Ziniu, et al.
Published: (2025)
by: Zhang, Ziniu, et al.
Published: (2025)
Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers
by: Kenneweg, Philip, et al.
Published: (2024)
by: Kenneweg, Philip, et al.
Published: (2024)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
by: Choi, Sehyun
Published: (2024)
by: Choi, Sehyun
Published: (2024)
Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning
by: Wang, Xubin, et al.
Published: (2026)
by: Wang, Xubin, et al.
Published: (2026)
In-Context Learning with Transformers: Softmax Attention Adapts to Function Lipschitzness
by: Collins, Liam, et al.
Published: (2024)
by: Collins, Liam, et al.
Published: (2024)
Second-Order Information Matters: Revisiting Machine Unlearning for Large Language Models
by: Gu, Kang, et al.
Published: (2024)
by: Gu, Kang, et al.
Published: (2024)
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024)
by: Liang, Haodong, et al.
Published: (2024)
Your Transformer is Secretly Linear
by: Razzhigaev, Anton, et al.
Published: (2024)
by: Razzhigaev, Anton, et al.
Published: (2024)
EPSVec: Efficient and Private Synthetic Data Generation via Dataset Vectors
by: Banayeeanzade, Amin, et al.
Published: (2026)
by: Banayeeanzade, Amin, et al.
Published: (2026)
On the Rate of Convergence of Kolmogorov-Arnold Network Regression Estimators
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Unlocking In-Context Learning for Natural Datasets Beyond Language Modelling
by: Bratulić, Jelena, et al.
Published: (2025)
by: Bratulić, Jelena, et al.
Published: (2025)
Boosting In-Context Learning in LLMs Through the Lens of Classical Supervised Learning
by: Gundem, Korel, et al.
Published: (2025)
by: Gundem, Korel, et al.
Published: (2025)
Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context Learning
by: Liu, Hui, et al.
Published: (2024)
by: Liu, Hui, et al.
Published: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
by: Gan, Woody Haosheng, et al.
Published: (2025)
by: Gan, Woody Haosheng, et al.
Published: (2025)
Second-Order Convergence in Private Stochastic Non-Convex Optimization
by: Tao, Youming, et al.
Published: (2025)
by: Tao, Youming, et al.
Published: (2025)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
by: Yan, Tianyi Lorena, et al.
Published: (2025)
by: Yan, Tianyi Lorena, et al.
Published: (2025)
CoCA: Fusing Position Embedding with Collinear Constrained Attention in Transformers for Long Context Window Extending
by: Zhu, Shiyi, et al.
Published: (2023)
by: Zhu, Shiyi, et al.
Published: (2023)
MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling
by: MiniCPM Team, et al.
Published: (2026)
by: MiniCPM Team, et al.
Published: (2026)
Is In-Context Learning Learning?
by: de Wynter, Adrian
Published: (2025)
by: de Wynter, Adrian
Published: (2025)
Similar Items
-
Convergent Evolution: How Different Language Models Learn Similar Number Representations
by: Fu, Deqing, et al.
Published: (2026) -
Transformers Learn Low Sensitivity Functions: Investigations and Implications
by: Vasudeva, Bhavya, et al.
Published: (2024) -
Pre-trained Large Language Models Use Fourier Features to Compute Addition
by: Zhou, Tianyi, et al.
Published: (2024) -
Transformers Provably Learn Algorithmic Solutions for Graph Connectivity, But Only with the Right Data
by: Ye, Qilin, et al.
Published: (2025) -
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)