Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Xingwu, Lu, Miao, Wu, Beining, Zou, Difan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025)
by: Li, Tianle, et al.
Published: (2025)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
by: Chen, Xingwu, et al.
Published: (2025)
by: Chen, Xingwu, et al.
Published: (2025)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
Scaling Laws for Precision in High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2026)
by: Zhang, Dechen, et al.
Published: (2026)
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
by: Chen, Xingwu, et al.
Published: (2024)
by: Chen, Xingwu, et al.
Published: (2024)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
by: Jiao, Difan, et al.
Published: (2026)
by: Jiao, Difan, et al.
Published: (2026)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
by: Jiang, Bingqing, et al.
Published: (2026)
by: Jiang, Bingqing, et al.
Published: (2026)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
Transformers Handle Endogeneity in In-Context Linear Regression
by: Liang, Haodong, et al.
Published: (2024)
by: Liang, Haodong, et al.
Published: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
by: Zuo, Yifei, et al.
Published: (2025)
by: Zuo, Yifei, et al.
Published: (2025)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
by: Wang, Xu, et al.
Published: (2025)
by: Wang, Xu, et al.
Published: (2025)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
by: Fu, Deqing, et al.
Published: (2023)
by: Fu, Deqing, et al.
Published: (2023)
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
by: Xie, Chengxing, et al.
Published: (2024)
by: Xie, Chengxing, et al.
Published: (2024)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
by: Tang, Xuan, et al.
Published: (2025)
by: Tang, Xuan, et al.
Published: (2025)
To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
by: Hao, Shugang, et al.
Published: (2025)
by: Hao, Shugang, et al.
Published: (2025)
Investigation into In-Context Learning Capabilities of Transformers
by: Chandrupatla, Rushil, et al.
Published: (2026)
by: Chandrupatla, Rushil, et al.
Published: (2026)
Transformers Don't In-Context Learn Least Squares Regression
by: Hill, Joshua, et al.
Published: (2025)
by: Hill, Joshua, et al.
Published: (2025)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
by: Li, Yanjie, et al.
Published: (2024)
by: Li, Yanjie, et al.
Published: (2024)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
by: Lin, Licong, et al.
Published: (2024)
by: Lin, Licong, et al.
Published: (2024)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
by: Cao, Hoang T. H., et al.
Published: (2026)
by: Cao, Hoang T. H., et al.
Published: (2026)
Towards Theoretical Understandings of Self-Consuming Generative Models
by: Fu, Shi, et al.
Published: (2024)
by: Fu, Shi, et al.
Published: (2024)
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
by: Hübotter, Jonas, et al.
Published: (2025)
by: Hübotter, Jonas, et al.
Published: (2025)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
by: Chen, Hao Mark, et al.
Published: (2025)
by: Chen, Hao Mark, et al.
Published: (2025)
Active Test-Time Adaptation: Theoretical Analyses and An Algorithm
by: Gui, Shurui, et al.
Published: (2024)
by: Gui, Shurui, et al.
Published: (2024)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
by: Goel, Ayush, et al.
Published: (2026)
by: Goel, Ayush, et al.
Published: (2026)
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
by: Yu, Chao, et al.
Published: (2025)
by: Yu, Chao, et al.
Published: (2025)
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
by: Wu, Jingfeng, et al.
Published: (2023)
by: Wu, Jingfeng, et al.
Published: (2023)
STGAN: Spatial-temporal Graph Autoregression Network for Pavement Distress Deterioration Prediction
by: Tong, Shilin, et al.
Published: (2025)
by: Tong, Shilin, et al.
Published: (2025)
Ister: Linear Transformer for Efficient Multivariate Time Series Forecasting
by: Cao, Fanpu, et al.
Published: (2024)
by: Cao, Fanpu, et al.
Published: (2024)
Explaining the Explainer: Understanding the Inner Workings of Transformer-based Symbolic Regression Models
by: van Breda, Arco, et al.
Published: (2026)
by: van Breda, Arco, et al.
Published: (2026)
Towards Understanding Layer Contributions in Tabular In-Context Learning Models
by: Balef, Amir Rezaei, et al.
Published: (2025)
by: Balef, Amir Rezaei, et al.
Published: (2025)
Investigating the Histogram Loss in Regression
by: Imani, Ehsan, et al.
Published: (2024)
by: Imani, Ehsan, et al.
Published: (2024)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
by: Zhang, Yi, et al.
Published: (2025)
by: Zhang, Yi, et al.
Published: (2025)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Polynormer: Polynomial-Expressive Graph Transformer in Linear Time
by: Deng, Chenhui, et al.
Published: (2024)
by: Deng, Chenhui, et al.
Published: (2024)
Similar Items
-
On the Robustness of Transformers against Context Hijacking for Linear Classification
by: Li, Tianle, et al.
Published: (2025) -
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
by: Chen, Xingwu, et al.
Published: (2025) -
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
by: Chen, Xingwu, et al.
Published: (2024) -
Learning under Quantization for High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2025) -
Scaling Laws for Precision in High-Dimensional Linear Regression
by: Zhang, Dechen, et al.
Published: (2026)