Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xingwu, Lu, Miao, Wu, Beining, Zou, Difan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
Learning under Quantization for High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
Scaling Laws for Precision in High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)
What Can Transformer Learn with Varying Depth? Case Studies on Sequence Learning Tasks
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
Understanding the Dynamics of Demonstration Conflict in In-Context Learning
von: Jiao, Difan, et al.
Veröffentlicht: (2026)
von: Jiao, Difan, et al.
Veröffentlicht: (2026)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
Structured Role-Aware Policy Optimization for Multimodal Reasoning
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
On the Memorization of Consistency Distillation for Diffusion Models
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
von: Jiang, Bingqing, et al.
Veröffentlicht: (2026)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
von: Zhang, Dechen, et al.
Veröffentlicht: (2025)
Transformers Handle Endogeneity in In-Context Linear Regression
von: Liang, Haodong, et al.
Veröffentlicht: (2024)
von: Liang, Haodong, et al.
Veröffentlicht: (2024)
Local Linear Attention: An Optimal Interpolation of Linear and Softmax Attention For Test-Time Regression
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
von: Zuo, Yifei, et al.
Veröffentlicht: (2025)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
On the Collapse Errors Induced by the Deterministic Sampler for Diffusion Models
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Superiority of Multi-Head Attention in In-Context Linear Regression
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
von: Cui, Yingqian, et al.
Veröffentlicht: (2024)
A Human-Like Reasoning Framework for Multi-Phases Planning Task with Large Language Models
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
von: Xie, Chengxing, et al.
Veröffentlicht: (2024)
A Convergence Analysis of Adaptive Optimizers under Floating-point Quantization
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
von: Hao, Shugang, et al.
Veröffentlicht: (2025)
von: Hao, Shugang, et al.
Veröffentlicht: (2025)
Investigation into In-Context Learning Capabilities of Transformers
von: Chandrupatla, Rushil, et al.
Veröffentlicht: (2026)
von: Chandrupatla, Rushil, et al.
Veröffentlicht: (2026)
Transformers Don't In-Context Learn Least Squares Regression
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
von: Hill, Joshua, et al.
Veröffentlicht: (2025)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
von: Li, Yanjie, et al.
Veröffentlicht: (2024)
von: Li, Yanjie, et al.
Veröffentlicht: (2024)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
von: Lin, Licong, et al.
Veröffentlicht: (2024)
von: Lin, Licong, et al.
Veröffentlicht: (2024)
Transformers Learn Robust In-Context Regression under Distributional Uncertainty
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
von: Cao, Hoang T. H., et al.
Veröffentlicht: (2026)
Towards Theoretical Understandings of Self-Consuming Generative Models
von: Fu, Shi, et al.
Veröffentlicht: (2024)
von: Fu, Shi, et al.
Veröffentlicht: (2024)
Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models
von: Hübotter, Jonas, et al.
Veröffentlicht: (2025)
von: Hübotter, Jonas, et al.
Veröffentlicht: (2025)
Rethinking Optimal Verification Granularity for Compute-Efficient Test-Time Scaling
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
von: Chen, Hao Mark, et al.
Veröffentlicht: (2025)
Active Test-Time Adaptation: Theoretical Analyses and An Algorithm
von: Gui, Shurui, et al.
Veröffentlicht: (2024)
von: Gui, Shurui, et al.
Veröffentlicht: (2024)
In-Context Learning in Linear vs. Quadratic Attention Models: An Empirical Study on Regression Tasks
von: Goel, Ayush, et al.
Veröffentlicht: (2026)
von: Goel, Ayush, et al.
Veröffentlicht: (2026)
Extending Test-Time Scaling: A 3D Perspective with Context, Batch, and Turn
von: Yu, Chao, et al.
Veröffentlicht: (2025)
von: Yu, Chao, et al.
Veröffentlicht: (2025)
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
STGAN: Spatial-temporal Graph Autoregression Network for Pavement Distress Deterioration Prediction
von: Tong, Shilin, et al.
Veröffentlicht: (2025)
von: Tong, Shilin, et al.
Veröffentlicht: (2025)
Ister: Linear Transformer for Efficient Multivariate Time Series Forecasting
von: Cao, Fanpu, et al.
Veröffentlicht: (2024)
von: Cao, Fanpu, et al.
Veröffentlicht: (2024)
Explaining the Explainer: Understanding the Inner Workings of Transformer-based Symbolic Regression Models
von: van Breda, Arco, et al.
Veröffentlicht: (2026)
von: van Breda, Arco, et al.
Veröffentlicht: (2026)
Towards Understanding Layer Contributions in Tabular In-Context Learning Models
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
von: Balef, Amir Rezaei, et al.
Veröffentlicht: (2025)
Investigating the Histogram Loss in Regression
von: Imani, Ehsan, et al.
Veröffentlicht: (2024)
von: Imani, Ehsan, et al.
Veröffentlicht: (2024)
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
von: Zhang, Yi, et al.
Veröffentlicht: (2025)
Understanding the Role of Training Data in Test-Time Scaling
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
von: Javanmard, Adel, et al.
Veröffentlicht: (2025)
Polynormer: Polynomial-Expressive Graph Transformer in Linear Time
von: Deng, Chenhui, et al.
Veröffentlicht: (2024)
von: Deng, Chenhui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025) -
Reshaping Reasoning in LLMs: A Theoretical Analysis of RL Training Dynamics through Pattern Selection
von: Chen, Xingwu, et al.
Veröffentlicht: (2025) -
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2024) -
Learning under Quantization for High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2025) -
Scaling Laws for Precision in High-Dimensional Linear Regression
von: Zhang, Dechen, et al.
Veröffentlicht: (2026)