Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Xingxuan, Wang, Haoran, Li, Jiansheng, Xue, Yuan, Guan, Shikai, Xu, Renzhe, Zou, Hao, Yu, Han, Cui, Peng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Error Slice Discovery via Manifold Compactness
von: Yu, Han, et al.
Veröffentlicht: (2025)
von: Yu, Han, et al.
Veröffentlicht: (2025)
Rethinking the Evaluation Protocol of Domain Generalization
von: Yu, Han, et al.
Veröffentlicht: (2023)
von: Yu, Han, et al.
Veröffentlicht: (2023)
PPA-Game: Characterizing and Learning Competitive Dynamics Among Online Content Creators
von: Xu, Renzhe, et al.
Veröffentlicht: (2024)
von: Xu, Renzhe, et al.
Veröffentlicht: (2024)
Generating Risky Samples with Conformity Constraints via Diffusion Models
von: Yu, Han, et al.
Veröffentlicht: (2025)
von: Yu, Han, et al.
Veröffentlicht: (2025)
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)
A Survey on Evaluation of Out-of-Distribution Generalization
von: Yu, Han, et al.
Veröffentlicht: (2024)
von: Yu, Han, et al.
Veröffentlicht: (2024)
Breaking the Quality-Privacy Tradeoff in Tabular Data Generation via In-Context Learning
von: Han, Xinyan, et al.
Veröffentlicht: (2026)
von: Han, Xinyan, et al.
Veröffentlicht: (2026)
Sample Weight Averaging for Stable Prediction
von: Yu, Han, et al.
Veröffentlicht: (2025)
von: Yu, Han, et al.
Veröffentlicht: (2025)
TFMLinker: Universal Link Predictor by Graph In-Context Learning with Tabular Foundation Models
von: Liao, Tianyin, et al.
Veröffentlicht: (2026)
von: Liao, Tianyin, et al.
Veröffentlicht: (2026)
ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
von: Yu, Han, et al.
Veröffentlicht: (2025)
von: Yu, Han, et al.
Veröffentlicht: (2025)
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
von: Li, Hongkang, et al.
Veröffentlicht: (2024)
Understanding the Generalization of Stochastic Gradient Adam in Learning Neural Networks
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
von: Tang, Xuan, et al.
Veröffentlicht: (2025)
A Knowledge-Informed Pretrained Model for Causal Discovery
von: Xu, Wenbo, et al.
Veröffentlicht: (2026)
von: Xu, Wenbo, et al.
Veröffentlicht: (2026)
COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
von: Li, Jiansheng, et al.
Veröffentlicht: (2025)
von: Li, Jiansheng, et al.
Veröffentlicht: (2025)
Accelerated Distributional Temporal Difference Learning with Linear Function Approximation
von: Jin, Kaicheng, et al.
Veröffentlicht: (2025)
von: Jin, Kaicheng, et al.
Veröffentlicht: (2025)
Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
von: Xing, Yue, et al.
Veröffentlicht: (2024)
von: Xing, Yue, et al.
Veröffentlicht: (2024)
To Theoretically Understand Transformer-Based In-Context Learning for Optimizing CSMA
von: Hao, Shugang, et al.
Veröffentlicht: (2025)
von: Hao, Shugang, et al.
Veröffentlicht: (2025)
On the Generalization Properties of Learning the Random Feature Models with Learnable Activation Functions
von: Ma, Zailin, et al.
Veröffentlicht: (2025)
von: Ma, Zailin, et al.
Veröffentlicht: (2025)
Towards the Effect of Examples on In-Context Learning: A Theoretical Case Study
von: He, Pengfei, et al.
Veröffentlicht: (2024)
von: He, Pengfei, et al.
Veröffentlicht: (2024)
How Transformers Utilize Multi-Head Attention in In-Context Learning? A Case Study on Sparse Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
von: Chen, Xingwu, et al.
Veröffentlicht: (2024)
Understanding Generalization and Forgetting in In-Context Continual Learning
von: Li, Guangyu, et al.
Veröffentlicht: (2026)
von: Li, Guangyu, et al.
Veröffentlicht: (2026)
T-Graphormer: Using Transformers for Spatiotemporal Forecasting
von: Bai, Hao Yuan, et al.
Veröffentlicht: (2025)
von: Bai, Hao Yuan, et al.
Veröffentlicht: (2025)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
von: Huang, Yixiao, et al.
Veröffentlicht: (2025)
Towards Understanding Transformers in Learning Random Walks
von: Shi, Wei, et al.
Veröffentlicht: (2025)
von: Shi, Wei, et al.
Veröffentlicht: (2025)
Revisiting Zeroth-Order Hessian Approximation: A Single-Step Policy Optimization Lens
von: Qiu, Junbin, et al.
Veröffentlicht: (2026)
von: Qiu, Junbin, et al.
Veröffentlicht: (2026)
Towards Theoretical Understanding of Transformer Test-Time Computing: Investigation on In-Context Linear Regression
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
von: Chen, Xingwu, et al.
Veröffentlicht: (2025)
Learning from Scratch: Structurally-masked Transformer for Next Generation Lib-free Simulation
von: Huang, Junlang, et al.
Veröffentlicht: (2025)
von: Huang, Junlang, et al.
Veröffentlicht: (2025)
Learning Cross-Domain Representations for Transferable Drug Perturbations on Single-Cell Transcriptional Responses
von: Liu, Hui, et al.
Veröffentlicht: (2024)
von: Liu, Hui, et al.
Veröffentlicht: (2024)
Noise May Contain Transferable Knowledge: Understanding Semi-supervised Heterogeneous Domain Adaptation from an Empirical Perspective
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
Generative Pre-Trained Transformer for Symbolic Regression Base In-Context Reinforcement Learning
von: Li, Yanjie, et al.
Veröffentlicht: (2024)
von: Li, Yanjie, et al.
Veröffentlicht: (2024)
Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer
von: Hsu, Alexander, et al.
Veröffentlicht: (2026)
von: Hsu, Alexander, et al.
Veröffentlicht: (2026)
Heterogeneous Data Game: Characterizing the Model Competition Across Multiple Data Sources
von: Xu, Renzhe, et al.
Veröffentlicht: (2025)
von: Xu, Renzhe, et al.
Veröffentlicht: (2025)
LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2025)
On the Robustness of Transformers against Context Hijacking for Linear Classification
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
In-Context Deep Learning via Transformer Models
von: Wu, Weimin, et al.
Veröffentlicht: (2024)
von: Wu, Weimin, et al.
Veröffentlicht: (2024)
Improving GBDT Performance on Imbalanced Datasets: An Empirical Study of Class-Balanced Loss Functions
von: Luo, Jiaqi, et al.
Veröffentlicht: (2024)
von: Luo, Jiaqi, et al.
Veröffentlicht: (2024)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
von: Li, Hongkang, et al.
Veröffentlicht: (2025)
von: Li, Hongkang, et al.
Veröffentlicht: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
von: Gu, Zhengyao, et al.
Veröffentlicht: (2025)
An Empirical Study on Context Length for Open-Domain Dialog Generation
von: Shen, Xinyi, et al.
Veröffentlicht: (2024)
von: Shen, Xinyi, et al.
Veröffentlicht: (2024)
TabGen-ICL: Residual-Aware In-Context Example Selection for Tabular Data Generation
von: Fang, Liancheng, et al.
Veröffentlicht: (2025)
von: Fang, Liancheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Error Slice Discovery via Manifold Compactness
von: Yu, Han, et al.
Veröffentlicht: (2025) -
Rethinking the Evaluation Protocol of Domain Generalization
von: Yu, Han, et al.
Veröffentlicht: (2023) -
PPA-Game: Characterizing and Learning Competitive Dynamics Among Online Content Creators
von: Xu, Renzhe, et al.
Veröffentlicht: (2024) -
Generating Risky Samples with Conformity Constraints via Diffusion Models
von: Yu, Han, et al.
Veröffentlicht: (2025) -
On the Out-Of-Distribution Generalization of Multimodal Large Language Models
von: Zhang, Xingxuan, et al.
Veröffentlicht: (2024)