Learning without training: The implicit dynamics of in-context learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dherin, Benoit, Munn, Michael, Mazzawi, Hanna, Wunder, Michael, Gonzalvo, Javier |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transmuting prompts into weights
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025)
Learning by solving differential equations
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
Deep Fusion: Efficient Network Training via Pre-trained Initializations
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2023)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2023)
How iteration order influences convergence and stability in deep learning
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
von: Munn, Michael, et al.
Veröffentlicht: (2024)
von: Munn, Michael, et al.
Veröffentlicht: (2024)
A Margin-based Multiclass Generalization Bound via Geometric Complexity
von: Munn, Michael, et al.
Veröffentlicht: (2024)
von: Munn, Michael, et al.
Veröffentlicht: (2024)
Grow, Don't Overwrite: Fine-tuning Without Forgetting
von: Adila, Dyah, et al.
Veröffentlicht: (2026)
von: Adila, Dyah, et al.
Veröffentlicht: (2026)
Equivalence of Context and Parameter Updates in Modern Transformer Blocks
von: Goldwaser, Adrian, et al.
Veröffentlicht: (2025)
von: Goldwaser, Adrian, et al.
Veröffentlicht: (2025)
On residual network depth
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
von: Dherin, Benoit, et al.
Veröffentlicht: (2025)
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2024)
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2024)
Language verY Rare for All
von: Merad, Ibrahim, et al.
Veröffentlicht: (2024)
von: Merad, Ibrahim, et al.
Veröffentlicht: (2024)
The broader spectrum of in-context learning
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
von: Lampinen, Andrew Kyle, et al.
Veröffentlicht: (2024)
CausalLM is not optimal for in-context learning
von: Ding, Nan, et al.
Veröffentlicht: (2023)
von: Ding, Nan, et al.
Veröffentlicht: (2023)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
von: Yano, Kazuki, et al.
Veröffentlicht: (2026)
von: Yano, Kazuki, et al.
Veröffentlicht: (2026)
Re-examining learning linear functions in context
von: Naim, Omar, et al.
Veröffentlicht: (2024)
von: Naim, Omar, et al.
Veröffentlicht: (2024)
Corridor Geometry in Gradient-Based Optimization
von: Dherin, Benoit, et al.
Veröffentlicht: (2024)
von: Dherin, Benoit, et al.
Veröffentlicht: (2024)
Understanding In-context Learning of Addition via Activation Subspaces
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
von: Hu, Xinyan, et al.
Veröffentlicht: (2025)
Density estimation with LLMs: a geometric investigation of in-context learning trajectories
von: Liu, Toni J. B., et al.
Veröffentlicht: (2024)
von: Liu, Toni J. B., et al.
Veröffentlicht: (2024)
Transparent Neighborhood Approximation for Text Classifier Explanation
von: Cai, Yi, et al.
Veröffentlicht: (2024)
von: Cai, Yi, et al.
Veröffentlicht: (2024)
In-context Learning in Presence of Spurious Correlations
von: Harutyunyan, Hrayr, et al.
Veröffentlicht: (2024)
von: Harutyunyan, Hrayr, et al.
Veröffentlicht: (2024)
Guideline Learning for In-context Information Extraction
von: Pang, Chaoxu, et al.
Veröffentlicht: (2023)
von: Pang, Chaoxu, et al.
Veröffentlicht: (2023)
In-context Learning and Gradient Descent Revisited
von: Deutch, Gilad, et al.
Veröffentlicht: (2023)
von: Deutch, Gilad, et al.
Veröffentlicht: (2023)
LLM Circuit Analyses Are Consistent Across Training and Scale
von: Tigges, Curt, et al.
Veröffentlicht: (2024)
von: Tigges, Curt, et al.
Veröffentlicht: (2024)
Breaking through the learning plateaus of in-context learning in Transformer
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
Adjoint sharding for very long context training of state space models
von: Xu, Xingzi, et al.
Veröffentlicht: (2025)
von: Xu, Xingzi, et al.
Veröffentlicht: (2025)
Linking In-context Learning in Transformers to Human Episodic Memory
von: Ji-An, Li, et al.
Veröffentlicht: (2024)
von: Ji-An, Li, et al.
Veröffentlicht: (2024)
Learning to Reason without External Rewards
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Using Pre-trained LLMs for Multivariate Time Series Forecasting
von: Wolff, Malcolm L., et al.
Veröffentlicht: (2025)
von: Wolff, Malcolm L., et al.
Veröffentlicht: (2025)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Implicit In-context Learning
von: Li, Zhuowei, et al.
Veröffentlicht: (2024)
von: Li, Zhuowei, et al.
Veröffentlicht: (2024)
Graph-based Molecular In-context Learning Grounded on Morgan Fingerprints
von: Al-Lawati, Ali, et al.
Veröffentlicht: (2025)
von: Al-Lawati, Ali, et al.
Veröffentlicht: (2025)
Reawakening knowledge: Anticipatory recovery from catastrophic interference via structured training
von: Yang, Yanlai, et al.
Veröffentlicht: (2024)
von: Yang, Yanlai, et al.
Veröffentlicht: (2024)
Towards Understanding the Relationship between In-context Learning and Compositional Generalization
von: Han, Sungjun, et al.
Veröffentlicht: (2024)
von: Han, Sungjun, et al.
Veröffentlicht: (2024)
Why does in-context learning fail sometimes? Evaluating in-context learning on open and closed questions
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
Learn or Recall? Revisiting Incremental Learning with Pre-trained Language Models
von: Zheng, Junhao, et al.
Veröffentlicht: (2023)
von: Zheng, Junhao, et al.
Veröffentlicht: (2023)
Polynomial Regression as a Task for Understanding In-context Learning Through Finetuning and Alignment
von: Wilcoxson, Max, et al.
Veröffentlicht: (2024)
von: Wilcoxson, Max, et al.
Veröffentlicht: (2024)
Large language models reorganize representational geometry during in-context learning
von: Xiong, Hua-Dong, et al.
Veröffentlicht: (2026)
von: Xiong, Hua-Dong, et al.
Veröffentlicht: (2026)
Transferable Post-training via Inverse Value Learning
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
von: Lu, Xinyu, et al.
Veröffentlicht: (2024)
Fine-tuning vs. In-context Learning in Large Language Models: A Formal Language Learning Perspective
von: Ghosh, Bishwamittra, et al.
Veröffentlicht: (2026)
von: Ghosh, Bishwamittra, et al.
Veröffentlicht: (2026)
Can LLMs Learn New Concepts Incrementally without Forgetting?
von: Zheng, Junhao, et al.
Veröffentlicht: (2024)
von: Zheng, Junhao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Transmuting prompts into weights
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2025) -
Learning by solving differential equations
von: Dherin, Benoit, et al.
Veröffentlicht: (2025) -
Deep Fusion: Efficient Network Training via Pre-trained Initializations
von: Mazzawi, Hanna, et al.
Veröffentlicht: (2023) -
How iteration order influences convergence and stability in deep learning
von: Dherin, Benoit, et al.
Veröffentlicht: (2025) -
The Impact of Geometric Complexity on Neural Collapse in Transfer Learning
von: Munn, Michael, et al.
Veröffentlicht: (2024)