Why Larger Language Models Do In-context Learning Differently?
Fuente:
arXiv
Saved in:
| Main Authors: | Shi, Zhenmei, Wei, Junyi, Xu, Zhuoyan, Liang, Yingyu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
by: Xu, Zhuoyan, et al.
Published: (2024)
by: Xu, Zhuoyan, et al.
Published: (2024)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
by: Xu, Zhuoyan, et al.
Published: (2024)
by: Xu, Zhuoyan, et al.
Published: (2024)
Can Language Models Compose Skills In-Context?
by: Liu, Zidong, et al.
Published: (2025)
by: Liu, Zidong, et al.
Published: (2025)
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Circuit Complexity Bounds for Visual Autoregressive Model
by: Ke, Yekun, et al.
Published: (2025)
by: Ke, Yekun, et al.
Published: (2025)
The Computational Limits of State-Space Models and Mamba via the Lens of Circuit Complexity
by: Chen, Yifang, et al.
Published: (2024)
by: Chen, Yifang, et al.
Published: (2024)
HSR-Enhanced Sparse Attention Acceleration
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
Beyond Linear Approximations: A Novel Pruning Approach for Attention Matrix
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Multi-Layer Transformers Gradient Can be Approximated in Almost Linear Time
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
by: Shi, Zhenmei, et al.
Published: (2024)
by: Shi, Zhenmei, et al.
Published: (2024)
A Tighter Complexity Analysis of SparseGPT
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
RoPE Attention Can Be Trained in Almost Linear Time
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Neural Algorithmic Reasoning for Hypergraphs with Looped Transformers
by: Huang, Zekai, et al.
Published: (2025)
by: Huang, Zekai, et al.
Published: (2025)
Conv-Basis: A New Paradigm for Efficient Attention Inference and Gradient Computation in Transformers
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Modern Hopfield Networks Require Chain-of-Thought to Solve $\mathsf{NC}^1$-Hard Problems
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Circuit Complexity Bounds for RoPE-based Transformer Architecture
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
On Fine-Grained I/O Complexity of Attention Backward Passes
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Time and Memory Trade-off of KV-Cache Compression in Tensor Transformer Decoding
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Bypassing the Exponential Dependency: Looped Transformers Efficiently Learn In-context by Multi-step Gradient Descent
by: Chen, Bo, et al.
Published: (2024)
by: Chen, Bo, et al.
Published: (2024)
When Can We Solve the Weighted Low Rank Approximation Problem in Truly Subquadratic Time?
by: Li, Chenyang, et al.
Published: (2025)
by: Li, Chenyang, et al.
Published: (2025)
Unraveling the Smoothness Properties of Diffusion Models: A Gaussian Mixture Perspective
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Curse of Attention: A Kernel-Based Perspective for Why Transformers Fail to Generalize on Time Series Forecasting and Beyond
by: Ke, Yekun, et al.
Published: (2024)
by: Ke, Yekun, et al.
Published: (2024)
RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
by: Cao, Yuefan, et al.
Published: (2025)
by: Cao, Yuefan, et al.
Published: (2025)
Exploring the Frontiers of Softmax: Provable Optimization, Applications in Diffusion Model, and Beyond
by: Cao, Yang, et al.
Published: (2024)
by: Cao, Yang, et al.
Published: (2024)
Out-of-distribution generalization via composition: a lens through induction heads in Transformers
by: Song, Jiajun, et al.
Published: (2024)
by: Song, Jiajun, et al.
Published: (2024)
Why Do Multilingual Reasoning Gaps Emerge in Reasoning Language Models?
by: Kang, Deokhyung, et al.
Published: (2025)
by: Kang, Deokhyung, et al.
Published: (2025)
Looped ReLU MLPs May Be All You Need as Practical Programmable Computers
by: Liang, Yingyu, et al.
Published: (2024)
by: Liang, Yingyu, et al.
Published: (2024)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Learning to Inference Adaptively for Multimodal Large Language Models
by: Xu, Zhuoyan, et al.
Published: (2025)
by: Xu, Zhuoyan, et al.
Published: (2025)
On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective
by: Li, Xiaoyu, et al.
Published: (2025)
by: Li, Xiaoyu, et al.
Published: (2025)
Kernel Regression in Structured Non-IID Settings: Theory and Implications for Denoising Score Learning
by: Zhang, Dechen, et al.
Published: (2025)
by: Zhang, Dechen, et al.
Published: (2025)
Probing the Decision Boundaries of In-context Learning in Large Language Models
by: Zhao, Siyan, et al.
Published: (2024)
by: Zhao, Siyan, et al.
Published: (2024)
Revisiting In-context Learning Inference Circuit in Large Language Models
by: Cho, Hakaze, et al.
Published: (2024)
by: Cho, Hakaze, et al.
Published: (2024)
In-context Pretraining: Language Modeling Beyond Document Boundaries
by: Shi, Weijia, et al.
Published: (2023)
by: Shi, Weijia, et al.
Published: (2023)
In-context Autoencoder for Context Compression in a Large Language Model
by: Ge, Tao, et al.
Published: (2023)
by: Ge, Tao, et al.
Published: (2023)
Learning to Interpret Weight Differences in Language Models
by: Goel, Avichal, et al.
Published: (2025)
by: Goel, Avichal, et al.
Published: (2025)
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
by: Liang, Yingyu, et al.
Published: (2025)
by: Liang, Yingyu, et al.
Published: (2025)
Similar Items
-
Do Large Language Models Have Compositional Ability? An Investigation into Limitations and Scalability
by: Xu, Zhuoyan, et al.
Published: (2024) -
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
by: Xu, Zhuoyan, et al.
Published: (2024) -
Can Language Models Compose Skills In-Context?
by: Liu, Zidong, et al.
Published: (2025) -
Training Tensor Attention Efficiently: From Cubic to Almost Linear Time
by: Cao, Yang, et al.
Published: (2024) -
Towards Infinite-Long Prefix in Transformer
by: Liang, Yingyu, et al.
Published: (2024)