Theoretical Understanding of In-Context Learning in Shallow Transformers with Unstructured Data
Fuente:
arXiv
Saved in:
| Main Authors: | Xing, Yue, Lin, Xiaofeng, Xu, Chenheng, Suh, Namjoon, Song, Qifan, Cheng, Guang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Survey on Statistical Theory of Deep Learning: Approximation, Training Dynamics, and Generative Models
by: Suh, Namjoon, et al.
Published: (2024)
by: Suh, Namjoon, et al.
Published: (2024)
CTSyn: A Foundation Model for Cross Tabular Data Generation
by: Lin, Xiaofeng, et al.
Published: (2024)
by: Lin, Xiaofeng, et al.
Published: (2024)
Better Representations via Adversarial Training in Pre-Training: A Theoretical Perspective
by: Xing, Yue, et al.
Published: (2024)
by: Xing, Yue, et al.
Published: (2024)
From Unstructured Data to In-Context Learning: Exploring What Tasks Can Be Learned and When
by: Wibisono, Kevin Christian, et al.
Published: (2024)
by: Wibisono, Kevin Christian, et al.
Published: (2024)
Discriminative Estimation of Total Variation Distance: A Fidelity Auditor for Generative Data
by: Tao, Lan, et al.
Published: (2024)
by: Tao, Lan, et al.
Published: (2024)
Generalization or Hallucination? Understanding Out-of-Context Reasoning in Transformers
by: Huang, Yixiao, et al.
Published: (2025)
by: Huang, Yixiao, et al.
Published: (2025)
Feature Resemblance: Towards a Theoretical Understanding of Analogical Reasoning in Transformers
by: Xu, Ruichen, et al.
Published: (2026)
by: Xu, Ruichen, et al.
Published: (2026)
Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges
by: Wu, Xiaofeng, et al.
Published: (2025)
by: Wu, Xiaofeng, et al.
Published: (2025)
Approximation of RKHS Functionals by Neural Networks
by: Zhou, Tian-Yi, et al.
Published: (2024)
by: Zhou, Tian-Yi, et al.
Published: (2024)
Advancing LLM Safe Alignment with Safety Representation Ranking
by: Du, Tianqi, et al.
Published: (2025)
by: Du, Tianqi, et al.
Published: (2025)
LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data
by: Zhang, Hanyu, et al.
Published: (2024)
by: Zhang, Hanyu, et al.
Published: (2024)
Mining Unstructured Medical Texts With Conformal Active Learning
by: Genari, Juliano, et al.
Published: (2025)
by: Genari, Juliano, et al.
Published: (2025)
A Theoretical Understanding of Chain-of-Thought: Coherent Reasoning and Error-Aware Demonstration
by: Cui, Yingqian, et al.
Published: (2024)
by: Cui, Yingqian, et al.
Published: (2024)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
by: Shao, Jintian, et al.
Published: (2025)
by: Shao, Jintian, et al.
Published: (2025)
Towards Better Understanding of In-Context Learning Ability from In-Context Uncertainty Quantification
by: Liu, Shang, et al.
Published: (2024)
by: Liu, Shang, et al.
Published: (2024)
Key Information Retrieval to Classify the Unstructured Data Content of Preferential Trade Agreements
by: Zhao, Jiahui, et al.
Published: (2024)
by: Zhao, Jiahui, et al.
Published: (2024)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Sequence-Level Leakage Risk of Training Data in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2024)
by: Tiwari, Trishita, et al.
Published: (2024)
Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers
by: Huang, Yiran, et al.
Published: (2026)
by: Huang, Yiran, et al.
Published: (2026)
Value Alignment from Unstructured Text
by: Padhi, Inkit, et al.
Published: (2024)
by: Padhi, Inkit, et al.
Published: (2024)
Transformers as Decision Makers: Provable In-Context Reinforcement Learning via Supervised Pretraining
by: Lin, Licong, et al.
Published: (2023)
by: Lin, Licong, et al.
Published: (2023)
SELF: Self-Extend the Context Length With Logistic Growth Function
by: Dang, Phat Thanh, et al.
Published: (2025)
by: Dang, Phat Thanh, et al.
Published: (2025)
LLM Safety Alignment is Divergence Estimation in Disguise
by: Haldar, Rajdeep, et al.
Published: (2025)
by: Haldar, Rajdeep, et al.
Published: (2025)
Short Data, Long Context: Distilling Positional Knowledge in Transformers
by: Huber, Patrick, et al.
Published: (2026)
by: Huber, Patrick, et al.
Published: (2026)
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Understanding Contextual Recall in Transformers: How Finetuning Enables In-Context Reasoning over Pretraining Knowledge
by: Vasudeva, Bhavya, et al.
Published: (2026)
by: Vasudeva, Bhavya, et al.
Published: (2026)
KITE: Kernelized and Information Theoretic Exemplars for In-Context Learning
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Adversarial Vulnerability as a Consequence of On-Manifold Inseparibility
by: Haldar, Rajdeep, et al.
Published: (2024)
by: Haldar, Rajdeep, et al.
Published: (2024)
InfiniteHiP: Extending Language Model Context Up to 3 Million Tokens on a Single GPU
by: Lee, Heejun, et al.
Published: (2025)
by: Lee, Heejun, et al.
Published: (2025)
End-To-End Causal Effect Estimation from Unstructured Natural Language Data
by: Dhawan, Nikita, et al.
Published: (2024)
by: Dhawan, Nikita, et al.
Published: (2024)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Folded Context Condensation in Path Integral Formalism for Infinite Context Transformers
by: Paeng, Won-Gi, et al.
Published: (2024)
by: Paeng, Won-Gi, et al.
Published: (2024)
f-GRPO and Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment
by: Haldar, Rajdeep, et al.
Published: (2026)
by: Haldar, Rajdeep, et al.
Published: (2026)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
What is Wrong with Perplexity for Long-context Language Modeling?
by: Fang, Lizhe, et al.
Published: (2024)
by: Fang, Lizhe, et al.
Published: (2024)
TimeAutoDiff: A Unified Framework for Generation, Imputation, Forecasting, and Time-Varying Metadata Conditioning of Heterogeneous Time Series Tabular Data
by: Suh, Namjoon, et al.
Published: (2024)
by: Suh, Namjoon, et al.
Published: (2024)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
by: Song, Woomin, et al.
Published: (2025)
by: Song, Woomin, et al.
Published: (2025)
TyphoFormer: Language-Augmented Transformer for Accurate Typhoon Track Forecasting
by: Li, Lincan, et al.
Published: (2025)
by: Li, Lincan, et al.
Published: (2025)
A Theoretical Understanding of Self-Correction through In-context Alignment
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning
by: Lee, Jaeseong, et al.
Published: (2024)
by: Lee, Jaeseong, et al.
Published: (2024)
Similar Items
-
A Survey on Statistical Theory of Deep Learning: Approximation, Training Dynamics, and Generative Models
by: Suh, Namjoon, et al.
Published: (2024) -
CTSyn: A Foundation Model for Cross Tabular Data Generation
by: Lin, Xiaofeng, et al.
Published: (2024) -
Better Representations via Adversarial Training in Pre-Training: A Theoretical Perspective
by: Xing, Yue, et al.
Published: (2024) -
From Unstructured Data to In-Context Learning: Exploring What Tasks Can Be Learned and When
by: Wibisono, Kevin Christian, et al.
Published: (2024) -
Discriminative Estimation of Total Variation Distance: A Fidelity Auditor for Generative Data
by: Tao, Lan, et al.
Published: (2024)