When and How Unlabeled Data Provably Improve In-Context Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yingcong, Chang, Xiangyu, Kara, Muti, Liu, Xiaofeng, Roy-Chowdhury, Amit, Oymak, Samet |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Benefits of Task-Specific Prompts for In-context Learning
by: Chang, Xiangyu, et al.
Published: (2025)
by: Chang, Xiangyu, et al.
Published: (2025)
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023)
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)
by: Li, Yingcong, et al.
Published: (2024)
Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
One-Layer Transformer Provably Learns One-Nearest Neighbor In Context
by: Li, Zihao, et al.
Published: (2024)
by: Li, Zihao, et al.
Published: (2024)
A Median Perspective on Unlabeled Data for Out-of-Distribution Detection
by: Abbas, Momin, et al.
Published: (2025)
by: Abbas, Momin, et al.
Published: (2025)
Q-Learning for Stochastic Control under General Information Structures and Non-Markovian Environments
by: Kara, Ali Devran, et al.
Published: (2023)
by: Kara, Ali Devran, et al.
Published: (2023)
Agentic Transformers Provably Learn to Search via Reinforcement Learning
by: Yang, Tong, et al.
Published: (2026)
by: Yang, Tong, et al.
Published: (2026)
Exploring the Robustness of In-Context Learning with Noisy Labels
by: Cheng, Chen, et al.
Published: (2024)
by: Cheng, Chen, et al.
Published: (2024)
Transformers Provably Learn Chain-of-Thought Reasoning with Length Generalization
by: Huang, Yu, et al.
Published: (2025)
by: Huang, Yu, et al.
Published: (2025)
When and Why SignSGD Outperforms SGD: A Theoretical Study Based on $\ell_1$-norm Lower Bounds
by: Tao, Hongyi, et al.
Published: (2026)
by: Tao, Hongyi, et al.
Published: (2026)
In-Context Learning Under Regime Change
by: Dudley, Carson, et al.
Published: (2026)
by: Dudley, Carson, et al.
Published: (2026)
Policy Gradient Methods for Risk-Sensitive Distributional Reinforcement Learning with Provable Convergence
by: Xiao, Minheng, et al.
Published: (2024)
by: Xiao, Minheng, et al.
Published: (2024)
Provable Offline Reinforcement Learning for Structured Cyclic MDPs
by: Lee, Kyungbok, et al.
Published: (2026)
by: Lee, Kyungbok, et al.
Published: (2026)
From Self-Attention to Markov Models: Unveiling the Dynamics of Generative Transformers
by: Ildiz, M. Emrullah, et al.
Published: (2024)
by: Ildiz, M. Emrullah, et al.
Published: (2024)
Advances and Challenges in Semantic Textual Similarity: A Comprehensive Survey
by: Kumar, Lokendra, et al.
Published: (2025)
by: Kumar, Lokendra, et al.
Published: (2025)
Provably Safe Generative Sampling with Constricting Barrier Functions
by: Gadginmath, Darshan, et al.
Published: (2026)
by: Gadginmath, Darshan, et al.
Published: (2026)
Provable Acceleration for Diffusion Models under Minimal Assumptions
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
Causal LLM Routing: End-to-End Regret Minimization from Observational Data
by: Tsiourvas, Asterios, et al.
Published: (2025)
by: Tsiourvas, Asterios, et al.
Published: (2025)
Variational Learning is Effective for Large Deep Networks
by: Shen, Yuesong, et al.
Published: (2024)
by: Shen, Yuesong, et al.
Published: (2024)
ACING: Actor-Critic for Instruction Learning in Black-Box LLMs
by: Kharrat, Salma, et al.
Published: (2024)
by: Kharrat, Salma, et al.
Published: (2024)
Reinforcement Learning from Human Feedback with Active Queries
by: Ji, Kaixuan, et al.
Published: (2024)
by: Ji, Kaixuan, et al.
Published: (2024)
SMiLE: Provably Enforcing Global Relational Properties in Neural Networks
by: Francobaldi, Matteo, et al.
Published: (2025)
by: Francobaldi, Matteo, et al.
Published: (2025)
Can Transformers Learn Optimal Filtering for Unknown Systems?
by: Balim, Haldun, et al.
Published: (2023)
by: Balim, Haldun, et al.
Published: (2023)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
by: Fernando, Heshan, et al.
Published: (2024)
by: Fernando, Heshan, et al.
Published: (2024)
Graphon Particle Systems, Part II: Dynamics of Distributed Stochastic Continuum Optimization
by: Chen, Yan, et al.
Published: (2024)
by: Chen, Yan, et al.
Published: (2024)
Provably Learning Diverse Features in Multi-View Data with Midpoint Mixup
by: Chidambaram, Muthu, et al.
Published: (2022)
by: Chidambaram, Muthu, et al.
Published: (2022)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Selective Attention: Enhancing Transformer through Principled Context Control
by: Zhang, Xuechen, et al.
Published: (2024)
by: Zhang, Xuechen, et al.
Published: (2024)
Correlated Noise Provably Beats Independent Noise for Differentially Private Learning
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
by: Choquette-Choo, Christopher A., et al.
Published: (2023)
Applications of 0-1 Neural Networks in Prescription and Prediction
by: Patil, Vrishabh, et al.
Published: (2024)
by: Patil, Vrishabh, et al.
Published: (2024)
A Method to Improve the Performance of Reinforcement Learning Based on the Y Operator for a Class of Stochastic Differential Equation-Based Child-Mother Systems
by: Yin, Cheng, et al.
Published: (2023)
by: Yin, Cheng, et al.
Published: (2023)
Neural-Rendezvous: Provably Robust Guidance and Control to Encounter Interstellar Objects
by: Tsukamoto, Hiroyasu, et al.
Published: (2022)
by: Tsukamoto, Hiroyasu, et al.
Published: (2022)
LISA: Layerwise Importance Sampling for Memory-Efficient Large Language Model Fine-Tuning
by: Pan, Rui, et al.
Published: (2024)
by: Pan, Rui, et al.
Published: (2024)
A Context-Free Smart Grid Model Using Complex System Approach
by: Amor, Soufian Ben, et al.
Published: (2025)
by: Amor, Soufian Ben, et al.
Published: (2025)
An Improved Last-Iterate Convergence Rate for Anchored Gradient Descent Ascent
by: Surina, Anja, et al.
Published: (2026)
by: Surina, Anja, et al.
Published: (2026)
Comparative Analysis of Optimization Strategies for K-means Clustering in Big Data Contexts: A Review
by: Mussabayev, Ravil, et al.
Published: (2023)
by: Mussabayev, Ravil, et al.
Published: (2023)
Multi-head Transformers Provably Learn Symbolic Multi-step Reasoning via Gradient Descent
by: Yang, Tong, et al.
Published: (2025)
by: Yang, Tong, et al.
Published: (2025)
Similar Items
-
Provable Benefits of Task-Specific Prompts for In-context Learning
by: Chang, Xiangyu, et al.
Published: (2025) -
Fine-grained Analysis of In-context Linear Estimation: Data, Architecture, and Beyond
by: Li, Yingcong, et al.
Published: (2024) -
Transformers as Support Vector Machines
by: Tarzanagh, Davoud Ataee, et al.
Published: (2023) -
Gating is Weighting: Understanding Gated Linear Attention through In-context Learning
by: Li, Yingcong, et al.
Published: (2025) -
Mechanics of Next Token Prediction with Self-Attention
by: Li, Yingcong, et al.
Published: (2024)