Test-Time Adaptation via Many-Shot Prompting: Benefits, Limits, and Pitfalls
Fuente:
arXiv
Saved in:
| Main Authors: | Upasani, Shubhangi, Wu, Chen, Rainton, Jay, Li, Bo, Thakker, Urmish, Hu, Changran, Zhang, Qizheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
by: Zhang, Qizheng, et al.
Published: (2025)
by: Zhang, Qizheng, et al.
Published: (2025)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025)
by: Wang, Guangtao, et al.
Published: (2025)
The Limits of Long-Context Reasoning in Automated Bug Fixing
by: Raju, Ravi, et al.
Published: (2026)
by: Raju, Ravi, et al.
Published: (2026)
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
by: Zhao, Xueliang, et al.
Published: (2024)
by: Zhao, Xueliang, et al.
Published: (2024)
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
by: Upasani, Shubhangi, et al.
Published: (2026)
by: Upasani, Shubhangi, et al.
Published: (2026)
Training Domain Draft Models for Speculative Decoding: Best Practices and Insights
by: Hong, Fenglu, et al.
Published: (2025)
by: Hong, Fenglu, et al.
Published: (2025)
SambaLingo: Teaching Large Language Models New Languages
by: Csaki, Zoltan, et al.
Published: (2024)
by: Csaki, Zoltan, et al.
Published: (2024)
Constructing Domain-Specific Evaluation Sets for LLM-as-a-judge
by: Raju, Ravi, et al.
Published: (2024)
by: Raju, Ravi, et al.
Published: (2024)
Rethinking Pruning Large Language Models: Benefits and Pitfalls of Reconstruction Error Minimization
by: Shin, Sungbin, et al.
Published: (2024)
by: Shin, Sungbin, et al.
Published: (2024)
Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective
by: Wang, Siwei, et al.
Published: (2025)
by: Wang, Siwei, et al.
Published: (2025)
Towards Reliable Latent Knowledge Estimation in LLMs: Zero-Prompt Many-Shot Based Factual Knowledge Extraction
by: Wu, Qinyuan, et al.
Published: (2024)
by: Wu, Qinyuan, et al.
Published: (2024)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
by: Zhang, Qizheng, et al.
Published: (2025)
by: Zhang, Qizheng, et al.
Published: (2025)
Many-Shot In-Context Learning
by: Agarwal, Rishabh, et al.
Published: (2024)
by: Agarwal, Rishabh, et al.
Published: (2024)
DynaPrompt: Dynamic Test-Time Prompt Tuning
by: Xiao, Zehao, et al.
Published: (2025)
by: Xiao, Zehao, et al.
Published: (2025)
Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Many-Shot Regurgitation (MSR) Prompting
by: Sonkar, Shashank, et al.
Published: (2024)
by: Sonkar, Shashank, et al.
Published: (2024)
Towards Compute-Optimal Many-Shot In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2025)
by: Golchin, Shahriar, et al.
Published: (2025)
Theoretical Benefit and Limitation of Diffusion Language Model
by: Feng, Guhao, et al.
Published: (2025)
by: Feng, Guhao, et al.
Published: (2025)
FAAST: Forward-Only Associative Learning via Closed-Form Fast Weights for Test-Time Supervised Adaptation
by: Bao, Guangsheng, et al.
Published: (2026)
by: Bao, Guangsheng, et al.
Published: (2026)
PromptAL: Sample-Aware Dynamic Soft Prompts for Few-Shot Active Learning
by: Xiang, Hui, et al.
Published: (2025)
by: Xiang, Hui, et al.
Published: (2025)
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
by: Yan, Kai, et al.
Published: (2025)
by: Yan, Kai, et al.
Published: (2025)
Distributed LLMs and Multimodal Large Language Models: A Survey on Advances, Challenges, and Future Directions
by: Amini, Hadi, et al.
Published: (2025)
by: Amini, Hadi, et al.
Published: (2025)
Scaling Laws for Many-Shot In-Context Learning with Self-Generated Annotations
by: Gu, Zhengyao, et al.
Published: (2025)
by: Gu, Zhengyao, et al.
Published: (2025)
Improving Zero-Shot Cross-Lingual Transfer via Progressive Code-Switching
by: Li, Zhuoran, et al.
Published: (2024)
by: Li, Zhuoran, et al.
Published: (2024)
Dynamic Prompt Fusion for Multi-Task and Cross-Domain Adaptation in LLMs
by: Hu, Xin, et al.
Published: (2025)
by: Hu, Xin, et al.
Published: (2025)
More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
by: Zhang, Xiaoqing, et al.
Published: (2025)
by: Zhang, Xiaoqing, et al.
Published: (2025)
The Surprising Effectiveness of Test-Time Training for Few-Shot Learning
by: Akyürek, Ekin, et al.
Published: (2024)
by: Akyürek, Ekin, et al.
Published: (2024)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
by: Zhang, Zhaowei, et al.
Published: (2025)
by: Zhang, Zhaowei, et al.
Published: (2025)
AdaRewriter: Unleashing the Power of Prompting-based Conversational Query Reformulation via Test-Time Adaptation
by: Lai, Yilong, et al.
Published: (2025)
by: Lai, Yilong, et al.
Published: (2025)
PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
by: Park, Junho, et al.
Published: (2026)
by: Park, Junho, et al.
Published: (2026)
Towards Few-Shot Adaptation of Foundation Models via Multitask Finetuning
by: Xu, Zhuoyan, et al.
Published: (2024)
by: Xu, Zhuoyan, et al.
Published: (2024)
Synthetic Document Question Answering in Hungarian
by: Li, Jonathan, et al.
Published: (2025)
by: Li, Jonathan, et al.
Published: (2025)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
by: Gu, Zhuohan, et al.
Published: (2026)
by: Gu, Zhuohan, et al.
Published: (2026)
Not How Many, But Which: Parameter Placement in Low-Rank Adaptation
by: Sehanobish, Arijit, et al.
Published: (2026)
by: Sehanobish, Arijit, et al.
Published: (2026)
Many-Shot In-Context Learning in Multimodal Foundation Models
by: Jiang, Yixing, et al.
Published: (2024)
by: Jiang, Yixing, et al.
Published: (2024)
How Many Human Judgments Are Enough? Feasibility Limits of Human Preference Evaluation
by: Lee, Wilson Y.
Published: (2026)
by: Lee, Wilson Y.
Published: (2026)
PromptFix: Few-shot Backdoor Removal via Adversarial Prompt Tuning
by: Zhang, Tianrong, et al.
Published: (2024)
by: Zhang, Tianrong, et al.
Published: (2024)
MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
by: Phan, Peter, et al.
Published: (2025)
by: Phan, Peter, et al.
Published: (2025)
Sparse Transformer with Local and Seasonal Adaptation for Multivariate Time Series Forecasting
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 Bits
by: Zhou, Zikai, et al.
Published: (2025)
by: Zhou, Zikai, et al.
Published: (2025)
Similar Items
-
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
by: Zhang, Qizheng, et al.
Published: (2025) -
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
by: Wang, Guangtao, et al.
Published: (2025) -
The Limits of Long-Context Reasoning in Automated Bug Fixing
by: Raju, Ravi, et al.
Published: (2026) -
SubgoalXL: Subgoal-based Expert Learning for Theorem Proving
by: Zhao, Xueliang, et al.
Published: (2024) -
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models
by: Upasani, Shubhangi, et al.
Published: (2026)