Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Liu, Lin, Ziqian, Lee, Kangwook, Papailiopoulos, Dimitris, Nowak, Robert |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Looped Transformers are Better at Learning Learning Algorithms
di: Yang, Liu, et al.
Pubblicazione: (2023)
di: Yang, Liu, et al.
Pubblicazione: (2023)
Dual Operating Modes of In-Context Learning
di: Lin, Ziqian, et al.
Pubblicazione: (2024)
di: Lin, Ziqian, et al.
Pubblicazione: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
di: Park, Jongho, et al.
Pubblicazione: (2024)
di: Park, Jongho, et al.
Pubblicazione: (2024)
In-Context Learning with Hypothesis-Class Guidance
di: Lin, Ziqian, et al.
Pubblicazione: (2025)
di: Lin, Ziqian, et al.
Pubblicazione: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
di: Yang, Seongjun, et al.
Pubblicazione: (2023)
di: Yang, Seongjun, et al.
Pubblicazione: (2023)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
di: Lee, Nayoung, et al.
Pubblicazione: (2025)
di: Lee, Nayoung, et al.
Pubblicazione: (2025)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
di: Xiong, Zheyang, et al.
Pubblicazione: (2024)
Variation Spaces for Multi-Output Neural Networks: Insights on Multi-Task Learning and Network Compression
di: Shenouda, Joseph, et al.
Pubblicazione: (2023)
di: Shenouda, Joseph, et al.
Pubblicazione: (2023)
ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
di: Zeng, Yuchen, et al.
Pubblicazione: (2025)
di: Zeng, Yuchen, et al.
Pubblicazione: (2025)
Understanding Task Vectors in In-Context Learning: Emergence, Functionality, and Limitations
di: Dong, Yuxin, et al.
Pubblicazione: (2025)
di: Dong, Yuxin, et al.
Pubblicazione: (2025)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
di: Lee, Chungpa, et al.
Pubblicazione: (2026)
di: Lee, Chungpa, et al.
Pubblicazione: (2026)
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
di: Kim, Junhyuck, et al.
Pubblicazione: (2024)
di: Kim, Junhyuck, et al.
Pubblicazione: (2024)
Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective
di: Han, Seungwook, et al.
Pubblicazione: (2024)
di: Han, Seungwook, et al.
Pubblicazione: (2024)
How Well Can Transformers Emulate In-context Newton's Method?
di: Giannou, Angeliki, et al.
Pubblicazione: (2024)
di: Giannou, Angeliki, et al.
Pubblicazione: (2024)
Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
di: Kim, Junhyuck, et al.
Pubblicazione: (2025)
di: Kim, Junhyuck, et al.
Pubblicazione: (2025)
Endless Terminals: Scaling RL Environments for Terminal Agents
di: Gandhi, Kanishk, et al.
Pubblicazione: (2026)
di: Gandhi, Kanishk, et al.
Pubblicazione: (2026)
Can MLLMs Perform Text-to-Image In-Context Learning?
di: Zeng, Yuchen, et al.
Pubblicazione: (2024)
di: Zeng, Yuchen, et al.
Pubblicazione: (2024)
Optimizing DDPM Sampling with Shortcut Fine-Tuning
di: Fan, Ying, et al.
Pubblicazione: (2023)
di: Fan, Ying, et al.
Pubblicazione: (2023)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
di: Kim, Jungtaek, et al.
Pubblicazione: (2026)
di: Kim, Jungtaek, et al.
Pubblicazione: (2026)
Wait, Wait, Wait... Why Do Reasoning Models Loop?
di: Pipis, Charilaos, et al.
Pubblicazione: (2025)
di: Pipis, Charilaos, et al.
Pubblicazione: (2025)
MEMENTO: Teaching LLMs to Manage Their Own Context
di: Kontonis, Vasilis, et al.
Pubblicazione: (2026)
di: Kontonis, Vasilis, et al.
Pubblicazione: (2026)
The Effects of Multi-Task Learning on ReLU Neural Network Functions
di: Nakhleh, Julia, et al.
Pubblicazione: (2024)
di: Nakhleh, Julia, et al.
Pubblicazione: (2024)
The Expressive Power of Low-Rank Adaptation
di: Zeng, Yuchen, et al.
Pubblicazione: (2023)
di: Zeng, Yuchen, et al.
Pubblicazione: (2023)
Sample More to Think Less: Group Filtered Policy Optimization for Concise Reasoning
di: Shrivastava, Vaishnavi, et al.
Pubblicazione: (2025)
di: Shrivastava, Vaishnavi, et al.
Pubblicazione: (2025)
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
di: Zeng, Thomas, et al.
Pubblicazione: (2025)
di: Zeng, Thomas, et al.
Pubblicazione: (2025)
Context and Diversity Matter: The Emergence of In-Context Learning in World Models
di: Wang, Fan, et al.
Pubblicazione: (2025)
di: Wang, Fan, et al.
Pubblicazione: (2025)
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
di: Mukherjee, Subhojyoti, et al.
Pubblicazione: (2024)
Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
di: Xie, Zixuan, et al.
Pubblicazione: (2026)
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
di: Ma, Qian, et al.
Pubblicazione: (2025)
di: Ma, Qian, et al.
Pubblicazione: (2025)
Efficient Active Learning with Abstention
di: Zhu, Yinglun, et al.
Pubblicazione: (2022)
di: Zhu, Yinglun, et al.
Pubblicazione: (2022)
CHAI: Clustered Head Attention for Efficient LLM Inference
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
di: Agarwal, Saurabh, et al.
Pubblicazione: (2024)
Looped Transformers for Length Generalization
di: Fan, Ying, et al.
Pubblicazione: (2024)
di: Fan, Ying, et al.
Pubblicazione: (2024)
Memorization Capacity for Additive Fine-Tuning with Small ReLU Networks
di: Sohn, Jy-yong, et al.
Pubblicazione: (2024)
di: Sohn, Jy-yong, et al.
Pubblicazione: (2024)
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
di: Bu, Dake, et al.
Pubblicazione: (2025)
di: Bu, Dake, et al.
Pubblicazione: (2025)
Towards Provable Emergence of In-Context Reinforcement Learning
di: Wang, Jiuqi, et al.
Pubblicazione: (2025)
di: Wang, Jiuqi, et al.
Pubblicazione: (2025)
Active Learning with Neural Networks: Insights from Nonparametric Statistics
di: Zhu, Yinglun, et al.
Pubblicazione: (2022)
di: Zhu, Yinglun, et al.
Pubblicazione: (2022)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
di: Kim, Sungnyun, et al.
Pubblicazione: (2025)
di: Kim, Sungnyun, et al.
Pubblicazione: (2025)
Strategy Coopetition Explains the Emergence and Transience of In-Context Learning
di: Singh, Aaditya K., et al.
Pubblicazione: (2025)
di: Singh, Aaditya K., et al.
Pubblicazione: (2025)
Emergence of In-Context Reinforcement Learning from Noise Distillation
di: Zisman, Ilya, et al.
Pubblicazione: (2023)
di: Zisman, Ilya, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Looped Transformers are Better at Learning Learning Algorithms
di: Yang, Liu, et al.
Pubblicazione: (2023) -
Dual Operating Modes of In-Context Learning
di: Lin, Ziqian, et al.
Pubblicazione: (2024) -
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
di: Park, Jongho, et al.
Pubblicazione: (2024) -
In-Context Learning with Hypothesis-Class Guidance
di: Lin, Ziqian, et al.
Pubblicazione: (2025) -
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
di: Yang, Seongjun, et al.
Pubblicazione: (2023)