Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Chungpa, Sohn, Jy-yong, Lee, Kangwook |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
Analysis of Using Sigmoid Loss for Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2024)
von: Lee, Chungpa, et al.
Veröffentlicht: (2024)
Memorization Capacity for Additive Fine-Tuning with Small ReLU Networks
von: Sohn, Jy-yong, et al.
Veröffentlicht: (2024)
von: Sohn, Jy-yong, et al.
Veröffentlicht: (2024)
On the Similarities of Embeddings in Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
von: Kim, Jungtaek, et al.
Veröffentlicht: (2026)
von: Kim, Jungtaek, et al.
Veröffentlicht: (2026)
Parameter-Efficient Fine-Tuning of State Space Models
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
Measuring Representational Shifts in Continual Learning: A Linear Transformation Perspective
von: Kim, Joonkyu, et al.
Veröffentlicht: (2025)
von: Kim, Joonkyu, et al.
Veröffentlicht: (2025)
ERD: A Framework for Improving LLM Reasoning for Cognitive Distortion Classification
von: Lim, Sehee, et al.
Veröffentlicht: (2024)
von: Lim, Sehee, et al.
Veröffentlicht: (2024)
Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning
von: Kwon, Jihoon, et al.
Veröffentlicht: (2026)
von: Kwon, Jihoon, et al.
Veröffentlicht: (2026)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025)
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025)
Buffer-based Gradient Projection for Continual Federated Learning
von: Dai, Shenghong, et al.
Veröffentlicht: (2024)
von: Dai, Shenghong, et al.
Veröffentlicht: (2024)
Deeper Insights Without Updates: The Power of In-Context Learning Over Fine-Tuning
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
von: Yin, Qingyu, et al.
Veröffentlicht: (2024)
Can MLLMs Perform Text-to-Image In-Context Learning?
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
The Expressive Power of Low-Rank Adaptation
von: Zeng, Yuchen, et al.
Veröffentlicht: (2023)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2023)
Alleviating Forgetfulness of Linear Attention by Hybrid Sparse Attention and Contextualized Learnable Token Eviction
von: He, Mutian, et al.
Veröffentlicht: (2025)
von: He, Mutian, et al.
Veröffentlicht: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
Optimizing DDPM Sampling with Shortcut Fine-Tuning
von: Fan, Ying, et al.
Veröffentlicht: (2023)
von: Fan, Ying, et al.
Veröffentlicht: (2023)
In-Context Learning and Fine-Tuning GPT for Argument Mining
von: Cabessa, Jérémie, et al.
Veröffentlicht: (2024)
von: Cabessa, Jérémie, et al.
Veröffentlicht: (2024)
Order-Independence Without Fine Tuning
von: McIlroy-Young, Reid, et al.
Veröffentlicht: (2024)
von: McIlroy-Young, Reid, et al.
Veröffentlicht: (2024)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
von: Li, He, et al.
Veröffentlicht: (2026)
von: Li, He, et al.
Veröffentlicht: (2026)
Scaling Laws for Forgetting When Fine-Tuning Large Language Models
von: Kalajdzievski, Damjan
Veröffentlicht: (2024)
von: Kalajdzievski, Damjan
Veröffentlicht: (2024)
ScoNe: Benchmarking Negation Reasoning in Language Models With Fine-Tuning and In-Context Learning
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
von: She, Jingyuan Selena, et al.
Veröffentlicht: (2023)
Forgetting to Forget: Attention Sink as A Gateway for Backdooring LLM Unlearning
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
von: Shang, Bingqi, et al.
Veröffentlicht: (2025)
Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
von: Wang, Huazheng, et al.
Veröffentlicht: (2025)
von: Wang, Huazheng, et al.
Veröffentlicht: (2025)
SEA: Sparse Linear Attention with Estimated Attention Mask
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
von: Lee, Heejun, et al.
Veröffentlicht: (2023)
Entropy-Adaptive Fine-Tuning: Resolving Confident Conflicts to Mitigate Forgetting
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
von: Diao, Muxi, et al.
Veröffentlicht: (2026)
Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
von: Lai, Song, et al.
Veröffentlicht: (2025)
von: Lai, Song, et al.
Veröffentlicht: (2025)
Fine-Tuning Language Models with Just Forward Passes
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)
von: Malladi, Sadhika, et al.
Veröffentlicht: (2023)
QEFT: Quantization for Efficient Fine-Tuning of LLMs
von: Lee, Changhun, et al.
Veröffentlicht: (2024)
von: Lee, Changhun, et al.
Veröffentlicht: (2024)
CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation
von: Fawi, Muhammad
Veröffentlicht: (2024)
von: Fawi, Muhammad
Veröffentlicht: (2024)
In-Context Fine-Tuning for Time-Series Foundation Models
von: Das, Abhimanyu, et al.
Veröffentlicht: (2024)
von: Das, Abhimanyu, et al.
Veröffentlicht: (2024)
ENTP: Encoder-only Next Token Prediction
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
von: Park, Juneyoung, et al.
Veröffentlicht: (2026)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
Forgetting Transformer: Softmax Attention with a Forget Gate
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model Agents
von: Fu, Yao, et al.
Veröffentlicht: (2024)
von: Fu, Yao, et al.
Veröffentlicht: (2024)
Dual Operating Modes of In-Context Learning
von: Lin, Ziqian, et al.
Veröffentlicht: (2024)
von: Lin, Ziqian, et al.
Veröffentlicht: (2024)
Improving Multi-lingual Alignment Through Soft Contrastive Learning
von: Park, Minsu, et al.
Veröffentlicht: (2024)
von: Park, Minsu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025) -
A Theoretical Framework for Preventing Class Collapse in Supervised Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2025) -
Analysis of Using Sigmoid Loss for Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2024) -
Memorization Capacity for Additive Fine-Tuning with Small ReLU Networks
von: Sohn, Jy-yong, et al.
Veröffentlicht: (2024) -
On the Similarities of Embeddings in Contrastive Learning
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)