The Convergence Gap: Instruction-Tuned Language Models Stabilize Later in the Forward Pass
Fuente:
arXiv
Saved in:
| Main Author: | Zhou, Yifan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
by: Yang, Yifan, et al.
Published: (2025)
by: Yang, Yifan, et al.
Published: (2025)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Instruction Tuning Changes How Upstream State Conditions Late Readout: A Cross-Patching Diagnostic
by: Zhou, Yifan
Published: (2026)
by: Zhou, Yifan
Published: (2026)
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)
by: Li, Tung-Ling, et al.
Published: (2025)
Test-Time Model Adaptation with Only Forward Passes
by: Niu, Shuaicheng, et al.
Published: (2024)
by: Niu, Shuaicheng, et al.
Published: (2024)
Beyond Anti-Forgetting: Multimodal Continual Instruction Tuning with Positive Forward Transfer
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
by: Chen, Tong, et al.
Published: (2024)
by: Chen, Tong, et al.
Published: (2024)
The Mirrored Influence Hypothesis: Efficient Data Influence Estimation by Harnessing Forward Passes
by: Ko, Myeongseob, et al.
Published: (2024)
by: Ko, Myeongseob, et al.
Published: (2024)
Instruction Tuning Chronologically Consistent Language Models
by: He, Songrun, et al.
Published: (2025)
by: He, Songrun, et al.
Published: (2025)
SAME: Stabilized Mixture-of-Experts for Multimodal Continual Instruction Tuning
by: Xie, Zhen-Hao, et al.
Published: (2026)
by: Xie, Zhen-Hao, et al.
Published: (2026)
Tuning-Free Bilevel Optimization: New Algorithms and Convergence Analysis
by: Yang, Yifan, et al.
Published: (2024)
by: Yang, Yifan, et al.
Published: (2024)
On Tuning Neural ODE for Stability, Consistency and Faster Convergence
by: Akhtar, Sheikh Waqas
Published: (2023)
by: Akhtar, Sheikh Waqas
Published: (2023)
DCFold: Efficient Protein Structure Generation with Single Forward Pass
by: Zhang, Zhe, et al.
Published: (2026)
by: Zhang, Zhe, et al.
Published: (2026)
Filter Images First, Generate Instructions Later: Pre-Instruction Data Selection for Visual Instruction Tuning
by: Safaei, Bardia, et al.
Published: (2025)
by: Safaei, Bardia, et al.
Published: (2025)
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
by: Xiao, Yuxin, et al.
Published: (2024)
by: Xiao, Yuxin, et al.
Published: (2024)
On the Convergence of Zeroth-Order Federated Tuning for Large Language Models
by: Ling, Zhenqing, et al.
Published: (2024)
by: Ling, Zhenqing, et al.
Published: (2024)
Constrained Particle Seeking: Solving Diffusion Inverse Problems with Just Forward Passes
by: Dou, Hongkun, et al.
Published: (2026)
by: Dou, Hongkun, et al.
Published: (2026)
A Differentiable Partially Observable Generalized Linear Model with Forward-Backward Message Passing
by: Li, Chengrui, et al.
Published: (2024)
by: Li, Chengrui, et al.
Published: (2024)
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
by: Cao, Yihan, et al.
Published: (2023)
by: Cao, Yihan, et al.
Published: (2023)
Semi-supervised Instruction Tuning for Large Language Models on Text-Attributed Graphs
by: Song, Zixing, et al.
Published: (2026)
by: Song, Zixing, et al.
Published: (2026)
Investigating the Multilingual Calibration Effects of Language Model Instruction-Tuning
by: Huang, Jerry, et al.
Published: (2026)
by: Huang, Jerry, et al.
Published: (2026)
Through the Gaps: Uncovering Tactical Line-Breaking Passes with Clustering
by: Karakuş, Oktay, et al.
Published: (2025)
by: Karakuş, Oktay, et al.
Published: (2025)
Training Large-Scale Optical Neural Networks with Two-Pass Forward Propagation
by: Ahmadnejad, Amirreza, et al.
Published: (2024)
by: Ahmadnejad, Amirreza, et al.
Published: (2024)
Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
by: Kolawole, Steven, et al.
Published: (2024)
by: Kolawole, Steven, et al.
Published: (2024)
EVA-0: Test-Time Model Evolution with Only Two Forward Passes per Sample
by: Chen, Guohao, et al.
Published: (2026)
by: Chen, Guohao, et al.
Published: (2026)
Instruction Tuning for Large Language Models: A Survey
by: Zhang, Shengyu, et al.
Published: (2023)
by: Zhang, Shengyu, et al.
Published: (2023)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Analyzing and Enhancing the Backward-Pass Convergence of Unrolled Optimization
by: Kotary, James, et al.
Published: (2023)
by: Kotary, James, et al.
Published: (2023)
Instructions as Backdoors: Backdoor Vulnerabilities of Instruction Tuning for Large Language Models
by: Xu, Jiashu, et al.
Published: (2023)
by: Xu, Jiashu, et al.
Published: (2023)
Federated Data-Efficient Instruction Tuning for Large Language Models
by: Qin, Zhen, et al.
Published: (2024)
by: Qin, Zhen, et al.
Published: (2024)
PassNet: Scaling Large Language Models for Graph Compiler Pass Generation
by: Liu, Yiqun, et al.
Published: (2026)
by: Liu, Yiqun, et al.
Published: (2026)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
by: Li, Kenneth, et al.
Published: (2024)
by: Li, Kenneth, et al.
Published: (2024)
From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction Tuning
by: Wu, Xuansheng, et al.
Published: (2023)
by: Wu, Xuansheng, et al.
Published: (2023)
A Unified Graph Language Model for Multi-Domain Multi-Task Graph Alignment Instruction Tuning
by: Chen, Haibo, et al.
Published: (2026)
by: Chen, Haibo, et al.
Published: (2026)
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning
by: Wu, Rujie, et al.
Published: (2026)
by: Wu, Rujie, et al.
Published: (2026)
Contrastive Instruction Tuning
by: Yan, Tianyi Lorena, et al.
Published: (2024)
by: Yan, Tianyi Lorena, et al.
Published: (2024)
Convergence of Message Passing Graph Neural Networks with Generic Aggregation On Large Random Graphs
by: Cordonnier, Matthieu, et al.
Published: (2023)
by: Cordonnier, Matthieu, et al.
Published: (2023)
Personalized Federated Instruction Tuning via Neural Architecture Search
by: Zhang, Pengyu, et al.
Published: (2024)
by: Zhang, Pengyu, et al.
Published: (2024)
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
by: Dima, George-Andrei, et al.
Published: (2025)
by: Dima, George-Andrei, et al.
Published: (2025)
Similar Items
-
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023) -
SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
by: Yang, Yifan, et al.
Published: (2025) -
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026) -
Instruction Tuning Changes How Upstream State Conditions Late Readout: A Cross-Patching Diagnostic
by: Zhou, Yifan
Published: (2026) -
Logit-Gap Steering: A Forward-Pass Diagnostic for Alignment Robustness
by: Li, Tung-Ling, et al.
Published: (2025)