Stabilizing LLM Supervised Fine-Tuning via Explicit Distributional Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xinyu, Sun, Changzhi, Wu, Yuanbin, Wang, Xiaoling |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Proximal Supervised Fine-Tuning
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024)
von: Sun, Hao
Veröffentlicht: (2024)
Logic-Regularized Verifier Elicits Reasoning from LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
von: Wang, Xinyu, et al.
Veröffentlicht: (2026)
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
von: Huang, Zeyu, et al.
Veröffentlicht: (2025)
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)
Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety
von: Shen, Qian, et al.
Veröffentlicht: (2026)
von: Shen, Qian, et al.
Veröffentlicht: (2026)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Parameter Efficient Quasi-Orthogonal Fine-Tuning via Givens Rotation
von: Ma, Xinyu, et al.
Veröffentlicht: (2024)
von: Ma, Xinyu, et al.
Veröffentlicht: (2024)
SRFT: A Single-Stage Method with Supervised and Reinforcement Fine-Tuning for Reasoning
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
von: Fu, Yuqian, et al.
Veröffentlicht: (2025)
Filter-then-Weight: Online Data Selection and Reweighting for LLM Fine-Tuning
von: Wang, Fangxin, et al.
Veröffentlicht: (2026)
von: Wang, Fangxin, et al.
Veröffentlicht: (2026)
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
von: Zou, Heming, et al.
Veröffentlicht: (2025)
von: Zou, Heming, et al.
Veröffentlicht: (2025)
Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning
von: Qian, Livia, et al.
Veröffentlicht: (2026)
von: Qian, Livia, et al.
Veröffentlicht: (2026)
Mitigating Training Imbalance in LLM Fine-Tuning via Selective Parameter Merging
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
von: Ju, Yiming, et al.
Veröffentlicht: (2024)
Towards Understanding Fine-Tuning Mechanisms of LLMs via Circuit Analysis
von: Wang, Xu, et al.
Veröffentlicht: (2025)
von: Wang, Xu, et al.
Veröffentlicht: (2025)
Fine-Tuning LLMs for Report Summarization: Analysis on Supervised and Unsupervised Data
von: Rallapalli, Swati, et al.
Veröffentlicht: (2025)
von: Rallapalli, Swati, et al.
Veröffentlicht: (2025)
IA2: Alignment with ICL Activations Improves Supervised Fine-Tuning
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
von: Mishra, Aayush, et al.
Veröffentlicht: (2025)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
von: Xia, Yuchen, et al.
Veröffentlicht: (2024)
Secure LLM Fine-Tuning via Safety-Aware Probing
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
von: Wu, Chengcan, et al.
Veröffentlicht: (2025)
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
von: Meng, Haoming, et al.
Veröffentlicht: (2026)
LLMs Meet Finance: Fine-Tuning Foundation Models for the Open FinLLM Leaderboard
von: Rao, Varun, et al.
Veröffentlicht: (2025)
von: Rao, Varun, et al.
Veröffentlicht: (2025)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning
von: Huang, Wei, et al.
Veröffentlicht: (2026)
von: Huang, Wei, et al.
Veröffentlicht: (2026)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning with Discrete Fourier Transform
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
von: Gao, Ziqi, et al.
Veröffentlicht: (2024)
Understanding Forgetting in LLM Supervised Fine-Tuning and Preference Learning -- A Convex Optimization Perspective
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
von: Fernando, Heshan, et al.
Veröffentlicht: (2024)
Enhancing LLM Reliability via Explicit Knowledge Boundary Modeling
von: Zheng, Hang, et al.
Veröffentlicht: (2025)
von: Zheng, Hang, et al.
Veröffentlicht: (2025)
MSSR: Memory-Aware Adaptive Replay for Continual LLM Fine-Tuning
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
von: Lu, Yiyang, et al.
Veröffentlicht: (2026)
Rethinking Parameter Sharing for LLM Fine-Tuning with Multiple LoRAs
von: Ban, Hao, et al.
Veröffentlicht: (2025)
von: Ban, Hao, et al.
Veröffentlicht: (2025)
Boosting Large Language Models with Mask Fine-Tuning
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Mingyuan, et al.
Veröffentlicht: (2025)
Crafting Efficient Fine-Tuning Strategies for Large Language Models
von: Oliver, Michael, et al.
Veröffentlicht: (2024)
von: Oliver, Michael, et al.
Veröffentlicht: (2024)
Parameter-Efficient Fine-Tuning for Foundation Models
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
von: Zhang, Dan, et al.
Veröffentlicht: (2025)
CURLoRA: Stable LLM Continual Fine-Tuning and Catastrophic Forgetting Mitigation
von: Fawi, Muhammad
Veröffentlicht: (2024)
von: Fawi, Muhammad
Veröffentlicht: (2024)
Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2025)
von: Ayala, Orlando Marquez, et al.
Veröffentlicht: (2025)
Aligning Large Language Models via Fine-grained Supervision
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
von: Xu, Dehong, et al.
Veröffentlicht: (2024)
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
von: Wang, Shuhe, et al.
Veröffentlicht: (2024)
AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air
von: Yang, Shiyi, et al.
Veröffentlicht: (2025)
von: Yang, Shiyi, et al.
Veröffentlicht: (2025)
EMORL: Ensemble Multi-Objective Reinforcement Learning for Efficient and Flexible LLM Fine-Tuning
von: Kong, Lingxiao, et al.
Veröffentlicht: (2025)
von: Kong, Lingxiao, et al.
Veröffentlicht: (2025)
LIFT the Veil for the Truth: Principal Weights Emerge after Rank Reduction for Reasoning-Focused Supervised Fine-Tuning
von: Liu, Zihang, et al.
Veröffentlicht: (2025)
von: Liu, Zihang, et al.
Veröffentlicht: (2025)
Upcycling Instruction Tuning from Dense to Mixture-of-Experts via Parameter Merging
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
von: Hui, Tingfeng, et al.
Veröffentlicht: (2024)
Prompting and Fine-Tuning of Small LLMs for Length-Controllable Telephone Call Summarization
von: Thulke, David, et al.
Veröffentlicht: (2024)
von: Thulke, David, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Proximal Supervised Fine-Tuning
von: Zhu, Wenhong, et al.
Veröffentlicht: (2025) -
Supervised Fine-Tuning as Inverse Reinforcement Learning
von: Sun, Hao
Veröffentlicht: (2024) -
Logic-Regularized Verifier Elicits Reasoning from LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2026) -
Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
von: Huang, Zeyu, et al.
Veröffentlicht: (2025) -
Supervised Fine-Tuning Needs to Unlock the Potential of Token Priority
von: Shen, Zhanming, et al.
Veröffentlicht: (2026)