Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xie, Shiming, Chen, Hong, Yu, Fred, Sun, Zeye, Wu, Xiuyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Minor DPO reject penalty to increase training robustness
von: Xie, Shiming, et al.
Veröffentlicht: (2024)
von: Xie, Shiming, et al.
Veröffentlicht: (2024)
How does fine-tuning improve sensorimotor representations in large language models?
von: Wu, Minghua, et al.
Veröffentlicht: (2026)
von: Wu, Minghua, et al.
Veröffentlicht: (2026)
PDC & DM-SFT: A Road for LLM SQL Bug-Fix Enhancing
von: Duan, Yiwen, et al.
Veröffentlicht: (2024)
von: Duan, Yiwen, et al.
Veröffentlicht: (2024)
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
von: Han, Yifu, et al.
Veröffentlicht: (2025)
von: Han, Yifu, et al.
Veröffentlicht: (2025)
Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
Order Doesn't Matter, But Reasoning Does: Training LLMs with Order-Centric Augmentation
von: He, Qianxi, et al.
Veröffentlicht: (2025)
von: He, Qianxi, et al.
Veröffentlicht: (2025)
Retention analysis of edited knowledge after fine-tuning
von: Wen, Fufang, et al.
Veröffentlicht: (2025)
von: Wen, Fufang, et al.
Veröffentlicht: (2025)
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration
von: Ma, Zerun, et al.
Veröffentlicht: (2026)
von: Ma, Zerun, et al.
Veröffentlicht: (2026)
Rethinking harmless refusals when fine-tuning foundation models
von: Pop, Florin, et al.
Veröffentlicht: (2024)
von: Pop, Florin, et al.
Veröffentlicht: (2024)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
von: Liu, Tao, et al.
Veröffentlicht: (2026)
von: Liu, Tao, et al.
Veröffentlicht: (2026)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
Panacea: Mitigating Harmful Fine-tuning for Large Language Models via Post-fine-tuning Perturbation
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
von: Wang, Yibo, et al.
Veröffentlicht: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Metaphor identification using large language models: A comparison of RAG, prompt engineering, and fine-tuning
von: Fuoli, Matteo, et al.
Veröffentlicht: (2025)
von: Fuoli, Matteo, et al.
Veröffentlicht: (2025)
Order Matters: Investigate the Position Bias in Multi-constraint Instruction Following
von: Zeng, Jie, et al.
Veröffentlicht: (2025)
von: Zeng, Jie, et al.
Veröffentlicht: (2025)
Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Following
von: Ren, Qingyu, et al.
Veröffentlicht: (2025)
von: Ren, Qingyu, et al.
Veröffentlicht: (2025)
Step-by-Step Mastery: Enhancing Soft Constraint Following Ability of Large Language Models
von: Ren, Qingyu, et al.
Veröffentlicht: (2025)
von: Ren, Qingyu, et al.
Veröffentlicht: (2025)
Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
von: Song, Yueqi, et al.
Veröffentlicht: (2025)
RLHF in an SFT Way: From Optimal Solution to Reward-Weighted Alignment
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
von: Du, Yuhao, et al.
Veröffentlicht: (2025)
Investigating the performance of Retrieval-Augmented Generation and fine-tuning for the development of AI-driven knowledge-based systems
von: Lakatos, Robert, et al.
Veröffentlicht: (2024)
von: Lakatos, Robert, et al.
Veröffentlicht: (2024)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
Continual SFT Matches Multimodal RLHF with Negative Supervision
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024)
von: Maharjan, Jenish, et al.
Veröffentlicht: (2024)
Using Optimal Transport as Alignment Objective for fine-tuning Multilingual Contextualized Embeddings
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2021)
von: Alqahtani, Sawsan, et al.
Veröffentlicht: (2021)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
von: Lu, Jun, et al.
Veröffentlicht: (2024)
von: Lu, Jun, et al.
Veröffentlicht: (2024)
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
von: Vargas, Francisco, et al.
Veröffentlicht: (2025)
von: Vargas, Francisco, et al.
Veröffentlicht: (2025)
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
von: Wu, Eric, et al.
Veröffentlicht: (2024)
von: Wu, Eric, et al.
Veröffentlicht: (2024)
Response-free item difficulty modelling for multiple-choice items with fine-tuned transformers: Component-wise representation and multi-task learning
von: Netík, Jan, et al.
Veröffentlicht: (2026)
von: Netík, Jan, et al.
Veröffentlicht: (2026)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
RandLoRA: Full-rank parameter-efficient fine-tuning of large models
von: Albert, Paul, et al.
Veröffentlicht: (2025)
von: Albert, Paul, et al.
Veröffentlicht: (2025)
Airavata: Introducing Hindi Instruction-tuned LLM
von: Gala, Jay, et al.
Veröffentlicht: (2024)
von: Gala, Jay, et al.
Veröffentlicht: (2024)
Improving LLM Safety and Helpfulness using SFT and DPO: A Study on OPT-350M
von: Pant, Piyush
Veröffentlicht: (2025)
von: Pant, Piyush
Veröffentlicht: (2025)
LoX: Low-Rank Extrapolation Robustifies LLM Safety Against Fine-tuning
von: Perin, Gabriel J., et al.
Veröffentlicht: (2025)
von: Perin, Gabriel J., et al.
Veröffentlicht: (2025)
A new approach for fine-tuning sentence transformers for intent classification and out-of-scope detection tasks
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhang, Tianyi, et al.
Veröffentlicht: (2024)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
von: Wang, Sudong, et al.
Veröffentlicht: (2026)
von: Wang, Sudong, et al.
Veröffentlicht: (2026)
RTLC -- Research, Teach-to-Learn, Critique: A three-stage prompting paradigm inspired by the Feynman Learning Technique that lifts LLM-as-judge accuracy on JudgeBench with no fine-tuning
von: Morandi, Andrea
Veröffentlicht: (2026)
von: Morandi, Andrea
Veröffentlicht: (2026)
SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
von: Shen, Han, et al.
Veröffentlicht: (2024)
von: Shen, Han, et al.
Veröffentlicht: (2024)
Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
von: Balabanov, Oleksandr, et al.
Veröffentlicht: (2024)
von: Balabanov, Oleksandr, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Minor DPO reject penalty to increase training robustness
von: Xie, Shiming, et al.
Veröffentlicht: (2024) -
How does fine-tuning improve sensorimotor representations in large language models?
von: Wu, Minghua, et al.
Veröffentlicht: (2026) -
PDC & DM-SFT: A Road for LLM SQL Bug-Fix Enhancing
von: Duan, Yiwen, et al.
Veröffentlicht: (2024) -
Reinforcement learning fine-tuning of language model for instruction following and math reasoning
von: Han, Yifu, et al.
Veröffentlicht: (2025) -
Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)