Gradients Must Earn Their Influence: Unifying SFT with Generalized Entropic Objectives
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zecheng, Liu, Deyuan, Li, Chunshan, Zhang, Yupeng, Zhao, Zhengyun, Chu, Dianhui, Wang, Bingning, Sui, Dianbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Beyond Confidence: The Rhythms of Reasoning in Generative Models
von: Liu, Deyuan, et al.
Veröffentlicht: (2026)
von: Liu, Deyuan, et al.
Veröffentlicht: (2026)
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
Mitigating Gender Bias in Code Large Language Models via Model Editing
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
von: Liu, Deyuan, et al.
Veröffentlicht: (2024)
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
von: Qin, Zhanyue, et al.
Veröffentlicht: (2025)
von: Qin, Zhanyue, et al.
Veröffentlicht: (2025)
Surrogate Signals from Format and Length: Reinforcement Learning for Solving Mathematical Problems without Ground Truth Answers
von: Xin, Rihui, et al.
Veröffentlicht: (2025)
von: Xin, Rihui, et al.
Veröffentlicht: (2025)
HBot: A Chatbot for Healthcare Applications in Traditional Chinese Medicine Based on Human Body 3D Visualization
von: Zhang, Bolin, et al.
Veröffentlicht: (2024)
von: Zhang, Bolin, et al.
Veröffentlicht: (2024)
Full-ECE: A Metric For Token-level Calibration on Large Language Models
von: Liu, Han, et al.
Veröffentlicht: (2024)
von: Liu, Han, et al.
Veröffentlicht: (2024)
EchoReview: Learning Peer Review from the Echoes of Scientific Citations
von: Zhang, Yinuo, et al.
Veröffentlicht: (2026)
von: Zhang, Yinuo, et al.
Veröffentlicht: (2026)
Implicit Reward as the Bridge: A Unified View of SFT and DPO Connections
von: Wang, Bo, et al.
Veröffentlicht: (2025)
von: Wang, Bo, et al.
Veröffentlicht: (2025)
Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillation
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
von: Zhao, Anhao, et al.
Veröffentlicht: (2026)
ProFit: Leveraging High-Value Signals in SFT via Probability-Guided Token Selection
von: Liu, Tao, et al.
Veröffentlicht: (2026)
von: Liu, Tao, et al.
Veröffentlicht: (2026)
FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
von: Fan, Qihang, et al.
Veröffentlicht: (2026)
Large Language Models for Planning: A Comprehensive and Systematic Survey
von: Cao, Pengfei, et al.
Veröffentlicht: (2025)
von: Cao, Pengfei, et al.
Veröffentlicht: (2025)
PDC & DM-SFT: A Road for LLM SQL Bug-Fix Enhancing
von: Duan, Yiwen, et al.
Veröffentlicht: (2024)
von: Duan, Yiwen, et al.
Veröffentlicht: (2024)
Calibrated Language Models Must Hallucinate
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023)
von: Kalai, Adam Tauman, et al.
Veröffentlicht: (2023)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
von: Kang, Feiyang, et al.
Veröffentlicht: (2025)
Towards a Unified View of Large Language Model Post-Training
von: Lv, Xingtai, et al.
Veröffentlicht: (2025)
von: Lv, Xingtai, et al.
Veröffentlicht: (2025)
DR.EHR: Dense Retrieval for Electronic Health Record with Knowledge Injection and Synthetic Data
von: Zhao, Zhengyun, et al.
Veröffentlicht: (2025)
von: Zhao, Zhengyun, et al.
Veröffentlicht: (2025)
UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024)
An Empirical Study of SFT-DPO Interaction and Parameterization in Small Language Models
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
von: Feng, Yuming, et al.
Veröffentlicht: (2026)
Continual SFT Matches Multimodal RLHF with Negative Supervision
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN
von: Xu, Yao, et al.
Veröffentlicht: (2025)
von: Xu, Yao, et al.
Veröffentlicht: (2025)
CodeUnlearn: Amortized Zero-Shot Machine Unlearning in Language Models Using Discrete Concept
von: Wu, YuXuan, et al.
Veröffentlicht: (2024)
von: Wu, YuXuan, et al.
Veröffentlicht: (2024)
LLaSA: Large Language and Structured Data Assistant
von: Xu, Yao, et al.
Veröffentlicht: (2024)
von: Xu, Yao, et al.
Veröffentlicht: (2024)
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
von: Li, Chuhan, et al.
Veröffentlicht: (2024)
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
von: Shin, Haebin, et al.
Veröffentlicht: (2025)
On Fairness of Unified Multimodal Large Language Model for Image Generation
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
Towards Objectively Benchmarking Social Intelligence for Language Agents at Action Level
von: Wang, Chenxu, et al.
Veröffentlicht: (2024)
von: Wang, Chenxu, et al.
Veröffentlicht: (2024)
Minor SFT loss for LLM fine-tune to increase performance and reduce model deviation
von: Xie, Shiming, et al.
Veröffentlicht: (2024)
von: Xie, Shiming, et al.
Veröffentlicht: (2024)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Deeper Insights into Learning Performance of Stochastic Configuration Networks
von: Yan, Xiufeng, et al.
Veröffentlicht: (2024)
von: Yan, Xiufeng, et al.
Veröffentlicht: (2024)
Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
von: Liu, Jiahao, et al.
Veröffentlicht: (2025)
von: Liu, Jiahao, et al.
Veröffentlicht: (2025)
How Open Must Language Models be to Enable Reliable Scientific Inference?
von: Michaelov, James A., et al.
Veröffentlicht: (2026)
von: Michaelov, James A., et al.
Veröffentlicht: (2026)
CollectiveSFT: Scaling Large Language Models for Chinese Medical Benchmark with Collective Instructions in Healthcare
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhu, Jingwei, et al.
Veröffentlicht: (2024)
SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
von: Limozin, Alexis, et al.
Veröffentlicht: (2026)
SkillBrew: Multi-Objective Curation of Skill Banks for LLM Agents
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
von: Hu, Wentao, et al.
Veröffentlicht: (2026)
PR-CAD: Progressive Refinement for Unified Controllable and Faithful Text-to-CAD Generation with Large Language Models
von: An, Jiyuan, et al.
Veröffentlicht: (2026)
von: An, Jiyuan, et al.
Veröffentlicht: (2026)
Enhancing Multilingual Counterfactual Generation through Alignment-as-Preference Optimization
von: Wang, Yilong, et al.
Veröffentlicht: (2026)
von: Wang, Yilong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Beyond Confidence: The Rhythms of Reasoning in Generative Models
von: Liu, Deyuan, et al.
Veröffentlicht: (2026) -
Checkpoint Merging via Bayesian Optimization in LLM Pretraining
von: Liu, Deyuan, et al.
Veröffentlicht: (2024) -
Mitigating Gender Bias in Code Large Language Models via Model Editing
von: Qin, Zhanyue, et al.
Veröffentlicht: (2024) -
Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
von: Liu, Deyuan, et al.
Veröffentlicht: (2024) -
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
von: Qin, Zhanyue, et al.
Veröffentlicht: (2025)